Abstract
The relative contributions of segmental versus prosodic factors to the perceived naturalness of synthetic speech was measured by transplanting prosody between natural speech and the output of a diphone synthesizer. A small corpus was created containing matched sentence pairs wherein one member of the pair was a natural utterance and the other was a synthetic utterance generated with diphone data from the same talker. Two additional sentences were formed from each sentence pair by transplanting the prosodic structure between the natural and synthetic members of each pair. In two listening experiments subjects were asked to (a) classify each sentence as “natural” or “synthetic, or (b) rate the naturalness of each sentence. Results showed that the prosodic information was more important than segmental information in both classification and ratings of naturalness.
| Original language | English |
|---|---|
| State | Published - 1998 |
| Event | 5th International Conference on Spoken Language Processing, ICSLP 1998 - Sydney, Australia Duration: 30 Nov 1998 → 4 Dec 1998 |
Conference
| Conference | 5th International Conference on Spoken Language Processing, ICSLP 1998 |
|---|---|
| Country/Territory | Australia |
| City | Sydney |
| Period | 30/11/98 → 4/12/98 |
Fingerprint
Dive into the research topics of 'PROSODIC VS. SEGMENTAL CONTRIBUTIONS TO NATURALNESS IN A DIPHONE SYNTHESIZER'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver