Skip to main navigation Skip to search Skip to main content

PROSODIC VS. SEGMENTAL CONTRIBUTIONS TO NATURALNESS IN A DIPHONE SYNTHESIZER

  • University of Delaware

Research output: Contribution to conferencePaperpeer-review

5 Scopus citations

Abstract

The relative contributions of segmental versus prosodic factors to the perceived naturalness of synthetic speech was measured by transplanting prosody between natural speech and the output of a diphone synthesizer. A small corpus was created containing matched sentence pairs wherein one member of the pair was a natural utterance and the other was a synthetic utterance generated with diphone data from the same talker. Two additional sentences were formed from each sentence pair by transplanting the prosodic structure between the natural and synthetic members of each pair. In two listening experiments subjects were asked to (a) classify each sentence as “natural” or “synthetic, or (b) rate the naturalness of each sentence. Results showed that the prosodic information was more important than segmental information in both classification and ratings of naturalness.

Original languageEnglish
StatePublished - 1998
Event5th International Conference on Spoken Language Processing, ICSLP 1998 - Sydney, Australia
Duration: 30 Nov 19984 Dec 1998

Conference

Conference5th International Conference on Spoken Language Processing, ICSLP 1998
Country/TerritoryAustralia
CitySydney
Period30/11/984/12/98

Fingerprint

Dive into the research topics of 'PROSODIC VS. SEGMENTAL CONTRIBUTIONS TO NATURALNESS IN A DIPHONE SYNTHESIZER'. Together they form a unique fingerprint.

Cite this