Segment selection method based on tonal validity evaluation using machine learning for concatenattve speech synthesis

Akihiro Yoshida, Hideyuki Mizuno, Kazunori Mano

研究成果: Conference contribution

4 引用 (Scopus)

抜粋

This paper proposes a speech segment selection method based on machine learning for concatenative speech synthesis systems. The proposed method has two novel features. One is its use of Support Vector Machine (SVM) to estimate the subjective correctness of pitch accent with respect to each accent phrase of possible candidate speech segments. The other is its use of a determination function to identify the best segment based on SVM output. The determination function involves two assessments; one counts the number of each sign of SVM output and the other compares the distance values. The sign of SVM output is generally used to classify target objects, but the distance SVM output also represents important information. An experiment that assesses SVM performance for Japanese accent validity shows that its accuracy is 81%. To confirm the effectiveness of the proposed segment selection method, preference tests are conducted. The test indicates that the proposed method can yield Japanese synthesized speech with more natural intonation than the conventional method that uses only target cost and concatenation cost.

元の言語English
ホスト出版物のタイトル2008 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP
ページ4617-4620
ページ数4
DOI
出版物ステータスPublished - 2008 9 17
イベント2008 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP - Las Vegas, NV, United States
継続期間: 2008 3 312008 4 4

出版物シリーズ

名前ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
ISSN(印刷物)1520-6149

Conference

Conference2008 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP
United States
Las Vegas, NV
期間08/3/3108/4/4

ASJC Scopus subject areas

  • Software
  • Signal Processing
  • Electrical and Electronic Engineering

フィンガープリント Segment selection method based on tonal validity evaluation using machine learning for concatenattve speech synthesis' の研究トピックを掘り下げます。これらはともに一意のフィンガープリントを構成します。

  • これを引用

    Yoshida, A., Mizuno, H., & Mano, K. (2008). Segment selection method based on tonal validity evaluation using machine learning for concatenattve speech synthesis. : 2008 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP (pp. 4617-4620). [4518685] (ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings). https://doi.org/10.1109/ICASSP.2008.4518685