Fetching the paper…
Reading the bibliography…
J. Hirschberg and J. Pierrehumbert, “The intonational structuring of discourse,” in 24th Annual Meeting of the Association for Computational Linguistics . New York, New York, USA: Association for Computational Linguistics, Jul. 1986, pp. 136–144. [Online]. Available: https://aclanthology.org/P86-1021
1986
Earlier work this paper cites.
J. Pierrehumbert and J. Hirschberg, The meaning of intonational contours in the interpretation of discourse , 01 1990
1990
Earlier work this paper cites.
M. Heldner and E. Strangert, “To what extent is perceived focus determined by f0-cues?” 5th European Conference on Speech Communication and Technology (Eurospeech 1997) , 1997
1997
Earlier work this paper cites.
T. Yoshimura, K. Tokuda, T. Masuko, T. Kobayashi, and T. Kitamura, “Duration modeling for HMM-based speech synthesis,” in Proc. 5th International Conference on Spoken Language Processing (ICSLP 1998) , 1998, p. paper 0939
1998
Earlier work this paper cites.
S.-H. Chen, S.-J. Chen, and C.-C. Kuo, “Perceptual distortion analysis and quality estimation of prosody-modified speech for TD-PSOLA,” in 2006 IEEE International Conference on Acoustics Speech and Signal Processing Proceedings , vol. 1. IEEE, 2006, pp. I–I
2006
Earlier work this paper cites.
V. Strom, A. Nenkova, R. Clark, Y. Vazquez-Alvarez, J. Brenier, S. King, and D. Jurafsky, “Modelling prominence and emphasis improves unit-selection synthesis,” 2007
2007
Earlier work this paper cites.
P. Boersma and D. Weenink, “Praat: doing phonetics by computer (version 5.1.13),” 2009. [Online]. Available: http://www.praat.org
2009
Earlier work this paper cites.
M. Breen, E. Fedorenko, M. Wagner, and E. Gibson, “Acoustic correlates of information structure.” Language and Cognitive Processes - LANG COGNITIVE PROCESS , vol. 25, pp. 1044–1098, 09 2010
2010
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz et al. , “The kaldi speech recognition toolkit,” in IEEE 2011 workshop on automatic speech recognition and understanding , no. CONF. IEEE Signal Processing Society, 2011
2011
Earlier work this paper cites.
B. Series, “Method for the subjective assessment of intermediate quality level of audio systems,” International Telecommunication Union Radiocommunication Assembly , 2014
2014
Earlier work this paper cites.
Y. Chen and R. Pan, “Automatic emphatic information extraction from aligned acoustic data and its application on sentence compression,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 31, no. 1, 2017
2017
Earlier work this paper cites.
A. Heba, T. Pellegrini, T. Jorquera, R. André-Obrecht, and J.-P. Lorré, “Lexical emphasis detection in spoken french using f-banks and neural networks,” in Statistical Language and Speech Processing: 5th International Conference, SLSP 2017, Le Mans, France, October 23–25, 2017, Proceedings 5 . Springer, 2017, pp. 241–249
2017
Earlier work this paper cites.
Q. T. Do, T. Toda, G. Neubig, S. Sakti, and S. Nakamura, “Preserving word-level emphasis in speech-to-speech translation,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 25, no. 3, pp. 544–556, 2017
2017
Cited alongside, same era.
S. Ö. Arık, M. Chrzanowski, A. Coates, G. Diamos, A. Gibiansky, Y. Kang, X. Li, J. Miller, A. Ng, J. Raiman et al. , “Deep voice: Real-time neural text-to-speech,” in International conference on machine learning . PMLR, 2017, pp. 195–204
2017
Cited alongside, same era.
M. Wang, Z. Wu, X. Wu, H. Meng, S. Kang, J. Jia, and L. Cai, “Emphatic speech synthesis and control based on characteristic transferring in end-to-end speech synthesis,” in 2018 First Asian Conference on Affective Computing and Intelligent Interaction (ACII Asia) . IEEE, 2018, pp. 1–6
2018
Cited alongside, same era.
Y. Mass, S. Shechtman, M. Mordechay, R. Hoory, O. Sar Shalom, G. Lev, and D. Konopnicki, “Word Emphasis Prediction for Expressive Text to Speech,” in Proc. Interspeech 2018 , 2018, pp. 2868–2872
S. Latif, I. Kim, I. Calapodescu, and L. Besacier, “Controlling prosody in end-to-end TTS: A case study on contrastive focus generation,” in Proceedings of the 25th Conference on Computational Natural Language Learning . Online: Association for Computational Linguistics, Nov. 2021, pp. 544–551. [Online]. Available: https://aclanthology.org/2021.conll-1.42
2021
Later among the works it cites.
L. Liu, J. Hu, Z. Wu, S. Yang, S. Yang, J. Jia, and H. Meng, “Controllable emphatic speech synthesis based on forward attention for expressive speech synthesis,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 410–414
2021
Later among the works it cites.
Y. Ren, C. Hu, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech 2: Fast and high-quality end-to-end text to speech,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=piLPYqxtWuA
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
S. Shechtman and M. Mordechay, “Emphatic speech prosody prediction with deep LSTM networks,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 5119–5123
2018
Cited alongside, same era.
L. Zhang, J. Jia, F. Meng, S. Zhou, W. Chen, C. Zhang, and R. Li, “Emphasis detection for voice dialogue applications using multi-channel convolutional bidirectional long short-term memory network,” in 2018 11th International Symposium on Chinese Spoken Language Processing (ISCSLP) , 2018, pp. 210–214
2018
Cited alongside, same era.
F. Charpentier and M. Stella, “Diphone synthesis using an overlap-add technique for speech waveforms concatenation,” in ICASSP’86. IEEE International Conference on Acoustics, Speech, and Signal Processing , vol. 11. IEEE, 1986, pp. 2015–2018
2018
Cited alongside, same era.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,” in 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2018, pp. 4779–4783
2018
Cited alongside, same era.
Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” Advances in neural information processing systems , vol. 32, 2019
2019
Cited alongside, same era.
M. Wagner, “Prosodic Focus,” in The Wiley Blackwell Companion to Semantics , 1st ed., D. Gutzmann, L. Matthewson, C. Meier, H. Rullmann, and T. Zimmermann, Eds. Wiley, Nov. 2020, pp. 1–75. [Online]. Available: https://onlinelibrary.wiley.com/doi/10.1002/9781118788516.sem133
2020
Cited alongside, same era.
C. Yu, H. Lu, N. Hu, M. Yu, C. Weng, K. Xu, P. Liu, D. Tuo, S. Kang, G. Lei, D. Su, and D. Yu, “DurIAN: Duration Informed Attention Network for Speech Synthesis,” in Proc. Interspeech 2020 , 2020, pp. 2027–2031. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-2968
2020
Cited alongside, same era.
Z. Hodari, A. Moinet, S. Karlapati, J. Lorenzo-Trueba, T. Merritt, A. Joly, A. Abbas, P. Karanasou, and T. Drugman, “Camp: a two-stage approach to modelling prosody in context,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6578–6582
2021
Later among the works it cites.
Y. Jiao, A. Gabryś, G. Tinchev, B. Putrycz, D. Korzekwa, and V. Klimkov, “Universal neural vocoding with parallel wavenet,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 6044–6048
2021
Later among the works it cites.
P. Makarov, S. A. Abbas, M. Lajszczak, A. Joly, S. Karlapati, A. Moinet, T. Drugman, and P. Karanasou, “Simple and effective multi-sentence TTS with expressive and coherent prosody,” in Interspeech 2022, 23rd Annual Conference of the International Speech Communication Association, Incheon, Korea, 18-22 September 2022 , H. Ko and J. H. L. Hansen, Eds. ISCA, 2022, pp. 3368–3372. [Online]. Available: https://doi.org/10.21437/Interspeech.2022-379
2022
Later among the works it cites.
S. Karlapati, P. Karanasou, M. Lajszczak, S. A. Abbas, A. Moinet, P. Makarov, R. Li, A. van Korlaar, S. Slangen, and T. Drugman, “Copycat2: A single model for multi-speaker TTS and many-to-many fine-grained prosody transfer,” in Interspeech 2022, 23rd Annual Conference of the International Speech Communication Association, Incheon, Korea, 18-22 September 2022 , H. Ko and J. H. L. Hansen, Eds. ISCA, 2022, pp. 3363–3367. [Online]. Available: https://doi.org/10.21437/Interspeech.2022-367
2022
Later among the works it cites.
S. A. Abbas, T. Merritt, A. Moinet, S. Karlapati, E. Muszynska, S. Slangen, E. Gatti, and T. Drugman, “Expressive, variable, and controllable duration modelling in TTS,” in Interspeech 2022, 23rd Annual Conference of the International Speech Communication Association, Incheon, Korea, 18-22 September 2022 , H. Ko and J. H. L. Hansen, Eds. ISCA, 2022, pp. 4546–4550. [Online]. Available: https://doi.org/10.21437/Interspeech.2022-384
2022
Later among the works it cites.
J. Effendi, Y. Virkar, R. Barra-Chicote, and M. Federico, “Duration modeling of neural TTS for automatic dubbing,” in ICASSP 2022 , 2022. [Online]. Available: https://www.amazon.science/publications/duration-modeling-of-neural-{TTS}-for-automatic-dubbing
2022
Later among the works it cites.
M. Lajszczak, A. Prasad, A. Van Korlaar, B. Bollepalli, A. Bonafonte, A. Joly, M. Nicolis, A. Moinet, T. Drugman, T. Wood et al. , “Distribution augmentation for low-resource expressive text-to-speech,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 8307–8311
2022
Later among the works it cites.
A. Graves and J. Schmidhuber, “Framewise phoneme classification with bidirectional lstm networks,” in Proceedings. 2005 IEEE International Joint Conference on Neural Networks, 2005. , vol. 4. IEEE, 2005, pp. 2047–2052
2052
Closest in time.