Fetching the paper…
Reading the bibliography…
State-of-the-art text-to-speech (TTS) systems have utilized pretrained language models (PLMs) to enhance prosody and create more natural-sounding speech.
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler, “Aligning Books and Movies: Towards Story-Like Visual Explanations by Watching Movies and Reading Books,” in 2015 IEEE International Conference on Computer Vision (ICCV) . Santiago, Chile: IEEE, Dec. 2015, pp. 19–27. [Online]. Available: http://ieeexplore.ieee.org/document/7410368/
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan, R. A. Saurous, Y. Agiomvrgiannakis, and Y. Wu, “Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram Predictions,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . Calgary, AB: IEEE, Apr. 2018, pp. 4779–4783. [Online]. Available: https://ieeexplore.ieee.org/document/8461368/
2018
Earlier work this paper cites.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding,” in Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP . Brussels, Belgium: Association for Computational Linguistics, 2018, pp. 353–355. [Online]. Available: http://aclweb.org/anthology/W18-5446
2018
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 4171–4186. [Online]. Available: https://aclanthology.org/N19-1423
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
T. Hayashi, S. Watanabe, T. Toda, K. Takeda, S. Toshniwal, and K. Livescu, “Pre-Trained Text Embeddings for Enhanced Text-to-Speech Synthesis,” in INTERSPEECH 2019 . ISCA, Sep. 2019, pp. 4430–4434. [Online]. Available: https://www.isca-speech.org/archive/interspeech_2019/hayashi19_interspeech.html
2019
Earlier work this paper cites.
A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov, “Unsupervised Cross-lingual Representation Learning at Scale,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Online: Association for Computational Linguistics, 2020, pp. 8440–8451. [Online]. Available: https://www.aclweb.org/anthology/2020.acl-main.747
2020
Cited alongside, same era.
K. Clark, M.-T. Luong, Q. V. Le, and C. D. Manning, “ELECTRA: Pre-training text encoders as discriminators rather than generators,” in International Conference on Learning Representations , 2020. [Online]. Available: https://openreview.net/forum?id=r1xMH1BtvB
2020
Cited alongside, same era.
Z. Dai, G. Lai, Y. Yang, and Q. V. Le, “Funnel-transformer: filtering out sequential redundancy for efficient language processing,” in Proceedings of the 34th International Conference on Neural Information Processing Systems , 2020, pp. 4271–4282
2020
Cited alongside, same era.
2020
Later among the works it cites.
L. Zhuang, L. Wayne, S. Ya, and Z. Jun, “A Robustly Optimized BERT Pre-training Approach with Post-training,” in Proceedings of the 20th Chinese National Conference on Computational Linguistics . Huhhot, China: Chinese Information Processing Society of China, Aug. 2021, pp. 1218–1227. [Online]. Available: https://aclanthology.org/2021.ccl-1.108
2021
Later among the works it cites.
G. Xu, W. Song, Z. Zhang, C. Zhang, X. He, and B. Zhou, “Improving Prosody Modelling with Cross-Utterance BERT Embeddings for End-to-End Speech Synthesis,” in ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . Toronto, ON, Canada: IEEE, Jun. 2021, pp. 6079–6083. [Online]. Available: https://ieeexplore.ieee.org/document/9414102/
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Iandola, A. Shaw, R. Krishna, and K. Keutzer, “SqueezeBERT: What can computer vision teach NLP about efficient neural networks?” in Proceedings of SustaiNLP: Workshop on Simple and Efficient Natural Language Processing . Online: Association for Computational Linguistics, 2020, pp. 124–135. [Online]. Available: https://www.aclweb.org/anthology/2020.sustainlp-1.17
2020
Cited alongside, same era.
Z. Sun, H. Yu, X. Song, R. Liu, Y. Yang, and D. Zhou, “MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Online: Association for Computational Linguistics, 2020, pp. 2158–2170. [Online]. Available: https://www.aclweb.org/anthology/2020.acl-main.195
2020
Cited alongside, same era.
Y. Xiao, L. He, H. Ming, and F. K. Soong, “Improving Prosody with Linguistic and BERT Derived Features in Multi-Speaker Based Mandarin Chinese Neural TTS,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . Barcelona, Spain: IEEE, May 2020, pp. 6704–6708. [Online]. Available: https://ieeexplore.ieee.org/document/9054337/
2020
Cited alongside, same era.
T. Kenter, M. Sharma, and R. Clark, “Improving the Prosody of RNN-Based English Text-To-Speech Synthesis by Incorporating a BERT Model,” in INTERSPEECH 2020 . ISCA, Oct. 2020, pp. 4412–4416. [Online]. Available: https://www.isca-speech.org/archive/interspeech_2020/kenter20_interspeech.html
2020
Cited alongside, same era.
2020
Cited alongside, same era.
X. Qiu, T. Sun, Y. Xu, Y. Shao, N. Dai, and X. Huang, “Pre-trained models for natural language processing: A survey,” Science China Technological Sciences , vol. 63, no. 10, pp. 1872–1897, Oct. 2020. [Online]. Available: https://link.springer.com/10.1007/s11431-020-1647-3
2020
Cited alongside, same era.
A. Rogers, O. Kovaleva, and A. Rumshisky, “A Primer in BERTology: What We Know About How BERT Works,” Transactions of the Association for Computational Linguistics , vol. 8, pp. 842–866, Dec. 2020. [Online]. Available: https://direct.mit.edu/tacl/article/96482
2020
Cited alongside, same era.
2021
Later among the works it cites.
X. Han, Z. Zhang, N. Ding, Y. Gu, X. Liu, Y. Huo, J. Qiu, Y. Yao, A. Zhang, L. Zhang, W. Han, M. Huang, Q. Jin, Y. Lan, Y. Liu, Z. Liu, Z. Lu, X. Qiu, R. Song, J. Tang, J.-R. Wen, J. Yuan, W. X. Zhao, and J. Zhu, “Pre-trained models: Past, present and future,” AI Open , vol. 2, pp. 225–250, 2021. [Online]. Available: https://linkinghub.elsevier.com/retrieve/pii/S2666651021000231
2021
Later among the works it cites.
P. Makarov, S. Ammar Abbas, M. Łajszczak, A. Joly, S. Karlapati, A. Moinet, T. Drugman, and P. Karanasou, “Simple and Effective Multi-sentence TTS with Expressive and Coherent Prosody,” in INTERSPEECH 2022 . ISCA, Sep. 2022, pp. 3368–3372. [Online]. Available: https://www.isca-speech.org/archive/interspeech_2022/makarov22_interspeech.html
2022
Later among the works it cites.
S. Karlapati, P. Karanasou, M. Łajszczak, S. Ammar Abbas, A. Moinet, P. Makarov, R. Li, A. Van Korlaar, S. Slangen, and T. Drugman, “CopyCat2: A Single Model for Multi-Speaker TTS and Many-to-Many Fine-Grained Prosody Transfer,” in INTERSPEECH 2022 . ISCA, Sep. 2022, pp. 3363–3367. [Online]. Available: https://www.isca-speech.org/archive/interspeech_2022/karlapati22_interspeech.html
2022
Later among the works it cites.
S. Casola, I. Lauriola, and A. Lavelli, “Pre-trained transformers: an empirical comparison,” Machine Learning with Applications , vol. 9, p. 100334, Sep. 2022. [Online]. Available: https://linkinghub.elsevier.com/retrieve/pii/S2666827022000445
2022
Later among the works it cites.
J. Li, T. Tang, Z. Gong, L. Yang, Z. Yu, Z. Chen, J. Wang, X. Zhao, and J.-R. Wen, “ElitePLM: An Empirical Study on General Language Ability Evaluation of Pretrained Language Models,” in Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies . Seattle, United States: Association for Computational Linguistics, 2022, pp. 3519–3539. [Online]. Available: https://aclanthology.org/2022.naacl-main.258
2022
Later among the works it cites.
V. Lialin, K. Zhao, N. Shivagunde, and A. Rumshisky, “Life after BERT: What do Other Muppets Understand about Language?” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . Dublin, Ireland: Association for Computational Linguistics, 2022, pp. 3180–3193. [Online]. Available: https://aclanthology.org/2022.acl-long.227
2022
Later among the works it cites.
A. Abbas, S. Karlapati, B. Schnell, P. Karanasou, M. Granero Moya, A. Nagaraj, A. Boustati, N. Peinelt, A. Moinet, and T. Drugman, “eCat: An End-to-End Model for Multi-Speaker TTS & Many-to-Many Fine-Grained Prosody Transfer,” in INTERSPEECH 2023 . ISCA, 2023
2023
Closest in time.