Fetching the paper…
Reading the bibliography…
We introduce EmphAssess, a prosodic benchmark designed to evaluate the capability of speech-to-speech models to encode and reproduce prosodic emphasis.
Direct speech-to-speech translation with a sequence-to-sequence model
Ye Jia, Ron J Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, and Yonghui Wu. 2019 · 1904
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al. 2019 · 1910
Earlier work this paper cites.
Prosody in the comprehension of spoken language: A literature review
Anne Cutler, Delphine Dahan, and Wilma Van Donselaar. 1997 · 1997
Earlier work this paper cites.
The perception of prosodic prominence
Jacques Terken and Dik Hermes. 2000 · 2000
Earlier work this paper cites.
The music of everyday speech: Prosody and discourse analysis
Ann Wennerstrom. 2001 · 2001
Earlier work this paper cites.
Masoud Jalili Sabet, Philipp Dufter, François Yvon, and Hinrich Schütze. 2020 · 2004
Earlier work this paper cites.
The detection of emphatic words using acoustic and lexical features
Jason M Brenier, Daniel M Cer, and Daniel Jurafsky. 2005 · 2005
Earlier work this paper cites.
Nltk: the natural language toolkit
Steven Bird. 2006 · 2006
Earlier work this paper cites.
Prosodic prominence and boundaries in sequence-to-sequence speech synthesis
Antti Suni, Sofoklis Kakouros, Martti Vainio, and Juraj Šimko. 2020 · 2006
Earlier work this paper cites.
Acoustic correlates of prosodic prominence for naiïve listeners of american english
Yoonsook Mo. 2008 · 2008
Earlier work this paper cites.
A study on the effect of prosodic emphasis transfer on overall speech translation quality
Andreas Tsiartas, Panayiotis G Georgiou, and Shrikanth S Narayanan. 2013 · 2013
Earlier work this paper cites.
Prosody and language comprehension
Delphine Dahan. 2015 · 2015
Earlier work this paper cites.
Preserving word-level emphasis in speech-to-speech translation
Quoc Truong Do, Tomoki Toda, Graham Neubig, Sakriani Sakti, and Satoshi Nakamura. 2016 · 2016
Earlier work this paper cites.
Lexical emphasis detection in spoken french using f-banks and neural networks
Abdelwahab Heba, Thomas Pellegrini, Tom Jorquera, Régine André-Obrecht, and Jean-Pierre Lorré. 2017 · 2017
Earlier work this paper cites.
Montreal forced aligner: Trainable text-speech alignment using kaldi
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger. 2017 · 2017
Cited alongside, same era.
Learning cross-lingual knowledge with multilingual blstm for emphasis detection with limited training data
Yishuang Ning, Zhiyong Wu, Runnan Li, Jia Jia, Mingxing Xu, Helen Meng, and Lianhong Cai. 2017 · 2017
Cited alongside, same era.
Sequence-to-sequence models for emphasis speech translation
Quoc Truong Do, Sakriani Sakti, and Satoshi Nakamura. 2018 · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018 · 2018
Cited alongside, same era.
Emphasis detection for voice dialogue applications using multi-channel convolutional bidirectional long short-term memory network
Long Zhang, Jia Jia, Fanbo Meng, Suping Zhou, Wei Chen, Cunjun Zhang, and Runnan Li. 2018 · 2018
Cited alongside, same era.
Translatotron 2: High-quality direct speech-to-speech translation with voice preservation
Ye Jia, Michelle Tadmor Ramanovich, Tal Remez, and Roi Pomerantz. 2022 · 2022
Later among the works it cites.
On the utility of self-supervised models for prosody-related tasks
Guan-Ting Lin, Chi-Luen Feng, Wei-Ping Huang, Yuan Tseng, Tzu-Han Lin, Chen-An Li, Hung-yi Lee, and Nigel G Ward. 2023 · 2022
Later among the works it cites.
Self-supervised speech representation learning: A review
Abdelrahman Mohamed, Hung-yi Lee, Lasse Borgholt, Jakob D Havtorn, Joakim Edin, Christian Igel, Katrin Kirchhoff, Shang-Wen Li, Karen Livescu, Lars Maaløe, et al. 2022 · 2022
Later among the works it cites.
Deep learning for prominence detection in children’s read speech
Mithilesh Vaidya, Kamini Sabu, and Preeti Rao. 2022 · 2022
Later among the works it cites.
Towards cross-language prosody transfer for dialog
Jonathan E Avila and Nigel G Ward. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli. 2020 · 2020
Cited alongside, same era.
Inferring emphasis for real voice data: an attentive multimodal neural network approach
Suping Zhou, Jia Jia, Long Zhang, Yanfeng Wang, Wei Chen, Fanbo Meng, Fei Yu, and Jialie Shen. 2020 · 2020
Cited alongside, same era.
Xls-r: Self-supervised cross-lingual speech representation learning at scale
Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, et al. 2021 · 2021
Cited alongside, same era.
Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021 · 2021
Cited alongside, same era.
Text-free prosody-aware generative spoken language modeling
Eugene Kharitonov, Ann Lee, Adam Polyak, Yossi Adi, Jade Copet, Kushal Lakhotia, Tu-Anh Nguyen, Morgane Rivière, Abdelrahman Mohamed, Emmanuel Dupoux, et al. 2021 · 2021
Cited alongside, same era.
On generative spoken language modeling from raw audio
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al. 2021 · 2021
Cited alongside, same era.
Textless speech-to-speech translation on real data
Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Yossi Adi, Juan Pino, Jiatao Gu, et al. 2021 · 2021
Cited alongside, same era.
Max Bain, Jaesung Huh, Tengda Han, and Andrew Zisserman. 2023 · 2023
Closest in time.
Seamlessm4t-massively multilingual & multimodal machine translation
Loïc Barrault, Yu-An Chung, Mariano Cora Meglioli, David Dale, Ning Dong, Paul-Ambroise Duquenne, Hady Elsahar, Hongyu Gong, Kevin Heffernan, John Hoffman, et al. 2023 · 2023
Closest in time.
Audiolm: a language modeling approach to audio generation
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, et al. 2023 · 2023
Closest in time.
Prosaudit, a prosodic benchmark for self-supervised speech models
Maureen de Seyssel, Marvin Lavechin, Hadrien Titeux, Arthur Thomas, Gwendal Virlet, Andrea Santos Revilla, Guillaume Wisniewski, Bogdan Ludusan, and Emmanuel Dupoux. 2023 · 2023
Closest in time.
Enhancing expressivity transfer in textless speech-to-speech translation
Jarod Duret, Benjamin O’Brien, Yannick Estève, and Titouan Parcollet. 2023 · 2023
Closest in time.
A holistic cascade system, benchmark, and human evaluation protocol for expressive speech-to-speech translation
Wen-Chin Huang, Benjamin Peloquin, Justine Kao, Changhan Wang, Hongyu Gong, Elizabeth Salesky, Yossi Adi, Ann Lee, and Peng-Jen Chen. 2023 · 2023
Closest in time.
Prosodic prominence across languages
D Robert Ladd and Amalia Arvaniti. 2023 · 2023
Closest in time.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2023 · 2023
Closest in time.
Audiopalm: A large language model that can speak and listen
Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al. 2023 · 2023
Closest in time.