Fetching the paper…
Reading the bibliography…
Speech-to-speech translation systems today do not adequately support use for dialog purposes.
A. Berry, “Spanish and American Turn-Taking Styles: A Comparative Study,” Education Resources Information Center, Tech. Rep. ED398747, 1994
1994
Earlier work this paper cites.
D. Ramírez Verdugo, “The nature and patterning of native and non-native intonation in the expression of certainty and uncertainty: Pragmatic effects,” Journal of Pragmatics , vol. 37, no. 12, pp. 2086–2115, 2005
2005
Earlier work this paper cites.
A. Rilliard, A. Allauzen, and P. B. de Mareüil, “Using Dynamic Time Warping to Compute Prosodic Similarity Measures,” in Interspeech , 2011
2011
Earlier work this paper cites.
T. Kano, S. Sakti, S. Takamichi, G. Neubig, T. Toda, and S. Nakamura, “A method for translation of paralinguistic information,” in Proceedings of the 9th International Workshop on Spoken Language Translation , 2012, pp. 158–163
2012
Earlier work this paper cites.
M. G. V. Farías, “A comparative analysis of intonation between Spanish and English speakers in tag questions, wh-questions, inverted questions, and repetition questions,” Revista Brasileira de Linguística Aplicada , vol. 13, no. 4, pp. 1061–1083, 2013
2013
Earlier work this paper cites.
L. Mary, A. Babu K. K, A. Joseph, and G. M. George, “Evaluation of mimicked speech using prosodic features,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing , 2013, pp. 7189–7193
2013
Earlier work this paper cites.
Q. T. Do, S. Sakti, G. Neubig, T. Toda, and S. Nakamura, “Improving translation of emphasis with pause prediction in speech-to-speech translation systems.” in IWSLT , 2015
2015
Earlier work this paper cites.
R. Cattoni, M. A. Di Gangi, L. Bentivogli, M. Negri, and M. Turchi, “MuST-C: A multilingual corpus for end-to-end speech translation,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , vol. 1. Association for Computational Linguistics, 2019, pp. 2012–2017
2017
Earlier work this paper cites.
N. G. Ward and P. Gallardo, “Non-Native Differences in Prosodic-Construction Use,” Dialogue & Discourse , vol. 8, no. 1, pp. 1–30, 2017
2017
Earlier work this paper cites.
G. Zárate-Sández, “Production of final boundary tones in declarative utterances by English-speaking learners of Spanish,” in Proceedings of the 9th International Conference on Speech Prosody. International Speech Communication Association (ISCA) Online Archive , 2018, pp. 927–31
2018
Earlier work this paper cites.
N. G. Ward, Prosodic Patterns in English Conversation . Cambridge University Press, 2019
2019
Earlier work this paper cites.
D. J. Liebling, M. Lahav, A. Evans, A. Donsbach, J. Holbrook, B. Smus, and L. Boran, “Unmet Needs and Opportunities for Mobile Translation AI,” in Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . Association for Computing Machinery, 2020, pp. 1–13
2020
Cited alongside, same era.
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber, “Common Voice: A Massively-Multilingual Speech Corpus,” in Proceedings of the 12th Language Resources and Evaluation Conference . European Language Resources Association, 2020, pp. 4218–4222
2020
Cited alongside, same era.
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert, “MLS: A Large-Scale Multilingual Dataset for Speech Research,” in Proc. Interspeech , 2020, pp. 2757–2761
2020
Cited alongside, same era.
M. Zanon Boito, W. Havard, M. Garnerin, É. Le Ferrand, and L. Besacier, “MaSS: A Large and Clean Multilingual Corpus of Sentence-aligned Spoken Utterances Extracted from the Bible,” in Proceedings of the 12th Language Resources and Evaluation Conference . European Language Resources Association, 2020, pp. 6486–6493
2022
Later among the works it cites.
A. Lee, P.-J. Chen, C. Wang, J. Gu, S. Popuri, X. Ma, A. Polyak, Y. Adi, Q. He, Y. Tang, J. Pino, and W.-N. Hsu, “Direct Speech-to-Speech Translation With Discrete Units,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2022, pp. 3327–3339
2022
Later among the works it cites.
A. Lee, H. Gong, P.-A. Duquenne, H. Schwenk, P.-J. Chen, C. Wang, S. Popuri, J. Pino, J. Gu, and W.-N. Hsu, “Textless speech-to-speech translation on real data,” in NAACL , 2022
2022
Later among the works it cites.
Q. Dong, F. Yue, T. Ko, M. Wang, Q. Bai, and Y. Zhang, “Leveraging Pseudo-labeled Data to Improve Direct Speech-to-Speech Translation,” in Proc. Interspeech , 2022, pp. 1781–1785
2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
C. Wang, A. Wu, J. Gu, and J. Pino, “CoVoST 2 and Massively Multilingual Speech-to-Text Translation,” in Interspeech , 2021, pp. 2247–2251
2021
Cited alongside, same era.
C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux, “VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . Association for Computational Linguistics, 2021, pp. 993–1003
2021
Cited alongside, same era.
E. Salesky, M. Wiesner, J. Bremerman, R. Cattoni, M. Negri, M. Turchi, D. W. Oard, and M. Post, “The Multilingual TEDx Corpus for Speech Recognition and Translation,” in Interspeech 2021 . ISCA, 2021, pp. 3655–3659
2021
Cited alongside, same era.
A. Öktem, M. Farrús, and A. Bonafonte, “Corpora compilation for prosody-informed speech processing,” Language Resources and Evaluation , vol. 55, no. 4, pp. 925–946, 2021
2021
Cited alongside, same era.
K. Doi, K. Sudoh, and S. Nakamura, “Large-Scale English-Japanese Simultaneous Interpretation Corpus: Construction and Analyses with Sentence-Aligned Data,” in Proceedings of the 18th International Conference on Spoken Language Translation (IWSLT) . Association for Computational Linguistics, 2021, pp. 226–235
2021
Cited alongside, same era.
C. Zhang, X. Tan, Y. Ren, T. Qin, K. Zhang, and T.-Y. Liu, “UWSpeech: Speech to Speech Translation for Unwritten Languages,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 16, 2021, pp. 14 319–14 327
2021
Cited alongside, same era.
Y. Jia, M. T. Ramanovich, T. Remez, and R. Pomerantz, “Translatotron 2: High-quality direct speech-to-speech translation with voice preservation,” in Proceedings of the 39th International Conference on Machine Learning , 2022, pp. 10 120–10 134
2022
Cited alongside, same era.
Later among the works it cites.
Y. Jia, Y. Ding, A. Bapna, C. Cherry, Y. Zhang, A. Conneau, and N. Morioka, “Leveraging unsupervised and weakly-supervised data to improve direct speech-to-speech translation,” in Proc. Interspeech , 2022, pp. 1721–1725
2022
Later among the works it cites.
Y. Jia, M. Tadmor Ramanovich, Q. Wang, and H. Zen, “CVSS Corpus and Massively Multilingual Speech-to-Speech Translation,” in Proceedings of the Thirteenth Language Resources and Evaluation Conference . European Language Resources Association, 2022, pp. 6691–6703
2022
Later among the works it cites.
N. G. Ward, J. E. Avila, and E. Rivas, “Dialogs Re-enacted Across Languages,” University of Texas at El Paso, Technical UTEP-CS-22-108, 2022
2022
Later among the works it cites.
N. Ward, A. Kirkland, M. Wlodarczak, and É. Székely, “Two pragmatic functions of breathy voice in American English conversation,” in 11th International Conference on Speech Prosody , 2022, pp. 82–86
2022
Later among the works it cites.
2023
Closest in time.
OpenAI, “Whisper,” 2023. [Online]. Available: https://github.com/openai/whisper
2023
Closest in time.