Fetching the paper…
Reading the bibliography…
We introduce CVSS, a massively multilingual-to-English speech-to-speech translation (S2ST) corpus, covering sentence-level parallel S2ST pairs from 21 languages into English.
JANUS: A speech-to-speech translation system using connectionist and symbolic processing strategies
Alex Waibel, Ajay N Jain, Arthur E McNair, Hiroaki Saito, Alexander G Hauptmann, and Joe Tebelskis · 1991
Earlier work this paper cites.
TIMIT acoustic phonetic continuous speech corpus, 1993
John S Garofolo, Lori F Lamel, William M Fisher, Jonathan G Fiscus, and David S Pallett · 1993
Earlier work this paper cites.
Normalization of non-standard words
Richard Sproat, Alan W Black, Stanley Chen, Shankar Kumar, Mari Ostendorf, and Christopher Richards · 2001
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Creating corpora for speech-to-speech translation
Genichiro Kikui, Eiichiro Sumita, Toshiyuki Takezawa, and Seiichi Yamamoto · 2003
Earlier work this paper cites.
The Fisher corpus: A resource for the next generations of speech-to-text
Christopher Cieri, David Miller, and Kevin Walker · 2004
Earlier work this paper cites.
CIAIR simultaneous interpretation corpus
Hitomi Tohyama, Shigeki Matsubara, Koichiro Ryu, N Kawaguch, and Yasuyoshi Inagaki · 2004
Earlier work this paper cites.
An approach to corpus-based interpreting studies: Developing EPIC
Claudio Bendazzoli, Annalisa Sandrelli, et al · 2005
Earlier work this paper cites.
Comparative study on corpora for speech translation
Genichiro Kikui, Seiichi Yamamoto, Toshiyuki Takezawa, and Eiichiro Sumita · 2006
Earlier work this paper cites.
Resources for new research directions in speaker recognition: the Mixer 3, 4 and 5 corpora
Christopher Cieri, Linda Corson, David Graff, and Kevin Walker · 2007
Earlier work this paper cites.
Speaker recognition: Building the Mixer 4 and 5 corpora
Linda Brandschain, Christopher Cieri, David Graff, Abby Neely, and Kevin Walker · 2008
Earlier work this paper cites.
Improved speech-to-text translation with the Fisher and Callhome Spanish–English speech translation corpus
Matt Post, Gaurav Kumar, Adam Lopez, Damianos Karakos, Chris Callison-Burch, and Sanjeev Khudanpur · 2013
Earlier work this paper cites.
Constructing a speech translation system using simultaneous interpretation data
Hiroaki Shimizu, Graham Neubig, Sakriani Sakti, Tomoki Toda, and Satoshi Nakamura · 2013
Earlier work this paper cites.
Collection of a simultaneous translation corpus for comparative analysis
Hiroaki Shimizu, Graham Neubig, Sakriani Sakti, Tomoki Toda, and Satoshi Nakamura · 2014
Earlier work this paper cites.
The Kestrel TTS text normalization system
Peter Ebden and Richard Sproat · 2015
Earlier work this paper cites.
LibriSpeech: an ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Interpretese vs. translationese: The uniqueness of human strategies in simultaneous interpretation
He He, Jordan Boyd-Graber, and Hal Daumé III · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Sequence-to-sequence models can directly translate foreign speech
Ron J. Weiss, Jan Chorowski, Navdeep Jaitly, Yonghui Wu, and Zhifeng Chen · 2017
Cited alongside, same era.
Transfer learning from speaker verification to multispeaker text-to-speech synthesis
Ye Jia, Yu Zhang, Ron J Weiss, Quan Wang, Jonathan Shen, Fei Ren, Zhifeng Chen, Patrick Nguyen, Ruoming Pang, Ignacio Lopez Moreno, and Yonghui Wu · 2018
Cited alongside, same era.
Efficient neural audio synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron van den Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post · 2018
Cited alongside, same era.
Generalized end-to-end loss for speaker verification
Jonathan Shen, Ye Jia, Mike Chrzanowski, Yu Zhang, Isaac Elias, Heiga Zen, and Yonghui Wu · 2020
Later among the works it cites.
CoVoST: A diverse multilingual speech-to-text translation corpus
Changhan Wang, Juan Pino, Anne Wu, and Jiatao Gu · 2020
Later among the works it cites.
XLS-R: Self-supervised cross-lingual speech representation learning at scale
Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick von Platen, Yatharth Saraf, Juan Pino, et al · 2021
Later among the works it cites.
Multimodal and multilingual embeddings for large-scale speech mining
Paul-Ambroise Duquenne, Hongyu Gong, and Holger Schwenk · 2021
Later among the works it cites.
Is" moby dick" a whale or a bird? named entities and terminology in speech translation
Marco Gaido, Susana Rodríguez, Matteo Negri, Luisa Bentivogli, and Marco Turchi · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li Wan, Quan Wang, Alan Papir, and Ignacio Lopez Moreno · 2018
Cited alongside, same era.
Pre-training on high-resource speech recognition improves low-resource speech-to-text translation
Sameer Bansal, Herman Kamper, Karen Livescu, Adam Lopez, and Sharon Goldwater · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Lingvo: A modular and scalable framework for sequence-to-sequence modeling
Jonathan Shen, Patrick Nguyen, Yonghui Wu, Zhifeng Chen, et al · 2019
Cited alongside, same era.
Speech-to-speech translation between untranscribed unknown languages
Andros Tjandra, Sakriani Sakti, and Satoshi Nakamura · 2019
Cited alongside, same era.
LibriTTS: A corpus derived from LibriSpeech for text-to-speech
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu · 2019
Cited alongside, same era.
Learning to speak fluently in a foreign language: Multilingual speech synthesis and cross-language voice cloning
Yu Zhang, Ron J Weiss, Heiga Zen, Yonghui Wu, Zhifeng Chen, RJ Skerry-Ryan, Ye Jia, Andrew Rosenberg, and Bhuvana Ramabhadran · 2019
Cited alongside, same era.
Later among the works it cites.
PnG BERT: Augmented BERT on phonemes and graphemes for neural TTS
Ye Jia, Heiga Zen, Jonathan Shen, Yu Zhang, and Yonghui Wu · 2021
Later among the works it cites.
Transformer-based direct speech-to-speech translation with transcoder
Takatomo Kano, Sakriani Sakti, and Satoshi Nakamura · 2021
Later among the works it cites.
Textless speech-to-speech translation on real data
Ann Lee, Hongyu Gong, Paul-Ambroise Duquenne, Holger Schwenk, Peng-Jen Chen, Changhan Wang, Sravya Popuri, Juan Pino, Jiatao Gu, and Wei-Ning Hsu · 2021
Later among the works it cites.
Multilingual speech translation with efficient finetuning of pretrained models
Xian Li, Changhan Wang, Yun Tang, Chau Tran, Yuqing Tang, Juan Pino, Alexei Baevski, Alexis Conneau, and Michael Auli · 2021
Later among the works it cites.
Direct simultaneous speech to speech translation
Xutai Ma, Hongyu Gong, Danni Liu, Ann Lee, Yun Tang, Peng-Jen Chen, Wei-Ning Hsu, Kenneth Heafield, Phillip Koehn, and Juan Pino · 2021
Later among the works it cites.
Dr-Vectors: Decision residual networks and an improved loss for speaker recognition
Jason Pelecanos, Quan Wang, and Ignacio Lopez Moreno · 2021
Later among the works it cites.
Optimally encoding inductive biases into the transformer improves end-to-end speech translation
Piyush Vyas, Anastasia Kuznetsova, and Donald S Williamson · 2021
Later among the works it cites.
UWSpeech: Speech to speech translation for unwritten languages
Chen Zhang, Xu Tan, Yi Ren, Tao Qin, Kejun Zhang, and Tie-Yan Liu · 2021
Later among the works it cites.
Translatotron 2: High-quality direct speech-to-speech translation with voice preservation
Ye Jia, Michelle Tadmor Ramanovich, Tal Remez, and Roi Pomerantz · 2022
Closest in time.
Direct speech-to-speech translation with discrete units
Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, and Wei-Ning Hsu · 2022
Closest in time.
Parameter-free attentive scoring for speaker verification
Jason Pelecanos, Quan Wang, Yiling Huang, and Ignacio Lopez Moreno · 2022
Closest in time.
Attentive temporal pooling for conformer-based streaming language identification in long-form speech
Quan Wang, Yang Yu, Jason Pelecanos, Yiling Huang, and Ignacio Lopez Moreno · 2022
Closest in time.