Fetching the paper…
Reading the bibliography…
We present two end-to-end models: Audio-to-Byte (A2B) and Byte-to-Audio (B2A), for multilingual speech recognition and synthesis.
“P. 800: Methods for subjective determination of transmission quality,”
ITUT Rec, · 1996
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Context-dependent acoustic modeling using graphemes for large vocabulary speech recognition,”
Stephan Kanthak and Hermann Ney, · 2002
Earlier work this paper cites.
“Grapheme based speech recognition,”
Mirjam Killer, Sebastian Stuker, and Tanja Schultz, · 2003
Earlier work this paper cites.
“A grapheme based speech recognition system for russian,”
Sebastian Stüker and Tanja Schultz, · 2004
Earlier work this paper cites.
Multilingual speech processing
Tanja Schultz and Katrin Kirchhoff, · 2006
Earlier work this paper cites.
“Current trends in multilingual speech processing,”
Hervé Bourlard, John Dines, Mathew Magimai-Doss, Philip N Garner, David Imseng, Petr Motlicek, Hui Liang, Lakshmi Saheer, and Fabio Valente, · 2011
Earlier work this paper cites.
“Comparing grapheme-based and phoneme-based speech recognition for afrikaans,”
Willem D Basson and Marelie H Davel, · 2012
Earlier work this paper cites.
Code-switching in conversation: Language, interaction and identity
Peter Auer, · 2013
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Earlier work this paper cites.
“Unicode-based graphemic systems for limited resource languages,”
Mark JF Gales, Kate M Knill, and Anton Ragni, · 2015
Earlier work this paper cites.
“Neural Machine Translation by Jointly Learning to Align and Translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2015
Cited alongside, same era.
“Fast and accurate recurrent neural network acoustic models for speech recognition,”
Haşim Sak, Andrew Senior, Kanishka Rao, and Françoise Beaufays, · 2015
Cited alongside, same era.
“Listen, Attend and Spell: A Neural Network for Large Vocabulary Conversational Speech Recognition ,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Cited alongside, same era.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2016
Cited alongside, same era.
“On Online Attention-based Speech Recognition and Joint Mandarin Character-Pinyin Training,”
William Chan and Ian Lane, · 2016
Cited alongside, same era.
“Tacotron: A fully end-to-end text-to-speech synthesis model,”
Yuxuan Wang, RJ Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, et al., · 2017
Later among the works it cites.
“A comparison of sequence-to-sequence models for speech recognition,”
Rohit Prabhavalkar, Kanishka Rao, Tara N Sainath, Bo Li, Leif Johnson, and Navdeep Jaitly, · 2017
Later among the works it cites.
“Generated of Large-scale Simulated Utterances in Virtual Rooms to Train Deep-Neural Networks for Far-field Speech Recognition in Google Home,”
C. Kim, A. Misra, K. Chin, T. Hughes, A. Narayanan, T. N. Sainath, and M. Bacchiani, · 2017
Later among the works it cites.
“State-of-the-art Speech Recognition With Sequence-to-Sequence Models,”
Chung-Cheng Chiu, Tara N. Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J. Weiss, Kanishka Rao, Ekaterina Gonina, Navdeep Jaitly, Bo Li, Jan Chorowsk, and Michiel Bacchiani, · 2018
Closest in time.
“Improved training of end-to-end attention models for speech recognition,”
Albert Zeyer, Kazuki Irie, Ralf Schluter, and Hermann Ney, · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hagen Soltau, Hank Liao, and Hasim Sak, · 2016
Cited alongside, same era.
“Multilingual language processing from bytes,”
Dan Gillick, Cliff Brunk, Oriol Vinyals, and Amarnag Subramanya, · 2016
Cited alongside, same era.
“Lower Frame Rate Neural Network Acoustic Models,”
Golan Pundak and Tara N Sainath, · 2016
Cited alongside, same era.
“Latent Sequence Decompositions,”
William Chan, Yu Zhang, Quoc Le, and Navdeep Jaitly, · 2017
Cited alongside, same era.
“Exploring Architectures, Data and Units for Streaming End-to-End Speech Recognition with RNN-Transducer,”
K. Rao, R. Prabhavalkar, and H. Sak, · 2017
Cited alongside, same era.
“Char2wav: End-to-end speech synthesis,”
Jose Sotelo, Soroush Mehri, Kundan Kumar, Joao Felipe Santos, Kyle Kastner, Aaron Courville, and Yoshua Bengio, · 2017
Cited alongside, same era.
Closest in time.
“Advancing Acoustic-to-Word CTC Model,”
Jinyu Li, Guoli Ye, Amit Das, Rui Zhao, and Yifan Gong, · 2018
Closest in time.
“Natural TTS synthesis by conditioning wavenet on mel spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, RJ Skerry-Ryan, et al., · 2018
Closest in time.
“Multi-dialect speech recognition with a single sequence-to-sequence model,”
Bo Li, Tara N Sainath, Khe Chai Sim, Michiel Bacchiani, Eugene Weinstein, Patrick Nguyen, Zhifeng Chen, Yonghui Wu, and Kanishka Rao, · 2018
Closest in time.
“Multilingual speech recognition with a single end-to-end model,”
Shubham Toshniwal, Tara N Sainath, Ron J Weiss, Bo Li, Pedro Moreno, Eugene Weinstein, and Kanishka Rao, · 2018
Closest in time.
“Efficient neural audio synthesis,”
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aäron van den Oord, Sander Dieleman, and Koray Kavukcuoglu, · 2018
Closest in time.