Fetching the paper…
Reading the bibliography…
Recent publications on automatic-speech-recognition (ASR) have a strong focus on attention encoder-decoder (AED) architectures which tend to suffer from over-fitting in low resource scenarios.
“Speech synthesis from short-time fourier transform magnitude and its application to speech processing,”
Daniel W. Griffin, Douglas S. Deadrick, and Jae S. Lim, · 1984
Earlier work this paper cites.
“The general use of tying in phoneme-based HMM speech recognisers,”
Steve J Young, · 1992
Earlier work this paper cites.
“Long short-term memory,”
Sepp Hochreiter and Jürgen Schmidhuber, · 1997
Earlier work this paper cites.
“Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, and Faustino Gomez, · 2006
Earlier work this paper cites.
“Joint-sequence models for grapheme-to-phoneme conversion,”
Maximilian Bisani and Hermann Ney, · 2008
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Earlier work this paper cites.
“RASR/NN: The RWTH neural network toolkit for speech recognition,”
Simon Wiesler, Alexander Richard, Pavel Golik, Ralf Schluter, and Hermann Ney, · 2014
Earlier work this paper cites.
“Dropout: A simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Earlier work this paper cites.
“Librispeech: An asr corpus based on public domain audio books.,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Zoneout: Regularizing RNNs by randomly preserving hidden activations,”
David Krueger, Tegan Maharaj, János Kramár, Mohammad Pezeshki, Nicolas Ballas, Nan Rosemary Ke, Anirudh Goyal, Yoshua Bengio, Aaron Courville, and Chris Pal, · 2016
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Earlier work this paper cites.
“Returnn: The RWTH extensible training framework for universal recurrent neural networks,”
Patrick Doetsch, Albert Zeyer, Paul Voigtlaender, Ilia Kulikov, Ralf Schlüter, and Hermann Ney, · 2017
Earlier work this paper cites.
“Natural TTS synthesis by conditioning wavenet on MEL spectrogram predictions,”
Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, Rif A. Saurous, Yannis Agiomvrgiannakis, and Yonghui Wu, · 2018
Earlier work this paper cites.
“Sisyphus, a workflow manager designed for machine translation and automatic speech recognition,”
Jan-Thorsten Peter, Eugen Beck, and Hermann Ney, · 2018
Earlier work this paper cites.
“Improved training of end-to-end attention models for speech recognition,”
Albert Zeyer, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2018
Earlier work this paper cites.
“Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,”
Yuxuan Wang, Daisy Stanton, Yu Zhang, RJ Skerry-Ryan, Eric Battenberg, Joel Shor, Ying Xiao, Fei Ren, Ye Jia, and Rif A. Saurous, · 2018
Cited alongside, same era.
“X-vectors: Robust dnn embeddings for speaker recognition,”
David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Cited alongside, same era.
“A comparative study on transformer vs RNN in speech applications,”
Shigeki Karita, Nanxin Chen, Tomoki Hayashi, Takaaki Hori, Hirofumi Inaguma, Ziyan Jiang, Masao Someki, Nelson Enrique Yalta Soplin, Ryuichi Yamamoto, Xiaofei Wang, Shinji Watanabe, Takenori Yoshimura, and Wangyou Zhang, · 2019
Cited alongside, same era.
“RWTH ASR systems for librispeech: Hybrid vs attention,”
Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
“Generating synthetic audio data for attention-based speech recognition systems,”
Nick Rossenbach, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2020
Later among the works it cites.
“You do not need more data: Improving end-to-end speech recognition by text-to-speech data augmentation,”
Aleksandr Laptev, Roman Korostik, Aleksey Svischev, Andrei Andrusenko, Ivan Medennikov, and Sergey Rybin, · 2020
Later among the works it cites.
“Self-training for end-to-end speech recognition,”
Jacob Kahn, Ann Lee, and Awni Hannun, · 2020
Later among the works it cites.
“Teacher-student training for robust tacotron-based TTS,”
Rui Liu, Berrak Sisman, Jingdong Li, Feilong Bao, Guanglai Gao, and Haizhou Li, · 2020
Later among the works it cites.
Jonathan Shen, Ye Jia, Mike Chrzanowski, Yu Zhang, Isaac Elias, Heiga Zen, and Yonghui Wu, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 2019
Cited alongside, same era.
“Semi-supervised sequence-to-sequence ASR using unpaired speech and text,”
Murali Karthick Baskar, Shinji Watanabe, Ramon Astudillo, Takaaki Hori, Lukáš Burget, and Jan Černocký, · 2019
Cited alongside, same era.
“Speech recognition with augmented synthesized speech,”
Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran, Ye Jia, Pedro Moreno, Yonghui Wu, and Zelin Wu, · 2019
Cited alongside, same era.
“LibriTTS: A corpus derived from librispeech for text-to-speech,”
Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu, · 2019
Cited alongside, same era.
“A density ratio approach to language model fusion in end-to-end automatic speech recognition,”
Erik McDermott, Hasim Sak, and Ehsan Variani, · 2019
Cited alongside, same era.
“A comparison of transformer and lstm encoder decoder models for asr,”
Albert Zeyer, Parnia Bahar, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2019
Cited alongside, same era.
“Language modeling with deep transformers,”
Kazuki Irie, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Cited alongside, same era.
“Developing RNN-t models surpassing high-performance hybrid models with customization capability,”
Jinyu Li, Rui Zhao, Zhong Meng, Yanqing Liu, Wenning Wei, Sarangarajan Parthasarathy, Vadim Mazalov, Zhenghao Wang, Lei He, Sheng Zhao, and Yifan Gong, · 2020
Later among the works it cites.
“Hybrid autoregressive transducer (HAT),”
Ehsan Variani, David Rybach, Cyril Allauzen, and Michael Riley, · 2020
Later among the works it cites.
“slimIPL: Language-model-free iterative pseudo-labeling,”
Tatiana Likhomanenko, Qiantong Xu, Jacob Kahn, Gabriel Synnaeve, and Ronan Collobert, · 2020
Later among the works it cites.
“Rnn-transducer with stateless prediction network,”
Mohammadreza Ghodsi, Xiaofeng Liu, James Apfel, Rodrigo Cabrera, and Eugene Weinstein, · 2020
Later among the works it cites.
“Eat: Enhanced ASR-TTS for self-supervised speech recognition,”
Murali Karthick Baskar, Lukas Burget, Shinji Watanabe, Ramon Fernandez Astudillo, and Jan Honza Cernocky, · 2021
Closest in time.
“Synthasr: Unlocking synthetic data for speech recognition,”
Amin Fazel, Wei Yang, Yulan Liu, Roberto Barra-Chicote, Yixiong Meng, Roland Maas, and Jasha Droppo, · 2021
Closest in time.
“Internal language model estimation for domain-adaptive end-to-end speech recognition,”
Zhong Meng, S. Parthasarathy, Eric Sun, Yashesh Gaur, Naoyuki Kanda, Liang Lu, Xie Chen, Rui Zhao, Jinyu Li, and Y. Gong, · 2021
Closest in time.
“Investigating methods to improve language model integration for attention-based encoder-decoder asr models,”
Mohammad Zeineldeen, Aleksandr Glushko, Wilfried Michel, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2021
Closest in time.
“Librispeech transducer model with internal language model prior correction,”
Albert Zeyer, André Merboldt, Wilfried Michel, Ralf Schlüter, and Hermann Ney, · 2021
Closest in time.
“Phoneme based neural transducer for large vocabulary speech recognition,”
Wei Zhou, Simon Berger, Ralf Schluter, and Hermann Ney, · 2021
Closest in time.