Fetching the paper…
Reading the bibliography…
We revisit self-training in the context of end-to-end speech recognition.
“Probability of error of some adaptive pattern-recognition machines,”
H Scudder, · 1965
Earlier work this paper cites.
“Unsupervised training of a speech recognizer: Recent experiments,”
Thomas Kemp and Alex Waibel, · 1999
Earlier work this paper cites.
“Confidence-measure-driven unsupervised incremental adaptation for hmm-based speech recognition,”
Delphine Charlet, · 2001
Earlier work this paper cites.
“NLTK: The natural language toolkit,”
Edward Loper and Steven Bird, · 2002
Earlier work this paper cites.
“Unsupervised training of acoustic models for large vocabulary continuous speech recognition,”
Frank Wessel and Hermann Ney, · 2004
Earlier work this paper cites.
“Unsupervised versus supervised training of acoustic models,”
Jeff Ma and Richard Schwartz, · 2008
Earlier work this paper cites.
“Semi-supervised training of deep neural networks,”
Karel Veselỳ, Mirko Hannemann, and Lukáš Burget, · 2013
Earlier work this paper cites.
“Learning phrase representations using RNN encoder-decoder for statistical machine translation,”
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“LibriSpeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2015
Cited alongside, same era.
“High quality agreement-based semi-supervised training data for acoustic modeling,”
Félix de Chaumont Quitry, Asa Oines, Pedro Moreno, and Eugene Weinstein, · 2016
Cited alongside, same era.
“Semi-supervised DNN training with word selection for ASR,”
Karel Veselỳ, Lukás Burget, and Jan Cernockỳ, · 2017
Cited alongside, same era.
“Towards better decoding and language model integration in sequence to sequence models,”
Jan Chorowski and Navdeep Jaitly, · 2017
Cited alongside, same era.
“SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,”
Taku Kudo and John Richardson, · 2018
Cited alongside, same era.
“Improved training of end-to-end attention models for speech recognition,”
Albert Zeyer, Kazuki Irie, Ralf Schlüter, and Hermann Ney, · 2018
Later among the works it cites.
“RWTH ASR systems for LibriSpeech: Hybrid vs attention,”
Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Closest in time.
“Sequence-to-sequence speech recognition with time-depth separable convolutions,”
Awni Hannun, Ann Lee, Qiantong Xu, and Ronan Collobert, · 2019
Closest in time.
“Cycle-consistency training for end-to-end speech recognition,”
Takaaki Hori, Ramon Astudillo, Tomoki Hayashi, Yu Zhang, Shinji Watanabe, and Jonathan Le Roux, · 2019
Closest in time.
“Semi-supervised sequence-to-sequence ASR using unpaired speech and text,”
Murali Karthick Baskar, Shinji Watanabe, Ramon Astudillo, Takaaki Hori, Lukáš Burget, and Jan Černockỳ, · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neil Zeghidour, Qiantong Xu, Vitaliy Liptchinsky, Nicolas Usunier, Gabriel Synnaeve, and Ronan Collobert, · 2018
Cited alongside, same era.
“Back-translation-style data augmentation for end-to-end ASR,”
Tomoki Hayashi, Shinji Watanabe, Yu Zhang, Tomoki Toda, Takaaki Hori, Ramon Astudillo, and Kazuya Takeda, · 2018
Cited alongside, same era.
“Semi-supervised end-to-end speech recognition,”
Shigeki Karita, Shinji Watanabe, Tomoharu Iwata, Atsunori Ogawa, and Marc Delcroix, · 2018
Cited alongside, same era.
“Wav2letter++: A fast open-source speech recognition system,”
Vineel Pratap, Awni Hannun, Qiantong Xu, Jeff Cai, Jacob Kahn, Gabriel Synnaeve, Vitaliy Liptchinsky, and Ronan Collobert, · 2019
Closest in time.
“Lessons from building acoustic models with a million hours of speech,”
Sree Hari Krishnan Parthasarathi and Nikko Strom, · 2019
Closest in time.
“Adversarial training of end-to-end speech recognition using a criticizing language model,”
Alexander H Liu, Hung-yi Lee, and Lin-shan Lee, · 2019
Closest in time.