Fetching the paper…
Reading the bibliography…
We employ a combination of recent developments in semi-supervised learning for automatic speech recognition to obtain state-of-the-art results on LibriSpeech utilizing the unlabeled audio of the Libri-Light dataset.
Jasper: An end-to-end convolutional neural acoustic model
Jason Li, Vitaly Lavrukhin, Boris Ginsburg, Ryan Leary, Oleksii Kuchaiev, Jonathan M Cohen, Huyen Nguyen, and Ravi Teja Gadde · 1904
Earlier work this paper cites.
End-to-end asr: from supervised to semi-supervised learning with modern architectures
Gabriel Synnaeve, Qiantong Xu, Jacob Kahn, Edouard Grave, Tatiana Likhomanenko, Vineel Pratap, Anuroop Sriram, Vitaliy Liptchinsky, and Ronan Collobert · 1911
Earlier work this paper cites.
Probability of error of some adaptive pattern-recognition machines
H Scudder · 1965
Earlier work this paper cites.
Unsupervised word sense disambiguation rivaling supervised methods
David Yarowsky · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Utilizing untranscribed training data to improve performance
George Zavaliagkos and Thomas Colthurst · 1998
Earlier work this paper cites.
Lightly supervised acoustic model training
Lori Lamel, Jean luc Gauvain, and Gilles Adda · 2000
Earlier work this paper cites.
Learning extraction patterns for subjective expressions
Ellen Riloff and Janyce Wiebe · 2003
Earlier work this paper cites.
Improved noisy student training for automatic speech recognition
Daniel S Park, Yu Zhang, Ye Jia, Wei Han, Chung-Cheng Chiu, Bo Li, Yonghui Wu, and Quoc V Le · 2005
Earlier work this paper cites.
Analysis of low-resource acoustic model self-training
Scott Novotney and Richard Schwartz · 2009
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves · 2012
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima · 2012
Earlier work this paper cites.
Deep neural network features and semi-supervised training for low resource speech recognition
Samuel Thomas, Michael L Seltzer, Kenneth Church, and Hynek Hermansky · 2013
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
On using monolingual corpora in neural machine translation
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loic Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2015
Earlier work this paper cites.
Fast and accurate recurrent neural network acoustic models for speech recognition
Haşim Sak, Andrew Senior, Kanishka Rao, and Françoise Beaufays · 2015
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Earlier work this paper cites.
Letter-based speech recognition with gated convnets
Vitaliy Liptchinsky, Gabriel Synnaeve, and Ronan Collobert · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Cited alongside, same era.
Neural network language modeling with letter-based features and importance sampling
Hainan Xu, Ke Li, Yiming Wang, Jian Wang, Shiyin Kang, Xie Chen, Daniel Povey, and Sanjeev Khudanpur · 2018
Cited alongside, same era.
Extracting domain invariant features by unsupervised learning for robust automatic speech recognition
Wei-Ning Hsu and James Glass · 2018
Cited alongside, same era.
Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech
Yu-An Chung and James Glass · 2018
Cited alongside, same era.
vq-wav2vec: Self-supervised learning of discrete speech representations
Alexei Baevski, Steffen Schneider, and Michael Auli · 2019
Later among the works it cites.
Revisiting self-training for neural sequence generation
Junxian He, Jiatao Gu, Jiajun Shen, and Marc’Aurelio Ranzato · 2019
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Libri-light: A benchmark for asr with limited or no supervision
J. Kahn, M. Riviere, W. Zheng, E. Kharitonov, Q. Xu, P.E. Mazare, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, and et al · 2020
Closest in time.
Big self-supervised models are strong semi-supervised learners
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Cited alongside, same era.
Specaugment: A simple data augmentation method for automatic speech recognition
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le · 2019
Cited alongside, same era.
Rwth asr systems for librispeech: Hybrid vs attention–w/o data augmentation
Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, and Hermann Ney · 2019
Cited alongside, same era.
State-of-the-art speech recognition using multi-stream self-attention with dilated 1d convolutions
Kyu J Han, Ramon Prieto, and Tao Ma · 2019
Cited alongside, same era.
0-1 phase transitions in sparse spiked matrix estimation
Jean Barbier and Nicolas Macris · 2019
Cited alongside, same era.
Semi-supervised training for end-to-end models via weak distillation
Bo Li, Tara N Sainath, Ruoming Pang, and Zelin Wu · 2019
Cited alongside, same era.
Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey Hinton · 2020
Closest in time.
Rethinking pre-training and self-training
Barret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui, Hanxiao Liu, Ekin D Cubuk, and Quoc V Le · 2020
Closest in time.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al · 2020
Closest in time.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Closest in time.
Self-training with noisy student improves imagenet classification
Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V Le · 2020
Closest in time.
Specaugment on large scale datasets
Daniel S Park, Yu Zhang, Chung-Cheng Chiu, Youzheng Chen, Bo Li, William Chan, Quoc V Le, and Yonghui Wu · 2020
Closest in time.
https://paperswithcode.com/sota/speech-recognition-on-librispeech-test-clean , October 8, 2020a
Speech recognition on librispeech test-clean · 2020
Closest in time.
https://paperswithcode.com/sota/speech-recognition-on-librispeech-test-other , October 8, 2020b
Speech recognition on librispeech test-other · 2020
Closest in time.
Wei Han, Zhengdong Zhang, Yu Zhang, Jiahui Yu, Chung-Cheng Chiu, James Qin, Anmol Gulati, Ruoming Pang, and Yonghui Wu · 2020
Closest in time.
Self-training for end-to-end speech recognition
Jacob Kahn, Ann Lee, and Awni Hannun · 2020
Closest in time.
Semi-supervised speech recognition via local prior matching
Wei-Ning Hsu, Ann Lee, Gabriel Synnaeve, and Awni Hannun · 2020
Closest in time.
Fixmatch: Simplifying semi-supervised learning with consistency and confidence
Kihyuk Sohn, David Berthelot, Chun-Liang Li, Zizhao Zhang, Nicholas Carlini, Ekin D Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel · 2020
Closest in time.
Deep contextualized acoustic representations for semi-supervised speech recognition
Shaoshi Ling, Yuzong Liu, Julian Salazar, and Katrin Kirchhoff · 2020
Closest in time.
Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss
Qian Zhang, Han Lu, Hasim Sak, Anshuman Tripathi, Erik McDermott, Stephen Koo, and Shankar Kumar · 2020
Closest in time.