Fetching the paper…
Reading the bibliography…
We introduce a new approach for speech pre-training named SPIRAL which works by learning denoising representation of perturbed data in a teacher-student framework.
Effectiveness of self-supervised pre-training for speech recognition, 2019
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed · 1911
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification
S. Chopra, R. Hadsell, and Y. LeCun · 2005
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Product quantization for nearest neighbor search
Herve Jégou, Matthijs Douze, and Cordelia Schmid · 2011
Earlier work this paper cites.
An investigation of deep neural networks for noise robust speech recognition
Michael L. Seltzer, Dong Yu, and Yongqiang Wang · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
The third ‘CHiME’ speech separation and recognition challenge: Dataset, task and baselines
Jon Barker, Ricard Marxer, Emmanuel Vincent, and Shinji Watanabe · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Audio augmentation for speech recognition
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Semi-supervised maximum mutual information training of deep neural network acoustic models
Vimal Manohar, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Improved deep metric learning with multi-class N-pair loss objective
Kihyuk Sohn · 2016
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Antti Tarvainen and Harri Valpola · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2018
Cited alongside, same era.
Learning noise-invariant representations for robust speech recognition
Davis Liang, Zhiheng Huang, and Zachary C Lipton · 2018
Cited alongside, same era.
Low latency acoustic modeling using temporal convolution and LSTMs
Vijayaditya Peddinti, Yiming Wang, Daniel Povey, and Sanjeev Khudanpur · 2018
Cited alongside, same era.
Unsupervised feature learning via non-parametric instance discrimination
Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin · 2018
Cited alongside, same era.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Self-training for end-to-end speech recognition
Jacob Kahn, Ann Lee, and Awni Hannun · 2020
Later among the works it cites.
Libri-light: A benchmark for ASR with limited or no supervision
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al · 2020
Later among the works it cites.
Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders
Andy T. Liu, Shu-wen Yang, Po-Han Chi, Po-chun Hsu, and Hung-yi Lee · 2020
Later among the works it cites.
Specaugment on large scale datasets
Daniel S. Park, Yu Zhang, Chung-Cheng Chiu, Youzheng Chen, Bo Li, William Chan, Quoc V. Le, and Yonghui Wu · 2020
Later among the works it cites.
Improved noisy student training for automatic speech recognition
Daniel S. Park, Yu Zhang, Ye Jia, Wei Han, Chung-Cheng Chiu, Bo Li, Yonghui Wu, and Quoc V. Le · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
RWTH ASR Systems for LibriSpeech: Hybrid vs Attention
Christoph Lüscher, Eugen Beck, Kazuki Irie, Markus Kitza, Wilfried Michel, Albert Zeyer, Ralf Schlüter, and Hermann Ney · 2019
Cited alongside, same era.
Wav2Letter++: A fast open-source speech recognition system
Vineel Pratap, Awni Hannun, Qiantong Xu, Jeff Cai, Jacob Kahn, Gabriel Synnaeve, Vitaliy Liptchinsky, and Ronan Collobert · 2019
Cited alongside, same era.
Representation learning with contrastive predictive coding, 2019
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2019
Cited alongside, same era.
vq-wav2vec: Self-supervised learning of discrete speech representations
Alexei Baevski, Steffen Schneider, and Michael Auli · 2020
Cited alongside, same era.
Semi-supervised ASR by end-to-end self-training
Yang Chen, Weiran Wang, and Chao Wang · 2020
Cited alongside, same era.
Generative pre-training for speech with autoregressive predictive coding
Yu-An Chung and James Glass · 2020
Cited alongside, same era.
Gabriel Synnaeve, Qiantong Xu, Jacob Kahn, Tatiana Likhomanenko, Edouard Grave, Vineel Pratap, Anuroop Sriram, Vitaliy Liptchinsky, and Ronan Collobert · 2020
Later among the works it cites.
End-to-end ASR: from supervised to semi-supervised learning with modern architectures
Gabriel Synnaeve, Qiantong Xu, Jacob Kahn, Tatiana Likhomanenko, Edouard Grave, Vineel Pratap, Anuroop Sriram, Vitaliy Liptchinsky, and Ronan Collobert · 2020
Later among the works it cites.
Unsupervised pre-training of bidirectional speech encoders via masked reconstruction
Weiran Wang, Qingming Tang, and Karen Livescu · 2020
Later among the works it cites.
Self-training with noisy student improves imagenet classification
Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V. Le · 2020
Later among the works it cites.
Iterative pseudo-labeling for speech recognition
Qiantong Xu, Tatiana Likhomanenko, Jacob Kahn, Awni Hannun, Gabriel Synnaeve, and Ronan Collobert · 2020
Later among the works it cites.
Pushing the limits of semi-supervised learning for automatic speech recognition
Yu Zhang, James Qin, Daniel S. Park, Wei Han, Chung-Cheng Chiu, Ruoming Pang, Quoc V. Le, and Yonghui Wu · 2020
Later among the works it cites.
Exploring simple Siamese representation learning
Xinlei Chen and Kaiming He · 2021
Later among the works it cites.
The People’s Speech: A large-scale diverse english speech recognition dataset for commercial usage
Daniel Galvez, Greg Diamos, Juan Manuel Ciro Torres, Keith Achorn, Anjali Gopi, David Kanter, Max Lam, Mark Mazumder, and Vijay Janapa Reddi · 2021
Later among the works it cites.
Hubert: How much can a bad teacher benefit ASR pre-training?
Wei-Ning Hsu, Yao-Hung Hubert Tsai, Benjamin Bolte, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Later among the works it cites.
ICASSP 2021 deep noise suppression challenge
Chandan KA Reddy, Harishchandra Dubey, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, and Sriram Srinivasan · 2021
Later among the works it cites.
Understanding self-supervised learning dynamics without contrastive pairs
Yuandong Tian, Xinlei Chen, and Surya Ganguli · 2021
Later among the works it cites.
Contrastive semi-supervised learning for asr
Alex Xiao, Christian Fuegen, and Abdelrahman Mohamed · 2021
Later among the works it cites.
Self-training and pre-training are complementary for speech recognition
Qiantong Xu, Alexei Baevski, Tatiana Likhomanenko, Paden Tomasello, Alexis Conneau, Ronan Collobert, Gabriel Synnaeve, and Michael Auli · 2021
Later among the works it cites.