Fetching the paper…
Reading the bibliography…
This paper is a study of performance-efficiency trade-offs in pre-trained models for automatic speech recognition (ASR).
vq-wav2vec: Self-supervised learning of discrete speech representations
Alexei Baevski, Steffen Schneider, and Michael Auli · 1910
Earlier work this paper cites.
Statistical theory of extreme values and some practical applications: a series of lectures , volume 33
Emil Julius Gumbel · 1954
Earlier work this paper cites.
Error bounds for convolutional codes and an asymptotically optimum decoding algorithm
A. Viterbi · 1967
Earlier work this paper cites.
Switchboard-1 release 2 ldc97s62
John Godfrey and Edward Holliman · 1993
Earlier work this paper cites.
2000 hub5 english evaluation speech ldc2002s09
Linguistic Data Consortium · 2002
Earlier work this paper cites.
Fisher english training speech parts 1 and 2 ldc200{4,5}s13
Christopher Cieri, David Graff, Owen Kimball, Dave Miller, and Kevin Walker · 2004
Earlier work this paper cites.
Fisher english training speech parts 1 and 2 transcripts ldc200{4,5}t19
Christopher Cieri, David Graff, Owen Kimball, Dave Miller, and Kevin Walker · 2004
Earlier work this paper cites.
Iterative pseudo-labeling for speech recognition
Qiantong Xu, Tatiana Likhomanenko, Jacob Kahn, Awni Hannun, Gabriel Synnaeve, and Ronan Collobert · 2005
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli · 2006
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, Santiago Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
2003 nist rich transcription evaluation data ldc2007s10
et al Fiscus, Jonathan G · 2007
Earlier work this paper cites.
Recurrent neural network based language model
Tomáš Mikolov, Martin Karafiát, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Self-training and pre-training are complementary for speech recognition
Qiantong Xu, Alexei Baevski, T. Likhomanenko, Paden Tomasello, Alexis Conneau, Ronan Collobert, Gabriel Synnaeve, and Michael Auli · 2010
Earlier work this paper cites.
Pushing the limits of semi-supervised learning for automatic speech recognition
Y. Zhang, James Qin, D. Park, Wei Han, C. Chiu, Ruoming Pang, Quoc V. Le, and Yonghui Wu · 2010
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely · 2011
Earlier work this paper cites.
Applying convolutional neural networks concepts to hybrid nn-hmm model for speech recognition
Ossama Abdel-Hamid, Abdel-rahman Mohamed, Hui Jiang, and Gerald Penn · 2012
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Chris J Maddison, Daniel Tarlow, and Tom Minka · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Cited alongside, same era.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals · 2016
Cited alongside, same era.
Wav2letter: an end-to-end convnet-based speech recognition system
Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve · 2016
Cited alongside, same era.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q. Weinberger · 2016
Cited alongside, same era.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, S. Gross, Francisco Massa, A. Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Z. Lin, N. Gimelshein, L. Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli · 2019
Later among the works it cites.
Speech-xlnet: Unsupervised acoustic model pretraining for self-attention networks
Xingchen Song, Guangsen Wang, Zhiyong Wu, Yiheng Huang, Dan Su, Dong Yu, and Helen Meng · 2019
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V. Le · 2019
Later among the works it cites.
Transformer-transducer: End-to-end speech recognition with self-attention
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A temporal coherence loss function for learning unsupervised acoustic embeddings
Gabriel Synnaeve and Emmanuel Dupoux · 2016
Cited alongside, same era.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross B. Girshick, P. Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Y. Jia, and Kaiming He · 2017
Cited alongside, same era.
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Unsupervised cross-modal alignment of speech and text embedding spaces
Yu-An Chung, Wei-Hung Weng, Schrasing Tong, and James Glass · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition
Linhao Dong, Shuang Xu, and Bo Xu · 2018
Cited alongside, same era.
Ching-Feng Yeh, Jay Mahadeokar, Kaustubh Kalgaonkar, Yongqiang Wang, Duc Le, Mahaveer Jain, Kjell Schubert, Christian Fuegen, and Michael L Seltzer · 2019
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey E. Hinton · 2020
Later among the works it cites.
Unsupervised cross-lingual representation learning for speech recognition
Alexis Conneau, Alexei Baevski, Ronan Collobert, Abdel rahman Mohamed, and Michael Auli · 2020
Later among the works it cites.
Reducing transformer depth on demand with structured dropout
Angela Fan, Edouard Grave, and Armand Joulin · 2020
Later among the works it cites.
Bootstrap your own latent: A new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, et al · 2020
Later among the works it cites.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, C. Chiu, Niki Parmar, Y. Zhang, Jiahui Yu, Wei Han, S. Wang, Z. Zhang, Yonghui Wu, and Ruoming Pang · 2020
Later among the works it cites.
Wei Han, Z. Zhang, Y. Zhang, Jiahui Yu, C. Chiu, James Qin, Anmol Gulati, Ruoming Pang, and Yonghui Wu · 2020
Later among the works it cites.
Deberta: Decoding-enhanced bert with disentangled attention
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen · 2020
Later among the works it cites.
Hubert: How much can a bad teacher benefit asr pre-training
Wei-Ning Hsu, Yao-Hung Hubert Tsai, Benjamin Bolte, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2020
Later among the works it cites.
Libri-light: A benchmark for asr with limited or no supervision
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux · 2020
Later among the works it cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Fast and accurate model scaling
Piotr Dollár, Mannat Singh, and Ross B. Girshick · 2021
Closest in time.
Psla: Improving audio event classification with pretraining, sampling, labeling, and aggregation
Yuan Gong, Yu-An Chung, and James Glass · 2021
Closest in time.
Robust wav2vec 2.0: Analyzing domain shift in self-supervised pre-training
Wei-Ning Hsu, Anuroop Sriram, Alexei Baevski, T. Likhomanenko, Qiantong Xu, Vineel Pratap, Jacob Kahn, Ann Lee, Ronan Collobert, Gabriel Synnaeve, and Michael Auli · 2021
Closest in time.
Byol for audio: Self-supervised learning for general-purpose audio representation
Daisuke Niizumi, Daiki Takeuchi, Yasunori Ohishi, Noboru Harada, and Kunio Kashino · 2021
Closest in time.
Emotion recognition from speech using wav2vec 2.0 embeddings
Leonardo Pepino, P. Riera, and L. Ferrer · 2021
Closest in time.
Contrastive learning of general-purpose audio representations
Aaqib Saeed, David Grangier, and Neil Zeghidour · 2021
Closest in time.
A deeper look at sheet music composer classification using self-supervised pretraining
Daniel Yang, Kevin Ji, and TJ Tsai · 2021
Closest in time.
Musicoder: A universal music-acoustic encoder based on transformer
Yilun Zhao and Jia Guo · 2021
Closest in time.