Fetching the paper…
Reading the bibliography…
We show for the first time that learning powerful representations from speech audio alone followed by fine-tuning on transcribed speech can outperform the best semi-supervised methods while being conceptually simpler.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 1907
Earlier work this paper cites.
Statistical theory of extreme values and some practical applications: a series of lectures , volume 33
E. J. Gumbel · 1954
Earlier work this paper cites.
The DARPA TIMIT Acoustic-Phonetic Continuous Speech Corpus CDROM
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, and N. L. Dahlgren · 1993
Earlier work this paper cites.
Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, and F. Gomez · 2006
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
M. G. A. Hyvärinen · 2010
Earlier work this paper cites.
Product quantization for nearest neighbor search
H. Jegou, M. Douze, and C. Schmid · 2011
Earlier work this paper cites.
Japanese and korean voice search
M. Schuster and K. Nakajima · 2012
Earlier work this paper cites.
A* sampling
C. J. Maddison, D. Tarlow, and T. Minka · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
Layer normalization
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Weinberger · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
E. Jang, S. Gu, and B. Poole · 2016
Earlier work this paper cites.
Ethnologue: Languages of the world, nineteenth edition
M. P. Lewis, G. F. Simon, and C. D. Fennig · 2016
Earlier work this paper cites.
Neural discrete representation learning
A. van den Oord, O. Vinyals, et al · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Adaptive input representations for neural language modeling
A. Baevski and M. Auli · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
The challenge of realistic music generation: modelling raw audio at scale
S. Dieleman, A. van den Oord, and K. Simonyan · 2018
Earlier work this paper cites.
Scaling neural machine translation
M. Ott, S. Edunov, D. Grangier, and M. Auli · 2018
Cited alongside, same era.
Deep contextualized word representations
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Cited alongside, same era.
Light gated recurrent units for speech recognition
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
A. van den Oord, Y. Li, and O. Vinyals · 2018
Cited alongside, same era.
Learning filterbanks from raw speech for phone recognition
N. Zeghidour, N. Usunier, I. Kokkinos, T. Schaiz, G. Synnaeve, and E. Dupoux · 2018
wav2vec: Unsupervised pre-training for speech recognition
S. Schneider, A. Baevski, R. Collobert, and M. Auli · 2019
Later among the works it cites.
Vqvae unsupervised unit discovery and multi-scale code2spec inverter for zerospeech challenge 2019
A. Tjandra, B. Sisman, M. Zhang, S. Sakti, H. Li, and S. Nakamura · 2019
Later among the works it cites.
Pay less attention with lightweight and dynamic convolutions
F. Wu, A. Fan, A. Baevski, Y. N. Dauphin, and M. Auli · 2019
Later among the works it cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
A. Baevski, S. Schneider, and M. Auli · 2020
Closest in time.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning representations by maximizing mutual information across views
P. Bachman, R. D. Hjelm, and W. Buchwalter · 2019
Cited alongside, same era.
Effectiveness of self-supervised pre-training for speech recognition
A. Baevski, M. Auli, and A. Mohamed · 2019
Cited alongside, same era.
Unsupervised speech representation learning using wavenet autoencoders
J. Chorowski, R. J. Weiss, S. Bengio, and A. van den Oord · 2019
Cited alongside, same era.
An unsupervised autoregressive model for speech representation learning
Y. Chung, W. Hsu, H. Tang, and J. R. Glass · 2019
Cited alongside, same era.
R. Eloff, A. Nortje, B. van Niekerk, A. Govender, L. Nortje, A. Pretorius, E. Van Biljon, E. van der Westhuizen, L. van Staden, and H. Kamper · 2019
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2019
Cited alongside, same era.
A. Fan, E. Grave, and A. Joulin · 2020
Closest in time.
Conformer: Convolution-augmented transformer for speech recognition
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang · 2020
Closest in time.
Contextnet: Improving convolutional neural networks for automatic speech recognition with global context
W. Han, Z. Zhang, Y. Zhang, J. Yu, C.-C. Chiu, J. Qin, A. Gulati, R. Pang, and Y. Wu · 2020
Closest in time.
Learning hierarchical discrete linguistic units from visually-grounded speech
D. Harwath, W.-N. Hsu, and J. Glass · 2020
Closest in time.
Libri-light: A benchmark for asr with limited or no supervision
J. Kahn et al · 2020
Closest in time.
Learning robust and multilingual speech representations
K. Kawakami, L. Wang, C. Dyer, P. Blunsom, and A. van den Oord · 2020
Closest in time.
A. Laptev, R. Korostik, A. Svischev, A. Andrusenko, I. Medennikov, and S. Rybin · 2020
Closest in time.
Improved noisy student training for automatic speech recognition
D. S. Park, Y. Zhang, Y. Jia, W. Han, C.-C. Chiu, B. Li, Y. Wu, and Q. V. Le · 2020
Closest in time.
Multi-task self-supervised learning for robust speech recognition
M. Ravanelli, J. Zhong, S. Pascual, P. Swietojanski, J. Monteiro, J. Trmal, and Y. Bengio · 2020
Closest in time.
Unsupervised pretraining transfers well across languages
M. Rivière, A. Joulin, P.-E. Mazaré, and E. Dupoux · 2020
Closest in time.
End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures
G. Synnaeve, Q. Xu, J. Kahn, T. Likhomanenko, E. Grave, V. Pratap, A. Sriram, V. Liptchinsky, and R. Collobert · 2020
Closest in time.
Unsupervised pre-training of bidirectional speech encoders via masked reconstruction
W. Wang, Q. Tang, and K. Livescu · 2020
Closest in time.
Iterative pseudo-labeling for speech recognition
Q. Xu, T. Likhomanenko, J. Kahn, A. Hannun, G. Synnaeve, and R. Collobert · 2020
Closest in time.
Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss
Q. Zhang, H. Lu, H. Sak, A. Tripathi, E. McDermott, S. Koo, and S. Kumar · 2020
Closest in time.