Fetching the paper…
Reading the bibliography…
We explore unsupervised pre-training for speech recognition by learning representations of raw audio.
Large vocabulary continuous speech recognition using htk
Philip C Woodland, Julian J Odell, Valtcho Valtchev, and Steve J Young · 1994
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Natural language processing (almost) from scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa · 2011
Earlier work this paper cites.
Scalable modified Kneser-Ney language model estimation
Kenneth Heafield, Ivan Pouzyrevsky, Jonathan H. Clark, and Philipp Koehn · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei A. Efros · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Earlier work this paper cites.
Wav2letter: an end-to-end convnet-based speech recognition system
Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve · 2016
Earlier work this paper cites.
SGDR: stochastic gradient descent with restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Show and tell: Lessons learned from the 2015 MS COCO image captioning challenge
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2016
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Cited alongside, same era.
Investigation of transfer learning for asr using lf-mmi trained neural networks
Pegah Ghahremani, Vimal Manohar, Hossein Hadian, Daniel Povey, and Sanjeev Khudanpur · 2017
Cited alongside, same era.
A segmental framework for fully-unsupervised large-vocabulary speech recognition
Herman Kamper, Aren Jansen, and Sharon Goldwater · 2017
Cited alongside, same era.
Transfer learning for speech recognition on a budget
Julius Kunze, Lous Kirsch, Ilia Kurenkov, Andreas Krug, Jens Johannsmeier, and Sebastian Stober · 2017
Cited alongside, same era.
Yi-Chen Chen, Chia-Hao Shen, Sung-Feng Huang, Hung-yi Lee, and Lin-Shan Lee · 2018
Cited alongside, same era.
Light gated recurrent units for speech recognition
Mirco Ravanelli, Philemon Brakel, Maurizio Omologo, and Yoshua Bengio · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Yuxin Wu and Kaiming He · 2018
Later among the works it cites.
Cloze-driven pretraining of self-attention networks
Alexei Baevski, Sergey Edunov, Yinhan Liu, Luke Zettlemoyer, and Michael Auli · 2019
Closest in time.
Unsupervised speech representation learning using wavenet autoencoders
Jan Chorowski, Ron J. Weiss, Samy Bengio, and Aäron van den Oord · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu-An Chung, Wei-Hung Weng, Schrasing Tong, and James R. Glass · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
End-to-end speech recognition using lattice-free mmi
Hossein Hadian, Hossein Sameti1, Daniel Povey, and Sanjeev Khudanpur · 2018
Cited alongside, same era.
Zheng Lian, Ya Li, Jianhua Tao, and Jian Huang · 2018
Cited alongside, same era.
wav2letter++: The fastest open-source speech recognition system
Vineel Pratap, Awni Hannun, Qiantong Xu, Jeff Cai, Jacob Kahn, Gabriel Synnaeve, Vitaliy Liptchinsky, and Ronan Collobert · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
Learning speaker representations with mutual information
Mirco Ravanelli and Yoshua Bengio · 2018
Cited alongside, same era.
Closest in time.
A fully differentiable beam search decoder
Ronan Collobert, Awni Hannun, and Gabriel Synnaeve · 2019
Closest in time.
Pre-trained language model representations for language generation
Sergey Edunov, Alexei Baevski, and Michael Auli · 2019
Closest in time.
Data-efficient image recognition with contrastive predictive coding
Olivier J. Hénaff, Ali Razavi, Carl Doersch, S. M. Ali Eslami, and Aäron van den Oord · 2019
Closest in time.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau · 2019
Closest in time.
Who needs words? lexicon-free speech recognition
Tatiana Likhomanenko, Gabriel Synnaeve, and Ronan Collobert · 2019
Closest in time.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Closest in time.
3d human pose estimation in video with temporal convolutions and semi-supervised training
Dario Pavllo, Christoph Feichtenhofer, David Grangier, and Michael Auli · 2019
Closest in time.