Fetching the paper…
Reading the bibliography…
We propose vq-wav2vec to learn discrete representations of audio segments through a wav2vec-style self-supervised context prediction task.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli · 1904
Earlier work this paper cites.
Statistical theory of extreme values and some practical applications: a series of lectures , volume 33
Emil Julius Gumbel · 1954
Earlier work this paper cites.
Vorbis i specification, 2004
C Montgomery · 2004
Earlier work this paper cites.
Product quantization for nearest neighbor search
Herve Jegou, Matthijs Douze, and Cordelia Schmid · 2011
Earlier work this paper cites.
Definition of the opus audio codec, 2012
Tim Terriberry and Koen Vos · 2012
Earlier work this paper cites.
Scalable modified Kneser-Ney language model estimation
Kenneth Heafield, Ivan Pouzyrevsky, Jonathan H. Clark, and Philipp Koehn · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
A* sampling
Chris J Maddison, Daniel Tarlow, and Tom Minka · 2014
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Earlier work this paper cites.
Wav2letter: an end-to-end convnet-based speech recognition system
Ronan Collobert, Christian Puhrsch, and Gabriel Synnaeve · 2016
Earlier work this paper cites.
ffmpeg tool software, 2016
FFmpeg Developers · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
SGDR: stochastic gradient descent with restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
The zero resource speech challenge 2015: Proposed approaches and results
Maarten Versteegh, Xavier Anguera, Aren Jansen, and Emmanuel Dupoux · 2016
Cited alongside, same era.
Investigation of transfer learning for asr using lf-mmi trained neural networks
Pegah Ghahremani, Vimal Manohar, Hossein Hadian, Daniel Povey, and Sanjeev Khudanpur · 2017
Cited alongside, same era.
Neural discrete representation learning
Aaron van den Oord, Oriol Vinyals, et al · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Speech2vec: A sequence-to-sequence framework for learning word embeddings from speech
Unsupervised speech representation learning using wavenet autoencoders
Jan Chorowski, Ron J. Weiss, Samy Bengio, and Aäron van den Oord · 2019
Closest in time.
An unsupervised autoregressive model for speech representation learning
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass · 2019
Closest in time.
A fully differentiable beam search decoder
Ronan Collobert, Awni Hannun, and Gabriel Synnaeve · 2019
Closest in time.
The zero resource speech challenge 2019: Tts without t
Ewan Dunbar, Robin Algayres, Julien Karadayi, Mathieu Bernard, Juan Benjumea, Xuan-Nga Cao, Lucie Miskic, Charlotte Dugrain, Lucas Ondel, Alan W Black, et al · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yu-An Chung and James Glass · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
End-to-end speech recognition using lattice-free mmi
Hossein Hadian, Hossein Sameti1, Daniel Povey, and Sanjeev Khudanpur · 2018
Cited alongside, same era.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli · 2018
Cited alongside, same era.
Light gated recurrent units for speech recognition
Mirco Ravanelli, Philemon Brakel, Maurizio Omologo, and Yoshua Bengio · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Yuxin Wu and Kaiming He · 2018
Cited alongside, same era.
Ryan Eloff, André Nortje, Benjamin van Niekerk, Avashna Govender, Leanne Nortje, Arnu Pretorius, Elan Van Biljon, Ewald van der Westhuizen, Lisa van Staden, and Herman Kamper · 2019
Closest in time.
On the choice of modeling unit for sequence-to-sequence speech recognition
Kazuki Irie, Rohit Prabhavalkar, Anjuli Kannan, Antoine Bruguier, David Rybach, and Patrick Nguyen · 2019
Closest in time.
Spanbert: Improving pre-training by representing and predicting spans
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S. Weld, Luke Zettlemoyer, and Omer Levy · 2019
Closest in time.
Who needs words? lexicon-free speech recognition
Tatiana Likhomanenko, Gabriel Synnaeve, and Ronan Collobert · 2019
Closest in time.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Closest in time.
Transformers with convolutional context for ASR
Abdelrahman Mohamed, Dmytro Okhonko, and Luke Zettlemoyer · 2019
Closest in time.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Closest in time.
Specaugment: A simple data augmentation method for automatic speech recognition, 2019
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le · 2019
Closest in time.
Vqvae unsupervised unit discovery and multi-scale code2spec inverter for zerospeech challenge 2019
Andros Tjandra, Berrak Sisman, Mingyang Zhang, Sakriani Sakti, Haizhou Li, and Satoshi Nakamura · 2019
Closest in time.