Fetching the paper…
Reading the bibliography…
Wav2vec 2.0 is a recently proposed self-supervised framework for speech representation learning.
“Statistical theory of extreme values and some practical applications,”
Emil Julius Gumbel, · 1954
Earlier work this paper cites.
“Visualizing data using t-sne,”
Laurens van der Maaten and Geoffrey Hinton, · 2008
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Bridging nonlinearities and stochastic regularizers with gaussian error linear units,”
Dan Hendrycks and Kevin Gimpel, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Voxceleb: A large-scale speaker identification dataset,”
Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2017
Earlier work this paper cites.
“Deep contextualized word representations,”
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer, · 2018
Earlier work this paper cites.
“Improving language understanding by generative pre-training,” 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever, · 2018
Earlier work this paper cites.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Earlier work this paper cites.
“Additive margin softmax for face verification,”
Feng Wang, Jian Cheng, Weiyang Liu, and Haijun Liu, · 2018
Cited alongside, same era.
“AP18-OLR challenge: Three tasks and their baselines,”
Zhiyuan Tang, Dong Wang, and Qing Chen, · 2018
Cited alongside, same era.
“Exploring the encoding layer and loss function in end-to-end speaker and language recognition system,”
Weicheng Cai, Jinkun Chen, and Ming Li, · 2018
Cited alongside, same era.
“Attentive statistics pooling for deep speaker embedding,”
Koji Okabe, Takafumi Koshinaka, and Koichi Shinoda, · 2018
Cited alongside, same era.
Zhanyu Ma, Hong Yu, Wei Chen, and Jun Guo, · 2018
Cited alongside, same era.
“BERT: pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Later among the works it cites.
“Learning representations by maximizing mutual information across views,”
Philip Bachman, R Devon Hjelm, and William Buchwalter, · 2019
Later among the works it cites.
“An unsupervised autoregressive model for speech representation learning,”
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James R. Glass, · 2019
Later among the works it cites.
“Improving transformer-based speech recognition using unsupervised pre-training,”
Dongwei Jiang, Xiaoning Lei, Wubo Li, Ne Luo, Yuxuan Hu, Wei Zou, and Xiangang Li, · 2019
Later among the works it cites.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jie Li, Xiaorui Wang, Yan Li, and Yuanyuan Zhao, · 2019
Cited alongside, same era.
“Large-scale multilingual speech recognition with a streaming end-to-end model,”
Anjuli Kannan, Arindrima Datta, Tara N. Sainath, Eugene Weinstein, Bhuvana Ramabhadran, Yonghui Wu, Ankur Bapna, Zhifeng Chen, and Seungji Lee, · 2019
Cited alongside, same era.
“Utterance-level aggregation for speaker recognition in the wild,”
Weidi Xie, Arsha Nagrani, Joon Son Chung, and Andrew Zisserman, · 2019
Cited alongside, same era.
“Attention based hybrid i-vector blstm model for language recognition.,”
Bharat Padi, Anand Mohan, and Sriram Ganapathy, · 2019
Cited alongside, same era.
“Interactive learning of teacher-student model for short utterance spoken language identification,”
Peng Shen, Xugang Lu, Sheng Li, and Hisashi Kawai, · 2019
Cited alongside, same era.
“Voxceleb: Large-scale speaker verification in the wild,”
Arsha Nagrani, Joon Son Chung, Weidi Xie, and Andrew Zisserman, · 2020
Closest in time.
“Data-efficient image recognition with contrastive predictive coding,”
Olivier J. Hénaff, · 2020
Closest in time.
“Multi-task self-supervised learning for robust speech recognition,”
Mirco Ravanelli, Jianyuan Zhong, Santiago Pascual, Pawel Swietojanski, Joao Monteiro, Jan Trmal, and Yoshua Bengio, · 2020
Closest in time.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Closest in time.