Fetching the paper…
Reading the bibliography…
Recent success in speech representation learning enables a new way to leverage unlabeled data to train speech recognition model.
“A Fast Learning Algorithm for Deep Nelief Nets,”
Geoffrey E Hinton, Simon Osindero, and Yee-Whye Teh, · 2006
Earlier work this paper cites.
“Greedy Layer-wise Training of Deep Networks,”
Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle, · 2007
Earlier work this paper cites.
“Deep Neural Networks for Acoustic Modeling in Speech Recognition,”
Geoffrey Hinton, Li Deng, Dong Yu, George Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Brian Kingsbury, et al., · 2012
Earlier work this paper cites.
“Semi-supervised Training of Deep Neural Networks,”
Karel Veselỳ, Mirko Hannemann, and Lukáš Burget, · 2013
Earlier work this paper cites.
“Learning Small-size DNN with Output-distribution-based Criteria,”
Jinyu Li, Rui Zhao, Jui-Ting Huang, and Yifan Gong, · 2014
Earlier work this paper cites.
“Distilling the Knowledge in a Neural Network,”
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean, · 2015
Earlier work this paper cites.
“EESEN: End-to-end Speech Recognition using Deep RNN Models and WFST-based Decoding,”
Yajie Miao, Mohammad Gowayyed, and Florian Metze, · 2015
Earlier work this paper cites.
“Semi-supervised training in deep learning acoustic model,”
Yan Huang, Yongqiang Wang, and Yifan Gong, · 2016
Earlier work this paper cites.
“Listen, Attend and Spell: A Neural Network for Large Vocabulary Conversational Speech Recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Neural Discrete Representation Learning,”
Aaron Van Den Oord, Oriol Vinyals, et al., · 2017
Earlier work this paper cites.
“Attention is All You Need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Categorical Reparameterization with Gumbel-Softmax,”
Eric Jang, Shixiang Gu, and Ben Poole, · 2017
Earlier work this paper cites.
“Exploring Architectures, Data and Units for Streaming End-to-end Speech Recognition with RNN-Transducer,”
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar, · 2017
Earlier work this paper cites.
“Hybrid CTC/attention Architecture for End-to-end Speech Recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi, · 2017
Cited alongside, same era.
“Deep Contextualized Word Representations,”
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer, · 2018
Cited alongside, same era.
“Improving Language Understanding by Generative Pre-training,” 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever, · 2018
Cited alongside, same era.
“Representation Learning with Contrastive Predictive Coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Cited alongside, same era.
“End-to-end ASR: from Supervised to Semi-supervised Learning With Modern Architectures,”
Gabriel Synnaeve, Qiantong Xu, Jacob Kahn, Edouard Grave, Tatiana Likhomanenko, Vineel Pratap, Anuroop Sriram, Vitaliy Liptchinsky, and Ronan Collobert, · 2019
“wav2vec 2.0: A Framework for Self-supervised Learning of Speech Representations,”
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Closest in time.
“TERA: Self-supervised Learning of Transformer Encoder Representation for Speech,”
Andy T Liu, Shang-Wen Li, and Hung-yi Lee, · 2020
Closest in time.
“vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations,”
Alexei Baevski, Steffen Schneider, and Michael Auli, · 2020
Closest in time.
“Improved Speech Representations with Multi-target Autoregressive Predictive Coding,”
Yu-An Chung and James Glass, · 2020
Closest in time.
“Speech-XLNet: Unsupervised Acoustic Model Pretraining for Self-attention Networks,”
Xingchen Song, Guangsen Wang, Zhiyong Wu, Yiheng Huang, Dan Su, Dong Yu, and Helen Meng, · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Cited alongside, same era.
“XLNet: Generalized Autoregressive Pretraining for Language Understanding,”
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le, · 2019
Cited alongside, same era.
“wav2vec: Unsupervised Pre-training for Speech Recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Cited alongside, same era.
“An Unsupervised Autoregressive Model for Speech Representation Learning,”
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass, · 2019
Cited alongside, same era.
“Improving Transformer-based Speech Recognition using Unsupervised Pre-training,”
Dongwei Jiang, Xiaoning Lei, Wubo Li, Ne Luo, Yuxuan Hu, Wei Zou, and Xiangang Li, · 2019
Cited alongside, same era.
“Transformers with Convolutional Context for ASR,”
Abdelrahman Mohamed, Dmytro Okhonko, and Luke Zettlemoyer, · 2019
Cited alongside, same era.
“Improving RNN Transducer Modeling for End-to-end Speech Recognition,”
Jinyu Li, Rui Zhao, Hu Hu, and Yifan Gong, · 2019
Cited alongside, same era.
“Unsupervised Pre-training of Bidirectional Speech Encoders via Masked Reconstruction,”
Weiran Wang, Qingming Tang, and Karen Livescu, · 2020
Closest in time.
“Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders,”
Andy T Liu, Shu-wen Yang, Po-Han Chi, Po-chun Hsu, and Hung-yi Lee, · 2020
Closest in time.
“Audio ALBERT: A Lite BERT for Self-supervised Learning of Audio Representation,”
Po-Han Chi, Pei-Hung Chung, Tsung-Han Wu, Chun-Cheng Hsieh, Shang-Wen Li, and Hung-yi Lee, · 2020
Closest in time.
“Towards Unsupervised Speech Recognition and Synthesis with Quantized Speech Representation Learning,”
Alexander H Liu, Tao Tu, Hung-yi Lee, and Lin-shan Lee, · 2020
Closest in time.
“Vector-Quantized Autoregressive Predictive Coding,”
Yu-An Chung, Hao Tang, and James Glass, · 2020
Closest in time.
“Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech,”
David Harwath, Wei-Ning Hsu, and James Glass, · 2020
Closest in time.
“Transformer-based Acoustic Modeling for Hybrid Speech Recognition,”
Yongqiang Wang, Abdelrahman Mohamed, Due Le, Chunxi Liu, Alex Xiao, Jay Mahadeokar, Hongzhao Huang, Andros Tjandra, Xiaohui Zhang, Frank Zhang, et al., · 2020
Closest in time.