Fetching the paper…
Reading the bibliography…
With 4.5 million hours of English speech from 10 different sources across 120 countries and models of up to 10 billion parameters, we explore the frontiers of scale for automatic speech recognition.
“The fisher corpus: A resource for the next generations of speech-to-text.,”
Christopher Cieri, David Miller, and Kevin Walker, · 2004
Earlier work this paper cites.
“Aphasiabank: Methods for studying discourse,”
Brian MacWhinney, Davida Fromm, Margaret Forbes, and Audrey Holland, · 2011
Earlier work this paper cites.
“Large scale deep neural network acoustic modeling with semi-supervised training data for youtube video transcription,”
Hank Liao, Erik McDermott, and Andrew Senior, · 2013
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
Tom Ko, Vijayaditya Peddinti, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Improving automatic recognition of aphasic speech with aphasia bank,”
Duc Le and Emily Mower Provost, · 2015
Earlier work this paper cites.
“Training deep nets with sublinear memory cost,”
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin, · 2016
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, et al., · 2017
Earlier work this paper cites.
“Accurate, large minibatch sgd: Training imagenet in 1 hour,”
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, et al., · 2017
Earlier work this paper cites.
“Automatic quantitative analysis of spontaneous aphasic speech,”
Duc Le, Keli Licata, and Emily Mower Provost, · 2018
Earlier work this paper cites.
“Roberta: A robustly optimized bert pretraining approach,”
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, et al., · 2019
Earlier work this paper cites.
“Megatron-lm: Training multi-billion parameter language models using model parallelism,”
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro, · 2019
Earlier work this paper cites.
“Transformer-transducer: End-to-end speech recognition with self-attention,”
Ching-Feng Yeh, Jay Mahadeokar, Kaustubh Kalgaonkar, Yongqiang Wang, Duc Le, Mahaveer Jain, et al., · 2019
Cited alongside, same era.
“From senones to chenones: Tied context-dependent graphemes for hybrid speech recognition,”
Duc Le, Xiaohui Zhang, Weiyi Zheng, Christian Fügen, Geoffrey Zweig, and Michael L. Seltzer, · 2019
Cited alongside, same era.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Cited alongside, same era.
“fairseq: A fast, extensible toolkit for sequence modeling,”
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, et al., · 2019
Cited alongside, same era.
“Language models are few-shot learners,”
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, et al., · 2020
Cited alongside, same era.
“Speechstew: Simply mix all available speech recognition data to train one large neural network,”
William Chan, Daniel Park, Chris Lee, Yu Zhang, Quoc Le, and Mohammad Norouzi, · 2021
Closest in time.
“Scaling end-to-end models for large-scale multilingual asr,”
Bo Li, Ruoming Pang, Tara N Sainath, Anmol Gulati, Yu Zhang, James Qin, Parisa Haghani, W Ronny Huang, and Min Ma, · 2021
Closest in time.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Closest in time.
“Fairscale: A general purpose modular pytorch library for high performance and large scale training,”
Mandeep Baines, Shruti Bhosale, Vittorio Caggiano, Naman Goyal, Siddharth Goyal, Myle Ott, et al., · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Pushing the limits of semi-supervised learning for automatic speech recognition,”
Yu Zhang, James Qin, Daniel S Park, Wei Han, Chung-Cheng Chiu, Ruoming Pang, et al., · 2020
Cited alongside, same era.
“Scaling laws for neural language models,”
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, et al., · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Transformer-based acoustic modeling for hybrid speech recognition,”
Yongqiang Wang, Abdelrahman Mohamed, Due Le, Chunxi Liu, Alex Xiao, Jay Mahadeokar, Hongzhao Huang, Andros Tjandra, Xiaohui Zhang, Frank Zhang, et al., · 2020
Cited alongside, same era.
“Common voice: A massively-multilingual speech corpus,”
Rosana Ardila, Megan Branson, Kelly Davis, Michael Kohler, Josh Meyer, Michael Henretty, Reuben Morais, Lindsay Saunders, Francis Tyers, and Gregor Weber, · 2020
Cited alongside, same era.
“Libri-light: A benchmark for asr with limited or no supervision,”
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, et al., · 2020
Cited alongside, same era.
“On layer normalization in the transformer architecture,”
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, et al., · 2020
Cited alongside, same era.
Jay Mahadeokar, Yuan Shangguan, Duc Le, Gil Keren, Hang Su, Thong Le, et al., · 2021
Closest in time.
“Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training,”
Wei-Ning Hsu, Anuroop Sriram, Alexei Baevski, Tatiana Likhomanenko, Qiantong Xu, Vineel Pratap, Jacob Kahn, Ann Lee, Ronan Collobert, Gabriel Synnaeve, and Michael Auli, · 2021
Closest in time.
“Rethinking Evaluation in ASR: Are Our Models Robust Enough?,”
Tatiana Likhomanenko, Qiantong Xu, Vineel Pratap, Paden Tomasello, Jacob Kahn, Gilad Avidov, Ronan Collobert, and Gabriel Synnaeve, · 2021
Closest in time.
“Scaling laws for acoustic models,”
Jasha Droppo and Oguz Elibol, · 2021
Closest in time.
Patrick K O’Neill, Vitaly Lavrukhin, Somshubra Majumdar, Vahid Noroozi, Yuekai Zhang, Oleksii Kuchaiev, et al., · 2021
Closest in time.
“Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,”
William Fedus, Barret Zoph, and Noam Shazeer, · 2021
Closest in time.
“Rezero is all you need: Fast convergence at large depth,”
Thomas Bachlechner, Bodhisattwa Prasad Majumder, Huanru Henry Mao, Garrison W Cottrell, and Julian McAuley, · 2021
Closest in time.
“Improving aphasic speech recognition by using novel semi-supervised learning methods on aphasiabank for english and spanish,”
Iván G Torre, Mónica Romero, and Aitor Álvarez, · 2021
Closest in time.