Fetching the paper…
Reading the bibliography…
We study the capabilities of speech processing systems trained simply to predict large amounts of transcripts of audio on the internet.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B · 1992
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
Algorithm 799: revolve: an implementation of checkpointing for the reverse or adjoint mode of computational differentiation
Griewank, A. and Walther, A · 2000
Earlier work this paper cites.
Multilingual speech processing
Schultz, T. and Kirchhoff, K · 2006
Earlier work this paper cites.
Massively multilingual asr: 50 languages, 1 model, 1 billion parameters
Pratap, V., Sriram, A., Tomasello, P., Hannun, A. Y., Liptchinsky, V., Synnaeve, G., and Collobert, R · 2007
Earlier work this paper cites.
Covost 2 and massively multilingual speech-to-text translation
Wang, C., Wu, A., and Pino, J · 2007
Earlier work this paper cites.
Deep belief networks for phone recognition
Mohamed, A.-r., Dahl, G., Hinton, G., et al · 2009
Earlier work this paper cites.
fairseq s2t: Fast speech-to-text modeling with fairseq
Wang, C., Tang, Y., Ma, X., Wu, A., Okhonko, D., and Pino, J · 2010
Earlier work this paper cites.
Multitask training with text data for end-to-end speech recognition
Wang, P., Sainath, T. N., and Weiss, R. J · 2010
Earlier work this paper cites.
Natural language processing (almost) from scratch
Collobert, R., Weston, J., Bottou, L., Karlen, M., Kavukcuoglu, K., and Kuksa, P · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E · 2011
Earlier work this paper cites.
Feature engineering in context-dependent deep neural networks for conversational speech transcription
Seide, F., Li, G., Chen, X., and Yu, D · 2011
Earlier work this paper cites.
Unbiased look at dataset bias
Torralba, A. and Efros, A. A · 2011
Earlier work this paper cites.
Mls: A large-scale multilingual dataset for speech research
Pratap, V., Xu, Q., Sriram, A., Synnaeve, G., and Collobert, R · 2012
Earlier work this paper cites.
Large scale deep neural network acoustic modeling with semi-supervised training data for youtube video transcription
Liao, H., McDermott, E., and Senior, A · 2013
Earlier work this paper cites.
The audio degradation toolbox and its application to robustness evaluation
Mauch, M. and Ewert, S · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y · 2013
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V · 2014
Earlier work this paper cites.
Deep speech 2: end-to-end speech recognition in english and mandarin. arxiv
Amodei, D., Anubhai, R., Battenberg, E., Case, C., Casper, J., Catanzaro, B., Chen, J., Chrzanowski, M., Coates, A., Diamos, G., et al · 2015
Earlier work this paper cites.
Multi-task sequence to sequence learning
Luong, M.-T., Le, Q. V., Sutskever, I., Vinyals, O., and Kaiser, L · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Panayotov, V., Chen, G., Povey, D., and Khudanpur, S · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Chen, T., Xu, B., Zhang, C., and Guestrin, C · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q · 2016
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Jia, R. and Liang, P · 2017
Earlier work this paper cites.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Johnson, M., Schuster, M., Le, Q. V., Krikun, M., Wu, Y., Chen, Z., Thorat, N., Viégas, F., Wattenberg, M., Corrado, G., et al · 2017
Cited alongside, same era.
Building machines that learn and think like people
Lake, B. M., Ullman, T. D., Tenenbaum, J. B., and Gershman, S. J · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Ted-lium 3: twice as much data and corpus repartition for experiments on speaker adaptation
Hernandez, F., Nguyen, V., Ghannay, S., Tomashenko, N. A., and Estève, Y · 2018
Cited alongside, same era.
The effect of natural distribution shift on question answering models
Miller, J., Krauth, K., Recht, B., and Schmidt, L · 2020
Later among the works it cites.
pandas-dev/pandas: Pandas, February 2020
pandas development team, T · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P. J., et al · 2020
Later among the works it cites.
Measuring robustness to natural distribution shifts in image classification
Taori, R., Dave, A., Shankar, V., Carlini, N., Recht, B., and Schmidt, L · 2020
Later among the works it cites.
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python
Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., Carey, C. J., Polat, İ., Feng, Y., Moore, E. W., VanderPlas, J., Laxalde, D., Perktold, J., Cimrman, R., Henriksen, I., Quintero, E. A., Harris, C. R., Archibald, A. M., Ribeiro, A. H., Pedregosa, F., van Mulbregt, P., and SciPy 1.0 Contributors · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring the limits of weakly supervised pretraining
Mahajan, D., Girshick, R., Ramanathan, V., He, K., Paluri, M., Li, Y., Bharambe, A., and Van Der Maaten, L · 2018
Cited alongside, same era.
The natural language decathlon: Multitask learning as question answering
McCann, B., Keskar, N. S., Xiong, C., and Socher, R · 2018
Cited alongside, same era.
Toward domain-invariant speech recognition via large scale training
Narayanan, A., Misra, A., Sim, K. C., Pundak, G., Tripathi, A., Elfeky, M., Haghani, P., Strohman, T., and Bacchiani, M · 2018
Cited alongside, same era.
Multilingual speech recognition with a single end-to-end model
Toshniwal, S., Sainath, T. N., Weiss, R. J., Li, B., Moreno, P. J., Weinstein, E., and Rao, K · 2018
Cited alongside, same era.
Strike (with) a pose: Neural networks are easily fooled by strange poses of familiar objects
Alcorn, M. A., Li, Q., Gong, Z., Wang, C., Mai, L., Ku, W.-S., and Nguyen, A · 2019
Cited alongside, same era.
Common voice: A massively-multilingual speech corpus
Ardila, R., Branson, M., Davis, K., Henretty, M., Kohler, M., Meyer, J., Morais, R., Saunders, L., Tyers, F. M., and Weber, G · 2019
Cited alongside, same era.
Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models
Barbu, A., Mayo, D., Alverio, J., Luo, W., Wang, C., Gutfreund, D., Tenenbaum, J., and Katz, B · 2019
Cited alongside, same era.
Later among the works it cites.
Chime-6 challenge: Tackling multispeaker speech recognition for unsegmented recordings
Watanabe, S., Mandel, M., Barker, J., Vincent, E., Arora, A., Chang, X., Khudanpur, S., Manohar, V., Povey, D., Raj, D., et al · 2020
Later among the works it cites.
Pushing the limits of semi-supervised learning for automatic speech recognition
Zhang, Y., Qin, J., Park, D. S., Han, W., Chiu, C.-C., Pang, R., Le, Q. V., and Wu, Y · 2020
Later among the works it cites.
XLS-R: Self-supervised cross-lingual speech representation learning at scale
Babu, A., Wang, C., Tjandra, A., Lakhotia, K., Xu, Q., Goyal, N., Singh, K., von Platen, P., Saraf, Y., Pino, J., et al · 2021
Later among the works it cites.
Unsupervised speech recognition
Baevski, A., Hsu, W.-N., Conneau, A., and Auli, M · 2021
Later among the works it cites.
SpeechStew: Simply mix all available speech recognition data to train one large neural network
Chan, W., Park, D., Lee, C., Zhang, Y., Le, Q., and Norouzi, M · 2021
Later among the works it cites.
Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio
Chen, G., Chai, S., Wang, G., Du, J., Zhang, W.-Q., Weng, C., Su, D., Povey, D., Trmal, J., Zhang, J., et al · 2021
Later among the works it cites.
Earnings-21: a practical benchmark for asr in the wild
Del Rio, M., Delworth, N., Westerman, R., Huang, M., Bhandari, N., Palakapilly, J., McNamara, Q., Dong, J., Zelasko, P., and Jetté, M · 2021
Later among the works it cites.
The people’s speech: A large-scale diverse english speech recognition dataset for commercial usage
Galvez, D., Diamos, G., Torres, J. M. C., Achorn, K., Gopi, A., Kanter, D., Lam, M., Mazumder, M., and Reddi, V. J · 2021
Later among the works it cites.
Scaling laws for neural machine translation
Ghorbani, B., Firat, O., Freitag, M., Bapna, A., Krikun, M., Garcia, X., Chelba, C., and Cherry, C · 2021
Later among the works it cites.
Contextualizing/s/retraction: Sibilant variation and change in washington dc african american language
Gunter, K., Vaughn, C., and Kendall, T · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Later among the works it cites.
SpeechBrain: A general-purpose speech toolkit, 2021
Ravanelli, M., Parcollet, T., Plantinga, P., Rouhe, A., Cornell, S., Lugosch, L., Subakan, C., Dawalatabad, N., Heba, A., Zhong, J., Chou, J.-C., Yeh, S.-L., Fu, S.-W., Liao, C.-F., Rastorgueva, E., Grondin, F., Aris, W., Na, H., Gao, Y., Mori, R. D., and Bengio, Y · 2021
Later among the works it cites.
Voxlingua107: a dataset for spoken language recognition
Valk, J. and Alumäe, T · 2021
Later among the works it cites.
Wang, C., Riviere, M., Lee, A., Wu, A., Talnikar, C., Haziza, D., Williamson, M., Pino, J., and Dupoux, E · 2021
Later among the works it cites.
Self-training and pre-training are complementary for speech recognition
Xu, Q., Baevski, A., Likhomanenko, T., Tomasello, P., Conneau, A., Collobert, R., Synnaeve, G., and Auli, M · 2021
Later among the works it cites.
Zhang, Y., Park, D. S., Han, W., Qin, J., Gulati, A., Shor, J., Jansen, A., Xu, Y., Huang, Y., Wang, S., et al · 2021
Later among the works it cites.
mslam: Massively multilingual joint pre-training for speech and text
Bapna, A., Cherry, C., Zhang, Y., Jia, Y., Johnson, M., Cheng, Y., Khanuja, S., Riesa, J., and Conneau, A · 2022
Closest in time.
Unispeech-sat: Universal speech representation learning with speaker aware pre-training
Chen, S., Wu, Y., Wang, C., Chen, Z., Chen, Z., Liu, S., Wu, J., Qian, Y., Wei, F., Li, J., et al · 2022
Closest in time.
Fleurs: Few-shot learning evaluation of universal representations of speech
Conneau, A., Ma, M., Khanuja, S., Zhang, Y., Axelrod, V., Dalmia, S., Riesa, J., Rivera, C., and Bapna, A · 2022
Closest in time.
The corpus of regional african american language
Kendall, T. and Farrington, C · 2022
Closest in time.
Using the output embedding to improve language models
Press, O. and Wolf, L · 2025
Closest in time.