Fetching the paper…
Reading the bibliography…
Recent advances in speech recognition and translation rely on hundreds of thousands of hours of Internet speech data.
M. Bisani and H. Ney, “Bootstrap estimates for confidence intervals in asr performance evaluation,” ICASSP , 2004
2004
Earlier work this paper cites.
D. Bahdanau, K. H. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in ICLR , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NeuRIPS , 2017
2017
Earlier work this paper cites.
T. Kudo and J. Richardson, “Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,” in EMNLP: System Demonstrations , 2018
2018
Earlier work this paper cites.
F. Hernandez, V. Nguyen, S. Ghannay, N. Tomashenko, and Y. Esteve, “Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,” in Speech and Computer: 20th International Conference, SPECOM , 2018
2018
Earlier work this paper cites.
C. D. Kim, B. Kim, H. Lee, and G. Kim, “AudioCaps: Generating captions for audios in the wild,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Minneapolis, Minnesota: Association for Computational Linguistics, 2019
2019
Earlier work this paper cites.
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu et al. , “Conformer: Convolution-augmented transformer for speech recognition,” in Interspeech , 2020
2020
Earlier work this paper cites.
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” in Conference on Language Resources and Evaluation (LREC ) , 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
P. Żelasko, D. Povey, J. Trmal, and S. Khudanpur, “Lhotse: a speech data representation library for the modern deep learning ecosystem,” in NeurIPS Data-Centric AI Workshop , 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux, “VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,” in Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , 2021, pp. 993–1003
D. Rekesh, N. R. Koluguri, S. Kriman, S. Majumdar, V. Noroozi, H. Huang, O. Hrinchuk, K. Puvvada, A. Kumar, J. Balam et al. , “Fast conformer with linearly scalable attention for efficient speech recognition,” in Automatic Speech Recognition and Understanding Workshop (ASRU) . IEEE, 2023, pp. 1–8
2023
Later among the works it cites.
K. Dhawan, K. Rekesh, and B. Ginsburg, “Unified model for code-switching speech recognition and language identification based on concatenated tokenizer,” in Proceedings of the 6th Workshop on Computational Approaches to Linguistic Code-Switching , 2023
2023
Later among the works it cites.
O. Hrinchuk, V. Bataev, E. Bakhturina, and B. Ginsburg, “NVIDIA NeMo offline speech translation systems for IWSLT 2023,” in IWSLT , E. Salesky, M. Federico, and M. Carpuat, Eds., 2023
2023
Later among the works it cites.
NVIDIA, “Stt european fastconformer hybrid transducer-ctc large pnc,” https://catalog.ngc.nvidia.com/orgs/nvidia/teams/nemo/models/stt_multilingual_fastconformer_hybrid_large_pc_blend_eu , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Peng, J. Tian, B. Yan, D. Berrebbi, X. Chang, X. Li, J. Shi, S. Arora, W. Chen, R. Sharma et al. , “Reproducing whisper-style training using an open-source toolkit and publicly available data,” in Automatic Speech Recognition and Understanding Workshop (ASRU) , 2023
2023
Cited alongside, same era.
“Lhotse shar: Storage format optimized for sequential i/o and modularity,” https://colab.research.google.com/github/lhotse-speech/lhotse/blob/master/examples/04-lhotse-shar.ipynb
Cited in the paper.
NVIDIA, “Megatron multilingual model,” https://catalog.ngc.nvidia.com/orgs/nvidia/teams/nemo/models/megatronnmt_en_any_500m
Cited in the paper.
——, “Megatron multilingual model,” https://catalog.ngc.nvidia.com/orgs/nvidia/teams/nemo/models/megatronnmt_any_en_500m
Cited in the paper.
2023
Later among the works it cites.
V. Srivastav, S. Majumdar, N. Koluguri, A. Moumen, S. Gandhi et al. , “Open automatic speech recognition leaderboard,” https://huggingface.co/spaces/hf-audio/open_asr_leaderboard , 2023
2023
Later among the works it cites.
A. Conneau, M. Ma, S. Khanuja, Y. Zhang, V. Axelrod, S. Dalmia, J. Riesa, C. Rivera, and A. Bapna, “Fleurs: Few-shot learning evaluation of universal representations of speech,” in Spoken Language Technology Workshop (SLT) , 2023, pp. 798–805
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Closest in time.
“webdataset,” https://webdataset.github.io/webdataset/ , accessed: 2024-03-08
2024
Closest in time.