Fetching the paper…
Reading the bibliography…
This technical report describes the MERaLiON-SpeechEncoder, a foundation model designed to support a wide range of downstream speech applications.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber · 2006
Earlier work this paper cites.
Librispeech: An ASR corpus based on public domain audio books
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
TED-LIUM 3: Twice as much data and corpus repartition for experiments on speaker adaptation
F. Hernandez, V. Nguyen, S. Ghannay, N. A. Tomashenko, and Y. Estève · 2018
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
T. Kudo and J. Richardson · 2018
Earlier work this paper cites.
ESPnet: End-to-end speech processing toolkit
S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, and T. Ochiai · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Building the Singapore English national speech corpus
J. X. Koh, A. Mislan, K. Khoo, B. Ang, W. Ang, C. Ng, and Y.-Y. Tan · 2019
Earlier work this paper cites.
Fairseq: A fast, extensible toolkit for sequence modeling
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli · 2019
Earlier work this paper cites.
SpecAugment: A simple data augmentation method for automatic speech recognition
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le · 2019
Earlier work this paper cites.
wav2vec: Unsupervised pre-training for speech recognition
S. Schneider, A. Baevski, R. Collobert, and M. Auli · 2019
Earlier work this paper cites.
Common Voice: A massively-multilingual speech corpus
R. Ardila, M. Branson, K. Davis, M. Kohler, J. Meyer, M. Henretty, R. Morais, L. Saunders, F. Tyers, and G. Weber · 2020
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
A. Baevski, H. Zhou, A. Mohamed, and M. Auli · 2020
Earlier work this paper cites.
Conformer: Convolution-augmented transformer for speech recognition
A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang · 2020
Cited alongside, same era.
Libri-light: A benchmark for ASR with limited or no supervision
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux · 2020
Cited alongside, same era.
MLS: A large-scale multilingual dataset for speech research
V. Pratap, Q. Xu, A. Sriram, G. Synnaeve, and R. Collobert · 2020
Cited alongside, same era.
The People’s Speech: A large-scale diverse English speech recognition dataset for commercial usage
D. Galvez, G. Diamos, J. Torres, K. Achorn, J. Cerón, A. Gopi, D. Kanter, M. Lam, M. Mazumder, and V. Janapa Reddi · 2021
Cited alongside, same era.
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units
W.-N. Hsu, B. Bolte, Y.-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed · 2021
Cited alongside, same era.
Improving self-supervised learning for speech recognition with intermediate layer supervision
C. Wang, Y. Wu, S. Chen, S. Liu, J. Li, Y. Qian, and Z. Yang · 2022
Later among the works it cites.
Torchaudio: Building blocks for audio and speech processing
Y.-Y. Yang, M. Hira, Z. Ni, A. Chourdia, A. Astafurov, C. Chen, C.-F. Yeh, C. Puhrsch, D. Pollack, D. Genzel, D. Greenberg, E. Z. Yang, J. Lian, J. Mahadeokar, J. Hwang, J. Chen, P. Goldsborough, P. Roy, S. Narenthiran, S. Watanabe, S. Chintala, V. Quenneville-Bélair, and Y. Shi · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever · 2023
Later among the works it cites.
ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
J. Shi, D. Berrebbi, W. Chen, E.-P. Hu, W.-P. Huang, H.-L. Chung, X. Chang, S.-W. Li, A. Mohamed, H. yi Lee, and S. Watanabe · 2023
Later among the works it cites.
Google USM: Scaling automatic speech recognition beyond 100 languages
Y. Zhang, W. Han, J. Qin, Y. Wang, A. Bapna, Z. Chen, N. Chen, B. Li, V. Axelrod, G. Wang, et al · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation
C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, and E. Dupoux · 2021
Cited alongside, same era.
SUPERB: Speech Processing Universal PERformance Benchmark
S. wen Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K. tik Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H. yi Lee · 2021
Cited alongside, same era.
data2vec: A general framework for self-supervised learning in speech, vision and language
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli · 2022
Cited alongside, same era.
Self-supervised learning with random-projection quantizer for speech recognition
C.-C. Chiu, J. Qin, Y. Zhang, J. Yu, and Y. Wu · 2022
Cited alongside, same era.
Self-supervised speech representation learning: A review
A. Mohamed, H.-y. Lee, L. Borgholt, J. D. Havtorn, J. Edin, C. Igel, K. Kirchhoff, S.-W. Li, K. Livescu, L. Maaløe, T. N. Sainath, and S. Watanabe · 2022
Cited alongside, same era.
SUPERB-SG: Enhanced Speech processing Universal PERformance Benchmark for Semantic and Generative Capabilities
H.-S. Tsai, H.-J. Chang, W.-C. Huang, Z. Huang, K. Lakhotia, S.-w. Yang, S. Dong, A. Liu, C.-I. Lai, J. Shi, X. Chang, P. Hall, H.-J. Chen, S.-W. Li, S. Watanabe, A. Mohamed, and H.-y. Lee · 2022
Cited alongside, same era.
GigaSpeech: An evolving, multi-domain ASR corpus with 10,000 hours of transcribed audio
G. Chen, S. Chai, G.-B. Wang, J. Du, W.-Q. Zhang, C. Weng, D. Su, D. Povey, J. Trmal, J. Zhang, M. Jin, S. Khudanpur, S. Watanabe, S. Zhao, W. Zou, X. Li, X. Yao, Y. Wang, Z. You, and Z. Yan
Cited in the paper.
Later among the works it cites.
Press release on Singapore’s national multimodal large language model programme, 2023
A*STAR · 2024
Closest in time.
MERaLiON-AudioLLM: Technical report
MERaLiON Team · 2024
Closest in time.
Whisper text normalization, 2024
OpenAI · 2024
Closest in time.
TEDLIUM ASR training receipe using Kaldi speech recognition toolkit, 2024
D. Povey · 2024
Closest in time.
Anatomy of industrial scale multilingual ASR
F. M. Ramirez, L. Chkhetiani, A. Ehrenberg, R. McHardy, R. Botros, Y. Khare, A. Vanzo, T. Peyash, G. Oexle, M. Liang, I. Sklyar, E. Fakhan, A. Efty, D. McCrystal, S. Flamini, D. Donato, and T. Yoshioka · 2024
Closest in time.
Seed-ASR: Understanding diverse speech and contexts with LLM-based speech recognition
Seed Team · 2024
Closest in time.