Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated unparalleled effectiveness in various NLP tasks, and integrating LLMs with automatic speech recognition (ASR) is becoming a mainstream paradigm.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in ICASSP , 2015
2015
Earlier work this paper cites.
H. Bu, J. Du, X. Na, B. Wu, and H. Zheng, “AISHELL-1: An open-source Mandarin speech corpus and a speech recognition baseline,” in O-COCOSDA , 2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is All you Need,” in Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
J. Du, X. Na, X. Liu, and H. Bu, “AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale,” CoRR , 2018
2018
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in NAACL-HLT , 2019
2019
Earlier work this paper cites.
D. Zhao, T. N. Sainath, D. Rybach, P. Rondon, D. Bhatia, B. Li, and R. Pang, “Shallow-Fusion End-to-End Contextual Biasing,” in Interspeech , 2019
2019
Earlier work this paper cites.
S. Deena, M. Hasan, M. Doulaty, O. Saz, and T. Hain, “Recurrent Neural Network Language Model Adaptation for Multi-Genre Broadcast Speech Recognition and Alignment,” ACM , 2019
2019
Earlier work this paper cites.
K. Irie, A. Zeyer, R. Schlüter, and H. Ney, “Language Modeling with Deep Transformers,” in Interspeech , 2019
2019
Earlier work this paper cites.
C. Shan, C. Weng, G. Wang, D. Su, M. Luo, D. Yu, and L. Xie, “Component Fusion: Learning Replaceable Language Model Component for End-to-end Speech Recognition System,” in ICASSP , 2019
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” in ICLR , 2019
2019
Earlier work this paper cites.
J. Zhang, T. He, S. Sra, and A. Jadbabaie, “Why Gradient Clipping Accelerates Training: A Theoretical Justification for Adaptivity,” in ICLR , 2020
2020
Earlier work this paper cites.
Z. Liu, K. Li, S. Bakshi, and F. Peng, “Private Language Model Adaptation for Speech Recognition,” CoRR , 2021
2021
Earlier work this paper cites.
W. Hsu, B. Bolte, Y. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, “HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,” ACM , 2021
2021
Earlier work this paper cites.
Y. Fu, L. Cheng, S. Lv, Y. Jv, Y. Kong, Z. Chen, Y. Hu, L. Xie, J. Wu, H. Bu, X. Xu, J. Du, and J. Chen, “AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario,” in Interspeech , 2021
2021
Cited alongside, same era.
D. Wu, B. Zhang, C. Yang, Z. Peng, W. Xia, X. Chen, and X. Lei, “U2++: Unified Two-pass Bidirectional End-to-end Model for Speech Recognition,” CoRR , 2021
2021
Cited alongside, same era.
M. Jung, O. Kwon, S. Seo, and S. Seo, “Blank Collapse: Compressing CTC emission for the faster decoding,” CoRR , 2022
2022
Cited alongside, same era.
B. Zhang, H. Lv, P. Guo, Q. Shao, C. Yang, L. Xie, X. Xu, H. Bu, X. Chen, C. Zeng, D. Wu, and Z. Peng, “WENETSPEECH: A 10000+ Hours Multi-Domain Mandarin Corpus for Speech Recognition,” in ICASSP , 2022
2022
Cited alongside, same era.
P. Dighe, Y. Su, S. Zheng, Y. Liu, V. Garg, X. Niu, and A. H. Tewfik, “Leveraging Large Language Models for Exploiting ASR Uncertainty,” CoRR , 2023
2023
Later among the works it cites.
R. Ma, M. Qian, P. Manakul, M. J. F. Gales, and K. Knill, “Can Generative Large Language Models Perform ASR Error Correction?” CoRR , 2023
2023
Later among the works it cites.
J. Wu, Y. Gaur, Z. Chen, L. Zhou, Y. Zhu, T. Wang, J. Li, S. Liu, B. Ren, L. Liu, and Y. Wu, “On Decoder-Only Architecture For Speech-to-Text and Large Language Model Integration,” in ASRU , 2023
2023
Later among the works it cites.
Y. Gong, H. Luo, A. H. Liu, L. Karlinsky, and J. R. Glass, “Listen, Think, and Understand,” CoRR , 2023
2023
Later among the works it cites.
C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, L. Lu, Z. Ma, and C. Zhang, “SALMONN: Towards Generic Hearing Abilities for Large Language Models,” CoRR , 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-Rank Adaptation of Large Language Models,” in ICLR , 2022
2022
Cited alongside, same era.
Zhifu Gao and Shiliang Zhang and Ian McLoughlin and Zhijie Yan, “Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition,” in Interspeech , 2022
2022
Cited alongside, same era.
B. Zhang, D. Wu, Z. Peng, X. Song, Z. Yao, H. Lv, L. Xie, C. Yang, F. Pan, and J. Niu, “WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit,” in Interspeech , 2022
2022
Cited alongside, same era.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample, “LLaMA: Open and Efficient Foundation Language Models,” CoRR , 2023
2023
Cited alongside, same era.
OpenAI, “GPT-4 Technical Report,” CoRR , 2023
2023
Cited alongside, same era.
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth et al. , “Gemini: a family of highly capable multimodal models,” CoRR , 2023
2023
Cited alongside, same era.
D. Zhang, S. Li, X. Zhang, J. Zhan, P. Wang, Y. Zhou, and X. Qiu, “SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities,” in EMNLP , 2023
2023
Cited alongside, same era.
R. Ma, X. Wu, J. Qiu, Y. Qin, H. Xu, P. Wu, and Z. Ma, “Internal Language Model Estimation Based Adaptive Language Model Fusion for Domain Adaptation,” in ICASSP , 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
Y. Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou, “Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models,” CoRR , 2023
2023
Later among the works it cites.
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust Speech Recognition via Large-Scale Weak Supervision,” in ICML , 2023
2023
Later among the works it cites.
S. Chen, Y. Wu, C. Wang, S. Liu, D. Tompkins, Z. Chen, W. Che, X. Yu, and F. Wei, “BEATs: Audio Pre-Training with Acoustic Tokenizers,” in ICML , 2023
2023
Later among the works it cites.
J. Li, D. Li, S. Savarese, and S. C. H. Hoi, “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,” in ICML , 2023
2023
Later among the works it cites.
R. Huang, M. Li, D. Yang, J. Shi, X. Chang, Z. Ye, Y. Wu, Z. Hong, J. Huang, J. Liu, Y. Ren, Y. Zou, Z. Zhao, and S. Watanabe, “AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head,” in AAAI , 2024
2024
Closest in time.
Z. Ma, G. Yang, Y. Yang, Z. Gao, J. Wang, Z. Du, F. Yu, Q. Chen, S. Zheng, S. Zhang, and X. Chen, “An Embarrassingly Simple Approach for LLM with Strong ASR Capacity,” CoRR , 2024
2024
Closest in time.
J. Du, J. Li, G. Chen, and W. Zhang, “SpeechColab Leaderboard: An Open-Source Platform for Automatic Speech Recognition Evaluation,” CoRR , 2024
2024
Closest in time.