Fetching the paper…
Reading the bibliography…
In recent years, Large Language Models (LLMs) have garnered significant attention from the research community due to their exceptional performance and generalization capabilities.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves et al., · 2006
Earlier work this paper cites.
“Audio augmentation for speech recognition,”
Tom Ko, Vijayaditya Peddinti, et al., · 2015
Earlier work this paper cites.
William Chan, Navdeep Jaitly, Quoc V. Le, and Oriol Vinyals, · 2015
Earlier work this paper cites.
“Deep context: End-to-end contextual speech recognition,”
Golan Pundak, Tara N. Sainath, et al., · 2018
Earlier work this paper cites.
“Shallow-fusion end-to-end contextual biasing,”
Ding Zhao, Tara N. Sainath, David Rybach, Pat Rondon, et al., · 2019
Earlier work this paper cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 2019
Earlier work this paper cites.
“fairseq: A fast, extensible toolkit for sequence modeling,”
Myle Ott, Sergey Edunov, Alexei Baevski, et al., · 2019
Earlier work this paper cites.
“Contextual RNN-T for open domain ASR,”
Mahaveer Jain, Gil Keren, Jay Mahadeokar, Geoffrey Zweig, Florian Metze, and Yatharth Saraf, · 2020
Earlier work this paper cites.
“Conformer: Convolution-augmented transformer for speech recognition,”
Anmol Gulati, James Qin, et al., · 2020
Earlier work this paper cites.
“G2g: Tts-driven pronunciation learning for graphemic hybrid asr,”
Duc Le, Thilo Koehler, Christian Fuegen, and Michael L. Seltzer, · 2020
Cited alongside, same era.
“Contextualized streaming end-to-end speech recognition with trie-based deep biasing and shallow fusion,”
Duc Le, Mahaveer Jain, Gil Keren, Suyoun Kim, et al., · 2021
Cited alongside, same era.
“Scaling asr improves zero and few shot learning,”
Alex Xiao, Weiyi Zheng, Gil Keren, et al., · 2021
Cited alongside, same era.
“Deep shallow fusion for rnn-t personalization,”
Duc Le, Gil Keren, Julian Chan, et al., · 2021
Cited alongside, same era.
“Roformer: Enhanced transformer with rotary position embedding,”
Jianlin Su, Yu Lu, et al., · 2021
Cited alongside, same era.
“Contextual adapters for personalized speech recognition in neural transducers,”
“Robust acoustic and semantic contextual biasing in neural transducers for speech recognition,”
Xuandi Fu, Kanthashree Mysore Sathyendra, Ankur Gandhe, et al., · 2023
Closest in time.
“Adaptive contextual biasing for transducer based streaming speech recognition,” 2023
Tianyi Xu, Zhanheng Yang, Kaixun Huang, et al., · 2023
Closest in time.
“On decoder-only architecture for speech-to-text and large language model integration,” 2023
Jian Wu, Yashesh Gaur, et al., · 2023
Closest in time.
“Audiopalm: A large language model that can speak and listen,”
Paul K. Rubenstein et al., · 2023
Closest in time.
“Prompting large language models with speech recognition abilities,” 2023
Yassir Fathullah, Chunyang Wu, Egor Lakomkin, et al., · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kanthashree Mysore Sathyendra, Thejaswi Muniyappa, Feng-Ju Chang, et al., · 2022
Cited alongside, same era.
“Lora: Low-rank adaptation of large language models,”
Edward J. Hu, Yelong Shen, et al., · 2022
Cited alongside, same era.
“Llama: Open and efficient foundation language models,” 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, et al., · 2023
Cited alongside, same era.
“Gpt-4 technical report,” 2023
OpenAI, · 2023
Cited alongside, same era.
Alec Radford, Jong Wook Kim, et al., · 2023
Closest in time.
“The case for 4-bit precision: k-bit inference scaling laws,”
Tim Dettmers and Luke Zettlemoyer, · 2023
Closest in time.
“Longnet: Scaling transformers to 1,000,000,000 tokens,” 2023
Jiayu Ding, Shuming Ma, et al., · 2023
Closest in time.