Fetching the paper…
Reading the bibliography…
The integration of Language Models (LMs) has proven to be an effective way to address domain shifts in speech recognition.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Ted-lium: an automatic speech recognition dedicated corpus.,”
Anthony Rousseau, Paul Deléglise, and Yannick Esteve, · 2012
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“On using monolingual corpora in neural machine translation,”
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loic Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“An analysis of incorporating an external language model into a sequence-to-sequence model,”
Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N Sainath, Zhijeng Chen, and Rohit Prabhavalkar, · 2018
Earlier work this paper cites.
“Cold fusion: Training seq2seq models together with language models,”
Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh, and Adam Coates, · 2018
Earlier work this paper cites.
“Language models are unsupervised multitask learners,”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al., · 2019
Earlier work this paper cites.
“Transformer-xl: Attentive language models beyond a fixed-length context,”
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G Carbonell, Quoc Le, and Ruslan Salakhutdinov, · 2019
Earlier work this paper cites.
“A density ratio approach to language model fusion in end-to-end automatic speech recognition,”
Erik McDermott, Hasim Sak, and Ehsan Variani, · 2019
Earlier work this paper cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Earlier work this paper cites.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2019
Earlier work this paper cites.
“Language models are few-shot learners,”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al., · 2020
Cited alongside, same era.
“Hybrid autoregressive transducer (HAT),”
Ehsan Variani, David Rybach, Cyril Allauzen, and Michael Riley, · 2020
Cited alongside, same era.
“Libri-light: A benchmark for asr with limited or no supervision,”
Jacob Kahn, Morgane Riviere, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al., · 2020
Cited alongside, same era.
“Adapting GPT, GPT-2 and BERT language models for speech recognition,”
Xianrui Zheng, Chao Zhang, and Philip C Woodland, · 2021
Cited alongside, same era.
“Internal language model estimation for domain-adaptive end-to-end speech recognition,”
Zhong Meng, Sarangarajan Parthasarathy, Eric Sun, Yashesh Gaur, Naoyuki Kanda, Liang Lu, Xie Chen, Rui Zhao, Jinyu Li, and Yifan Gong, · 2021
Cited alongside, same era.
“ASR adaptation for e-commerce chatbots using cross-utterance context and multi-task language modeling,”
Ashish Shenoy, Sravan Bodapati, and Katrin Kirchhoff, · 2021
Later among the works it cites.
“Robust speech recognition via large-scale weak supervision,”
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever, · 2022
Later among the works it cites.
“Recent advances in end-to-end automatic speech recognition,”
Jinyu Li, · 2022
Later among the works it cites.
“Factorized neural transducer for efficient language model adaptation,”
Xie Chen, Zhong Meng, Sarangarajan Parthasarathy, and Jinyu Li, · 2022
Later among the works it cites.
“Residual language model for end-to-end speech recognition,”
Emiru Tsunoo, Yosuke Kashiwagi, Chaitanya Narisetty, and Shinji Watanabe, · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Internal language model training for domain-adaptive end-to-end speech recognition,”
Zhong Meng, Naoyuki Kanda, Yashesh Gaur, Sarangarajan Parthasarathy, Eric Sun, Liang Lu, Xie Chen, Jinyu Li, and Yifan Gong, · 2021
Cited alongside, same era.
“Internal language model adaptation with text-only data for end-to-end speech recognition,”
Zhong Meng, Yashesh Gaur, Naoyuki Kanda, Jinyu Li, Xie Chen, Yu Wu, and Yifan Gong, · 2021
Cited alongside, same era.
“Domain prompts: Towards memory and compute efficient domain adaptation of asr systems,”
Saket Dingliwal, Ashish Shenoy, Sravan Bodapati, Ankur Gandhe, Ravi Teja Gadde, and Katrin Kirchhoff, · 2021
Cited alongside, same era.
Patrick K O’Neill, Vitaly Lavrukhin, Somshubra Majumdar, Vahid Noroozi, Yuekai Zhang, Oleksii Kuchaiev, Jagadeesh Balam, Yuliya Dovzhenko, Keenan Freyberg, Michael D Shulman, et al., · 2021
Cited alongside, same era.
“Hubert: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“Prefix-tuning: Optimizing continuous prompts for generation,”
Xiang Lisa Li and Percy Liang, · 2021
Cited alongside, same era.
“Adapting long context nlm for asr rescoring in conversational agents,”
Ashish Shenoy, Sravan Bodapati, Monica Sunkara, Srikanth Ronanki, and Katrin Kirchhoff, · 2021
Cited alongside, same era.
“Flamingo: a visual language model for few-shot learning,”
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al., · 2022
Later among the works it cites.
“Fast and accurate factorized neural transducer for text adaption of end-to-end speech recognition models,”
Rui Zhao, Jian Xue, Partha Parthasarathy, Veljko Miljanic, and Jinyu Li, · 2023
Closest in time.
“Adaptable end-to-end asr models using replaceable internal lms and residual softmax,”
Keqi Deng and Philip C Woodland, · 2023
Closest in time.
OpenAI, · 2023
Closest in time.
“LLaMA: Open and efficient foundation language models,”
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al., · 2023
Closest in time.
“Minigpt-4: Enhancing vision-language understanding with advanced large language models,”
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny, · 2023
Closest in time.
“LLaMA-adapter: Efficient fine-tuning of language models with zero-init attention,”
Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao, · 2023
Closest in time.