Fetching the paper…
Reading the bibliography…
Recent advances in large language models (LLMs) have promoted generative error correction (GER) for automatic speech recognition (ASR), which leverages the rich linguistic knowledge and powerful reasoning ability of LLMs to improve recognition results.
The aurora experimental framework for the performance evaluation of speech recognition systems under noisy conditions
Hans-Günter Hirsch and David Pearce · 2000
Earlier work this paper cites.
Subjective comparison of speech enhancement algorithms
Yi Hu and Philipos C Loizou · 2006
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Freesound technical demo
Frederic Font, Gerard Roma, and Xavier Serra · 2013
Earlier work this paper cites.
The voice bank corpus: Design, collection and data analysis of a large regional accent speech database
Christophe Veaux, Junichi Yamagishi, and Simon King · 2013
Earlier work this paper cites.
The rats collection: Supporting hlt research with degraded audio data
David Graff, Kevin Walker, Stephanie M Strassel, Xiaoyi Ma, Karen Jones, and Ann Sawyer · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
An overview of noise-robust automatic speech recognition
Jinyu Li, Li Deng, Yifan Gong, and Reinhold Haeb-Umbach · 2014
Earlier work this paper cites.
Bidirectional recurrent neural network language models for automatic speech recognition
Ebru Arisoy, Abhinav Sethy, Bhuvana Ramabhadran, and Stanley Chen · 2015
Earlier work this paper cites.
Robust automatic speech recognition: a bridge to practical applications , chapter 1, pp. 1–20
Jinyu Li, Li Deng, Reinhold Haeb-Umbach, and Yifan Gong · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Investigating rnn-based speech enhancement methods for noise-robust text-to-speech
Cassia Valentini-Botinhao, Xin Wang, Shinji Takaki, and Junichi Yamagishi · 2016
Earlier work this paper cites.
The 4th chime speech separation and recognition challenge
Emmanuel Vincent, Shinji Watanabe, Jon Barker, and Ricard Marxer · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Mutual information neural estimation
Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeshwar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, and Devon Hjelm · 2018
Earlier work this paper cites.
Learning word vectors for 157 languages
Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and Tomas Mikolov · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2018
Earlier work this paper cites.
Espnet: End-to-end speech processing toolkit
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, et al · 2018
Earlier work this paper cites.
Metricgan: Generative adversarial networks based black-box metric scores optimization for speech enhancement
Szu-Wei Fu, Chien-Feng Liao, Yu Tsao, and Shou-De Lin · 2019
Cited alongside, same era.
A spelling correction model for end-to-end speech recognition
Jinxi Guo, Tara N Sainath, and Ron J Weiss · 2019
Cited alongside, same era.
Speech recognition with no speech or with noisy speech
Gautam Krishna, Co Tran, Jianguo Yu, and Ahmed H Tewfik · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Cited alongside, same era.
Effective sentence scoring method using bert for speech recognition
Joonbo Shin, Yoonhyung Lee, and Kyomin Jung · 2019
Cited alongside, same era.
Language models are few-shot learners
Spatial-channel token distillation for vision mlps
Yanxi Li, Xinghao Chen, Minjing Dong, Yehui Tang, Yunhe Wang, and Chang Xu · 2022
Later among the works it cites.
Introducing chatgpt
OpenAI · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Later among the works it cites.
Prompting large language models with speech recognition abilities
Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Junteng Jia, Yuan Shangguan, Ke Li, Jinxi Guo, Wenhan Xiong, Jay Mahadeokar, Ozlem Kalinli, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Deliberation model based two-pass end-to-end speech recognition
Ke Hu, Tara N Sainath, Ruoming Pang, and Rohit Prabhavalkar · 2020
Cited alongside, same era.
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Fastcorrect 2: Fast error correction on multiple candidates for automatic speech recognition
Yichong Leng, Xu Tan, Rui Wang, Linchen Zhu, Jin Xu, Wenjie Liu, Linquan Liu, Tao Qin, Xiang-Yang Li, Edward Lin, et al · 2021
Cited alongside, same era.
Unsupervised noise adaptive speech enhancement by discriminator-constrained optimal transport
Hsin-Yi Lin, Huan-Hsin Tseng, Xugang Lu, and Yu Tsao · 2021
Cited alongside, same era.
Dual application of speech enhancement for automatic speech recognition
Ashutosh Pandey, Chunxi Liu, Yun Wang, and Yatharth Saraf · 2021
Cited alongside, same era.
Trapping llm hallucinations using tagged context prompts
Philip Feldman, James R Foulds, and Shimei Pan · 2023
Later among the works it cites.
Llama-adapter v2: Parameter-efficient visual instruction model
Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, et al · 2023
Later among the works it cites.
Scaling up deliberation for multilingual asr
Ke Hu, Bo Li, and Tara N Sainath · 2023
Later among the works it cites.
Rao Ma, Mark JF Gales, Kate Knill, and Mengjie Qian · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Enhancing speaker diarization with large language models: A contextual beam search approach
Tae Jin Park, Kunal Dhawan, Nithin Koluguri, and Jagadeesh Balam · 2023
Later among the works it cites.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Later among the works it cites.
Whispering llama: A cross-modal generative error correction framework for speech recognition
Srijith Radhakrishnan, Chao-Han Yang, Sumeer Khan, Rohit Kumar, Narsis Kiani, David Gomez-Cabrero, and Jesper Tegnér · 2023
Later among the works it cites.
Can whisper perform speech-based in-context learning
Siyin Wang, Chao-Han Huck Yang, Ji Wu, and Chao Zhang · 2023
Later among the works it cites.
Low-rank adaptation of large language model rescoring for parameter-efficient speech recognition
Yu Yu, Chao-Han Huck Yang, Jari Kolehmainen, Prashanth G Shivakumar, Yile Gu, Sungho Ryu, Roger Ren, Qi Luo, Aditya Gourav, I-Fan Chen, et al · 2023
Later among the works it cites.
Diarizationlm: Speaker diarization post-processing with large language models
Quan Wang, Yiling Huang, Guanlong Zhao, Evan Clark, Wei Xia, and Hank Liao · 2024
Closest in time.