Fetching the paper…
Reading the bibliography…
Recent studies have successfully shown that large language models (LLMs) can be successfully used for generative error correction (GER) on top of the automatic speech recognition (ASR) output.
A description of a parametrically controlled modular structure for speech processing
N Dixon and H Silverman · 1975
Earlier work this paper cites.
Continuous speech recognition by statistical methods
Frederick Jelinek · 1976
Earlier work this paper cites.
The ATIS spoken language systems pilot corpus
Charles T. Hemphill, John J. Godfrey, and George R. Doddington · 1990
Earlier work this paper cites.
The design for the wall street journal-based csr corpus
Douglas B Paul and Janet Baker · 1992
Earlier work this paper cites.
Efficient general lattice generation and rescoring
Andrej Ljolje, Fernando Pereira, and Michael Riley · 1999
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Discriminative training of language models for speech recognition
Hong-Kwang Jeff Kuo, Eric Fosler-Lussier, Hui Jiang, and Chin-Hui Lee · 2002
Earlier work this paper cites.
Model-based feature enhancement with uncertainty decoding for noise robust asr
Veronique Stouten, Patrick Wambacq, et al · 2006
Earlier work this paper cites.
Csr-i (wsj0) complete
John Garofalo, David Graff, Doug Paul, and David Pallett · 2007
Earlier work this paper cites.
Speech recognition with weighted finite-state transducers
Mehryar Mohri, Fernando Pereira, and Michael Riley · 2008
Earlier work this paper cites.
On-the-fly lattice rescoring for real-time automatic speech recognition
Haşim Sak, Murat Saraclar, and Tunga Güngör · 2010
Earlier work this paper cites.
Bayesian learning for neural networks , volume 118
Radford M Neal · 2012
Earlier work this paper cites.
Fusion of multiple uncertainty estimators and propagators for noise robust asr
Dung T Tran, Emmanuel Vincent, and Denis Jouvet · 2014
Earlier work this paper cites.
On using monolingual corpora in neural machine translation
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loic Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2015
Earlier work this paper cites.
Musan: A music, speech, and noise corpus
David Snyder, Guoguo Chen, and Daniel Povey · 2015
Earlier work this paper cites.
Towards better decoding and language model integration in sequence to sequence models
Jan Chorowski and Navdeep Jaitly · 2016
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
The 4th chime speech separation and recognition challenge
Emmanuel Vincent, Shinji Watanabe, Jon Barker, and Ricard Marxer · 2016
Earlier work this paper cites.
Cold fusion: Training seq2seq models together with language models
Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh, and Adam Coates · 2017
Earlier work this paper cites.
An analysis of incorporating an external language model into a sequence-to-sequence model
Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N Sainath, Zhijeng Chen, and Rohit Prabhavalkar · 2018
Earlier work this paper cites.
End-to-end contextual speech recognition using class language models and a token passing decoder
Zhehuai Chen, Mahaveer Jain, Yongqiang Wang, Michael L Seltzer, and Christian Fuegen · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Cited alongside, same era.
Two-pass end-to-end speech recognition
Tara N Sainath, Ruoming Pang, David Rybach, Yanzhang He, Rohit Prabhavalkar, Wei Li, Mirkó Visontai, Qiao Liang, Trevor Strohman, Yonghui Wu, et al · 2019
Cited alongside, same era.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli · 2019
Cited alongside, same era.
Distilling the knowledge of bert for sequence-to-sequence asr
Hayato Futami, Hirofumi Inaguma, Sei Ueno, Masato Mimura, Shinsuke Sakai, and Tatsuya Kawahara · 2020
Cited alongside, same era.
Does my multimodal model learn cross-modal interactions? it’s harder to tell than you might think!
Modality-specific learning rates for effective multimodal additive late-fusion
Yiqun Yao and Rada Mihalcea · 2022
Later among the works it cites.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Later among the works it cites.
Pengi: An audio language model for audio tasks
Soham Deshmukh, Benjamin Elizalde, Rita Singh, and Huaming Wang · 2023
Later among the works it cites.
Prompting large language models with speech recognition abilities
Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Junteng Jia, Yuan Shangguan, Ke Li, Jinxi Guo, Wenhan Xiong, Jay Mahadeokar, Ozlem Kalinli, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jack Hessel and Lillian Lee · 2020
Cited alongside, same era.
Deliberation model based two-pass end-to-end speech recognition
Ke Hu, Tara N Sainath, Ruoming Pang, and Rohit Prabhavalkar · 2020
Cited alongside, same era.
Uncertainty estimation in autoregressive structured prediction
Andrey Malinin and Mark Gales · 2020
Cited alongside, same era.
Calibrating deep neural networks using focal loss
Jishnu Mukhoti, Viveka Kulharia, Amartya Sanyal, Stuart Golodetz, Philip Torr, and Puneet Dokania · 2020
Cited alongside, same era.
Multimodal learning with incomplete modalities by knowledge distillation
Qi Wang, Liang Zhan, Paul Thompson, and Jiayu Zhou · 2020
Cited alongside, same era.
Don’t just blame over-parametrization for over-confidence: Theoretical analysis of calibration in binary classification
Yu Bai, Song Mei, Huan Wang, and Caiming Xiong · 2021
Cited alongside, same era.
Multi-domain knowledge distillation via uncertainty-matching for end-to-end asr models
Ho-Gyeong Kim, Min-Joong Lee, Hoshik Lee, Tae Gyoon Kang, Jihyun Lee, Eunho Yang, and Sung Ju Hwang · 2021
Cited alongside, same era.
Layoutxlm: Multimodal pre-training for multilingual visually-rich document understanding
Yiheng Xu, Tengchao Lv, Lei Cui, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, and Furu Wei · 2021
Cited alongside, same era.
Jiaming Han, Renrui Zhang, Wenqi Shao, Peng Gao, Peng Xu, Han Xiao, Kaipeng Zhang, Chris Liu, Song Wen, Ziyu Guo, et al · 2023
Later among the works it cites.
Low-resource music genre classification with cross-modal neural model reprogramming
Yun-Ning Hung, Chao-Han Huck Yang, Pin-Yu Chen, and Alexander Lerch · 2023
Later among the works it cites.
Yuxuan Lei, Dingkang Yang, Mingcheng Li, Shunli Wang, Jiawei Chen, and Lihua Zhang · 2023
Later among the works it cites.
Softcorrect: Error correction with soft detection for automatic speech recognition
Yichong Leng, Xu Tan, Wenjie Liu, Kaitao Song, Rui Wang, Xiang-Yang Li, Tao Qin, Ed Lin, and Tie-Yan Liu · 2023
Later among the works it cites.
Mimic-it: Multi-modal in-context instruction tuning
Bo Li, Yuanhan Zhang, Liangyu Chen, Jinghao Wang, Fanyi Pu, Jingkang Yang, Chunyuan Li, and Ziwei Liu · 2023
Later among the works it cites.
Vision transformers are parameter-efficient audio-visual learners
Yan-Bo Lin, Yi-Lin Sung, Jie Lei, Mohit Bansal, and Gedas Bertasius · 2023
Later among the works it cites.
Reducing language confusion for code-switching speech recognition with token-level language diarization
Hexin Liu, Haihua Xu, Leibny Paola Garcia, Andy WH Khong, Yi He, and Sanjeev Khudanpur · 2023
Later among the works it cites.
Macaw-llm: Multi-modal language modeling with image, audio, video, and text integration
Chenyang Lyu, Minghao Wu, Longyue Wang, Xinting Huang, Bingshuai Liu, Zefeng Du, Shuming Shi, and Zhaopeng Tu · 2023
Later among the works it cites.
Gpt-4 technical report
R OpenAI · 2023
Later among the works it cites.
Kosmos-2: Grounding multimodal large language models to the world
Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao, Shaohan Huang, Shuming Ma, and Furu Wei · 2023
Later among the works it cites.
Whispering llama: A cross-modal generative error correction framework for speech recognition
Srijith Radhakrishnan, Chao-Han Yang, Sumeer Khan, Rohit Kumar, Narsis Kiani, David Gomez-Cabrero, and Jesper Tegnér · 2023
Later among the works it cites.
An empirical study of multimodal model merging
Yi-Lin Sung, Linjie Li, Kevin Lin, Zhe Gan, Mohit Bansal, and Lijuan Wang · 2023
Later among the works it cites.
Caption anything: Interactive image description with diverse multimodal controls
Teng Wang, Jinrui Zhang, Junjie Fei, Yixiao Ge, Hao Zheng, Yunlong Tang, Zhe Li, Mingqi Gao, Shanshan Zhao, Ying Shan, et al · 2023
Later among the works it cites.
Generative speech recognition error correction with large language models and task-activating prompting
Chao-Han Huck Yang, Yile Gu, Yi-Chieh Liu, Shalini Ghosh, Ivan Bulyko, and Andreas Stolcke · 2023
Later among the works it cites.
Invariant training 2d-3d joint hard samples for few-shot point cloud recognition
Xuanyu Yi, Jiajun Deng, Qianru Sun, Xian-Sheng Hua, Joo-Hwee Lim, and Hanwang Zhang · 2023
Later among the works it cites.
Low-rank adaptation of large language model rescoring for parameter-efficient speech recognition
Yu Yu, Chao-Han Huck Yang, Jari Kolehmainen, Prashanth G Shivakumar, Yile Gu, Sungho Ryu Roger Ren, Qi Luo, Aditya Gourav, I-Fan Chen, Yi-Chieh Liu, et al · 2023
Later among the works it cites.