Fetching the paper…
Reading the bibliography…
Collecting audio-text pairs is expensive; however, it is much easier to access text-only data.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Speech recognition with deep recurrent neural networks,”
Alex Graves, Abdel-Rahman Mohamed, and Geoffrey Hinton, · 2013
Earlier work this paper cites.
“Deep speech: Scaling up end-to-end speech recognition,”
Awni Hannun, Carl Case, Jared Casper, Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta, Adam Coates, et al., · 2014
Earlier work this paper cites.
“Attention-based models for speech recognition,”
Jan K. Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“LibriSpeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Hybrid CTC/attention architecture for end-to-end speech recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R. Hershey, and Tomoki Hayashi, · 2017
Earlier work this paper cites.
“Multi-modal data augmentation for end-to-end asr,”
Adithya Renduchintala, Shuoyang Ding, Matthew Wiesner, and Shinji Watanabe, · 2018
Earlier work this paper cites.
“Language modeling with deep transformers,”
Kazuki Irie, Albert Zeyer, Ralf Schlüter, and Hermann Ney, · 2019
Earlier work this paper cites.
“Independent language modeling architecture for end-to-end ASR,”
Haihua Xu, Yerbolat Khassanov, Zhiping Zeng, Eng Siong Chng, Chongjia Ni, Bin Ma, Haizhou Li, et al., · 2020
Earlier work this paper cites.
“Language models are few-shot learners,”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al., · 2020
Cited alongside, same era.
“Recent developments on espnet toolkit boosted by conformer,”
Pengcheng Guo, Florian Boyer, Xuankai Chang, Tomoki Hayashi, Yosuke Higuchi, et al., · 2020
Cited alongside, same era.
“Internal language model estimation for domain-adaptive end-to-end speech recognition,”
Zhong Meng, Sarangarajan Parthasarathy, Eric Sun, Yashesh Gaur, Naoyuki Kanda, Liang Lu, Xie Chen, Rui Zhao, Jinyu Li, and Yifan Gong, · 2021
Cited alongside, same era.
“Multitask training with text data for end-to-end speech recognition,”
Peidong Wang, Tara N. Sainath, and Ron J. Weiss, · 2021
Cited alongside, same era.
“CTC-based compression for direct speech translation,”
Marco Gaido, Mauro Cettolo, Matteo Negri, and Marco Turchi, · 2021
Cited alongside, same era.
“Speechprompt v2: Prompt tuning for speech classification tasks,”
Kai-Wei Chang, Yu-Kai Wang, Hua Shen, Iu-thing Kang, Wei-Cheng Tseng, Shang-Wen Li, and Hung-yi Lee, · 2023
Closest in time.
“SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,”
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu, · 2023
Closest in time.
“AudioPaLM: A large language model that can speak and listen,”
Paul K Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, et al., · 2023
Closest in time.
“Prompting large language models with speech recognition abilities,”
Yassir Fathullah, Chunyang Wu, Egor Lakomkin, Junteng Jia, Yuan Shangguan, Ke Li, Jinxi Guo, Wenhan Xiong, Jay Mahadeokar, Ozlem Kalinli, et al., · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“FSR: Accelerating the inference process of transducer-based models by applying fast-skip regularization,”
Zhengkun Tian, Jiangyan Yi, Ye Bai, Jianhua Tao, Shuai Zhang, and Zhengqi Wen, · 2021
Cited alongside, same era.
“Residual language model for end-to-end speech recognition,”
Emiru Tsunoo, Yosuke Kashiwagi, Chaitanya Prasad Narisetty, and Shinji Watanabe, · 2022
Cited alongside, same era.
“MAESTRO: Matched speech text representations through modality matching,”
Zhehuai Chen, Yu Zhang, Andrew Rosenberg, Bhuvana Ramabhadran, Pedro J. Moreno, Ankur Bapna, and Heiga Zen, · 2022
Cited alongside, same era.
“PaLM: Scaling language modeling with pathways,”
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al., · 2022
Cited alongside, same era.
Jian Wu, Yashesh Gaur, Zhuo Chen, Long Zhou, Yimeng Zhu, Tianrui Wang, Jinyu Li, Shujie Liu, Bo Ren, Linquan Liu, et al., · 2023
Closest in time.
“Integrating pretrained ASR and LM to perform sequence generation for spoken language understanding,”
Siddhant Arora, Hayato Futami, Yosuke Kashiwagi, Emiru Tsunoo, Brian Yan, and Shinji Watanabe, · 2023
Closest in time.
“Accelerating RNN-T training and inference using CTC guidance,”
Yongqiang Wang, Zhehuai Chen, Chengjian Zheng, Yu Zhang, Wei Han, and Parisa Haghani, · 2023
Closest in time.
“Librispeech transducer model with internal language model prior correction,”
Albert Zeyer, André Merboldt, Wilfried Michel, Ralf Schlüter, and Hermann Ney, · 2056
Closest in time.