Fetching the paper…
Reading the bibliography…
Modern automatic speech recognition (ASR) model is required to accurately transcribe diverse speech signals (from different domains, languages, accents, etc) given the specific contextual information in various application scenarios.
Switchboard: Telephone speech corpus for research and development
John J Godfrey, Edward C Holliman, and Jane McDaniel · 1992
Earlier work this paper cites.
The ami meeting corpus
Wessel Kraaij, Thomas Hain, Mike Lincoln, and Wilfried Post · 2005
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves · 2012
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
the language atlas of china
Chinese Academy of Social Sciences and City University of Hong Kong · 2012
Earlier work this paper cites.
Eesen: End-to-end speech recognition using deep rnn models and wfst-based decoding
Yajie Miao, Mohammad Gowayyed, and Florian Metze · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
End-to-end attention-based large vocabulary speech recognition
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio · 2016
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals · 2016
Earlier work this paper cites.
Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng · 2017
Earlier work this paper cites.
Hybrid ctc/attention architecture for end-to-end speech recognition
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Aishell-2: Transforming mandarin asr research into industrial scale
Jiayu Du, Xingyu Na, Xuechen Liu, and Hui Bu · 2018
Earlier work this paper cites.
Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation
François Hernandez, Vincent Nguyen, Sahar Ghannay, Natalia Tomashenko, and Yannick Esteve · 2018
Earlier work this paper cites.
Minimum word error rate training for attention-based sequence-to-sequence models
Rohit Prabhavalkar, Tara N Sainath, Yonghui Wu, Patrick Nguyen, Zhifeng Chen, Chung-Cheng Chiu, and Anjuli Kannan · 2018
Earlier work this paper cites.
Streaming end-to-end speech recognition for mobile devices
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Shallow-fusion end-to-end contextual biasing
Ding Zhao, Tara N Sainath, David Rybach, Pat Rondon, Deepti Bhatia, Bo Li, and Ruoming Pang · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Cif: Continuous integrate-and-fire for end-to-end speech recognition
Linhao Dong and Bo Xu · 2020
Cited alongside, same era.
Conformer: Convolution-augmented transformer for speech recognition
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2023
Later among the works it cites.
Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models
Yunfei Chu, Jin Xu, Xiaohuan Zhou, Qian Yang, Shiliang Zhang, Zhijie Yan, Chang Zhou, and Jingren Zhou · 2023
Later among the works it cites.
Fleurs: Few-shot learning evaluation of universal representations of speech
Alexis Conneau, Min Ma, Simran Khanuja, Yu Zhang, Vera Axelrod, Siddharth Dalmia, Jason Riesa, Clara Rivera, and Ankur Bapna · 2023
Later among the works it cites.
Prompting large language models for zero-shot domain adaptation in speech recognition
Y. Li, Y. Wu, J. Li, and S. Liu · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Mls: A large-scale multilingual dataset for speech research
Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve, and Ronan Collobert · 2020
Cited alongside, same era.
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units
W. Hsu, B. Bolte, Y. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed · 2021
Cited alongside, same era.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al · 2022
Cited alongside, same era.
Self-supervised learning with random-projection quantizer for speech recognition
Chung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu, and Yonghui Wu · 2022
Cited alongside, same era.
Zhifu Gao, Shiliang Zhang, Ian McLoughlin, and Zhijie Yan · 2022
Cited alongside, same era.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Later among the works it cites.
AudioPaLM: A large language model that can speak and listen
P.K. Rubenstein, C. Asawaroengchai, D.D. Nguyen, et al · 2023
Later among the works it cites.
Salmonn: Towards generic hearing abilities for large language models
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
On decoder-only architecture for speech-to-text and large language model integration
J. Wu, Y. Gaur, Z. Chen, et al · 2023
Later among the works it cites.
Google usm: Scaling automatic speech recognition beyond 100 languages
Yu Zhang, Wei Han, James Qin, Yongqiang Wang, Ankur Bapna, Zhehuai Chen, Nanxin Chen, Bo Li, Vera Axelrod, Gary Wang, et al · 2023
Later among the works it cites.
Salm: Speech-augmented language model with in-context learning for speech recognition and translation
Zhehuai Chen, He Huang, Andrei Andrusenko, Oleksii Hrinchuk, Krishna C Puvvada, Jason Li, Subhankar Ghosh, Jagadeesh Balam, and Boris Ginsburg · 2024
Closest in time.
Unveiling the potential of llm-based asr on chinese open-source datasets
Xuelong Geng, Tianyi Xu, Kun Wei, Bingsheng Mu, Hongfei Xue, He Wang, Yangze Li, Pengcheng Guo, Yuhang Dai, Longhao Li, et al · 2024
Closest in time.
Audiogpt: Understanding and generating speech, music, sound, and talking head
Rongjie Huang, Mingze Li, Dongchao Yang, Jiatong Shi, Xuankai Chang, Zhenhui Ye, Yuning Wu, Zhiqing Hong, Jiawei Huang, Jinglin Liu, et al · 2024
Closest in time.
Anatomy of industrial scale multilingual asr
Francis McCann Ramirez, Luka Chkhetiani, Andrew Ehrenberg, Robert McHardy, Rami Botros, Yash Khare, Andrea Vanzo, Taufiquzzaman Peyash, Gabriel Oexle, Michael Liang, et al · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Closest in time.