Fetching the paper…
Reading the bibliography…
Advancements in deep neural networks have allowed automatic speech recognition (ASR) systems to attain human parity on several publicly available clean speech datasets.
The atis spoken language systems pilot corpus
Charles T Hemphill, John J Godfrey, and George R Doddington · 1990
Earlier work this paper cites.
Switchboard: Telephone speech corpus for research and development
John J Godfrey, Edward C Holliman, and Jane McDaniel · 1992
Earlier work this paper cites.
The design for the wall street journal-based csr corpus
Douglas B Paul and Janet Baker · 1992
Earlier work this paper cites.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
Error correction and learner perceptions in l2 spanish writing
Terri A Greenslade and J César Félix-Brasdefer · 2006
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Better evaluation for grammatical error correction
Daniel Dahlmeier and Hwee Tou Ng · 2012
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Alex Graves · 2012
Earlier work this paper cites.
Search results based n-best hypothesis rescoring with maximum entropy classification
Fuchun Peng, Scott Roy, Ben Shahshahani, and Françoise Beaufays · 2013
Earlier work this paper cites.
An overview of noise-robust automatic speech recognition
Jinyu Li, Li Deng, Yifan Gong, and Reinhold Haeb-Umbach · 2014
Earlier work this paper cites.
Bidirectional recurrent neural network language models for automatic speech recognition
Ebru Arisoy, Abhinav Sethy, Bhuvana Ramabhadran, and Stanley Chen · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals · 2016
Earlier work this paper cites.
Towards better decoding and language model integration in sequence to sequence models
Jan Chorowski and Navdeep Jaitly · 2016
Earlier work this paper cites.
Exploiting n-best hypotheses to improve an smt approach to grammatical error correction
Duc Tam Hoang, Shamil Chollampatt, and Hwee Tou Ng · 2016
Earlier work this paper cites.
The 4th chime speech separation and recognition challenge
Emmanuel Vincent, Shinji Watanabe, Jon Barker, and Ricard Marxer · 2016
Earlier work this paper cites.
Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline
Hui Bu, Jiayu Du, Xingyu Na, Bengu Wu, and Hao Zheng · 2017
Earlier work this paper cites.
Asr in classroom today: Automatic visualization of conceptual network in science classrooms
Daniela Caballero, Roberto Araya, Hanna Kronholm, Jouni Viiri, André Mansikkaniemi, Sami Lehesvuori, Tuomas Virtanen, and Mikko Kurimo · 2017
Earlier work this paper cites.
Lip reading sentences in the wild
Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2017
Earlier work this paper cites.
Cold fusion: Training seq2seq models together with language models
Anuroop Sriram, Heewoo Jun, Sanjeev Satheesh, and Adam Coates · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deliberation networks: Sequence generation beyond one-pass decoding
Yingce Xia, Fei Tian, Lijun Wu, Jianxin Lin, Tao Qin, Nenghai Yu, and Tie-Yan Liu · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Speech-transformer: a no-recurrence sequence-to-sequence model for speech recognition
Linhao Dong, Shuang Xu, and Bo Xu · 2018
Earlier work this paper cites.
Achieving human parity on automatic chinese to english news translation
Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mu Li, et al · 2018
Earlier work this paper cites.
Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation
François Hernandez, Vincent Nguyen, Sahar Ghannay, Natalia Tomashenko, and Yannick Esteve · 2018
Earlier work this paper cites.
An analysis of incorporating an external language model into a sequence-to-sequence model
Anjuli Kannan, Yonghui Wu, Patrick Nguyen, Tara N Sainath, Zhijeng Chen, and Rohit Prabhavalkar · 2018
Earlier work this paper cites.
A comparison of techniques for language model integration in encoder-decoder speech recognition
Shubham Toshniwal, Anjuli Kannan, Chung-Cheng Chiu, Yonghui Wu, Tara N Sainath, and Karen Livescu · 2018
Earlier work this paper cites.
Espnet: End-to-end speech processing toolkit
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, et al · 2018
Earlier work this paper cites.
Improved training of end-to-end attention models for speech recognition
Albert Zeyer, Kazuki Irie, Ralf Schlüter, and Hermann Ney · 2018
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M Tyers, and Gregor Weber · 2019
Earlier work this paper cites.
Bert for joint intent classification and slot filling
Qian Chen, Zhu Zhuo, and Wen Wang · 2019
Earlier work this paper cites.
An empirical study of efficient asr rescoring with transformers
Hongzhao Huang and Fuchun Peng · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Julian Salazar, Davis Liang, Toan Q Nguyen, and Katrin Kirchhoff · 2019
Earlier work this paper cites.
Component fusion: Learning replaceable language model component for end-to-end speech recognition system
Changhao Shan, Chao Weng, Guangsen Wang, Dan Su, Min Luo, Dong Yu, and Lei Xie · 2019
Earlier work this paper cites.
Effective sentence scoring method using bert for speech recognition
Joonbo Shin, Yoonhyung Lee, and Kyomin Jung · 2019
Cited alongside, same era.
Spoken language intent detection using confusion2vec
Prashanth Gurunath Shivakumar, Mu Yang, and Panayiotis Georgiou · 2019
Cited alongside, same era.
Automatic spelling correction with transformer for ctc-based end-to-end speech recognition
Shiliang Zhang, Ming Lei, and Zhijie Yan · 2019
Cited alongside, same era.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Armen Aghajanyan, Luke Zettlemoyer, and Sonal Gupta · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Bart based semantic correction for mandarin automatic speech recognition system
Yun Zhao, Xuerui Yang, Jinchao Wang, Yongyu Gao, Chao Yan, and Yuanfu Zhou · 2021
Later among the works it cites.
Data distributional properties drive emergent in-context learning in transformers
Stephanie Chan, Adam Santoro, Andrew Lampinen, Jane Wang, Aaditya Singh, Pierre Richemond, James McClelland, and Felix Hill · 2022
Later among the works it cites.
Noise-robust speech recognition with 10 minutes unparalleled in-domain data
Chen Chen, Nana Hou, Yuchen Hu, Shashank Shirol, and Eng Siong Chng · 2022
Later among the works it cites.
Wavlm: Large-scale self-supervised pre-training for full stack speech processing
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al · 2022
Later among the works it cites.
Prompt-augmented linear probing: Scaling beyond the limit of few-shot in-context learners
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Audio-attention discriminative language model for asr rescoring
Ankur Gandhe and Ariya Rastrow · 2020
Cited alongside, same era.
Conformer: Convolution-augmented transformer for speech recognition
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, et al · 2020
Cited alongside, same era.
Inserting punctuation to asr output in a real-time production environment
Pavel Hlubík, Martin Španěl, Marek Boháč, and Lenka Weingartová · 2020
Cited alongside, same era.
Deliberation model based two-pass end-to-end speech recognition
Ke Hu, Tara N Sainath, Ruoming Pang, and Rohit Prabhavalkar · 2020
Cited alongside, same era.
Libri-light: A benchmark for asr with limited or no supervision
Jacob Kahn, Morgane Riviere, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al · 2020
Cited alongside, same era.
Speech technology for healthcare: Opportunities, challenges, and state of the art
Siddique Latif, Junaid Qadir, Adnan Qayyum, Muhammad Usama, and Shahzad Younis · 2020
Cited alongside, same era.
Improving spoken language understanding by exploiting asr n-best hypotheses
Mingda Li, Weitong Ruan, Xinyue Liu, Luca Soldaini, Wael Hamza, and Chengwei Su · 2020
Cited alongside, same era.
Hyunsoo Cho, Hyuhng Joon Kim, Junyeob Kim, Sang-Woo Lee, Sang-goo Lee, Kang Min Yoo, and Taeuk Kim · 2022
Later among the works it cites.
Error correction in asr using sequence-to-sequence models
Samrat Dutta, Shreyansh Jain, Ayush Maheshwari, Souvik Pal, Ganesh Ramakrishnan, and Preethi Jyothi · 2022
Later among the works it cites.
Vision-language pre-training: Basics, recent advances, and future trends
Zhe Gan, Linjie Li, Chunyuan Li, Lijuan Wang, Zicheng Liu, Jianfeng Gao, et al · 2022
Later among the works it cites.
Yukun Huang, Yanda Chen, Zhou Yu, and Kathleen McKeown · 2022
Later among the works it cites.
Can language models learn from explanations in context?
Andrew K Lampinen, Ishita Dasgupta, Stephanie CY Chan, Kory Matthewson, Michael Henry Tessler, Antonia Creswell, James L McClelland, Jane X Wang, and Felix Hill · 2022
Later among the works it cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, et al · 2022
Later among the works it cites.
Mask the correct tokens: An embarrassingly simple approach for error correction
Kai Shen, Yichong Leng, Xu Tan, Siliang Tang, Yuan Zhang, Wenjie Liu, and Edward Lin · 2022
Later among the works it cites.
Effect and analysis of large-scale language model rescoring on competitive asr systems
Takuma Udagawa, Masayuki Suzuki, Gakuto Kurata, Nobuyasu Itoh, and George Saon · 2022
Later among the works it cites.
Deliberation of streaming rnn-transducer by non-autoregressive decoding
Weiran Wang, Ke Hu, and Tara N Sainath · 2022
Later among the works it cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Later among the works it cites.
Automatic speech recognition in german: A detailed error analysis
Johannes Wirth and Rene Peinl · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Later among the works it cites.
Speechprompt v2: Prompt tuning for speech classification tasks
Kai-Wei Chang, Yu-Kai Wang, Hua Shen, Iu-thing Kang, Wei-Cheng Tseng, Shang-Wen Li, and Hung-yi Lee · 2023
Closest in time.
How to estimate model transferability of pre-trained speech models?
Zih-Ching Chen, Chao-Han Huck Yang, Bo Li, Yu Zhang, Nanxin Chen, Shou-Yiin Chang, Rohit Prabhavalkar, Hung-yi Lee, and Tara N Sainath · 2023
Closest in time.
Putting natural in natural language processing
Grzegorz Chrupała · 2023
Closest in time.
Scaling up deliberation for multilingual asr
Ke Hu, Bo Li, and Tara N Sainath · 2023
Closest in time.
Llm-adapters: An adapter family for parameter-efficient fine-tuning of large language models
Zhiqiang Hu, Yihuai Lan, Lei Wang, Wanyu Xu, Ee-Peng Lim, Roy Ka-Wei Lee, Lidong Bing, and Soujanya Poria · 2023
Closest in time.
Gpt-4 vs. gpt-3.5: A concise showdown
Anis Koubaa · 2023
Closest in time.
Softcorrect: Error correction with soft detection for automatic speech recognition
Yichong Leng, Xu Tan, Wenjie Liu, Kaitao Song, Rui Wang, Xiang-Yang Li, Tao Qin, Ed Lin, and Tie-Yan Liu · 2023
Closest in time.
Rao Ma, Mark JF Gales, Kate Knill, and Mengjie Qian · 2023
Closest in time.
Whispering llama: A cross-modal generative error correction framework for speech recognition
Srijith Radhakrishnan, Chao-Han Huck Yang, Sumeer Ahmad Khan, Rohit Kumar, Narsis A Kiani, David Gomez-Cabrero, and Jesper N Tegner · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization
Yidong Wang, Zhuohao Yu, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, et al · 2023
Closest in time.
Speechgen: Unlocking the generative power of speech language models with prompts
Haibin Wu, Kai-Wei Chang, Yuan-Kuei Wu, and Hung-yi Lee · 2023
Closest in time.
Chao-Han Huck Yang, Yile Gu, Yi-Chieh Liu, Shalini Ghosh, Ivan Bulyko, and Andreas Stolcke · 2023
Closest in time.
From english to more languages: Parameter-efficient model reprogramming for cross-lingual speech recognition
Chao-Han Huck Yang, Bo Li, Yu Zhang, Nanxin Chen, Rohit Prabhavalkar, Tara N Sainath, and Trevor Strohman · 2023
Closest in time.
Explanation selection using unlabeled data for in-context learning
Xi Ye and Greg Durrett · 2023
Closest in time.
Llama-adapter: Efficient fine-tuning of language models with zero-init attention
Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hongsheng Li, Peng Gao, and Yu Qiao · 2023
Closest in time.
What makes good examples for visual in-context learning?
Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu · 2023
Closest in time.