Fetching the paper…
Reading the bibliography…
The lack of speech data annotated with labels required for spoken language understanding (SLU) is often a major hurdle in building end-to-end (E2E) systems that can directly process speech inputs.
“The ATIS spoken language systems pilot corpus,”
Charles T Hemphill, John J Godfrey, and George R Doddington, · 1990
Earlier work this paper cites.
“Language model estimation for optimizing end-to-end performance of a natural language call routing system,”
Vaibhava Goel, Hong-Kwang J. Kuo, Sabine Deligne, and Cheng Wu, · 2005
Earlier work this paper cites.
“Sequence transduction with recurrent neural networks,”
Alex Graves, · 2012
Earlier work this paper cites.
“Exploring ASR-free end-to-end modeling to improve spoken language understanding in a cloud-based dialog system,”
Yao Qian, Rutuja Ubale, Vikram Ramanaryanan, Patrick Lange, David Suendermann-Oeft, Keelan Evanini, and Eugene Tsuprun, · 2017
Earlier work this paper cites.
“Exploring architectures, data and units for streaming end-to-end speech recognition with RNN-transducer,”
Kanishka Rao, Haşim Sak, and Rohit Prabhavalkar, · 2017
Earlier work this paper cites.
“Towards end-to-end spoken language understanding,”
Dmitriy Serdyuk, Yongqiang Wang, Christian Fuegen, Anuj Kumar, Baiyang Liu, and Yoshua Bengio, · 2018
Earlier work this paper cites.
“From audio to semantics: Approaches to end-to-end spoken language understanding,”
Parisa Haghani, Arun Narayanan, Michiel Bacchiani, Galen Chuang, Neeraj Gaur, Pedro Moreno, Rohit Prabhavalkar, Zhongdi Qu, and Austin Waters, · 2018
Earlier work this paper cites.
“Spoken language understanding without speech recognition,”
Yuan-Ping Chen, Ryan Price, and Srinivas Bangalore, · 2018
Earlier work this paper cites.
“End-to-end named entity and semantic concept extraction from speech,”
Sahar Ghannay, Antoine Caubrière, Yannick Estève, Nathalie Camelin, Edwin Simonnet, Antoine Laurent, and Emmanuel Morin, · 2018
Earlier work this paper cites.
“Speech model pre-training for end-to-end spoken language understanding,”
Loren Lugosch, Mirco Ravanelli, Patrick Ignoto, Vikrant Singh Tomar, and Yoshua Bengio, · 2019
Earlier work this paper cites.
“Curriculum-based transfer learning for an effective end-to-end spoken language understanding and domain portability,”
Antoine Caubrière, Natalia Tomashenko, Antoine Laurent, Emmanuel Morin, Nathalie Camelin, and Yannick Estève, · 2019
Cited alongside, same era.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al., · 2019
Cited alongside, same era.
“Improving RNN transducer modeling for end-to-end speech recognition,”
Jinyu Li, Rui Zhao, Hu Hu, and Yifan Gong, · 2019
Cited alongside, same era.
“Joint speech recognition and speaker diarization via sequence transduction,”
Laurent El Shafey, Hagen Soltau, and Izhak Shafran, · 2019
Cited alongside, same era.
“Sequence noise injected training for end-to-end speech recognition,”
George Saon, Zoltán Tüske, Kartik Audhkhasi, and Brian Kingsbury, · 2019
Cited alongside, same era.
“End-to-end neural transformer based spoken language understanding,”
Martin Radfar, Athanasios Mouchtaris, and Siegfried Kunzmann, · 2020
Later among the works it cites.
“Improving end-to-end speech-to-intent classification with Reptile,”
Yusheng Tian and Philip John Gorinski, · 2020
Later among the works it cites.
“Large-scale transfer learning for low-resource spoken language understanding,”
Xueli Jia, Jianzong Wang, Zhiyong Zhang, Ning Cheng, and Jing Xiao, · 2020
Later among the works it cites.
“End-to-end spoken language understanding without full transcripts,”
Hong-Kwang J. Kuo, Zoltán Tüske, Samuel Thomas, Yinghui Huang, Kartik Audhkhasi, Brian Kingsbury, Gakuto Kurata, Zvi Kons, Ron Hoory, and Luis Lastras, · 2020
Later among the works it cites.
“End-to-end architectures for ASR-free spoken language understanding,”
Elisavet Palogiannidi, Ioannis Gkinis, George Mastrapas, Petr Mizera, and Themos Stafylakis, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Cited alongside, same era.
“Leveraging unpaired text data for training end-to-end speech-to-intent systems,”
Yinghui Huang, Hong-Kwang Kuo, Samuel Thomas, Zvi Kons, Kartik Audhkhasi, Brian Kingsbury, Ron Hoory, and Michael Picheny, · 2020
Cited alongside, same era.
“Using speech synthesis to train end-to-end spoken language understanding models,”
Loren Lugosch, Brett H Meyer, Derek Nowrouzezahrai, and Mirco Ravanelli, · 2020
Cited alongside, same era.
“Improved end-to-end spoken utterance classification with a self-attention acoustic classifier,”
Ryan Price, Mahnoosh Mehrabani, and Srinivas Bangalore, · 2020
Cited alongside, same era.
Mohammadreza Ghodsi, Xiaofeng Liu, James Apfel, Rodrigo Cabrera, and Eugene Weinstein, · 2020
Later among the works it cites.
“Harpervalleybank: A domain-specific spoken dialog corpus,”
Mike Wu, Jonathan Nafziger, Anthony Scodary, and Andrew Maas, · 2020
Later among the works it cites.
“Rnn transducer models for spoken language understanding,”
Samuel Thomas, Hong-Kwang J Kuo, George Saon, Zoltán Tüske, Brian Kingsbury, Gakuto Kurata, Zvi Kons, and Ron Hoory, · 2021
Later among the works it cites.
“Advancing RNN transducer technology for speech recognition,”
George Saon, Zoltán Tüske, Daniel Bolanos, and Brian Kingsbury, · 2021
Later among the works it cites.