Fetching the paper…
Reading the bibliography…
Much recent work on Spoken Language Understanding (SLU) is limited in at least one of three ways: models were trained on oracle text input and neglected ASR errors, models were trained to predict only intents without the slot values, or models were trained on a large amount of in-house data.
“The atis spoken language systems pilot corpus,”
Charles T Hemphill, John J Godfrey, and George R Doddington, · 1990
Earlier work this paper cites.
“Tina: A natural language system for spoken language applications,”
Stephanie Seneff, · 1992
Earlier work this paper cites.
“What is left to be understood in atis?,”
Gokhan Tur, Dilek Hakkani-Tür, and Larry Heck, · 2010
Earlier work this paper cites.
“Multilingual spoken-language understanding in the mit voyager system,”
James Glass, Giovanni Flammia, David Goodine, Michael Phillips, Joseph Polifroni, Shinsuke Sakai, Stephanie Seneff, and Victor Zue, · 2010
Earlier work this paper cites.
“End-to-end learning of semantic role labeling using recurrent neural networks,”
Jie Zhou and Wei Xu, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“Encoder-decoder with focus-mechanism for sequence labelling based spoken language understanding,”
Su Zhu and Kai Yu, · 2017
Earlier work this paper cites.
“Hybrid ctc/attention architecture for end-to-end speech recognition,”
Shinji Watanabe, Takaaki Hori, Suyoun Kim, John R Hershey, and Tomoki Hayashi, · 2017
Earlier work this paper cites.
Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, et al., · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Earlier work this paper cites.
“From audio to semantics: Approaches to end-to-end spoken language understanding,”
Parisa Haghani, Arun Narayanan, Michiel Bacchiani, Galen Chuang, Neeraj Gaur, Pedro Moreno, Rohit Prabhavalkar, Zhongdi Qu, and Austin Waters, · 2018
Earlier work this paper cites.
“Towards end-to-end spoken language understanding,”
Dmitriy Serdyuk, Yongqiang Wang, Christian Fuegen, Anuj Kumar, Baiyang Liu, and Yoshua Bengio, · 2018
Cited alongside, same era.
“End-to-end named entity extraction from speech,”
Sahar Ghannay, Antoine Caubriere, Yannick Esteve, Antoine Laurent, and Emmanuel Morin, · 2018
Cited alongside, same era.
Taku Kudo and John Richardson, · 2018
Cited alongside, same era.
“Representation learning with contrastive predictive coding,”
Aaron van den Oord, Yazhe Li, and Oriol Vinyals, · 2018
Cited alongside, same era.
“Is atis too shallow to go deeper for benchmarking spoken language understanding models?,”
Frédéric Béchet and Christian Raymond, · 2018
“A scalable noisy speech dataset and online subjective test framework,”
Chandan KA Reddy, Ebrahim Beyrami, Jamie Pool, Ross Cutler, Sriram Srinivasan, and Johannes Gehrke, · 2019
Later among the works it cites.
“Specaugment: A simple data augmentation method for automatic speech recognition,”
Daniel S Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D Cubuk, and Quoc V Le, · 2019
Later among the works it cites.
“Speechbert: Cross-modal pre-trained language model for end-to-end spoken question answering,”
Yung-Sung Chuang, Chi-Liang Liu, and Hung-Yi Lee, · 2019
Later among the works it cites.
“Learning asr-robust contextualized embeddings for spoken language understanding,”
Chao-Wei Huang and Yun-Nung Chen, · 2020
Closest in time.
“Large-scale unsupervised pre-training for end-to-end spoken language understanding,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Gunrock: A social bot for complex and engaging long conversations,”
Dian Yu, Michelle Cohn, Yi Mang Yang, Chun-Yen Chen, Weiming Wen, Jiaping Zhang, Mingyang Zhou, Kevin Jesse, Austin Chau, Antara Bhowmick, et al., · 2019
Cited alongside, same era.
“Speech model pre-training for end-to-end spoken language understanding,”
Loren Lugosch, Mirco Ravanelli, Patrick Ignoto, Vikrant Singh Tomar, and Yoshua Bengio, · 2019
Cited alongside, same era.
“Bert for joint intent classification and slot filling,”
Qian Chen, Zhu Zhuo, and Wen Wang, · 2019
Cited alongside, same era.
“Recent advances in end-to-end spoken language understanding,”
Natalia Tomashenko, Antoine Caubrière, Yannick Estève, Antoine Laurent, and Emmanuel Morin, · 2019
Cited alongside, same era.
“An unsupervised autoregressive model for speech representation learning,”
Yu-An Chung, Wei-Ning Hsu, Hao Tang, and James Glass, · 2019
Cited alongside, same era.
“wav2vec: Unsupervised pre-training for speech recognition,”
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli, · 2019
Cited alongside, same era.
“Mitigating the impact of speech recognition errors on spoken question answering by adversarial domain adaptation,”
Chia-Hsuan Lee, Yun-Nung Chen, and Hung-Yi Lee, · 2019
Cited alongside, same era.
Pengwei Wang, Liangchen Wei, Yong Cao, Jinghui Xie, and Zaiqing Nie, · 2020
Closest in time.
“Speech to text adaptation: Towards an efficient cross-modal distillation,”
Won Ik Cho, Donghyun Kwak, Jiwon Yoon, and Nam Soo Kim, · 2020
Closest in time.
“End-to-end neural transformer based spoken language understanding,”
Martin Radfar, Athanasios Mouchtaris, and Siegfried Kunzmann, · 2020
Closest in time.
“Speech to semantics: Improve asr and nlu jointly via all-neural interfaces,”
Milind Rao, Anirudh Raju, Pranav Dheram, Bach Bui, and Ariya Rastrow, · 2020
Closest in time.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Closest in time.
“Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,”
Andy T Liu, Shu-wen Yang, Po-Han Chi, Po-chun Hsu, and Hung-yi Lee, · 2020
Closest in time.
“Style attuned pre-training and parameter efficient fine-tuning for spoken language understanding,”
Jin Cao, Jun Wang, Wael Hamza, Kelly Vanee, and Shang-Wen Li, · 2020
Closest in time.