Fetching the paper…
Reading the bibliography…
Conventional spoken language understanding systems consist of two main components: an automatic speech recognition module that converts audio to a transcript, and a natural language understanding module that transforms the resulting text (or top N hypotheses) into a set of domains, intents, and arguments.
“Multitask learning,”
Rich Caruana, · 1997
Earlier work this paper cites.
“Long short-term memory,”
Jürgen Schmidhuber and Sepp Hochreiter, · 1997
Earlier work this paper cites.
“Bidirectional recurrent neural networks,”
Mike Schuster and Kuldip K Paliwal, · 1997
Earlier work this paper cites.
“Introduction to the conll-2003 shared task: Language-independent named entity recognition,”
Erik F. Tjong Kim Sang and Fien De Meulder, · 2003
Earlier work this paper cites.
“Beyond asr 1-best: Using word confusion networks in spoken language understanding,”
Dilek Hakkani-Tür, Frédéric Béchet, Giuseppe Riccardi, and Gokhan Tur, · 2006
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
Spoken language understanding: Systems for extracting semantic information from speech
Gokhan Tur and Renato De Mori, · 2011
Earlier work this paper cites.
“A reranking approach for recognition and classification of speech input in conversational dialogue systems,”
Fabrizio Morbini, Kartik Audhkhasi, Ron Artstein, Maarten Van Segbroeck, Kenji Sagae, Panayiotis Georgiou, David R Traum, and Shri Narayanan, · 2012
Earlier work this paper cites.
“Sequence to sequence learning with neural networks,”
Ilya Sutskever, Oriol Vinyals, and Quoc V Le, · 2014
Earlier work this paper cites.
“Learning phrase representations using rnn encoder–decoder for statistical machine translation,”
Kyunghyun Cho, Bart van Merriënboer, Çağlar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Multi-task sequence to sequence learning,”
Minh-Thang Luong, Quoc V Le, Ilya Sutskever, Oriol Vinyals, and Lukasz Kaiser, · 2015
Cited alongside, same era.
“Fast and accurate recurrent neural network acoustic models for speech recognition,”
Hasim Sak, Andrew W. Senior, Kanishka Rao, and Françoise Beaufays, · 2015
Cited alongside, same era.
“Scheduled sampling for sequence prediction with recurrent neural networks,”
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer, · 2015
Cited alongside, same era.
“Multi-domain joint semantic frame parsing using bi-directional rnn-lstm.,”
Dilek Hakkani-Tür, Gökhan Tür, Asli Celikyilmaz, Yun-Nung Chen, Jianfeng Gao, Li Deng, and Ye-Yi Wang, · 2016
Cited alongside, same era.
“Latticernn: Recurrent neural networks over lattices.,”
Faisal Ladhak, Ankur Gandhe, Markus Dreyer, Lambert Mathias, Ariya Rastrow, and Björn Hoffmeister, · 2016
Cited alongside, same era.
“English conversational telephone speech recognition by humans and machines,”
George Saon, Gakuto Kurata, Tom Sercu, Kartik Audhkhasi, Samuel Thomas, Dimitrios Dimitriadis, Xiaodong Cui, Bhuvana Ramabhadran, Michael Picheny, Lynn-Li Lim, Bergul Roomi, and Phil Hall, · 2017
Later among the works it cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
C.C. Chiu, T.N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R.J. Weiss, K. Rao, K. Gonina, and N. Jaitly, · 2017
Later among the works it cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, · 2017
Later among the works it cites.
“Optimizing expected word error rate via sampling for speech recognition,”
Matt Shannon, · 2017
Later among the works it cites.
“Minimum word error rate training for attention-based sequence-to-sequence models,”
Rohit Prabhavalkar, Tara N Sainath, Yonghui Wu, Patrick Nguyen, Zhifeng Chen, Chung-Cheng Chiu, and Anjuli Kannan, · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al., · 2016
Cited alongside, same era.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Cited alongside, same era.
“Attention-based recurrent neural network models for joint intent detection and slot filling,”
Bing Liu and Ian Lane, · 2016
Cited alongside, same era.
“Tensorflow: a system for large-scale machine learning.,”
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al., · 2016
Cited alongside, same era.
“Acoustic modeling for google home,”
Bo Li, Tara Sainath, Arun Narayanan, Joe Caroselli, Michiel Bacchiani, Ananya Misra, Izhak Shafran, Hasim Sak, Golan Pundak, Kean Chin, et al., · 2017
Cited alongside, same era.
“Onenet: Joint domain, intent, slot prediction for spoken language understanding,”
Young-Bum Kim, Sungjin Lee, and Karl Stratos, · 2017
Cited alongside, same era.
“Comparing human and machine errors in conversational speech transcription,”
Andreas Stolcke and Jasha Droppo, · 2017
Cited alongside, same era.
Later among the works it cites.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Katya Gonina, et al., · 2017
Later among the works it cites.
“In-datacenter performance analysis of a tensor processing unit,”
Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, et al., · 2017
Later among the works it cites.
“https://ai.googleblog.com/2018/05/duplex-ai-system-for-natural-conversation.html,”
“Google duplex: An AI system for accomplishing real-world tasks over the phone,” · 2018
Closest in time.
“Towards end-to-end spoken language understanding,”
Dmitriy Serdyuk, Yongqiang Wang, Christian Fuegen, Anuj Kumar, Baiyang Liu, and Yoshua Bengio, · 2018
Closest in time.
“Incorporating asr errors with attention-based, jointly trained rnn for intent detection and slot filling,”
Raphael Schumann and Pongtep Angkititrakul, · 2018
Closest in time.
“Spoken language understanding without speech recognition,”
Yuan-Ping Chen, Ryan Price, and Srinivas Bangalore, · 2018
Closest in time.
“State-of-the-art speech recognition with sequence-to-sequence models,”
Chung-Cheng Chiu, Tara Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J. Weiss, Kanishka Rao, Katya Gonina, Navdeep Jaitly, Bo Li, Jan Chorowski, and Michiel Bacchiani, · 2018
Closest in time.