Fetching the paper…
Reading the bibliography…
Whereas conventional spoken language understanding (SLU) systems map speech to text, and then text to intent, end-to-end SLU systems map speech directly to intent through a single trainable model.
A. L. Gorin, G. Riccardi, and J. H. Wright, “How may I help you?”
1997
Earlier work this paper cites.
R. Caruana, “Multitask learning,”
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,”
1998
Earlier work this paper cites.
Y.-Y. Wang, L. Deng, and A. Acero, “Spoken language understanding,”
2005
Earlier work this paper cites.
S. Fernández, A. Graves, and J. Schmidhuber, “Sequence Labelling in Structured Domains with Hierarchical Recurrent Neural Networks,”
2007
Earlier work this paper cites.
G. Tur and R. D. Mori,
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,”
2012
Earlier work this paper cites.
G. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large vocabulary speech recognition,”
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
A. Graves and N. Jaitly, “Towards End-to-End Speech Recognition with Recurrent Neural Networks,”
2014
Earlier work this paper cites.
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?”
2014
Earlier work this paper cites.
G. Mesnil, Y. Dauphin, K. Yao, Y. Bengio, L. Deng, D. Hakkani-tur, and X. He, “Using Recurrent Neural Networks for Slot Filling in Spoken Language Understanding,”
2015
Earlier work this paper cites.
D. Amodei, R. Anubhai, E. Battenberg, C. Case, J. Casper, B. Catanzaro, J. Chen, M. Chrzanowski, A. Coates, G. Diamos
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,”
2015
Earlier work this paper cites.
A. M. Dai and Q. V. Le, “Semi-supervised sequence learning,”
2015
Earlier work this paper cites.
D. Wang and T. Zheng, “Transfer learning for speech and language processing,”
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “LibriSpeech: An ASR corpus based on public domain audio books,”
2015
Cited alongside, same era.
I. Goodfellow, Y. Bengio, and A. Courville,
2016
Cited alongside, same era.
2016
Cited alongside, same era.
Y. Qian, R. Ubale, V. Ramanarayanan, and P. Lange, “Exploring ASR-free end-to-end modeling to improve spoken language understanding in a cloud-based dialog system,”
2017
Cited alongside, same era.
2018
Later among the works it cites.
V. Renkens and H. Van hamme, “Capsule Networks for Low Resource Spoken Language Understanding,”
2018
Later among the works it cites.
J. Howard and S. Ruder, “Universal language model fine-tuning for text classification,”
2018
Later among the works it cites.
A. Radford and T. Salimans, “Improving Language Understanding by Generative Pre-Training,”
2018
Later among the works it cites.
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Kunze, L. Kirsch, I. Kurenkov, A. Krug, J. Johannsmeier, and S. Stober, “Transfer Learning for Speech Recognition on a Budget,”
2017
Cited alongside, same era.
P. Ghahremani, V. Manohar, H. Hadian, D. Povey, and S. Khudanpur, “Investigation of transfer learning for ASR using LF-MMI trained neural networks,”
2017
Cited alongside, same era.
B. Li, T. N. Sainath, A. Narayanan, J. Caroselli, M. Bacchiani, A. Misra, I. Shafran, H. Sak, G. Pundak, K. K. Chin
2017
Cited alongside, same era.
S. Sabour, N. Frosst, and G. E. Hinton, “Dynamic routing between capsules,”
2017
Cited alongside, same era.
K. Audhkhasi, B. Ramabhadran, G. Saon, M. Picheny, and D. Nahamoo, “Direct Acoustics-to-Word Models for English Conversational Speech Recognition,”
2017
Cited alongside, same era.
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal Forced Aligner: Trainable text-speech alignment using Kaldi,”
2017
Cited alongside, same era.
2018
Cited alongside, same era.
P. Warden, “Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition,”
2018
Later among the works it cites.
2018
Later among the works it cites.
K. Audhkhasi, B. Kingsbury, B. Ramabhadran, G. Saon, and M. Picheny, “Building competitive direct acoustics-to-word models for english conversational speech recognition,”
2018
Later among the works it cites.
R. Sanabria and F. Metze, “Hierarchical multitask learning with CTC,”
2018
Later among the works it cites.
2018
Later among the works it cites.
M. Ravanelli and Y. Bengio, “Speaker recognition from raw waveform with SincNet,”
2018
Later among the works it cites.
——, “Interpretable convolutional filters with SincNet,”
2018
Later among the works it cites.
V. Sanh, T. Wolf, and S. Ruder, “A hierarchical multi-task approach for learning embeddings from semantic tasks,”
2019
Closest in time.