Fetching the paper…
Reading the bibliography…
End-to-end spoken language understanding (SLU) models are a class of model architectures that predict semantics directly from speech.
L. Hubert and P. Arabie, “Comparing partitions,”
1985
Earlier work this paper cites.
D. M. Blei, A. Y. Ng, and M. I. Jordan, “Latent dirichlet allocation,”
2003
Earlier work this paper cites.
C.-J. Lee, S.-K. Jung, K.-D. Kim, D.-H. Lee, and G. G.-B. Lee, “Recent approaches to dialog management for spoken dialog systems,”
2010
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath
2012
Earlier work this paper cites.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in
2013
Earlier work this paper cites.
P. Xu and R. Sarikaya, “Contextual domain classification in spoken language understanding systems using recurrent neural network,” in
2014
Earlier work this paper cites.
R. Sarikaya, G. E. Hinton, and A. Deoras, “Application of deep belief networks for natural language understanding,”
2014
Cited alongside, same era.
S. Ravuri and A. Stolcke, “Recurrent neural network and lstm models for lexical utterance classification,” in
2015
Cited alongside, same era.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an asr corpus based on public domain audio books,” in
2015
Cited alongside, same era.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in
2016
Cited alongside, same era.
P. Haghani, A. Narayanan, M. Bacchiani, G. Chuang, N. Gaur, P. Moreno, R. Prabhavalkar, Z. Qu, and A. Waters, “From audio to semantics: Approaches to end-to-end spoken language understanding,” in
2018
Cited alongside, same era.
D. Serdyuk, Y. Wang, C. Fuegen, A. Kumar, B. Liu, and Y. Bengio, “Towards end-to-end spoken language understanding,” in
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
2019
Later among the works it cites.
Picovoice, “Speech-to-intent benchmark,” https://github.com/Picovoice/speech-to-intent-benchmark, 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y.-P. Chen, R. Price, and S. Bangalore, “Spoken language understanding without speech recognition,” in
2018
Cited alongside, same era.
2020
Closest in time.