Fetching the paper…
Reading the bibliography…
End-to-end (E2E) spoken language understanding (SLU) systems predict utterance semantics directly from speech using a single model.
P. Price, “Evaluation of spoken language systems: The ATIS domain,” in Speech and Natural Language: Proceedings of a Workshop Held at Hidden Valley, Pennsylvania, June 24-27, 1990 , 1990
1990
Earlier work this paper cites.
G. Tur and R. De Mori, Spoken language understanding: Systems for extracting semantic information from speech . John Wiley & Sons, 2011
2011
Earlier work this paper cites.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath et al. , “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal processing magazine , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
A. Graves, A. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2013, pp. 6645–6649
2013
Earlier work this paper cites.
P. Xu and R. Sarikaya, “Contextual domain classification in spoken language understanding systems using recurrent neural network,” in 2014 IEEE ICASSP . IEEE, 2014, pp. 136–140
2014
Earlier work this paper cites.
R. Sarikaya, G. E. Hinton, and A. Deoras, “Application of deep belief networks for natural language understanding,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 22, no. 4, pp. 778–784, 2014
2014
Earlier work this paper cites.
Y.-C. Tam, Y. Lei, J. Zheng, and W. Wang, “ASR error detection using recurrent neural network language model and complementary asr,” in 2014 IEEE ICASSP . IEEE, 2014, pp. 2312–2316
2014
Earlier work this paper cites.
S. Ravuri and A. Stolcke, “Recurrent neural network and LSTM models for lexical utterance classification,” in Sixteenth Annual Conference of the International Speech Communication Association , 2015
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in 2015 IEEE ICASSP . IEEE, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
D. Bahdanau, J. Chorowski, D. Serdyuk, P. Brakel, and Y. Bengio, “End-to-end attention-based large vocabulary speech recognition,” in 2016 IEEE ICASSP . IEEE, 2016, pp. 4945–4949
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
X. Yang, Y.-N. Chen, D. Hakkani-Tür, P. Crook, X. Li, J. Gao, and L. Deng, “End-to-end joint learning of natural language understanding and dialogue manager,” in 2017 IEEE ICASSP . IEEE, 2017, pp. 5690–5694
2017
Earlier work this paper cites.
P. Haghani, A. Narayanan, M. Bacchiani, G. Chuang, N. Gaur, P. Moreno, R. Prabhavalkar, Z. Qu, and A. Waters, “From audio to semantics: Approaches to end-to-end spoken language understanding,” in 2018 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2018, pp. 720–726
2018
Cited alongside, same era.
Y.-P. Chen, R. Price, and S. Bangalore, “Spoken language understanding without speech recognition,” in 2018 IEEE ICASSP . IEEE, 2018, pp. 6189–6193
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
L. Lugosch, B. H. Meyer, D. Nowrouzezahrai, and M. Ravanelli, “Using speech synthesis to train end-to-end spoken language understanding models,” in 2020 IEEE ICASSP . IEEE, 2020, pp. 8499–8503
2020
Later among the works it cites.
E. Palogiannidi, I. Gkinis, G. Mastrapas, P. Mizera, and T. Stafylakis, “End-to-end architectures for ASR-free spoken language understanding,” in 2020 IEEE ICASSP . IEEE, 2020, pp. 7974–7978
2020
Later among the works it cites.
M. Radfar, A. Mouchtaris, and S. Kunzmann, “End-to-End Neural Transformer Based Spoken Language Understanding,” in Proc. Interspeech 2020 . ISCA, 2020, pp. 866–870
2020
Later among the works it cites.
Y. Tian and P. J. Gorinski, “Improving end-to-end speech-to-intent classification with Reptile,” Proc. Interspeech 2020 , pp. 891–895, 2020
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C.-C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, E. Gonina et al. , “State-of-the-art speech recognition with sequence-to-sequence models,” in 2018 IEEE ICASSP . IEEE, 2018, pp. 4774–4778
2018
Cited alongside, same era.
J. Howard and S. Ruder, “Universal language model fine-tuning for text classification,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 328–339
2018
Cited alongside, same era.
R. Voleti, J. M. Liss, and V. Berisha, “Investigating the effects of word substitution errors on sentence embeddings,” in 2019 IEEE ICASSP . IEEE, 2019, pp. 7315–7319
2019
Cited alongside, same era.
M. Moore, M. Saxon, H. Venkateswara, V. Berisha, and S. Panchanathan, “Say what? a dataset for exploring the error patterns that two ASR engines make.” in INTERSPEECH , 2019, pp. 2528–2532
2019
Cited alongside, same era.
A. Raghuvanshi, V. Ramakrishnan, V. Embar, L. Carroll, and K. Raghunathan, “Entity resolution for noisy ASR transcripts,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP): System Demonstrations , 2019, pp. 61–66
2019
Cited alongside, same era.
L. Lugosch, M. Ravanelli, P. Ignoto, V. S. Tomar, and Y. Bengio, “Speech model pre-training for end-to-end spoken language understanding,” Proc. Interspeech 2019 , pp. 814–818, 2019
2019
Cited alongside, same era.
N. Tomashenko, A. Caubrière, Y. Estève, A. Laurent, and E. Morin, “Recent advances in end-to-end spoken language understanding,” in International Conference on Statistical Language and Speech Processing . Springer, 2019, pp. 44–55
2019
Cited alongside, same era.
H. Wang, S. Dong, Y. Liu, J. Logan, A. K. Agrawal, and Y. Liu, “ASR error correction with augmented transformer for entity retrieval,” Proc. Interspeech 2020 , pp. 1550–1554, 2020
2020
Cited alongside, same era.
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
J. P. McKenna, S. Choudhary, M. Saxon, G. P. Strimel, and A. Mouchtaris, “Semantic Complexity in End-to-End Spoken Language Understanding,” in Proc. Interspeech 2020 . ISCA, 2020, pp. 4273–4277
2020
Later among the works it cites.
H.-K. J. Kuo, Z. Tüske, S. Thomas, Y. Huang, K. Audhkhasi, B. Kingsbury, G. Kurata, Z. Kons, R. Hoory, and L. Lastras, “End-to-End Spoken Language Understanding Without Full Transcripts,” in Proc. Interspeech 2020 . ISCA, 2020, pp. 906–910
2020
Later among the works it cites.
M. Rao, A. Raju, P. Dheram, B. Bui, and A. Rastrow, “Speech to semantics: Improve ASR and NLU jointly via all-neural interfaces,” Proc. Interspeech 2020 , Oct 2020
2020
Later among the works it cites.
T. Wolf, J. Chaumond, L. Debut, V. Sanh, C. Delangue, A. Moi, P. Cistac, M. Funtowicz, J. Davison, S. Shleifer et al. , “Transformers: State-of-the-art natural language processing,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations , 2020, pp. 38–45
2020
Later among the works it cites.