Fetching the paper…
Reading the bibliography…
We propose a novel deliberation-based approach to end-to-end (E2E) spoken language understanding (SLU), where a streaming automatic speech recognition (ASR) model produces the first-pass hypothesis and a second-pass natural language understanding (NLU) component generates the semantic parse by conditioning on both ASR's text and audio embeddings.
A. Graves, “Sequence transduction with recurrent neural networks,” in ICML Representation Learning Workshop , 2012
2012
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proc. NIPS , 2017
2017
Earlier work this paper cites.
A. See, P. J. Liu, and C. D. Manning, “Get to the point: Summarization with pointer-generator networks,” in Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2017
2017
Earlier work this paper cites.
P. Haghani, A. Narayanan, M. A. U. Bacchiani, G. Chuang, N. Gaur, P. J. M. Mengibar, D. Qu, R. Prabhavalkar, and A. Waters, “From audio to semantics: Approaches to end-to-end spoken language understanding,” in Proc. SLT , 2018
2018
Earlier work this paper cites.
D. Serdyuk, Y. Wang, C. Fuegen, A. Kumar, B. Liu, and Y. Bengio, “Towards end-to-end spoken language understanding,” 2018
2018
Earlier work this paper cites.
S. Gupta, R. Shah, M. Mohit, A. Kumar, and M. Lewis, “Semantic Parsing for Task Oriented Dialog using Hierarchical Representations,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2018
2018
Earlier work this paper cites.
T. Kudo, “Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates,” in Proc. ACL , 2018
2018
Earlier work this paper cites.
T. Kudo and J. Richardson, “SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing,” in Proc. EMNLP: System Demonstrations , 2018
2018
Earlier work this paper cites.
D. Park, W. Chan, Y. Zhang, C. Chiu, B. Zoph, E. Cubuk, and Q. Le, “SpecAugment: A simple data augmentation method for automatic speech recognition,” in Proc. INTERSPEECH , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
M. Rao, A. Raju, P. Dheram, B. Bui, and A. Rastrow, “Speech to semantics: Improve asr and nlu jointly via all-neural interfaces,” in Proc. INTERSPEECH , 2020
2020
Cited alongside, same era.
A. Aghajanyan, J. Maillard, A. Shrivastava, K. Diedrick, M. Haeger, H. Li, Y. Mehdad, V. Stoyanov, A. Kumar, M. Lewis, and S. Gupta, “Conversational Semantic Parsing,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , 2020
2020
M. Radfar, A. Mouchtaris, S. Kunzmann, and A. Rastrow, “Fans: Fusing asr and nlu for on-device slu,” in Interspeech , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
J. Mahadeokar, Y. Shangguan, D. Le, G. Keren, H. Su, T. Le, C. Yeh, C. Fuegen, and M. L. Seltzer, “Alignment Restricted Streaming Recurrent Neural Network Transducer,” in Proc. SLT , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
S. Rongali, L. Soldaini, E. Monti, and W. Hamza, “Don’t Parse, Generate! A Sequence to Sequence Architecture for Task-Oriented Semantic Parsing,” in Proceedings of the Web Conference (WWW) , 2020
2020
Cited alongside, same era.
“Deliberation model based two-pass end-to-end speech recognition,” in Proc. ICASSP , K. Hu, R. Prabhavalkar, R. Pang, and T. Sainath, Eds., 2020
2020
Cited alongside, same era.
X. Chen, A. Ghoshal, Y. Mehdad, L. Zettlemoyer, and S. Gupta, “Low-Resource Domain Adaptation for Compositional Task-Oriented Semantic Parsing,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020
2020
Cited alongside, same era.
N. Potdar, A. R. Avila, C. Xing, D. Wang, Y. Cao, and X. Chen, “A streaming end-to-end framework for spoken language understanding,” in IJCAI , 2021
2021
Cited alongside, same era.
P. Tomasello, A. Shrivastava, D. Lazar, P.-C. Hsu, D. Le, A. Sagar, A. Elkahky, J. Copet, W.-N. Hsu, Y. Mordechay, R. Algayres, T. A. Nguyen, E. Dupoux, L. Zettlemoyer, and A. Mohamed, “STOP: A dataset for Spoken Task Oriented Semantic Parsing,” in CoRR
Cited in the paper.
D. Le, M. Jain, G. Keren, S. Kim, Y. Shi, J. Mahadeokar, J. Chan, Y. Shangguan, C. Fuegen, O. Kalinli, Y. Saraf, and M. L. Seltzer, “Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion,” in Proc. Interspeech , 2021, pp. 1772–1776
2021
Later among the works it cites.
A. Babu, A. Shrivastava, A. Aghajanyan, A. Aly, A. Fan, and M. Ghazvininej, “Non-Autoregressive Semantic Parsing for Compositional Task-Oriented Dialog,” in Proceedings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) , 2021
2021
Later among the works it cites.
A. Shrivastava, P. Chuang, A. Babu, S. Desai, A. Arora, A. Zotov, and A. Aly, “Span Pointer Networks for Non-Autoregressive Task-Oriented Semantic Parsing,” in Proceedings of the Findings of the Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2021
2021
Later among the works it cites.
Y. Shi, C. Wu, D. Wang, A. Xiao, J. Mahadeokar, X. Zhang, C. Liu, K. Li, Y. Shangguan, V. Nagaraja et al. , “Streaming transformer transducer based speech recognition using non-causal convolution,” Proc. ICASSP , 2022
2022
Closest in time.