Fetching the paper…
Reading the bibliography…
Most End-to-End (E2E) SLU networks leverage the pre-trained ASR networks but still lack the capability to understand the semantics of utterances, crucial for the SLU task.
“Splat: Speech-language joint pre-training for spoken language understanding,”
Y.-A. Chung, C. Zhu, and M. Zeng, · 1907
Earlier work this paper cites.
Spoken language understanding: Systems for extracting semantic information from speech
G. Tur and R. De Mori, · 2011
Earlier work this paper cites.
“Using recurrent neural networks for slot filling in spoken language understanding,”
G. Mesnil, Y. Dauphin, K. Yao, Y. Bengio, L. Deng, D. Hakkani-Tur, X. He, L. Heck, G. Tur, D. Yu, et al., · 2014
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
G. Hinton, O. Vinyals, and J. Dean, · 2015
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
W. Chan, N. Jaitly, Q. Le, and O. Vinyals, · 2016
Earlier work this paper cites.
“Categorical reparametrization with gumble-softmax,”
E. Jang, S. Gu, and B. Poole, · 2017
Earlier work this paper cites.
“Attention is all you need,”
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, · 2017
Earlier work this paper cites.
“Towards end-to-end spoken language understanding,”
D. Serdyuk, Y. Wang, C. Fuegen, A. Kumar, B. Liu, and Y. Bengio, · 2018
Earlier work this paper cites.
“Speech Model Pre-Training for End-to-End Spoken Language Understanding,”
L. Lugosch, M. Ravanelli, P. Ignoto, V. S. Tomar, and Y. Bengio, · 2019
Cited alongside, same era.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, · 2019
Cited alongside, same era.
“Roberta: A robustly optimized bert pretraining approach,”
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, · 2019
Cited alongside, same era.
“Speech to Text Adaptation: Towards an Efficient Cross-Modal Distillation,”
W. I. Cho, D. Kwak, J. W. Yoon, and N. S. Kim, · 2020
Cited alongside, same era.
“Tie your embeddings down: Cross-modal latent spaces for end-to-end spoken language understanding,”
B. Agrawal, M. Müller, M. Radfar, S. Choudhary, A. Mouchtaris, and S. Kunzmann, · 2020
“Semantic complexity in end-to-end spoken language understanding,”
J. P. McKenna, S. Choudhary, M. Saxon, G. P. Strimel, and A. Mouchtaris, · 2020
Later among the works it cites.
“Using speech synthesis to train end-to-end spoken language understanding models,”
L. Lugosch, B. H. Meyer, D. Nowrouzezahrai, and M. Ravanelli, · 2020
Later among the works it cites.
“Two-stage textual knowledge distillation to speech encoder for spoken language understanding,”
S. Kim, G. Kim, S. Shin, and S. Lee, · 2021
Closest in time.
“St-bert: Cross-modal language model pre-training for end-to-end spoken language understanding,”
M. Kim, G. Kim, S.-W. Lee, and J.-W. Ha, · 2021
Closest in time.
“Do as i mean, not as i say: Sequence loss training for spoken language understanding,”
M. Rao, P. Dheram, G. Tiwari, A. Raju, J. Droppo, A. Rastrow, and A. Stolcke, · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Leveraging unpaired text data for training end-to-end speech-to-intent systems,”
Y. Huang, H.-K. Kuo, S. Thomas, Z. Kons, K. Audhkhasi, B. Kingsbury, R. Hoory, and M. Picheny, · 2020
Cited alongside, same era.
“Towards semi-supervised semantics understanding from speech,”
C.-I. Lai, J. Cao, S. Bodapati, and S.-W. Li, · 2020
Cited alongside, same era.
“Slurp: A spoken language understanding resource package,”
E. Bastianelli, A. Vanzo, P. Swietojanski, and V. Rieser, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, · 2020
Cited alongside, same era.
“Semi-supervised spoken language understanding via self-supervised speech and language model pretraining,”
C.-I. Lai, Y.-S. Chuang, H.-Y. Lee, S.-W. Li, and J. Glass, · 2021
Closest in time.
“End-to-end spoken language understanding for generalized voice assistants,”
M. Saxon, S. Choudhary, J. P. McKenna, and A. Mouchtaris, · 2021
Closest in time.
“Top-down attention in end-to-end spoken language understanding,”
Y. Chen, W. Lu, A. Mottini, L. E. Li, J. Droppo, Z. Du, and B. Zeng, · 2021
Closest in time.
“Speech-language pre-training for end-to-end spoken language understanding,”
Y. Qian, X. Bianv, Y. Shi, N. Kanda, L. Shen, Z. Xiao, and M. Zeng, · 2021
Closest in time.