Fetching the paper…
Reading the bibliography…
End-to-end models are an attractive new approach to spoken language understanding (SLU) in which the meaning of an utterance is inferred directly from the raw audio without employing the standard pipeline composed of a separately trained speech recognizer and natural language understanding module.
“Noise and the reality gap: The use of simulation in evolutionary robotics,”
Nick Jakobi, Phil Husbands, and Inman Harvey, · 1995
Earlier work this paper cites.
“Towards end-to-end speech recognition with recurrent neural networks,”
Alex Graves and Navdeep Jaitly, · 2014
Earlier work this paper cites.
“Empirical evaluation of gated recurrent neural networks on sequence modeling,”
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“LibriSpeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2015
Earlier work this paper cites.
“Exploring ASR-free end-to-end modeling to improve spoken language understanding in a cloud-based dialog system,”
Yao Qian, Rutuja Ubale, Vikram Ramanarayanan, and Patrick Lange, · 2017
Earlier work this paper cites.
“CSTR VCTK corpus: English multi-speaker corpus for CSTR voice cloning toolkit,”
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al., · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“From Audio to Semantics: Approaches to end-to-end spoken language understanding,”
Parisa Haghani, Arun Narayanan, Michiel Bacchiani, Galen Chuang, Neeraj Gaur, Pedro Moreno, Rohit Prabhavalkar, Zhongdi Qu, and Austin Waters, · 2018
Cited alongside, same era.
“Towards end-to-end spoken language understanding,”
Dmitriy Serdyuk, Yongqiang Wang, Christian Fuegen, Anuj Kumar, Baiyang Liu, and Yoshua Bengio, · 2018
Cited alongside, same era.
“Spoken language understanding without speech recognition,”
Yuan-Ping Chen, Ryan Price, and Srinivas Bangalore, · 2018
Cited alongside, same era.
“Training neural speech recognition systems with synthetic speech augmentation,”
Jason Li, Ravi Gadde, Boris Ginsburg, and Vitaly Lavrukhin, · 2018
Cited alongside, same era.
“Investigating backtranslation in neural machine translation,”
Alberto Poncelas, Dimitar Shterionov, Andy Way, Gideon Maillette de Buy Wenniger, and Peyman Passban, · 2018
Cited alongside, same era.
“End-to-end spoken language understanding: Bootstrapping in low resource scenarios,”
Swapnil Bhosale, Imran Sheikh, Sri Harsha Dumpala, and Sunil Kumar Kopparapu, · 2019
Closest in time.
“Curriculum-based transfer learning for an effective end-to-end spoken language understanding and domain portability,”
Antoine Caubrière, Natalia Tomashenko, Antoine Laurent, Emmanuel Morin, Nathalie Camelin, and Yannick Estève, · 2019
Closest in time.
“Speech recognition with augmented synthesized speech,”
Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran, Ye Jia, Pedro Moreno, Yonghui Wu, and Zelin Wu, · 2019
Closest in time.
“Improving performance of end-to-end ASR on numeric sequences,”
Cal Peyser, Hao Zhang, Tara N. Sainath, and Zelin Wu, · 2019
Closest in time.
“Streaming end-to-end speech recognition for mobile devices,”
Yanzhang He, Tara N. Sainath, Rohit Prabhavalkar, Ian McGraw, Raziel Alvarez, Ding Zhao, David Rybach, Anjuli Kannan, Yonghui Wu, Ruoming Pang, et al., · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“VoiceLoop: Voice fitting and synthesis via a phonological loop,”
Yaniv Taigman, Lior Wolf, Adam Polyak, and Eliya Nachmani, · 2018
Cited alongside, same era.
“Recent advances in end-to-end spoken language understanding,”
Natalia Tomashenko, Antoine Caubriere, Yannick Esteve, Antoine Laurent, and Emmanuel Morin, · 2019
Cited alongside, same era.
“Speech model pre-training for end-to-end spoken language understanding,”
Loren Lugosch, Mirco Ravanelli, Patrick Ignoto, Vikrant Singh Tomar, and Yoshua Bengio, · 2019
Closest in time.
“Spoken language understanding on the edge,”
Alaa Saade, Alice Coucke, Alexandre Caulier, Joseph Dureau, Adrien Ball, Théodore Bluche, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, and Maël Primet, · 2019
Closest in time.