Fetching the paper…
Reading the bibliography…
Collecting sufficient labeled data for spoken language understanding (SLU) is expensive and time-consuming.
“A post-processing system to yield reduced word error rates: Recognizer output voting error reduction (rover),”
Jonathan G Fiscus, · 1997
Earlier work this paper cites.
“Dialogue act modeling for automatic tagging and recognition of conversational speech,”
Andreas Stolcke, Klaus Ries, Noah Coccaro, Elizabeth Shriberg, Rebecca Bates, et al., · 2000
Earlier work this paper cites.
“Bleu: a method for automatic evaluation of machine translation,”
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu, · 2002
Earlier work this paper cites.
“IEMOCAP: Interactive emotional dyadic motion capture database,”
C. Busso, M. Bulut, C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. Changa, S. Lee, and S. Narayanan, · 2008
Earlier work this paper cites.
“Dropout: A simple way to prevent neural networks from overfitting,”
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“XSEDE: Accelerating scientific discovery,”
J. Towns, T. Cockerill, M. Dahan, I. Foster, K. Gaither, A. Grimshaw, V. Hazlewood, S. Lathrop, D. Lifka, G. D. Peterson, R. Roskies, J. R. Scott, and N. Wilkins-Diehr, · 2014
Earlier work this paper cites.
“Bridges: a uniquely flexible HPC resource for new communities and data analytics,”
Nicholas A Nystrom, Michael J Levine, Ralph Z Roskies, and J Ray Scott, · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
Earlier work this paper cites.
“Joint CTC/attention decoding for end-to-end speech recognition,”
Takaaki Hori, Shinji Watanabe, and John R Hershey, · 2017
Earlier work this paper cites.
Alice Coucke, Alaa Saade, Adrien Ball, Théodore Bluche, Alexandre Caulier, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, Thibaut Lavril, et al., · 2018
Earlier work this paper cites.
“Spoken language understanding on the edge,”
Alaa Saade, Alice Coucke, Alexandre Caulier, Joseph Dureau, Adrien Ball, Théodore Bluche, David Leroy, Clément Doumouro, Thibault Gisselbrecht, Francesco Caltagirone, et al., · 2018
Earlier work this paper cites.
“Lexico-acoustic neural-based models for dialog act classification,”
Daniel Ortega and Ngoc Thang Vu, · 2018
Earlier work this paper cites.
“From audio to semantics: Approaches to end-to-end spoken language understanding,”
Parisa Haghani, Arun Narayanan, Michiel Bacchiani, Galen Chuang, Neeraj Gaur, Pedro Moreno, Rohit Prabhavalkar, Zhongdi Qu, and Austin Waters, · 2018
Earlier work this paper cites.
“Towards end-to-end spoken language understanding,”
Dmitriy Serdyuk, Yongqiang Wang, Christian Fuegen, Anuj Kumar, Baiyang Liu, and Yoshua Bengio, · 2018
Earlier work this paper cites.
“ESPnet: End-to-end speech processing toolkit,”
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, Adithya Renduchintala, and Tsubasa Ochiai, · 2018
Earlier work this paper cites.
“Gunrock: A social bot for complex and engaging long conversations,”
Dian Yu, Michelle Cohn, Yi Mang Yang, Chun-Yen Chen, Weiming Wen, et al., · 2019
Earlier work this paper cites.
“Speech model pre-training for end-to-end spoken language understanding,”
Loren Lugosch, Mirco Ravanelli, Patrick Ignoto, Vikrant Singh Tomar, and Yoshua Bengio, · 2019
Cited alongside, same era.
“Speech model pre-training for end-to-end spoken language understanding,”
Loren Lugosch, Mirco Ravanelli, Patrick Ignoto, Vikrant Singh Tomar, and Yoshua Bengio, · 2019
Cited alongside, same era.
“BERT: pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2019
Cited alongside, same era.
“Two-pass end-to-end speech recognition,”
Tara N. Sainath, Ruoming Pang, David Rybach, Yanzhang He, Rohit Prabhavalkar, Wei Li, Mirkó Visontai, Qiao Liang, Trevor Strohman, Yonghui Wu, Ian McGraw, and Chung-Cheng Chiu, · 2019
Cited alongside, same era.
“PyTorch: An imperative style, high-performance deep learning library,”
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al., · 2019
“Semantic complexity in end-to-end spoken language understanding,”
Joseph P. McKenna, Samridhi Choudhary, Michael Saxon, Grant P. Strimel, and Athanasios Mouchtaris, · 2020
Later among the works it cites.
“Earnings-21: A practical benchmark for ASR in the wild,”
Miguel Del Rio, Natalie Delworth, Ryan Westerman, Michelle Huang, Nishchal Bhandari, Joseph Palakapilly, Quinten McNamara, Joshua Dong, Piotr Żelasko, and Miguel Jetté, · 2021
Later among the works it cites.
“SUPERB: Speech Processing Universal PERformance Benchmark,”
Shu wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, et al., · 2021
Later among the works it cites.
“SLUE: New benchmark tasks for spoken language understanding evaluation on natural speech,”
Suwon Shon, Ankita Pasad, Felix Wu, Pablo Brusco, Yoav Artzi, Karen Livescu, and Kyu J Han, · 2021
Later among the works it cites.
“DeBERTa: decoding-enhanced BERT with disentangled attention,”
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“fairseq: A fast, extensible toolkit for sequence modeling,”
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli, · 2019
Cited alongside, same era.
“SpecAugment: A simple data augmentation method for automatic speech recognition,”
Daniel S. Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D. Cubuk, and Quoc V. Le, · 2019
Cited alongside, same era.
“When does label smoothing help?,”
Rafael Müller, Simon Kornblith, and Geoffrey E Hinton, · 2019
Cited alongside, same era.
“SLURP: A spoken language understanding resource package,”
Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski, and Verena Rieser, · 2020
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Tie your embeddings down: Cross-modal latent spaces for end-to-end spoken language understanding,”
Bhuvan Agrawal, Markus Müller, Martin Radfar, Samridhi Choudhary, Athanasios Mouchtaris, and Siegfried Kunzmann, · 2020
Cited alongside, same era.
“SPLAT: Speech-language joint pre-training for spoken language understanding,”
Yu-An Chung, Chenguang Zhu, and Michael Zeng, · 2020
Cited alongside, same era.
“Semi-supervised spoken language understanding via self-supervised speech and language model pretraining,”
Cheng-I Lai, Yung-Sung Chuang, Hung-Yi Lee, Shang-Wen Li, and James Glass, · 2021
Later among the works it cites.
“On the use of external data for spoken named entity recognition,”
Ankita Pasad, Felix Wu, Suwon Shon, Karen Livescu, and Kyu J Han, · 2021
Later among the works it cites.
“SpeechBERT: An audio-and-text jointly learned language model for end-to-end spoken question answering,”
Yung-Sung Chuang, Chi-Liang Liu, Hung yi Lee, and Lin shan Lee, · 2021
Later among the works it cites.
“TERA: Self-supervised learning of transformer encoder representation for speech,”
Andy T Liu, Shang-Wen Li, and Hung-yi Lee, · 2021
Later among the works it cites.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Later among the works it cites.
“WavLM: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2021
Later among the works it cites.
“GigaSpeech: An evolving, multi-domain ASR corpus with 10,000 hours of transcribed audio,”
Guoguo Chen, Shuzhou Chai, Guanbo Wang, Jiayu Du, Wei-Qiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, et al., · 2021
Later among the works it cites.
“SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition,”
Patrick K. O’Neill, Vitaly Lavrukhin, Somshubra Majumdar, Vahid Noroozi, Yuekai Zhang, et al., · 2021
Later among the works it cites.
“An exploration of self-supervised pretrained representations for end-to-end speech recognition,”
Xuankai Chang, Takashi Maekaku, Pengcheng Guo, Jing Shi, Yen-Ju Lu, Aswin Shanmugam Subramanian, Tianzi Wang, Shu-wen Yang, Yu Tsao, Hung-yi Lee, et al., · 2021
Later among the works it cites.
Changhan Wang, Morgane Riviere, Ann Lee, Anne Wu, Chaitanya Talnikar, Daniel Haziza, Mary Williamson, Juan Pino, and Emmanuel Dupoux, · 2021
Later among the works it cites.
“Rethinking end-to-end evaluation of decomposable tasks: A case study on spoken language understanding,”
Siddhant Arora, Alissa Ostapenko, Vijay Viswanathan, Siddharth Dalmia, Florian Metze, Shinji Watanabe, and Alan W. Black, · 2021
Later among the works it cites.
“ESPnet-SLU: Advancing spoken language understanding through ESPnet,”
Siddhant Arora, Siddharth Dalmia, Pavel Denisov, Xuankai Chang, Yushi Ueda, Yifan Peng, Yuekai Zhang, Sujay Kumar, Karthik Ganesan, Brian Yan, et al., · 2022
Closest in time.
“Two-pass low latency end-to-end spoken language understanding,”
Siddhant Arora, Siddharth Dalmia, Xuankai Chang, Brian Yan, Alan Black, and Shinji Watanabe, · 2022
Closest in time.