Fetching the paper…
Reading the bibliography…
Spoken language understanding (SLU) tasks are usually solved by first transcribing an utterance with automatic speech recognition (ASR) and then feeding the output to a text-based model.
“Librispeech: an ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2015
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Learning problem-agnostic speech representations from multiple self-supervised tasks,”
Santiago Pascual, Mirco Ravanelli, Joan Serra, Antonio Bonafonte, and Yoshua Bengio, · 2019
Earlier work this paper cites.
“Speech model pre-training for end-to-end spoken language understanding,”
Loren Lugosch, Mirco Ravanelli, Patrick Ignoto, Vikrant Singh Tomar, and Yoshua Bengio, · 2019
Earlier work this paper cites.
“Electra: Pre-training text encoders as discriminators rather than generators,”
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning, · 2019
Earlier work this paper cites.
“Huggingface’s transformers: State-of-the-art natural language processing,”
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al., · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,”
Andy T Liu, Shu-wen Yang, Po-Han Chi, Po-chun Hsu, and Hung-yi Lee, · 2020
Cited alongside, same era.
“Libri-light: A benchmark for asr with limited or no supervision,”
Jacob Kahn, Morgane Rivière, Weiyi Zheng, Evgeny Kharitonov, Qiantong Xu, Pierre-Emmanuel Mazaré, Julien Karadayi, Vitaliy Liptchinsky, Ronan Collobert, Christian Fuegen, et al., · 2020
Cited alongside, same era.
“CoVoST 2 and massively multilingual speech-to-text translation,”
Changhan Wang, Anne Wu, and Juan Pino, · 2020
Cited alongside, same era.
“Common voice: A massively-multilingual speech corpus,”
Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M. Tyers, and Gregor Weber, · 2020
“Superb: Speech processing universal performance benchmark,”
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y Lin, Andy T Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, et al., · 2021
Closest in time.
“Unsupervised speech recognition,”
Alexei Baevski, Wei-Ning Hsu, Alexis Conneau, and Michael Auli, · 2021
Closest in time.
“Robust wav2vec 2.0: Analyzing domain shift in self-supervised pre-training,”
Wei-Ning Hsu, Anuroop Sriram, Alexei Baevski, Tatiana Likhomanenko, Qiantong Xu, Vineel Pratap, Jacob Kahn, Ann Lee, Ronan Collobert, Gabriel Synnaeve, et al., · 2021
Closest in time.
“On scaling contrastive representations for low-resource speech recognition,”
Lasse Borgholt, Tycho MS Tax, Jakob D Havtorn, Lars Maaløe, and Christian Igel, · 2021
Closest in time.
“Layer-wise analysis of a self-supervised speech representation model,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Semi-supervised spoken language understanding via self-supervised speech and language model pretraining,”
Cheng-I Lai, Yung-Sung Chuang, Hung-Yi Lee, Shang-Wen Li, and James Glass, · 2021
Cited alongside, same era.
Ankita Pasad, Ju-Chieh Chou, and Karen Livescu, · 2021
Closest in time.
“Speech-language pre-training for end-to-end spoken language understanding,”
Yao Qian, Ximo Bianv, Yu Shi, Naoyuki Kanda, Leo Shen, Zhen Xiao, and Michael Zeng, · 2021
Closest in time.