Fetching the paper…
Reading the bibliography…
Recent studies find existing self-supervised speech encoders contain primarily acoustic rather than semantic information.
“OntoNotes: the 90% solution”
Eduard Hovy et al · 2006
Earlier work this paper cites.
“Visualizing data using t-SNE.”
Laurens Van and Geoffrey Hinton · 2008
Earlier work this paper cites.
“Deep residual learning for image recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“Learning multiple visual domains with residual adapters”
Sylvestre-Alvise Rebuffi, Hakan Bilen and Andrea Vedaldi · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Earlier work this paper cites.
“What does bert look at? an analysis of bert’s attention”
Kevin Clark, Urvashi Khandelwal, Omer Levy and Christopher Manning · 2019
Earlier work this paper cites.
“Parameter-efficient transfer learning for NLP”
Neil Houlsby et al · 2019
Earlier work this paper cites.
Mike Lewis et al · 2019
Earlier work this paper cites.
“Roberta: A robustly optimized bert pretraining approach”
Yinhan Liu et al · 2019
Earlier work this paper cites.
“Speech model pre-training for end-to-end spoken language understanding”
Loren Lugosch et al · 2019
Earlier work this paper cites.
“wav2vec: Unsupervised pre-training for speech recognition”
Steffen Schneider, Alexei Baevski, Ronan Collobert and Michael Auli · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed and Michael Auli · 2020
Earlier work this paper cites.
“SLURP: A spoken language understanding resource package”
Emanuele Bastianelli, Andrea Vanzo, Pawel Swietojanski and Verena Rieser · 2020
Cited alongside, same era.
“Longformer: The long-document transformer”
Iz Beltagy, Matthew Peters and Arman Cohan · 2020
Cited alongside, same era.
“Splat: Speech-language joint pre-training for spoken language understanding”
Yu-An Chung, Chenguang Zhu and Michael Zeng · 2020
Cited alongside, same era.
“Exploring wav2vec 2.0 on speaker verification and language identification”
Zhiyun Fan, Meng Li, Shiyu Zhou and Bo Xu · 2020
Cited alongside, same era.
“Generative adversarial networks”
Ian Goodfellow et al · 2020
“Layer-wise analysis of a self-supervised speech representation model”
Ankita Pasad, Ju-Chieh Chou and Karen Livescu · 2021
Later among the works it cites.
“Speech-language pre-training for end-to-end spoken language understanding”
Yao Qian et al · 2021
Later among the works it cites.
“Do as i mean, not as i say: Sequence loss training for spoken language understanding”
Milind Rao et al · 2021
Later among the works it cites.
“Superb: Speech processing universal performance benchmark”
Shu-wen Yang et al · 2021
Later among the works it cites.
“Tie your embeddings down: Cross-modal latent spaces for end-to-end spoken language understanding”
Bhuvan Agrawal et al · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Multilingual speech translation with efficient finetuning of pretrained models”
Xian Li et al · 2020
Cited alongside, same era.
“Towards unsupervised speech recognition and synthesis with quantized speech representation learning”
Alexander Liu, Tao Tu, Hung-yi Lee and Lin-shan Lee · 2020
Cited alongside, same era.
“Unsupervised speech recognition”
Alexei Baevski, Wei-Ning Hsu, Alexis Conneau and Michael Auli · 2021
Cited alongside, same era.
“Speak or chat with me: End-to-end spoken language understanding system with flexible inputs”
Sujeong Cha et al · 2021
Cited alongside, same era.
“HuBERT: How much can a bad teacher benefit ASR pre-training?”
Wei-Ning Hsu et al · 2021
Cited alongside, same era.
“St-bert: Cross-modal language model pre-training for end-to-end spoken language understanding”
Minjeong Kim, Gyuwan Kim, Sang-Woo Lee and Jung-Woo Ha · 2021
Cited alongside, same era.
“Lightweight adapter tuning for multilingual speech translation”
Hang Le et al · 2021
Cited alongside, same era.
Tzu-hsun Feng et al · 2022
Closest in time.
“Deliberation Model for On-Device Spoken Language Understanding”
Duc Le et al · 2022
Closest in time.
“DUAL: Discrete Spoken Unit Adaptive Learning for Textless Spoken Question Answering.”
Guan-Ting Lin et al · 2022
Closest in time.
“Towards End-to-end Unsupervised Speech Recognition”
Alexander Liu, Wei-Ning Hsu, Michael Auli and Alexei Baevski · 2022
Closest in time.
“On Compressing Sequences for Self-Supervised Speech Models”
Yen Meng et al · 2022
Closest in time.
“Integration of pre-trained networks with continuous token interface for end-to-end spoken language understanding”
Seunghyun Seo, Donghyun Kwak and Bowon Lee · 2022
Closest in time.
“Slue: New benchmark tasks for spoken language understanding evaluation on natural speech”
Suwon Shon et al · 2022
Closest in time.