Fetching the paper…
Reading the bibliography…
Self-supervised learning (SSL) has led to great strides in speech processing.
“GigaSpeech: An Evolving, Multi-Domain ASR Corpus with 10,000 Hours of Transcribed Audio”
Guoguo Chen et al · 1965
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks”
Alex Graves et al · 2006
Earlier work this paper cites.
“ImageNet: A large-scale hierarchical image database”
Jia Deng et al · 2009
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books”
Vassil Panayotov et al · 2015
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“Montreal Forced Aligner: Trainable Text-Speech Alignment Using Kaldi”
Michael McAuliffe et al · 2017
Earlier work this paper cites.
“Hybrid CTC/Attention Architecture for End-to-End Speech Recognition”
Shinji Watanabe et al · 2017
Earlier work this paper cites.
“GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding”
Alex Wang et al · 2018
Earlier work this paper cites.
“Deep clustering for unsupervised learning of visual features”
Mathilde Caron et al · 2018
Earlier work this paper cites.
“ESPnet: End-to-End Speech Processing Toolkit”
Shinji Watanabe et al · 2018
Earlier work this paper cites.
“SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems”
Alex Wang et al · 2019
Earlier work this paper cites.
“wav2vec: Unsupervised Pre-Training for Speech Recognition”
Steffen Schneider et al · 2019
Earlier work this paper cites.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin et al · 2019
Earlier work this paper cites.
“Speech Model Pre-Training for End-to-End Spoken Language Understanding”
Loren Lugosch et al · 2019
Earlier work this paper cites.
“fairseq: A Fast, Extensible Toolkit for Sequence Modeling”
Myle Ott et al · 2019
Cited alongside, same era.
“wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations”
Alexei Baevski et al · 2020
Cited alongside, same era.
“Language Models are Few-Shot Learners”
Tom Brown et al · 2020
Cited alongside, same era.
“Array programming with NumPy”
Charles. Harris et al · 2020
Cited alongside, same era.
“Libri-Light: A Benchmark for ASR with Limited or No Supervision”
J. Kahn et al · 2020
Cited alongside, same era.
“SUPERB: Speech Processing Universal PERformance Benchmark”
Shu-wen Yang et al · 2021
Cited alongside, same era.
“FitHuBERT: Going Thinner and Deeper for Knowledge Distillation of Speech Self-Supervised Models”
Yeonghyeon Lee et al · 2022
Later among the works it cites.
“Deep versus Wide: An Analysis of Student Architectures for Task-Agnostic Knowledge Distillation of Self-Supervised Speech Models”
Takanori Ashihara et al · 2022
Later among the works it cites.
“PaLM: Scaling language modeling with pathways”
Aakanksha Chowdhery et al · 2022
Later among the works it cites.
“GPT-NeoX-20B: An Open-Source Autoregressive Language Model”
Sidney Black et al · 2022
Later among the works it cites.
“BLOOM: A 176b-parameter open-access multilingual language model”
Teven Scao et al · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“An Exploration of Self-Supervised Pretrained Representations for End-to-End Speech Recognition”
Xuankai Chang et al · 2021
Cited alongside, same era.
“HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units”
Wei-Ning Hsu et al · 2021
Cited alongside, same era.
“PARP: Prune, Adjust and Re-Prune for Self-Supervised Speech Recognition”
Cheng-I Lai et al · 2021
Cited alongside, same era.
“Beyond the imitation game: Quantifying and extrapolating the capabilities of language models”
Aarohi Srivastava et al · 2022
Cited alongside, same era.
“Self-Supervised Speech Representation Learning: A Review”
Abdelrahman Mohamed et al · 2022
Cited alongside, same era.
“WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing”
Sanyuan Chen et al · 2022
Cited alongside, same era.
“Torchaudio: Building blocks for audio and speech processing”
Yao-Yuan Yang et al · 2022
Later among the works it cites.
“Understanding the Role of Self Attention for Efficient Speech Recognition”
Kyuhong Shim, Jungwook Choi and Wonyong Sung · 2022
Later among the works it cites.
“Supervision-Guided Codebooks for Masked Prediction in Speech Pre-training”
Chengyi Wang et al · 2022
Later among the works it cites.
“Biased Self-supervised learning for ASR”
Florian Kreyssig et al · 2022
Later among the works it cites.
“Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and Understanding”, 2022
Yifan Peng et al · 2022
Later among the works it cites.
“LightHuBERT: Lightweight and Configurable Speech Representation Learning with Once-for-All Hidden-Unit BERT”
Rui Wang et al · 2022
Later among the works it cites.
“Structured Pruning of Self-Supervised Pre-trained Models for Speech Recognition and Understanding”
Yifan Peng et al · 2023
Closest in time.
“E-Branchformer: Branchformer with Enhanced Merging for Speech Recognition”
Kwangyoun Kim et al · 2023
Closest in time.
“ASBERT: ASR-Specific Self-Supervised Learning with Self-Training”
Hyung Kim et al · 2023
Closest in time.