Fetching the paper…
Reading the bibliography…
Compressing self-supervised models has become increasingly necessary, as self-supervised models become larger.
“Multiframe deep neural networks for acoustic modeling,”
Vincent Vanhoucke, Matthieu Devin, and Georg Heigold, · 2013
Earlier work this paper cites.
“Word embeddings for speech recognition,”
Samy Bengio and Georg Heigold, · 2014
Earlier work this paper cites.
“Distilling the knowledge in a neural network,”
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean, · 2015
Earlier work this paper cites.
“EESEN: End-to-end speech recognition using deep RNN models and WFST-based decoding,”
Yajie Miao, Mohammad Gowayyed, and Florian Metze, · 2015
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,”
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals, · 2016
Earlier work this paper cites.
“End-to-end attention-based large vocabulary speech recognition,”
Dzmitry Bahdanau, Jan Chorowski, Dmitriy Serdyuk, Philemon Brakel, and Yoshua Bengio, · 2016
Earlier work this paper cites.
“Deep convolutional acoustic word embeddings using word-pair side information,”
Herman Kamper, Weiran Wang, and Karen Livescu, · 2016
Earlier work this paper cites.
“Neural machine translation of rare words with subword units,”
Rico Sennrich, Barry Haddow, and Alexandra Birch, · 2016
Earlier work this paper cites.
“Neural architecture search with reinforcement learning,”
Barret Zoph and Quoc V. Le, · 2017
Earlier work this paper cites.
“Joint CTC-attention based end-to-end speech recognition using multi-task learning,”
Suyoun Kim, Takaaki Hori, and Shinji Watanabe, · 2017
Earlier work this paper cites.
“Very deep convolutional networks for end-to-end speech recognition,”
Yu Zhang, William Chan, and Navdeep Jaitly, · 2017
Earlier work this paper cites.
“Advances in Joint CTC-Attention Based End-to-End Speech Recognition with a Deep CNN Encoder and RNN-LM,”
Takaaki Hori, Shinji Watanabe, Yu Zhang, and William Chan, · 2017
Earlier work this paper cites.
“Montreal forced aligner: Trainable text-speech alignment using Kaldi,”
Michael McAuliffe, Michaela Socolof, Sarah Mihuc, Michael Wagner, and Morgan Sonderegger, · 2017
Earlier work this paper cites.
“A time-restricted self-attention layer for ASR,”
Daniel Povey, Hossein Hadian, Pegah Ghahremani, Ke Li, and Sanjeev Khudanpur, · 2018
Cited alongside, same era.
“SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing,”
Taku Kudo and John Richardson, · 2018
Cited alongside, same era.
“The lottery ticket hypothesis: Finding sparse, trainable neural networks,”
Jonathan Frankle and Michael Carbin, · 2019
Cited alongside, same era.
“fairseq: A fast, extensible toolkit for sequence modeling,”
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli, · 2019
Cited alongside, same era.
“Audio-linguistic embeddings for spoken sentences,”
Albert Haque, Michele Guo, Prateek Verma, and Li Fei-Fei, · 2019
Cited alongside, same era.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
“Segmental contrastive predictive coding for unsupervised word segmentation,”
Saurabhchand Bhati, Jesús Villalba, Piotr Żelasko, Laureano Moro-Velázquez, and Najim Dehak, · 2021
Later among the works it cites.
“Variable-rate discrete representation learning,”
Sander Dieleman, Charlie Nash, Jesse Engel, and Karen Simonyan, · 2021
Later among the works it cites.
“On generative spoken language modeling from raw audio,”
Kushal Lakhotia, Eugene Kharitonov, Wei-Ning Hsu, Yossi Adi, Adam Polyak, Benjamin Bolte, Tu-Anh Nguyen, Jade Copet, Alexei Baevski, Abdelrahman Mohamed, et al., · 2021
Later among the works it cites.
“Towards unsupervised phone and word segmentation using self-supervised vector-quantized neural networks,”
Herman Kamper and Benjamin van Niekerk, · 2021
Later among the works it cites.
“Unsupervised speech recognition,”
Alexei Baevski, Wei-Ning Hsu, Alexis Conneau, and Michael Auli, · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Cited alongside, same era.
“DeCoAR 2.0: Deep contextualized acoustic representations with vector quantization,”
Shaoshi Ling and Yuzong Liu, · 2020
Cited alongside, same era.
“Longformer: The long-document transformer,”
Iz Beltagy, Matthew E. Peters, and Arman Cohan, · 2020
Cited alongside, same era.
“Big bird: Transformers for longer sequences,”
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al., · 2020
Cited alongside, same era.
“CIF: Continuous integrate-and-fire for end-to-end speech recognition,”
Linhao Dong and Bo Xu, · 2020
Cited alongside, same era.
“HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,”
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed, · 2021
Cited alongside, same era.
“SUPERB: Speech Processing Universal PERformance Benchmark,”
Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko-tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, and Hung-yi Lee, · 2021
Cited alongside, same era.
“WavLM: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2022
Closest in time.
“data2vec: A general framework for self-supervised learning in speech, vision and language,”
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli, · 2022
Closest in time.
“Efficient transformers: A survey,”
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler, · 2022
Closest in time.
“Performance-efficiency trade-offs in unsupervised pre-training for speech recognition,”
Felix Wu, Kwangyoun Kim, Jing Pan, Kyu J. Han, Kilian Q. Weinberger, and Yoav Artzi, · 2022
Closest in time.
“On-demand compute reduction with stochastic wav2vec 2.0,”
Apoorv Vyas, Wei-Ning Hsu, Michael Auli, and Alexei Baevski, · 2022
Closest in time.
“FitHuBERT: Going Thinner and Deeper for Knowledge Distillation of Speech Self-Supervised Models,”
Yeonghyeon Lee, Kangwook Jang, Jahyun Goo, Youngmoon Jung, and Hoi Rin Kim, · 2022
Closest in time.
“DistilHuBERT: Speech representation learning by layer-wise distillation of hidden-unit BERT,”
Heng-Jui Chang, Shu-wen Yang, and Hung-yi Lee, · 2022
Closest in time.
“Towards end-to-end unsupervised speech recognition,”
Alexander H. Liu, Wei-Ning Hsu, Michael Auli, and Alexei Baevski, · 2022
Closest in time.