Fetching the paper…
Reading the bibliography…
Wav2vec 2.0 (W2V2) has shown impressive performance in automatic speech recognition (ASR).
Discriminative training for large vocabulary speech recognition
Daniel Povey, · 2005
Earlier work this paper cites.
“Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,”
Alex Graves, Santiago Fernández, Faustino Gomez, and Jürgen Schmidhuber, · 2006
Earlier work this paper cites.
“Sequence-discriminative training of deep neural networks.,”
Karel Veselỳ, Arnab Ghoshal, Lukás Burget, and Daniel Povey, · 2013
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Purely sequence-trained neural networks for asr based on lattice-free mmi.,”
Daniel Povey, Vijayaditya Peddinti, Daniel Galvez, Pegah Ghahremani, Vimal Manohar, Xingyu Na, Yiming Wang, and Sanjeev Khudanpur, · 2016
Earlier work this paper cites.
“End-to-end speech recognition using lattice-free mmi.,”
Hossein Hadian, Hossein Sameti, Daniel Povey, and Sanjeev Khudanpur, · 2018
Earlier work this paper cites.
“Semi-supervised training of acoustic models using lattice-free mmi,”
Vimal Manohar, Hossein Hadian, Daniel Povey, and Sanjeev Khudanpur, · 2018
Earlier work this paper cites.
“Tinybert: Distilling bert for natural language understanding,”
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu, · 2019
Earlier work this paper cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Earlier work this paper cites.
“On semi-supervised lf-mmi training of acoustic models with limited data,”
Imran Sheikh, Emmanuel Vincent, and Irina Illina, · 2020
Cited alongside, same era.
Yu-An Chung, Yu Zhang, Wei Han, Chung-Cheng Chiu, James Qin, Ruoming Pang, and Yonghui Wu, · 2021
Cited alongside, same era.
“Wavlm: Large-scale self-supervised pre-training for full stack speech processing,”
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al., · 2021
Cited alongside, same era.
“Improving hybrid ctc/attention end-to-end speech recognition with pretrained acoustic and language models,”
Keqi Deng, Songjun Cao, Yike Zhang, and Long Ma, · 2021
Cited alongside, same era.
“Efficient knowledge distillation for rnn-transducer models,”
Sankaran Panchapagesan, Daniel S Park, Chung-Cheng Chiu, Yuan Shangguan, Qiao Liang, and Alexander Gruenstein, · 2021
“Comparing ctc and lfmmi for out-of-domain adaptation of wav2vec 2.0 acoustic model,”
Apoorv Vyas, Srikanth Madikeri, and Hervé Bourlard, · 2021
Later among the works it cites.
“Developing real-time streaming transformer transducer for speech recognition on large-scale dataset,”
Xie Chen, Yu Wu, Zhenghao Wang, Shujie Liu, and Jinyu Li, · 2021
Later among the works it cites.
“Unsupervised speech recognition,”
Alexei Baevski, Wei-Ning Hsu, Alexis Conneau, and Michael Auli, · 2021
Later among the works it cites.
“Gigaspeech: An evolving, multi-domain asr corpus with 10,000 hours of transcribed audio,”
Guoguo Chen, Shuzhou Chai, Guanbo Wang, Jiayu Du, Wei-Qiang Zhang, Chao Weng, Dan Su, Daniel Povey, Jan Trmal, Junbo Zhang, et al., · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Shrinking bigfoot: Reducing wav2vec 2.0 footprint,”
Zilun Peng, Akshay Budhkar, Ilana Tuil, Jason Levy, Parinaz Sobhani, Raphael Cohen, and Jumana Nassour, · 2021
Cited alongside, same era.
“Knowledge distillation for neural transducers from large self-supervised pre-trained models,”
Xiaoyu Yang, Qiujia Li, and Philip C Woodland, · 2021
Cited alongside, same era.
“Improving streaming transformer based asr under a framework of self-supervised learning,”
Songjun Cao, Yueteng Kang, Yanzhe Fu, Xiaoshuo Xu, Sining Sun, Yike Zhang, and Long Ma, · 2021
Cited alongside, same era.
Nauman Dawalatabad, Tushar Vatsal, Ashutosh Gupta, Sungsoo Kim, Shatrughan Singh, Dhananjaya Gowda, and Chanwoo Kim, · 2022
Later among the works it cites.
Rui Wang, Qibing Bai, Junyi Ao, Long Zhou, Zhixiang Xiong, Zhihua Wei, Yu Zhang, Tom Ko, and Haizhou Li, · 2022
Later among the works it cites.
“Distilhubert: Speech representation learning by layer-wise distillation of hidden-unit bert,”
Heng-Jui Chang, Shu-wen Yang, and Hung-yi Lee, · 2022
Later among the works it cites.
“Fithubert: Going thinner and deeper for knowledge distillation of speech self-supervised learning,”
Yeonghyeon Lee, Kangwook Jang, Jahyun Goo, Youngmoon Jung, and Hoirin Kim, · 2022
Later among the works it cites.