Fetching the paper…
Reading the bibliography…
We present a simple and effective self-supervised learning approach for speech recognition.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, H., Mohamed, A., and Auli, M · 2006
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Sequence transduction with recurrent neural networks
Graves, A · 2012
Earlier work this paper cites.
Japanese and Korean voice search
Schuster, M. and Nakajima, K · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Feature learning with raw-waveform cldnns for voice activity detection
Zazo Candil, R., Sainath, T. N., Simko, G., and Parada, C · 2016
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Neural discrete representation learning, 2018
van den Oord, A., Vinyals, O., and Kavukcuoglu, K · 2018
Earlier work this paper cites.
Effectiveness of self-supervised pre-training for speech recognition
Baevski, A., Auli, M., and Mohamed, A · 2019
Cited alongside, same era.
wav2vec: Unsupervised pre-training for speech recognition
Schneider, S., Baevski, A., Collobert, R., and Auli, M · 2019
Cited alongside, same era.
Lingvo: a modular and scalable framework for sequence-to-sequence modeling, 2019
Shen, J., Nguyen, P., Wu, Y., Chen, Z., and et al · 2019
Cited alongside, same era.
Unsupervised cross-lingual representation learning for speech recognition
Conneau, A., Baevski, A., Collobert, R., Mohamed, A., and Auli, M · 2020
Cited alongside, same era.
Conformer: Convolution-augmented transformer for speech recognition, 2020
Gulati, A., Qin, J., Chiu, C.-C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., Wu, Y., and Pang, R · 2020
Xls-r: Self-supervised cross-lingual speech representation learning at scale
Babu, A., Wang, C., Tjandra, A., Lakhotia, K., Xu, Q., Goyal, N., Singh, K., von Platen, P., Saraf, Y., Pino, J., et al · 2021
Later among the works it cites.
Joint unsupervised and supervised training for multilingual asr
Bai, J., Li, B., Zhang, Y., Bapna, A., Siddhartha, N., Sim, K. C., and Sainath, T. N · 2021
Later among the works it cites.
Beit: Bert pre-training of image transformers, 2021
Bao, H., Dong, L., and Wei, F · 2021
Later among the works it cites.
Chung, Y.-A., Zhang, Y., Han, W., Chiu, C.-C., Qin, J., Pang, R., and Wu, Y · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Libri-light: A benchmark for ASR with limited or no supervision
Kahn, J., Rivière, M., Zheng, W., Kharitonov, E., Xu, Q., Mazaré, P.-E., Karadayi, J., Liptchinsky, V., Collobert, R., Fuegen, C., Likhomanenko, T., Synnaeve, G., Joulin, A., Mohamed, A., and Dupoux, E · 2020
Cited alongside, same era.
Mls: A large-scale multilingual dataset for speech research
Pratap, V., Xu, Q., Sriram, A., Synnaeve, G., and Collobert, R · 2020
Cited alongside, same era.
A streaming on-device end-to-end model surpassing server-side conventional model quality and latency
Sainath, T. N., He, Y., Li, B., Narayanan, A., Pang, R., Bruguier, A., Chang, S.-y., Li, W., Alvarez, R., Chen, Z., and et al · 2020
Cited alongside, same era.
Pushing the limits of semi-supervised learning for automatic speech recognition
Zhang, Y., Qin, J., Park, D. S., Han, W., Chiu, C.-C., Pang, R., Le, Q. V., and Wu, Y · 2020
Cited alongside, same era.
vq-wav2vec: Self-supervised learning of discrete speech representations
Baevski, A., Schneider, S., and Auli, M
Cited in the paper.
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2021
Later among the works it cites.
HuBERT: Self-supervised speech representation learning by masked prediction of hidden units
Hsu, W.-N., Bolte, B., Tsai, Y.-H. H., Lakhotia, K., Salakhutdinov, R., and Mohamed, A · 2021
Later among the works it cites.
Scaling end-to-end models for large-scale multilingual asr
Li, B., Pang, R., Sainath, T. N., Gulati, A., Zhang, Y., Qin, J., Haghani, P., Huang, W. R., and Ma, M · 2021
Later among the works it cites.
FastEmit: Low-latency Streaming ASR with Sequence-level Emission Regularization
Yu, J., Chiu, C.-C., Li, B., et al · 2021
Later among the works it cites.
Zhang, Y., Daniel Park, S., Han, W., Qin, J., Gulati, A., Shor, J., Jansen, A., Xu, Y., Huang, Y., Wang, S., Zhou, Z., Li, B., Ma, M., Chan, W., Yu, J., Wang, Y., Cao, L., Sim, K. C., Ramabhadran, B., Sainath, T. N., Beaufays, F., Chen, Z., Le, Q. V., Chiu, C.-C., Pang, R., and Wu, Y · 2021
Later among the works it cites.