Fetching the paper…
Reading the bibliography…
We introduce a self-supervised speech pre-training method called TERA, which stands for Transformer Encoder Representations from Alteration.
1905
Earlier work this paper cites.
1907
Earlier work this paper cites.
K.-F. Lee and H.-W. Hon, “Speaker-independent phone recognition using hidden markov models,” IEEE Transactions on Acoustics, Speech, and Signal Processing , vol. 37, no. 11, pp. 1641–1648, 1989
1989
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, “DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1,” NASA STI/Recon Technical Report N, p. 27403, Feb. 1993
1993
Earlier work this paper cites.
M. J. Gales, “Maximum likelihood linear transformations for hmm-based speech recognition,” Computer speech & language , vol. 12, no. 2, pp. 75–98, 1998
1998
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning , 2006, pp. 369–376
2006
Earlier work this paper cites.
S. Li, L. Li, Q. Hong, and L. Liu, “Improving Transformer-Based Speech Recognition with Unsupervised Pre-Training and Multi-Task Semantic Knowledge Learning,” in Interspeech 2020 , 2020, pp. 5006–5010. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-2007
2007
Earlier work this paper cites.
T. Mikolov, M. Karafiát, L. Burget, J. Černockỳ, and S. Khudanpur, “Recurrent neural network based language model,” in Eleventh annual conference of the international speech communication association , 2010
2010
Earlier work this paper cites.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The kaldi speech recognition toolkit,” in ASRU , 2011
2011
Earlier work this paper cites.
C. Lopes and F. Perdigao, “Phone recognition on the timit database,” Speech Technologies/Book , vol. 1, pp. 285–302, 2011
2011
Earlier work this paper cites.
2013
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in ICASSP , 2015
2015
Earlier work this paper cites.
L. Tóth, “Phone recognition with hierarchical convolutional deep maxout networks,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2015, no. 1, pp. 1–13, 2015
2015
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , ser. NIPS’17. Red Hook, NY, USA: Curran Associates Inc., 2017, p. 6000–6010
2017
Earlier work this paper cites.
P. Warden, “Speech commands: A public dataset for single-word speech recognition.” Dataset available online , 2017. [Online]. Available: http://download.tensorflow.org/data/speech_commands_v0.01.tar.gz
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers) , 2018
2018
Earlier work this paper cites.
N. Zeghidour, N. Usunier, I. Kokkinos, T. Schaiz, G. Synnaeve, and E. Dupoux, “Learning filterbanks from raw speech for phone recognition,” ICASSP 2018 , Apr 2018
2018
Earlier work this paper cites.
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio, “Light gated recurrent units for speech recognition,” IEEE Transactions on Emerging Topics in Computational Intelligence , vol. 2, no. 2, p. 92–102, Apr 2018
2018
Earlier work this paper cites.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition,” Interspeech , 2019
2019
Earlier work this paper cites.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An Unsupervised Autoregressive Model for Speech Representation Learning,” in Proc. Interspeech 2019 , 2019, pp. 146–150. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2019-1473
2019
Cited alongside, same era.
J. Chorowski, R. J. Weiss, S. Bengio, and A. van den Oord, “Unsupervised speech representation learning using wavenet autoencoders,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 27, no. 12, p. 2041–2053, Dec 2019
2019
Cited alongside, same era.
A. T. Liu, P.-c. Hsu, and H.-Y. Lee, “Unsupervised end-to-end learning of discrete linguistic units for voice conversion,” Interspeech , Sep 2019
2019
Cited alongside, same era.
F. de Chaumont Quitry, M. Tagliasacchi, and D. Roblek, “Learning audio representations via phase prediction,” 2019
2019
Cited alongside, same era.
Y.-A. Chung and J. Glass, “Improved speech representations with multi-target autoregressive predictive coding,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Online: Association for Computational Linguistics, Jul. 2020, pp. 2353–2358. [Online]. Available: https://www.aclweb.org/anthology/2020.acl-main.213
2020
Closest in time.
Y.-A. Chung, H. Tang, and J. Glass, “Vector-quantized autoregressive predictive coding,” Interspeech 2020 , pp. 3760–3764, 2020
2020
Closest in time.
S. Ling, Y. Liu, J. Salazar, and K. Kirchhoff, “Deep contextualized acoustic representations for semi-supervised speech recognition,” in ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2020, pp. 6429–6433
2020
Closest in time.
M. Tagliasacchi, B. Gfeller, F. d. C. Quitry, and D. Roblek, “Pre-training audio representations with self-supervision,” IEEE Signal Processing Letters , vol. 27, pp. 600–604, 2020
2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Pascual, M. Ravanelli, J. Serrà, A. Bonafonte, and Y. Bengio, “Learning problem-agnostic speech representations from multiple self-supervised tasks,” Interspeech 2019 , Sep 2019
2019
Cited alongside, same era.
C. Sun, X. Qiu, Y. Xu, and X. Huang, “How to fine-tune bert for text classification?” Chinese Computational Linguistics , p. 194–206, 2019
2019
Cited alongside, same era.
A. Chronopoulou, C. Baziotis, and A. Potamianos, “An embarrassingly simple approach for transfer learning from pretrained language models,” Proceedings of the 2019 Conference of the North , 2019
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 4171–4186. [Online]. Available: https://www.aclweb.org/anthology/N19-1423
2019
Cited alongside, same era.
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “Specaugment: A simple data augmentation method for automatic speech recognition,” Interspeech 2019 , Sep 2019
2019
Cited alongside, same era.
N.-Q. Pham, T.-S. Nguyen, J. Niehues, M. Müller, and A. Waibel, “Very Deep Self-Attention Networks for End-to-End Speech Recognition,” in Interspeech 2019 , 2019, pp. 66–70. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2019-2702
2019
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in International Conference on Learning Representations , 2019. [Online]. Available: https://openreview.net/forum?id=Bkg6RiCqY7
2019
Cited alongside, same era.
Closest in time.
S. Khurana, A. Laurent, W.-N. Hsu, J. Chorowski, A. Łańcucki, R. Marxer, and J. Glass, “A Convolutional Deep Markov Model for Unsupervised Speech Representation Learning,” in Interspeech 2020 , Shanghai, China, Oct. 2020. [Online]. Available: https://hal.archives-ouvertes.fr/hal-02912029
2020
Closest in time.
A. T. Liu, S.-w. Yang, P.-H. Chi, P.-c. Hsu, and H.-y. Lee, “Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,” ICASSP 2020 , May 2020
2020
Closest in time.
P.-H. Chi, P.-H. Chung, T.-H. Wu, C.-C. Hsieh, S.-W. Li, and H. yi Lee, “Audio albert: A lite bert for self-supervised learning of audio representation,” in SLT 2020 , 2020
2020
Closest in time.
W. Wang, Q. Tang, and K. Livescu, “Unsupervised pre-training of bidirectional speech encoders via masked reconstruction,” ICASSP 2020 , May 2020
2020
Closest in time.
X. Song, G. Wang, Y. Huang, Z. Wu, D. Su, and H. Meng, “Speech-XLNet: Unsupervised Acoustic Model Pretraining for Self-Attention Networks,” in Interspeech 2020 , 2020, pp. 3765–3769. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-1511
2020
Closest in time.
L. Liu and Y. Huang, “Masked pre-trained encoder base on joint ctc-transformer,” 2020
2020
Closest in time.
A. H. Liu, Y.-A. Chung, and J. Glass, “Non-autoregressive predictive coding for learning speech representations from local dependencies,” 2020
2020
Closest in time.
A. T. Liu and Y. Shu-wen, “The S3PRL toolkit: Self-supervised speech pre-training and representation learning,” 2020. [Online]. Available: https://github.com/s3prl/s3prl
2020
Closest in time.
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut, “Albert: A lite bert for self-supervised learning of language representations,” in International Conference on Learning Representations , 2020. [Online]. Available: https://openreview.net/forum?id=H1eA7AEtvS
2020
Closest in time.
S. wen Yang, A. T. Liu, and H. yi Lee, “Understanding Self-Attention of Self-Supervised Audio Transformers,” in Proc. Interspeech 2020 , 2020, pp. 3785–3789. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-2231
2020
Closest in time.
P. Wang, L. Wei, Y. Cao, J. Xie, and Z. Nie, “Large-scale unsupervised pre-training for end-to-end spoken language understanding,” in ICASSP , 2020
2020
Closest in time.
S. Ling, J. Salazar, Y. Liu, and K. Kirchhoff, “BERTphone: Phonetically-aware Encoder Representations for Utterance-level Speaker and Language Recognition,” in Odyssey 2020 The Speaker and Language Recognition Workshop , 2020
2020
Closest in time.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P. E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen, T. Likhomanenko, G. Synnaeve, A. Joulin, A. Mohamed, and E. Dupoux, “Libri-light: A benchmark for asr with limited or no supervision,” in ICASSP 2020 , 2020, pp. 7669–7673, https://github.com/facebookresearch/libri-light
2020
Closest in time.
D. Jiang, W. Li, R. Zhang, M. Cao, N. Luo, Y. Han, W. Zou, and X. Li, “A further study of unsupervised pre-training for transformer based speech recognition,” in Submitted to International Conference on Learning Representations , 2021, under review. [Online]. Available: https://openreview.net/forum?id=hrpSB_rzQTU
2021
Closest in time.
H. Wu, A. T. Liu, and H. yi Lee, “Defense for Black-Box Attacks on Anti-Spoofing Models by Self-Supervised Learning,” in Proc. Interspeech 2020 , 2020, pp. 3780–3784. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-2026
2026
Closest in time.