Fetching the paper…
Reading the bibliography…
Self-supervised learning (SSL) has proven vital for advancing research in natural language processing (NLP) and computer vision (CV).
C. Busso et al. , “Iemocap: Interactive emotional dyadic motion capture database,” Language resources and evaluation , vol. 42, no. 4, pp. 335–359, 2008
2008
Earlier work this paper cites.
T. Giorgino, “Computing and visualizing dynamic time warping alignments in R: The dtw package,” Journal of Statistical Software , vol. 31, no. 7, pp. 1–24, 2009
2009
Earlier work this paper cites.
K. Heafield, “Kenlm: Faster and smaller language model queries,” in Proceedings of the sixth workshop on statistical machine translation , 2011, pp. 187–197
2011
Earlier work this paper cites.
L. J. Rodriguez-Fuentes, A. Varona, M. Penagarikano, G. Bordel, and M. Diez, “Gtts-ehu systems for quesst at mediaeval 2014,” in MediaEval , 2014
2014
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An ASR corpus based on public domain audio books,” in ICASSP , 2015, pp. 5206–5210
2015
Earlier work this paper cites.
X. Anguera, L. Rodriguez-Fuentes, A. Buzo, F. Metze, I. Szöke, and M. Penagarikano, “Quesst2014: Evaluating query-by-example speech search in a zero-resource setting with real-life queries,” in ICASSP , 2015, pp. 5833–5837
2015
Earlier work this paper cites.
P. Warden, “Speech commands: A public dataset for single-word speech recognition.” Dataset available online , 2017
2017
Earlier work this paper cites.
M. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in NAACL , 2018, pp. 2227–2237
2018
Earlier work this paper cites.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “GLUE: A multi-task benchmark and analysis platform for natural language understanding,” in EMNLP , 2018, pp. 353–355
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in ICASSP , 2018, pp. 5329–5333
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in NAACL , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
P. Goyal, D. Mahajan, A. Gupta, and I. Misra, “Scaling and benchmarking self-supervised visual representation learning,” in ICCV , 2019, pp. 6391–6400
2019
Earlier work this paper cites.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An Unsupervised Autoregressive Model for Speech Representation Learning,” in Interspeech , 2019, pp. 146–150
2019
Cited alongside, same era.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition.” in Interspeech , 2019
2019
Cited alongside, same era.
S. Pascual, M. Ravanelli, J. Serrà, A. Bonafonte, and Y. Bengio, “Learning problem-agnostic speech representations from multiple self-supervised tasks,” in Interspeech , 2019, pp. 161–165
2019
Cited alongside, same era.
L. Lugosch, M. Ravanelli, P. Ignoto, V. S. Tomar, and Y. Bengio, “Speech model pre-training for end-to-end spoken language understanding,” in Interspeech , 2019, pp. 814–818
2019
Cited alongside, same era.
N. Tomashenko et al. , “Recent advances in end-to-end spoken language understanding,” in International Conference on Statistical Language and Speech Processing , 2019, pp. 44–55
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” in NeurIPS , 2020
2020
Later among the works it cites.
M. Ravanelli, J. Zhong, S. Pascual, P. Swietojanski, J. Monteiro, J. Trmal, and Y. Bengio, “Multi-task self-supervised learning for robust speech recognition,” in ICASSP , 2020, pp. 6989–6993
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
V. Pratap et al. , “Wav2letter++: A fast open-source speech recognition system,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2019, pp. 6460–6464
2019
Cited alongside, same era.
Y. Fujita, N. Kanda, S. Horiguchi, K. Nagamatsu, and S. Watanabe, “End-to-end neural speaker diarization with permutation-free objectives,” in Interspeech , 2019, pp. 4300–4304
2019
Cited alongside, same era.
S. Gururangan, A. Marasović, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith, “Don’t stop pretraining: Adapt language models to domains and tasks,” in ACL , 2020, pp. 8342–8360
2020
Cited alongside, same era.
A. Newell and J. Deng, “How useful is self-supervised pretraining for visual tasks?” in CVPR , 2020
2020
Cited alongside, same era.
A. T. Liu, S.-w. Yang, P.-H. Chi, P.-c. Hsu, and H.-y. Lee, “Mockingjay: Unsupervised speech representation learning with deep bidirectional transformer encoders,” ICASSP , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
T. A. Nguyen et al. , “The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,” in NeurIPS Workshop on Self-Supervised Learning for Speech and Audio Processing , 2020
2020
Later among the works it cites.
J. Shor, A. Jansen, R. Maor, O. Lang, O. Tuval, F. de Chaumont Quitry, M. Tagliasacchi, I. Shavitt, D. Emanuel, and Y. Haviv, “Towards learning a universal non-semantic representation of speech,” in Interspeech , 2020, pp. 140–144
2020
Later among the works it cites.
A. Nagrani, J. S. Chung, W. Xie, and A. Zisserman, “Voxceleb: Large-scale speaker verification in the wild,” Computer Speech & Language , vol. 60, p. 101027, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
Y.-A. Chung, H. Tang, and J. Glass, “Vector-quantized autoregressive predictive coding,” in Interspeech , 2020, pp. 3760–3764
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Rivière, A. Joulin, P.-E. Mazaré, and E. Dupoux, “Unsupervised pretraining transfers well across languages,” in ICASSP , 2020, pp. 7414–7418
2020
Later among the works it cites.
C.-I. Lai, Y.-S. Chuang, H.-Y. Lee, S.-W. Li, and J. Glass, “Semi-supervised spoken language understanding via self-supervised speech and language model pretraining,” in ICASSP , 2021
2021
Closest in time.