Fetching the paper…
Reading the bibliography…
Inspired by the humans' cognitive ability to generalise knowledge and skills, Self-Supervised Learning (SSL) targets at discovering general representations from large-scale data without requiring human annotations, which is an expensive and time consuming task.
A. Kolesnikov, X. Zhai, and L. Beyer, “Revisiting self-supervised visual representation learning,” in Proc. ICCV , Long Beach, CA, USA, 2019, pp. 1920–1929
1929
Earlier work this paper cites.
H. B. Barlow et al. , “Possible principles underlying the transformation of sensory messages,” Sensory communication , vol. 1, no. 01, 1961
1961
Earlier work this paper cites.
J. Piaget, “Part I: Cognitive development in children: Piaget development and learning,” Journal of research in science teaching , vol. 2, no. 3, pp. 176–186, 1964
1964
Earlier work this paper cites.
W. F. Brewer and G. V. Nakamura, “The nature and functions of schemas,” Center for the Study of Reading Technical Report , no. 325, p. 52 pages, 1984
1984
Earlier work this paper cites.
R. Baillargeon and J. DeVos, “Object permanence in young infants: Further evidence,” Child development , vol. 62, no. 6, pp. 1227–1246, 1991
1991
Earlier work this paper cites.
D. N. Perkins, G. Salomon et al. , “Transfer of learning,” International encyclopedia of education , vol. 2, pp. 6452–6457, 1992
1992
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, “DARPA TIMIT acoustic-phonetic continous speech corpus CD-ROM. NIST speech disc 1-1.1,” NASA STI/Recon Technical Report , vol. 93, p. 27403, 1993
1993
Earlier work this paper cites.
B. J. Wadsworth, Piagetś theory of cognitive and affective development: Foundations of constructivism , 1996
1996
Earlier work this paper cites.
W. Huitt and J. Hummel, “Piagetś theory of cognitive development,” Educational psychology interactive , vol. 3, no. 2, pp. 1–5, 2003
2003
Earlier work this paper cites.
S. Chopra, R. Hadsell, and Y. LeCun, “Learning a similarity metric discriminatively, with application to face verification,” in Proc. CVPR , vol. 1, San Diego, CA, USA, 2005, pp. 539–546
2005
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in Proc. ICML , Pittsburgh,PA,USA, 2006, pp. 369–376
2006
Earlier work this paper cites.
R. Raina, A. Battle, H. Lee, B. Packer, and A. Y. Ng, “Self-taught learning: Transfer learning from unlabeled data,” in Proc. ICML , Sydney, Australia, 2007, pp. 759–766
2007
Earlier work this paper cites.
M. Gutmann and A. Hyvärinen, “Noise-contrastive estimation: A new estimation principle for unnormalized statistical models,” in Proc. AISTATS , Sardinia, Italy, 2010, pp. 297–304
2010
Earlier work this paper cites.
H. Jegou, M. Douze, and C. Schmid, “Product quantization for nearest neighbor search,” IEEE transactions on pattern analysis and machine intelligence , vol. 33, no. 1, pp. 117–128, 2010
2010
Earlier work this paper cites.
P. Baldi, “Autoencoders, unsupervised learning, and deep architectures,” in Proceedings of ICML workshop on unsupervised and transfer learning , Bellevue, Washington, USA, 2012, pp. 37–49
2012
Earlier work this paper cites.
Y. Bengio, A. Courville, and P. Vincent, “Representation learning: A review and new perspectives,” IEEE transactions on pattern analysis and machine intelligence , vol. 35, no. 8, pp. 1798–1828, 2013
2013
Earlier work this paper cites.
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” in Proc. ICLR , Scottsdale, AZ, USA, 2013, p. 12 pages
2013
Earlier work this paper cites.
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in Proc. NeurIPS , Sierra Nevada, USA, 2013, pp. 3111–3119
2013
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” Nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Earlier work this paper cites.
C. Doersch, A. Gupta, and A. A. Efros, “Unsupervised visual representation learning by context prediction,” in Proc. ICCV , Santiago, Chile, 2015, pp. 1422–1430
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proc. ICML , Lille, France, 2015, pp. 448–456
2015
Earlier work this paper cites.
F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embedding for face recognition and clustering,” in Proc. CVPR , Boston, MA, USA, 2015, pp. 815–823
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: an ASR corpus based on public domain audio books,” in Proc. ICASSP , Brisbane, Australia, 2015, pp. 5206–5210
2015
Earlier work this paper cites.
G. W. Oesterdiekhoff, “Child and ancient man: How to define their commonalities and differences,” American Journal of Psychology , vol. 129, no. 3, pp. 295–312, 2016
2016
Earlier work this paper cites.
M. Noroozi and P. Favaro, “Unsupervised learning of visual representations by solving jigsaw puzzles,” in Proc. ECCV , Amsterdam, Netherlands, 2016, pp. 69–84
2016
Earlier work this paper cites.
K. Sohn, “Improved deep metric learning with multi-class n-pair loss objective,” in Proc. NeurIPS , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett, Eds., Barcelona, Spain, 2016, p. 9 pages
2016
Earlier work this paper cites.
I. Misra, C. L. Zitnick, and M. Hebert, “Shuffle and learn: Unsupervised learning using temporal order verification,” in Proc. ECCV , Amsterdam, Netherlands, 2016, pp. 527–544
2016
Earlier work this paper cites.
D. Harwath, A. Torralba, and J. R. Glass, “Unsupervised learning of spoken language with visual context,” p. 1866–1874, 2016
2016
Earlier work this paper cites.
E. Shelhamer, P. Mahmoudieh, M. Argus, and T. Darrell, “Loss is its own reward: Self-supervision for reinforcement learning,” p. 4 pages, 2017
2017
Earlier work this paper cites.
G. Larsson, M. Maire, and G. Shakhnarovich, “Colorization as a proxy task for visual understanding,” Proc. CVPR , pp. 840–849, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
C.-Y. Wu, R. Manmatha, A. J. Smola, and P. Krahenbuhl, “Sampling matters in deep embedding learning,” in Proc. ICCV , Venice, Italy, 2017, pp. 2840–2848
2017
Earlier work this paper cites.
A. van den Oord, O. Vinyals, and K. Kavukcuoglu, “Neural discrete representation learning,” in Proc. NeurIPS , Long Beach, CA, USA, 2017, pp. 6309–6318
2017
Earlier work this paper cites.
J. Bradbury, S. Merity, C. Xiong, and R. Socher, “Quasi-recurrent neural networks,” Proc. ICLR , p. 12 pages, 2017
2017
Earlier work this paper cites.
D. Yu, M. Kolbæk, Z.-H. Tan, and J. Jensen, “Permutation invariant training of deep models for speaker-independent multi-talker speech separation,” in Proc. ICASSP , New Orleans, LA, USA, 2017, pp. 241–245
2017
Earlier work this paper cites.
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” in Proc. ICCV , Venice,Italy, 2017, pp. 609–617
2017
Earlier work this paper cites.
P. Sermanet, C. Lynch, Y. Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain, “Time-contrastive networks: Self-supervised learning from video,” in Proc. ICRA , Brisbane, Australia, 2018, pp. 1134–1141
2018
Earlier work this paper cites.
Y.-A. Chung, W.-H. Weng, S. Tong, and J. Glass, “Unsupervised cross-modal alignment of speech and text embedding spaces,” Proc. NeurIPS , vol. 31, pp. 7354–7364, 2018
2018
Earlier work this paper cites.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The sound of pixels,” in Proc. ECCV , Munich, Germany, 2018, pp. 570–586
2018
Earlier work this paper cites.
N. Komodakis and S. Gidaris, “Unsupervised representation learning by predicting image rotations,” in Proc. ICLR , Salt Lake City, Utah, USA, 2018, p. 16 pages
2018
Earlier work this paper cites.
D. Dwibedi, J. Tompson, C. Lynch, and P. Sermanet, “Learning actionable representations from visual observations,” in Proc. IROS , Madrid,Spain, 2018, pp. 1577–1584
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. Gidaris, P. Singh, and N. Komodakis, “Unsupervised representation learning by predicting image rotations,” in Proc. ICLR , Vancouver, Canada, 2018, p. 16 pages
2018
Earlier work this paper cites.
Y.-A. Chung and J. Glass, “Speech2Vec: A sequence-to-sequence framework for learning word embeddings from speech,” in Proc. INTERSPEECH , Hyderabad, India, 2018, pp. 811–815
2018
Earlier work this paper cites.
M. Caron, P. Bojanowski, A. Joulin, and M. Douze, “Deep clustering for unsupervised learning of visual features,” in Proc. ECCV , Munich,Germany, 2018, pp. 132–149
2018
Earlier work this paper cites.
M. Noroozi, A. Vinjimoor, P. Favaro, and H. Pirsiavash, “Boosting self-supervised learning via knowledge transfer,” in Proc. CVPR , Salt Lake City, Utah, USA, 2018, pp. 9359–9367
2018
Earlier work this paper cites.
M. Ravanelli and Y. Bengio, “Speaker recognition from raw waveform with SincNet,” Proc. SLT , pp. 1021–1028, 2018
2018
Earlier work this paper cites.
D. Wang and J. Chen, “Supervised speech separation based on deep learning: An overview,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 26, no. 10, pp. 1702–1726, 2018
2018
Earlier work this paper cites.
H.-S. Choi, J.-H. Kim, J. Huh, A. Kim, J.-W. Ha, and K. Lee, “Phase-aware speech enhancement with deep complex u-net,” in Proc. ICLR , Vancouver, Canada, 2018, p. 20 pages
2018
Earlier work this paper cites.
——, “Objects that sound,” in Proc. ECCV , Munich, Germany, 2018, pp. 435–451
2018
Earlier work this paper cites.
B. Korbar, D. Tran, and L. Torresani, “Cooperative learning of audio and video models from self-supervised synchronization,” in Proc. NeurIPS , 2018, p. 7774–7785
2018
Earlier work this paper cites.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in Proc. ECCV , Munich, Germany, 2018, pp. 631–648
2018
Earlier work this paper cites.
A. Nagrani, S. Albanie, and A. Zisserman, “Learnable PINs: Cross-modal embeddings for person identity,” in Proc. ECCV , Munich, Germany, 2018, pp. 71–88
2018
Earlier work this paper cites.
M. Alvi, A. Zisserman, and C. Nellåker, “Turning a blind eye: Explicit removal of biases and variation from deep neural network embeddings,” in Proc. ECCV , Munich, Germany, 2018, pp. 556–572
2018
Earlier work this paper cites.
D. Harwath, A. Recasens, D. Surís, G. Chuang, A. Torralba, and J. Glass, “Jointly discovering visual objects and spoken words from raw sensory input,” in Proc. ECCV , Munich, Germany, 2018, pp. 649–665
2018
Earlier work this paper cites.
J. Shen, R. Pang, R. J. Weiss, M. Schuster, N. Jaitly, Z. Yang, Z. Chen, Y. Zhang, Y. Wang, R. Skerrv-Ryan et al. , “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions,” in Proc. ICASSP , 2018, pp. 4779–4783
2018
Earlier work this paper cites.
A. Owens, J. Wu, J. H. Mcdermott, W. T. Freeman, and A. Torralba, “Learning sight from sound: Ambient sound provides supervision for visual learning,” Int. J. Comput. Vision , vol. 126, no. 10, p. 1120–1137, 2018
2018
Earlier work this paper cites.
Z. Wu, Y. Xiong, S. X. Yu, and D. Lin, “Unsupervised feature learning via non-parametric instance discrimination,” in Proc. CVPR , Salt Lake City, UT, USA, 2018, pp. 3733–3742
2018
Cited alongside, same era.
N. Saunshi, O. Plevrakis, S. Arora, M. Khodak, and H. Khandeparkar, “A theoretical analysis of contrastive unsupervised representation learning,” in Proc. ICML , Long Beach, CA, USA, 2019, pp. 5628–5637
2019
Cited alongside, same era.
D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple augmentation method for automatic speech recognition,” in Proc. INTERSPEECH , Graz, Austria, 2019, pp. 2613–2617
2019
Cited alongside, same era.
M. Ravanelli and Y. Bengio, “Learning speaker representations with mutual information,” in Proc. INTERSPEECH , Graz, Austria, 2019, pp. 1153–1157
2019
Cited alongside, same era.
T. Afouras, A. Owens, J. S. Chung, and A. Zisserman, “Self-supervised learning of audio-visual objects from video,” in Proc. ECCV , 2020, pp. 208–224
2020
Later among the works it cites.
A. Shukla, S. Petridis, and M. Pantic, “Learning speech representations from raw audio by joint audiovisual self-supervision,” in Proc. ICML , 2020, p. 8 pages
2020
Later among the works it cites.
A. Shukla, K. Vougioukas, P. Ma, S. Petridis, and M. Pantic, “Visually guided self supervised learning of speech representations,” in Proc. ICASSP , Barcelona,Spain, 2020, pp. 6299–6303
2020
Later among the works it cites.
J.-B. Alayrac, A. Recasens, R. Schneider, R. Arandjelović, J. Ramapuram, J. De Fauw, L. Smaira, S. Dieleman, and A. Zisserman, “Self-supervised multi modal versatile networks,” in Proc. NeurIPS , 2020, p. 13 pages
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. D. Hjelm, A. Fedorov, S. Lavoie-Marchildon, K. Grewal, P. Bachman, A. Trischler, and Y. Bengio, “Learning deep representations by mutual information estimation and maximization,” in Proc. ICLR , Vancouver, Canada, 2019, p. 24 pages
2019
Cited alongside, same era.
M. Tschannen, J. Djolonga, P. K. Rubenstein, S. Gelly, and M. Lucic, “On mutual information maximization for representation learning,” in Proc. ICLR , New Orleans, LA, USA, 2019
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL , Minneapolis, MN, USA, 2019, pp. 4171–4186
2019
Cited alongside, same era.
C. Zhuang, A. L. Zhai, and D. Yamins, “Local aggregation for unsupervised learning of visual embeddings,” in Proc. ICCV , Seoul, South Korea, 2019, pp. 6002–6012
2019
Cited alongside, same era.
Y.-A. Chung, W.-N. Hsu, H. Tang, and J. Glass, “An Unsupervised Autoregressive Model for Speech Representation Learning,” in Proc. INTERSPEECH , Graz, Austria, 2019, pp. 146–150
2019
Cited alongside, same era.
S. Schneider, A. Baevski, R. Collobert, and M. Auli, “wav2vec: Unsupervised pre-training for speech recognition,” in Proc. INTERSPEECH , Graz, Austria, 2019, pp. 3465–3469
2019
Cited alongside, same era.
A. Baevski, S. Schneider, and M. Auli, “vq-wav2vec: Self-supervised learning of discrete speech representations,” in Proc. ICLR , New Orleans, LA, USA, 2019, p. 12 pages
2019
Cited alongside, same era.
S. Pascual, M. Ravanelli, J. Serrà, A. Bonafonte, and Y. Bengio, “Learning problem-agnostic speech representations from multiple self-supervised tasks,” in Proc. INTERSPEECH , Graz, Austria, 2019, pp. 161–165
2019
Cited alongside, same era.
X. Favory, K. Drossos, T. Virtanen, and X. Serra, “COALA: Co-aligned autoencoders for learning semantically enriched audio representations,” Proc. ICML , p. 8 pages, 2020
2020
Later among the works it cites.
S. Khurana, A. Laurent, and J. Glass, “Cstnet: Contrastive speech translation network for self-supervised speech representation learning,” 2020
2020
Later among the works it cites.
S. Siriwardhana, A. Reis, R. Weerasekera, and S. Nanayakkara, “Jointly fine-tuning “bert-like” self supervised models to improve multimodal speech emotion recognition,” in Proc. INTERSPEECH , Shanghai,China, 10 2020, pp. 3755–3759
2020
Later among the works it cites.
H. Nguyen, F. Bougares, N. Tomashenko, Y. Estève, and laurent besacier, “Investigating self-supervised pre-training for end-to-end speech translation,” in Proc. ICML , 2020, p. 7 pages
2020
Later among the works it cites.
J. Engel, R. Swavely, L. H. Hantrakul, A. Roberts, and C. Hawthorne, “Self-supervised pitch detection by inverse audio synthesis,” in Proc. ICML , 2020, p. 9 pages
2020
Later among the works it cites.
“The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,” in Proc. NeurIPS , 2020
2020
Later among the works it cites.
“LeBenchmark: A reproducible framework for assessing self-supervised representation learning from speech,” in Proc. INTERSPEECH , 2020
2020
Later among the works it cites.
J. Kahn, M. Rivière, W. Zheng, E. Kharitonov, Q. Xu, P.-E. Mazaré, J. Karadayi, V. Liptchinsky, R. Collobert, C. Fuegen et al. , “Libri-light: A benchmark for asr with limited or no supervision,” in Proc. ICASSP , Barcelona,Spain, 2020, pp. 7669–7673
2020
Later among the works it cites.
K. Clark, M.-T. Luong, Q. V. Le, and C. D. Manning, “ELECTRA: Pre-training text encoders as discriminators rather than generators,” in Proc. ICLR , 2020, p. 18 pages
2020
Later among the works it cites.
X. Liu, F. Zhang, Z. Hou, L. Mian, Z. Wang, J. Zhang, and J. Tang, “Self-supervised learning: Generative or contrastive,” IEEE Transactions on Knowledge and Data Engineering , p. 20 pages, 2021
2021
Later among the works it cites.
Y. Bansal, G. Kaplun, and B. Barak, “For self-supervised learning, rationality implies generalization, provably,” in Proc. ICLR , Vienna, Austria, 2021, p. 25 pages
2021
Later among the works it cites.
J. Teng and W. Huang, “Can pretext-based self-supervised learning be boosted by downstream data? a theoretical analysis,” CoRR , 2021
2021
Later among the works it cites.
J. D. Lee, Q. Lei, N. Saunshi, and J. Zhuo, “Predicting what you already know helps: Provable self-supervised learning,” in Proc. ICLR , Vienna, Austria, 2021, p. 30 pages
2021
Later among the works it cites.
F. Wang and H. Liu, “Understanding the behaviour of contrastive loss,” in Proc. CVPR , Nashville, TN, USA, 2021, pp. 2495–2504
2021
Later among the works it cites.
A. Jaiswal, A. R. Babu, M. Z. Zadeh, D. Banerjee, and F. Makedon, “A survey on contrastive self-supervised learning,” Technologies , vol. 9, no. 2, p. 22 pages, 2021
2021
Later among the works it cites.
C. Tosh, A. Krishnamurthy, and D. Hsu, “Contrastive learning, multi-view redundancy, and linear models,” in Proc. ALT , 2021, pp. 1179–1206
2021
Later among the works it cites.
L. Wu, H. Lin, C. Tan, Z. Gao, and S. Z. Li, “Self-supervised learning on graphs: Contrastive, generative, or predictive,” IEEE Transactions on Knowledge and Data Engineering , p. 20 pages, 2021
2021
Later among the works it cites.
S. Liu, G. Keren, E. Parada-Cabaleiro, and B. Schuller, “N-HANS: A neural network-based toolkit for in-the-wild audio enhancement,” Multimedia Tools and Applications , vol. 80, pp. 28 365–28 389, 2021
2021
Later among the works it cites.
H. Al-Tahan and Y. Mohsenzadeh, “CLAR: Contrastive learning of auditory representations,” in Proc. AISTATS , 2021, pp. 2530–2538
2021
Later among the works it cites.
A. Saeed, D. Grangier, and N. Zeghidour, “Contrastive learning of general-purpose audio representations,” in Proc. ICASSP , Toronto, Canada, 2021, pp. 3875–3879
2021
Later among the works it cites.
E. Fonseca, D. Ortego, K. McGuinness, N. E. O’Connor, and X. Serra, “Unsupervised contrastive learning of sound event representations,” in Proc. ICASSP , Toronto, Canada, 2021, pp. 371–375
2021
Later among the works it cites.
X. Chen, S. Xie, and K. He, “An empirical study of training self-supervised vision transformers,” in Proc. ICCV , 2021, pp. 9640–9649
2021
Later among the works it cites.
A. H. Liu, Y.-A. Chung, and J. Glass, “Non-autoregressive predictive coding for learning speech representations from local dependencies,” in Proc. INTERSPEECH , Brno, Czechia, 2021, pp. 3730–3734
2021
Later among the works it cites.
J. Zbontar, L. Jing, I. Misra, Y. LeCun, and S. Deny, “Barlow twins: Self-supervised learning via redundancy reduction,” p. 11 pages, 2021
2021
Later among the works it cites.
X. Chen and K. He, “Exploring simple siamese representation learning,” in Proc. ICCV , Montreal, Canada, 2021, pp. 15 750–15 758
2021
Later among the works it cites.
A. N. Carr, Q. Berthet, M. Blondel, O. Teboul, and N. Zeghidour, “Self-supervised learning of audio representations from permutations with differentiable ranking,” IEEE Signal Processing Letters , vol. 28, pp. 708–712, 2021
2021
Later among the works it cites.
Y. Tian, X. Chen, and S. Ganguli, “Understanding self-supervised learning dynamics without contrastive pairs,” in Proc. ICML , 2021, pp. 10 268–10 278
2021
Later among the works it cites.
S. Liu, J. Han, E. Puyal, S. Kontaxis, S. Sun, P. Locatelli, J. Dineley, F. Pokorny, G. Costa, L. Leocani, A. Guerrero, C. Nos, A. Zabalza, P. Soerensen, M. Buron, M. Magyari, Y. Ranjan, Z. Rashid, P. Conde, and R.-C. Consortium, “Fitbeat: COVID-19 estimation based on wristband heart rate using a contrastive convolutional auto-encoder,” Pattern Recognition , vol. 123, p. 108403, 2021
2021
Later among the works it cites.
D. Niizumi, D. Takeuchi, Y. Ohishi, N. Harada, and K. Kashino, “BYOL for audio: Self-supervised learning for general-purpose audio representation,” pp. 1–8, 2021
2021
Later among the works it cites.
F. Gontier, V. Lostanlen, M. Lagrange, N. Fortin, C. Lavandier, and J.-F. Petiot, “Polyphonic training set synthesis improves self-supervised urban sound classification,” The Journal of the Acoustical Society of America , vol. 149, no. 6, pp. 4309–4326, 2021
2021
Later among the works it cites.
A. T. Liu, S.-W. Li, and H.-y. Lee, “TERA: Self-supervised learning of transformer encoder representation for speech,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 29, pp. 2351–2366, 2021
2021
Later among the works it cites.
P.-H. Chi, P.-H. Chung, T.-H. Wu, C.-C. Hsieh, Y.-H. Chen, S.-W. Li, and H.-y. Lee, “Audio albert: A lite bert for self-supervised learning of audio representation,” in Proc. SLT , 2021, pp. 344–350
2021
Later among the works it cites.
J. Bai, W. Wang, Y. Zhou, and C. Xiong, “Representation learning for sequence data with deep autoencoding predictive components,” in Proc. ICLR , 2021. [Online]. Available: https://openreview.net/forum?id=Naqw7EHIfrv
2021
Later among the works it cites.
E. Kharitonov, M. Rivière, G. Synnaeve, L. Wolf, P.-E. Mazaré, M. Douze, and E. Dupoux, “Data augmenting contrastive learning of speech representations in the time domain,” in Proc. SLT , 2021, pp. 215–222
2021
Later among the works it cites.
2021
Later among the works it cites.
W.-N. Hsu, A. Sriram, A. Baevski, T. Likhomanenko, Q. Xu, V. Pratap, J. Kahn, A. Lee, R. Collobert, G. Synnaeve, and M. Auli, “Robust wav2vec 2.0: Analyzing domain shift in self-supervised pre-training,” in Proc. INTERSPEECH , Brno, Czech Republic, 2021, pp. 721–725
2021
Later among the works it cites.
A. Babu, C. Wang, A. Tjandra, K. Lakhotia, Q. Xu, N. Goyal, K. Singh, P. von Platen, Y. Saraf, J. Pino, A. Baevski, A. Conneau, and M. Auli, “XLS-R: Self-supervised cross-lingual speech representation learning at scale,” 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Baevski, W.-N. Hsu, A. Conneau, and M. Auli, “Unsupervised speech recognition,” in Proc. NeurIPS , 2021, p. 15 pages
2021
Later among the works it cites.
Y. Qiu, R. Wang, S. Singh, Z. Ma, and F. Hou, “Self-supervised learning based phone-fortified speech enhancement,” in Proc. INTERSPEECH , Brno, Czech Republic, 2021, pp. 211–215
2021
Later among the works it cites.
S.-F. Huang, S.-P. Chuang, D.-R. Liu, Y.-C. Chen, G.-P. Yang, and H.-y. Lee, “Stabilizing label assignment for speech separation by self-supervised pre-training,” in Proc. INTERSPEECH , 08 2021, pp. 3056–3060
2021
Later among the works it cites.
A. Sivaraman, S. Kim, and M. Kim, “Personalized speech enhancement through self-supervised data augmentation and purification,” in Proc. INTERSPEECH , 08 2021, pp. 2676–2680
2021
Later among the works it cites.
J. Zhang, X. Xu, F. Shen, H. Lu, X. Liu, and H. T. Shen, “Enhancing audio-visual association with self-supervised curriculum learning,” in Proc. AAAI Conference on Artificial Intelligence , 2021, pp. 3351–3359
2021
Later among the works it cites.
W.-N. Hsu, D. F. Harwath, C. Song, and J. R. Glass, “Text-free image-to-speech synthesis using learned segmental units,” in Proc. ACL/IJCNLP , 2021, p. 25 pages
2021
Later among the works it cites.
P. Morgado, N. Vasconcelos, and I. Misra, “Audio-visual instance discrimination with cross-modal agreement,” in Proc. CVPR , Nashville, TN, USA, 2021, pp. 12 475–12 486
2021
Later among the works it cites.
E. Tzinis, S. Wisdom, A. Jansen, S. Hershey, T. Remez, D. Ellis, and J. R. Hershey, “Into the wild with audioscope: Unsupervised audio-visual separation of on-screen sounds,” in Proc. ICLR , 2021, p. 9 pages
2021
Later among the works it cites.
A. Shukla, S. Petridis, and M. Pantic, “Does visual self-supervision improve learning of speech representations for emotion recognition,” IEEE Transactions on Affective Computing , 2021
2021
Later among the works it cites.
H. Akbari, L. Yuan, R. Qian, W.-H. Chuang, S.-F. Chang, Y. Cui, and B. Gong, “VATT: Transformers for multimodal self-supervised learning from raw video, audio and text,” in Proc. NeurIPS , 2021, p. 20 pages
2021
Later among the works it cites.
“SUPERB: Speech processing universal performance benchmark,” in Proc. INTERSPEECH , 2021
2021
Later among the works it cites.
J. L. Suárez, S. García, and F. Herrera, “A tutorial on distance metric learning: Mathematical foundations, algorithms, experimental analysis, prospects and challenges,” Neurocomputing , vol. 425, pp. 300–322, 2021
2021
Later among the works it cites.
C. Wang, Y. Wu, Y. Qian, K. Kumatani, S. Liu, F. Wei, M. Zeng, and X. Huang, “Unispeech: Unified speech representation learning with labeled and unlabeled data,” in Proceedings of the 38th Proc. ICML, ICML 2021, 18-24 July 2021, Virtual Event , ser. Proceedings of Machine Learning Research, vol. 139. PMLR, 2021, pp. 10 937–10 947
2021
Later among the works it cites.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in Proc. ICML , Lille, France, 2015, pp. 2048–2057
2057
Closest in time.