Fetching the paper…
Reading the bibliography…
Through solving pretext tasks, self-supervised learning leverages unlabeled data to extract useful latent representations replacing traditional input features in the downstream task.
1904
Earlier work this paper cites.
1907
Earlier work this paper cites.
K. Wei, R. Iyer, and J. Bilmes, “Submodularity in data subset selection and active learning,” in Proceedings of the 32nd International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, F. Bach and D. Blei, Eds., vol. 37. Lille, France: PMLR, 07–09 Jul 2015, pp. 1954–1963. [Online]. Available: https://proceedings.mlr.press/v37/wei15.html
1963
Earlier work this paper cites.
J. Garofolo, L. Lamel, W. Fisher, J. Fiscus, D. Pallett, N. Dahlgren, and V. Zue, “Timit acoustic-phonetic continuous speech corpus,” Linguistic Data Consortium , 11 1992
1992
Earlier work this paper cites.
H. Hermansky, N. Morgan, A. Bayya, and P. Kohn, “Rasta-plp speech analysis technique,” vol. 1, 04 1992, pp. 121 – 124 vol.1
1992
Earlier work this paper cites.
J. Herre, E. Allamanche, and O. Hellmuth, “Robust matching of audio signals using spectral flatness features,” in Proceedings of the 2001 IEEE Workshop on the Applications of Signal Processing to Audio and Acoustics (Cat. No. 01TH8575) . IEEE, 2001, pp. 127–130
2001
Earlier work this paper cites.
I. Guyon, J. Weston, S. Barnhill, and V. Vapnik, “Gene selection for cancer classification using support vector machines,” Machine Learning , vol. 46, pp. 389–422, 01 2002
2002
Earlier work this paper cites.
I. Guyon and A. Elisseeff, “An introduction of variable and feature selection,” J. Machine Learning Research Special Issue on Variable and Feature Selection , vol. 3, pp. 1157 – 1182, 01 2003
2003
Earlier work this paper cites.
P. Murphy and O. Akande, “Cepstrum-Based Harmonics-to-Noise Ratio Measurement in Voiced Speech,” in Nonlinear Speech Modeling and Applications , G. Chollet, A. Esposito, M. Faundez-Zanuy, and M. Marinaro, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2005, pp. 199–218
2005
Earlier work this paper cites.
H. Peng, F. Long, and C. Ding, “Feature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 27, no. 8, pp. 1226–1238, 2005
2005
Earlier work this paper cites.
S. Essid, “Classification automatique des signaux audio-fréquences: reconnaissance des instruments de musique,” Ph.D. dissertation, Université Pierre et Marie Curie-Paris VI, 2005
2005
Earlier work this paper cites.
2006
Earlier work this paper cites.
M. Yuan and Y. Lin, “Model selection and estimation in regression with grouped variables,” Journal of the Royal Statistical Society Series B , vol. 68, pp. 49–67, 02 2006
2006
Earlier work this paper cites.
S. Sonnenburg, G. Rätsch, C. Schäfer, and B. Schölkopf, “Large scale multiple kernel learning,” J. Mach. Learn. Res. , vol. 7, p. 1531–1565, Dec. 2006
2006
Earlier work this paper cites.
J. Sundberg and M. Nordenberg, “Effects of vocal loudness variation on spectrum balance as reflected by the alpha measure of long-term-average spectra of speech,” The Journal of the Acoustical Society of America , vol. 120, pp. 453–7, 08 2006
2006
Earlier work this paper cites.
S. Ioffe, “Probabilistic Linear Discriminant Analysis,” in Computer Vision – ECCV 2006 , A. Leonardis, H. Bischof, and A. Pinz, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2006, pp. 531–542
2006
Earlier work this paper cites.
A. Rakotomamonjy, F. Bach, S. Canu, and Y. Grandvalet, “More efficiency in multiple kernel learning,” Proceedings of the 24th International Con- ference on Machine Learning (ICML) , vol. 227, 01 2007
2007
Earlier work this paper cites.
A. Gretton, K. Fukumizu, C. H. Teo, L. Song, B. Schölkopf, and A. Smola, “A kernel statistical test of independence,” 01 2007
2007
Earlier work this paper cites.
B. Schuller, B. Vlasenko, R. Minguez, G. Rigoll, and A. Wendemuth, “Comparing one and two-stage acoustic modeling in the recognition of emotion in speech,” in 2007 IEEE Workshop on Automatic Speech Recognition Understanding (ASRU) , 2007, pp. 596–600
2007
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, E. A. Kazemzadeh, E. M. Provost, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: interactive emotional dyadic motion capture database,” Language Resources and Evaluation , vol. 42, pp. 335–359, 2008
2008
Earlier work this paper cites.
2010
Earlier work this paper cites.
2010
Earlier work this paper cites.
B. Mathieu, S. Essid, T. Fillon, J. Prado, and G. Richard, “Yaafe, an easy to use and efficient audio feature extraction software,” in Proceedings of the 11th International Society for Music Information Retrieval Conference , Utrecht, The Netherlands, August 9-13 2010, pp. 441–446, http://ismir2010.ismir.net/proceedings/ismir2010-75.pdf
2010
Earlier work this paper cites.
M. Carlin, S. Thomas, A. Jansen, and H. Hermansky, “Rapid evaluation of speech representations for spoken term discovery,” 01 2011, pp. 821–824
2011
Earlier work this paper cites.
A. Graves, “Connectionist temporal classification,” in Supervised Sequence Labelling with Recurrent Neural Networks . Springer, 2012, pp. 61–93
2012
Earlier work this paper cites.
T. Schatz, V. Peddinti, F. Bach, A. Jansen, H. Hermansky, and E. Dupoux, “Evaluating speech features with the Minimal-Pair ABX task: Analysis of the classical MFC/PLP pipeline,” in INTERSPEECH 2013 : 14th Annual Conference of the International Speech Communication Association , Lyon, France, Aug. 2013, pp. 1–5
2013
Earlier work this paper cites.
D. Renshaw, H. Kamper, A. Jansen, and S. Goldwater, “A comparison of neural network methods for unsupervised representation learning on the zero resource speech challenge,” in INTERSPEECH , 2015
2015
Earlier work this paper cites.
Z. Chen, S. Watanabe, H. Erdogan, and J. Hershey, “Speech enhancement and recognition using multi-task learning of long short-term memory recurrent neural networks,” in INTERSPEECH , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
E. Loweimi, M. Doulaty, J. Barker, and T. Hain, “Long-term statistical feature extraction from speech signal and its application in emotion recognition,” 11 2015
2015
Earlier work this paper cites.
V. Peddinti, D. Povey, and S. Khudanpur, “A time delay neural network architecture for efficient modeling of long temporal contexts,” in INTERSPEECH , 2015
2015
Earlier work this paper cites.
D. Snyder, D. Garcia-Romero, and D. Povey, “Time Delay Deep Neural Network-Based Universal Background Models for Speaker Recognition,” in 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) , 2015, pp. 92–97. [Online]. Available: https://app.dimensions.ai/details/publication/pub.1093368586
2015
Earlier work this paper cites.
V. Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Librispeech: An asr corpus based on public domain audio books,” 04 2015, pp. 5206–5210
2015
Earlier work this paper cites.
C. Doersch, A. Gupta, and A. A. Efros, “Unsupervised visual representation learning by context prediction,” 2016
2016
Cited alongside, same era.
A. F. T. Martins and R. F. Astudillo, “From softmax to sparsemax: A sparse model of attention and multi-label classification,” in Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48 , ser. ICML’16. JMLR.org, 2016, p. 1614–1623
2016
Cited alongside, same era.
M. Noroozi and P. Favaro, “Unsupervised learning of visual representations by solving jigsaw puzzles,” 2017
2017
Cited alongside, same era.
C. Doersch, A. Zisserman, and Deepmind, “Multi-task Self-Supervised Visual Learning,” Tech. Rep., 2017
2017
Cited alongside, same era.
M. Ravanelli, J. Zhong, S. Pascual, P. Swietojanski, J. Monteiro, J. Trmal, and Y. Bengio, “Multi-task self-supervised learning for robust speech recognition,” 2020
2020
Later among the works it cites.
R. Algayres, M. S. Zaiem, B. Sagot, and E. Dupoux, “Evaluating the reliability of acoustic speech embeddings,” in INTERSPEECH 2020 - Annual Conference of the International Speech Communication Association , Shanghai / Vitrtual, China, Oct. 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
A. Nagrani, J. S. Chung, and A. Zisserman, “Voxceleb: A large-scale speaker identification dataset,” Interspeech 2017 , Aug 2017
2017
Cited alongside, same era.
M. McAuliffe, M. Socolof, S. Mihuc, M. Wagner, and M. Sonderegger, “Montreal forced aligner: Trainable text-speech alignment using kaldi,” 08 2017, pp. 498–502
2017
Cited alongside, same era.
R. Arandjelovic and A. Zisserman, “Objects that sound,” in Proceedings of the European Conference on Computer Vision (ECCV) , September 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
S. Khurana, A. Laurent, W.-N. Hsu, J. Chorowski, A. Lancucki, R. Marxer, and J. Glass, “A convolutional deep markov model for unsupervised speech representation learning,” 2020
2020
Later among the works it cites.
A. Saeed, D. Grangier, and N. Zeghidour, “Contrastive Learning of General-Purpose Audio Representations,” oct 2020
2020
Later among the works it cites.
A. Shukla, S. Petridis, and M. Pantic, “Learning speech representations from raw audio by joint audiovisual self-supervision,” 07 2020
2020
Later among the works it cites.
J. D. Lee, Q. Lei, N. Saunshi, and J. Zhuo, “Predicting what you already know helps: Provable self-supervised learning,” 2020
2020
Later among the works it cites.
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in Proceedings of the 37th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, H. D. III and A. Singh, Eds., vol. 119. PMLR, 13–18 Jul 2020, pp. 1597–1607. [Online]. Available: http://proceedings.mlr.press/v119/chen20j.html
2020
Later among the works it cites.
M. Tschannen, J. Djolonga, P. K. Rubenstein, S. Gelly, and M. Lucic, “On mutual information maximization for representation learning,” in 8th International Conference on Learning Representations (ICLR) , Apr. 2020. [Online]. Available: https://openreview.net/forum?id=rkxoh24FPH
2020
Later among the works it cites.
Y. Tian, C. Sun, B. Poole, D. Krishnan, C. Schmid, and P. Isola, “What Makes for Good Views for Contrastive Learning?” in Advances in Neural Information Processing Systems , H. Larochelle, M. Ranzato, R. Hadsell, M. F. Balcan, and H. Lin, Eds., vol. 33. Curran Associates, Inc., 2020, pp. 6827–6839. [Online]. Available: https://proceedings.neurips.cc/paper/2020/file/4c2e5eaae9152079b9e95845750bb9ab-Paper.pdf
2020
Later among the works it cites.
M. Gump, W.-N. Hsu, and J. Glass, “Unsupervised Methods for Evaluating Speech Representations,” 2020. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-2990
2020
Later among the works it cites.
R. Ardila, M. Branson, K. Davis, M. Henretty, M. Kohler, J. Meyer, R. Morais, L. Saunders, F. M. Tyers, and G. Weber, “Common voice: A massively-multilingual speech corpus,” 2020
2020
Later among the works it cites.
A. Baevski, S. Schneider, and M. Auli, “vq-wav2vec: Self-supervised learning of discrete speech representations,” 2020
2020
Later among the works it cites.
W.-N. Hsu, Y.-H. H. Tsai, B. Bolte, R. Salakhutdinov, and A. Mohamed, “Hubert: How much can a bad teacher benefit asr pre-training?” in ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2021, pp. 6533–6537
2021
Closest in time.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y. Qian, Y. Qian, J. Wu, M. Zeng, and F. Wei, “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” 2021
2021
Closest in time.
C. Wang, Y. Wu, Y. Qian, K. Kumatani, S. Liu, F. Wei, M. Zeng, and X. Huang, “Unispeech: Unified speech representation learning with labeled and unlabeled data,” 2021
2021
Closest in time.
H.-H. Wu, C.-C. Kao, Q. Tang, M. Sun, B. McFee, J. P. Bello, and C. Wang, “Multi-task self-supervised pre-training for music classification,” 2021
2021
Closest in time.
S. Zaiem, T. Parcollet, and S. Essid, “Conditional independence for pretext task selection in self-supervised speech representation learning,” 2021
2021
Closest in time.
M. Ravanelli, T. Parcollet, A. Rouhe, P. Plantinga, E. Rastorgueva, L. Lugosch, N. Dawalatabad, C. Ju-Chieh, A. Heba, F. Grondin, W. Aris, C.-F. Liao, S. Cornell, S.-L. Yeh, H. Na, Y. Gao, S.-W. Fu, C. Subakan, R. De Mori, and Y. Bengio, “Speechbrain,” https://github.com/speechbrain/speechbrain
2021
Closest in time.
Z. Fan, M. Li, S. Zhou, and B. Xu, “Exploring wav2vec 2.0 on speaker verification and language identification,” 2021
2021
Closest in time.
2021
Closest in time.
S. Yang, P.-H. Chi, Y.-S. Chuang, C.-I. J. Lai, K. Lakhotia, Y. Y. Lin, A. T. Liu, J. Shi, X. Chang, G.-T. Lin, T.-H. Huang, W.-C. Tseng, K. tik Lee, D.-R. Liu, Z. Huang, S. Dong, S.-W. Li, S. Watanabe, A. Mohamed, and H. yi Lee, “Superb: Speech processing universal performance benchmark,” 2021
2021
Closest in time.
S. Evain, H. Nguyen, H. Le, M. Z. Boito, S. Mdhaffar, S. Alisamir, Z. Tong, N. Tomashenko, M. Dinarelli, T. Parcollet, A. Allauzen, Y. Esteve, B. Lecouteux, F. Portet, S. Rossato, F. Ringeval, D. Schwab, and L. Besacier, “Lebenchmark: A reproducible framework for assessing self-supervised representation learning from speech,” 2021
2021
Closest in time.
T. Xiao, X. Wang, A. A. Efros, and T. Darrell, “What should not be contrastive in contrastive learning,” 2021
2021
Closest in time.
Y. Li, R. Pogodin, D. J. Sutherland, and A. Gretton, “Self-supervised learning with kernel dependence maximization,” 2021
2021
Closest in time.
K. Lakhotia, E. Kharitonov, W.-N. Hsu, Y. Adi, A. Polyak, B. Bolte, T.-A. Nguyen, J. Copet, A. Baevski, A. Mohamed, and E. Dupoux, “Generative spoken language modeling from raw audio,” 2021
2021
Closest in time.
S. Chen, Y. Wu, C. Wang, Z. Chen, Z. Chen, S. Liu, J. Wu, Y. Qian, F. Wei, J. Li, and X. Yu, “Unispeech-sat: Universal speech representation learning with speaker aware pre-training,” 2021
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
S. Sadhu, D. He, C.-W. Huang, S. H. Mallidi, M. Wu, A. Rastrow, A. Stolcke, J. Droppo, and R. Maas, “wav2vec-C: A Self-Supervised Model for Speech Representation Learning,” in Proc. Interspeech 2021 , 2021, pp. 711–715
2021
Closest in time.
2021
Closest in time.
R. Iyer, N. Khargonkar, J. Bilmes, and H. Asnani, “Generalized submodular information measures: Theoretical properties, examples, optimization algorithms, and applications,” IEEE Transactions on Information Theory , vol. 68, no. 2, pp. 752–781, 2022
2022
Closest in time.