Fetching the paper…
Reading the bibliography…
Self-supervised learning has attracted plenty of recent research interest.
A. Kolesnikov, X. Zhai, and L. Beyer, “Revisiting self-supervised visual representation learning,” in CVPR , 2019, pp. 1920–1929
1929
Earlier work this paper cites.
H. McGurk and J. MacDonald, “Hearing lips and seeing voices,” Nature , vol. 264, no. 5588, p. 746, 1976
1976
Earlier work this paper cites.
A. Varga and H. J. Steeneken, “Assessment for automatic speech recognition: Ii. noisex-92: A database and an experiment to study the effect of additive noise on speech recognition systems,” Speech communication , vol. 12, no. 3, pp. 247–251, 1993
1993
Earlier work this paper cites.
C. Bregler, M. Covell, and M. Slaney, “Video rewrite: Driving visual speech with audio,” in Proceedings of the 24th annual conference on Computer graphics and interactive techniques , 1997, pp. 353–360
1997
Earlier work this paper cites.
T. Ezzat, G. Geiger, and T. Poggio, “Trainable videorealistic speech animation,” ACM Transactions on Graphics (TOG) , vol. 21, no. 3, pp. 388–398, 2002
2002
Earlier work this paper cites.
C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Language resources and evaluation , vol. 42, no. 4, p. 335, 2008
2008
Earlier work this paper cites.
J. F. Cohn, T. S. Kruez, I. Matthews, Y. Yang, M. H. Nguyen, M. T. Padilla, F. Zhou, and F. De la Torre, “Detecting depression from facial actions and vocal prosody,” in 2009 3rd International Conference on Affective Computing and Intelligent Interaction and Workshops . IEEE, 2009, pp. 1–7
2009
Earlier work this paper cites.
F. Ringeval, A. Sonderegger, J. Sauer, and D. Lalanne, “Introducing the recola multimodal corpus of remote collaborative and affective interactions,” in 2013 10th IEEE international conference and workshops on automatic face and gesture recognition (FG) . IEEE, 2013, pp. 1–8
2013
Earlier work this paper cites.
F. Eyben, F. Weninger, F. Gross, and B. Schuller, “Recent developments in opensmile, the munich open-source multimedia feature extractor,” in Proceedings of the 21st ACM international conference on Multimedia , 2013, pp. 835–838
2013
Earlier work this paper cites.
H. Cao, D. Cooper, M. Keutmann, R. Gur, A. Nenkova, and R. Verma, “Crema-d: Crowd-sourced emotional multimodal actors dataset,” IEEE transactions on affective computing , vol. 5, no. 4, pp. 377–390, 2014
2014
Earlier work this paper cites.
C. Doersch, A. Gupta, and A. Efros, “Unsupervised visual representation learning by context prediction,” in ICCV , 2015, pp. 1422–1430
2015
Earlier work this paper cites.
S. Petridis and M. Pantic, “Prediction-based audiovisual fusion for classification of non-linguistic vocalisations,” IEEE Transactions on Affective Computing , vol. 7, no. 1, pp. 45–58, 2015
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in MICCAI , 2015, pp. 234–241
2015
Earlier work this paper cites.
J. Chung and A. Zisserman, “Lip reading in the wild,” in ACCV , 2016
2016
Earlier work this paper cites.
B. Fernando, H. Bilen, E. Gavves, and S. Gould, “Self-supervised video representation learning with odd-one-out networks,” in CVPR , 2017, pp. 3636–3645
2017
Earlier work this paper cites.
S. Suwajanakorn, S. M. Seitz, and I. Kemelmacher-Shlizerman, “Synthesizing obama: learning lip sync from audio,” ACM Transactions on Graphics (TOG) , vol. 36, no. 4, pp. 1–13, 2017
2017
Earlier work this paper cites.
A. Shukla, S. S. Gullapuram, H. Katti, K. Yadati, M. Kankanhalli, and R. Subramanian, “Evaluating content-centric vs. user-centric ad affect recognition,” in Proceedings of the 19th ACM International Conference on Multimodal Interaction , 2017, pp. 402–410
2017
Earlier work this paper cites.
A. Shukla, S. S. Gullapuram, H. Katti, K. Yadati, M. Kankanhalli, and R. Subramanian, “Affect recognition in ads with application to computational advertising,” in Proceedings of the 25th ACM international conference on Multimedia , 2017, pp. 1148–1156
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in NAACL-HLT , 2018, pp. 2227–2237
2018
Earlier work this paper cites.
S. Gidaris, P. Singh, and N. Komodakis, “Unsupervised representation learning by predicting image rotations,” in ICLR , 2018
2018
Earlier work this paper cites.
M. Caron, P. Bojanowski, A. Joulin, and M. Douze, “Deep clustering for unsupervised learning of visual features,” in ECCV , 2018, pp. 132–149
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Owens and A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” ECCV , 2018
2018
Earlier work this paper cites.
A. Owens, J. Wu, J. McDermott, W. Freeman, and A. Torralba, “Learning sight from sound: Ambient sound provides supervision for visual learning,” IJCV , vol. 126, no. 10, pp. 1120–1137, 2018
2018
Cited alongside, same era.
S. Petridis, T. Stafylakis, P. Ma, F. Cai, G. Tzimiropoulos, and M. Pantic, “End-to-end audiovisual speech recognition,” in 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2018, pp. 6548–6552
2018
Cited alongside, same era.
A. Shukla, H. Katti, M. Kankanhalli, and R. Subramanian, “Looking beyond a clever narrative: Visual context and attention are primary drivers of affect in video advertisements,” in Proceedings of the 20th ACM International Conference on Multimodal Interaction , 2018, pp. 210–219
2018
Cited alongside, same era.
A. B. Zadeh, P. P. Liang, S. Poria, E. Cambria, and L.-P. Morency, “Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 2236–2246
2019
Later among the works it cites.
J. Chorowski, R. J. Weiss, S. Bengio, and A. van den Oord, “Unsupervised speech representation learning using wavenet autoencoders,” IEEE/ACM transactions on audio, speech, and language processing , vol. 27, no. 12, pp. 2041–2053, 2019
2019
Later among the works it cites.
H. Pham, P. Liang, T. Manzini, L. Morency, and B. Póczos, “Found in translation: Learning robust joint representations by cyclic translations between modalities,” in AAAI , vol. 33, 2019, pp. 6892–6899
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
K. Vougioukas, S. Petridis, and M. Pantic, “End-to-end speech-driven facial animation with temporal gans,” Proceedings of the British Conference on Machine Vision (BMVC) , 2018
2018
Cited alongside, same era.
S. Livingstone and F. Russo, “The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,” PloS one , vol. 13, no. 5, p. e0196391, 2018
2018
Cited alongside, same era.
N. Majumder, D. Hazarika, A. Gelbukh, E. Cambria, and S. Poria, “Multimodal sentiment analysis using hierarchical fusion with context modeling,” Knowledge-Based Systems , vol. 161, pp. 124–133, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
R. Arandjelovic and A. Zisserman, “Objects that sound,” in ECCV , 2018, pp. 435–451
2018
Cited alongside, same era.
B. Korbar, D. Tran, and L. Torresani, “Cooperative learning of audio and video models from self-supervised synchronization,” in NeurIPS , 2018, pp. 7763–7774
2018
Cited alongside, same era.
M. Ravanelli, P. Brakel, M. Omologo, and Y. Bengio, “Light gated recurrent units for speech recognition,” IEEE Transactions on Emerging Topics in Computational Intelligence , vol. 2, no. 2, pp. 92–102, 2018
2018
Cited alongside, same era.
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. Freeman, and M. Rubinstein, “Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation,” SIGGRAPH , 2018
2018
Cited alongside, same era.
D. Kollias, P. Tzirakis, M. A. Nicolaou, A. Papaioannou, G. Zhao, B. Schuller, I. Kotsia, and S. Zafeiriou, “Deep affect prediction in-the-wild: Aff-wild database and challenge, deep architectures, and beyond,” International Journal of Computer Vision , pp. 1–23, 2019
2019
Later among the works it cites.
K. Vougioukas, S. Petridis, and M. Pantic, “Realistic speech-driven facial animation with gans,” International Journal of Computer Vision , pp. 1–16, 2019
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-performance deep learning library,” in Advances in Neural Information Processing Systems 32 , 2019, pp. 8024–8035
2019
Later among the works it cites.
J. Kossaifi, R. Walecki, Y. Panagakis, J. Shen, M. Schmitt, F. Ringeval, J. Han, V. Pandit, A. Toisoul, B. W. Schuller et al. , “Sewa db: A rich database for audio-visual emotion and sentiment research in the wild,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2019
2019
Later among the works it cites.
K. Vougioukas, P. Ma, S. Petridis, and M. Pantic, “Video-driven speech reconstruction using generative adversarial networks,” Proc. Interspeech 2019 , pp. 4125–4129, 2019
2019
Later among the works it cites.
2020
Closest in time.
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in CVPR , 2020, pp. 9729–9738
2020
Closest in time.
I. Misra and L. v. d. Maaten, “Self-supervised learning of pretext-invariant representations,” in CVPR , 2020, pp. 6707–6717
2020
Closest in time.
Q. Xie, M.-T. Luong, E. Hovy, and Q. V. Le, “Self-training with noisy student improves imagenet classification,” in CVPR , 2020, pp. 10 687–10 698
2020
Closest in time.
A. Kumar and V. K. Ithapu, “Secost:: Sequential co-supervision for large scale weakly labeled audio event detection,” in ICASSP 2020 . IEEE, 2020, pp. 666–670
2020
Closest in time.
M. Rivière, A. Joulin, P.-E. Mazaré, and E. Dupoux, “Unsupervised pretraining transfers well across languages,” ICASSP , 2020
2020
Closest in time.
A. Piergiovanni, A. Angelova, and M. S. Ryoo, “Evolving losses for unsupervised video representation learning,” CVPR , 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.
A. Shukla, S. S. Gullapuram, H. Katti, M. Kankanhalli, S. Winkler, and R. Subramanian, “Recognition of advertisement emotions with application to computational advertising,” IEEE Transactions on Affective Computing , 2020
2020
Closest in time.
A. Shukla, K. Vougioukas, P. Ma, S. Petridis, and M. Pantic, “Visually guided self supervised learning of speech representations,” Proceedings of the International Conference on Acoustics Speech and Signal Processing (ICASSP) , 2020
2020
Closest in time.
A. Shukla, S. Petridis, and M. Pantic, “Visual self-supervision by facial reconstruction for speech representation learning,” Sight and Sound Workshop, CVPR , 2020
2020
Closest in time.
A. Shukla, S. Petridis, and M. Pantic, “Learning speech representations from raw audio by joint audiovisual self-supervision,” Workshop on Self-supervision in Audio and Speech, ICML , 2020
2020
Closest in time.