Fetching the paper…
Reading the bibliography…
Our objective is to transform a video into a set of discrete audio-visual objects using self-supervised learning.
Hershey, J., Movellan, J.: Audio-vision: Locating sounds via audio-visual synchrony. In: NeurIPS. vol. 12 (1999)
1999
Earlier work this paper cites.
Cutler, R., Davis, L.: Look who’s talking: Speaker detection using video and audio correlation. In: 2000 IEEE International Conference on Multimedia and Expo. ICME2000. Proceedings. Latest Advances in the Fast Changing World of Multimedia (Cat. No. 00TH8532). vol. 3, pp. 1589–1592. IEEE (2000)
2000
Earlier work this paper cites.
Fisher III, J.W., Darrell, T., Freeman, W.T., Viola, P.A.: Learning joint statistical models for audio-visual fusion and segregation. In: NeurIPS (2000)
2000
Earlier work this paper cites.
Rix, A.W., Beerends, J.G., Hollier, M.P., Hekstra, A.P.: Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs. In: Proc. ICASSP. vol. 2, pp. 749–752. IEEE (2001)
2001
Earlier work this paper cites.
Févotte, C., Gribonval, R., Vincent, E.: BSS EVAL toolbox user guide. IRISA Technical Report 1706. http://www.irisa.fr/metiss/bss eval/. (2005)
2005
Earlier work this paper cites.
Kidron, E., Schechner, Y.Y., Elad, M.: Pixels that sound. In: Proc. CVPR (2005)
2005
Earlier work this paper cites.
Barzelay, Z., Schechner, Y.Y.: Harmony in motion. In: 2007 IEEE Conference on Computer Vision and Pattern Recognition (2007)
2007
Earlier work this paper cites.
Izadinia, H., Saleemi, I., Shah, M.: Multimodal analysis for identification and segmentation of moving-sounding objects. IEEE Transactions on Multimedia 15
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
Doersch, C., Gupta, A., Efros, A.A.: Unsupervised visual representation learning by context prediction. In: Proc. ICCV. pp. 1422–1430 (2015)
2015
Earlier work this paper cites.
Pfister, T., Charles, J., Zisserman, A.: Flowing convnets for human pose estimation in videos. In: Proc. ICCV (2015)
2015
Earlier work this paper cites.
Wang, X., Gupta, A.: Unsupervised learning of visual representations using videos. In: Proc. ICCV. pp. 2794–2802 (2015)
2015
Earlier work this paper cites.
Chakravarty, P., Tuytelaars, T.: Cross-modal supervision for learning active speaker detection in video. In: Proc. ECCV (2016)
2016
Earlier work this paper cites.
Chung, J.S., Zisserman, A.: Out of time: automated lip sync in the wild. In: Workshop on Multi-view Lip-reading, ACCV (2016)
2016
Earlier work this paper cites.
Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., Berg, A.C.: SSD: Single shot multibox detector. In: Proc. ECCV. pp. 21–37. Springer (2016)
2016
Earlier work this paper cites.
Owens, A., Isola, P., McDermott, J., Torralba, A., Adelson, E.H., Freeman, W.T.: Visually indicated sounds. In: Computer Vision and Pattern Recognition (CVPR) (2016)
2016
Earlier work this paper cites.
Arandjelović, R., Zisserman, A.: Look, listen and learn. In: Proc. ICCV (2017)
2017
Earlier work this paper cites.
Arandjelovic, R., Zisserman, A.: Objects that sound. In: Proc. ECCV (2017)
2017
Earlier work this paper cites.
Gadde, R., Jampani, V., Gehler, P.V.: Semantic video cnns through representation warping. In: Proc. ICCV. pp. 4463–4472 (2017)
2017
Earlier work this paper cites.
Gebru, I.D., Ba, S., Li, X., Horaud, R.: Audio-visual speaker diarization based on spatiotemporal bayesian fusion. IEEE PAMI (2017)
2017
Earlier work this paper cites.
Nagrani, A., Chung, J.S., Zisserman, A.: VoxCeleb: a large-scale speaker identification dataset. In: INTERSPEECH (2017)
2017
Cited alongside, same era.
Afouras, T., Chung, J.S., Zisserman, A.: The conversation: Deep audio-visual speech enhancement. In: INTERSPEECH (2018)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Chung, J.S., Nagrani, A., Zisserman, A.: VoxCeleb2: Deep speaker recognition. In: INTERSPEECH (2018)
2018
Cited alongside, same era.
Ephrat, A., Mosseri, I., Lang, O., Dekel, T., Wilson, K., Hassidim, A., Freeman, W.T., Rubinstein, M.: Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation. ACM Transactions on Graphics (TOG) 37
Deng, J., Guo, J., Yuxiang, Z., Yu, J., Kotsia, I., Zafeiriou, S.: Retinaface: Single-stage dense face localisation in the wild. In: arxiv (2019)
2019
Later among the works it cites.
Dutta, A., Zisserman, A.: The VIA annotation software for images, audio and video. In: Proceedings of the 27th ACM International Conference on Multimedia. MM ’19, ACM, New York, NY, USA (2019)
2019
Later among the works it cites.
Gan, C., Zhao, H., Chen, P., Cox, D., Torralba, A.: Self-supervised moving vehicle tracking with stereo sound. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 7053–7062 (2019)
2019
Later among the works it cites.
Gao, R., Grauman, K.: 2.5d visual sound. In: CVPR (2019)
2019
Later among the works it cites.
Gao, R., Grauman, K.: Co-separating sounds of visual objects. arXiv preprint arXiv:1904.07750 (2019)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
Gabbay, A., Ephrat, A., Halperin, T., Peleg, S.: Seeing through noise: Visually driven speaker separation and enhancement. In: Proc. ICASSP. pp. 3051–3055. IEEE (2018)
2018
Cited alongside, same era.
Gao, R., Feris, R.S., Grauman, K.: Learning to separate object sounds by watching unlabeled video. In: Proc. ECCV (2018)
2018
Cited alongside, same era.
Harwath, D., Recasens, A., Surís, D., Chuang, G., Torralba, A., Glass, J.: Jointly discovering visual objects and spoken words from raw sensory input. In: Proceedings of the European conference on computer vision (ECCV). pp. 649–665 (2018)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Korbar, B., Tran, D., Torresani, L.: Co-training of audio and video representations from self-supervised temporal synchronization. CoRR (2018)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
Owens, A., Efros, A.A.: Audio-visual scene analysis with self-supervised multisensory features. Proc. ECCV (2018)
2018
Cited alongside, same era.
2019
Later among the works it cites.
Han, T., Xie, W., Zisserman, A.: Video representation learning by dense predictive coding. In: Workshop on Large Scale Holistic Video Understanding, ICCV (2019)
2019
Later among the works it cites.
Hu, D., Nie, F., Li, X.: Deep multimodal clustering for unsupervised audiovisual learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (June 2019)
2019
Later among the works it cites.
2019
Later among the works it cites.
Rouditchenko, A., Zhao, H., Gan, C., McDermott, J., Torralba, A.: Self-supervised audio-visual co-segmentation. In: Proc. ICASSP. pp. 2357–2361. IEEE (2019)
2019
Later among the works it cites.
Shahid, M., Beyan, C., Murino, V.: Voice activity detection by upper body motion analysis and unsupervised domain adaptation. In: The IEEE International Conference on Computer Vision (ICCV) Workshops (Oct 2019)
2019
Later among the works it cites.
2019
Later among the works it cites.
Zhao, H., Gan, C., Ma, W.C., Torralba, A.: The sound of motions. Proc. ICCV (2019)
2019
Later among the works it cites.
Chen, T., Kornblith, S., Norouzi, M., Hinton, G.: A simple framework for contrastive learning of visual representations. ICML (2020)
2020
Closest in time.
Han, T., Xie, W., Zisserman, A.: Memory-augmented dense predictive coding for video representation learning. In: ECCV (2020)
2020
Closest in time.
He, K., Fan, H., Wu, Y., Xie, S., Girshick, R.: Momentum contrast for unsupervised visual representation learning. CVPR (2020)
2020
Closest in time.
Hénaff, O.J., Srinivas, A., De Fauw, J., Razavi, A., Doersch, C., Eslami, S., Oord, A.v.d.: Data-efficient image recognition with contrastive predictive coding. ICML (2020)
2020
Closest in time.
2020
Closest in time.
Misra, I., van der Maaten, L.: Self-supervised learning of pretext-invariant representations. In: CVPR (2020)
2020
Closest in time.
Nagrani, A., Chung, J.S., Albanie, S., Zisserman, A.: Disentangled speech embeddings using cross-modal self-supervision. In: Proc. ICASSP. pp. 6829–6833. IEEE (2020)
2020
Closest in time.
Ramaswamy, J., Das, S.: See the sound, hear the pixels. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) (March 2020)
2020
Closest in time.