Fetching the paper…
Reading the bibliography…
We propose a self-supervised method for learning representations based on spatial audio-visual correspondences in egocentric videos.
Audio vision: Using audio-visual synchrony to locate sounds
John Hershey and Javier Movellan · 1999
Earlier work this paper cites.
Learning joint statistical models for audio-visual fusion and segregation
John W Fisher III, Trevor Darrell, William Freeman, and Paul Viola · 2000
Earlier work this paper cites.
Real-time speaker localization and speech separation by audio-visual integration
Kazuhiro Nakadai, Ken-ichi Hidai, Hiroshi G Okuno, and Hiroaki Kitano · 2002
Earlier work this paper cites.
Blind separation of speech mixtures via time-frequency masking
Özgür Yılmaz and Scott Rickard · 2004
Earlier work this paper cites.
Hello! my name is… buffy” – automatic naming of characters in tv video
Mark Everingham, Josef Sivic, and Andrew Zisserman · 2006
Earlier work this paper cites.
Supervised and semi-supervised separation of sounds from single-channel mixtures
Paris Smaragdis, Bhiksha Raj, and Madhusudana Shashanka · 2007
Earlier work this paper cites.
Monaural sound source separation by nonnegative matrix factorization with temporal continuity and sparseness criteria
Tuomas Virtanen · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Source-filter based clustering for monaural blind source separation
Martin Spiertz and Volker Gnann · 2009
Earlier work this paper cites.
Under-determined reverberant audio source separation using a full-rank spatial covariance model
Ngoc QK Duong, Emmanuel Vincent, and Rémi Gribonval · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and A. Ng · 2011
Earlier work this paper cites.
The kaldi speech recognition toolkit
Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, Jan Silovsky, Georg Stemmer, and Karel Vesely · 2011
Earlier work this paper cites.
Deep learning for monaural speech separation
P. Huang, M. Kim, M. Hasegawa-Johnson, and P. Smaragdis · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation
Olaf Ronneberger, Philipp Fischer, and Thomas Brox · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Look, listen and learn
Relja Arandjelović and Andrew Zisserman · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Motion informed audio source separation
S. Parekh, S. Essid, A. Ozerov, N. Q. K. Duong, P. Pérez, and G. Richard · 2017
Earlier work this paper cites.
Audio-visual object localization and separation using low-rank and sparsity
Jie Pu, Yannis Panagakis, Stavros Petridis, and Maja Pantic · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
The conversation: Deep audio-visual speech enhancement
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Cited alongside, same era.
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein · 2018
Cited alongside, same era.
Learning to separate object sounds by watching unlabeled video
Ruohan Gao, Rogerio Feris, and Kristen Grauman · 2018
Cited alongside, same era.
Cooperative learning of audio and video models from self-supervised synchronization
Bruno Korbar, Du Tran, and Lorenzo Torresani · 2018
Cited alongside, same era.
Audio-visual speech inpainting with deep learning
Giovanni Morrone, Daniel Michelsanti, Zheng-Hua Tan, and Jesper Jensen · 2021
Later among the works it cites.
Audio-visual floorplan reconstruction
Senthil Purushwalkam, Sebastia Vicenc Amengual Gari, Vamsi Krishna Ithapu, Carl Schissler, Philip Robinson, Abhinav Gupta, and Kristen Grauman · 2021
Later among the works it cites.
Localize to binauralize: Audio spatialization from visual sound source localization
Kranthi Kumar Rachavarapu, Aakanksha, Vignesh Sundaresha, and A. N. Rajagopalan · 2021
Later among the works it cites.
Is someone speaking? exploring long-term temporal features for audio-visual active speaker detection
Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian, Mike Zheng Shou, and Haizhou Li · 2021
Later among the works it cites.
Unicon: Unified context network for robust active speaker detection
Yuanhang Zhang, Susan Liang, Shuang Yang, Xiao Liu, Zhongqin Wu, Shiguang Shan, and Xilin Chen · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Self-supervised generation of spatial audio for 360°video
Pedro Morgado, Nuno Nvasconcelos, Timothy Langlois, and Oliver Wang · 2018
Cited alongside, same era.
Audio-visual scene analysis with self-supervised multisensory features
Andrew Owens and Alexei A Efros · 2018
Cited alongside, same era.
Sdr – half-baked or well done?
Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R. Hershey · 2018
Cited alongside, same era.
Charades-ego: A large-scale dataset of paired third and first person videos
Gunnar A. Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari · 2018
Cited alongside, same era.
Self-supervised moving vehicle tracking with stereo sound
Chuang Gan, Hang Zhao, Peihao Chen, David Cox, and Antonio Torralba · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Vision-infused deep audio inpainting
Hang Zhou, Ziwei Liu, Xudong Xu, Ping Luo, and Xiaogang Wang · 2019
Cited alongside, same era.
Mae-ast: Masked autoencoding audio spectrogram transformer
Alan Baade, Puyuan Peng, and David F. Harwath · 2022
Later among the works it cites.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Jian Ma, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, and Michael Wray · 2022
Later among the works it cites.
Masked autoencoders as spatiotemporal learners
Christoph Feichtenhofer, Yanghao Li, Kaiming He, et al · 2022
Later among the works it cites.
Audiovisual masked autoencoders
Mariana-Iuliana Georgescu, Eduardo Fonseca, Radu Tudor Ionescu, Mario Lucic, Cordelia Schmid, and Anurag Arnab · 2022
Later among the works it cites.
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
Egocentric deep multi-channel audio-visual active speaker localization
Hao Jiang, Calvin Murdock, and Vamsi Krishna Ithapu · 2022
Later among the works it cites.
Active audio-visual separation of dynamic sound sources
Sagnik Majumder and Kristen Grauman · 2022
Later among the works it cites.
Few-shot audio-visual learning of environment acoustics
Sagnik Majumder, Changan Chen, Ziad Al-Halah, and Kristen Grauman · 2022
Later among the works it cites.
Learning long-term spatial-temporal graphs for active speaker detection
Kyle Min, Sourya Roy, Subarna Tripathi, Tanaya Guha, and Somdeb Majumdar · 2022
Later among the works it cites.
Localizing visual sounds the easy way
Shentong Mo and Pedro Morgado · 2022
Later among the works it cites.
Self-supervised learning for audio-visual relationships of videos with stereo sounds
Tomoya Sato, Yusuke Sugano, and Yoichi Sato · 2022
Later among the works it cites.
VideoMAE: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang · 2022
Later among the works it cites.
Self-supervised learning of audio representations from audio-visual data using spatial alignment
Shanshan Wang, Archontis Politis, Annamaria Mesaros, and Tuomas Virtanen · 2022
Later among the works it cites.
Camera pose estimation and localization with active audio sensing
Karren Yang, Michael Firman, Eric Brachmann, and Clément Godard · 2022
Later among the works it cites.
Contrastive audio-visual masked autoencoder
Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R. Glass · 2023
Closest in time.
Epic-sounds: A large-scale dataset of actions that sound
Jaesung Huh, Jacob Chalk, Evangelos Kazakos, Dima Damen, and Andrew Zisserman · 2023
Closest in time.
Chat2map: Efficient scene mapping from multi-ego conversations
Sagnik Majumder, Hao Jiang, Pierre Moulon, Ethan Henderson, Paul Calamia, Kristen Grauman, and Vamsi Krishna Ithapu · 2023
Closest in time.
Egocentric auditory attention localization in conversations
Fiona Ryan, Hao Jiang, Abhinav Shukla, James M Rehg, and Vamsi Krishna Ithapu · 2023
Closest in time.