Fetching the paper…
Reading the bibliography…
Sounds originate from object motions and vibrations of surrounding air.
A computational approach to edge detection
J. Canny · 1986
Earlier work this paper cites.
The merging of the senses
B. E. Stein and M. A. Meredith · 1993
Earlier work this paper cites.
Audio vision: Using audio-visual synchrony to locate sounds
J. R. Hershey and J. R. Movellan · 2000
Earlier work this paper cites.
Learning joint statistical models for audio-visual fusion and segregation
J. W. Fisher III, T. Darrell, W. T. Freeman, and P. A. Viola · 2001
Earlier work this paper cites.
Non-negative matrix factorization for polyphonic music transcription
P. Smaragdis and J. C. Brown · 2003
Earlier work this paper cites.
The cocktail party problem
S. Haykin and Z. Chen · 2005
Earlier work this paper cites.
Pixels that sound
E. Kidron, Y. Y. Schechner, and M. Elad · 2005
Earlier work this paper cites.
On space-time interest points
I. Laptev · 2005
Earlier work this paper cites.
Harmony in motion
Z. Barzelay and Y. Y. Schechner · 2007
Earlier work this paper cites.
Monaural sound source separation by nonnegative matrix factorization with temporal continuity and sparseness criteria
T. Virtanen · 2007
Earlier work this paper cites.
A spatio-temporal descriptor based on 3d-gradients
A. Klaser, M. Marszałek, and C. Schmid · 2008
Earlier work this paper cites.
Nonnegative matrix and tensor factorizations: applications to exploratory multi-way data analysis and blind source separation
A. Cichocki, R. Zdunek, A. H. Phan, and S.-i. Amari · 2009
Earlier work this paper cites.
The cocktail party problem
J. H. McDermott · 2009
Earlier work this paper cites.
Multimodal deep learning
J. Ngiam, A. Khosla, M. Kim, J. Nam, H. Lee, and A. Y. Ng · 2011
Earlier work this paper cites.
Action recognition by dense trajectories
H. Wang, A. Kläser, C. Schmid, and C.-L. Liu · 2011
Earlier work this paper cites.
Multimodal analysis for identification and segmentation of moving-sounding objects
H. Izadinia, I. Saleemi, and M. Shah · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
H. Wang and C. Schmid · 2013
Earlier work this paper cites.
mir_eval: A transparent implementation of common mir metrics
C. Raffel, B. McFee, E. J. Humphrey, J. Salamon, O. Nieto, D. Liang, D. P. Ellis, and C. C. Raffel · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Cited alongside, same era.
U-net: Convolutional networks for biomedical image segmentation
O. Ronneberger, P. Fischer, and T. Brox · 2015
Cited alongside, same era.
Deep karaoke: Extracting vocals from musical mixtures using a convolutional deep neural network
A. J. Simpson, G. Roma, and M. D. Plumbley · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Cited alongside, same era.
Soundnet: Learning sound representations from unlabeled video
Y. Aytar, C. Vondrick, and A. Torralba · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Visual to sound: Generating natural sound for videos in the wild
Y. Zhou, Z. Wang, C. Fang, T. Bui, and T. L. Berg · 2017
Later among the works it cites.
A. Ephrat, I. Mosseri, O. Lang, T. Dekel, K. Wilson, A. Hassidim, W. T. Freeman, and M. Rubinstein · 2018
Later among the works it cites.
End-to-end learning of motion representation for video understanding
L. Fan, W. Huang, S. E. Chuang Gan, B. Gong, and J. Huang · 2018
Later among the works it cites.
Geometry-guided CNN for self-supervised video representation learning
C. Gan, B. Gong, K. Liu, H. Su, and L. J. Guibas · 2018
Later among the works it cites.
Learning to separate object sounds by watching unlabeled video
R. Gao, R. Feris, and K. Grauman · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep clustering: Discriminative embeddings for segmentation and separation
J. R. Hershey, Z. Chen, J. Le Roux, and S. Watanabe · 2016
Cited alongside, same era.
Ambient sound provides supervision for visual learning
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba · 2016
Cited alongside, same era.
Look, listen and learn
R. Arandjelovic and A. Zisserman · 2017
Cited alongside, same era.
R. Arandjelović and A. Zisserman · 2017
Cited alongside, same era.
Y. Bian, C. Gan, X. Liu, F. Li, X. Long, Y. Li, H. Qi, J. Zhou, S. Wen, and Y. Lin · 2017
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
R. Gao and K. Grauman · 2018
Later among the works it cites.
Co-training of audio and video representations from self-supervised temporal synchronization
B. Korbar, D. Tran, and L. Torresani · 2018
Later among the works it cites.
Attention clusters: Purely attention based local feature integration for video classification
X. Long, C. Gan, G. de Melo, J. Wu, X. Liu, and S. Wen · 2018
Later among the works it cites.
Self-supervised generation of spatial audio for 360 video
P. Morgado, N. Nvasconcelos, T. Langlois, and O. Wang · 2018
Later among the works it cites.
Seeing voices and hearing faces: Cross-modal biometric matching
A. Nagrani, S. Albanie, and A. Zisserman · 2018
Later among the works it cites.
Audio-visual scene analysis with self-supervised multisensory features
A. Owens and A. A. Efros · 2018
Later among the works it cites.
Learning to localize sound source in visual scenes
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. S. Kweon · 2018
Later among the works it cites.
Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume
D. Sun, X. Yang, M.-Y. Liu, and J. Kautz · 2018
Later among the works it cites.
Videos as space-time region graphs
X. Wang and A. Gupta · 2018
Later among the works it cites.
The sound of pixels
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba · 2018
Later among the works it cites.
Trajectory convolution for action recognition
Y. Zhao, Y. Xiong, and D. Lin · 2018
Later among the works it cites.
Creating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications
B. Li, X. Liu, K. Dinesh, Z. Duan, and G. Sharma · 2019
Closest in time.
Talking face generation by adversarially disentangled audio-visual representation
H. Zhou, Y. Liu, Z. Liu, P. Luo, and X. Wang · 2019
Closest in time.