Fetching the paper…
Reading the bibliography…
Visual events are usually accompanied by sounds in our daily lives.
B. F. Skinner, “”Superstition” in the pigeon.”
1948
Earlier work this paper cites.
B. Jones and B. Kabanoff, “Eye movements in auditory space perception,”
1975
Earlier work this paper cites.
B. R. Shelton and C. L. Searle, “The influence of vision on the absolute identification of sound-source position,”
1980
Earlier work this paper cites.
W. W. Gaver, “What in the world do we hear?: An ecological approach to auditory event perception,”
1993
Earlier work this paper cites.
D. R. Perrott, J. Cisneros, R. L. McKinley, and W. R. D’Angelo, “Aurally aided visual search under virtual and free-field listening conditions,”
1997
Earlier work this paper cites.
J. R. Hershey and J. R. Movellan, “Audio vision: Using audio-visual synchrony to locate sounds,” in
1999
Earlier work this paper cites.
R. S. Bolia, W. R. D’Angelo, and R. L. McKinley, “Aurally aided visual search in three-dimensional space,”
1999
Earlier work this paper cites.
J. W. Fisher III, T. Darrell, W. T. Freeman, and P. A. Viola, “Learning joint statistical models for audio-visual fusion and segregation,” in
2001
Earlier work this paper cites.
H. L. Van Trees,
2002
Earlier work this paper cites.
M. Corbetta and G. L. Shulman, “Control of goal-directed and stimulus-driven attention in the brain,”
2002
Earlier work this paper cites.
E. Kidron, Y. Y. Schechner, and M. Elad, “Pixels that sound,” in
2005
Earlier work this paper cites.
Z. Barzelay and Y. Y. Schechner, “Harmony in motion,” in
2007
Earlier work this paper cites.
B. E. Stein and T. R. Stanford, “Multisensory integration: current issues from the perspective of the single neuron,”
2008
Earlier work this paper cites.
P. Majdak, M. J. Goupell, and B. Laback, “3-d localization of virtual sound sources: Effects of visual environment, pointing method, and training,”
2010
Earlier work this paper cites.
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes (voc) challenge,”
2010
Earlier work this paper cites.
H. Izadinia, I. Saleemi, and M. Shah, “Multimodal analysis for identification and segmentation of moving-sounding objects,”
2013
Earlier work this paper cites.
A. Van den Oord, S. Dieleman, and B. Schrauwen, “Deep content-based music recommendation,” in
2013
Earlier work this paper cites.
S. Shalev-Shwartz and S. Ben-David,
2014
Earlier work this paper cites.
M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in
2014
Earlier work this paper cites.
A. Zunino, M. Crocco, S. Martelli, A. Trucco, A. Del Bue, and V. Murino, “Seeing the sound: A new multimodal imaging device for computer vision,” in
2015
Cited alongside, same era.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in
2015
Cited alongside, same era.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in
2015
Cited alongside, same era.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in
2015
Cited alongside, same era.
E. Hoffer and N. Ailon, “Deep metric learning using triplet network,” in
2015
Cited alongside, same era.
S.-H. Chou, Y.-C. Chen, K.-H. Zeng, H.-N. Hu, J. Fu, and M. Sun, “Self-view grounding given a narrated
2017
Later among the works it cites.
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. So Kweon, “Learning to localize sound source in visual scenes,” in
2018
Later among the works it cites.
R. Arandjelovic and A. Zisserman, “Objects that sound,” in
2018
Later among the works it cites.
R. Gao, R. Feris, and K. Grauman, “Learning to separate object sounds by watching unlabeled video,”
2018
Later among the works it cites.
H. Zhao, C. Gan, A. Rouditchenko, C. Vondrick, J. McDermott, and A. Torralba, “The sound of pixels,” in
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Abadi et al., “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015. [Online]. Available:
2015
Cited alongside, same era.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in
2015
Cited alongside, same era.
A. Owens, J. Wu, J. McDermott, W. Freeman, and A. Torralba, “Ambient sound provides supervision for visual learning,” in
2016
Cited alongside, same era.
Y. Aytar, C. Vondrick, and A. Torralba, “Soundnet: Learning sound representations from unlabeled video,” in
2016
Cited alongside, same era.
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” in
2016
Cited alongside, same era.
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L. Li, “Yfcc100m: The new data in multimedia research,” in
2016
Cited alongside, same era.
Y.-C. Su, D. Jayaraman, and K. Grauman, “Pano2vid: Automatic cinematography for watching
2016
Cited alongside, same era.
2018
Later among the works it cites.
D. Harwath, A. Recasens, D. Surís, G. Chuang, A. Torralba, and J. Glass, “Jointly discovering visual objects and spoken words from raw sensory input,”
2018
Later among the works it cites.
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,”
2018
Later among the works it cites.
Y. Tian, J. Shi, B. Li, Z. Duan, and C. Xu, “Audio-visual event localization in unconstrained videos,” in
2018
Later among the works it cites.
Y. Zhou, Z. Wang, C. Fang, T. Bui, and T. L. Berg, “Visual to sound: Generating natural sound for videos in the wild,” in
2018
Later among the works it cites.
C. Kim, H. V. Shin, T.-H. O. Oh, A. Kaspar, M. Elgharib, and W. Matusik, “On learning associations of faces and voices,” in
2018
Later among the works it cites.
A. Owens, J. Wu, J. H. McDermott, W. T. Freeman, and A. Torralba, “Learning sight from sound: Ambient sound provides supervision for visual learning,”
2018
Later among the works it cites.
B. Korbar, D. Tran, and L. Torresani, “Cooperative learning of audio and video models from self-supervised synchronization,”
2018
Later among the works it cites.
T. Afouras, J. S. Chung, and A. Zisserman, “The conversation: Deep audio-visual speech enhancement,” in
2018
Later among the works it cites.
H.-T. Cheng, C.-H. Chao, J.-D. Dong, H.-K. Wen, T.-L. Liu, and M. Sun, “Cube padding for weakly-supervised saliency prediction in
2018
Later among the works it cites.
P. Morgado, N. Vasconcelos, T. Langlois, and O. Wang, “Self-supervised generation of spatial audio for
2018
Later among the works it cites.
R. Gao and K. Grauman, “2.5d visual sound,” in
2019
Closest in time.