Fetching the paper…
Reading the bibliography…
Humans can robustly recognize and localize objects by integrating visual and auditory cues.
Wallach, H.: The role of head movements and vestibular and visual cues in sound localization. Journal of Experimental Psychology 27
1940
Earlier work this paper cites.
Thurlow, W.R., Mangels, J.W., Runge, P.S.: Head movements during sound localizationtd. The Journal of the Acoustical Society of America 489
1967
Earlier work this paper cites.
Griffin, D., Jae Lim: Signal estimation from modified short-time fourier transform. In: ICASSP ’83. IEEE International Conference on Acoustics, Speech, and Signal Processing. vol. 8, pp. 804–807 (April 1983). https://doi.org/10.1109/ICASSP.1983.1172092
1983
Earlier work this paper cites.
Fendrich, R.: The merging of the senses. Journal of Cognitive Neuroscience 5
1993
Earlier work this paper cites.
Computational auditory scene analysis. Computer Speech and Language 8
1994
Earlier work this paper cites.
Rosenblum, L.D., Gordon, M.S., Jarquin, L.: Echolocating distance by moving and stationary listeners. Ecological Psychology 12
2000
Earlier work this paper cites.
Ernst, M.O., Bülthoff, H.H.: Merging the senses into a robust percept. Trends in Cognitive Sciences 8
2004
Earlier work this paper cites.
Barzelay, Z., Schechner, Y.Y.: Harmony in motion. In: IEEE Conference on Computer Vision and Pattern Recognition (2007)
2007
Earlier work this paper cites.
Russell, B.C., Torralba, A., Murphy, K.P., Freeman, W.T.: Labelme: a database and web-based tool for image annotation. International journal of computer vision 77
2008
Earlier work this paper cites.
Urmson, C., Anhalt, J., Bae, H., Bagnell, J.A.D., Baker, C.R., Bittner, R.E., Brown, T., Clark, M.N., Darms, M., Demitrish, D., Dolan, J.M., Duggins, D., Ferguson, D., Galatali, T., Geyer, C.M., Gittleman, M., Harbaugh, S., Hebert, M., Howard, T., Kolski, S., Likhachev, M., Litkouhi, B., Kelly, A., McNaughton, M., Miller, N., Nickolaou, J., Peterson, K., Pilnick, B., Rajkumar, R., Rybski, P., Sadekar, V., Salesky, B., Seo, Y.W., Singh, S., Snider, J.M., Struble, J.C., Stentz, A.T., Taylor, M., Whittaker, W.R.L., Wolkowicki, Z., Zhang, W., Ziglar, J.: Autonomous driving in urban environments: Boss and the urban challenge. Journal of Field Robotics Special Issue on the 2007 DARPA Urban Challenge, Part I 25
2008
Earlier work this paper cites.
Fazenda, B., Hidajat Atmoko, Fengshou Gu, Luyang Guan, Ball, A.: Acoustic based safety emergency vehicle detection for intelligent transport systems. In: ICCAS-SICE (2009)
2009
Earlier work this paper cites.
Saxena, A., Ng, A.Y.: Learning sound location from a single microphone. In: IEEE International Conference on Robotics and Automation (2009)
2009
Earlier work this paper cites.
Brutzer, S., Höferlin, B., Heidemann, G.: Evaluation of background subtraction techniques for video surveillance. In: CVPR (2011)
2011
Earlier work this paper cites.
Antonacci, F., Filos, J., Thomas, M.R.P., Habets, E.A.P., Sarti, A., Naylor, P.A., Tubaro, S.: Inference of room geometry from acoustic impulse responses. IEEE Transactions on Audio, Speech, and Language Processing 20
2012
Earlier work this paper cites.
Huang, W., Alem, L., Livingston, M.A.: Human factors in augmented reality environments. Springer Science and Business Media (2012)
2012
Earlier work this paper cites.
Dokmanic, I., Parhizkar, R., Walther, A., Lu, Y.M., Vetterli, M.: Acoustic echoes reveal room shape. Proceedings of the National Academy of Sciences 110
2013
Earlier work this paper cites.
Geiger, A., Lenz, P., Stiller, C., Urtasun, R.: Vision meets robotics: The kitti dataset. International Journal of Robotics Research (IJRR) (2013)
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
McAnally, K.I., Martin, R.L.: Sound localization with head movement: implications for 3-d audio displays. Frontiers in Neuroscience 210
2014
Cited alongside, same era.
Salamon, J., Jacoby, C., Bello, J.P.: A dataset and taxonomy for urban sound research. In: ACM International Conference on Multimedia (2014)
2014
Cited alongside, same era.
Tiete, J., Domínguez, F., da Silva, B., Segers, L., Steenhaut, K., Touhafi, A.: Soundcompass: A distributed mems microphone array-based sensor for sound source localization. In: Sensors (2014)
2014
Cited alongside, same era.
Argentieri, S., Danès, P., Souères, P.: A survey on sound source localization in robotics: From binaural to array processing methods. Computer Speech and Language 34
2015
Cited alongside, same era.
Wang, P., Shen, X., Lin, Z., Cohen, S., Price, B., Yuille, A.L.: Towards unified depth and semantic prediction from a single image. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 2800–2809 (2015)
Arandjelović, R., Zisserman, A.: Objects that sound. In: ECCV (2018)
2018
Later among the works it cites.
Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., Adam, H.: Encoder-decoder with atrous separable convolution for semantic image segmentation. In: Proceedings of the European conference on computer vision (ECCV). pp. 801–818 (2018)
2018
Later among the works it cites.
Gao, R., Feris, R., Grauman, K.: Learning to separate object sounds by watching unlabeled video. In: ECCV (2018)
2018
Later among the works it cites.
Hecker, S., Dai, D., Van Gool, L.: End-to-end learning of driving models with surround-view cameras and route planners. In: The European Conference on Computer Vision (ECCV) (September 2018)
2018
Later among the works it cites.
Li, D., Langlois, T.R., Zheng, C.: Scene-aware audio for 360° videos. ACM Trans. Graph. 37
2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
Ye, M., Yu Zhang, Yang, R., Manocha, D.: 3d reconstruction in the presence of glasses by acoustic and stereo fusion. In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2015)
2015
Cited alongside, same era.
Aytar, Y., Vondrick, C., Torralba, A.: Soundnet: Learning sound representations from unlabeled video. In: Advances in Neural Information Processing Systems (2016)
2016
Cited alongside, same era.
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., Schiele, B.: The cityscapes dataset for semantic urban scene understanding. In: Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2016)
2016
Cited alongside, same era.
Mousavian, A., Pirsiavash, H., Košecká, J.: Joint semantic segmentation and depth estimation with deep convolutional networks. In: 2016 Fourth International Conference on 3D Vision (3DV). pp. 611–619. IEEE (2016)
2016
Cited alongside, same era.
Owens, A., Isola, P., McDermott, J., Torralba, A., Adelson, E.H., Freeman, W.T.: Visually indicated sounds. In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2016)
2016
Cited alongside, same era.
Owens, A., Wu, J., McDermott, J.H., Freeman, W.T., Torralba, A.: Ambient sound provides supervision for visual learning. In: Leibe, B., Matas, J., Sebe, N., Welling, M. (eds.) ECCV (2016)
2016
Cited alongside, same era.
Arandjelovic, R., Zisserman, A.: Look, listen and learn. In: The IEEE International Conference on Computer Vision (ICCV) (2017)
2017
Cited alongside, same era.
Later among the works it cites.
Owens, A., Efros, A.A.: Audio-visual scene analysis with self-supervised multisensory features. In: The European Conference on Computer Vision (ECCV) (September 2018)
2018
Later among the works it cites.
Senocak, A., Oh, T.H., Kim, J., Yang, M.H., So Kweon, I.: Learning to localize sound source in visual scenes. In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (June 2018)
2018
Later among the works it cites.
Zhao, H., Gan, C., Rouditchenko, A., Vondrick, C., McDermott, J., Torralba, A.: The sound of pixels. In: The European Conference on Computer Vision (ECCV) (September 2018)
2018
Later among the works it cites.
Seeing and Hearing Egocentric Actions: How Much Can We Learn? (2019)
2019
Later among the works it cites.
Delmerico, J., Mintchev, S., Giusti, A., Gromov, B., Melo, K., Horvat, T., Cadena, C., Hutter, M., Ijspeert, A., Floreano, D., Gambardella, L., Siegwart, R., Scaramuzza, D.: The current state and future outlook of rescue robotics. Journal of Field Robotics 36
2019
Later among the works it cites.
Gan, C., Zhao, H., Chen, P., Cox, D., Torralba, A.: Self-supervised moving vehicle tracking with stereo sound. In: The IEEE International Conference on Computer Vision (ICCV) (October 2019)
2019
Later among the works it cites.
Gao, R., Grauman, K.: 2.5 d visual sound. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 324–333 (2019)
2019
Later among the works it cites.
Gao, R., Grauman, K.: Co-separating sounds of visual objects. In: The IEEE International Conference on Computer Vision (ICCV) (October 2019)
2019
Later among the works it cites.
Godard, C., Aodha, O.M., Firman, M., Brostow, G.J.: Digging into self-supervised monocular depth estimation. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 3828–3838 (2019)
2019
Later among the works it cites.
Irie, G., Ostrek, M., Wang, H., Kameoka, H., Kimura, A., Kawanishi, T., Kashino, K.: Seeing through sounds: Predicting visual semantic segmentation results from multichannel audio signals. In: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 3961–3964. IEEE (2019)
2019
Later among the works it cites.
Simeoni, M.M.J.A., Kashani, S., Hurley, P., Vetterli, M.: Deepwave: A recurrent neural-network for real-time acoustic imaging p. 38 (2019)
2019
Later among the works it cites.
Zhao, H., Gan, C., Ma, W.C., Torralba, A.: The sound of motions. In: The IEEE International Conference on Computer Vision (ICCV) (October 2019)
2019
Later among the works it cites.