Fetching the paper…
Reading the bibliography…
Moving around in the world is naturally a multisensory experience, but today's embodied agents are deaf---restricted to solely their visual perception of the environment.
1903
Earlier work this paper cites.
Tolman, E.C.: Cognitive maps in rats and men. Psychological review (1948)
1948
Earlier work this paper cites.
Egan, M.D., Quirt, J., Rousseau, M.: Architectural acoustics (1989)
1989
Earlier work this paper cites.
Kuttruff, K.H.: Auralization of impulse responses modeled on the basis of ray-tracing results. Journal of the Audio Engineering Society (1993)
1993
Earlier work this paper cites.
Veach, E., Guibas, L.: Bidirectional estimators for light transport. In: Photorealistic Rendering Techniques (1995)
1995
Earlier work this paper cites.
Thinus-Blanc, C., Gaunet, F.: Representation of space in blind persons: vision as a spatial sense? Psychological bulletin (1997)
1997
Earlier work this paper cites.
Lessard, N., Paré, M., Lepore, F., Lassonde, M.: Early-blind human subjects localize sound sources better than sighted subjects. Nature (1998)
1998
Earlier work this paper cites.
Nakadai, K., Nakamura, K.: Sound source localization and separation. Wiley Encyclopedia of Electrical and Electronics Engineering (1999)
1999
Earlier work this paper cites.
RoÈder, B., Teder-SaÈlejaÈrvi, W., Sterr, A., RoÈsler, F., Hillyard, S.A., Neville, H.J.: Improved auditory spatial tuning in blind humans. Nature (1999)
1999
Earlier work this paper cites.
Hershey, J.R., Movellan, J.R.: Audio vision: Using audio-visual synchrony to locate sounds. In: NeurIPS (2000)
2000
Earlier work this paper cites.
Nakadai, K., Lourens, T., Okuno, H.G., Kitano, H.: Active audition for humanoid. In: AAAI (2000)
2000
Earlier work this paper cites.
Nakadai, K., Okuno, H.G., Kitano, H.: Epipolar geometry based sound localization and extraction for humanoid audition. In: IROS Workshops. IEEE (2001)
2001
Earlier work this paper cites.
Daniel, J.: Spatial sound encoding including near field effect: Introducing distance coding filters and a viable, new ambisonic format. In: Audio Engineering Society Conference: 23rd International Conference: Signal Processing in Audio Recording and Reproduction. Audio Engineering Society (2003)
2003
Earlier work this paper cites.
Wood, J., Magennis, M., Arias, E.F.C., Gutierrez, T., Graupp, H., Bergamasco, M.: The design and evaluation of a computer game for the blind in the grab haptic audio virtual environment. Proceedings of Eurohpatics (2003)
2003
Earlier work this paper cites.
Hartley, R., Zisserman, A.: Multiple view geometry in computer vision. Cambridge University Press (2004)
2004
Earlier work this paper cites.
Voss, P., Lassonde, M., Gougoux, F., Fortin, M., Guillemot, J.P., Lepore, F.: Early-and late-onset blind individuals show supra-normal auditory abilities in far-space. Current Biology (2004)
2004
Earlier work this paper cites.
Gougoux, F., Zatorre, R.J., Lassonde, M., Voss, P., Lepore, F.: A functional neuroimaging study of sound localization: visual cortex activity predicts performance in early-blind individuals. PLoS biology (2005)
2005
Earlier work this paper cites.
Thrun, S., Burgard, W., Fox, D.: Probabilistic robotics. MIT Press (2005)
2005
Earlier work this paper cites.
Qin, J., Cheng, J., Wu, X., Xu, Y.: A learning based approach to audio surveillance in household environment. International Journal of Information Acquisition (2006)
2006
Earlier work this paper cites.
Fortin, M., Voss, P., Lord, C., Lassonde, M., Pruessner, J., Saint-Amour, D., Rainville, C., Lepore, F.: Wayfinding in the blind: larger hippocampal volume and supranormal spatial navigation. Brain (2008)
2008
Earlier work this paper cites.
van der Maaten, L., Hinton, G.: Visualizing high-dimensional data using t-sne. Journal of Machine Learning Research 9
2008
Earlier work this paper cites.
Merabet, L., Sanchez, J.: Audio-based navigation using virtual environments: combining technology and neuroscience. AER Journal: Research and Practice in Visual Impairment and Blindness (2009)
2009
Earlier work this paper cites.
Wu, X., Gong, H., Chen, P., Zhong, Z., Xu, Y.: Surveillance robot utilizing video and audio information. Journal of Intelligent and Robotic Systems (2009)
2009
Earlier work this paper cites.
Yoshida, T., Nakadai, K., Okuno, H.G.: Automatic speech recognition improved by two-layered audio-visual integration for robot audition. In: 2009 9th IEEE-RAS International Conference on Humanoid Robots. pp. 604–609. IEEE (2009)
2009
Earlier work this paper cites.
Merabet, L.B., Pascual-Leone, A.: Neural reorganization following sensory loss: the opportunity of change. Nature Reviews Neuroscience (2010)
2010
Earlier work this paper cites.
Georgiev, I.: Implementing vertex connection and merging. Technical Re-port. Saarland University. (2012)
2012
Earlier work this paper cites.
Connors, E.C., Yazzolino, L.A., Sánchez, J., Merabet, L.B.: Development of an audio-based virtual gaming environment to assist with navigation skills in the blind. JoVE (Journal of Visualized Experiments) (2013)
2013
Earlier work this paper cites.
Romano, J.M., Brindza, J.P., Kuchenbecker, K.J.: Ros open-source audio recognizer: Roar environmental sound detection tools for robot programming. Autonomous robots (2013)
2013
Earlier work this paper cites.
Wymann, B., Espié, E., Guionneau, C., Dimitrakakis, C., Coulom, R., Sumner, A.: Torcs, the open racing car simulator (2013), http://www.torcs.org
2013
Earlier work this paper cites.
Picinali, L., Afonso, A., Denis, M., Katz, B.: Exploration of architectural spaces by blind people using auditory virtual reality for the construction of spatial knowledge. International Journal of Human-Computer Studies 72
2014
Earlier work this paper cites.
Viciana-Abad, R., Marfil, R., Perez-Lorenzo, J., Bandera, J., Romero-Garces, A., Reche-Lopez, P.: Audio-visual perception system for a humanoid robotic head. Sensors (2014)
2014
Earlier work this paper cites.
Wang, Y., Kapadia, M., Huang, P., Kavan, L., Badler, N.: Sound localization and multi-modal steering for autonomous virtual agents. In: Symposium on Interactive 3D Graphics and Games (2014)
2014
Earlier work this paper cites.
Alameda-Pineda, X., Horaud, R.: Vision-guided robot hearing. The International Journal of Robotics Research (2015)
2015
Earlier work this paper cites.
Alameda-Pineda, X., Staiano, J., Subramanian, R., Batrinca, L., Ricci, E., Lepri, B., Lanz, O., Sebe, N.: Salsa: A novel dataset for multimodal group behavior analysis. IEEE transactions on pattern analysis and machine intelligence (2015)
2015
Earlier work this paper cites.
Chung, J., Kastner, K., Dinh, L., Goel, K., Courville, A.C., Bengio, Y.: A recurrent latent variable model for sequential data. In: NeurIPS (2015)
2015
Earlier work this paper cites.
Ekstrom, A.D.: Why vision is important to how we navigate. Hippocampus (2015)
2015
Earlier work this paper cites.
Gebru, I.D., Ba, S., Evangelidis, G., Horaud, R.: Tracking the active speaker based on a joint audio-visual observation model. In: Proceedings of the IEEE International Conference on Computer Vision Workshops. pp. 15–21 (2015)
2015
Earlier work this paper cites.
Rafaely, B.: Fundamentals of Spherical Array Processing. Springer (2015)
2015
Earlier work this paper cites.
Savioja, L., Svensson, U.P.: Overview of geometrical room acoustic modeling techniques. The Journal of the Acoustical Society of America (2015)
2015
Cited alongside, same era.
2015
Cited alongside, same era.
Ammirato, P., Poirson, P., Park, E., Kosecka, J., Berg, A.: A dataset for developing and benchmarking active vision. In: ICRA (2016)
2016
Cited alongside, same era.
Cao, C., Ren, Z., Schissler, C., Manocha, D., Zhou, K.: Interactive sound propagation with bidirectional path tracing. ACM Transactions on Graphics (TOG) (2016)
2016
Cited alongside, same era.
Johnson, M., Hofmann, K., Hutton, T., Bignell, D.: The malmo platform for artificial intelligence experimentation. In: Intl. Joint Conference on AI (2016)
Haarnoja, T., Zhou, A., Abbeel, P., Levine, S.: Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In: ICML (2018)
2018
Later among the works it cites.
Henriques, J.F., Vedaldi, A.: Mapnet: An allocentric spatial memory for mapping environments. In: CVPR (2018)
2018
Later among the works it cites.
Jayaraman, D., Grauman, K.: End-to-end policy learning for active visual categorization. TPAMI (2018)
2018
Later among the works it cites.
Locher, B., Piquerez, A., Habermacher, M., Ragettli, M., Röösli, M., Brink, M., Cajochen, C., Vienneau, D., Foraster, M., Müller, U., et al.: Differences between outdoor and indoor sound levels for open, tilted, and closed windows. International journal of environmental research and public health (2018)
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Kempka, M., Wydmuch, M., Runc, G., Toczek, J., Jakowski, W.: Vizdoom: A doom-based ai research platform for visual reinforce- ment learning. In: Proc. IEEE Conf. on Computational Intelligence and Games (2016)
2016
Cited alongside, same era.
Kuttruff, H.: Room acoustics. CRC Press (2016)
2016
Cited alongside, same era.
Lerer, A., Gross, S., Fergus, R.: Learning physical intuition of block towers by example. In: ICML (2016)
2016
Cited alongside, same era.
Owens, A., Isola, P., McDermott, J., Torralba, A., Adelson, E.H., Freeman, W.T.: Visually indicated sounds. In: CVPR (2016)
2016
Cited alongside, same era.
Owens, A., Wu, J., McDermott, J.H., Freeman, W.T., Torralba, A.: Ambient sound provides supervision for visual learning. In: ECCV (2016)
2016
Cited alongside, same era.
Yusuf Aytar, Carl Vondrick, A.T.: Learning sound representations from unlabeled video. In: NeurIPS (2016)
2016
Cited alongside, same era.
Armeni, I., Sax, A., Zamir, A.R., Savarese, S.: Joint 2D-3D-Semantic Data for Indoor Scene Understanding. ArXiv e-prints (Feb 2017)
2017
Cited alongside, same era.
2018
Later among the works it cites.
Morgado, P., Nvasconcelos, N., Langlois, T., Wang, O.: Self-supervised generation of spatial audio for 360 video. In: NeurIPS (2018)
2018
Later among the works it cites.
Owens, A., Efros, A.A.: Audio-visual scene analysis with self-supervised multisensory features. In: ECCV (2018)
2018
Later among the works it cites.
Savinov, N., Dosovitskiy, A., Koltun, V.: Semi-parametric topological memory for navigation. In: ICLR (2018)
2018
Later among the works it cites.
Senocak, A., Oh, T.H., Kim, J., Yang, M.H., So Kweon, I.: Learning to localize sound source in visual scenes. In: CVPR (2018)
2018
Later among the works it cites.
Tian, Y., Shi, J., Li, B., Duan, Z., Xu, C.: Audio-visual event localization in unconstrained videos. In: ECCV (2018)
2018
Later among the works it cites.
Xia, F., Zamir, A.R., He, Z., Sax, A., Malik, J., Savarese, S.: Gibson env: Real-world perception for embodied agents. In: CVPR (2018)
2018
Later among the works it cites.
Zaunschirm, M., Schörkhuber, C., Höldrich, R.: Binaural rendering of ambisonic signals by head-related impulse response time alignment and a diffuseness constraint. The Journal of the Acoustical Society of America (2018)
2018
Later among the works it cites.
Zhao, H., Gan, C., Rouditchenko, A., Vondrick, C., McDermott, J., Torralba, A.: The sound of pixels. In: ECCV (2018)
2018
Later among the works it cites.
Zhou, Y., Wang, Z., Fang, C., Bui, T., Berg, T.L.: Visual to sound: Generating natural sound for videos in the wild. In: CVPR (2018)
2018
Later among the works it cites.
Chen, H., Suhr, A., Misra, D., Snavely, N., Artzi, Y.: Touchdown: Natural language navigation and spatial reasoning in visual street environments. In: CVPR (2019)
2019
Closest in time.
Gao, R., Grauman, K.: 2.5 d visual sound. In: CVPR (2019)
2019
Closest in time.
Gao, R., Grauman, K.: Co-separating sounds of visual objects. In: ICCV (2019)
2019
Closest in time.
Gordon, D., Kadian, A., Parikh, D., Hoffman, J., Batra, D.: Splitnet: Sim2sim and task2task transfer for embodied visual navigation. ICCV (2019)
2019
Closest in time.
Jain, U., Weihs, L., Kolve, E., Rastegari, M., Lazebnik, S., Farhadi, A., Schwing, A.G., Kembhavi, A.: Two body problem: Collaborative visual task completion. In: CVPR (2019), first two authors contributed equally
2019
Closest in time.
2019
Closest in time.
Manolis Savva*, Abhishek Kadian*, Oleksandr Maksymets*, Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., Parikh, D., Batra, D.: Habitat: A Platform for Embodied AI Research. In: ICCV (2019)
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
Thomason, J., Gordon, D., Bisk, Y.: Shifting the baseline: Single modality performance on visual navigation & qa. In: NAACL-HLT (2019)
2019
Closest in time.
Wijmans, E., Datta, S., Maksymets, O., Das, A., Gkioxari, G., Lee, S., Essa, I., Parikh, D., Batra, D.: Embodied Question Answering in Photorealistic Environments with Point Cloud Perception. In: CVPR (2019)
2019
Closest in time.
Wortsman, M., Ehsani, K., Rastegari, M., Farhadi, A., Mottaghi, R.: Learning to learn how to learn: Self-adaptive visual navigation using meta-learning. In: CVPR (2019)
2019
Closest in time.
2019
Closest in time.
Wu, Y., Wu, Y., Tamar, A., Russell, S., Gkioxari, G., Tian, Y.: Bayesian relational memory for semantic visual navigation. ICCV (2019)
2019
Closest in time.
2019
Closest in time.
Zotter, F., Frank, M.: Ambisonics: A Practical 3D Audio Theory for Recording, Studio Production, Sound Reinforcement and Virtual Reality. Springer (2019)
2019
Closest in time.
Chaplot, D.S., Gupta, S., Gupta, A., Salakhutdinov, R.: Learning to explore using active neural mapping. In: ICLR (2020)
2020
Closest in time.
Das, A., Carnevale, F., Merzic, H., Rimell, L., Schneider, R., Abramson, J., Hung, A., Ahuja, A., Clark, S., Wayne, G., et al.: Probing emergent semantics in predictive agents via question answering. In: ICML (2020)
2020
Closest in time.
Gan, C., Zhang, Y., Wu, J., Gong, B., Tenenbaum, J.: Look, listen, and act: Towards audio-visual embodied navigation. In: ICRA (2020)
2020
Closest in time.
Gao, R., Chen, C., Al-Halah, Z., Schissler, C., Grauman, K.: VisualEchoes: Spatial image representation learning through echolocation. In: ECCV (2020)
2020
Closest in time.
Jain, U., Weihs, L., Kolve, E., Farhadi, A., Lazebnik, S., Kembhavi, A., Schwing, A.G.: A cordial sync: Going beyond marginal policies for multi-agent embodied tasks. In: ECCV (2020), first two authors contributed equally
2020
Closest in time.
Wijmans, E., Kadian, A., Morcos, A., Lee, S., Essa, I., Parikh, D., Savva, M., Batra, D.: Decentralized distributed ppo: Solving pointgoal navigation. In: ICLR (2020)
2020
Closest in time.