Fetching the paper…
Reading the bibliography…
In audio-visual navigation, an agent intelligently travels through a complex, unmapped 3D environment using both sights and sounds to find a sound source (e.g., a phone ringing in another room).
Early-blind human subjects localize sound sources better than sighted subjects
Nadia Lessard, Michael Paré, Franco Lepore, and Maryse Lassonde · 1998
Earlier work this paper cites.
Sound source localization and separation
Kazuhiro Nakadai and Keisuke Nakamura · 1999
Earlier work this paper cites.
Audio vision: Using audio-visual synchrony to locate sounds
John R Hershey and Javier R Movellan · 2000
Earlier work this paper cites.
Active audition for humanoid
Kazuhiro Nakadai, Tino Lourens, Hiroshi G Okuno, and Hiroaki Kitano · 2000
Earlier work this paper cites.
Probabilistic robotics
Sebastian Thrun · 2002
Earlier work this paper cites.
A functional neuroimaging study of sound localization: visual cortex activity predicts performance in early-blind individuals
Frédéric Gougoux, Robert J Zatorre, Maryse Lassonde, Patrice Voss, and Franco Lepore · 2005
Earlier work this paper cites.
Visual simultaneous localization and mapping: a survey
Jorge Fuentes-Pacheco, José Ruíz Ascencio, and Juan M. Rendón-Mancha · 2012
Earlier work this paper cites.
Acoustic echoes reveal room shape
I. Dokmanic, R. Parhizkar, A. Walther, Y. Lu, and M. Vetterli · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Sound localization and multi-modal steering for autonomous virtual agents
Y. Wang, M. Kapadia, P. Huang, L. Kavan, and N. Badler · 2014
Earlier work this paper cites.
Robust reconstruction of indoor scenes
Sungjoon Choi, Qian-Yi Zhou, and Vladlen Koltun · 2015
Earlier work this paper cites.
A recurrent latent variable model for sequential data
Junyoung Chung, Kyle Kastner, Laurent Dinh, Kratarth Goel, Aaron C Courville, and Yoshua Bengio · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Deep learning of structured environments for robot search
J. A. Caley, N. R. J. Lawrance, and G. A. Hollinger · 2016
Earlier work this paper cites.
Learning to navigate in complex environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andrew J Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, et al · 2016
Earlier work this paper cites.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup · 2017
Earlier work this paper cites.
Matterport3d: Learning from rgb-d data in indoor environments
Angel Chang, Angela Dai, Thomas Funkhouser, Maciej Halber, Matthias Niessner, Manolis Savva, Shuran Song, Andy Zeng, and Yinda Zhang · 2017
Cited alongside, same era.
Localization of sound sources in robotics: A review
Caleb Rascon and Ivan Meza · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
On evaluation of embodied navigation agents
Peter Anderson, Angel Chang, Devendra Singh Chaplot, Alexey Dosovitskiy, Saurabh Gupta, Vladlen Koltun, Jana Kosecka, Jitendra Malik, Roozbeh Mottaghi, Manolis Savva, et al · 2018
Cited alongside, same era.
Objects that sound
Relja Arandjelovic and Andrew Zisserman · 2018
Cited alongside, same era.
Scene memory transformer for embodied agents in long-horizon tasks
Kuan Fang, Alexander Toshev, Li Fei-Fei, and Silvio Savarese · 2019
Later among the works it cites.
Infobot: Transfer and exploration via the information bottleneck
Anirudh Goyal, Riashat Islam, Daniel Strouse, Zafarali Ahmed, Matthew Botvinick, Hugo Larochelle, Yoshua Bengio, and Sergey Levine · 2019
Later among the works it cites.
Learning multi-level hierarchies with hindsight
Andrew Levy, George Konidaris, Robert Platt, and Kate Saenko · 2019
Later among the works it cites.
Hrl4in: Hierarchical reinforcement learning for interactive navigation with mobile manipulators
Chengshu Li, Fei Xia, Roberto Martin-Martin, and Silvio Savarese · 2019
Later among the works it cites.
Benchmarking classic and learned navigation in complex 3d environments
Dmytro Mishkin, Alexey Dosovitskiy, and Vladlen Koltun · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mapnet: An allocentric spatial memory for mapping environments
Joao F Henriques and Andrea Vedaldi · 2018
Cited alongside, same era.
Driving policy transfer via modularity and abstraction
Matthias Müller, Alexey Dosovitskiy, Bernard Ghanem, and Vladlen Koltun · 2018
Cited alongside, same era.
Data-efficient hierarchical reinforcement learning
Ofir Nachum, Shixiang Shane Gu, Honglak Lee, and Sergey Levine · 2018
Cited alongside, same era.
Semi-parametric topological memory for navigation
Nikolay Savinov, Alexey Dosovitskiy, and Vladlen Koltun · 2018
Cited alongside, same era.
Alexander Sax, Bradley Emi, Amir R Zamir, Leonidas Guibas, Silvio Savarese, and Jitendra Malik · 2018
Cited alongside, same era.
Learning to localize sound source in visual scenes
Arda Senocak, Tae-Hyun Oh, Junsik Kim, Ming-Hsuan Yang, and In So Kweon · 2018
Cited alongside, same era.
Learning over subgoals for efficient navigation of structured, unknown environments
Gregory J. Stein, Christopher Bradley, and Nicholas Roy · 2018
Cited alongside, same era.
Habitat: A Platform for Embodied AI Research
Manolis Savva, Abhishek Kadian, Oleksandr Maksymets, Yili Zhao, Erik Wijmans, Bhavana Jain, Julian Straub, Jia Liu, Vladlen Koltun, Jitendra Malik, Devi Parikh, and Dhruv Batra · 2019
Later among the works it cites.
The replica dataset: A digital replica of indoor spaces
Julian Straub, Thomas Whelan, Lingni Ma, Yufan Chen, Erik Wijmans, Simon Green, Jakob J Engel, Raul Mur-Artal, Carl Ren, Shobhit Verma, et al · 2019
Later among the works it cites.
Learning spatial common sense with geometry-aware recurrent networks
Hsiao-Yu Fish Tung, Ricson Cheng, and Katerina Fragkiadaki · 2019
Later among the works it cites.
SoundSpaces: Audio-visual navigation in 3D environments
Changan Chen, Unnat Jain, Carl Schissler, Sebastia Vicenc Amengual Gari, Ziad Al-Halah, Vamsi Krishna Ithapu, Philip Robinson, and Kristen Grauman · 2020
Closest in time.
Batvision: Learning to see 3d spatial layout with two ears
Jesper Christensen, Sascha Hornauer, and Stella Yu · 2020
Closest in time.
Look, listen, and act: Towards audio-visual embodied navigation
Chuang Gan, Yiwei Zhang, Jiajun Wu, Boqing Gong, and Joshua B Tenenbaum · 2020
Closest in time.
VisualEchoes: Spatial image representation learning through echolocation
Ruohan Gao, Changan Chen, Ziad Al-Halah, Carl Schissler, and Kristen Grauman · 2020
Closest in time.
Ego-topo: Environment affordances from egocentric video
Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, and Kristen Grauman · 2020
Closest in time.
Hierarchical foresight: Self-supervised learning of long-horizon tasks via visual subgoal generation
Suraj Nair and Chelsea Finn · 2020
Closest in time.
Spatial action maps for mobile manipulation
Jimmy Wu, Xingyuan Sun, Andy Zeng, Shuran Song, Johnny Lee, Szymon Rusinkiewicz, and Thomas Funkhouser · 2020
Closest in time.