Fetching the paper…
Reading the bibliography…
We introduce the novel task of interactive scene exploration, wherein robots autonomously explore environments and produce an action-conditioned scene graph (ACSG) that captures the structure of the underlying environment.
In Proc. of IEEE Workshop on Computer Vision , 1985
Active perception vs. passive perception · 1985
Earlier work this paper cites.
Active perception
R. Bajcsy · 1988
Earlier work this paper cites.
Active perception and reinforcement learning
S. D. Whitehead and D. H. Ballard · 1990
Earlier work this paper cites.
A frontier-based approach for autonomous exploration
B. Yamauchi · 1997
Earlier work this paper cites.
Information gain-based exploration using rao-blackwellized particle filters
C. Stachniss, G. Grisetti, and W. Burgard · 2005
Earlier work this paper cites.
Active sensor planning for multiview vision tasks
S. Chen, Y. F. Li, W. Wang, and J. Zhang · 2008
Earlier work this paper cites.
Active perception: Interactive manipulation for improving object detection
Q. V. Le, A. Saxena, and A. Y. Ng · 2008
Earlier work this paper cites.
Information-theoretic planning with trajectory optimization for dense 3d mapping
B. Charrow, G. Kahn, S. Patil, S. Liu, K. Goldberg, P. Abbeel, N. Michael, and V. Kumar · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bernstein, and L. Fei-Fei · 2015
Earlier work this paper cites.
Receding horizon "next-best-view" planner for 3d exploration
A. Bircher, M. Kamel, K. Alexis, H. Oleynikova, and R. Siegwart · 2016
Earlier work this paper cites.
Visual relationship detection with language priors
C. Lu, R. Krishna, M. Bernstein, and L. Fei-Fei · 2016
Earlier work this paper cites.
Speeding-up robot exploration by exploiting background information
S. Oßwald, M. Bennewitz, W. Burgard, and C. Stachniss · 2016
Earlier work this paper cites.
Learning to poke by poking: Experiential learning of intuitive physics
P. Agrawal, A. V. Nair, P. Abbeel, J. Malik, and S. Levine · 2016
Earlier work this paper cites.
Uncertainty-aware receding horizon exploration and mapping using aerial robots
C. Papachristos, S. Khattak, and K. Alexis · 2017
Earlier work this paper cites.
Scene graph generation by iterative message passing
D. Xu, Y. Zhu, C. B. Choy, and L. Fei-Fei · 2017
Earlier work this paper cites.
Visual translation embedding network for visual relation detection
H. Zhang, Z. Kyaw, S.-F. Chang, and T.-S. Chua · 2017
Earlier work this paper cites.
Robot exploration in unknown cluttered environments when dealing with uncertainty
F. Niroui, B. Sprenger, and G. Nejat · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Earlier work this paper cites.
Learning to push by grasping: Using multiple tasks for effective learning
L. Pinto and A. Gupta · 2017
Earlier work this paper cites.
Building kinematic and dynamic models of articulated objects with multi-modal interactive perception
R. Martín-Martín and O. Brock · 2017
Earlier work this paper cites.
Mapping instructions to actions in 3d environments with visual goal prediction
D. Misra, A. Bennett, V. Blukis, E. Niklasson, M. Shatkhin, and Y. Artzi · 2018
Earlier work this paper cites.
Visual representations for semantic target driven navigation
A. Mousavian, A. Toshev, M. Fiser, J. Kosecka, and J. Davidson · 2018
Earlier work this paper cites.
Graph r-cnn for scene graph generation
J. Yang, J. Lu, S. Lee, D. Batra, and D. Parikh · 2018
Earlier work this paper cites.
Embodied question answering
A. Das, S. Datta, G. Gkioxari, S. Lee, D. Parikh, and D. Batra · 2018
Earlier work this paper cites.
Deep reinforcement learning robot for search and rescue applications: Exploration in unknown cluttered environments
F. Niroui, K. Zhang, Z. Kashino, and G. Nejat · 2019
Earlier work this paper cites.
3d scene graph: A structure for unified semantics, 3d space, and camera
I. Armeni, Z.-Y. He, J. Gwak, A. R. Zamir, M. Fischer, J. Malik, and S. Savarese · 2019
Earlier work this paper cites.
Clevrer: Collision events for video representation and reasoning
K. Yi, C. Gan, Y. Li, P. Kohli, J. Wu, A. Torralba, and J. B. Tenenbaum · 2019
Earlier work this paper cites.
Large-scale study of curiosity-driven learning
Y. Burda, H. Edwards, D. Pathak, A. Storkey, T. Darrell, and A. A. Efros · 2019
Earlier work this paper cites.
Learning affordance landscapes for interaction exploration in 3d environments
T. Nagarajan and K. Grauman · 2020
Earlier work this paper cites.
Interactive gibson benchmark: A benchmark for interactive navigation in cluttered environments
F. Xia, W. B. Shen, C. Li, P. Kasimbeg, M. E. Tchapmi, A. Toshev, R. Martín-Martín, and S. Savarese · 2020
Earlier work this paper cites.
Localize, assemble, and predicate: Contextual object proposal embedding for visual relation detection
R. Wu, K. Xu, C. Liu, N. Zhuang, and Y. Mu · 2020
Earlier work this paper cites.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox · 2020
Earlier work this paper cites.
Rearrangement: A challenge for embodied ai
D. Batra, A. X. Chang, S. Chernova, A. J. Davison, J. Deng, V. Koltun, S. Levine, J. Malik, I. Mordatch, R. Mottaghi, et al · 2020
Cited alongside, same era.
Relmogen: Leveraging motion generation in reinforcement learning for mobile manipulation
F. Xia, C. Li, R. Martín-Martín, O. Litany, A. Toshev, and S. Savarese · 2020
Cited alongside, same era.
The arches space-analogue demonstration mission: Towards heterogeneous teams of autonomous robots for collaborative scientific sampling in planetary exploration
M. J. Schuster, M. G. Müller, S. G. Brunner, H. Lehner, P. Lehner, R. Sakagami, A. Dömel, L. Meyer, B. Vodermayer, R. Giubilato, et al · 2020
Cited alongside, same era.
Where to map? iterative rover-copter path planning for mars exploration
T. Sasaki, K. Otsu, R. Thakker, S. Haesaert, and A.-a. Agha-mohammadi · 2020
Cited alongside, same era.
Autonomous exploration under uncertainty via deep reinforcement learning on graphs
Look before you leap: Unveiling the power of gpt-4v in robotic vision-language planning
Y. Hu, F. Lin, T. Zhang, L. Yi, and Y. Gao · 2023
Later among the works it cites.
Foundations of spatial perception for robotics: Hierarchical representations and real-time systems
N. Hughes, Y. Chang, S. Hu, R. Talak, R. Abdulhai, J. Strader, and L. Carlone · 2023
Later among the works it cites.
Indoor and outdoor 3d scene graph generation via language-enabled spatial ontologies
J. Strader, N. Hughes, W. Chen, A. Speranzon, and L. Carlone · 2023
Later among the works it cites.
Learning reusable manipulation strategies
J. Mao, T. Lozano-Pérez, J. B. Tenenbaum, and L. P. Kaelbling · 2023
Later among the works it cites.
Programmatically grounded, compositionally generalizable robotic manipulation
R. Wang, J. Mao, J. Hsu, H. Zhao, J. Wu, and Y. Gao · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
F. Chen, J. D. Martin, Y. Huang, J. Wang, and B. Englot · 2020
Cited alongside, same era.
Active perception and representation for robotic manipulation
Y. Zaky, G. Paruthi, B. Tripp, and J. Bergstra · 2020
Cited alongside, same era.
Task-oriented active perception and planning in environments with partially known semantics
M. Ghasemi, E. Bulgur, and U. Topcu · 2020
Cited alongside, same era.
Auxiliary tasks and exploration enable objectgoal navigation
J. Ye, D. Batra, A. Das, and E. Wijmans · 2021
Cited alongside, same era.
Visual room rearrangement
L. Weihs, M. Deitke, A. Kembhavi, and R. Mottaghi · 2021
Cited alongside, same era.
Manipulathor: A framework for visual object manipulation
K. Ehsani, W. Han, A. Herrasti, E. VanderBilt, L. Weihs, E. Kolve, A. Kembhavi, and R. Mottaghi · 2021
Cited alongside, same era.
Interesting object, curious agent: Learning task-agnostic exploration
S. Parisi, V. Dean, D. Pathak, and A. Gupta · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Cited alongside, same era.
Compositional Diffusion-Based Continuous Constraint Solvers
Z. Yang, J. Mao, Y. Du, J. Wu, J. B. Tenenbaum, T. Lozano-Pérez, and L. P. Kaelbling · 2023
Later among the works it cites.
Composable part-based manipulation
W. Liu, J. Mao, J. Hsu, T. Hermans, A. Garg, and J. Wu · 2023
Later among the works it cites.
Interactive visual task learning for robots
W. Gu, A. Sah, and N. Gopalan · 2023
Later among the works it cites.
Esc: Exploration with soft commonsense constraints for zero-shot object navigation
K.-Q. Zhou, K. Zheng, C. Pryor, Y. Shen, H. Jin, L. Getoor, and X. Wang · 2023
Later among the works it cites.
Peanut: predicting and navigating to unseen targets
A. J. Zhai and S. Wang · 2023
Later among the works it cites.
Vlfm: Vision-language frontier maps for zero-shot semantic navigation
N. Yokoyama, S. Ha, D. Batra, J. Wang, and B. Bucher · 2023
Later among the works it cites.
Scientific exploration of challenging planetary analog environments with a team of legged robots
P. Arm, G. Waibel, J. Preisig, T. Tuna, R. Zhou, V. Bickel, G. Ligeza, T. Miki, F. Kehl, H. Kolvenbach, et al · 2023
Later among the works it cites.
Ditto in the house: Building articulation models of indoor scenes through interactive perception
C.-C. Hsu, Z. Jiang, and Y. Zhu · 2023
Later among the works it cites.
Perceiving unseen 3d objects by poking the objects
L. Chen, Y. Song, H. Bao, and X. Zhou · 2023
Later among the works it cites.
Sam-rl: Sensing-aware model-based reinforcement learning via differentiable physics-based simulation and rendering
J. Lv, Y. Feng, C. Zhang, S. Zhao, L. Shao, and C. Lu · 2023
Later among the works it cites.
Active-perceptive motion generation for mobile manipulation
S. Jauhri, S. Lueth, and G. Chalvatzaki · 2023
Later among the works it cites.
R. OpenAI · 2023
Later among the works it cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence · 2023
Later among the works it cites.
Gpt-4v(ision) system card
OpenAI · 2023
Later among the works it cites.
The dawn of lmms: Preliminary explorations with gpt-4v(ision)
Z. Yang, L. Li, K. Lin, J. Wang, C.-C. Lin, Z. Liu, and L. Wang · 2023
Later among the works it cites.
Leveraging pre-trained large language models to construct and utilize world models for model-based task planning
L. Guan, K. Valmeekam, S. Sreedharan, and S. Kambhampati · 2023
Later among the works it cites.
Llm-grounder: Open-vocabulary 3d visual grounding with large language model as an agent
J. Yang, X. Chen, S. Qian, N. Madaan, M. Iyengar, D. F. Fouhey, and J. Chai · 2023
Later among the works it cites.
Think, act, and ask: Open-world interactive personalized robot navigation
Y. Dai, R. Peng, S. Li, and J. Chai · 2023
Later among the works it cites.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2023
Later among the works it cites.
Distilled feature fields enable few-shot language-guided manipulation
W. Shen, G. Yang, A. Yu, J. Wong, L. P. Kaelbling, and P. Isola · 2023
Later among the works it cites.
Grounding dino: Marrying dino with grounded pre-training for open-set object detection
S. Liu, Z. Zeng, T. Ren, F. Li, H. Zhang, J. Yang, C. Li, J. Yang, H. Su, J. Zhu, et al · 2023
Later among the works it cites.
Segment anything in high quality
L. Ke, M. Ye, M. Danelljan, Y. Liu, Y.-W. Tai, C.-K. Tang, and F. Yu · 2023
Later among the works it cites.
Conceptfusion: Open-set multimodal 3d mapping
K. M. Jatavallabhula, A. Kuwajerwala, Q. Gu, M. Omama, T. Chen, S. Li, G. Iyer, S. Saryazdi, N. Keetha, A. Tewari, et al · 2023
Later among the works it cites.
Ovir-3d: Open-vocabulary 3d instance retrieval without training on 3d data
S. Lu, H. Chang, E. P. Jing, A. Boularias, and K. Bekris · 2023
Later among the works it cites.
Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning
Q. Gu, A. Kuwajerwala, S. Morin, K. Jatavallabhula, B. Sen, A. Agarwal, C. Rivera, W. Paul, K. Ellis, R. Chellappa, C. Gan, C. de Melo, J. Tenenbaum, A. Torralba, F. Shkurti, and L. Paull · 2023
Later among the works it cites.
Task and motion planning in hierarchical 3d scene graphs
A. Ray, C. Bradley, L. Carlone, and N. Roy · 2024
Closest in time.
Llms can’t plan, but can help planning in llm-modulo frameworks
S. Kambhampati, K. Valmeekam, L. Guan, K. Stechly, M. Verma, S. Bhambri, L. Saldyt, and A. Murthy · 2024
Closest in time.