Fetching the paper…
Reading the bibliography…
Moving objects to find a fully-occluded target object, known as mechanical search, is a challenging problem in robotics.
A robot exploration and mapping strategy based on a semantic hierarchy of spatial representations
B. Kuipers and Y.-T. Byun · 1991
Earlier work this paper cites.
Divergence measures based on the shannon entropy
J. Lin · 1991
Earlier work this paper cites.
Integration of representation into goal-driven behavior-based robots
J. Maja · 1992
Earlier work this paper cites.
On the relative complexity of active vs. passive visual search
J. K. Tsotsos · 1992
Earlier work this paper cites.
Robotic vehicles for planetary exploration
B. H. Wilcox · 1992
Earlier work this paper cites.
Using intermediate objects to improve the efficiency of visual search
L. E. Wixson and D. H. Ballard · 1994
Earlier work this paper cites.
Integrating grid-based and topological maps for mobile robot navigation
S. Thrun and A. Bücken · 1996
Earlier work this paper cites.
Spatial learning for navigation in dynamic environments
B. Yamauchi and R. Beer · 1996
Earlier work this paper cites.
A frontier-based approach for autonomous exploration
B. Yamauchi · 1997
Earlier work this paper cites.
Learning metric-topological maps for indoor mobile robot navigation
S. Thrun · 1998
Earlier work this paper cites.
Adaptive mobile robot navigation and mapping
H. J. S. Feder, J. J. Leonard, and C. M. Smith · 1999
Earlier work this paper cites.
Walk the talk: Connecting language, knowledge, and action in route instructions
M. MacMahon, B. Stankiewicz, and B. Kuipers · 2006
Earlier work this paper cites.
Utilizing object-object and object-scene context when planning to find things
T. Kollar and N. Roy · 2009
Earlier work this paper cites.
Toward understanding natural language directions
T. Kollar, S. Tellex, D. Roy, and N. Roy · 2010
Earlier work this paper cites.
Following directions using statistical machine translation
C. Matuszek, D. Fox, and K. Koscher · 2010
Earlier work this paper cites.
Learning to interpret natural language navigation instructions from observations
D. Chen and R. Mooney · 2011
Earlier work this paper cites.
Understanding natural language commands for robotic navigation and mobile manipulation
S. Tellex, T. Kollar, S. Dickerson, M. Walter, A. Banerjee, S. Teller, and N. Roy · 2011
Earlier work this paper cites.
Using the web to interactively learn to find objects
M. Samadi, T. Kollar, and M. Veloso · 2012
Earlier work this paper cites.
Object reading: Text recognition for object recognition
S. Karaoglu, J. Gemert, and T. Gevers · 2012
Earlier work this paper cites.
Interactive environment exploration in clutter
M. Gupta, T. Rühr, M. Beetz, and G. S. Sukhatme · 2013
Earlier work this paper cites.
Manipulation-based active search for occluded objects
L. L. Wong, L. P. Kaelbling, and T. Lozano-Pérez · 2013
Earlier work this paper cites.
Active visual object search in unknown environments using uncertain semantics
A. Aydemir, A. Pronobis, M. Göbelbecker, and P. Jensfelt · 2013
Earlier work this paper cites.
Imitation learning for natural language direction following through unknown environments
F. Duvallet, T. Kollar, and A. Stentz · 2013
Earlier work this paper cites.
Learning to parse natural language commands to a robot control system
C. Matuszek, E. Herbst, L. Zettlemoyer, and D. Fox · 2013
Earlier work this paper cites.
Object search by manipulation
M. R. Dogar, M. C. Koval, A. Tallavajhula, and S. S. Srinivasa · 2014
Earlier work this paper cites.
A natural language planner interface for mobile manipulators
T. M. Howard, S. Tellex, and N. Roy · 2014
Earlier work this paper cites.
Learning models for following natural language directions in unknown environments
S. Hemachandra, F. Duvallet, T. M. Howard, N. Roy, A. Stentz, and M. R. Walter · 2015
Earlier work this paper cites.
Learning to interpret natural language commands through human-robot dialog
J. Thomason, S. Zhang, R. J. Mooney, and P. Stone · 2015
Earlier work this paper cites.
Act to see and see to act: Pomdp planning for objects search in clutter
J. K. Li, D. Hsu, and W. S. Lee · 2016
Earlier work this paper cites.
Inferring maps and behaviors from natural language instructions
F. Duvallet, M. R. Walter, T. Howard, S. Hemachandra, J. Oh, S. Teller, N. Roy, and A. Stentz · 2016
Earlier work this paper cites.
Listen, attend, and walk: Neural mapping of navigational instructions to action sequences
H. Mei, M. Bansal, and M. Walter · 2016
Earlier work this paper cites.
Tell me dave: Context-sensitive grounding of natural language to manipulation instructions
D. K. Misra, J. Sung, K. Lee, and A. Saxena · 2016
Earlier work this paper cites.
Target-driven visual navigation in indoor scenes using deep reinforcement learning
Y. Zhu, R. Mottaghi, E. Kolve, J. J. Lim, A. Gupta, L. Fei-Fei, and A. Farhadi · 2017
Earlier work this paper cites.
Mapping instructions and visual observations to actions with reinforcement learning
D. K. Misra, J. Langford, and Y. Artzi · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
P. Anderson, Q. Wu, D. Teney, J. Bruce, M. Johnson, N. Sünderhauf, I. Reid, S. Gould, and A. Van Den Hengel · 2018
Earlier work this paper cites.
Visual semantic navigation using scene priors
W. Yang, X. Wang, A. Farhadi, A. Gupta, and R. Mottaghi · 2018
Earlier work this paper cites.
Speaker-follower models for vision-and-language navigation
D. Fried, R. Hu, V. Cirik, A. Rohrbach, J. Andreas, L.-P. Morency, T. Berg-Kirkpatrick, K. Saenko, D. Klein, and T. Darrell · 2018
Earlier work this paper cites.
Interactive visual grounding of referring expressions for human-robot interaction
M. Shridhar and D. Hsu · 2018
Earlier work this paper cites.
Efficient grounding of abstract spatial concepts for natural language interaction with robot platforms
R. Paul, J. Arkin, D. Aksaray, N. Roy, and T. M. Howard · 2018
Cited alongside, same era.
Mechanical search: Multi-step retrieval of a target object occluded by clutter
M. Danielczuk, A. Kurenkov, A. Balakrishna, M. Matl, D. Wang, R. Martin-Martin, A. Garg, S. Savarese, and K. Goldberg · 2019
Cited alongside, same era.
Online planning for target object search in clutter under partial observability
Y. Xiao, S. Katt, A. ten Pas, S. Chen, and C. Amato · 2019
Cited alongside, same era.
Splitnet: Sim2sim and task2task transfer for embodied visual navigation
D. Gordon, A. Kadian, D. Parikh, J. Hoffman, and D. Batra · 2019
Cited alongside, same era.
Learning to learn how to learn: Self-adaptive visual navigation using meta-learning
M. Wortsman, K. Ehsani, M. Rastegari, A. Farhadi, and R. Mottaghi · 2019
Cited alongside, same era.
Socratic models: Composing zero-shot multimodal reasoning with language
A. Zeng, A. Wong, S. Welker, K. Choromanski, F. Tombari, A. Purohit, M. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke, et al · 2022
Later among the works it cites.
Inner monologue: Embodied reasoning through planning with language models
W. H. et al · 2022
Later among the works it cites.
Progprompt: Generating situated robot task plans using large language models
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg · 2022
Later among the works it cites.
LM-nav: Robotic navigation with large pre-trained models of language, vision, and action
D. Shah, B. Osiński, brian ichter, and S. Levine · 2022
Later among the works it cites.
Open-vocabulary object detection via vision and language knowledge distillation, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language as an abstraction for hierarchical deep reinforcement learning
Y. Jiang, S. S. Gu, K. P. Murphy, and C. Finn · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Holm: Hallucinating objects with language models for referring expression recognition in partially-observed scenes
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Cited alongside, same era.
Language conditioned imitation learning over unstructured data
C. Lynch and P. Sermanet · 2020
Cited alongside, same era.
Learning to explore using active neural slam
D. S. Chaplot, D. Gandhi, S. Gupta, A. Gupta, and R. Salakhutdinov · 2020
Cited alongside, same era.
Semantic visual navigation by watching youtube videos
M. Chang, A. Gupta, and S. Gupta · 2020
Cited alongside, same era.
Object goal navigation using goal-oriented semantic exploration
D. S. Chaplot, D. P. Gandhi, A. Gupta, and R. R. Salakhutdinov · 2020
Cited alongside, same era.
X. Gu, T.-Y. Lin, W. Kuo, and Y. Cui · 2022
Later among the works it cites.
Detecting twenty-thousand classes using image-level supervision, 2022
X. Zhou, R. Girdhar, A. Joulin, P. Krähenbühl, and I. Misra · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou · 2022
Later among the works it cites.
Optimal shelf arrangement to minimize robot retrieval time
L. Y. Chen, H. Huang, M. Danielczuk, J. Ichnowski, and K. Goldberg · 2022
Later among the works it cites.
Memory-augmented reinforcement learning for image-goal navigation
L. Mezghan, S. Sukhbaatar, T. Lavril, O. Maksymets, D. Batra, P. Bojanowski, and K. Alahari · 2022
Later among the works it cites.
Zero experience required: Plug & play modular transfer learning for semantic visual navigation
Z. Al-Halah, S. K. Ramakrishnan, and K. Grauman · 2022
Later among the works it cites.
Procthor: Large-scale embodied ai using procedural generation
M. Deitke, E. VanderBilt, A. Herrasti, L. Weihs, K. Ehsani, J. Salvador, W. Han, E. Kolve, A. Kembhavi, and R. Mottaghi · 2022
Later among the works it cites.
Semantic abstraction: Open-world 3d scene understanding from 2d vision-language models
H. Ha and S. Song · 2022
Later among the works it cites.
Language-driven semantic segmentation
B. Li, K. Q. Weinberger, S. Belongie, V. Koltun, and R. Ranftl · 2022
Later among the works it cites.
Openscene: 3d scene understanding with open vocabularies
S. Peng, K. Genova, C. Jiang, A. Tagliasacchi, M. Pollefeys, T. Funkhouser, et al · 2022
Later among the works it cites.
Clip-fields: Weakly supervised semantic fields for robotic memory
N. M. M. Shafiullah, C. Paxton, L. Pinto, S. Chintala, and A. Szlam · 2022
Later among the works it cites.
Open-vocabulary queryable scene representations for real world planning
B. Chen, F. Xia, B. Ichter, K. Rao, K. Gopalakrishnan, M. S. Ryoo, A. Stone, and D. Kappler · 2022
Later among the works it cites.
Simple open-vocabulary object detection with vision transformers
M. Minderer, A. Gritsenko, A. Stone, M. Neumann, D. Weissenborn, A. Dosovitskiy, A. Mahendran, A. Arnab, M. Dehghani, Z. Shen, et al · 2022
Later among the works it cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Later among the works it cites.
Correcting robot plans with natural language feedback
P. Sharma, B. Sundaralingam, V. Blukis, C. Paxton, T. Hermans, A. Torralba, J. Andreas, and D. Fox · 2022
Later among the works it cites.
Text and code embeddings by contrastive pre-training
A. N. et al · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
S. Z. et al · 2022
Later among the works it cites.
Visual language maps for robot navigation
C. Huang, O. Mees, A. Zeng, and W. Burgard · 2023
Closest in time.
Real-world robot learning with masked visual pre-training
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell · 2023
Closest in time.
Open-world object manipulation using pre-trained vision-language models
A. Stone, T. Xiao, Y. Lu, K. Gopalakrishnan, K.-H. Lee, Q. Vuong, P. Wohlhart, B. Zitkovich, F. Xia, C. Finn, et al · 2023
Closest in time.
Open-vocabulary queryable scene representations for real world planning, 2023
B. Chen, F. Xia, B. Ichter, K. Rao, K. Gopalakrishnan, M. S. Ryoo, A. Stone, and D. Kappler · 2023
Closest in time.
Lerf: Language embedded radiance fields
J. Kerr, C. M. Kim, K. Goldberg, A. Kanazawa, and M. Tancik · 2023
Closest in time.
Open-vocabulary semantic segmentation with mask-adapted clip, 2023
F. Liang, B. Wu, X. Dai, K. Li, Y. Zhao, H. Zhang, P. Zhang, P. Vajda, and D. Marculescu · 2023
Closest in time.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence · 2023
Closest in time.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick · 2023
Closest in time.
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Closest in time.
Conceptfusion: Open-set multimodal 3d mapping
K. M. Jatavallabhula, A. Kuwajerwala, Q. Gu, M. Omama, T. Chen, S. Li, G. Iyer, S. Saryazdi, N. Keetha, A. Tewari, et al · 2023
Closest in time.
No, to the right: Online language corrections for robotic manipulation via shared autonomy
Y. Cui, S. Karamcheti, R. Palleti, N. Shivakumar, P. Liang, and D. Sadigh · 2023
Closest in time.
Grounded decoding: Guiding text generation with grounded models for robot control
W. Huang, F. Xia, D. Shah, D. Driess, A. Zeng, Y. Lu, P. Florence, I. Mordatch, S. Levine, K. Hausman, et al · 2023
Closest in time.
Google product category
G. M. Center · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Visual instruction tuning, 2023
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Closest in time.
Gpt-4v(ision) system card, 2023
OpenAI · 2023
Closest in time.