Fetching the paper…
Reading the bibliography…
Robotic manipulation in open-world settings requires not only task execution but also the ability to detect and learn from failures.
Children’s critical thinking when learning from others
G. D. Heyman · 2008
Earlier work this paper cites.
Learning by trial and error
H. P. Young · 2009
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo, and J. P. Bigham · 2018
Earlier work this paper cites.
Human trust after robot mistakes: Study of the effects of different forms of robot communication
S. Ye, G. Neville, M. Schrum, M. Gombolay, S. Chernova, and A. Howard · 2019
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
A. Gupta, P. Dollar, and R. Girshick · 2019
Earlier work this paper cites.
Towards vqa models that can read
A. Singh, V. Natarjan, M. Shah, Y. Jiang, X. Chen, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
Childhood as a solution to explore–exploit tensions
A. Gopnik · 2020
Earlier work this paper cites.
Pddlstream: Integrating symbolic planners and blackbox samplers via optimistic adaptive planning
C. R. Garrett, T. Lozano-Pérez, and L. P. Kaelbling · 2020
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Maniskill: Generalizable manipulation skill benchmark with large-scale demonstrations
T. Mu, Z. Ling, F. Xiang, D. Yang, X. Li, S. Tao, Z. Huang, Z. Jia, and H. Su · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang · 2022
Earlier work this paper cites.
A survey of embodied ai: From simulators to research tasks
J. Duan, S. Yu, H. L. Tan, H. Zhu, and C. Tan · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
P. Lu, S. Mishra, T. Xia, L. Qiu, K.-W. Chang, S.-C. Zhu, O. Tafjord, P. Clark, and A. Kalyan · 2022
Earlier work this paper cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al · 2023
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Adding conditional control to text-to-image diffusion models
L. Zhang, A. Rao, and M. Agrawala · 2023
Cited alongside, same era.
Newton: Are large language models capable of physical reasoning?
Y. R. Wang, J. Duan, D. Fox, and S. Srinivasa · 2023
Cited alongside, same era.
Eureka: Human-level reward design via coding large language models
Y. J. Ma, W. Liang, G. Wang, D.-A. Huang, O. Bastani, D. Jayaraman, Y. Zhu, L. Fan, and A. Anandkumar · 2023
Cited alongside, same era.
Voxposer: Composable 3d value maps for robotic manipulation with language models
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei · 2023
Cited alongside, same era.
Improved baselines with visual instruction tuning, 2023
H. Liu, C. Li, Y. Li, and Y. J. Lee · 2023
Cited alongside, same era.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
Hello gpt-4o, May 2024
OpenAI · 2024
Closest in time.
Robopoint: A vision-language model for spatial affordance prediction for robotics
W. Yuan, J. Duan, V. Blukis, W. Pumacay, R. Krishna, A. Murali, A. Mousavian, and D. Fox · 2024
Closest in time.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
B. Chen, Z. Xu, S. Kirmani, B. Ichter, D. Driess, P. Florence, D. Sadigh, L. Guibas, and F. Xia · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Khanna, E. Yadollahi, M. Björkman, I. Leite, and C. Smith · 2023
Cited alongside, same era.
Scaling up and distilling down: Language-guided robot skill acquisition
H. Ha, P. Florence, and S. Song · 2023
Cited alongside, same era.
Gensim: Generating robotic simulation tasks via large language models
L. Wang, Y. Ling, Z. Yuan, M. Shridhar, C. Bao, Y. Qin, B. Wang, H. Xu, and X. Wang · 2023
Cited alongside, same era.
Vision-language models as success detectors
Y. Du, K. Konyushkova, M. Denil, A. Raju, J. Landon, F. Hill, N. de Freitas, and S. Cabi · 2023
Cited alongside, same era.
Mimicgen: A data generation system for scalable robot learning using human demonstrations
A. Mandlekar, S. Nasiriany, B. Wen, I. Akinola, Y. Narang, L. Fan, Y. Zhu, and D. Fox · 2023
Cited alongside, same era.
Toward general-purpose robots via foundation models: A survey and meta-analysis
Y. Hu, Q. Xie, V. Jain, J. Francis, J. Patrikar, N. Keetha, S. Kim, Y. Xie, T. Zhang, Z. Zhao, et al · 2023
Cited alongside, same era.
Foundation models in robotics: Applications, challenges, and the future
R. Firoozi, J. Tucker, S. Tian, A. Majumdar, J. Sun, W. Liu, Y. Zhu, S. Song, A. Kapoor, K. Hausman, et al · 2023
Cited alongside, same era.
Y. J. Ma, W. Liang, H.-J. Wang, S. Wang, Y. Zhu, L. Fan, O. Bastani, and D. Jayaraman · 2024
Closest in time.
Trust the proc3s: Solving long-horizon robotics problems with llms and constraint satisfaction, 2024
A. Curtis, N. Kumar, J. Cao, T. Lozano-Pérez, and L. P. Kaelbling · 2024
Closest in time.
Copa: General robotic manipulation through spatial constraints of parts with foundation models
H. Huang, F. Lin, Y. Hu, S. Wang, and Y. Gao · 2024
Closest in time.
Manipulate-anything: Automating real-world robots using vision-language models
J. Duan, W. Yuan, W. Pumacay, Y. R. Wang, K. Ehsani, D. Fox, and R. Krishna · 2024
Closest in time.
Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation
W. Huang, C. Wang, Y. Li, R. Zhang, and L. Fei-Fei · 2024
Closest in time.
Replan: Robotic replanning with perception and language models
M. Skreta, Z. Zhou, J. L. Yuan, K. Darvish, A. Aspuru-Guzik, and A. Garg · 2024
Closest in time.
Intervengen: Interventional data generation for robust and data-efficient robot imitation learning
R. Hoque, A. Mandlekar, C. Garrett, K. Goldberg, and D. Fox · 2024
Closest in time.
Decomposing the generalization gap in imitation learning for visual robotic manipulation
A. Xie, L. Lee, T. Xiao, and C. Finn · 2024
Closest in time.
The colosseum: A benchmark for evaluating generalization for robotic manipulation
W. Pumacay, I. Singh, J. Duan, R. Krishna, J. Thomason, and D. Fox · 2024
Closest in time.
Deep generative models in robotics: A survey on learning from multimodal demonstrations
J. Urain, A. Mandlekar, Y. Du, M. Shafiullah, D. Xu, K. Fragkiadaki, G. Chalvatzaki, and J. Peters · 2024
Closest in time.
Moka: Open-vocabulary robotic manipulation through mark-based visual prompting
F. Liu, K. Fang, P. Abbeel, and S. Levine · 2024
Closest in time.
Llara: Supercharging robot learning data for vision-language policy
X. Li, C. Mata, J. Park, K. Kahatapitiya, Y. S. Jang, J. Shang, K. Ranasinghe, R. Burgert, M. Cai, Y. J. Lee, et al · 2024
Closest in time.
Octopi: Object property reasoning with large tactile-language models
S. Yu, K. Lin, A. Xiao, J. Duan, and H. Soh · 2024
Closest in time.
Droid: A large-scale in-the-wild robot manipulation dataset
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, 2024
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.