Fetching the paper…
Reading the bibliography…
Solving complex long-horizon robotic manipulation problems requires sophisticated high-level planning capabilities, the ability to reason about the physical world, and reactively choose appropriate motor skills.
A review of robot learning for manipulation: Challenges, representations, and algorithms, 2020
Kroemer, O., Niekum, S., and Konidaris, G · 1907
Earlier work this paper cites.
Driess, D., Ha, J.-S., and Toussaint, M · 2006
Earlier work this paper cites.
Denoising diffusion probabilistic models, 2020
Ho, J., Jain, A., and Abbeel, P · 2006
Earlier work this paper cites.
Learning compositional models of robot skills for task and motion planning, 2021
Wang, Z., Garrett, C. R., Kaelbling, L. P., and Lozano-Pérez, T · 2006
Earlier work this paper cites.
Integrated task and motion planning, 2020a
Garrett, C. R., Chitnis, R., Holladay, R., Kim, B., Silver, T., Kaelbling, L. P., and Lozano-Pérez, T · 2010
Earlier work this paper cites.
Hierarchical task and motion planning in the now
Kaelbling, L. P. and Lozano-Pérez, T · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Score-based generative modeling through stochastic differential equations, 2021
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2011
Earlier work this paper cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Earlier work this paper cites.
Hg-dagger: Interactive imitation learning with human experts
Kelly, M., Sidrane, C., Driggs-Campbell, K., and Kochenderfer, M. J · 2018
Earlier work this paper cites.
Toward next-generation learned robot manipulation
Cui, J. and Trinkle, J · 2021
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models, 2021
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2021
Earlier work this paper cites.
Instructpix2pix: Learning to follow image editing instructions
Brooks, T., Holynski, A., and Efros, A. A · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Self-rag: Learning to retrieve, generate, and critique through self-reflection
Asai, A., Wu, Z., Wang, Y., Sil, A., and Hajishirzi, H · 2023
Cited alongside, same era.
Qwen-vl: A versatile vision-language model for understanding, generation, and retrieval
Bai, J., Bai, S., Du, S., Han, S., Liu, P., et al · 2023
Cited alongside, same era.
Zero-shot robotic manipulation with pretrained image-editing diffusion models, 2023
Black, K., Nakamoto, M., Atreya, P., Walke, H., Finn, C., Kumar, A., and Levine, S · 2023
Cited alongside, same era.
Pali-x: On scaling up a multilingual vision and language model
Physically grounded vision-language models for robotic manipulation, 2024
Gao, J., Sarkar, B., Xia, F., Xiao, T., Wu, J., Ichter, B., Majumdar, A., and Sadigh, D · 2024
Later among the works it cites.
Introducing gemini: Our largest and most capable ai model
Google · 2024
Later among the works it cites.
Large language models cannot self-correct reasoning yet, 2024
Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., and Zhou, D · 2024
Later among the works it cites.
How far is video generation from world model: A physical law perspective, 2024
Kang, B., Yue, Y., Lu, R., Lin, Z., Zhao, Y., Wang, K., Huang, G., and Feng, J · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, X., Dai, J., Li, X., Peng, B., Singh, M., Tao, S., Wang, X., Wang, Y., Xia, Y., et al · 2023
Cited alongside, same era.
Palm-e: An embodied multimodal language model
Driess, D., Black, A., Kataoka, H., Tsurumine, Y., Koyama, Y., Mansard, N., Fox, D., Choromanski, K., Ichter, B., Hausman, K., et al · 2023
Cited alongside, same era.
Du, Y., Yang, M., Florence, P., Xia, F., Wahid, A., Ichter, B., Sermanet, P., Yu, T., Abbeel, P., Tenenbaum, J. B., Kaelbling, L., Zeng, A., and Tompson, J · 2023
Cited alongside, same era.
Look before you leap: Unveiling the power of gpt-4v in robotic vision-language planning, 2023
Hu, Y., Lin, F., Zhang, T., Yi, L., and Gao, Y · 2023
Cited alongside, same era.
Voxposer: Composable 3d value maps for robotic manipulation with language models, 2023
Huang, W., Wang, C., Zhang, R., Li, Y., Wu, J., and Fei-Fei, L · 2023
Cited alongside, same era.
Visual instruction tuning, 2023
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Cited alongside, same era.
Pan, L., Saxon, M., Xu, W., Nathani, D., Wang, X., and Wang, W. Y · 2023
Cited alongside, same era.
Self-instruct: Aligning language models with self-generated instructions, 2023
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2023
Cited alongside, same era.
Li, B., Zhang, Y., Guo, D., Zhang, R., Li, F., Zhang, H., Zhang, K., Zhang, P., Li, Y., Liu, Z., and Li, C · 2024
Later among the works it cites.
Multistage cable routing through hierarchical imitation learning
Luo, J., Xu, C., Geng, X., Feng, G., Fang, K., Tan, L., Schaal, S., and Levine, S · 2024
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al · 2024
Later among the works it cites.
Pivot: Iterative visual prompting elicits actionable knowledge for vlms, 2024
Nasiriany, S., Xia, F., Yu, W., Xiao, T., Liang, J., Dasgupta, I., Xie, A., Driess, D., Wahid, A., Xu, Z., Vuong, Q., Zhang, T., Lee, T.-W. E., Lee, K.-H., Xu, P., Kirmani, S., Zhu, Y., Zeng, A., Hausman, K., Heess, N., Finn, C., Levine, S., and Ichter, B · 2024
Later among the works it cites.
Self-reflection in llm agents: Effects on problem-solving performance
Renze, M. and Guven, E · 2024
Later among the works it cites.
Yell at your robot: Improving on-the-fly from language corrections, 2024
Shi, L. X., Hu, Z., Zhao, T. Z., Sharma, A., Pertsch, K., Luo, J., Levine, S., and Finn, C · 2024
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2024
Later among the works it cites.
Gpt-4v (ision) for robotics: Multimodal task planning from human demonstration
Wake, N., Kanehira, A., Sasabuchi, K., Takamatsu, J., and Ikeuchi, K · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K · 2024
Later among the works it cites.
Do generative video models learn physical principles from watching videos?, 2025
Motamed, S., Culp, L., Swersky, K., Jaini, P., and Geirhos, R · 2025
Closest in time.
Exact: Teaching ai agents to explore with reflective-mcts and exploratory learning, 2025
Yu, X., Peng, B., Vajipey, V., Cheng, H., Galley, M., Gao, J., and Yu, Z · 2025
Closest in time.