Fetching the paper…
Reading the bibliography…
Robotic real-world reinforcement learning (RL) with vision-language-action (VLA) models is bottlenecked by sparse, handcrafted rewards and inefficient exploration.
Reboot: Reuse data for bootstrapping efficient real-world dexterous manipulation
Z. Hu, A. Rovinsky, J. Luo, V. Kumar, A. Gupta, and S. Levine · 1949
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Learning to describe differences between pairs of similar images
H. Jhamtani and T. Berg-Kirkpatrick · 2018
Earlier work this paper cites.
Robonet: Large-scale multi-robot learning
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn · 2019
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
Vip: Towards universal visual reward and representation via value-implicit pre-training
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang · 2022
Earlier work this paper cites.
Dexterous manipulation from images: Autonomous real-world rl via substep guidance
K. Xu, Z. Hu, R. Doshi, A. Rovinsky, V. Kumar, A. Gupta, and S. Levine · 2022
Earlier work this paper cites.
Instructpix2pix: Learning to follow image editing instructions
T. Brooks, A. Holynski, and A. A. Efros · 2023
Earlier work this paper cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2023
Earlier work this paper cites.
Rh20t: A comprehensive robotic dataset for learning diverse skills in one-shot
H.-S. Fang, H. Fang, Z. Tang, J. Liu, C. Wang, J. Wang, H. Zhu, and C. Lu · 2023
Earlier work this paper cites.
Deep rl at scale: Sorting waste in office buildings with a fleet of mobile manipulators
A. Herzog, K. Rao, K. Hausman, Y. Lu, P. Wohlhart, M. Yan, J. Lin, M. G. Arenas, T. Xiao, D. Kappler, et al · 2023
Earlier work this paper cites.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2023
Earlier work this paper cites.
Liv: Language-image representations and rewards for robotic control
Y. J. Ma, V. Kumar, A. Zhang, O. Bastani, and D. Jayaraman · 2023
Earlier work this paper cites.
Alan: Autonomously exploring robotic agents in the real world
R. Mendonca, S. Bahl, and D. Pathak · 2023
Earlier work this paper cites.
N. M. M. Shafiullah, A. Rai, H. Etukuru, Y. Liu, I. Misra, S. Chintala, and L. Pinto · 2023
Earlier work this paper cites.
Demonstrating a walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning
L. Smith, I. Kostrikov, and S. Levine · 2023
Earlier work this paper cites.
Bridgedata v2: A dataset for robot learning at scale
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du, et al · 2023
Earlier work this paper cites.
Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking
H. Bharadhwaj, J. Vakil, M. Sharma, A. Gupta, S. Tulsiani, and V. Kumar · 2024
Earlier work this paper cites.
On-robot reinforcement learning with goal-contrastive rewards
O. Biza, T. Weng, L. Sun, K. Schmeckpeper, T. Kelestemur, Y. J. Ma, R. Platt, J.-W. van de Meent, and L. L. Wong · 2024
Earlier work this paper cites.
π 0 \pi_{0} : A vision-language-action flow model for general robot control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Cited alongside, same era.
Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
B. Chen, Z. Xu, S. Kirmani, B. Ichter, D. Sadigh, L. Guibas, and F. Xia · 2024
Cited alongside, same era.
Aligniql: Policy alignment in implicit q-learning through constrained optimization
L. He, L. Shen, J. Tan, and X. Wang · 2024
Cited alongside, same era.
Droid: A large-scale in-the-wild robot manipulation dataset
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y. Chen, K. Ellis, et al · 2024
Cited alongside, same era.
Gr00t n1: An open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al · 2025
Closest in time.
Agibot world colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems
Q. Bu, J. Cai, L. Chen, X. Cui, Y. Ding, S. Feng, X. He, X. Huang, et al · 2025
Closest in time.
Conrft: A reinforced fine-tuning method for vla models via consistency policy
Y. Chen, S. Tian, S. Liu, Y. Zhou, H. Li, and D. Zhao · 2025
Closest in time.
Graspvla: a grasping foundation model pre-trained on billion-scale synthetic action data
S. Deng, M. Yan, S. Wei, H. Ma, Y. Yang, J. Chen, Z. Zhang, T. Yang, X. Zhang, H. Cui, et al · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn · 2024
Cited alongside, same era.
Practice makes perfect: Planning to learn skill parameter policies
N. Kumar, T. Silver, W. McClinton, L. Zhao, S. Proulx, T. Lozano-Pérez, L. P. Kaelbling, and J. Barry · 2024
Cited alongside, same era.
Data scaling laws in imitation learning for robotic manipulation
F. Lin, Y. Hu, P. Sheng, C. Wen, J. You, and Y. Gao · 2024
Cited alongside, same era.
Vision language models are in-context value learners
Y. J. Ma, J. Hejna, C. Fu, D. Shah, J. Liang, Z. Xu, S. Kirmani, P. Xu, D. Driess, T. Xiao, et al · 2024
Cited alongside, same era.
Continuously improving mobile manipulation with autonomous real-world rl
R. Mendonca, E. Panov, B. Bucher, J. Wang, and D. Pathak · 2024
Cited alongside, same era.
Learning a diffusion model policy from rewards via q-score matching
M. Psenka, A. Escontrela, P. Abbeel, and Y. Ma · 2024
Cited alongside, same era.
Diffusion policy policy optimization
A. Z. Ren, J. Lidard, L. L. Ankile, A. Simeonov, P. Agrawal, A. Majumdar, B. Burchfiel, H. Dai, and M. Simchowitz · 2024
Cited alongside, same era.
Robovqa: Multimodal long-horizon reasoning for robotics
P. Sermanet, T. Ding, J. Zhao, F. Xia, D. Dwibedi, K. Gopalakrishnan, C. Chan, G. Dulac-Arnold, S. Maddineni, N. J. Joshi, et al · 2024
Cited alongside, same era.
Z. Geng, M. Deng, X. Bai, J. Z. Kolter, and K. He · 2025
Closest in time.
Egodex: Learning dexterous manipulation from large-scale egocentric video
R. Hoque, P. Huang, D. J. Yoon, M. Sivapurapu, and J. Zhang · 2025
Closest in time.
π \pi 0. 5: a vision-language-action model with open-world generalization, 2025
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, et al · 2025
Closest in time.
A forget-and-grow strategy for deep reinforcement learning scaling in continuous control
Z. Kang, C. Hu, Y. Luo, Z. Yuan, R. Zheng, and H. Xu · 2025
Closest in time.
Fine-tuning vision-language-action models: Optimizing speed and success
M. J. Kim, C. Finn, and P. Liang · 2025
Closest in time.
Reinforcement learning with action chunking
Q. Li, Z. Zhou, and S. Levine · 2025
Closest in time.
Robofac: A comprehensive framework for robotic failure analysis and correction
W. Lu, M. Ye, Z. Ye, R. Tao, S. Yang, and B. Zhao · 2025
Closest in time.
Fmb: a functional manipulation benchmark for generalizable robotic learning
J. Luo, C. Xu, F. Liu, L. Tan, Z. Lin, J. Wu, P. Abbeel, and S. Levine · 2025
Closest in time.
Flow-based policy for online reinforcement learning
L. Lv, Y. Li, Y. Luo, F. Sun, T. Kong, J. Xu, and X. Ma · 2025
Closest in time.
S. Park, Q. Li, and S. Levine · 2025
Closest in time.
Modeling fine-grained hand-object dynamics for egocentric video representation learning
B. Pei, Y. Huang, J. Xu, G. Chen, Y. He, L. Yang, Y. Wang, W. Xie, Y. Qiao, F. Wu, et al · 2025
Closest in time.
Fast: Efficient action tokenization for vision-language-action models
K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine · 2025
Closest in time.
Gemini robotics: Bringing ai into the physical world
G. R. Team, S. Abeyruwan, J. Ainslie, J.-B. Alayrac, M. G. Arenas, T. Armstrong, A. Balakrishna, R. Baruch, M. Bauza, M. Blokzijl, et al · 2025
Closest in time.
W. Wang, J. Song, C. Liu, J. Ma, S. Feng, J. Wang, Y. Jiang, K. Chen, S. Zhan, Y. Wang, et al · 2025
Closest in time.