Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) is widely used to produce robust robotic manipulation policies, but fine-tuning vision-language-action (VLA) models with RL can be unstable due to inaccurate value estimates and sparse supervision at intermediate steps.
D. A. Pomerleau, “Alvinn: An autonomous land vehicle in a neural network,”
1988
Earlier work this paper cites.
R. M. Neal, “Slice sampling,”
2003
Earlier work this paper cites.
S. Ross, G. Gordon, and D. Bagnell, “A reduction of imitation learning and structured prediction to no-regret online learning,” in
2011
Earlier work this paper cites.
L. N. Smith, “Cyclical learning rates for training neural networks.” IEEE, 2017, pp. 464–472
2017
Earlier work this paper cites.
M. Kelly, C. Sidrane, K. Driggs-Campbell, and M. J. Kochenderfer, “Hg-dagger: Interactive imitation learning with human experts,” in
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
Anthony Brohan and others, “Rt-1: Robotics transformer for real-world control at scale,” in
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
H. Liu, S. Nasiriany, L. Zhang, Z. Bao, and Y. Zhu, “Robot learning on the job: Human-in-the-loop autonomy and learning during deployment,”
2022
Earlier work this paper cites.
E. Zelikman, Y. Wu, J. Mu, and N. Goodman, “Star: Bootstrapping reasoning with reasoning,”
2022
Earlier work this paper cites.
Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le, “Flow matching for generative modeling,”
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du,
2023
Cited alongside, same era.
O. X.-E. Collaboration
2023
Cited alongside, same era.
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,”
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Chebotar, Q. Vuong, K. Hausman, F. Xia, Y. Lu, A. Irpan, A. Kumar, T. Yu, A. Herzog, K. Pertsch,
2023
Cited alongside, same era.
2024
Later among the works it cites.
M. Zare, P. M. Kebria, A. Khosravi, and S. Nahavandi, “A survey of imitation learning: Algorithms, recent developments, and challenges,”
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
T. L. Team
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky, “
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
J. Luo, Z. Hu, C. Xu, Y. L. Tan, J. Berg, A. Sharma, S. Schaal, C. Finn, A. Gupta, and S. Levine, “Serl: A software suite for sample-efficient robotic reinforcement learning,” in
2024
Cited alongside, same era.
Q. Bu, J. Cai, L. Chen, X. Cui, Y. Ding, S. Feng, X. He, X. Huang,
2025
Closest in time.
Q. Li, Z. Zhou, and S. Levine, “Reinforcement learning with action chunking,”
2025
Closest in time.
P. Wu, Y. Shentu, Q. Liao, D. Jin, M. Guo, K. Sreenath, X. Lin, and P. Abbeel, “Robocopilot: Human-in-the-loop interactive imitation learning for robot manipulation,” 2025
2025
Closest in time.
Z. Hu, R. Wu, N. Enock, J. Li, R. Kadakia, Z. Erickson, and A. Kumar, “Rac: Robot learning for long-horizon tasks by scaling recovery and correction,” 2025
2025
Closest in time.
S. Park, Q. Li, and S. Levine, “Flow q-learning,”
2025
Closest in time.
2025
Closest in time.