Fetching the paper…
Reading the bibliography…
Robotic control policies learned from human demonstrations have achieved impressive results in many real-world applications.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Deterministic policy gradient algorithms
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller · 2014
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2015
Earlier work this paper cites.
End to end learning for self-driving cars
M. Bojarski · 2016
Earlier work this paper cites.
Density estimation using real nvp
L. Dinh, J. Sohl-Dickstein, and S. Bengio · 2016
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
T. Zhang, Z. McCarthy, O. Jow, D. Lee, X. Chen, K. Goldberg, and P. Abbeel · 2018
Earlier work this paper cites.
Vision-based multi-task manipulation for inexpensive robots using end-to-end learning from demonstration
R. Rahmatizadeh, P. Abolghasemi, L. Bölöni, and S. Levine · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Language-conditioned imitation learning for robot manipulation tasks
S. Stepputtis, J. Campbell, M. Phielipp, S. Lee, C. Baral, and H. Ben Amor · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
J. Ho, A. Jain, and P. Abbeel · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2020
Earlier work this paper cites.
Parrot: Data-driven behavioral priors for reinforcement learning
A. Singh, H. Liu, G. Zhou, A. Yu, N. Rhinehart, and S. Levine · 2020
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Earlier work this paper cites.
Awac: Accelerating online reinforcement learning with offline datasets
A. Nair, A. Gupta, M. Dalal, and S. Levine · 2020
Earlier work this paper cites.
What matters in learning from offline human demonstrations for robot manipulation
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y. Zhu, and R. Martín-Martín · 2021
Earlier work this paper cites.
Accelerating robotic reinforcement learning via parameterized action primitives
M. Dalal, D. Pathak, and R. R. Salakhutdinov · 2021
Earlier work this paper cites.
Offline reinforcement learning with implicit q-learning
I. Kostrikov, A. Nair, and S. Levine · 2021
Earlier work this paper cites.
Is pessimism provably efficient for offline rl?
Y. Jin, Z. Yang, and Z. Wang · 2021
Earlier work this paper cites.
Bellman-consistent pessimism for offline reinforcement learning
T. Xie, C.-A. Cheng, N. Jiang, P. Mineiro, and A. Agarwal · 2021
Earlier work this paper cites.
What matters in learning from offline human demonstrations for robot manipulation
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y. Zhu, and R. Martín-Martín · 2021
Earlier work this paper cites.
Stable-baselines3: Reliable reinforcement learning implementations
A. Raffin, A. Hill, A. Gleave, A. Kanervisto, M. Ernestus, and N. Dormann · 2021
Earlier work this paper cites.
Behavior transformers: Cloning k k modes with one stone
N. M. Shafiullah, Z. Cui, A. A. Altanzaya, and L. Pinto · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
Learning language-conditioned robot behavior from offline data and crowd-sourced annotation
S. Nair, E. Mitchell, K. Chen, S. Savarese, C. Finn, et al · 2022
Earlier work this paper cites.
From play to policy: Conditional behavior generation from uncurated robot data
Z. J. Cui, Y. Wang, N. M. M. Shafiullah, and L. Pinto · 2022
Earlier work this paper cites.
Planning with diffusion for flexible behavior synthesis
M. Janner, Y. Du, J. B. Tenenbaum, and S. Levine · 2022
Earlier work this paper cites.
Is conditional generative modeling all you need for decision-making?
A. Ajay, Y. Du, A. Gupta, J. Tenenbaum, T. Jaakkola, and P. Agrawal · 2022
Cited alongside, same era.
Diffusion policies as an expressive policy class for offline reinforcement learning
Z. Wang, J. J. Hunt, and M. Zhou · 2022
Cited alongside, same era.
Offline reinforcement learning via high-fidelity generative behavior modeling
H. Chen, C. Lu, C. Ying, H. Su, and J. Zhu · 2022
Cited alongside, same era.
Flow matching for generative modeling
Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le · 2022
Cited alongside, same era.
Flow straight and fast: Learning to generate and transfer data with rectified flow
π 0 \pi_{0} : A vision-language-action flow model for general robot control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Later among the works it cites.
Openvla: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al · 2024
Later among the works it cites.
Juicer: Data-efficient imitation learning for robotic assembly
L. Ankile, A. Simeonov, I. Shenfeld, and P. Agrawal · 2024
Later among the works it cites.
3d diffusion policy: Generalizable visuomotor policy learning via simple 3d representations
Y. Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu · 2024
Later among the works it cites.
Nomad: Goal masked diffusion policies for navigation and exploration
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Liu, C. Gong, and Q. Liu · 2022
Cited alongside, same era.
Building normalizing flows with stochastic interpolants
M. S. Albergo and E. Vanden-Eijnden · 2022
Cited alongside, same era.
The franka emika robot: A reference platform for robotics research and education
S. Haddadin, S. Parusel, L. Johannsmeier, S. Golz, S. Gabl, F. Walch, M. Sabaghian, C. Jähne, L. Hausperger, and S. Haddadin · 2022
Cited alongside, same era.
Rt-trajectory: Robotic task generalization via hindsight trajectory sketches
J. Gu, S. Kirmani, P. Wohlhart, Y. Lu, M. G. Arenas, K. Rao, W. Yu, C. Fu, K. Gopalakrishnan, Z. Xu, et al · 2023
Cited alongside, same era.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song · 2023
Cited alongside, same era.
Zero-shot robotic manipulation with pretrained image-editing diffusion models
K. Black, M. Nakamoto, P. Atreya, H. Walke, C. Finn, A. Kumar, and S. Levine · 2023
Cited alongside, same era.
Guided flows for generative modeling and decision making
Q. Zheng, M. Le, N. Shaul, Y. Lipman, A. Grover, and R. T. Chen · 2023
Cited alongside, same era.
Hierarchical diffusion for offline decision making
W. Li, X. Wang, B. Jin, and H. Zha · 2023
Cited alongside, same era.
A. Sridhar, D. Shah, C. Glossop, and S. Levine · 2024
Later among the works it cites.
The ingredients for robotic diffusion transformers
S. Dasari, O. Mees, S. Zhao, M. K. Srirama, and S. Levine · 2024
Later among the works it cites.
Diffusion policy policy optimization
A. Z. Ren, J. Lidard, L. L. Ankile, A. Simeonov, P. Agrawal, A. Majumdar, B. Burchfiel, H. Dai, and M. Simchowitz · 2024
Later among the works it cites.
Policy agnostic rl: Offline rl and online rl fine-tuning of any class and backbone
M. S. Mark, T. Gao, G. G. Sampaio, M. K. Srirama, A. Sharma, C. Finn, and A. Kumar · 2024
Later among the works it cites.
Simple hierarchical planning with diffusion
C. Chen, F. Deng, K. Kawaguchi, C. Gulcehre, and S. Ahn · 2024
Later among the works it cites.
Poco: Policy composition from and for heterogeneous robot learning
L. Wang, J. Zhao, Y. Du, E. H. Adelson, and R. Tedrake · 2024
Later among the works it cites.
Diffusion-based reinforcement learning via q-weighted variational policy optimization
S. Ding, K. Hu, Z. Zhang, K. Ren, W. Zhang, J. Yu, J. Wang, and Y. Shi · 2024
Later among the works it cites.
Diffusion policies for out-of-distribution generalization in offline reinforcement learning
S. E. Ada, E. Oztop, and E. Ugur · 2024
Later among the works it cites.
Entropy-regularized diffusion policy with q-ensembles for offline reinforcement learning
R. Zhang, Z. Luo, J. Sjölund, T. Schön, and P. Mattsson · 2024
Later among the works it cites.
Aligniql: Policy alignment in implicit q-learning through constrained optimization
L. He, L. Shen, J. Tan, and X. Wang · 2024
Later among the works it cites.
Steering your generalists: Improving robotic foundation models via value guidance
M. Nakamoto, O. Mees, A. Kumar, and S. Levine · 2024
Later among the works it cites.
Diffusion-dice: In-sample diffusion guidance for offline reinforcement learning
L. Mao, H. Xu, X. Zhan, W. Zhang, and A. Zhang · 2024
Later among the works it cites.
Learning multimodal behaviors from scratch with diffusion policy gradient
S. Li, R. Krohn, T. Chen, A. Ajay, P. Agrawal, and G. Chalvatzaki · 2024
Later among the works it cites.
From imitation to refinement–residual rl for precise assembly
L. Ankile, A. Simeonov, I. Shenfeld, M. Torne, and P. Agrawal · 2024
Later among the works it cites.
Policy decorator: Model-agnostic online refinement for large policy model
X. Yuan, T. Mu, S. Tao, Y. Fang, M. Zhang, and H. Su · 2024
Later among the works it cites.
Reno: Enhancing one-step text-to-image models through reward-based noise optimization
L. Eyring, S. Karthik, K. Roth, A. Dosovitskiy, and Z. Akata · 2024
Later among the works it cites.
The lottery ticket hypothesis in denoising: Towards semantic-driven initialization
J. Mao, X. Wang, and K. Aizawa · 2024
Later among the works it cites.
Generating images of rare concepts using pre-trained diffusion models
D. Samuel, R. Ben-Ari, S. Raviv, N. Darshan, and G. Chechik · 2024
Later among the works it cites.
A noise is worth diffusion guidance
D. Ahn, J. Kang, S. Lee, J. Min, M. Kim, W. Jang, H. Cho, S. Paul, S. Kim, E. Cha, et al · 2024
Later among the works it cites.
Y. Lipman, M. Havasi, P. Holderrieth, N. Shaul, M. Le, B. Karrer, R. T. Chen, D. Lopez-Paz, H. Ben-Hamu, and I. Gat · 2024
Later among the works it cites.
Ogbench: Benchmarking offline goal-conditioned rl
S. Park, K. Frans, B. Eysenbach, and S. Levine · 2024
Later among the works it cites.
Serl: A software suite for sample-efficient robotic reinforcement learning
J. Luo, Z. Hu, C. Xu, Y. L. Tan, J. Berg, A. Sharma, S. Schaal, C. Finn, A. Gupta, and S. Levine · 2024
Later among the works it cites.
Kimi k1. 5: Scaling reinforcement learning with llms
K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao, et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al · 2025
Closest in time.
Gr00t n1: An open foundation model for generalist humanoid robots
J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. Fan, Y. Fang, D. Fox, F. Hu, S. Huang, et al · 2025
Closest in time.
S. Park, Q. Li, and S. Levine · 2025
Closest in time.
Energy-weighted flow matching for offline reinforcement learning
S. Zhang, W. Zhang, and Q. Gu · 2025
Closest in time.