Fetching the paper…
Reading the bibliography…
We introduce ReWiND, a framework for learning robot manipulation tasks solely from language instructions without per-task demonstrations.
Algorithms for inverse reinforcement learning
A. Y. Ng and S. J. Russell · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. Maas, J. A. Bagnell, and A. K. Dey · 2008
Earlier work this paper cites.
Guided cost learning: Deep inverse optimal control via policy optimization
C. Finn, S. Levine, and P. Abbeel · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
D. Sadigh, A. D. Dragan, S. S. Sastry, and S. A. Seshia · 2017
Earlier work this paper cites.
Reset-free guided policy search: Efficient deep reinforcement learning with stochastic initial states
W. Montgomery, A. Ajay, C. Finn, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Active reward learning from critiques
Y. Cui and S. Niekum · 2018
Earlier work this paper cites.
Learning from physical human corrections, one feature at a time
A. Bajcsy, D. P. Losey, M. K. O’Malley, and A. D. Dragan · 2018
Earlier work this paper cites.
Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic · 2019
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, A. Narayan, H. Shively, A. Bellathur, K. Hausman, C. Finn, and S. Levine · 2019
Earlier work this paper cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
X. B. Peng, A. Kumar, G. Zhang, and S. Levine · 2019
Earlier work this paper cites.
Active preference-based gaussian process regression for reward learning
E. Biyik, N. Huynh, M. J. Kochenderfer, and D. Sadigh · 2020
Earlier work this paper cites.
Scaling data-driven robotics with reward sketching and batch reinforcement learning
S. Cabi, S. G. Colmenarejo, A. Novikov, K. Konyushkova, S. Reed, R. Jeong, K. Zolna, Y. Aytar, D. Budden, M. Vecerik, et al · 2020
Earlier work this paper cites.
Accelerating reinforcement learning with learned skill priors
K. Pertsch, Y. Lee, and J. J. Lim · 2020
Earlier work this paper cites.
Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training
K. Lee, L. Smith, and P. Abbeel · 2021
Earlier work this paper cites.
Learning multimodal rewards from rankings
V. Myers, E. Biyik, N. Anari, and D. Sadigh · 2021
Earlier work this paper cites.
Learning reward functions from scale feedback
N. Wilde, E. Biyik, D. Sadigh, and S. L. Smith · 2021
Earlier work this paper cites.
Learning generalizable robotic reward functions from ”in-the-wild” human videos
A. S. Chen, S. Nair, and C. Finn · 2021
Earlier work this paper cites.
Reset-free reinforcement learning via multi-task learning: Learning dexterous manipulation behaviors without human intervention
A. Gupta, J. Yu, T. Z. Zhao, V. Kumar, A. Rovinsky, K. Xu, T. Devlin, and S. Levine · 2021
Earlier work this paper cites.
Deep reinforcement learning at the edge of the statistical precipice
R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. Bellemare · 2021
Earlier work this paper cites.
BC-z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2021
Earlier work this paper cites.
Few-shot preference learning for human-in-the-loop rl
J. Hejna and D. Sadigh · 2022
Earlier work this paper cites.
Can foundation models perform zero-shot task specification for robot manipulation?
Y. Cui, S. Niekum, A. Gupta, V. Kumar, and A. Rajeswaran · 2022
Earlier work this paper cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
L. Fan, G. Wang, Y. Jiang, A. Mandlekar, Y. Yang, H. Zhu, A. Tang, D.-A. Huang, Y. Zhu, and A. Anandkumar · 2022
Earlier work this paper cites.
Offline reinforcement learning with implicit q-learning
I. Kostrikov, A. Nair, and S. Levine · 2022
Cited alongside, same era.
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
D. Damen, H. Doughty, G. M. Farinella, A. Furnari, J. Ma, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray · 2022
Cited alongside, same era.
Bottom-up skill discovery from unsegmented demonstrations for long-horizon robot manipulation
Y. Zhu, P. Stone, and Y. Zhu · 2022
Cited alongside, same era.
Roboclip: One demonstration is enough to learn robot policies
S. A. Sontakke, J. Zhang, S. Arnold, K. Pertsch, E. Biyik, D. Sadigh, C. Finn, and L. Itti · 2023
Cited alongside, same era.
Reward design with language models
M. Kwon, S. M. Xie, K. Bullard, and D. Sadigh · 2023
Cited alongside, same era.
Language instructed reinforcement learning for human-ai coordination
Reinforcement learning with foundation priors: Let embodied agent efficiently learn on its own
W. Ye, Y. Zhang, H. Weng, X. Gu, S. Wang, T. Zhang, M. Wang, P. Abbeel, and Y. Gao · 2024
Later among the works it cites.
Contrast sets for evaluating language-guided robot policies
A. Anwar, R. Gupta, and J. Thomason · 2024
Later among the works it cites.
Dinov2: Learning robust visual features without supervision, 2024
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y. Huang, S.-W. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski · 2024
Later among the works it cites.
Real-world offline reinforcement learning from vision language model feedback
S. Venkataraman, Y. Wang, Z. Wang, Z. Erickson, and D. Held · 2024
Later among the works it cites.
Sprint: Scalable policy pre-training via language instruction relabeling
J. Zhang, K. Pertsch, J. Zhang, and J. J. Lim · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Hu and D. Sadigh · 2023
Cited alongside, same era.
Language to rewards for robotic skill synthesis
W. Yu, N. Gileadi, C. Fu, S. Kirmani, K.-H. Lee, M. Gonzalez Arenas, H.-T. Lewis Chiang, T. Erez, L. Hasenclever, J. Humplik, B. Ichter, T. Xiao, P. Xu, A. Zeng, T. Zhang, N. Heess, D. Sadigh, J. Tan, Y. Tassa, and F. Xia · 2023
Cited alongside, same era.
Do embodied agents dream of pixelated sheep?: Embodied decision making using language guided world modelling
K. Nottingham, P. Ammanabrolu, A. Suhr, Y. Choi, H. Hajishirzi, S. Singh, and R. Fox · 2023
Cited alongside, same era.
Lift: Unsupervised reinforcement learning with foundation models as teachers
T. Nam, J. Lee, J. Zhang, S. J. Hwang, J. J. Lim, and K. Pertsch · 2023
Cited alongside, same era.
Bootstrap your own skills: Learning to solve new tasks with large language model guidance
J. Zhang, J. Zhang, K. Pertsch, Z. Liu, X. Ren, M. Chang, S.-H. Sun, and J. J. Lim · 2023
Cited alongside, same era.
Liv: Language-image representations and rewards for robotic control
Y. J. Ma, W. Liang, V. Som, V. Kumar, A. Zhang, O. Bastani, and D. Jayaraman · 2023
Cited alongside, same era.
Bridgedata v2: A dataset for robot learning at scale
H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du, A. Lee, K. Fang, C. Finn, and S. Levine · 2023
Cited alongside, same era.
Later among the works it cites.
Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning
J. Luo, C. Xu, J. Wu, and S. Levine · 2024
Later among the works it cites.
Gemini: A family of highly capable multimodal models
G. Team · 2024
Later among the works it cites.
Lerobot: State-of-the-art machine learning for real-world robotics in pytorch
R. Cadene, S. Alibert, A. Soare, Q. Gallouedec, A. Zouitine, and T. Wolf · 2024
Later among the works it cites.
OpenVLA: An open-source vision-language-action model
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn · 2024
Later among the works it cites.
Steering your generalists: Improving robotic foundation models via value guidance
M. Nakamoto, O. Mees, A. Kumar, and S. Levine · 2024
Later among the works it cites.
TAIL: Task-specific adapters for imitation learning with large pretrained models
Z. Liu, J. Zhang, K. Asadi, Y. Liu, D. Zhao, S. Sabach, and R. Fakoor · 2024
Later among the works it cites.
So you think you can scale up autonomous robot data collection?
S. Mirchandani, S. Belkhale, J. Hejna, E. Choi, M. S. Islam, and D. Sadigh · 2024
Later among the works it cites.
Autonomous improvement of instruction following skills via foundation models
Z. Zhou, P. Atreya, A. Lee, H. R. Walke, O. Mees, and S. Levine · 2024
Later among the works it cites.
V-former: Offline RL with temporally-extended actions, 2024
J. Wu, S. Park, Z. Lin, J. Luo, and S. Levine · 2024
Later among the works it cites.
Hamster: Hierarchical action models for open-world robot manipulation
Y. Li, Y. Deng, J. Zhang, J. Jang, M. Memmel, C. R. Garrett, F. Ramos, D. Fox, A. Li, A. Gupta, and A. Goyal · 2025
Closest in time.
Video-language critic: Transferable reward functions for language-conditioned robotics
M. Alakuijala, R. McLean, I. Woungang, N. Farsad, S. Kaski, P. Marttinen, and K. Yuan · 2025
Closest in time.
Mile: Model-based intervention learning
Y. Korkmaz and E. Bıyık · 2025
Closest in time.
Robotic-clip: Fine-tuning clip on action data for robotic applications
N. Nguyen, M. N. Vu, T. D. Ta, B. Huang, T. Vo, N. Le, and A. Nguyen · 2025
Closest in time.
Vision language models are in-context value learners
Y. J. Ma, J. Hejna, A. Wahid, C. Fu, D. Shah, J. Liang, Z. Xu, S. Kirmani, P. Xu, D. Driess, T. Xiao, J. Tompson, O. Bastani, D. Jayaraman, W. Yu, T. Zhang, D. Sadigh, and F. Xia · 2025
Closest in time.
Subtask-aware visual reward learning from segmented demonstrations
C. Kim, M. Heo, D. Lee, J. Shin, H. Lee, J. J. Lim, and K. Lee · 2025
Closest in time.
VICtor: Learning hierarchical vision-instruction correlation rewards for long-horizon manipulation
K.-H. Hung, P.-C. Lo, J.-F. Yeh, H.-Y. Hsu, Y.-T. Chen, and W. H. Hsu · 2025
Closest in time.
A taxonomy for evaluating generalist robot policies
J. Gao, S. Belkhale, S. Dasari, A. Balakrishna, D. Shah, and D. Sadigh · 2025
Closest in time.
Steering your diffusion policy with latent space reinforcement learning
A. Wagenmaker, M. Nakamoto, Y. Zhang, S. Park, W. Yagoub, A. Nagabandi, A. Gupta, and S. Levine · 2025
Closest in time.
Improving vision-language-action model with online reinforcement learning
Y. Guo, J. Zhang, X. Chen, X. Ji, Y.-J. Wang, Y. Hu, and J. Chen · 2025
Closest in time.
Vla-rl: Towards masterful and general robotic manipulation with scalable reinforcement learning
G. Lu, W. Guo, C. Zhang, Y. Zhou, H. Jiang, Z. Gao, Y. Tang, and Z. Wang · 2025
Closest in time.
Expo: Stable reinforcement learning with expressive policies
P. Dong, Q. Li, D. Sadigh, and C. Finn · 2025
Closest in time.
ConRFT: A reinforced fine-tuning method for vla models via consistency policy
Y. Chen, S. Tian, S. Liu, Y. Zhou, H. Li, and D. Zhao · 2025
Closest in time.
Efficient evaluation of multi-task robot policies with active experiment selection
A. Anwar, R. Gupta, Z. Merchant, S. Ghosh, W. Neiswanger, and J. Thomason · 2025
Closest in time.
Efficient online reinforcement learning fine-tuning need not retain offline data
Z. Zhou, A. Peng, Q. Li, S. Levine, and A. Kumar · 2025
Closest in time.
Reinforcement learning with action chunking
Q. Li, Z. Zhou, and S. Levine · 2025
Closest in time.