Fetching the paper…
Reading the bibliography…
Image and video generative models that are pre-trained on Internet-scale data can greatly increase the generalization capacity of robot learning systems.
J. Schmidhuber, “Learning to generate sub-goals for action sequences,” in Artificial neural networks , 1991, pp. 967–972
1991
Earlier work this paper cites.
P. Dayan and G. E. Hinton, “Feudal reinforcement learning,” Advances in neural information processing systems , vol. 5, 1992
1992
Earlier work this paper cites.
L. P. Kaelbling, “Learning to achieve goals,” in IJCAI , vol. 2. Citeseer, 1993, pp. 1094–8
1993
Earlier work this paper cites.
R. S. Sutton, D. Precup, and S. Singh, “Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning,” Artificial intelligence , vol. 112, no. 1-2, pp. 181–211, 1999
1999
Earlier work this paper cites.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in International conference on machine learning . PMLR, 2015, pp. 2256–2265
2015
Earlier work this paper cites.
T. Schaul, D. Horgan, K. Gregor, and D. Silver, “Universal value function approximators,” in International conference on machine learning . PMLR, 2015, pp. 1312–1320
2015
Earlier work this paper cites.
T. D. Kulkarni, K. Narasimhan, A. Saeedi, and J. Tenenbaum, “Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation,” Advances in neural information processing systems , vol. 29, 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. Pieter Abbeel, and W. Zaremba, “Hindsight experience replay,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
P.-L. Bacon, J. Harb, and D. Precup, “The option-critic architecture,” in Proceedings of the AAAI conference on artificial intelligence , vol. 31, no. 1, 2017
2017
Earlier work this paper cites.
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu, “Feudal networks for hierarchical reinforcement learning,” in International conference on machine learning . PMLR, 2017, pp. 3540–3549
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel, “Domain randomization for transferring deep neural networks from simulation to the real world,” International Conference on Intelligent Robots and Systems , 2017
2017
Earlier work this paper cites.
R. Goyal, S. E. Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, and et al., “The” something something” video database for learning and evaluating visual common sense,” in IEEE international conference on computer vision (ICCV) , 2017
2017
Earlier work this paper cites.
O. Nachum, S. S. Gu, H. Lee, and S. Levine, “Data-efficient hierarchical reinforcement learning,” Advances in neural information processing systems , vol. 31, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville, “Film: Visual reasoning with a general conditioning layer,” in Proceedings of the AAAI conference on artificial intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn, “Robonet: Large-scale multi-robot learning,” in Conference on Robot Learning (CoRL) , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
B. Eysenbach, R. R. Salakhutdinov, and S. Levine, “Search on the replay buffer: Bridging planning and reinforcement learning,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Nasiriany, V. Pong, S. Lin, and S. Levine, “Planning with goal-conditioned policies,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
Z. Huang, F. Liu, and H. Su, “Mapping state space using landmarks for universal goal reaching,” Advances in Neural Information Processing Systems , vol. 32, 2019
2019
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems , vol. 33, pp. 6840–6851, 2020
2020
Earlier work this paper cites.
S. Tellex, N. Gopalan, H. Kress-Gazit, and C. Matuszek, “Robots that use language,” Annual Review of Control, Robotics, and Autonomous Systems , vol. 3, pp. 25–55, 2020
2020
Earlier work this paper cites.
S. Stepputtis, J. Campbell, M. Phielipp, S. Lee, C. Baral, and H. Ben Amor, “Language-conditioned imitation learning for robot manipulation tasks,” Advances in Neural Information Processing Systems , vol. 33, pp. 13 139–13 150, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. Mandlekar, F. Ramos, B. Boots, S. Savarese, L. Fei-Fei, A. Garg, and D. Fox, “Iris: Implicit reinforcement without interaction at scale for learning control from offline robot manipulation data,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2020, pp. 4414–4420
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
C. Lynch, M. Khansari, T. Xiao, V. Kumar, J. Tompson, S. Levine, and P. Sermanet, “Learning latent plans from play,” in Conference on Robot Learning (CoRL) . PMLR, 2020, pp. 1113–1132
2020
Earlier work this paper cites.
T. Zhang, S. Guo, T. Tan, X. Hu, and F. Chen, “Generating adjacency-constrained subgoals in hierarchical reinforcement learning,” Advances in neural information processing systems , vol. 33, pp. 21 579–21 590, 2020
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard, “Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,” in IEEE Robotics and Automation Letters (RAL) , 2021
2021
Cited alongside, same era.
S. K. S. Ghasemipour, D. Schuurmans, and S. S. Gu, “Emaq: Expected-max q-learning operator for simple yet effective offline and online rl,” in International Conference on Machine Learning . PMLR, 2021, pp. 3682–3691
2021
Cited alongside, same era.
2021
Cited alongside, same era.
A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian et al. , “Do as i can, not as i say: Grounding language in robotic affordances,” in Conference on robot learning . PMLR, 2023, pp. 287–318
2023
Later among the works it cites.
K. Lin, C. Agia, T. Migimatsu, M. Pavone, and J. Bohg, “Text2motion: From natural language instructions to feasible plans,” Autonomous Robots , vol. 47, no. 8, pp. 1345–1365, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Nair, E. Mitchell, K. Chen, B. Ichter, S. Savarese, and C. Finn, “Learning language-conditioned robot behavior from offline data and crowd-sourced annotation,” Conference on Robot Learning (CoRL) , 2021
2021
Cited alongside, same era.
K. Pertsch, Y. Lee, and J. Lim, “Accelerating reinforcement learning with learned skill priors,” in Conference on robot learning . PMLR, 2021, pp. 188–204
2021
Cited alongside, same era.
E. Chane-Sane, C. Schmid, and I. Laptev, “Goal-conditioned reinforcement learning with imagined subgoals,” in International Conference on Machine Learning . PMLR, 2021, pp. 1430–1440
2021
Cited alongside, same era.
C. Hoang, S. Sohn, J. Choi, W. Carvalho, and H. Lee, “Successor feature landmarks for long-horizon goal-conditioned reinforcement learning,” Advances in neural information processing systems , vol. 34, pp. 26 963–26 975, 2021
2021
Cited alongside, same era.
J. Kim, Y. Seo, and J. Shin, “Landmark-guided subgoal generation in hierarchical reinforcement learning,” Advances in neural information processing systems , vol. 34, pp. 28 336–28 349, 2021
2021
Cited alongside, same era.
L. Zhang, G. Yang, and B. C. Stadie, “World model as a graph: Learning latent landmarks for planning,” in International conference on machine learning . PMLR, 2021, pp. 12 611–12 620
2021
Cited alongside, same era.
2021
Cited alongside, same era.
P. Dhariwal and A. Nichol, “Diffusion models beat gans on image synthesis,” Advances in neural information processing systems , vol. 34, pp. 8780–8794, 2021
2021
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
O. Mees, J. Borja-Diaz, and W. Burgard, “Grounding language with visual affordances over unstructured data,” in Proceedings of the IEEE International Conference on Robotics and Automation (ICRA) , London, UK, 2023
2023
Later among the works it cites.
E. Rosete-Beas, O. Mees, G. Kalweit, J. Boedecker, and W. Burgard, “Latent plans for task-agnostic offline reinforcement learning,” in Conference on Robot Learning . PMLR, 2023, pp. 1838–1849
2023
Later among the works it cites.
2023
Later among the works it cites.
K. Fang, P. Yin, A. Nair, H. R. Walke, G. Yan, and S. Levine, “Generalization with lossy affordances: Leveraging broad offline data for learning visuomotor tasks,” in Conference on Robot Learning . PMLR, 2023, pp. 106–117
2023
Later among the works it cites.
T. Brooks, A. Holynski, and A. A. Efros, “Instructpix2pix: Learning to follow image editing instructions,” in Conference on Computer Vision and Pattern Recognition (CVPR) , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
Y. Du, S. Yang, B. Dai, H. Dai, O. Nachum, J. Tenenbaum, D. Schuurmans, and P. Abbeel, “Learning universal policies via text-guided video generation,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
A. Ajay, S. Han, Y. Du, S. Li, A. Gupta, T. Jaakkola, J. Tenenbaum, L. Kaelbling, A. Srivastava, and P. Agrawal, “Compositional foundation models for hierarchical planning,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
J. Gao, K. Hu, G. Xu, and H. Xu, “Can pre-trained text-to-image models generate visual goals for reinforcement learning?” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y. Tan, L. Y. Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine, “Octo: An open-source generalist robot policy,” in Proceedings of Robotics: Science and Systems , Delft, Netherlands, 2024
2024
Closest in time.
R. Doshi, H. Walke, O. Mees, S. Dasari, and S. Levine, “Scaling cross-embodied learning: One policy for manipulation, navigation, locomotion and aviation,” in Conference on Robot Learning , 2024
2024
Closest in time.
M. Zawalski, W. Chen, K. Pertsch, O. Mees, C. Finn, and S. Levine, “Robotic control via embodied chain-of-thought reasoning,” in Conference on Robot Learning , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Z. Zhou, P. Atreya, A. Lee, H. Walke, O. Mees, and S. Levine, “Autonomous improvement of instruction following skills via foundation models,” in Conference on Robot Learning , 2024
2024
Closest in time.
M. Nakamoto, O. Mees, A. Kumar, and S. Levine, “Steering your generalists: Improving robotic foundation models via value guidance,” Conference on Robot Learning (CoRL) , 2024
2024
Closest in time.
2024
Closest in time.
A. Z. Ren, J. Clark, A. Dixit, M. Itkina, A. Majumdar, and D. Sadigh, “Explore until confident: Efficient exploration for embodied question answering,” in Robotics Science and Systems (RSS) , 2024
2024
Closest in time.
V. Myers, B. C. Zheng, O. Mees, S. Levine, and K. Fang, “Policy adaptation via language optimization: Decomposing tasks for few-shot imitation,” in Conference on Robot Learning , 2024
2024
Closest in time.
N. Hirose, C. Glossop, A. Sridhar, D. Shah, O. Mees, and S. Levine, “Lelan: Learning a language-conditioned navigation policy from in-the-wild video,” in Conference on Robot Learning , 2024
2024
Closest in time.
S. Park, D. Ghosh, B. Eysenbach, and S. Levine, “Hiql: Offline goal-conditioned rl with latent states as actions,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
S. Fujimoto, D. Meger, and D. Precup, “Off-policy deep reinforcement learning without exploration,” in International conference on machine learning . PMLR, 2019, pp. 2052–2062
2062
Closest in time.