Fetching the paper…
Reading the bibliography…
A long-standing goal in AI is to develop agents capable of solving diverse tasks across a range of environments, including those never seen during training.
Robonet: Large-scale multi-robot learning
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn · 1910
Earlier work this paper cites.
Experience-embedded visual foresight
L. Yen-Chen, M. Bauza, and P. Isola · 1911
Earlier work this paper cites.
Optimal control-1950 to 1985
A. E. Bryson · 1996
Earlier work this paper cites.
Model predictive control: past, present and future
M. Morari and J. H. Lee · 1999
Earlier work this paper cites.
Predictive coding for locally-linear control
R. Shu, T. Nguyen, Y. Chow, T. Pham, K. Than, M. Ghavamzadeh, S. Ermon, and H. H. Bui · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
D. Ernst, P. Geurts, and L. Wehenkel · 2005
Earlier work this paper cites.
A generalized iterative lqg method for locally-optimal feedback control of constrained nonlinear stochastic systems
E. Todorov and W. Li · 2005
Earlier work this paper cites.
A generalized iterative lqg method for locally-optimal feedback control of constrained nonlinear stochastic systems
E. Todorov and W. Li · 2005
Earlier work this paper cites.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, T. X. Ma, S. Ermon, J. Zou, and C. Finn · 2005
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2006
Earlier work this paper cites.
Receding horizon differential dynamic programming
Y. Tassa, T. Erez, and W. Smart · 2007
Earlier work this paper cites.
Modeling and optimal control of human-like running
G. Schultz and K. Mombaur · 2009
Earlier work this paper cites.
Dynamic programming and optimal control: Volume I , volume 4
D. Bertsekas · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
K. Cho · 2014
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
M. Watter, J. T. Springenberg, J. Boedecker, and M. Riedmiller · 2015
Earlier work this paper cites.
Model predictive path integral control using covariance variable importance sampling
G. Williams, A. Aldrich, and E. Theodorou · 2015
Earlier work this paper cites.
Deep spatial autoencoders for visuomotor learning
C. Finn, X. Y. Tan, Y. Duan, T. Darrell, S. Levine, and P. Abbeel · 2016
Earlier work this paper cites.
Optimization-based locomotion planning, estimation, and control design for the atlas humanoid robot
S. Kuindersma, R. Deits, M. Fallon, A. Valenzuela, H. Dai, F. Permenter, T. Koolen, P. Marion, and R. Tedrake · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, et al · 2016
Earlier work this paper cites.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. Pieter Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
Successor features for transfer in reinforcement learning
A. Barreto, W. Dabney, R. Munos, J. J. Hunt, T. Schaul, H. P. van Hasselt, and D. Silver · 2017
Earlier work this paper cites.
Deep visual foresight for planning robot motion
C. Finn and S. Levine · 2017
Earlier work this paper cites.
Robust locally-linear controllable embedding
E. Banijamali, R. Shu, M. Ghavamzadeh, H. Bui, and A. Ghodsi · 2018
Earlier work this paper cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
F. Ebert, C. Finn, S. Dasari, A. Xie, A. Lee, and S. Levine · 2018
Earlier work this paper cites.
Learning latent 423 dynamics for planning from pixels
D. Hafner, T. Lillicrap, I. Fischer, R. Villegas, D. Ha, H. Lee, and J. Davidson · 2018
Cited alongside, same era.
State representation learning for control: An overview
T. Lesort, N. Díaz-Rodríguez, J.-F. Goudou, and D. Filliat · 2018
Cited alongside, same era.
Learning dexterous in-hand manipulation. corr abs/1808.00177 (2018)
M. A. OpenAI, B. Baker, M. Chociej, R. Józefowicz, B. McGrew, J. W. Pachocki, J. Pachocki, A. Petron, M. Plappert, G. Powell, et al · 2018
Cited alongside, same era.
Learning what you can do before doing anything
O. Rybkin, K. Pertsch, K. G. Derpanis, K. Daniilidis, and A. Jaegle · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
R. S. Sutton · 2018
Cited alongside, same era.
Temporal difference learning for model predictive control
N. Hansen, X. Wang, and H. Su · 2022
Later among the works it cites.
Example-based offline reinforcement learning without rewards
K. Hatch, T. Yu, R. Rafailov, and C. Finn · 2022
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
I. Kostrikov, A. Nair, and S. Levine · 2022
Later among the works it cites.
Investigating compounding prediction errors in learned dynamics models
N. Lambert, K. Pister, and R. Calandra · 2022
Later among the works it cites.
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27
Y. LeCun · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The laplacian in rl: Learning representations with efficient approximations
Y. Wu, G. Tucker, and O. Nachum · 2018
Cited alongside, same era.
Reinforcement Learning and Optimal Control
D. P. Bertsekas · 2019
Cited alongside, same era.
Learning to reach goals via iterated supervised learning
D. Ghosh, A. Gupta, A. Reddy, J. Fu, C. Devin, B. Eysenbach, and S. Levine · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Cited alongside, same era.
Solar: Deep structured representations for model-based reinforcement learning
M. Zhang, S. Vikram, L. Smith, P. Abbeel, M. J. Johnson, and S. Levine · 2019
Cited alongside, same era.
Model-predictive control via cross-entropy and gradient-based optimization
H. Bharadhwaj, K. Xie, and F. Shkurti · 2020
Cited alongside, same era.
The importance of pessimism in fixed-dataset policy optimization
J. Buckman, C. Gelada, and M. G. Bellemare · 2020
Cited alongside, same era.
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta · 2022
Later among the works it cites.
Joint embedding predictive architectures focus on slow features
V. Sobal, J. SV, S. Jalagam, N. Carion, K. Cho, and Y. LeCun · 2022
Later among the works it cites.
Does zero-shot reinforcement learning exist?
A. Touati, J. Rapin, and Y. Ollivier · 2022
Later among the works it cites.
Reachability-aware laplacian representation in reinforcement learning
K. Wang, K. Zhou, J. Feng, B. Hooi, and X. Wang · 2022
Later among the works it cites.
Don’t change the algorithm, change the data: Exploratory data for offline reinforcement learning
D. Yarats, D. Brandfonbrener, H. Liu, M. Laskin, P. Abbeel, A. Lazaric, and L. Pinto · 2022
Later among the works it cites.
How to leverage unlabeled data in offline reinforcement learning
T. Yu, A. Kumar, Y. Chebotar, K. Hausman, C. Finn, and S. Levine · 2022
Later among the works it cites.
Light-weight probing of unsupervised representations for reinforcement learning
W. Zhang, A. GX-Chen, V. Sobal, Y. LeCun, and N. Carion · 2022
Later among the works it cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al · 2023
Later among the works it cites.
Mastering diverse domains through world models
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap · 2023
Later among the works it cites.
Td-mpc2: Scalable, robust world models for continuous control
N. Hansen, H. Su, and X. Wang · 2023
Later among the works it cites.
The provable benefits of unsupervised data sharing for offline reinforcement learning
H. Hu, Y. Yang, Q. Zhao, and C. Zhang · 2023
Later among the works it cites.
Dinov2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al · 2023
Later among the works it cites.
Learning by reconstruction produces uninformative features for perception
R. Balestriero and Y. LeCun · 2024
Later among the works it cites.
π 0 \pi_{0} : A vision-language-action flow model for general robot control
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, et al · 2024
Later among the works it cites.
Closing the gap between td learning and supervised learning–a generalisation point of view
R. Ghugare, M. Geist, G. Berseth, and B. Eysenbach · 2024
Later among the works it cites.
Unsupervised-to-online reinforcement learning
J. Kim, S. Park, and S. Levine · 2024
Later among the works it cites.
How jepa avoids noisy features: The implicit bias of deep linear self distillation networks
E. Littwin, O. Saremi, M. Advani, V. Thilak, P. Nakkiran, C. Huang, and J. Susskind · 2024
Later among the works it cites.
Octo: An open-source generalist robot policy
Octo Model Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, C. Xu, J. Luo, T. Kreiman, Y. Tan, L. Y. Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine · 2024
Later among the works it cites.
A survey of imitation learning: Algorithms, recent developments, and challenges
M. Zare, P. M. Kebria, A. Khosravi, and S. Nahavandi · 2024
Later among the works it cites.
Dino-wm: World models on pre-trained visual features enable zero-shot planning
G. Zhou, H. Pan, Y. LeCun, and L. Pinto · 2024
Later among the works it cites.