Fetching the paper…
Reading the bibliography…
Several recent works have proposed a class of algorithms for the offline reinforcement learning (RL) problem that we will refer to as return-conditioned supervised learning (RCSL).
Learning to achieve goals
L. P. Kaelbling · 1993
Earlier work this paper cites.
Theory of classification: A survey of some recent advances
S. Boucheron, O. Bousquet, and G. Lugosi · 2005
Earlier work this paper cites.
Convexity, classification, and risk bounds
P. L. Bartlett, M. I. Jordan, and J. D. McAuliffe · 2006
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. P. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Constrained policy optimization
J. Achiam, D. Held, A. Tamar, and P. Abbeel · 2017
Earlier work this paper cites.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. Pieter Abbeel, and W. Zaremba · 2017
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
R. S. Sutton and A. G. Barto · 2018
Earlier work this paper cites.
Y. Tassa, Y. Doron, A. Muldal, T. Erez, Y. Li, D. d. L. Casas, D. Budden, A. Abdolmaleki, J. Merel, A. Lefrancq, et al · 2018
Cited alongside, same era.
Information-theoretic considerations in batch reinforcement learning
J. Chen and N. Jiang · 2019
Cited alongside, same era.
Learning to reach goals via iterated supervised learning
D. Ghosh, A. Gupta, A. Reddy, J. Fu, C. Devin, B. Eysenbach, and S. Levine · 2019
Cited alongside, same era.
A. Kumar, X. B. Peng, and S. Levine · 2019
Cited alongside, same era.
Reinforcement learning upside down: Don’t predict rewards–just map them to actions
J. Schmidhuber · 2019
A minimalist approach to offline reinforcement learning
S. Fujimoto and S. S. Gu · 2021
Later among the works it cites.
Generalized decision transformer for offline hindsight information matching
H. Furuta, Y. Matsuo, and S. S. Gu · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
M. Janner, Q. Li, and S. Levine · 2021
Later among the works it cites.
Jaxrl: Implementations of reinforcement learning algorithms in jax., 10 2021
I. Kostrikov · 2021
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
I. Kostrikov, A. Nair, and S. Levine · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Training agents using upside-down reinforcement learning
R. K. Srivastava, P. Shyam, F. Mutz, W. Jaśkowski, and J. Schmidhuber · 2019
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
Flax: A neural network library and ecosystem for JAX, 2020
J. Heek, A. Levskaya, A. Oliver, M. Ritter, B. Rondepierre, A. Steiner, and M. van Zee · 2020
Cited alongside, same era.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Cited alongside, same era.
Rvs: What is essential for offline rl via supervised learning?
S. Emmons, B. Eysenbach, I. Kostrikov, and S. Levine · 2021
Cited alongside, same era.
What are the statistical limits of offline rl with linear function approximation?, 2020a
R. Wang, D. P. Foster, and S. M. Kakade
Cited in the paper.
Bellman-consistent pessimism for offline reinforcement learning
T. Xie, C.-A. Cheng, N. Jiang, P. Mineiro, and A. Agarwal · 2021
Later among the works it cites.
Distributional Reinforcement Learning
M. G. Bellemare, W. Dabney, and M. Rowland · 2022
Closest in time.
You can’t count on luck: Why decision transformers fail in stochastic environments
K. Paster, S. McIlraith, and J. Ba · 2022
Closest in time.
Dichotomy of control: Separating what you can control from what you cannot
M. Yang, D. Schuurmans, P. Abbeel, and O. Nachum · 2022
Closest in time.
Upside-down reinforcement learning can diverge in stochastic environments with episodic resets, 2022
M. Štrupl, F. Faccio, D. R. Ashley, J. Schmidhuber, and R. K. Srivastava · 2022
Closest in time.