Fetching the paper…
Reading the bibliography…
Model-free reinforcement learning algorithms have exhibited great potential in solving single-task sequential decision-making problems with high-dimensional observations and long horizons, but are known to be hard to generalize across tasks.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
R. S. Sutton · 1990
Earlier work this paper cites.
Q-learning
C. J. C. H. Watkins and P. Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Learning to achieve goals
L. P. Kaelbling · 1993
Earlier work this paper cites.
Actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 1999
Earlier work this paper cites.
Random projection in dimensionality reduction: applications to image and text data
E. Bingham and H. Mannila · 2001
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi and B. Recht · 2007
Earlier work this paper cites.
Weighted sums of random kitchen sinks: Replacing minimization with randomization in learning
A. Rahimi and B. Recht · 2008
Earlier work this paper cites.
PILCO: A model-based and data-efficient approach to policy search
M. P. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Horde: a scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
R. S. Sutton, J. Modayil, M. Delp, T. Degris, P. M. Pilarski, A. White, and D. Precup · 2011
Earlier work this paper cites.
Direct trajectory optimization of rigid body dynamical systems through contact
M. Posa and R. Tedrake · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. A. Riedmiller · 2013
Earlier work this paper cites.
CHOMP: covariant hamiltonian optimization for motion planning
M. Zucker, N. D. Ratliff, A. D. Dragan, M. Pivtoraiko, M. Klingensmith, C. M. Dellin, J. A. Bagnell, and S. S. Srinivasa · 2013
Earlier work this paper cites.
Learning complex neural network policies with trajectory optimization
S. Levine and V. Koltun · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. P. Lillicrap, T. Erez, and Y. Tassa · 2015
Earlier work this paper cites.
Trust region policy optimization
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz · 2015
Earlier work this paper cites.
Improving multi-step prediction of learned time series models
A. Venkatraman, M. Hebert, and J. A. Bagnell · 2015
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Cited alongside, same era.
Aggressive driving with model predictive path integral control
G. Williams, P. Drews, B. Goldfain, J. M. Rehg, and E. A. Theodorou · 2016
Cited alongside, same era.
Hindsight experience replay
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
Successor features for transfer in reinforcement learning
A. Barreto, W. Dabney, R. Munos, J. J. Hunt, T. Schaul, H. P. van Hasselt, and D. Silver · 2017
Cited alongside, same era.
Multi-task deep reinforcement learning with popart
M. Hessel, H. Soyer, L. Espeholt, W. Czarnecki, S. Schmitt, and H. van Hasselt · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Later among the works it cites.
Towards a unified analysis of random Fourier features
Z. Li, J.-F. Ton, D. Oglic, and D. Sejdinovic · 2019
Later among the works it cites.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
O. Nachum, Y. Chow, B. Dai, and L. Li · 2019
Later among the works it cites.
Skew-fit: State-covering self-supervised reinforcement learning
V. H. Pong, M. Dalal, S. Lin, A. Nair, S. Bahl, and S. Levine · 2019
Later among the works it cites.
Learning to combat compounding-error in model-based reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Finn, P. Abbeel, and S. Levine · 2017
Cited alongside, same era.
Generalization properties of learning with random features
A. Rudi and L. Rosasco · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm
D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. P. Lillicrap, K. Simonyan, and D. Hassabis · 2017
Cited alongside, same era.
Distral: Robust multitask reinforcement learning
Y. W. Teh, V. Bapst, W. M. Czarnecki, J. Quan, J. Kirkpatrick, R. Hadsell, N. Heess, and R. Pascanu · 2017
Cited alongside, same era.
Model predictive path integral control: From theory to parallel computation
G. Williams, A. Aldrich, and E. A. Theodorou · 2017
Cited alongside, same era.
Deep reinforcement learning with successor features for navigation across similar environments
J. Zhang, J. T. Springenberg, J. Boedecker, and W. Burgard · 2017
Cited alongside, same era.
C. Xiao, Y. Wu, C. Ma, D. Schuurmans, and M. Müller · 2019
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2019
Later among the works it cites.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. P. Lillicrap, J. Ba, and M. Norouzi · 2020
Later among the works it cites.
Gamma-models: Generative temporal difference learning for infinite-horizon prediction
M. Janner, I. Mordatch, and S. Levine · 2020
Later among the works it cites.
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
Successor features combine elements of model-free and model-based reinforcement learning
L. Lehnert and M. L. Littman · 2020
Later among the works it cites.
Model-based reinforcement learning: A survey
T. M. Moerland, J. Broekens, and C. M. Jonker · 2020
Later among the works it cites.
Gendice: Generalized offline estimation of stationary values
R. Zhang, B. Dai, L. Li, and D. Schuurmans · 2020
Later among the works it cites.
Psiphi-learning: Reinforcement learning with demonstrations using successor features and inverse temporal difference learning
A. Filos, C. Lyle, Y. Gal, S. Levine, N. Jaques, and G. Farquhar · 2021
Later among the works it cites.
Learning to reach goals via iterated supervised learning
D. Ghosh, A. Gupta, A. Reddy, J. Fu, C. M. Devin, B. Eysenbach, and S. Levine · 2021
Later among the works it cites.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale
D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Later among the works it cites.
Model-based reinforcement learning via latent-space collocation
O. Rybkin, C. Zhu, A. Nagabandi, K. Daniilidis, I. Mordatch, and S. Levine · 2021
Later among the works it cites.
Investigating compounding prediction errors in learned dynamics models
N. O. Lambert, K. S. J. Pister, and R. Calandra · 2022
Later among the works it cites.
Bundled gradients through contact via randomized smoothing
H. J. T. Suh, T. Pang, and R. Tedrake · 2022
Later among the works it cites.