Fetching the paper…
Reading the bibliography…
This position paper proposes a fresh look at Reinforcement Learning (RL) from the perspective of data-efficiency.
R. Fakoor, P. Chaudhari, S. Soatto, and A. J. Smola · 1910
Earlier work this paper cites.
Imagined value gradients: Model-based policy optimization with transferable latent dynamics models
A. Byravan, J. T. Springenberg, A. Abdolmaleki, R. Hafner, M. Neunert, T. Lampe, N. Y. Siegel, N. Heess, and M. A. Riedmiller · 1910
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. P. Lillicrap, J. Ba, and M. Norouzi · 1912
Earlier work this paper cites.
Self-improving reactive agents based on reinforcement learning, planning and teaching
L.-J. Lin · 1992
Earlier work this paper cites.
Least-squares temporal difference learning
J. A. Boyan · 1999
Earlier work this paper cites.
Least-squares policy iteration
M. G. Lagoudakis and R. Parr · 2003
Earlier work this paper cites.
Reinforcement learning on explicitly specified time-scales
R. Schoknecht and M. Riedmiller · 2003
Earlier work this paper cites.
Tree-based batch mode reinforcement learning
D. Ernst, P. Geurts, and L. Wehenkel · 2005
Earlier work this paper cites.
Neural fitted q iteration – first experiences with a data efficient neural reinforcement learning method
M. Riedmiller · 2005
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2005
Earlier work this paper cites.
Simple sensor intentions for exploration
T. Hertweck, M. A. Riedmiller, M. Bloesch, J. T. Springenberg, N. Y. Siegel, M. Wulfmeier, R. Hafner, and N. Heess · 2005
Earlier work this paper cites.
Planning to explore via self-supervised world models
R. Sekar, O. Rybkin, K. Daniilidis, P. Abbeel, D. Hafner, and D. Pathak · 2005
Earlier work this paper cites.
Learning to drive in 20 minutes
M. Riedmiller, M. Montemerlo, and H. Dahlkamp · 2007
Earlier work this paper cites.
Learning to dribble on a real robot by success and failure
M. Riedmiller, R. Hafner, S. Lange, and M. Lauer · 2008
Earlier work this paper cites.
Towards general and autonomous learning of core skills: A case study in locomotion
R. Hafner, T. Hertweck, P. Klöppner, M. Bloesch, M. Neunert, M. Wulfmeier, S. Tunyasuvunakool, N. Heess, and M. A. Riedmiller · 2008
Earlier work this paper cites.
Toward the fundamental limits of imitation learning
N. Rajaraman, L. F. Yang, J. Jiao, and K. Ramchandran · 2009
Earlier work this paper cites.
COG: connecting new skills to past experience with offline reinforcement learning
A. Singh, A. Yu, J. Yang, J. Zhang, A. Kumar, and S. Levine · 2010
Earlier work this paper cites.
Learning dexterous manipulation from suboptimal experts
R. Jeong, J. T. Springenberg, J. Kay, D. Zheng, Y. Zhou, A. Galashov, N. Heess, and F. Nori · 2010
Earlier work this paper cites.
Local search for policy iteration in continuous control
J. T. Springenberg, N. Heess, D. J. Mankowitz, J. Merel, A. Byravan, A. Abdolmaleki, J. Kay, J. Degrave, J. Schrittwieser, Y. Tassa, J. Buchli, D. Belov, and M. A. Riedmiller · 2010
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
J. Schmidhuber · 2010
Earlier work this paper cites.
Horde : A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction categories and subject descriptors
R. Sutton, J. Modayil, M. Delp, T. Degris, P. Pilarski, A. White, and D. Precup · 2011
Earlier work this paper cites.
Batch reinforcement learning
S. Lange, T. Gabel, and M. Riedmiller · 2012
Earlier work this paper cites.
Bayesian reinforcement learning
N. Vlassis, M. Ghavamzadeh, S. Mannor, and P. Poupart · 2012
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
N. Heess, G. Wayne, D. Silver, T. Lillicrap, T. Erez, and Y. Tassa · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
S. Gu, T. P. Lillicrap, I. Sutskever, and S. Levine · 2016
Cited alongside, same era.
Bayesian reinforcement learning: A survey
M. Ghavamzadeh, S. Mannor, J. Pineau, and A. Tamar · 2016
Cited alongside, same era.
Why does hierarchy (sometimes) work so well in reinforcement learning?, 2019
O. Nachum, H. Tang, X. Lu, S. Gu, H. Lee, and S. Levine · 2019
Later among the works it cites.
Neural probabilistic motor primitives for humanoid control
J. Merel, L. Hasenclever, A. Galashov, A. Ahuja, V. Pham, G. Wayne, Y. W. Teh, and N. Heess · 2019
Later among the works it cites.
Learning latent plans from play, 2019
C. Lynch, M. Khansari, T. Xiao, V. Kumar, J. Tompson, S. Levine, and P. Sermanet · 2019
Later among the works it cites.
Probabilistic planning with sequential monte carlo methods
A. Piché, V. Thomas, C. Ibrahim, Y. Bengio, and C. Pal · 2019
Later among the works it cites.
Deep exploration via randomized value functions, 2019
I. Osband, B. V. Roy, D. Russo, and Z. Wen · 2019
Later among the works it cites.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. G. Bellemare, S. Srinivasan, G. Ostrovski, T. Schaul, D. Saxton, and R. Munos · 2016
Cited alongside, same era.
M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba · 2017
Cited alongside, same era.
A. Levy, R. P. Jr., and K. Saenko · 2017
Cited alongside, same era.
DDCO: discovery of deep continuous options forrobot learning from demonstrations
S. Krishnan, R. Fox, I. Stoica, and K. Goldberg · 2017
Cited alongside, same era.
Robust imitation of diverse behaviors
Z. Wang, J. Merel, S. E. Reed, G. Wayne, N. de Freitas, and N. Heess · 2017
Cited alongside, same era.
Variational intrinsic control
K. Gregor, D. J. Rezende, and D. Wierstra · 2017
Cited alongside, same era.
Stochastic neural networks for hierarchical reinforcement learning
C. Florensa, Y. Duan, and P. Abbeel · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
Later among the works it cites.
Keep doing what worked: Behavior modelling priors for offline reinforcement learning
N. Siegel, J. T. Springenberg, F. Berkenkamp, A. Abdolmaleki, M. Neunert, T. Lampe, R. Hafner, N. Heess, and M. Riedmiller · 2020
Later among the works it cites.
Rl unplugged: A suite of benchmarks for offline reinforcement learning
C. Gulcehre, Z. Wang, A. Novikov, T. Paine, S. Gómez, K. Zolna, R. Agarwal, J. S. Merel, D. J. Mankowitz, C. Paduraru, G. Dulac-Arnold, J. Li, M. Norouzi, M. Hoffman, N. Heess, and N. de Freitas · 2020
Later among the works it cites.
Critic regularized regression, 2020
Z. Wang, A. Novikov, K. Zolna, J. T. Springenberg, S. Reed, B. Shahriari, N. Siegel, J. Merel, C. Gulcehre, N. Heess, and N. de Freitas · 2020
Later among the works it cites.
Awac: Accelerating online reinforcement learning with offline datasets
A. Nair, A. Gupta, M. Dalal, and S. Levine · 2020
Later among the works it cites.
Compositional transfer in hierarchical reinforcement learning
M. Wulfmeier, A. Abdolmaleki, R. Hafner, J. Tobias Springenberg, M. Neunert, N. Siegel, T. Hertweck, T. Lampe, N. Heess, and M. Riedmiller · 2020
Later among the works it cites.
Scaling data-driven robotics with reward sketching and batch reinforcement learning, 2020
S. Cabi, S. G. Colmenarejo, A. Novikov, K. Konyushkova, S. Reed, R. Jeong, K. Zolna, Y. Aytar, D. Budden, M. Vecerik, O. Sushkov, D. Barker, J. Scholz, M. Denil, N. de Freitas, and Z. Wang · 2020
Later among the works it cites.
Advantage weighted regression: Simple and scalable off-policy reinforcement learning, 2020
X. B. Peng, A. Kumar, G. Zhang, and S. Levine · 2020
Later among the works it cites.
Behavior priors for efficient reinforcement learning, 2020
D. Tirumala, A. Galashov, H. Noh, L. Hasenclever, R. Pascanu, J. Schwarz, G. Desjardins, W. M. Czarnecki, A. Ahuja, Y. W. Teh, and N. Heess · 2020
Later among the works it cites.
Data-efficient hindsight off-policy option learning, 2020
M. Wulfmeier, D. Rao, R. Hafner, T. Lampe, A. Abdolmaleki, T. Hertweck, M. Neunert, D. Tirumala, N. Siegel, N. Heess, and M. Riedmiller · 2020
Later among the works it cites.
Mastering atari, go, chess and shogi by planning with a learned model
J. Schrittwieser, I. Antonoglou, T. Hubert, K. Simonyan, L. Sifre, S. Schmitt, A. Guez, E. Lockhart, D. Hassabis, T. Graepel, T. P. Lillicrap, and D. Silver · 2020
Later among the works it cites.
Importance weighted policy learning and adaption
A. Galashov, J. Sygnowski, G. Desjardins, J. Humplik, L. Hasenclever, R. Jeong, Y. W. Teh, and N. Heess · 2020
Later among the works it cites.
Temporally-extended ϵ \epsilon -greedy exploration, 2020
W. Dabney, G. Ostrovski, and A. Barreto · 2020
Later among the works it cites.
Towards general and autonomous learning of core skills: A case study in locomotion, 2020
R. Hafner, T. Hertweck, P. Klöppner, M. Bloesch, M. Neunert, M. Wulfmeier, S. Tunyasuvunakool, N. Heess, and M. Riedmiller · 2020
Later among the works it cites.
On multi-objective policy optimization as a tool for reinforcement learning
A. Abdolmaleki, S. H. Huang, G. Vezzani, B. Shahriari, J. T. Springenberg, S. Mishra, D. TB, A. Byravan, K. Bousmalis, A. Gyorgy, et al · 2021
Closest in time.
Mt-opt: Continuous multi-task robotic reinforcement learning at scale, 2021
D. Kalashnikov, J. Varley, Y. Chebotar, B. Swanson, R. Jonschkowski, C. Finn, S. Levine, and K. Hausman · 2021
Closest in time.
{OPAL}: Offline primitive discovery for accelerating offline reinforcement learning
A. Ajay, A. Kumar, P. Agrawal, S. Levine, and O. Nachum · 2021
Closest in time.
Actionable models: Unsupervised offline reinforcement learning of robotic skills, 2021
Y. Chebotar, K. Hausman, Y. Lu, T. Xiao, D. Kalashnikov, J. Varley, A. Irpan, B. Eysenbach, R. Julian, C. Finn, and S. Levine · 2021
Closest in time.
Reinforcement learning, bit by bit
X. Lu, B. V. Roy, V. Dwaracherla, M. Ibrahimi, I. Osband, and Z. Wen · 2021
Closest in time.