Fetching the paper…
Reading the bibliography…
Reasoning about the future -- understanding how decisions in the present time affect outcomes in the future -- is one of the central challenges for reinforcement learning (RL), especially in highly-stochastic or partially observable environments.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Richard S. Sutton · 1991
Earlier work this paper cites.
Learning without state-estimation in partially observable markovian decision processes
Satinder P. Singh · 1994
Earlier work this paper cites.
Analysis of temporal-difference learning with function approximation
John N. Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
A reinforcement learning algorithm in partially observable environments using short-term memory
Nobuo Suematsu and Akira Hayashi · 1999
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
The information bottleneck method
Naftali Tishby, Fernando C. Pereira, and William Bialek · 1999
Earlier work this paper cites.
The optimal reward baseline for gradient-based reinforcement learning
Lex Weaver and Nigel Tao · 2001
Earlier work this paper cites.
Manifold embeddings for model-based reinforcement learning under partial observability
Keith Bush and Joelle Pineau · 2009
Earlier work this paper cites.
Mental models and human reasoning
Philip N. Johnson-Laird · 2010
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Learning stochastic recurrent networks, 2015
Justin Bayer and Christian Osendorfer · 2015
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer · 2015
Earlier work this paper cites.
Professor forcing: A new algorithm for training recurrent networks
Alex M Lamb, Anirudh Goyal, Ying Zhang, Saizheng Zhang, Aaron C Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Earlier work this paper cites.
Reverse curriculum generation for reinforcement learning
Carlos Florensa, David Held, Markus Wulfmeier, Michael Zhang, and Pieter Abbeel · 2017
Earlier work this paper cites.
Z-forcing: Training stochastic recurrent networks
Anirudh Goyal, Alessandro Sordoni, Marc-Alexandre Côté, Nan Rosemary Ke, and Yoshua Bengio · 2017
Earlier work this paper cites.
Deep variational bayes filters: Unsupervised learning of state space models from raw data
M. Karl, Maximilian Sölch, J. Bayer, and P. V. D. Smagt · 2017
Earlier work this paper cites.
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Cited alongside, same era.
Count-based exploration with neural density models
Georg Ostrovski, Marc G. Bellemare, Aäron van den Oord, and Rémi Munos · 2017
Cited alongside, same era.
Learning model-based planning from scratch, 2017
Razvan Pascanu, Yujia Li, Oriol Vinyals, Nicolas Heess, Lars Buesing, Sebastien Racanière, David Reichert, Théophane Weber, Daan Wierstra, and Peter Battaglia · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Imagination-augmented agents for deep reinforcement learning
Sébastien Racanière, Theophane Weber, David Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adrià Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, Razvan Pascanu, Peter Battaglia, Demis Hassabis, David Silver, and Daan Wierstra · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Modeling the long term future in model-based reinforcement learning
Nan Rosemary Ke, Amanpreet Singh, Ahmed Touati, Anirudh Goyal, Yoshua Bengio, Devi Parikh, and Dhruv Batra · 2019
Later among the works it cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Later among the works it cites.
Algaedice: Policy gradient from arbitrary experience, 2019
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 2019
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning, 2020
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Value-driven hindsight modelling
Arthur Guez, Fabio Viola, Theophane Weber, Lars Buesing, Steven Kapturowski, Doina Precup, David Silver, and Nicolas Heess · 2020
Later among the works it cites.
Behaviour suite for reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Sample-efficient reinforcement learning with stochastic ensemble value expansion
Jacob Buckman, Danijar Hafner, George Tucker, Eugene Brevdo, and Honglak Lee · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke van Hoof, and David Meger · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Deep Dyna-Q: Integrating planning for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, and Kam-Fai Wong · 2018
Cited alongside, same era.
On the information bottleneck theory of deep learning
Andrew Michael Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan Daniel Tracey, and David Daniel Cox · 2018
Cited alongside, same era.
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepesvári, Satinder Singh, Benjamin Van Roy, Richard Sutton, David Silver, and Hado van Hasselt · 2020
Later among the works it cites.
Reinforcement learning upside down: Don’t predict rewards – just map them to actions, 2020
Juergen Schmidhuber · 2020
Later among the works it cites.
Benchmarking model-based reinforcement learning, 2020
Tingwu Wang, Xuchan Bao, Ignasi Clavera, Jerrick Hoang, Yeming Wen, Eric Langlois, Shunshi Zhang, Guodong Zhang, Pieter Abbeel, and Jimmy Ba · 2020
Later among the works it cites.
Behavior regularized offline reinforcement learning, 2020
Yifan Wu, George Tucker, and Ofir Nachum · 2020
Later among the works it cites.
Model based reinforcement learning for atari
Łukasz Kaiser, Mohammad Babaeizadeh, Piotr Miłos, Błażej Osiński, Roy H Campbell, Konrad Czechowski, Dumitru Erhan, Chelsea Finn, Piotr Kozakowski, Sergey Levine, Afroz Mohiuddin, Ryan Sepassi, George Tucker, and Henryk Michalewski · 2020
Later among the works it cites.
Recurrent independent mechanisms
Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, Sergey Levine, Yoshua Bengio, and Bernhard Schölkopf · 2021
Closest in time.
Reinforcement learning as one big sequence modeling problem
Michael Janner, Qiyang Li, and Sergey Levine · 2021
Closest in time.
Counterfactual credit assignment in model-free reinforcement learning
Thomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor, Alaa Saade, Anna Harutyunyan, Will Dabney, Thomas S Stepleton, Nicolas Heess, Arthur Guez, Eric Moulines, Marcus Hutter, Lars Buesing, and Remi Munos · 2021
Closest in time.
Posterior value functions: Hindsight baselines for policy gradient methods
Chris Nota, Philip Thomas, and Bruno C. Da Silva · 2021
Closest in time.
Data-efficient hindsight off-policy option learning, 2021
Markus Wulfmeier, Dushyant Rao, Roland Hafner, Thomas Lampe, Abbas Abdolmaleki, Tim Hertweck, Michael Neunert, Dhruva Tirumala, Noah Siegel, Nicolas Heess, and Martin Riedmiller · 2021
Closest in time.
Representation matters: Offline pretraining for sequential decision making, 2021
Mengjiao Yang and Ofir Nachum · 2021
Closest in time.
{BRAC}+: Going deeper with behavior regularized offline reinforcement learning, 2021
Chi Zhang, Sanmukh Rao Kuppannagari, and Viktor Prasanna · 2021
Closest in time.