Fetching the paper…
Reading the bibliography…
Learning policies on data synthesized by models can in principle quench the thirst of reinforcement learning algorithms for large amounts of real experience, which is often costly to acquire.
The relative importance of heredity and environment in determining the piebald pattern of guinea-pigs
Sewall Wright · 1920
Earlier work this paper cites.
Counterfactual probabilities: Computational methods, bounds and applications
Alexander Balke and Judea Pearl · 1994
Earlier work this paper cites.
Counterfactual thinking
Neal J Roese · 1997
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Varieties of causal intervention
Kevin B Korb, Lucas R Hope, Ann E Nicholson, and Karl Axnick · 2004
Earlier work this paper cites.
Using inaccurate models in reinforcement learning
Pieter Abbeel, Morgan Quigley, and Andrew Y Ng · 2006
Earlier work this paper cites.
Causality: Models, Reasoning and Inference
Judea Pearl · 2009
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Marc Toussaint · 2009
Earlier work this paper cites.
A survey of Monte Carlo tree search methods
Cameron B Browne, Edward Powley, Daniel Whitehouse, Simon M Lucas, Peter I Cowling, Philipp Rohlfshagen, Stephen Tavener, Diego Perez, Spyridon Samothrakis, and Simon Colton · 2012
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X Charles, D Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Earlier work this paper cites.
Auto-encoding variational Bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Guided policy search
Sergey Levine and Vladlen Koltun · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Learning neural network policies with guided policy search under unknown dynamics
Sergey Levine and Pieter Abbeel · 2014
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Cited alongside, same era.
Model Regularization for Stable Sample Rollouts
Erik Talvitie · 2014
Cited alongside, same era.
Convolutional LSTM network: A machine learning approach for precipitation nowcasting
Shi Xingjian, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo · 2015
Later among the works it cites.
Onur Atan, William R Zame, Qiaojun Feng, and Mihaela van der Schaar · 2016
Later among the works it cites.
Towards conceptual compression
Karol Gregor, Frederic Besse, Danilo Jimenez Rezende, Ivo Danihelka, and Daan Wierstra · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare · 2016
Later among the works it cites.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Pieter Abbeel, and Wojciech Zaremba · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Jimenez Rezende, and Daan Wierstra · 2015
Cited alongside, same era.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Tim Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2015
Cited alongside, same era.
The dependence of effective planning horizon on model accuracy
Nan Jiang, Alex Kulesza, Satinder Singh, and Richard Lewis · 2015
Cited alongside, same era.
Counterfactual estimation and optimization of click metrics in search engines: A case study
Lihong Li, Shunbao Chen, Jim Kleban, and Ankur Gupta · 2015
Cited alongside, same era.
Counterfactual risk minimization: Learning from logged bandit feedback
Adith Swaminathan and Thorsten Joachims · 2015
Cited alongside, same era.
Thomas Nedelec, Nicolas Le Roux, and Vianney Perchet · 2017
Later among the works it cites.
Elements of Causal Inference - Foundations and Learning Algorithms
J. Peters, D. Janzing, and B. Schölkopf · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Sébastien Racanière, Théophane Weber, David Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adrià Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al · 2017
Later among the works it cites.
IMPALA: Scalable distributed Deep-RL with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Closest in time.
The Book of Why: The New Science of Cause and Effect
Judea Pearl and Dana Mackenzie · 2018
Closest in time.