Fetching the paper…
Reading the bibliography…
Humans achieve efficient learning by relying on prior knowledge about the structure of naturally occurring tasks.
Function optimization using connectionist reinforcement learning algorithms
Ronald J Williams and Jing Peng · 1991
Earlier work this paper cites.
Learning from demonstration
Stefan Schaal · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Malcolm J. A. Strens · 2000
Earlier work this paper cites.
Optimal Learning: Computational procedures for Bayes-adaptive Markov decision processes
Michael O’Gordon Duff · 2002
Earlier work this paper cites.
A bayesian approach to unsupervised one-shot learning of object categories
Li Fei-Fei et al · 2003
Earlier work this paper cites.
Using rewards for belief state updates in partially observable markov decision processes
Masoumeh T Izadi and Doina Precup · 2005
Earlier work this paper cites.
An analytic solution to discrete bayesian reinforcement learning
Pascal Poupart, Nikos Vlassis, Jesse Hoey, and Kevin Regan · 2006
Earlier work this paper cites.
Multi-task reinforcement learning: a hierarchical bayesian approach
Aaron Wilson, Alan Fern, Soumya Ray, and Prasad Tadepalli · 2007
Earlier work this paper cites.
Bayes-optimal reinforcement learning for discrete uncertainty domains
Emma Brunskill · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Further optimal regret bounds for thompson sampling
Shipra Agrawal and Navin Goyal · 2013
Earlier work this paper cites.
Amortized inference in probabilistic reasoning
Samuel Gershman and Noah Goodman · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Timothy Lillicrap, Tom Erez, and Yuval Tassa · 2015
Earlier work this paper cites.
Rl 2 : Fast reinforcement learning via slow reinforcement learning, 2016
Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Earlier work this paper cites.
Inference networks for sequential monte carlo in graphical models
Brooks Paige and Frank Wood · 2016
Cited alongside, same era.
Relevant sparse codes with variational information bottleneck
Matthew Chalk, Olivier Marre, and Gasper Tkacik · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Remi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Cited alongside, same era.
Conditional neural processes
Marta Garnelo, Dan Rosenbaum, Christopher Maddison, Tiago Ramalho, David Saxton, Murray Shanahan, Yee Whye Teh, Danilo Rezende, and SM Ali Eslami · 2018
Later among the works it cites.
Marta Garnelo, Jonathan Schwarz, Dan Rosenbaum, Fabio Viola, Danilo J Rezende, SM Eslami, and Yee Whye Teh · 2018
Later among the works it cites.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Later among the works it cites.
Deep variational reinforcement learning for pomdps
Maximilian Igl, Luisa Zintgraf, Tuan Anh Le, Frank Wood, and Shimon Whiteson · 2018
Later among the works it cites.
Learning by playing solving sparse reward tasks from scratch
Martin Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom van de Wiele, Vlad Mnih, Nicolas Heess, and Jost Tobias Springenberg · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Emergence of locomotion behaviours in rich environments
Nicolas Heess, Srinivasan Sriram, Jay Lemmon, Josh Merel, Greg Wayne, Yuval Tassa, Tom Erez, Ziyu Wang, Ali Eslami, Martin Riedmiller, et al · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Why is posterior sampling better than optimism for reinforcement learning?
Ian Osband and Benjamin Van Roy · 2017
Cited alongside, same era.
Grounded language learning in a simulated 3d world
Karl Moritz Hermann, Felix Hill, Simon Green, Fumin Wang, Ryan Faulkner, Hubert Soyer, David Szepesvari, Wojciech Marian Czarnecki, Max Jaderberg, Denis Teplyashin, et al · 2017
Cited alongside, same era.
Reverse curriculum generation for reinforcement learning
Carlos Florensa, David Held, Markus Wulfmeier, Michael Zhang, and Pieter Abbeel · 2017
Cited alongside, same era.
Deep variational information bottleneck
Alexander A Alemi, Ian Fischer, Joshua V Dillon, and Kevin Murphy · 2017
Cited alongside, same era.
Later among the works it cites.
Meta-reinforcement learning of structured exploration strategies
Abhishek Gupta, Russell Mendonca, YuXuan Liu, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Recasting gradient-based meta-learning as hierarchical bayes
Erin Grant, Chelsea Finn, Sergey Levine, Trevor Darrell, and Thomas Griffiths · 2018
Later among the works it cites.
Bayesian model-agnostic meta-learning
Jaesik Yoon, Taesup Kim, Ousmane Dia, Sungwoong Kim, Yoshua Bengio, and Sungjin Ahn · 2018
Later among the works it cites.
Neural belief states for partially observed domains
Pol Moreno, Jan Humplik, George Papamakarios, Lars Buesing, Nicolas Heess, and Théophane Weber · 2018
Later among the works it cites.
Unsupervised predictive memory in a goal-directed agent
Greg Wayne, Chia-Chun Hung, David Amos, Mehdi Mirza, Arun Ahuja, Agnieszka Grabska-Barwinska, Jack Rae, Piotr Mirowski, Joel Z Leibo, Adam Santoro, et al · 2018
Later among the works it cites.
Learning longer-term dependencies in rnns with auxiliary losses
Trieu H Trinh, Andrew M Dai, Minh-Thang Luong, and Quoc V Le · 2018
Later among the works it cites.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Deirdre Quillen, Chelsea Finn, and Sergey Levine · 2019
Closest in time.
Emergent coordination through competition
Siqi Liu, Guy Lever, Nicholas Heess, Josh Merel, Saran Tunyasuvunakool, and Thore Graepel · 2019
Closest in time.
Russell Mendonca, Abhishek Gupta, Rosen Kralev, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2019
Closest in time.
Neural predictive belief representations, 2019
Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Bernardo Avila Pires, and Rémi Munos · 2019
Closest in time.
Learning dynamics model in reinforcement learning by incorporating the long term future
Nan Rosemary Ke, Amanpreet Singh, Ahmed Touati, Anirudh Goyal, Yoshua Bengio, Devi Parikh, and Dhruv Batra · 2019
Closest in time.