Fetching the paper…
Reading the bibliography…
Many real world tasks exhibit rich structure that is repeated across different parts of the state space or in time.
Infobot: Transfer and exploration via the information bottleneck
Anirudh Goyal, Riashat Islam, Daniel Strouse, Zafarali Ahmed, Matthew Botvinick, Hugo Larochelle, Yoshua Bengio, and Sergey Levine · 1901
Earlier work this paper cites.
Rational choice and the structure of the environment
H.A. Simon · 1956
Earlier work this paper cites.
Maximum likelihood from incomplete data via the EM algorithm
A. P. Dempster, N. M. Laird, and D. B. Rubin · 1977
Earlier work this paper cites.
Function optimization using connectionist reinforcement learning algorithms
R. J. Williams and J. Peng · 1991
Earlier work this paper cites.
Learning in graphical models
Radford M. Neal and Geoffrey E. Hinton · 1999
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen J. Wright · 2006
Earlier work this paper cites.
Linearly-solvable markov decision problems
Emanuel Todorov · 2007
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Marc Toussaint · 2009
Earlier work this paper cites.
Relative entropy policy search
Jan Peters, Katharina Mülling, and Yasemin Altün · 2010
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Brian D. Ziebart · 2010
Earlier work this paper cites.
Information, utility & bounded rationality
Pedro A. Ortega and Daniel A. Braun · 2011
Earlier work this paper cites.
The information theory of decision and action
Naftali Tishby and Daniel Polani · 2011
Earlier work this paper cites.
Optimal control as a graphical model inference problem
Hilbert J. Kappen, Vicenç Gómez, and Manfred Opper · 2012
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2012
Earlier work this paper cites.
Trading Value and Information in MDPs , pp. 57–74
Jonathan Rubin, Ohad Shamir, and Naftali Tishby · 2012
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
Susanne Still and Doina Precup · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Variational policy search via trajectory optimization
Sergey Levine and Vladlen Koltun · 2013
Cited alongside, same era.
Thermodynamics as a theory of decision-making with information-processing costs
Pedro A. Ortega and Daniel A. Braun · 2013
Cited alongside, same era.
Stochastic backpropagation and approximate inference in deep generative models
Danilo Jimenez Rezende, Shakir Mohamed, and Daan Wierstra · 2014
Cited alongside, same era.
G-learning: Taming the noise in reinforcement learning via soft updates
Roy Fox, Ari Pakman, and Naftali Tishby · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Alexander A. Alemi, Ben Poole, Ian Fischer, Joshua V. Dillon, Rif A. Saurous, and Kevin Murphy · 2017
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Emergence of locomotion behaviours in rich environments
Nicolas Heess, Srinivasan Sriram, Jay Lemmon, Josh Merel, Greg Wayne, Yuval Tassa, Tom Erez, Ziyu Wang, Ali Eslami, Martin Riedmiller, et al · 2017
Later among the works it cites.
Population based training of neural networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M. Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, Chrisantha Fernando, and Koray Kavukcuoglu · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Tim Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Deep variational information bottleneck
Alexander A. Alemi, Ian Fischer, Joshua V. Dillon, and Kevin Murphy · 2016
Cited alongside, same era.
Charles Beattie, Joel Z. Leibo, Denis Teplyashin, Tom Ward, Marcus Wainwright, Heinrich Küttler, Andrew Lefrancq, Simon Green, Víctor Valdés, Amir Sadik, Julian Schrittwieser, Keith Anderson, Sarah York, Max Cant, Adam Cain, Adrian Bolton, Stephen Gaffney, Helen King, Demis Hassabis, Shane Legg, and Stig Petersen · 2016
Cited alongside, same era.
Path integral guided policy search
Yevgen Chebotar, Mrinal Kalakrishnan, Ali Yahya, Adrian Li, Stefan Schaal, and Sergey Levine · 2016
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Later among the works it cites.
Distral: Robust multitask reinforcement learning
Yee Whye Teh, Victor Bapst, Wojciech Marian Czarnecki, John Quan, James Kirkpatrick, Raia Hadsell, Nicolas Heess, and Razvan Pascanu · 2017
Later among the works it cites.
A unified bellman equation for causal information and value in markov decision processes
Stas Tiomkin and Naftali Tishby · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Rémi Munos, Nicolas Heess, and Martin A. Riedmiller · 2018
Later among the works it cites.
Mix & match agent curricula for reinforcement learning
Wojciech Marian Czarnecki, Siddhant M. Jayakumar, Max Jaderberg, Leonard Hasenclever, Yee Whye Teh, Nicolas Heess, Simon Osindero, and Razvan Pascanu · 2018
Later among the works it cites.
Scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Later among the works it cites.
Divide-and-conquer reinforcement learning
Dibya Ghosh, Avi Singh, Aravind Rajeswaran, Vikash Kumar, and Sergey Levine · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Learning an embedding space for transferable robot skills
Karol Hausman, Jost Tobias Springenberg, Ziyu Wang, Nicolas Heess, and Martin Riedmiller · 2018
Later among the works it cites.
Mental labour
Wouter Kool and Matthew Botvinick · 2018
Later among the works it cites.
Kickstarting deep reinforcement learning
Simon Schmitt, Jonathan J. Hudson, Augustin Zidek, Simon Osindero, Carl Doersch, Wojciech M. Czarnecki, Joel Z. Leibo, Heinrich Kuttler, Andrew Zisserman, Karen Simonyan, and S. M. Ali Eslami · 2018
Later among the works it cites.
Deep mutual learning
Ying Zhang, Tao Xiang, Timothy Hospedales, and Huchuan Lu · 2018
Later among the works it cites.