Fetching the paper…
Reading the bibliography…
Intrinsically motivated reinforcement learning aims to address the exploration challenge for sparse-reward tasks.
A possibility for implementing curiosity and boredom in model-building neural controllers
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Optimal payoff functions for members of collectives
David H Wolpert and Kagan Tumer · 2002
Earlier work this paper cites.
Coordination in multiagent reinforcement learning: A bayesian approach
Georgios Chalkiadakis and Craig Boutilier · 2003
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Nuttapong Chentanez, Andrew G Barto, and Satinder P Singh · 2005
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Pierre-Yves Oudeyer, Frdric Kaplan, and Verena V Hafner · 2007
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
Pierre-Yves Oudeyer and Frederic Kaplan · 2009
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
Thomas Jaksch, Ronald Ortner, and Peter Auer · 2010
Earlier work this paper cites.
Coordinated multi-agent reinforcement learning in networked distributed pomdps
Chongjie Zhang and Victor Lesser · 2011
Earlier work this paper cites.
An overview of recent progress in the study of distributed multi-agent coordination
Yongcan Cao, Wenwu Yu, Wei Ren, and Guanrong Chen · 2012
Earlier work this paper cites.
A tutorial on variational bayesian inference
Charles W Fox and Stephen J Roberts · 2012
Earlier work this paper cites.
Game theory and multi-agent reinforcement learning
Ann Nowé, Peter Vrancx, and Yann-Michaël De Hauwere · 2012
Earlier work this paper cites.
Trading value and information in mdps
Jonathan Rubin, Ohad Shamir, and Naftali Tishby · 2012
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
Susanne Still and Doina Precup · 2012
Earlier work this paper cites.
Intrinsic motivation and reinforcement learning
Andrew G Barto · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Changing the environment based on empowerment as intrinsic motivation
Christoph Salge, Cornelius Glackin, and Daniel Polani · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo Jimenez Rezende · 2015
Cited alongside, same era.
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Cited alongside, same era.
Learning to communicate with deep multi-agent reinforcement learning
Jakob Foerster, Ioannis Alexandros Assael, Nando de Freitas, and Shimon Whiteson · 2016
Cited alongside, same era.
Information maximizing exploration with a latent dynamics model
Trevor Barron, Oliver Obst, and Heni Ben Amor · 2018
Later among the works it cites.
Coordinated exploration in concurrent reinforcement learning
Maria Dimakopoulou and Benjamin Van Roy · 2018
Later among the works it cites.
Scalable coordinated exploration in concurrent reinforcement learning
Maria Dimakopoulou, Ian Osband, and Benjamin Van Roy · 2018
Later among the works it cites.
Model-based value estimation for efficient model-free reinforcement learning
Vladimir Feinberg, Alvin Wan, Ion Stoica, Michael I Jordan, Joseph E Gonzalez, and Sergey Levine · 2018
Later among the works it cites.
Counterfactual multi-agent policy gradients
Jakob N Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Cited alongside, same era.
Learning to navigate in complex environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andrew J Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, et al · 2016
Cited alongside, same era.
On lower bounds for regret in reinforcement learning
Ian Osband and Benjamin Van Roy · 2016
Cited alongside, same era.
Training agent for first-person shooter game with actor-critic curriculum learning
Yuxin Wu and Yuandong Tian · 2016
Cited alongside, same era.
Optimistic posterior sampling for reinforcement learning: worst-case regret bounds
Shipra Agrawal and Randy Jia · 2017
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Cited alongside, same era.
Openai baselines
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu, and Peter Zhokhov · 2017
Cited alongside, same era.
Meta-reinforcement learning of structured exploration strategies
Abhishek Gupta, Russell Mendonca, YuXuan Liu, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Inequity aversion improves cooperation in intertemporal social dilemmas
Edward Hughes, Joel Z Leibo, Matthew Phillips, Karl Tuyls, Edgar Dueñez-Guzman, Antonio García Castañeda, Iain Dunning, Tina Zhu, Kevin McKee, Raphael Koster, et al · 2018
Later among the works it cites.
Intrinsic social motivation via causal influence in multi-agent rl
Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Caglar Gulcehre, Pedro A Ortega, DJ Strouse, Joel Z Leibo, and Nando de Freitas · 2018
Later among the works it cites.
Is q-learning provably efficient?
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck, and Michael I Jordan · 2018
Later among the works it cites.
Credit assignment for collective multiagent rl with global rewards
Duc Thien Nguyen, Akshat Kumar, and Hoong Chuin Lau · 2018
Later among the works it cites.
Openai five
OpenAI · 2018
Later among the works it cites.
Zero-shot visual imitation
Deepak Pathak, Parsa Mahmoudieh, Guanghao Luo, Pulkit Agrawal, Dian Chen, Yide Shentu, Evan Shelhamer, Jitendra Malik, Alexei A Efros, and Trevor Darrell · 2018
Later among the works it cites.
Prosocial learning agents solve generalized stag hunts better than selfish ones
Alexander Peysakhovich and Adam Lerer · 2018
Later among the works it cites.
Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning
Tabish Rashid, Mikayel Samvelyan, Christian Schroeder Witt, Gregory Farquhar, Jakob Foerster, and Shimon Whiteson · 2018
Later among the works it cites.
Learning by playing-solving sparse reward tasks from scratch
Martin Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom Van de Wiele, Volodymyr Mnih, Nicolas Heess, and Jost Tobias Springenberg · 2018
Later among the works it cites.
Learning to share and hide intentions using information regularization
DJ Strouse, Max Kleiman-Weiner, Josh Tenenbaum, Matt Botvinick, and David J Schwab · 2018
Later among the works it cites.
Building generalizable agents with a realistic and rich 3d environment
Yi Wu, Yuxin Wu, Georgia Gkioxari, and Yuandong Tian · 2018
Later among the works it cites.
Emi: Exploration with mutual information
Yeonwoo Jeong Sergey Levine Hyun Oh Song Hyoungseok Kim, Jaekyeom Kim · 2019
Closest in time.
Coordinated exploration via intrinsic rewards for multi-agent reinforcement learning
Shariq Iqbal and Fei Sha · 2019
Closest in time.