Fetching the paper…
Reading the bibliography…
Intrinsic rewards can improve exploration in reinforcement learning, but the exploration process may suffer from instability caused by non-stationary reward shaping and strong dependency on hyperparameters.
Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning . 1928–1937
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016 · 1937
Earlier work this paper cites.
Dynamic programming and Markov processes
Ronald A Howard. 1964 · 1964
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins. 1989 · 1989
Earlier work this paper cites.
Curious model-building control systems
Jürgen Schmidhuber. 1991 · 1991
Earlier work this paper cites.
Reinforcement learning for robots using neural networks
Long-Ji Lin. 1992 · 1992
Earlier work this paper cites.
Intrinsically motivated reinforcement learning. In Advances on Neural Information Processing Systems . 1281–1288
Nuttapong Chentanez, Andrew G Barto, and Satinder P Singh. 2005 · 2005
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Pierre-Yves Oudeyer, Frdric Kaplan, and Verena V Hafner. 2007 · 2007
Earlier work this paper cites.
Near-optimal hashing algorithms for approximate nearest neighbor in high dimensions
Alexandr Andoni and Piotr Indyk. 2008 · 2008
Earlier work this paper cites.
How can we define intrinsic motivation. In Conference on Epigenetic Robotics , Vol. 5. 29–31
Pierre-Yves Oudeyer, Frederic Kaplan, et al · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for Markov decision processes
Alexander L Strehl and Michael L Littman. 2008 · 2008
Earlier work this paper cites.
What is intrinsic motivation? A typology of computational approaches
Pierre-Yves Oudeyer and Frederic Kaplan. 2009 · 2009
Earlier work this paper cites.
Off-policy actor-critic. In International Conference on Machine Learning
Thomas Degris, Martha White, and Richard S Sutton. 2012 · 2012
Earlier work this paper cites.
Intrinsic motivation and reinforcement learning
Andrew G Barto. 2013 · 2013
Earlier work this paper cites.
Deterministic policy gradient algorithms. In International Conference on Machine Learning . PMLR, 387–395
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller. 2014 · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis. 2015 · 2015
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation. In Advances on Neural Information Processing Systems . 1471–1479
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos. 2016 · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning. In International Conference on Learning Representations
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2016 · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning. In Advances on Neural Information Processing Systems , Vol. 29
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc G Bellemare. 2016 · 2016
Cited alongside, same era.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard. 2019 · 2019
Later among the works it cites.
Benchmarking bonus-based exploration methods on the arcade learning environment
Adrien Ali Taïga, William Fedus, Marlos C Machado, Aaron Courville, and Marc G Bellemare. 2019 · 2019
Later among the works it cites.
Behavior regularized offline reinforcement learning
Yifan Wu, George Tucker, and Ofir Nachum. 2019 · 2019
Later among the works it cites.
Shared Experience Actor-Critic for Multi-Agent Reinforcement Learning. In Advances in Neural Information Processing Systems , Vol. 33. 10707–10717
Filippos Christianos, Lukas Schäfer, and Stefano V. Albrecht. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Georg Ostrovski, Marc G Bellemare, Aäron van den Oord, and Rémi Munos. 2017 · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction. In International Conference on Machine Learning
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Cited alongside, same era.
# Exploration: A study of count-based exploration for deep reinforcement learning. In Advances on Neural Information Processing Systems . 2753–2762
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, OpenAI Xi Chen, Yan Duan, John Schulman, Filip DeTurck, and Pieter Abbeel. 2017 · 2017
Cited alongside, same era.
Large-scale study of curiosity-driven learning
Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, and Alexei A Efros. 2018 · 2018
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures. In International Conference on Machine Learning . PMLR, 1407–1416
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Vlad Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods. In International Conference on Machine Learning . PMLR, 1587–1596
Scott Fujimoto, Herke Van Hoof, and David Meger. 2018 · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning . PMLR, 1861–1870
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Cited alongside, same era.
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. 2020 · 2020
Later among the works it cites.
Behaviour suite for reinforcement learning. In International Conference on Learning Representations
Ian Osband, Yotam Doron, Matteo Hessel, John Aslanides, Eren Sezener, Andre Saraiva, Katrina McKinney, Tor Lattimore, Csaba Szepesvari, Satinder Singh, et al · 2020
Later among the works it cites.
RIDE: Rewarding impact-driven exploration for procedurally-generated environments. In International Conference on Learning Representations
Roberta Raileanu and Tim Rocktäschel. 2020 · 2020
Later among the works it cites.
Optimistic Exploration even with a Pessimistic Initialisation. In International Conference on Learning Representations
Tabish Rashid, Bei Peng, Wendelin Böhmer, and Shimon Whiteson. 2020 · 2020
Later among the works it cites.
Deep reinforcement learning at the edge of the statistical precipice. In Advances on Neural Information Processing Systems
Rishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron Courville, and Marc G Bellemare. 2021 · 2021
Closest in time.
Adversarially Guided Actor-Critic. In International Conference on Learning Representations
Yannis Flet-Berliac, Johan Ferret, Olivier Pietquin, Philippe Preux, and Matthieu Geist. 2021 · 2021
Closest in time.
Benchmarking Multi-Agent Deep Reinforcement Learning Algorithms in Cooperative Tasks. In Advances in Neural Information Processing Systems, Track on Datasets and Benchmarks
Georgios Papoudakis, Filippos Christianos, Lukas Schäfer, and Stefano V. Albrecht. 2021 · 2021
Closest in time.
Decoupled Exploration and Exploitation Policies for Sample-Efficient Reinforcement Learning
William F Whitney, Michael Bloesch, Jost Tobias Springenberg, Abbas Abdolmaleki, and Martin Riedmiller. 2021 · 2021
Closest in time.
Robust On-Policy Data Collection for Data-Efficient Policy Evaluation. In Offline Reinforcement Learning Workshop at Neural Information Processing Systems Conference
Rujie Zhong, Josiah P. Hanna, Lukas Schäfer, and Stefano V. Albrecht. 2021 · 2021
Closest in time.
Deep Reinforcement Learning with Double Q-Learning. In AAAI Conference on Artificial Intelligence . AAAI Press, 2094–2100
Hado van Hasselt, Arthur Guez, and David Silver. 2016 · 2094
Closest in time.