Fetching the paper…
Reading the bibliography…
The exploration problem is one of the main challenges in deep reinforcement learning (RL).
Is deep reinforcement learning really superhuman on atari?
Marin Toromanoff, Émilie Wirbel, and Fabien Moutarde · 1908
Earlier work this paper cites.
Dynamic programming and markov processes
Ronald A Howard · 1960
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter; Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Reinforcement learning - an introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Jiajun Fan, He Ba, Xian Guo, and Jianye Hao · 2011
Earlier work this paper cites.
On upper-confidence bound policies for switching bandit problems
Aurélien Garivier and Eric Moulines · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G. Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael I. Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc G. Bellemare · 2016
Earlier work this paper cites.
Dueling network architectures for deep reinforcement learning
Ziyu Wang, Tom Schaul, Matteo Hessel, Hado van Hasselt, Marc Lanctot, and Nando de Freitas · 2016
Earlier work this paper cites.
Population based training of neural networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M. Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, Chrisantha Fernando, and Koray Kavukcuoglu · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
IMPALA: scalable distributed deep-rl with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Rémi Munos, Karen Simonyan, Volodymyr Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, and David Silver · 2018
Cited alongside, same era.
Learning to trade with deep actor critic methods
Off-policy actor-critic with shared experience replay
Simon Schmitt, Matteo Hessel, and Karen Simonyan · 2020
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Later among the works it cites.
A review for deep reinforcement learning in atari: Benchmarks, challenges, and solutions
Jiajun Fan · 2021
Later among the works it cites.
Mastering atari with discrete world models
Danijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, and Jimmy Ba · 2021
Later among the works it cites.
Going beyond linear transformers with recurrent fast weight programmers
Kazuki Irie, Imanol Schlag, Róbert Csordás, and Jürgen Schmidhuber · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jinke Li, Ruonan Rao, and Jun Shi · 2018
Cited alongside, same era.
Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents
Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew J. Hausknecht, and Michael Bowling · 2018
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2019
Cited alongside, same era.
Recurrent experience replay in distributed reinforcement learning
Steven Kapturowski, Georg Ostrovski, John Quan, Rémi Munos, and Will Dabney · 2019
Cited alongside, same era.
Cascade multi-head attention networks for action recognition
Jiaze Wang, Xiaojiang Peng, and Yu Qiao · 2019
Cited alongside, same era.
Agent57: Outperforming the atari human benchmark
Adrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, and Charles Blundell · 2020
Cited alongside, same era.
Never give up: Learning directed exploration strategies
Adrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo, Bilal Piot, Steven Kapturowski, Olivier Tieleman, Martín Arjovsky, Alexander Pritzel, Andrew Bolt, and Charles Blundell · 2020
Cited alongside, same era.
Effective diversity in population based reinforcement learning
Jack Parker-Holder, Aldo Pacchiano, Krzysztof Marcin Choromanski, and Stephen J. Roberts · 2020
Cited alongside, same era.
Category-level 6d object pose estimation via cascaded relation and recurrent reconstruction networks
Jiaze Wang, Kai Chen, and Qi Dou · 2021
Later among the works it cites.
Generalized data distribution iteration
Jiajun Fan and Changnan Xiao · 2022
Later among the works it cites.
Steven Kapturowski, Víctor Campos, Ray Jiang, Nemanja Rakićević, Hado van Hasselt, Charles Blundell, and Adrià Puigdomènech Badia · 2022
Later among the works it cites.
ESCM2: entire space counterfactual multi-task model for post-click conversion rate estimation
Hao Wang, Tai-Wei Chang, Tianqiao Liu, Jianmin Huang, Zhichao Chen, Chao Yu, Ruopeng Li, and Wei Chu · 2022
Later among the works it cites.
Disentangled contrastive learning for social recommendation
Jiahao Wu, Wenqi Fan, Jingfan Chen, Shengcai Liu, Qing Li, and Ke Tang · 2022
Later among the works it cites.
Dr-label: Improving gnn models for catalysis systems by label deconstruction and reconstruction
Bowen Wang, Chen Liang, Jiaze Wang, Furui Liu, Shaogang Hao, Dong Li, Jianye Hao, Guangyong Chen, Xiaolong Zou, and Pheng-Ann Heng · 2023
Closest in time.