Fetching the paper…
Reading the bibliography…
In the traditional view of reinforcement learning, the agent's goal is to find an optimal policy that maximizes its expected sum of rewards.
Continual learning in reinforcement environments
Mark Bishop Ring · 1994
Earlier work this paper cites.
A theory of universal artificial intelligence based on algorithmic complexity
Marcus Hutter · 2000
Earlier work this paper cites.
On the role of tracking in stationary environments
Richard S. Sutton, Anna Koop, and David Silver · 2007
Earlier work this paper cites.
Online bandit learning against an adaptive adversary: from regret to policy regret
Raman Arora, Ofer Dekel, and Ambuj Tewari · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin A. Riedmiller, Andreas Kirkeby Fidjeland, Georg Ostrovski, Stig Petersen, Charlie Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Christopher J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Earlier work this paper cites.
An emphatic approach to the problem of off-policy temporal-difference learning
Richard S. Sutton, A. Rupam Mahmood, and Martha White · 2016
Cited alongside, same era.
Deepstack: Expert-level artificial intelligence in heads-up no-limit poker
Matej Moravčík, Martin Schmid, Neil Burch, Viliam Lisỳ, Dustin Morrill, Nolan Bard, Trevor Davis, Kevin Waugh, Michael Johanson, and Michael Bowling · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Policy regret in repeated games
Raman Arora, Michael Dinitz, Teodor V. Marinov, and Mehryar Mohri · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Cited alongside, same era.
Towards continual reinforcement learning: A review and perspectives
Khimya Khetarpal, Matthew Riemer, Irina Rish, and Doina Precup · 2022
Later among the works it cites.
The partially observable history process
Dustin Morrill, Amy R. Greenwald, and Michael Bowling · 2022
Later among the works it cites.
Outracing champion Gran Turismo drivers with deep reinforcement learning
Peter R Wurman, Samuel Barrett, Kenta Kawamoto, James MacGlashan, Kaushik Subramanian, Thomas J Walsh, Roberto Capobianco, Alisa Devlic, Franziska Eckert, Florian Fuchs, et al · 2022
Later among the works it cites.
Settling the reward hypothesis
Michael Bowling, John D. Martin, David Abel, and Will Dabney · 2023
Later among the works it cites.
Continual learning as computationally constrained reinforcement learning
Saurabh Kumar, Henrik Marklund, Ashish Rao, Yifan Zhu, Hong Jun Jeon, Yueyang Liu, and Benjamin Van Roy · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marc G Bellemare, Salvatore Candido, Pablo Samuel Castro, Jun Gong, Marlos C Machado, Subhodeep Moitra, Sameera S Ponda, and Ziyu Wang · 2020
Cited alongside, same era.
Simple agent, complex environment: Efficient reinforcement learning with agent states
Shi Dong, Benjamin Van Roy, and Zhengyuan Zhou · 2022
Cited alongside, same era.
The 37 implementation details of proximal policy optimization
Shengyi Huang, Rousslan Fernand Julien Dossa, Antonin Raffin, Anssi Kanervisto, and Weixun Wang · 2022
Cited alongside, same era.
A definition of continual reinforcement learning
David Abel, André Barreto, Benjamin Van Roy, Doina Precup, Hado van Hasselt, and Satinder Singh
Cited in the paper.
Three dogmas of reinforcement learning
David Abel, Mark K. Ho, and Anna Harutyunyan
Cited in the paper.
Efficient deviation types and learning for hindsight rationality in extensive-form games
Dustin Morrill, Ryan D’Orazio, Marc Lanctot, Reca Sarfati, James Wright, Amy Greenwald, and Michael Bowling
Cited in the paper.
Hindsight and sequential rationality of correlated play
Dustin Morrill, Ryan D’Orazio, Reca Sarfati, Marc Lanctot, James R. Wright, Amy Greenwald, and Michael Bowling
Cited in the paper.
Jelly bean world: A testbed for never-ending learning
Emmanouil Antonios Platanios, Abulhair Saparov, and Tom Mitchell · 2023
Later among the works it cites.
The big world hypothesis and its ramifications for artificial intelligence
Khurram Javed and Richard S. Sutton · 2024
Later among the works it cites.