Fetching the paper…
Reading the bibliography…
Modern reinforcement learning has been conditioned by at least three dogmas.
Two dogmas of empiricism
Willard Van Orman Quine · 1951
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Le comportement de l’homme rationnel devant le risque: critique des postulats et axiomes de l’école américaine
Maurice Allais · 1953
Earlier work this paper cites.
Theory of Games and Economic Behavior
John von Neumann and Oskar Morgenstern · 1953
Earlier work this paper cites.
A behavioral model of rational choice
Herbert A Simon · 1955
Earlier work this paper cites.
A Markovian decision process
Richard Bellman · 1957
Earlier work this paper cites.
Stationary ordinal utility and impatience
Tjalling C Koopmans · 1960
Earlier work this paper cites.
GPS, a program that simulates human thought
Allen Newell and Herbert Alexander Simon · 1961
Earlier work this paper cites.
Utility theory without the completeness axiom
Robert J Aumann · 1962
Earlier work this paper cites.
Discrete dynamic programming
David Blackwell · 1962
Earlier work this paper cites.
The structure of scientific revolutions
Thomas S Kuhn · 1962
Earlier work this paper cites.
The function of dogma in scientific research
Thomas S Kuhn · 1963
Earlier work this paper cites.
Heuristic DENDRAL: A program for generating explanatory hypotheses
Bruce Buchanan, Georgia Sutherland, and Edward A Feigenbaum · 1969
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
Vladimir N Vapnik and Aleksei Y Chervonenkis · 1971
Earlier work this paper cites.
Risk-sensitive Markov decision processes
Ronald A Howard and James E Matheson · 1972
Earlier work this paper cites.
Artificial intelligence: A general survey
James Lighthill, Stuart Sutherland, Roger Needham, and Christopher Longuet-Higgins · 1973
Earlier work this paper cites.
Ordinal dynamic programming
Matthew J Sobel · 1975
Earlier work this paper cites.
"Expected Utility" analysis without the independence axiom
Mark J Machina · 1982
Earlier work this paper cites.
A theory of the learnable
Leslie G Valiant · 1984
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
The intentional stance
Daniel C Dennett · 1989
Earlier work this paper cites.
Learning from delayed rewards
Christopher J.C.H. Watkins · 1989
Earlier work this paper cites.
Expected utility hypothesis
Mark J Machina · 1990
Earlier work this paper cites.
Q Q -learning
Christopher J.C.H. Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Acting optimally in partially observable stochastic domains
Anthony R Cassandra, Leslie Pack Kaelbling, and Michael L Littman · 1994
Earlier work this paper cites.
Unified theories of cognition
Allen Newell · 1994
Earlier work this paper cites.
Continual learning in reinforcement environments
Mark B Ring · 1994
Earlier work this paper cites.
Provably bounded-optimal agents
Stuart J Russell and Devika Subramanian · 1994
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik · 1995
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Stuart J Russell and Peter Norvig · 1995
Earlier work this paper cites.
Temporal difference learning and TD-gammon
Gerald Tesauro et al · 1995
Earlier work this paper cites.
Intelligent agents: Theory and practice
Michael Wooldridge and Nicholas R Jennings · 1995
Earlier work this paper cites.
Average reward reinforcement learning: Foundations, algorithms, and empirical results
Sridhar Mahadevan · 1996
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
John Tsitsiklis and Benjamin Van Roy · 1996
Earlier work this paper cites.
Child: A first step towards continual learning
Mark B Ring · 1997
Earlier work this paper cites.
The extended mind
Andy Clark and David Chalmers · 1998
Earlier work this paper cites.
Machines, plants and animals: the origins of agency
Fred I Dretske · 1999
Earlier work this paper cites.
A theory of universal artificial intelligence based on algorithmic complexity
Marcus Hutter · 2000
Earlier work this paper cites.
Self-optimizing and Pareto-optimal policies in general environments based on Bayes-mixtures
Marcus Hutter · 2002
Earlier work this paper cites.
Risk-sensitive reinforcement learning
Oliver Mihatsch and Ralph Neuneier · 2002
Earlier work this paper cites.
Efficient solution algorithms for factored MDPs
Carlos Guestrin, Daphne Koller, Ronald Parr, and Shobha Venkataraman · 2003
Earlier work this paper cites.
Universal artificial intelligence: Sequential decisions based on algorithmic probability
Marcus Hutter · 2004
Earlier work this paper cites.
The reward hypothesis, 2004
Richard S Sutton · 2004
Earlier work this paper cites.
Toward a formal framework for continual learning
Mark B Ring · 2005
Earlier work this paper cites.
A proposal for the Dartmouth summer research project on Artificial Intelligence, August 31, 1955
John McCarthy, Marvin L Minsky, Nathaniel Rochester, and Claude E Shannon · 2006
Cited alongside, same era.
PAC model-free reinforcement learning
Alexander L Strehl, Lihong Li, Eric Wiewiora, John Langford, and Michael L Littman · 2006
Cited alongside, same era.
Linearly-solvable Markov decision problems
Emanuel Todorov · 2006
Cited alongside, same era.
The epoch-greedy algorithm for contextual multi-armed bandits
John Langford and Tong Zhang · 2007
Cited alongside, same era.
Computer science as empirical inquiry: Symbols and search
Allen Newell and Herbert A Simon · 2007
Cited alongside, same era.
Mechanisms of ecological rationality: heuristics and environments that make
Peter M Todd and Gerd Gigerenzer · 2007
Cited alongside, same era.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz · 2017
Later among the works it cites.
The Barbados 2018 list of open issues in continual learning
Tom Schaul, Hado van Hasselt, Joseph Modayil, Martha White, Adam White, Pierre-Luc Bacon, Jean Harb, Shibl Mourad, Marc Bellemare, and Doina Precup · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto · 2018
Later among the works it cites.
A strongly asymptotically optimal agent in general environments
Michael K Cohen, Elliot Catt, and Marcus Hutter · 2019
Later among the works it cites.
Provably efficient RL with rich observations via latent state decoding
Simon Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudik, and John Langford · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An object-oriented representation for efficient reinforcement learning
Carlos Diuk, Andre Cohen, and Michael L Littman · 2008
Cited alongside, same era.
Autonomous transfer for reinforcement learning
Matthew E Taylor, Gregory Kuhlmann, and Peter Stone · 2008
Cited alongside, same era.
Defining agency: Individuality, normativity, asymmetry, and spatio-temporality in action
Xabier E Barandiaran, Ezequiel Di Paolo, and Marieke Rohde · 2009
Cited alongside, same era.
The free-energy principle: a unified brain theory?
Karl J Friston · 2010
Cited alongside, same era.
Asymptotically optimal agents
Tor Lattimore and Marcus Hutter · 2011
Cited alongside, same era.
A unified framework for resource-bounded autonomous agents interacting with unknown environments
Pedro A Ortega · 2011
Cited alongside, same era.
Hyperbolic discounting and learning over multiple horizons
William Fedus, Carles Gelada, Yoshua Bengio, Marc G Bellemare, and Hugo Larochelle · 2019
Later among the works it cites.
A brief history of artificial intelligence: On the past, present, and future of artificial intelligence
Michael Haenlein and Andreas Kaplan · 2019
Later among the works it cites.
On value functions and the agent-environment boundary
Nan Jiang · 2019
Later among the works it cites.
Batch policy learning under constraints
Hoang Le, Cameron Voloshin, and Yisong Yue · 2019
Later among the works it cites.
Rethinking the discount factor in reinforcement learning: A decision theoretic approach
Silviu Pitis · 2019
Later among the works it cites.
Model-based RL in contextual decision processes: PAC bounds and exponential improvements over model-free approaches
Wen Sun, Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2019
Later among the works it cites.
Cybernetics or Control and Communication in the Animal and the Machine
Norbert Wiener · 2019
Later among the works it cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Action and perception as divergence minimization
Danijar Hafner, Pedro A Ortega, Jimmy Ba, Thomas Parr, Karl J Friston, and Nicolas Heess · 2020
Later among the works it cites.
What is an agent?
Anna Harutyunyan · 2020
Later among the works it cites.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
On the expressivity of Markov reward
David Abel, Will Dabney, Anna Harutyunyan, Mark K. Ho, Michael L Littman, Doina Precup, and Satinder Singh · 2021
Later among the works it cites.
Constrained Markov decision processes
Eitan Altman · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Later among the works it cites.
The Alignment Problem: Machine Learning and Human Values , pp. 130–131
Brian Christian · 2021
Later among the works it cites.
Multi-agent reinforcement learning with temporal logic specifications
Lewis Hammond, Alessandro Abate, Julian Gutierrez, and Michael Wooldridge · 2021
Later among the works it cites.
Reinforcement learning, bit by bit
Xiuyuan Lu, Benjamin Van Roy, Vikranth Dwaracherla, Morteza Ibrahimi, Ian Osband, and Zheng Wen · 2021
Later among the works it cites.
Benefits of assistance over reward learning, 2021
Rohin Shah, Pedro Freire, Neel Alex, Rachel Freedman, Dmitrii Krasheninnikov, Lawrence Chan, Michael D Dennis, Pieter Abbeel, Anca Dragan, and Stuart Russell · 2021
Later among the works it cites.
Reward is enough for convex MDPs
Tom Zahavy, Brendan O’Donoghue, Guillaume Desjardins, and Satinder Singh · 2021
Later among the works it cites.
Simple agent, complex environment: Efficient reinforcement learning with agent states
Shi Dong, Benjamin Van Roy, and Zhengyuan Zhou · 2022
Later among the works it cites.
Reward machines: Exploiting reward function structure in reinforcement learning
Rodrigo Toro Icarte, Toryn Q Klassen, Richard Valenzano, and Sheila A McIlraith · 2022
Later among the works it cites.
Towards continual reinforcement learning: A review and perspectives
Khimya Khetarpal, Matthew Riemer, Irina Rish, and Doina Precup · 2022
Later among the works it cites.
Time to take embodiment seriously
John D Martin · 2022
Later among the works it cites.
Challenging common assumptions in convex reinforcement learning
Mirco Mutti, Riccardo De Santi, Piersilvio De Bartolomeis, and Marcello Restelli · 2022
Later among the works it cites.
Utility theory for sequential decision making
Mehran Shakerinava and Siamak Ravanbakhsh · 2022
Later among the works it cites.
The quest for a common model of the intelligent decision maker
Richard S Sutton · 2022
Later among the works it cites.
The evolution of agency: Behavioral organization from lizards to humans
Michael Tomasello · 2022
Later among the works it cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Later among the works it cites.
Prediction and control in continual reinforcement learning
Nishanth Anand and Doina Precup · 2023
Later among the works it cites.
The parts of an imperfect agent
Sara Aronowitz · 2023
Later among the works it cites.
Distributional reinforcement learning
Marc G Bellemare, Will Dabney, and Mark Rowland · 2023
Later among the works it cites.
Settling the reward hypothesis
Michael Bowling, John D Martin, David Abel, and Will Dabney · 2023
Later among the works it cites.
Discovering agents
Zachary Kenton, Ramana Kumar, Sebastian Farquhar, Jonathan Richens, Matt MacDermott, and Tom Everitt · 2023
Later among the works it cites.
Continual learning as computationally constrained reinforcement learning
Saurabh Kumar, Henrik Marklund, Ashish Rao, Yifan Zhu, Hong Jun Jeon, Yueyang Liu, and Benjamin Van Roy · 2023
Later among the works it cites.
Convex reinforcement learning in finite trials
Mirco Mutti, Riccardo De Santi, Piersilvio De Bartolomeis, and Marcello Restelli · 2023
Later among the works it cites.
MAESTRO: Open-ended environment design for multi-agent reinforcement learning
Mikayel Samvelyan, Akbir Khan, Michael D Dennis, Minqi Jiang, Jack Parker-Holder, Jakob Nicolaus Foerster, Roberta Raileanu, and Tim Rocktäschel · 2023
Later among the works it cites.
Near-minimax-optimal risk-sensitive reinforcement learning with CVaR
Kaiwen Wang, Nathan Kallus, and Wen Sun · 2023
Later among the works it cites.
Robust agents learn causal world models
Jonathan Richens and Tom Everitt · 2024
Closest in time.