Fetching the paper…
Reading the bibliography…
Reward is the driving force for reinforcement-learning agents.
Theory of Games and Economic Behavior
John von Neumann and Oskar Morgenstern · 1953
Earlier work this paper cites.
Representation of a preference ordering by a numerical function
Gerard Debreu · 1954
Earlier work this paper cites.
Stationary ordinal utility and impatience
Tjalling C. Koopmans · 1960
Earlier work this paper cites.
Preference order dynamic programming
L. G. Mitten · 1974
Earlier work this paper cites.
Ordinal dynamic programming
Matthew J. Sobel · 1975
Earlier work this paper cites.
A new polynomial-time algorithm for linear programming
Narendra Karmarkar · 1984
Earlier work this paper cites.
Notes on the Theory of Choice
David Kreps · 1988
Earlier work this paper cites.
Interactions between learning and evolution
David Ackley and Michael L. Littman · 1992
Earlier work this paper cites.
Reward functions for accelerated learning
Maja J. Mataric · 1994
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Stuart J. Russell and Peter Norvig · 1994
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y. Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y. Ng, Stuart J. Russell, et al · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Pieter Abbeel and Andrew Y. Ng · 2004
Earlier work this paper cites.
The reward hypothesis, 2004
Richard S. Sutton · 2004
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Satinder Singh, Andrew G. Barto, and Nuttapong Chentanez · 2005
Earlier work this paper cites.
Apprenticeship learning using linear programming
Umar Syed, Michael Bowling, and Robert E. Schapire · 2008
Earlier work this paper cites.
Reinforcement learning or active inference?
Karl J. Friston, Jean Daunizeau, and Stefan J. Kiebel · 2009
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The TAMER framework
W. Bradley Knox and Peter Stone · 2009
Earlier work this paper cites.
Where do rewards come from?
Satinder Singh, Richard L Lewis, and Andrew G Barto · 2009
Earlier work this paper cites.
The free-energy principle: a unified brain theory?
Karl J. Friston · 2010
Earlier work this paper cites.
On separating agent designer goals from agent goals: Breaking the preferences–parameters confound, 2010
Satinder Singh, Richard L. Lewis, Jonathan Sorg, Andrew G. Barto, and Akram Helou · 2010
Earlier work this paper cites.
Reward design via online gradient ascent
Jonathan Sorg, Richard L. Lewis, and Satinder Singh · 2010
Earlier work this paper cites.
The Optimal Reward Problem: Designing Effective Reward for Bounded Agents
Jonathan Sorg · 2011
Cited alongside, same era.
Axioms for rational reinforcement learning
Peter Sunehag and Marcus Hutter · 2011
Cited alongside, same era.
Markov decision processes with ordinal rewards: Reference point-based preferences
Paul Weng · 2011
Cited alongside, same era.
A Bayesian approach for policy learning from trajectory preference queries
Aaron Wilson, Alan Fern, and Prasad Tadepalli · 2012
Cited alongside, same era.
The steady-state control problem for Markov decision processes
Sundararaman Akshay, Nathalie Bertrand, Serge Haddad, and Loic Helouet · 2013
Cited alongside, same era.
Discounting axioms imply risk neutrality
Matthew J. Sobel · 2013
Cited alongside, same era.
Using reward machines for high-level task specification and decomposition in reinforcement learning
Rodrigo Toro Icarte, Toryn Klassen, Richard Valenzano, and Sheila McIlraith · 2018
Later among the works it cites.
Building safe artificial intelligence: specification, robustness, and assurance, 2018
Pedro A. Ortega, Vishal Maini, and the DeepMind Safety Team · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 2018
Later among the works it cites.
Teaching multiple tasks to an RL agent using LTL
Rodrigo Toro Icarte, Toryn Q. Klassen, Richard Valenzano, and Sheila A. McIlraith · 2018
Later among the works it cites.
Learning to parse natural language to grounded reward functions with weak supervision
Edward C. Williams, Nakul Gopalan, Mine Rhee, and Stefanie Tellex · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning and the reward engineering principle
Daniel Dewey · 2014
Cited alongside, same era.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 2014
Cited alongside, same era.
Grounding English commands to reward functions
James MacGlashan, Monica Babes-Vroman, Marie desJardins, Michael L. Littman, Smaranda Muresan, Shawn Squire, Stefanie Tellex, Dilip Arumugam, and Lei Yang · 2015
Cited alongside, same era.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell · 2016
Cited alongside, same era.
Convergent actor critic by humans
James MacGlashan, Michael L. Littman, David L. Roberts, Robert Loftin, Bei Peng, and Matthew E. Taylor · 2016
Cited alongside, same era.
Repeated inverse reinforcement learning
Kareem Amin, Nan Jiang, and Satinder Singh · 2017
Cited alongside, same era.
William Fedus, Carles Gelada, Yoshua Bengio, Marc G. Bellemare, and Hugo Larochelle · 2019
Later among the works it cites.
Rethinking the discount factor in reinforcement learning: A decision theoretic approach
Silviu Pitis · 2019
Later among the works it cites.
Action and perception as divergence minimization
Danijar Hafner, Pedro A. Ortega, Jimmy Ba, Thomas Parr, Karl J. Friston, and Nicolas Heess · 2020
Later among the works it cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Hong Jun Jeon, Smitha Milli, and Anca Dragan · 2020
Later among the works it cites.
A composable specification language for reinforcement learning tasks
Kishor Jothimurugan, Rajeev Alur, and Osbert Bastani · 2020
Later among the works it cites.
Maximum reward formulation in reinforcement learning
Sai Krishna Gottipati, Yashaswi Pathak, Rohan Nuttall, Raviteja Chunduru, Ahmed Touati, Sriram Ganapathi Subramanian, Matthew E. Taylor, and Sarath Chandar · 2020
Later among the works it cites.
REALab: An embedded perspective on tampering
Ramana Kumar, Jonathan Uesato, Richard Ngo, Tom Everitt, Victoria Krakovna, and Shane Legg · 2020
Later among the works it cites.
Dueling posterior sampling for preference-based reinforcement learning
Ellen Novoseller, Yibing Wei, Yanan Sui, Yisong Yue, and Joel Burdick · 2020
Later among the works it cites.
Constrained MDPs and the reward hypothesis, 2020
Csaba Szepesvári · 2020
Later among the works it cites.
A Boolean task algebra for reinforcement learning
Geraud Nangue Tasse, Steven James, and Benjamin Rosman · 2020
Later among the works it cites.
Preference-based reinforcement learning with finite-time guarantees
Yichong Xu, Ruosong Wang, Lin Yang, Aarti Singh, and Artur Dubrawski · 2020
Later among the works it cites.
What can learned intrinsic rewards capture?
Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu, Manuel Kroiss, Hado van Hasselt, David Silver, and Satinder Singh · 2020
Later among the works it cites.
The Alignment Problem: Machine Learning and Human Values , pages 130–131
Brian Christian · 2021
Closest in time.
Multi-agent reinforcement learning with temporal logic specifications
Lewis Hammond, Alessandro Abate, Julian Gutierrez, and Michael Wooldridge · 2021
Closest in time.
Benefits of assistance over reward learning, 2021
Rohin Shah, Pedro Freire, Neel Alex, Rachel Freedman, Dmitrii Krasheninnikov, Lawrence Chan, Michael D. Dennis, Pieter Abbeel, Anca Dragan, and Stuart Russell · 2021
Closest in time.
Reward is enough
David Silver, Satinder Singh, Doina Precup, and Richard S. Sutton · 2021
Closest in time.