Fetching the paper…
Reading the bibliography…
The reward hypothesis posits that, "all of what we mean by goals and purposes can be well thought of as maximization of the expected value of the cumulative sum of a received scalar signal (reward)." We aim to fully settle this hypothesis.
Truth and probability
Ramsey, F. P · 1926
Earlier work this paper cites.
Multidimensional utilities (rev)
Hausner, M · 1953
Earlier work this paper cites.
Theory of Games and Economic Behavior
von Neumann, J. and Morgenstern, O · 1953
Earlier work this paper cites.
Specimen theoriae novae de mensura sortis (trans. in 1954 as exposition of a new theory on the measurement of risk)
Bernoulli, D · 1954
Earlier work this paper cites.
Stationary ordinal utility and impatience
Koopmans, T. C · 1960
Earlier work this paper cites.
Utility theory without the completeness axiom
Aumann, R. J · 1962
Earlier work this paper cites.
Stationary utility and time perspective
Koopmans, T. C., Diamond, P. A., and Williamson, R. E · 1964
Earlier work this paper cites.
Ordinal dynamic programming
Sobel, M. J · 1975
Earlier work this paper cites.
Temporal resolution of uncertainty and dynamic choice theory
Kreps, D. M. and Porteus, E. L · 1978
Earlier work this paper cites.
The psychology of preferences
Kahneman, D. and Tversky, A · 1982
Earlier work this paper cites.
Judgment under uncertainty: Heuristics and biases
Kahneman, D., Slovic, S. P., Slovic, P., and Tversky, A · 1982
Earlier work this paper cites.
Mental models: Towards a cognitive science of language, inference, and consciousness
Johnson-Laird, P. N · 1983
Earlier work this paper cites.
Extensional versus intuitive reasoning: The conjunction fallacy in probability judgment
Tversky, A. and Kahneman, D · 1983
Earlier work this paper cites.
On effectively computable realizations of choice functions: Dedicated to professors kenneth j. arrow and anil nerode
Lewis, A. A · 1985
Earlier work this paper cites.
Expected utility hypothesis
Machina, M. J · 1990
Earlier work this paper cites.
Rationality, computability, and complexity
Rustem, B. and Velupillai, K · 1990
Earlier work this paper cites.
Decisions with multiple objectives: preferences and value trade-offs
Keeney, R. L., Raiffa, H., and Meyer, R. F · 1993
Earlier work this paper cites.
Acting optimally in partially observable stochastic domains
Cassandra, A. R., Kaelbling, L. P., and Littman, M. L · 1994
Cited alongside, same era.
Average reward reinforcement learning: Foundations, algorithms, and empirical results
Mahadevan, S · 1996
Cited alongside, same era.
Multi-criteria reinforcement learning
Gábor, Z., Kalmár, Z., and Szepesvári, C · 1998
Cited alongside, same era.
What is artificial intelligence
McCarthy, J · 1998
Cited alongside, same era.
Computable preference and utility
Richter, M. K. and Wong, K.-C · 1999
Cited alongside, same era.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S., et al · 2000
Cited alongside, same era.
The reward hypothesis
A survey of preference-based reinforcement learning methods
Wirth, C., Akrour, R., Neumann, G., Fürnkranz, J., et al · 2017
Later among the works it cites.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Later among the works it cites.
Hyperbolic discounting and learning over multiple horizons
Fedus, W., Gelada, C., Bengio, Y., Bellemare, M. G., and Larochelle, H · 2019
Later among the works it cites.
Rethinking the discount factor in reinforcement learning: A decision theoretic approach
Pitis, S · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sutton, R. S · 2004
Cited alongside, same era.
Where do rewards come from?
Singh, S., Lewis, R. L., and Barto, A. G · 2009
Cited alongside, same era.
The st. petersburg paradox
Martin, R · 2011
Cited alongside, same era.
Axioms for rational reinforcement learning
Sunehag, P. and Hutter, M · 2011
Cited alongside, same era.
The sample-complexity of general reinforcement learning
Lattimore, T., Hutter, M., and Sunehag, P · 2013
Cited alongside, same era.
Theory of General Reinforcement Learning
Lattimore, T · 2014
Cited alongside, same era.
Constrained MDPs and the reward hypothesis
Szepesvári, C · 2020
Later among the works it cites.
On the expressivity of Markov reward
Abel, D., Dabney, W., Harutyunyan, A., Ho, M. K., Littman, M. L., Precup, D., and Singh, S · 2021
Later among the works it cites.
Simple agent, complex environment: Efficient reinforcement learning with agent states
Dong, S., Van Roy, B., and Zhou, Z · 2021
Later among the works it cites.
Reinforcement learning, bit by bit
Lu, X., Van Roy, B., Dwaracherla, V., Ibrahimi, M., Osband, I., and Wen, Z · 2021
Later among the works it cites.
Abstractions of General Reinforcement Learning
Majeed, S. J · 2021
Later among the works it cites.
Reward is enough
Silver, D., Singh, S., Precup, D., and Sutton, R. S · 2021
Later among the works it cites.
Expressing non-Markov reward to a Markov agent
Abel, D., Barreto, A., Bowling, M., Dabney, W., Hansen, S., Harutyunyan, A., Ho, M. K., Kumar, R., Littman, M. L., Precup, D., and Satinder, S · 2022
Closest in time.
On the expressivity of multidimensional Markov reward
Miura, S · 2022
Closest in time.
Rational multi-objective agents must admit non-markov reward representations
Pitis, S., Bailey, D., and Ba, J · 2022
Closest in time.
Utility theory for sequential decision making
Shakerinava, M. and Ravanbakhsh, S · 2022
Closest in time.
Distributional Reinforcement Learning
Bellemare, M. G., Dabney, W., and Rowland, M · 2023
Closest in time.