Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) typically defines a discount factor as part of the Markov Decision Process.
A note on measurement of utility
Paul A Samuelson · 1937
Earlier work this paper cites.
Myopia and inconsistency in dynamic utility maximization
Robert Henry Strotz · 1955
Earlier work this paper cites.
A markovian decision process
Richard Bellman · 1957
Earlier work this paper cites.
On a routing problem
Richard Bellman · 1958
Earlier work this paper cites.
Specious reward: a behavioral theory of impulsiveness and impulse control
George Ainslie · 1975
Earlier work this paper cites.
Preference reversal and self control: Choice as a function of reward amount and delay
Leonard Green, Ewin B Fisher, Steven Perlow, and Lisa Sherman · 1981
Earlier work this paper cites.
Probability and delay of reinforcement as factors in discrete-trial choice
James E Mazur · 1985
Earlier work this paper cites.
An adjusting procedure for studying delayed reinforcement
James E Mazur · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
Rule-injection hints as a means of improving network performance and learning time
Steven C Suddarth and YL Kergosien · 1990
Earlier work this paper cites.
Picoeconomics: The strategic interaction of successive motivational states within the person
George Ainslie · 1992
Earlier work this paper cites.
Scaling reinforcement learning algorithms by learning variable temporal resolution models
Satinder P Singh · 1992
Earlier work this paper cites.
Q-learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E Hinton · 1993
Earlier work this paper cites.
Markov decision models with weighted discounted criteria
Eugene A Feinberg and Adam Shwartz · 1994
Earlier work this paper cites.
Temporal discounting and preference reversals in choice between delayed outcomes
Leonard Green, Nathanael Fristoe, and Joel Myerson · 1994
Earlier work this paper cites.
Neuro-dynamic programming: an overview
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Discounting of delayed rewards: Models of individual choice
Joel Myerson and Leonard Green · 1995
Earlier work this paper cites.
Td models: Modeling the world at a mixture of time scales
Richard S Sutton · 1995
Earlier work this paper cites.
Neuro-dynamic programming , volume 5
Dimitri P Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
Out of control: Visceral influences on behavior
George Loewenstein · 1996
Earlier work this paper cites.
A framework for mesencephalic dopamine systems based on predictive hebbian learning
P Read Montague, Peter Dayan, and Terrence J Sejnowski · 1996
Earlier work this paper cites.
Rate of temporal discounting decreases with amount of reward
Leonard Green, Joel Myerson, and Edward McFadden · 1997
Earlier work this paper cites.
Normative and descriptive models of decision making: time discounting and risk sensitivity
Alex Kacelnik · 1997
Cited alongside, same era.
Choice, delay, probability, and conditioned reinforcement
James E Mazur · 1997
Cited alongside, same era.
Adaptive critic designs
Danil V Prokhorov and Donald C Wunsch · 1997
Cited alongside, same era.
A neural substrate of prediction and reward
Wolfram Schultz, Peter Dayan, and P Read Montague · 1997
Cited alongside, same era.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Cited alongside, same era.
On hyperbolic discounting and uncertain hazard rates
Peter D Sozou · 1998
Cited alongside, same era.
Neural models of temporal discounting
A David Redish and Zeb Kurth-Nelson · 2010
Later among the works it cites.
Time consistent discounting
Tor Lattimore and Marcus Hutter · 2011
Later among the works it cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, Adam White, and Doina Precup · 2011
Later among the works it cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Later among the works it cites.
Expressing tasks robustly via multiple discount factors
Ashley Edwards, Michael L Littman, and Charles L Isbell · 2015
Later among the works it cites.
How to discount deep reinforcement learning: Towards new dynamic strategies
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1998
Cited alongside, same era.
Constrained dynamic programming with two discount factors: Applications and an algorithm
Eugene A Feinberg and Adam Shwartz · 1999
Cited alongside, same era.
Catastrophic forgetting in connectionist networks
Robert M French · 1999
Cited alongside, same era.
Behavioral considerations suggest an average reward td model of the dopamine system
Nathaniel D Daw and David S Touretzky · 2000
Cited alongside, same era.
Time discounting and time preference: A critical review
Shane Frederick, George Loewenstein, and Ted O’donoghue · 2002
Cited alongside, same era.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Cited alongside, same era.
Vincent François-Lavet, Raphael Fonteneau, and Damien Ernst · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Later among the works it cites.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Learning to navigate in complex environments
Piotr Mirowski, Razvan Pascanu, Fabio Viola, Hubert Soyer, Andrew J Ballard, Andrea Banino, Misha Denil, Ross Goroshin, Laurent Sifre, Koray Kavukcuoglu, et al · 2016
Later among the works it cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Later among the works it cites.
Playing fps games with deep reinforcement learning
Guillaume Lample and Devendra Singh Chaplot · 2017
Later among the works it cites.
Average reward optimization with multiple discounting reinforcement learners
Chris Reinke, Eiji Uchibe, and Kenji Doya · 2017
Later among the works it cites.
Unifying task specification in reinforcement learning
Martha White · 2017
Later among the works it cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Later among the works it cites.
Dopamine: A research framework for deep reinforcement learning
Pablo Samuel Castro, Subhodeep Moitra, Carles Gelada, Saurabh Kumar, and Marc G. Bellemare · 2018
Later among the works it cites.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2018
Later among the works it cites.
Revisiting the Arcade Learning Environment: Evaluation protocols and open problems for general agents
Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2018
Later among the works it cites.
Openai five
OpenAI · 2018
Later among the works it cites.
Generalizing value estimation over timescal
Craig Sherstan, James MacGlashan, and Patrick M. Pilarski · 2018
Later among the works it cites.
Many-goals reinforcement learning
Vivek Veeriah, Junhyuk Oh, and Satinder Singh · 2018
Later among the works it cites.
Meta-gradient reinforcement learning
Zhongwen Xu, Hado van Hasselt, and David Silver · 2018
Later among the works it cites.
Rethinking the Discount Factor in Reinforcement Learning: A Decision Theoretic Approach
Silviu Pitis · 2019
Closest in time.
Separating value functions across time-scales
Joshua Romoff, Peter Henderson, Ahmed Touati, Yann Ollivier, Emma Brunskill, and Joelle Pineau · 2019
Closest in time.