Fetching the paper…
Reading the bibliography…
How do we formalize the challenge of credit assignment in reinforcement learning? Common intuition would draw attention to reward sparsity as a key contributor to difficult credit assignment and traditional heuristics would look to temporal recency for the solution, calling upon the classic eligibility trace.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
A markovian decision process
Richard Bellman · 1957
Earlier work this paper cites.
Coding theorems for a discrete source with a fidelity criterion
Claude E. Shannon · 1959
Earlier work this paper cites.
Steps toward artificial intelligence
Marvin Minsky · 1961
Earlier work this paper cites.
Rate Distortion Theory: A Mathematical Basis for Data Compression
Toby Berger · 1971
Earlier work this paper cites.
The variance of discounted markov decision processes
Matthew J Sobel · 1982
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Richard S Sutton · 1985
Earlier work this paper cites.
Discounted mdp’s: Distribution functions and exponential utility maximization
Kun-Jen Chung and Matthew J Sobel · 1987
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Causality, feedback and directed information
James Massey · 1990
Earlier work this paper cites.
Markov Decision Processes—Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Reinforcement learning with replacing eligibility traces
Satinder P Singh and Richard S Sutton · 1996
Earlier work this paper cites.
Directed information for channels with feedback
Gerhard Kramer · 1998
Earlier work this paper cites.
Introduction to reinforcement learning
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart J Russell · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S Sutton, and Satinder P Singh · 2000
Earlier work this paper cites.
Capacity of continuous channels with memory via directed information neural estimator
Ziv Aharoni, Dor Tsur, Ziv Goldfeld, and Haim Henry Permuter · 2003
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Nuttapong Chentanez, Andrew G Barto, and Satinder P Singh · 2005
Earlier work this paper cites.
Conservation of mutual and directed information
James L Massey and Peter C Massey · 2005
Earlier work this paper cites.
A markov decision approach to feedback channel capacity
Sekhar Tatikonda · 2005
Earlier work this paper cites.
Towards a unified theory of state abstraction for MDPs
Lihong Li, Thomas J. Walsh, and Michael L. Littman · 2006
Earlier work this paper cites.
Linearly-solvable markov decision problems
Emanuel Todorov · 2007
Earlier work this paper cites.
Ziv Aharoni, Oron Sabag, and Haim Henri Permuter · 2008
Earlier work this paper cites.
On directed information and gambling
Haim H Permuter, Young-Han Kim, and Tsachy Weissman · 2008
Cited alongside, same era.
The capacity of channels with feedback
Sekhar Tatikonda and Sanjoy Mitter · 2008
Cited alongside, same era.
Robot trajectory optimization using approximate inference
Marc Toussaint · 2009
Cited alongside, same era.
Nonparametric return distribution approximation for reinforcement learning
Tetsuro Morimura, Masashi Sugiyama, Hisashi Kashima, Hirotaka Hachiya, and Toshiyuki Tanaka · 2010
Cited alongside, same era.
Td_gamma: Re-evaluating complex backups in temporal difference learning
George Konidaris, Scott Niekum, and Philip S Thomas · 2011
Cited alongside, same era.
Information, utility and bounded rationality
Daniel Alexander Ortega and Pedro Alejandro Braun · 2011
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Later among the works it cites.
A unified bellman equation for causal information and value in markov decision processes
Stas Tiomkin and Naftali Tishby · 2017
Later among the works it cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Information theory of decisions and actions
Naftali Tishby and Daniel Polani · 2011
Cited alongside, same era.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2012
Cited alongside, same era.
Optimal control as a graphical model inference problem
Hilbert J Kappen, Vicenç Gómez, and Manfred Opper · 2012
Cited alongside, same era.
Parametric return density estimation for reinforcement learning
Tetsuro Morimura, Masashi Sugiyama, Hisashi Kashima, Hirotaka Hachiya, and Toshiyuki Tanaka · 2012
Cited alongside, same era.
Trading value and information in mdps
Jonathan Rubin, Ohad Shamir, and Naftali Tishby · 2012
Cited alongside, same era.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Brian D Ziebart · 2012
Cited alongside, same era.
Matteo Hessel, Joseph Modayil, Hado van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Gheshlaghi Azar, and David Silver · 2018
Later among the works it cites.
Sparse attentive backtracking: Temporal credit assignment through reminding
Nan Rosemary Ke, Anirudh Goyal, Olexa Bilaniuk, Jonathan Binas, Michael C Mozer, Chris Pal, and Yoshua Bengio · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Later among the works it cites.
Information-theoretic privacy for smart metering systems with a rechargeable battery
Simon Li, Ashish Khisti, and Aditya Mahajan · 2018
Later among the works it cites.
Learning by playing solving sparse reward tasks from scratch
Martin Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom Wiele, Vlad Mnih, Nicolas Heess, and Jost Tobias Springenberg · 2018
Later among the works it cites.
Learning to optimize via information-directed sampling
Daniel Russo and Benjamin Van Roy · 2018
Later among the works it cites.
State abstraction as compression in apprenticeship learning
David Abel, Dilip Arumugam, Kavosh Asadi, Yuu Jinnai, Michael L Littman, and Lawson LS Wong · 2019
Later among the works it cites.
Rudder: Return decomposition for delayed rewards
Jose A Arjona-Medina, Michael Gillhofer, Michael Widrich, Thomas Unterthiner, Johannes Brandstetter, and Sepp Hochreiter · 2019
Later among the works it cites.
Credit assignment as a proxy for transfer in reinforcement learning
Johan Ferret, Raphaël Marinier, Matthieu Geist, and Olivier Pietquin · 2019
Later among the works it cites.
Information asymmetry in kl-regularized rl
Alexandre Galashov, Siddhant M Jayakumar, Leonard Hasenclever, Dhruva Tirumala, Jonathan Schwarz, Guillaume Desjardins, Wojciech M Czarnecki, Yee Whye Teh, Razvan Pascanu, and Nicolas Heess · 2019
Later among the works it cites.
Hindsight credit assignment
Anna Harutyunyan, Will Dabney, Thomas Mesnard, Mohammad Gheshlaghi Azar, Bilal Piot, Nicolas Heess, Hado P van Hasselt, Gregory Wayne, Satinder Singh, Doina Precup, et al · 2019
Later among the works it cites.
Optimizing agent behavior over long time scales by transporting value
Chia-Chun Hung, Timothy Lillicrap, Josh Abramson, Yan Wu, Mehdi Mirza, Federico Carnevale, Arun Ahuja, and Greg Wayne · 2019
Later among the works it cites.
Emi: Exploration with mutual information
Hyoungseok Kim, Jaekyeom Kim, Yeonwoo Jeong, Sergey Levine, and Hyun Oh Song · 2019
Later among the works it cites.
Deep exploration via randomized value functions
Ian Osband, Benjamin Van Roy, Daniel J Russo, and Zheng Wen · 2019
Later among the works it cites.
Separating value functions across time-scales
Joshua Romoff, Peter Henderson, Ahmed Touati, Emma Brunskill, Joelle Pineau, and Yann Ollivier · 2019
Later among the works it cites.
Exploiting hierarchy for learning and transfer in kl-regularized rl
Dhruva Tirumala, Hyeonwoo Noh, Alexandre Galashov, Leonard Hasenclever, Arun Ahuja, Greg Wayne, Razvan Pascanu, Yee Whye Teh, and Nicolas Heess · 2019
Later among the works it cites.
Keeping your distance: Solving sparse reward tasks using self-balancing shaped rewards
Alexander Trott, Stephan Zheng, Caiming Xiong, and Richard Socher · 2019
Later among the works it cites.
Hado van Hasselt, Sephora Madjiheurem, Matteo Hessel, David Silver, André Barreto, and Diana Borsa · 2020
Later among the works it cites.