Fetching the paper…
Reading the bibliography…
Many practical applications of reinforcement learning require agents to learn from sparse and delayed rewards.
Dynamic programming
Richard Bellman · 1957
Earlier work this paper cites.
Counterspeculation, auctions, and competitive sealed tenders
William Vickrey · 1961
Earlier work this paper cites.
Multipart pricing of public goods
Edward H Clarke · 1971
Earlier work this paper cites.
Incentives in teams
Theodore Groves · 1973
Earlier work this paper cites.
Optimal auction design
Roger B Myerson · 1981
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard S Sutton · 1988
Earlier work this paper cites.
Reward functions for accelerated learning
Maja J Mataric · 1994
Earlier work this paper cites.
Bias plus variance decomposition for zero-one loss functions
Ron Kohavi and David Wolpert · 1996
Earlier work this paper cites.
Stochastic analysis and control of real-time systems with random time delays
Johan Nilsson, Bo Bernhardsson, and Björn Wittenmark · 1998
Earlier work this paper cites.
Learning to drive a bicycle using reinforcement learning and shaping
Jette Randløv and Preben Alstrøm · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart J Russell · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng and Stuart Russell · 2000
Earlier work this paper cites.
Markov decision processes with delays and asynchronous cost collection
Konstantinos V Katsikopoulos and Sascha E Engelbrecht · 2003
Earlier work this paper cites.
Advanced Sampling Theory With Applications: How Michael”” Selected”” Amy , volume 2
Sarjinder Singh · 2003
Earlier work this paper cites.
Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
András Antos, Csaba Szepesvári, and Rémi Munos · 2008
Earlier work this paper cites.
Biasing approximate dynamic programming with a lower discount factor
Marek Petrik and Bruno Scherrer · 2008
Earlier work this paper cites.
Effects of feedback delay on learning
Hazhir Rahmandad, Nelson Repenning, and John Sterman · 2009
Earlier work this paper cites.
Learning and planning in environments with delayed feedback
Thomas J Walsh, Ali Nouri, Lihong Li, and Michael L Littman · 2009
Earlier work this paper cites.
Control delay in reinforcement learning for real-time dynamic systems: a memoryless approach
Erik Schuitema, Lucian Buşoniu, Robert Babuška, and Pieter Jonker · 2010
Earlier work this paper cites.
Reward design via online gradient ascent
Jonathan Sorg, Richard L Lewis, and Satinder Singh · 2010
Earlier work this paper cites.
An empirical study of potential-based reward shaping and advice in complex, multi-agent systems
Sam Devlin, Daniel Kudenko, and Marek Grześ · 2011
Earlier work this paper cites.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X Charles, D Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Earlier work this paper cites.
TEXPLORE: real-time sample-efficient reinforcement learning for robots
Todd Hester and Peter Stone · 2013
Earlier work this paper cites.
Optimal jamming using delayed learning
Saidhiraj Amuru and R Michael Buehrer · 2014
Earlier work this paper cites.
Reinforcement learning and the reward engineering principle
Daniel Dewey · 2014
Earlier work this paper cites.
The dependence of effective planning horizon on model accuracy
Nan Jiang, Alex Kulesza, Satinder Singh, and Richard Lewis · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Deep learning for reward design to improve Monte Carlo tree search in ATARI games
Xiaoxiao Guo, Satinder P Singh, Richard L Lewis, and Honglak Lee · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Hyperbolic discounting and learning over multiple horizons
William Fedus, Carles Gelada, Yoshua Bengio, Marc G Bellemare, and Hugo Larochelle · 2019
Later among the works it cites.
Learning self-imitating diverse policies
Tanmay Gangwani, Qiang Liu, and Jian Peng · 2019
Later among the works it cites.
Hindsight credit assignment
Anna Harutyunyan, Will Dabney, Thomas Mesnard, Mohammad Gheshlaghi Azar, Bilal Piot, Nicolas Heess, Hado P van Hasselt, Gregory Wayne, Satinder Singh, Doina Precup, et al · 2019
Later among the works it cites.
Sequence modeling of temporal credit assignment for episodic reinforcement learning
Yang Liu, Yunan Luo, Yuanyi Zhong, Xi Chen, Qiang Liu, and Jian Peng · 2019
Later among the works it cites.
QTRAN: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Resource management with deep reinforcement learning
Hongzi Mao, Mohammad Alizadeh, Ishai Menache, and Srikanth Kandula · 2016
Cited alongside, same era.
Model-free preference-based reinforcement learning
Christian Wirth, Johannes Fürnkranz, and Gerhard Neumann · 2016
Cited alongside, same era.
Hindsight experience replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
A benchmark environment motivated by industrial control problems
Daniel Hein, Stefan Depeweg, Michel Tokic, Steffen Udluft, Alexander Hentschel, Thomas A Runkler, and Volkmar Sterzing · 2017
Cited alongside, same era.
Training agent for first-person shooter game with actor-critic curriculum learning
Yuxin Wu and Yuandong Tian · 2017
Cited alongside, same era.
Playing FPS games with environment-aware hierarchical reinforcement learning
Shihong Song, Jiayi Weng, Hang Su, Dong Yan, Haosheng Zou, and Jun Zhu · 2019
Later among the works it cites.
Deep coordination graphs
Wendelin Böhmer, Vitaly Kurin, and Shimon Whiteson · 2020
Later among the works it cites.
Rna secondary structure prediction by learning unrolled algorithms
Xinshi Chen, Yu Li, Ramzan Umarov, Xin Gao, and Le Song · 2020
Later among the works it cites.
Learning guidance rewards with trajectory-space smoothing
Tanmay Gangwani, Yuan Zhou, and Jian Peng · 2020
Later among the works it cites.
Gradient-free online learning in continuous games with delayed rewards
Amélie Héliou, Panayotis Mertikopoulos, and Zhengyuan Zhou · 2020
Later among the works it cites.
Deep reinforcement learning for autonomous internet of things: Model, applications and challenges
Lei Lei, Yue Tan, Kan Zheng, Shiwen Liu, Kuan Zhang, and Xuemin Shen · 2020
Later among the works it cites.
Align-RUDDER: Learning from few demonstrations by reward redistribution
Vihang P Patil, Markus Hofmarcher, Marius-Constantin Dinu, Matthias Dorfer, Patrick M Blies, Johannes Brandstetter, Jose A Arjona-Medina, and Sepp Hochreiter · 2020
Later among the works it cites.
Shapley Q-value: a local reward approach to solve global reward games
Jianhong Wang, Yuan Zhang, Tae-Kyun Kim, and Yunjie Gu · 2020
Later among the works it cites.
What can learned intrinsic rewards capture?
Zeyu Zheng, Junhyuk Oh, Matteo Hessel, Zhongwen Xu, Manuel Kroiss, Hado Van Hasselt, David Silver, and Satinder Singh · 2020
Later among the works it cites.
On the expressivity of Markov reward
David Abel, Will Dabney, Anna Harutyunyan, Mark K Ho, Michael Littman, Doina Precup, and Satinder Singh · 2021
Closest in time.
Reinforcement learning with random delays
Yann Bouteiller, Simon Ramstedt, Giovanni Beltrame, Christopher Pal, and Jonathan Binas · 2021
Closest in time.
On the theory of reinforcement learning with once-per-episode feedback
Niladri Chatterji, Aldo Pacchiano, Peter Bartlett, and Michael Jordan · 2021
Closest in time.
Reinforcement learning with trajectory feedback
Yonathan Efroni, Nadav Merlis, and Shie Mannor · 2021
Closest in time.
Off-policy reinforcement learning with delayed rewards
Beining Han, Zhizhou Ren, Zuofan Wu, Yuan Zhou, and Jian Peng · 2021
Closest in time.
Convergence proof for actor-critic methods applied to PPO and RUDDER
Markus Holzleitner, Lukas Gruber, José Arjona-Medina, Johannes Brandstetter, and Sepp Hochreiter · 2021
Closest in time.
PEBBLE: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training
Kimin Lee, Laura Smith, and Pieter Abbeel · 2021
Closest in time.
Revisiting state augmentation methods for reinforcement learning with stochastic delays
Somjit Nath, Mayank Baranwal, and Harshad Khadilkar · 2021
Closest in time.
Synthetic returns for long-term credit assignment
David Raposo, Sam Ritter, Adam Santoro, Greg Wayne, Theophane Weber, Matt Botvinick, Hado van Hasselt, and Francis Song · 2021
Closest in time.
Bandit learning with delayed impact of actions
Wei Tang, Chien-Ju Ho, and Yang Liu · 2021
Closest in time.
Learning to represent action values as a hypergraph on the action vertices
Arash Tavakoli, Mehdi Fatemi, and Petar Kormushev · 2021
Closest in time.
Pairwise weights for temporal credit assignment
Zeyu Zheng, Risto Vuorio, Richard Lewis, and Satinder Singh · 2021
Closest in time.