Fetching the paper…
Reading the bibliography…
In many finite horizon episodic reinforcement learning (RL) settings, it is desirable to optimize for the undiscounted return - in settings like Atari, for instance, the goal is to collect the most points while staying alive in the long run.
Hyperbolic discounting and learning over multiple horizons
Fedus, W., Fatemi, M., Bengio, Y., Bellemare, M. G., and Larochelle, H · 1902
Earlier work this paper cites.
A markovian decision process
Bellman, R · 1957
Earlier work this paper cites.
Temporal credit assignment in reinforcement learning
Sutton, R. S · 1984
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Sutton, R. S · 1988
Earlier work this paper cites.
Markov decision models with weighted discounted criteria
Feinberg, E. A. and Shwartz, A · 1994
Earlier work this paper cites.
Neuro-dynamic programming: an overview
Bertsekas, D. P. and Tsitsiklis, J. N · 1995
Earlier work this paper cites.
Td models: Modeling the world at a mixture of time scales
Sutton, R. S · 1995
Earlier work this paper cites.
Adaptive critic designs
Prokhorov, D. V. and Wunsch, D. C · 1997
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Decision boundary partitioning: Variable resolution model-free reinforcement learning
Reynolds, S. I · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Dietterich, T. G · 2000
Earlier work this paper cites.
Bias-variance error bounds for temporal difference updates
Kearns, M. J. and Singh, S. P · 2000
Earlier work this paper cites.
Actor-critic algorithms
Konda, V. R. and Tsitsiklis, J. N · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Sutton, R. S., McAllester, D. A., Singh, S. P., and Mansour, Y · 2000
Cited alongside, same era.
Discovering hierarchy in reinforcement learning with hexq
Hengst, B · 2002
Cited alongside, same era.
Near-optimal reinforcement learning in polynomial time
Kearns, M. and Singh, S · 2002
Cited alongside, same era.
Q-cut—dynamic discovery of sub-goals in reinforcement learning
Menache, I., Mannor, S., and Shimkin, N · 2002
Cited alongside, same era.
Learning rates for q-learning
Even-Dar, E. and Mansour, Y · 2003
Cited alongside, same era.
Q-decomposition for reinforcement learning agents
Russell, S. J. and Zimdars, A · 2003
Cited alongside, same era.
Unifying count-based exploration and intrinsic motivation
Bellemare, M. G., Srinivasan, S., Ostrovski, G., Schaul, T., Saxton, D., and Munos, R · 2016
Later among the works it cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Later among the works it cites.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Later among the works it cites.
Average reward optimization with multiple discounting reinforcement learners
Reinke, C., Uchibe, E., and Doya, K · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Solving large mdps quickly with partitioned value iteration
Wingate, D · 2004
Cited alongside, same era.
P3vi: A partitioned, prioritized, parallel value iterator
Wingate, D. and Seppi, K. D · 2004
Cited alongside, same era.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D · 2011
Cited alongside, same era.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Cited alongside, same era.
How to discount deep reinforcement learning: Towards new dynamic strategies
François-Lavet, V., Fonteneau, R., and Ernst, D · 2015
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2015
Cited alongside, same era.
van Seijen, H., Fatemi, M., Romoff, J., Laroche, R., Barnes, T., and Tsang, J · 2017
Later among the works it cites.
How many random seeds? statistical power analysis in deep reinforcement learning experiments
Colas, C., Sigaud, O., and Oudeyer, P.-Y · 2018
Later among the works it cites.
Pytorch implementations of reinforcement learning algorithms
Kostrikov, I · 2018
Later among the works it cites.
Openai five
OpenAI · 2018
Later among the works it cites.
The machine learning reproducibility checklist (version 1.0), 2018
Pineau, J · 2018
Later among the works it cites.
Generalizing value estimation over timescale
Sherstan, C., MacGlashan, J., and Pilarski, P. M · 2018
Later among the works it cites.
Meta-gradient reinforcement learning
Xu, Z., van Hasselt, H., and Silver, D · 2018
Later among the works it cites.