Fetching the paper…
Reading the bibliography…
A softmax operator applied to a set of values acts somewhat like the maximization function and somewhat like an average.
A Markovian decision process
Bellman, Richard · 1957
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Sutton, Richard S · 1990
Earlier work this paper cites.
The role of exploration in learning control
Thrun, Sebastian B · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 1992
Earlier work this paper cites.
Statistical analysis of the power sum of multiple correlated log-normal components
Safak, Aysel · 1993
Earlier work this paper cites.
When the best move isn’t optimal: Q-learning with exploration
John, George H · 1994
Earlier work this paper cites.
Markov Decision Processes—Discrete Stochastic Dynamic Programming
Puterman, Martin L · 1994
Earlier work this paper cites.
On-line Q-learning using connectionist systems
Rummery, G. A. and Niranjan, M · 1994
Earlier work this paper cites.
Experimental evidence on players’ models of other players
Stahl, Dale O. and Wilson, Paul W · 1994
Earlier work this paper cites.
Stable function approximation in dynamic programming
Gordon, Geoffrey J · 1995
Earlier work this paper cites.
A generalized reinforcement-learning model: Convergence and applications
Littman, Michael L. and Szepesvári, Csaba · 1996
Earlier work this paper cites.
Algorithms for Sequential Decision Making
Littman, Michael Lederman · 1996
Earlier work this paper cites.
Bayesian Q-learning
Dearden, Richard, Friedman, Nir, and Russell, Stuart · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Sutton, Richard S. and Barto, Andrew G · 1998
Cited alongside, same era.
Gradient descent for general reinforcement learning
Baird, Leemon and Moore, Andrew W · 1999
Cited alongside, same era.
Algorithms for inverse reinforcement learning
Ng, Andrew Y. and Russell, Stuart · 2000
Cited alongside, same era.
Convergence results for single-step on-policy reinforcement-learning algorithms
Singh, Satinder, Jaakkola, Tommi, Littman, Michael L., and Szepesvári, Csaba · 2000
Cited alongside, same era.
Reinforcement learning with function approximation converges to a region, 2001
Gordon, Geoffrey J · 2001
Cited alongside, same era.
A convergent form of approximate policy iteration
Perkins, Theodore J and Precup, Doina · 2002
Cited alongside, same era.
A theoretical and empirical analysis of Expected Sarsa
Van Seijen, Harm, Van Hasselt, Hado, Whiteson, Shimon, and Wiering, Marco · 2009
Later among the works it cites.
Relative entropy policy search
Peters, Jan, Mülling, Katharina, and Altun, Yasemin · 2010
Later among the works it cites.
Beyond equilibrium: Predicting human behavior in normal-form games
Wright, James R. and Leyton-Brown, Kevin · 2010
Later among the works it cites.
Apprenticeship learning about multiple intentions
Babes, Monica, Marivate, Vukosi N., Littman, Michael L., and Subramanian, Kaushik · 2011
Later among the works it cites.
Trading value and information in mdps
Rubin, Jonathan, Shamir, Ohad, and Tishby, Naftali · 2012
Later among the works it cites.
Algorithms for minimization without derivatives
Brent, Richard P · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convex optimization
Boyd, S.P. and Vandenberghe, L · 2004
Cited alongside, same era.
Elements of Information Theory
Cover, T.M. and Thomas, J.A · 2006
Cited alongside, same era.
Linearly-solvable markov decision problems
Todorov, Emanuel · 2006
Cited alongside, same era.
Goal inference as inverse planning
Baker, Chris L, Tenenbaum, Joshua B, and Saxe, Rebecca R · 2007
Cited alongside, same era.
Apprenticeship learning using inverse reinforcement learning and gradient methods
Neu, Gergely and Szepesvári, Csaba · 2007
Cited alongside, same era.
Bayesian inverse reinforcement learning
Ramachandran, Deepak and Amir, Eyal · 2007
Cited alongside, same era.
Kingma, Diederik and Ba, Jimmy · 2014
Later among the works it cites.
Algorithms for multi-armed bandit problems
Kuleshov, Volodymyr and Precup, Doina · 2014
Later among the works it cites.
A Practical Guide to Averaging Functions
Beliakov, Gleb, Sola, Humberto Bustince, and Sánchez, Tomasa Calvo · 2016
Closest in time.
Openai gym, 2016
Brockman, Greg, Cheung, Vicki, Pettersson, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech · 2016
Closest in time.
Taming the noise in reinforcement learning via soft updates
Fox, Roy, Pakman, Ari, and Tishby, Naftali · 2016
Closest in time.
Theano: A python framework for fast computation of mathematical expressions
Team, The Theano Development, Al-Rfou, Rami, Alain, Guillaume, Almahairi, Amjad, Angermueller, Christof, Bahdanau, Dzmitry, Ballas, Nicolas, Bastien, Frédéric, Bayer, Justin, Belikov, Anatoly, et al · 2016
Closest in time.