Fetching the paper…
Reading the bibliography…
Bandit and reinforcement learning (RL) problems can often be framed as optimization problems where the goal is to maximize average performance while having access only to stochastic estimates of the true gradient.
Estimation of particle transmission by random sampling
Herman Kahn and Theodore E Harris · 1951
Earlier work this paper cites.
Some aspects of the sequential design of experiments
Herbert Robbins · 1952
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Sample mean based index policies with o (log n) regret for the multi-armed bandit problem
Rajeev Agrawal · 1995
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
The nonstochastic multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert E Schapire · 2002
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter L Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
A tutorial on the cross-entropy method
Pieter-Tjerk De Boer, Dirk P Kroese, Shie Mannor, and Reuven Y Rubinstein · 2005
Earlier work this paper cites.
Monte Carlo strategies in scientific computing
Jun S Liu · 2008
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
Jan Peters and Stefan Schaal · 2008
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Earlier work this paper cites.
Evaluation and analysis of the performance of the exp3 algorithm in stochastic environments
Yevgeny Seldin, Csaba Szepesvári, Peter Auer, and Yasin Abbasi-Yadkori · 2012
Earlier work this paper cites.
Evaluation and analysis of the performance of the exp3 algorithm in stochastic environments
Yevgeny Seldin, Csaba Szepesvári, Peter Auer, and Yasin Abbasi-Yadkori · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Cited alongside, same era.
python-ternary: Ternary plots in python
Marc Harper and Bryan Weinstein · 2015
Cited alongside, same era.
Variance reduced stochastic gradient descent with neighbors
Thomas Hofmann, Aurelien Lucchi, Simon Lacoste-Julien, and Brian McWilliams · 2015
Cited alongside, same era.
Explore no more: Improved high-probability regret bounds for non-stochastic bandits
Gergely Neu · 2015
Cited alongside, same era.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Shixiang Gu, Timothy Lillicrap, Zoubin Ghahramani, Richard E Turner, and Sergey Levine · 2016
Cited alongside, same era.
Deep reinforcement learning and the deadly triad
Hado van Hasselt, Yotam Doron, Florian Strub, Matteo Hessel, Nicolas Sonnerat, and Joseph Modayil · 2018
Later among the works it cites.
Variance reduction for policy gradient with action-dependent factorized baselines
Cathy Wu, Aravind Rajeswaran, Yan Duan, Vikash Kumar, Alexandre M Bayen, Sham Kakade, Igor Mordatch, and Pieter Abbeel · 2018
Later among the works it cites.
Optimality and approximation with policy gradient methods in markov decision processes
Alekh Agarwal, Sham M Kakade, Jason D Lee, and Gaurav Mahajan · 2019
Later among the works it cites.
Provably efficient exploration in policy optimization
Qi Cai, Zhuoran Yang, Chi Jin, and Zhaoran Wang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc G. Bellemare · 2016
Cited alongside, same era.
Backpropagation through the void: Optimizing control variates for black-box gradient estimation
Will Grathwohl, Dami Choi, Yuhuai Wu, Geoffrey Roeder, and David Duvenaud · 2017
Cited alongside, same era.
Data-efficient policy evaluation through behavior policy search
Josiah P Hanna, Philip S Thomas, Peter Stone, and Scott Niekum · 2017
Cited alongside, same era.
Action-depedent control variates for policy optimization via stein’s identity
Hao Liu, Yihao Feng, Yi Mao, Dengyong Zhou, Jian Peng, and Qiang Liu · 2017
Cited alongside, same era.
A tutorial on thompson sampling
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
Importance sampling policy evaluation with an estimated behavior policy
Josiah P Hanna, Scott Niekum, and Peter Stone · 2018
Cited alongside, same era.
Paavo Parmas and Masashi Sugiyama · 2019
Later among the works it cites.
Ray interference: a source of plateaus in deep reinforcement learning
Tom Schaul, Diana Borsa, Joseph Modayil, and Razvan Pascanu · 2019
Later among the works it cites.
Pc-pg: Policy cover directed exploration for provable policy gradient learning
Alekh Agarwal, Mikael Henaff, Sham Kakade, and Wen Sun · 2020
Closest in time.
Trajectory-wise control variates for variance reduction in policy gradient methods
Ching-An Cheng, Xinyan Yan, and Byron Boots · 2020
Closest in time.
Optimistic policy optimization with bandit feedback
Yonathan Efroni, Lior Shani, Aviv Rosenberg, and Shie Mannor · 2020
Closest in time.
Neural replicator dynamics: Multiagent learning via hedging policy gradients
Daniel Hennes, Dustin Morrill, Shayegan Omidshafiei, Rémi Munos, Julien Perolat, Marc Lanctot, Audrunas Gruslys, Jean-Baptiste Lespiau, Paavo Parmas, Edgar Duéñez-Guzmán, et al · 2020
Closest in time.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Closest in time.
On the global convergence rates of softmax policy gradient methods
Jincheng Mei, Chenjun Xiao, Csaba Szepesvari, and Dale Schuurmans · 2020
Closest in time.