Fetching the paper…
Reading the bibliography…
Most bandit policies are designed to either minimize regret in any problem instance, making very few assumptions about the underlying environment, or in a Bayesian sense, assuming a prior distribution over environment parameters.
Empirical Bayes regret minimization
Chih-Wei Hsu, Branislav Kveton, Ofer Meshi, Martin Mladenov, and Csaba Szepesvari · 1904
Earlier work this paper cites.
Meta-learning of sequential strategies
Pedro Ortega, Jane Wang, Mark Rowland, Tim Genewein, Zeb Kurth-Nelson, Razvan Pascanu, Nicolas Heess, Joel Veness, Alexander Pritzel, Pablo Sprechmann, Siddhant Jayakumar, Tom McGrath, Kevin Miller, Mohammad Gheshlaghi Azar, Ian Osband, Neil Rabinowitz, Andras Gyorgy, Silvia Chiappa, Simon Osindero, Yee Whye Teh, Hado van Hasselt, Nando de Freitas, Matthew Botvinick, and Shane Legg · 1905
Earlier work this paper cites.
Global optimality guarantees for policy gradient methods
Jalaj Bhandari and Daniel Russo · 1906
Earlier work this paper cites.
Optimality and approximation with policy gradient methods in Markov decision processes
Alekh Agarwal, Sham Kakade, Jason Lee, and Gaurav Mahajan · 1908
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R. Thompson · 1933
Earlier work this paper cites.
Stochastic optimization
V. M. Aleksandrov, V. I. Sysoyev, and V. V. Shemeneva · 1968
Earlier work this paper cites.
Bandit processes and dynamic allocation indices
John Gittins · 1979
Earlier work this paper cites.
Bandit Problems: Sequential Allocation of Experiments
Donald Berry and Bert Fristedt · 1985
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
T. L. Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Learning representations by back-propagating errors
David Rumelhart, Geoffrey Hinton, and Ronald Williams · 1986
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
Richard Sutton · 1988
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald Williams · 1992
Earlier work this paper cites.
Gambling in a rigged casino: The adversarial multi-armed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, Yoav Freund, and Robert Schapire · 1995
Earlier work this paper cites.
Explanation-Based Neural Network Learning - A Lifelong Learning Approach
Sebastian Thrun · 1996
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jurgen Schmidhuber · 1997
Earlier work this paper cites.
Theoretical models of learning to learn
Jonathan Baxter · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard Sutton and Andrew Barto · 1998
Earlier work this paper cites.
Lifelong learning algorithms
Sebastian Thrun · 1998
Earlier work this paper cites.
A model of inductive bias learning
Jonathan Baxter · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
Jonathan Baxter and Peter Bartlett · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Provable meta-learning of linear representations
Nilesh Tripuraneni, Chi Jin, and Michael Jordan · 2002
Earlier work this paper cites.
Convex Optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Variance reduction techniques for gradient estimates in reinforcement learning
Evan Greensmith, Peter Bartlett, and Jonathan Baxter · 2004
Earlier work this paper cites.
Multi-armed bandit algorithms and empirical evaluation
Joannes Vermorel and Mehryar Mohri · 2005
Earlier work this paper cites.
Policy gradient optimization of Thompson sampling policies
Seungki Min, Ciamac Moallemi, and Daniel Russo · 2006
Cited alongside, same era.
Geometric variance reduction in Markov chains: Application to value function and gradient estimation
Remi Munos · 2006
Cited alongside, same era.
Differentiable linear bandit algorithm
Kaige Yang and Laura Toni · 2006
Cited alongside, same era.
UCI machine learning repository, 2007
A. Asuncion and D. J. Newman · 2007
Cited alongside, same era.
Stochastic linear optimization under bandit feedback
Varsha Dani, Thomas Hayes, and Sham Kakade · 2008
Cited alongside, same era.
The epoch-greedy algorithm for multi-armed bandits with side information
Learning to optimize via posterior sampling
Daniel Russo and Benjamin Van Roy · 2014
Later among the works it cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Later among the works it cites.
Policy gradient reinforcement learning without regret
Travis Dick · 2015
Later among the works it cites.
Efficient learning in large-scale combinatorial semi-bandits
Zheng Wen, Branislav Kveton, and Azin Ashkan · 2015
Later among the works it cites.
RL 2 : Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Later among the works it cites.
An information-theoretic analysis of Thompson sampling
Daniel Russo and Benjamin Van Roy · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Langford and Tong Zhang · 2008
Cited alongside, same era.
Learning diverse rankings with multi-armed bandits
Filip Radlinski, Robert Kleinberg, and Thorsten Joachims · 2008
Cited alongside, same era.
Exploration-exploitation tradeoff using variance estimates in multi-armed bandits
Jean-Yves Audibert, Remi Munos, and Csaba Szepesvari · 2009
Cited alongside, same era.
Curriculum learning
Yoshua Bengio, Jerome Louradour, Ronan Collobert, and Jason Weston · 2009
Cited alongside, same era.
UCB revisited: Improved regret bounds for the stochastic multi-armed bandit problem
Peter Auer and Ronald Ortner · 2010
Cited alongside, same era.
Parametric bandits: The generalized linear case
Sarah Filippi, Olivier Cappe, Aurelien Garivier, and Csaba Szepesvari · 2010
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert Schapire · 2010
Cited alongside, same era.
Later among the works it cites.
Boltzmann exploration done right
Nicolo Cesa-Bianchi, Claudio Gentile, Gabor Lugosi, and Gergely Neu · 2017
Later among the works it cites.
Multi-task learning for contextual bandits
Aniket Anand Deshmukh, Urun Dogan, and Clayton Scott · 2017
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, and Demis Hassabis · 2017
Later among the works it cites.
Dopamine
Marc Bellemare, Pablo Castro, Carles Gelada, Saurabh Kumar, and Subhodeep Moitra · 2018
Later among the works it cites.
Probabilistic model-agnostic meta-learning
Chelsea Finn, Kelvin Xu, and Sergey Levine · 2018
Later among the works it cites.
TopRank: A practical algorithm for online stochastic ranking
Tor Lattimore, Branislav Kveton, Shuai Li, and Csaba Szepesvari · 2018
Later among the works it cites.
Action-dependent control variates for policy optimization via Stein’s identity
Hao Liu, Yihao Feng, Yi Mao, Dengyong Zhou, Jian Peng, and Qiang Liu · 2018
Later among the works it cites.
A simple neural attentive meta-learner
Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel · 2018
Later among the works it cites.
Deep Bayesian bandits showdown: An empirical comparison of Bayesian deep networks for Thompson sampling
Carlos Riquelme, George Tucker, and Jasper Snoek · 2018
Later among the works it cites.
A tutorial on Thompson sampling
Daniel Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, and Zheng Wen · 2018
Later among the works it cites.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
Generalize across tasks: Efficient algorithms for linear representation learning
Brian Bullins, Elad Hazan, Adam Kalai, and Roi Livni · 2019
Later among the works it cites.
Bandit Algorithms
Tor Lattimore and Csaba Szepesvari · 2019
Later among the works it cites.
Differentiable meta-learning of bandit policies
Craig Boutilier, Chih-Wei Hsu, Branislav Kveton, Martin Mladenov, Csaba Szepesvari, and Manzil Zaheer · 2020
Closest in time.
Meta-learning with stochastic linear bandits
Leonardo Cella, Alessandro Lazaric, and Massimiliano Pontil · 2020
Closest in time.
Structure adaptive algorithms for stochastic bandits
Remy Degenne, Han Shao, and Wouter Koolen · 2020
Closest in time.
On the global convergence rates of softmax policy gradient methods
Jincheng Mei, Chenjun Xiao, Csaba Szepesvari, and Dale Schuurmans · 2020
Closest in time.
TensorFlow
tf · 2020
Closest in time.
A novel confidence-based algorithm for structured bandits
Andrea Tirinzoni, Alessandro Lazaric, and Marcello Restelli · 2020
Closest in time.
Graphical models meet bandits: A variational Thompson sampling approach
Tong Yu, Branislav Kveton, Zheng Wen, Ruiyi Zhang, and Ole Mengshoel · 2020
Closest in time.