Fetching the paper…
Reading the bibliography…
Reinforcement learning agents have demonstrated remarkable achievements in simulated environments.
“On the likelihood that one unknown probability exceeds another in view of the evidence of two samples”
William Thompson · 1933
Earlier work this paper cites.
“On the theory of apportionment”
William Thompson · 1935
Earlier work this paper cites.
“A definition of subjective probability”
Francis Anscombe and Robert Aumann · 1963
Earlier work this paper cites.
“Information value theory”
Ronald Howard · 1966
Earlier work this paper cites.
“A dynamic allocation index for the sequential design of experiments”
John Gittins · 1974
Earlier work this paper cites.
“Learning to Control”, 1976
Ian. Witten · 1976
Earlier work this paper cites.
“An adaptive optimal controller for discrete-time Markov environments”
Ian. Witten · 1977
Earlier work this paper cites.
“A dynamic allocation index for the discounted multiarmed bandit problem”
John Gittins and David Jones · 1979
Earlier work this paper cites.
“The Hedonistic Neuron: A Theory of Memory, Learning, and Intelligence”
A. Klopf · 1982
Earlier work this paper cites.
“Neuronlike adaptive elements that can solve difficult learning control problems”
Andrew Barto, Richard Sutton and Charles Anderson · 1983
Earlier work this paper cites.
“Temporal Credit Assignment in Reinforcement Learning”, 1984
Richard Sutton · 1984
Earlier work this paper cites.
“Asymptotically efficient adaptive allocation rules”
Tze Lai and Herbert Robbins · 1985
Earlier work this paper cites.
“Adaptive treatment allocation and the multi-armed bandit problem”
Tze Lai · 1987
Earlier work this paper cites.
“Learning to predict by the methods of temporal differences”
Richard Sutton · 1988
Earlier work this paper cites.
“Learning from delayed rewards”, 1989
Christopher Watkins · 1989
Earlier work this paper cites.
“Gain adaptation beats least squares”
Richard Sutton · 1992
Earlier work this paper cites.
“Practical issues in temporal difference learning”
Gerald Tesauro · 1992
Earlier work this paper cites.
“TD-Gammon, a self-teaching backgammon program, achieves master-level play”
Gerald Tesauro · 1994
Earlier work this paper cites.
“Instance-based utile distinctions for reinforcement learning with hidden state”
R McCallum · 1995
Earlier work this paper cites.
“Neuro-dynamic programming”
Dimitri Bertsekas and John Tsitsiklis · 1996
Earlier work this paper cites.
“Some upper bounds for relative entropy and applications”
S.S. Dragomir, M.L. Scholz and J. Sunde · 2000
Earlier work this paper cites.
“A Bayesian Framework for Reinforcement Learning”
Malcolm.. Strens · 2000
Earlier work this paper cites.
“Finite-time analysis of the multiarmed bandit problem”
Peter Auer, Nicolo Cesa-Bianchi and Paul Fischer · 2002
Earlier work this paper cites.
“Approximately Optimal Approximate Reinforcement Learning”
Sham Kakade and John Langford · 2002
Earlier work this paper cites.
“Near-Optimal Reinforcement Learning in Polynomial Time”
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
“Predictive Representations of State”
Michael Littman, Richard Sutton and Satinder Singh · 2002
Earlier work this paper cites.
“R-Max - a General Polynomial Time Algorithm for near-Optimal Reinforcement Learning”
Ronen. Brafman and Moshe Tennenholtz · 2003
Earlier work this paper cites.
“Optimal learning: Computational procedures for Bayes-adaptive Markov decision processes”, 2003
Michael. Duff · 2003
Earlier work this paper cites.
“Bayes meets Bellman: The Gaussian process approach to temporal difference learning”
Yaakov Engel, Shie Mannor and Ron Meir · 2003
Earlier work this paper cites.
“An adaptive sampling algorithm for solving Markov decision processes”
Hyeong Chang, Michael Fu, Jiaqiao Hu and Steven Marcus · 2005
Earlier work this paper cites.
“Reinforcement learning with Gaussian processes”
Yaakov Engel, Shie Mannor and Ron Meir · 2005
Earlier work this paper cites.
“Efficient selectivity and backup operators in Monte-Carlo tree search”
Rémi Coulom · 2006
Earlier work this paper cites.
“Elements of Information Theory”
Thomas Cover and Joy Thomas · 2006
Earlier work this paper cites.
“Bandit based Monte-Carlo planning”
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
“Universal Algorithmic Intelligence: A Mathematical Top → \rightarrow Down Approach”
Marcus Hutter · 2007
Cited alongside, same era.
“Near-optimal Regret Bounds for Reinforcement Learning”
Thomas Jaksch, Ronald Ortner and Peter Auer · 2010
Cited alongside, same era.
“Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction”
Richard Sutton et al · 2011
Cited alongside, same era.
“Bayesian Learning via Stochastic Gradient Langevin Dynamics”
M. Welling and Y.. Teh · 2011
Cited alongside, same era.
“Regret analysis of stochastic and nonstochastic multi-armed bandit problems”
Sébastien Bubeck and Nicolo Cesa-Bianchi · 2012
Cited alongside, same era.
“Optimal learning”
Warren Powell and Ilya Ryzhov · 2012
“Variational Bayesian reinforcement learning with regret bounds”
Brendan O’Donoghue · 2018
Later among the works it cites.
“The uncertainty Bellman equation and exploration”
Brendan O’Donoghue, Ian Osband, Remi Munos and Volodymyr Mnih · 2018
Later among the works it cites.
“Randomized prior functions for deep reinforcement learning”
Ian Osband, John Aslanides and Albin Cassirer · 2018
Later among the works it cites.
“Learning to optimize via information-directed sampling”
Daniel Russo and Benjamin Van · 2018
Later among the works it cites.
“A Tutorial on Thompson Sampling”
Daniel. Russo et al · 2018
Later among the works it cites.
“Reinforcement learning: An introduction”
Richard Sutton and Andrew Barto · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“The Knowledge Gradient Algorithm for a General Class of Online Learning Problems”
Ilya. Ryzhov, Warren. Powell and Peter. Frazier · 2012
Cited alongside, same era.
“Bayesian reinforcement learning”
Nikos Vlassis, Mohammad Ghavamzadeh, Shie Mannor and Pascal Poupart · 2012
Cited alongside, same era.
“Q-learning for history-based reinforcement learning”
Mayank Daswani, Peter Sunehag and Marcus Hutter · 2013
Cited alongside, same era.
“Playing Atari With Deep Reinforcement Learning”
Volodymyr Mnih et al · 2013
Cited alongside, same era.
“(More) Efficient Reinforcement Learning via Posterior Sampling”
Ian Osband, Daniel Russo and Benjamin Van · 2013
Cited alongside, same era.
“Feature reinforcement learning: state of the art”
Mayank Daswani, Peter Sunehag and Marcus Hutter · 2014
Cited alongside, same era.
“Reinforcement learning and optimal control”
Dimitri Bertsekas · 2019
Later among the works it cites.
“Large-Scale Study of Curiosity-Driven Learning”
Yuri Burda et al · 2019
Later among the works it cites.
“On the performance of Thompson sampling on logistic bandits”
Shi Dong, Tengyu Ma and Benjamin Van · 2019
Later among the works it cites.
“Quantifying information and uncertainty”
Alexander Frankel and Emir Kamenica · 2019
Later among the works it cites.
“An information-theoretic approach to minimax regret in partial monitoring”
Tor Lattimore and Csaba Szepesvári · 2019
Later among the works it cites.
“Information-Theoretic Confidence Bounds for Reinforcement Learning”
Xiuyuan Lu and Benjamin Van · 2019
Later among the works it cites.
“Information-Directed Exploration for Deep Reinforcement Learning”
Nikolay Nikolov, Johannes Kirschner, Felix Berkenkamp and Andreas Krause · 2019
Later among the works it cites.
“Deep Exploration via Randomized Value Functions”
Ian Osband, Benjamin Van, Daniel Russo and Zheng Wen · 2019
Later among the works it cites.
“Discovery of Useful Questions as Auxiliary Tasks”
Vivek Veeriah et al · 2019
Later among the works it cites.
“Connections between mirror descent, Thompson sampling and the information ratio”
Julian Zimmert and Tor Lattimore · 2019
Later among the works it cites.
“First-Order Bayesian Regret Analysis of Thompson Sampling”
Sébastien Bubeck and Mark Sellke · 2020
Later among the works it cites.
“Hypermodels for Exploration”
Vikranth Dwaracherla et al · 2020
Later among the works it cites.
“Temporal Difference Uncertainties as a Signal for Exploration”
Sebastian Flennerhag et al · 2020
Later among the works it cites.
“Provably efficient reinforcement learning with linear function approximation”
Chi Jin, Zhuoran Yang, Zhaoran Wang and Michael Jordan · 2020
Later among the works it cites.
“Information directed sampling for linear partial monitoring”
Johannes Kirschner, Tor Lattimore and Andreas Krause · 2020
Later among the works it cites.
“Asymptotically Optimal Information-Directed Sampling”
Johannes Kirschner, Tor Lattimore, Claire Vernade and Csaba Szepesvári · 2020
Later among the works it cites.
“Mirror Descent and the Information Ratio”
Tor Lattimore and András György · 2020
Later among the works it cites.
“Information-directed sampling for reinforcement learning”, 2020
Xiuyuan Lu · 2020
Later among the works it cites.
“Behaviour Suite for Reinforcement Learning”
Ian Osband et al · 2020
Later among the works it cites.
“Satisficing in Time-Sensitive Bandit Learning”, 2020
Daniel Russo and Benjamin Van · 2020
Later among the works it cites.
“Mastering Atari, go, chess and shogi by planning with a learned model”
Julian Schrittwieser et al · 2020
Later among the works it cites.
“Deciding What to Learn: A Rate-Distortion Approach”, 2021
Dilip Arumugam and Benjamin Van · 2021
Closest in time.
“A Bit Better? Quantifying Information for Bandit Learning”, 2021
Adithya Devraj, Kuang Xu and Benjamin Van · 2021
Closest in time.
“Langevin DQN”, 2021
Vikranth Dwaracherla and Benjamin Van · 2021
Closest in time.
“The Value Equivalence Principle for Model-Based Reinforcement Learning”
Christopher Grimm, Andre Barreto, Satinder Singh and David Silver · 2021
Closest in time.
“Information-Directed Sampling - Frequentist Analysis and Applications”
Johannes Kirschner · 2021
Closest in time.
“Neural networks with late-phase weights”
Johannes Oswald et al · 2021
Closest in time.
“A Note on Reinforcement Learning, Bit by Bit”, 2023
Chao Qin · 2023
Closest in time.