Fetching the paper…
Reading the bibliography…
We study a theory of reinforcement learning (RL) in which the learner receives binary feedback only once at the end of an episode.
“Simple statistical gradient-following algorithms for connectionist reinforcement learning”
Ronald Williams · 1992
Earlier work this paper cites.
“Algorithms for inverse reinforcement learning”
Andrew Ng and Stuart Russell · 2000
Earlier work this paper cites.
“Introduction to algorithms”
Thomas Cormen, Charles Leiserson, Ronald Rivest and Clifford Stein · 2009
Earlier work this paper cites.
“Parametric bandits: The generalized linear case”
Sarah Filippi, Olivier Cappe, Aurélien Garivier and Csaba Szepesvári · 2010
Earlier work this paper cites.
“Near-optimal regret bounds for reinforcement learning.”
Thomas Jaksch, Ronald Ortner and Peter Auer · 2010
Earlier work this paper cites.
“Improved algorithms for linear stochastic bandits”
Yasin Abbasi-Yadkori, Dávid Pál and Csaba Szepesvári · 2011
Earlier work this paper cites.
“Preference-based policy learning”
Riad Akrour, Marc Schoenauer and Michele Sebag · 2011
Earlier work this paper cites.
“Freedman’s inequality for matrix martingales”
Joel Tropp · 2011
Earlier work this paper cites.
“Preference-based reinforcement learning: a formal framework and a policy iteration algorithm”
Johannes Fürnkranz, Eyke Hüllermeier, Weiwei Cheng and Sang-Hyeun Park · 2012
Earlier work this paper cites.
“(More) Efficient Reinforcement Learning via Posterior Sampling”
Ian Osband, Daniel Russo and Benjamin Van · 2013
Earlier work this paper cites.
“Programming by feedback”
Riad Akrour, Marc Schoenauer, Michèle Sebag and Jean-Christophe Souplet · 2014
Earlier work this paper cites.
“Thompson sampling for learning parameterized Markov decision processes”
Aditya Gopalan and Shie Mannor · 2015
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“Human-level control through deep reinforcement learning”
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei Rusu, Joel Veness, Marc Bellemare, Alex Graves, Martin Riedmiller, Andreas Fidjeland and Georg Ostrovski · 2015
Earlier work this paper cites.
“End-to-end training of deep visuomotor policies”
Sergey Levine, Chelsea Finn, Trevor Darrell and Pieter Abbeel · 2016
Cited alongside, same era.
“On lower bounds for regret in reinforcement learning”
Ian Osband and Benjamin Van · 2016
Cited alongside, same era.
“Minimax regret bounds for reinforcement learning”
Mohammad Azar, Ian Osband and Rémi Munos · 2017
Cited alongside, same era.
“Deep Reinforcement Learning from Human Preferences”
Paul Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg and Dario Amodei · 2017
Cited alongside, same era.
“Unifying PAC and regret: uniform PAC bounds for episodic reinforcement learning”
Christoph Dann, Tor Lattimore and Emma Brunskill · 2017
Cited alongside, same era.
“Why is posterior sampling better than optimism for reinforcement learning?”
“Improved optimistic algorithms for logistic bandits”
Louis Faury, Marc Abeille, Clément Calauzènes and Olivier Fercoq · 2020
Later among the works it cites.
“Reward-free exploration for reinforcement learning”
Chi Jin, Akshay Krishnamurthy, Max Simchowitz and Tiancheng Yu · 2020
Later among the works it cites.
“Bandit algorithms”
Tor Lattimore and Csaba Szepesvári · 2020
Later among the works it cites.
“Dueling posterior sampling for preference-based reinforcement learning”
Ellen Novoseller, Yibing Wei, Yanan Sui, Yisong Yue and Joel Burdick · 2020
Later among the works it cites.
“On optimism in model-based reinforcement learning”
Aldo Pacchiano, Philip Ball, Jack Parker-Holder, Krzysztof Choromanski and Stephen Roberts · 2020
Later among the works it cites.
“Improved protein structure prediction using potentials from deep learning”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ian Osband and Benjamin Van · 2017
Cited alongside, same era.
“Mastering the game of Go without human knowledge”
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai and Adrian Bolton · 2017
Cited alongside, same era.
“A survey of preference-based reinforcement learning methods”
Christian Wirth, Riad Akrour, Gerhard Neumann and Johannes Fürnkranz · 2017
Cited alongside, same era.
“Is Q Q -learning provably efficient?”
Chi Jin, Zeyuan Allen-Zhu, Sebastien Bubeck and Michael Jordan · 2018
Cited alongside, same era.
“On the performance of Thompson sampling on logistic bandits”
Shi Dong, Tengyu Ma and Benjamin Van · 2019
Cited alongside, same era.
“Tight Regret Bounds for Model-Based Reinforcement Learning with Greedy Policies”
Yonathan Efroni, Nadav Merlis, Mohammad Ghavamzadeh and Shie Mannor · 2019
Cited alongside, same era.
“Non-asymptotic gap-dependent regret bounds for tabular MDPs”
Max Simchowitz and Kevin Jamieson · 2019
Cited alongside, same era.
Andrew Senior, Richard Evans, John Jumper, James Kirkpatrick, Laurent Sifre, Tim Green, Chongli Qin, Augustin Žídek, Alexander Nelson and Alex Bridgland · 2020
Later among the works it cites.
“Preference-based reinforcement learning with finite-time guarantees”
Yichong Xu, Ruosong Wang, Lin Yang, Aarti Singh and Artur Dubrawski · 2020
Later among the works it cites.
“Provably efficient reward-agnostic navigation with linear value iteration”
Andrea Zanette, Alessandro Lazaric, Mykel Kochenderfer and Emma Brunskill · 2020
Later among the works it cites.
“Online Markov Decision Processes with Aggregate Bandit Feedback”
Alon Cohen, Haim Kaplan, Tomer Koren and Yishay Mansour · 2021
Closest in time.
“Reinforcement learning with trajectory feedback”
Yonathan Efroni, Nadav Merlis and Shie Mannor · 2021
Closest in time.
“Confidence-Budget Matching for Sequential Budgeted Learning”
Yonathan Efroni, Nadav Merlis, Aadirupa Saha and Shie Mannor · 2021
Closest in time.
“Time-uniform, nonparametric, nonasymptotic confidence sequences”
Steven Howard, Aaditya Ramdas, Jon McAuliffe and Jasjeet Sekhon · 2021
Closest in time.
“Self-concordant analysis of generalized linear bandits with forgetting”
Yoan Russac, Louis Faury, Olivier Cappé and Aurélien Garivier · 2021
Closest in time.