Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert E. Schapire · 2014
Later among the works it cites.
Reinforcement and imitation learning via interactive no-regret learning
Original
Stephane Ross and J Andrew Bagnell · 2014
Later among the works it cites.
Model-based reinforcement learning and the Eluder dimension
Ian Osband and Benjamin Van Roy · 2014
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Later among the works it cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Later among the works it cites.
Learning to search better than your teacher
Kai-Wei Chang, Akshay Krishnamurthy, Alekh Agarwal, Hal Daume III, and John Langford · 2015
Later among the works it cites.
Efficient PAC-optimal exploration in concurrent, continuous state MDPs with delayed updates
Jason Pazis and Ronald Parr · 2016
Later among the works it cites.
PAC reinforcement learning with rich observations
Akshay Krishnamurthy, Alekh Agarwal, and John Langford · 2016
Later among the works it cites.
The Malmo Platform for artificial intelligence experimentation
Matthew Johnson, Katja Hofmann, Tim Hutton, and David Bignell · 2016
Later among the works it cites.
A PAC RL algorithm for episodic POMDPs
Zhaohan Daniel Guo, Shayan Doroudi, and Emma Brunskill · 2016
Later among the works it cites.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E. Schapire · 2017
Later among the works it cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Later among the works it cites.
Unifying PAC and regret: Uniform PAC bounds for episodic reinforcement learning
Christoph Dann, Tor Lattimore, and Emma Brunskill · 2017
Later among the works it cites.
Efficient reinforcement learning in deterministic systems with value function generalization
Zheng Wen and Benjamin Van Roy · 2017
Later among the works it cites.
Sbeed: Convergent reinforcement learning with nonlinear function approximation
Bo Dai, Albert Shaw, Lihong Li, Lin Xiao, Niao He, Zhen Liu, Jianshu Chen, and Le Song · 2018
Closest in time.