Reducing reinforcement learning to KWIK online regression
Lihong Li and Michael L Littman · 2010
Later among the works it cites.
Efficient optimal learning for contextual bandits
Miroslav Dudik, Daniel Hsu, Satyen Kale, Nikos Karampatziakis, John Langford, Lev Reyzin, and Tong Zhang · 2011
Later among the works it cites.
Contextual bandit learning with predictable rewards
Alekh Agarwal, Miroslav Dudík, Satyen Kale, John Langford, and Robert E Schapire · 2012
Later among the works it cites.
Competing with an infinite set of models in reinforcement learning
Phuong Nguyen, Odalric-Ambrym Maillard, Daniil Ryabko, and Ronald Ortner · 2013
Later among the works it cites.
Sample complexity of episodic fixed-horizon reinforcement learning
Christoph Dann and Emma Brunskill · 2015
Later among the works it cites.
Abstraction selection in model-based reinforcement learning
Nan Jiang, Alex Kulesza, and Satinder Singh · 2015
Later among the works it cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Beattie Charles, Sadik Amir, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Later among the works it cites.
Reinforcement learning of POMDPs using spectral methods
Kamyar Azizzadenesheli, Alessandro Lazaric, and Animashree Anandkumar · 2016
Closest in time.
Efficient PAC-optimal exploration in concurrent, continuous state MDPs with delayed updates
Jason Pazis and Ronald Parr · 2016
Closest in time.