Fetching the paper…
Reading the bibliography…
This paper addresses the problem of policy selection in domains with abundant logged data, but with a restricted interaction budget.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup, Richard S. Sutton, and Satinder P. Singh · 2000
Earlier work this paper cites.
Gaussian Processes for Machine Learning (Adaptive Computation and Machine Learning)
Carl Edward Rasmussen and Christopher K. I. Williams · 2005
Earlier work this paper cites.
Active preference learning with discrete choice data
E. Brochu, N. de Freitas, and A. Ghosh · 2007
Earlier work this paper cites.
Automatic gait optimization with gaussian process regression
Daniel J Lizotte, Tao Wang, Michael H Bowling, and Dale Schuurmans · 2007
Earlier work this paper cites.
Active reinforcement learning
Arkady Epshteyn, Adam Vogel, and Gerald DeJong · 2008
Earlier work this paper cites.
A Bayesian exploration-exploitation approach for optimal online sensing and planning with a visually guided mobile robot
R. Martinez–Cantin, N. de Freitas, E. Brochu, J. Castellanos, and A. Doucet · 2009
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham M Kakade, and Matthias Seeger · 2009
Earlier work this paper cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
Richard S Sutton, Hamid Reza Maei, Doina Precup, Shalabh Bhatnagar, David Silver, Csaba Szepesvári, and Eric Wiewiora · 2009
Earlier work this paper cites.
Efficient exploration through active learning for value function approximation in reinforcement learning
Takayuki Akiyama, Hirotaka Hachiya, and Masashi Sugiyama · 2010
Earlier work this paper cites.
Best arm identification in multi-armed bandits
Jean-Yves Audibert, Sébastien Bubeck, and Rémi Munos · 2010
Earlier work this paper cites.
A Bayesian interactive optimization approach to procedural animation design
E. Brochu, T. Brochu, and N. de Freitas · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Fitted policy search: Direct policy search using a batch reinforcement learning approach
Martino Migliavacca, Alessio Pecorino, Matteo Pirotta, Marcello Restelli, and Andrea Bonarini · 2010
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham Kakade, and Matthias Seeger · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
Sequential model-based optimization for general algorithm configuration
F. Hutter, H. H. Hoos, and K. Leyton-Brown · 2011
Earlier work this paper cites.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang · 2011
Earlier work this paper cites.
Hybrid batch bayesian optimization
Javad Azimi, Ali Jalali, and Xiaoli Zhang Fern · 2012
Earlier work this paper cites.
On Bayesian upper confidence bounds for bandit problems
Emilie Kaufmann, Olivier Cappé, and Aurélien Garivier · 2012
Earlier work this paper cites.
Contextual gaussian process bandit optimization
Andreas Krause and Cheng S Ong · 2012
Earlier work this paper cites.
Practical Bayesian optimization of machine learning algorithms
J. Snoek, H. Larochelle, and R. P. Adams · 2012
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Active reward learning
Christian Daniel, Malte Viering, Jan Metz, Oliver Kroemer, and Jan Peters · 2014
Earlier work this paper cites.
On correlation and budget constraints in model-based bandit optimization with application to automatic machine learning
Matthew Hoffman, Bobak Shahriari, and Nando de Freitas · 2014
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al · 2015
Cited alongside, same era.
Taking the human out of the loop: A review of Bayesian optimization
Bobak Shahriari, Kevin Swersky, Ziyu Wang, Ryan P Adams, and Nando De Freitas · 2015
Cited alongside, same era.
Safe exploration for optimization with gaussian processes
Yanan Sui, Alkis Gotovos, Joel Burdick, and Andreas Krause · 2015
Cited alongside, same era.
Personalized ad recommendation systems for life-time value optimization with guarantees
Georgios Theocharous, Philip S Thomas, and Mohammad Ghavamzadeh · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Scaling data-driven robotics with reward sketching and batch reinforcement learning
Serkan Cabi, Sergio Gómez Colmenarejo, Alexander Novikov, Ksenia Konyushkova, Scott Reed, Rae Jeong, Konrad Zolna, Yusuf Aytar, David Budden, Mel Vecerik, et al · 2020
Later among the works it cites.
D4RL: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
RL Unplugged: A suite of benchmarks for offline reinforcement learning
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Tom Le Paine, Sergio Gomez Colmenarejo, Konrad Zolna, Rishabh Agarwal, Josh Merel, Daniel Mankowitz, Cosmin Paduraru, et al · 2020
Later among the works it cites.
Jean Harb, Tom Schaul, Doina Precup, and Pierre-Luc Bacon · 2020
Later among the works it cites.
Acme: A research framework for distributed reinforcement learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Remi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
An emphatic approach to the problem of off-policy temporal-difference learning
Richard S Sutton, A Rupam Mahmood, and Martha White · 2016
Cited alongside, same era.
Deep reinforcement learning with double Q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
Distributed distributional deterministic policy gradients
Gabriel Barth-Maron, Matthew W Hoffman, David Budden, Will Dabney, Dan Horgan, TB Dhruva, Alistair Muldal, Nicolas Heess, and Timothy Lillicrap · 2018
Cited alongside, same era.
Bayesian optimization in AlphaGo
Yutian Chen, Aja Huang, Ziyu Wang, Ioannis Antonoglou, Julian Schrittwieser, David Silver, and Nando de Freitas · 2018
Cited alongside, same era.
Matt Hoffman, Bobak Shahriari, John Aslanides, Gabriel Barth-Maron, Feryal Behbahani, Tamara Norman, Abbas Abdolmaleki, Albin Cassirer, Fan Yang, Kate Baumli, et al · 2020
Later among the works it cites.
Morel: Model-based offline reinforcement learning
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims · 2020
Later among the works it cites.
Active reinforcement learning: Observing rewards at a cost
David Krueger, Jan Leike, Owain Evans, and John Salvatier · 2020
Later among the works it cites.
Conservative Q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Later among the works it cites.
Chip placement with deep reinforcement learning
Azalia Mirhoseini, Anna Goldie, Mustafa Yazgan, Joe Jiang, Ebrahim Songhori, Shen Wang, Young-Joon Lee, Eric Johnson, Omkar Pathak, Sungmin Bae, et al · 2020
Later among the works it cites.
Hyperparameter selection for offline reinforcement learning
Tom Le Paine, Cosmin Paduraru, Andrea Michi, Caglar Gulcehre, Konrad Zolna, Alexander Novikov, Ziyu Wang, and Nando de Freitas · 2020
Later among the works it cites.
Effective diversity in population based reinforcement learning
Jack Parker-Holder, Aldo Pacchiano, Krzysztof M Choromanski, and Stephen J Roberts · 2020
Later among the works it cites.
Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, et al · 2020
Later among the works it cites.
Keep doing what worked: Behavior modelling priors for offline reinforcement learning
Noah Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki, Michael Neunert, Thomas Lampe, Roland Hafner, Nicolas Heess, and Martin Riedmiller · 2020
Later among the works it cites.
dm_control: Software and tasks for continuous control
Yuval Tassa, Saran Tunyasuvunakool, Alistair Muldal, Yotam Doron, Siqi Liu, Steven Bohez, Josh Merel, Tom Erez, Timothy Lillicrap, and Nicolas Heess · 2020
Later among the works it cites.
Minimax weight and Q-function learning for off-policy evaluation
Masatoshi Uehara, Jiawei Huang, and Nan Jiang · 2020
Later among the works it cites.
Critic regularized regression
Ziyu Wang, Alexander Novikov, Konrad Żołna, Jost Tobias Springenberg, Scott Reed, Bobak Shahriari, Noah Siegel, Josh Merel, Caglar Gulcehre, Nicolas Heess, and Nando de Freitas · 2020
Later among the works it cites.
Offline policy selection under uncertainty
Mengjiao Yang, Bo Dai, Ofir Nachum, George Tucker, and Dale Schuurmans · 2020
Later among the works it cites.
Off-policy evaluation via the regularized lagrangian
Mengjiao Yang, Ofir Nachum, Bo Dai, Lihong Li, and Dale Schuurmans · 2020
Later among the works it cites.
MOPO: Model-based offline policy optimization
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma · 2020
Later among the works it cites.
Reverb: A framework for experience replay
Albin Cassirer, Gabriel Barth-Maron, Eugene Brevdo, Sabela Ramos, Toby Boyd, Thibault Sottiaux, and Manuel Kroiss · 2021
Closest in time.
Benchmarks for deep off-policy evaluation
Justin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker, ziyu wang, Alexander Novikov, Mengjiao Yang, Michael R Zhang, Yutian Chen, Aviral Kumar, Cosmin Paduraru, Sergey Levine, and Thomas Paine · 2021
Closest in time.
Regularized behavior value estimation
Caglar Gulcehre, Sergio Gómez Colmenarejo, Ziyu Wang, Jakub Sygnowski, Thomas Paine, Konrad Zolna, Yutian Chen, Matthew Hoffman, Razvan Pascanu, and Nando de Freitas · 2021
Closest in time.
Beyond pick-and-place: Tackling robotic stacking of diverse shapes
Alex X. Lee, Coline Manon Devin, Yuxiang Zhou, Thomas Lampe, Konstantinos Bousmalis, Jost Tobias Springenberg, Arunkumar Byravan, Abbas Abdolmaleki, Nimrod Gileadi, David Khosid, Claudio Fantacci, Jose Enrique Chen, Akhil Raju, Rae Jeong, Michael Neunert, Antoine Laurens, Stefano Saliceti, Federico Casarini, Martin Riedmiller, Raia Hadsell, and Francesco Nori · 2021
Closest in time.
Autoregressive dynamics models for offline policy evaluation and optimization
Michael R Zhang, Tom Le Paine, Ofir Nachum, Cosmin Paduraru, George Tucker, Ziyu Wang, and Mohammad Norouzi · 2021
Closest in time.