Fetching the paper…
Reading the bibliography…
In many practical applications of RL, it is expensive to observe state transitions from the environment.
A possibility for implementing curiosity and boredom in model-building neural controllers
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Bayesian nonlinear modeling for the prediction competition
David JC MacKay et al · 1994
Earlier work this paper cites.
Bayesian experimental design: A review
Kathryn Chaloner and Isabella Verdinelli · 1995
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Radford M Neal · 1995
Earlier work this paper cites.
Gaussian processes for regression
Christopher KI Williams and Carl Edward Rasmussen · 1996
Earlier work this paper cites.
Bayesian q-learning
Richard Dearden, Nir Friedman, and Stuart Russell · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S. Sutton and Andrew G. Barto · 1998
Earlier work this paper cites.
Model-based bayesian exploration
Richard Dearden, Nir Friedman, and David Andre · 1999
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large markov decision processes
Michael Kearns, Yishay Mansour, and Andrew Y Ng · 2002
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade · 2003
Earlier work this paper cites.
Using computational fluid dynamics for aerodynamics, 2006
Antony Jameson and Massimiliano Fatica · 2006
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi, Benjamin Recht, et al · 2007
Earlier work this paper cites.
Bayes-adaptive pomdps
Stephane Ross, Brahim Chaib-draa, and Joelle Pineau · 2007
Earlier work this paper cites.
Probabilistic planning for robotic exploration
Trey Smith · 2007
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
J Zico Kolter and Andrew Y Ng · 2009
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Peter Deisenroth and Carl Edward Rasmussen · 2011
Earlier work this paper cites.
Information collection on a graph
Ilya O Ryzhov and Warren B Powell · 2011
Earlier work this paper cites.
Efficient bayes-adaptive reinforcement learning using sample-based search
Arthur Guez, David Silver, and Peter Dayan · 2012
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Manuel Lopes, Tobias Lang, Marc Toussaint, and Pierre-yves Oudeyer · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Minimax pac bounds on the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Hilbert J Kappen · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Learning to optimize via information-directed sampling
Daniel Russo and Benjamin Van Roy · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Taking the human out of the loop: A review of bayesian optimization
Bobak Shahriari, Kevin Swersky, Ziyu Wang, Ryan P Adams, and Nando De Freitas · 2015
Cited alongside, same era.
Offline contextual bayesian optimization
Ian Char, Youngseog Chung, Willie Neiswanger, Kirthevasan Kandasamy, Andrew O Nelson, Mark Boyer, Egemen Kolemen, and Jeff Schneider · 2019
Later among the works it cites.
Information-directed exploration for deep reinforcement learning
Nikolay Nikolov, Johannes Kirschner, Felix Berkenkamp, and Andreas Krause · 2019
Later among the works it cites.
Self-supervised exploration via disagreement
Deepak Pathak, Dhiraj Gandhi, and Abhinav Gupta · 2019
Later among the works it cites.
Bayesian exploration for approximate dynamic programming
Ilya O Ryzhov, Martijn RK Mes, Warren B Powell, and Gerald van den Berg · 2019
Later among the works it cites.
Model-based active exploration
Pranav Shyam, Wojciech Jaśkowski, and Faustino Gomez · 2019
Later among the works it cites.
Model-based reinforcement learning with a generative model is minimax optimal
Alekh Agarwal, Sham Kakade, and Lin F. Yang · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unifying count-based exploration and intrinsic motivation
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos · 2016
Cited alongside, same era.
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Bayesian reinforcement learning: A survey
Mohammad Ghavamzadeh, Shie Mannor, Joelle Pineau, and Aviv Tamar · 2016
Cited alongside, same era.
Deep exploration via bootstrapped dqn
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
UCB and infogain exploration via q q -ensembles
Richard Y. Chen, Szymon Sidor, Pieter Abbeel, and John Schulman · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Later among the works it cites.
Ready policy one: World building through active learning
Philip Ball, Jack Parker-Holder, Aldo Pacchiano, Krzysztof Choromanski, and Stephen Roberts · 2020
Later among the works it cites.
Actively learning gaussian process dynamics
Mona Buisson-Fenet, Friedrich Solowjow, and Sebastian Trimpe · 2020
Later among the works it cites.
Breaking the sample size barrier in model-based reinforcement learning with a generative model
Gen Li, Yuting Wei, Yuejie Chi, Yuantao Gu, and Yuxin Chen · 2020
Later among the works it cites.
Neural dynamical systems: Balancing structure and flexibility in physical prediction
Viraj Mehta, Ian Char, Willie Neiswanger, Youngseog Chung, Andrew Oakleigh Nelson, Mark D Boyer, Egemen Kolemen, and Jeff Schneider · 2020
Later among the works it cites.
Sample-efficient cross-entropy method for real-time planning
Cristina Pinneri, Shambhuraj Sawant, Sebastian Blaes, Jan Achterhold, Joerg Stueckler, Michal Rolinek, and Georg Martius · 2020
Later among the works it cites.
Exploring model-based planning with policy networks
Tingwu Wang and Jimmy Ba · 2020
Later among the works it cites.
Efficiently sampling functions from gaussian process posteriors
James Wilson, Viacheslav Borovitskiy, Alexander Terenin, Peter Mostowsky, and Marc Deisenroth · 2020
Later among the works it cites.
Data-driven profile prediction for DIII-d
J. Abbate, R. Conlin, and E. Kolemen · 2021
Closest in time.
Explore the context: Optimal data collection for context-conditional dynamics models
Jan Achterhold and Joerg Stueckler · 2021
Closest in time.
The value of information when deciding what to learn, 2021
Dilip Arumugam and Benjamin Van Roy · 2021
Closest in time.
DiSECt: A Differentiable Simulation Engine for Autonomous Robotic Cutting
Eric Heiden, Miles Macklin, Yashraj S Narang, Dieter Fox, Animesh Garg, and Fabio Ramos · 2021
Closest in time.
Information directed reward learning for reinforcement learning
David Lindner, Matteo Turchetta, Sebastian Tschiatschek, Kamil Ciosek, and Andreas Krause · 2021
Closest in time.
Bayesian algorithm execution: Estimating computable properties of black-box functions using mutual information
Willie Neiswanger, Ke Alexander Wang, and Stefano Ermon · 2021
Closest in time.
Mbrl-lib: A modular library for model-based reinforcement learning
Luis Pineda, Brandon Amos, Amy Zhang, Nathan O. Lambert, and Roberto Calandra · 2021
Closest in time.