Fetching the paper…
Reading the bibliography…
Trading off exploration and exploitation in an unknown environment is key to maximising expected return during learning.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
A problem in the sequential design of experiments
Richard Bellman · 1956
Earlier work this paper cites.
Bayesian decision problems and Markov chains
James John Martin · 1967
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
A bayesian framework for reinforcement learning
Malcolm Strens · 2000
Earlier work this paper cites.
R-max-a general polynomial time algorithm for near-optimal reinforcement learning
Ronen I Brafman and Moshe Tennenholtz · 2002
Earlier work this paper cites.
Optimal Learning: Computational procedures for Bayes-adaptive Markov decision processes
Michael O’Gordon Duff and Andrew Barto · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
An analytic solution to discrete bayesian reinforcement learning
Pascal Poupart, Nikos Vlassis, Jesse Hoey, and Kevin Regan · 2006
Earlier work this paper cites.
Near-bayesian exploration in polynomial time
J Zico Kolter and Andrew Y Ng · 2009
Earlier work this paper cites.
Learning is planning: near bayes-optimal reinforcement learning via monte-carlo tree search
John Asmuth and Michael L Littman · 2011
Earlier work this paper cites.
Bayesian policy search with policy priors
David Wingate, Noah D Goodman, Daniel M Roy, Leslie P Kaelbling, and Joshua B Tenenbaum · 2011
Earlier work this paper cites.
Bayes-optimal reinforcement learning for discrete uncertainty domains
Emma Brunskill · 2012
Earlier work this paper cites.
Efficient bayes-adaptive reinforcement learning using sample-based search
Arthur Guez, David Silver, and Peter Dayan · 2012
Earlier work this paper cites.
Variance-based rewards for approximate bayesian reinforcement learning
Jonathan Sorg, Satinder Singh, and Richard L Lewis · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Acting optimally in partially observable stochastic domains
Anthony R. Cassandra, Leslie Pack Kaelbling, and Michael L. Littman · 2013
Earlier work this paper cites.
Scalable and efficient bayes-adaptive reinforcement learning based on monte-carlo tree search
Arthur Guez, David Silver, and Peter Dayan · 2013
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2014
Earlier work this paper cites.
Trading off scientific knowledge and user learning with multi-armed bandits
Yun-En Liu, Travis Mandel, Emma Brunskill, and Zoran Popovic · 2014
Earlier work this paper cites.
Bayesian reinforcement learning: A survey
Mohammad Ghavamzadeh, Shie Mannor, Joelle Pineau, Aviv Tamar, et al · 2015
Cited alongside, same era.
Contextual markov decision processes
Assaf Hallak, Dotan Di Castro, and Shie Mannor · 2015
Cited alongside, same era.
Hidden parameter markov decision processes: A semiparametric regression approach for discovering latent task parametrizations
Finale Doshi-Velez and George Konidaris · 2016
Cited alongside, same era.
RL 2 : Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Cited alongside, same era.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Cited alongside, same era.
Reptile: a scalable metalearning algorithm
Alex Nichol and John Schulman · 2018
Later among the works it cites.
Randomized prior functions for deep reinforcement learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Later among the works it cites.
Efficient transfer learning and online adaptation with latent variable models for continuous control
Christian F Perez, Felipe Petroski Such, and Theofanis Karaletsos · 2018
Later among the works it cites.
Meta reinforcement learning with latent variable gaussian processes
Steindór Sæmundsson, Katja Hofmann, and Marc Peter Deisenroth · 2018
Later among the works it cites.
Some considerations on learning to explore via meta-reinforcement learning
Bradly C Stadie, Ge Yang, Rein Houthooft, Xi Chen, Yan Duan, Yuhuai Wu, Pieter Abbeel, and Ilya Sutskever · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
One-shot imitation learning
Yan Duan, Marcin Andrychowicz, Bradly Stadie, OpenAI Jonathan Ho, Jonas Schneider, Ilya Sutskever, Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu · 2017
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Cited alongside, same era.
Robust and efficient transfer learning with hidden parameter markov decision processes
Taylor W Killian, Samuel Daulton, George Konidaris, and Finale Doshi-Velez · 2017
Cited alongside, same era.
A simple neural attentive meta-learner
Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Sebastian Tschiatschek, Kai Arulkumaran, Jan Stühmer, and Katja Hofmann · 2018
Later among the works it cites.
Direct policy transfer via hidden parameter markov decision processes
Jiayu Yao, Taylor Killian, George Konidaris, and Finale Doshi-Velez · 2018
Later among the works it cites.
Reinforcement learning with action-derived rewards for chemotherapy and clinical trial dosing regimen selection
Gregory Yauney and Pratik Shah · 2018
Later among the works it cites.
Decoupling dynamics and reward for transfer learning
Amy Zhang, Harsh Satija, and Joelle Pineau · 2018
Later among the works it cites.
Vpe: Variational policy embedding for transfer reinforcement learning
Isac Arnekvist, Danica Kragic, and Johannes A Stork · 2019
Closest in time.
Meta-amortized variational inference and learning
Kristy Choi, Mike Wu, Noah Goodman, and Stefano Ermon · 2019
Closest in time.
Meta-learning probabilistic inference for prediction
Jonathan Gordon, John Bronskill, Matthias Bauer, Sebastian Nowozin, and Richard E Turner · 2019
Closest in time.
Meta reinforcement learning as task inference
Jan Humplik, Alexandre Galashov, Leonard Hasenclever, Pedro A Ortega, Yee Whye Teh, and Nicolas Heess · 2019
Closest in time.
Deep variational reinforcement learning for pomdps
Maximilian Igl, Luisa Zintgraf, Tuan Anh Le, Frank Wood, and Shimon Whiteson · 2019
Closest in time.
Meta reinforcement learning with task embedding and shared policy
Lin Lan, Zhenguo Li, Xiaohong Guan, and Pinghui Wang · 2019
Closest in time.
Bayesian policy optimization for model uncertainty
Gilwoo Lee, Brian Hou, Aditya Mandalika, Jeongseok Lee, and Siddhartha S Srinivasa · 2019
Closest in time.
Contextual markov decision processes using generalized linear models
Aditya Modi and Ambuj Tewari · 2019
Closest in time.
Meta-learning of sequential strategies
Pedro A Ortega, Jane X Wang, Mark Rowland, Tim Genewein, Zeb Kurth-Nelson, Razvan Pascanu, Nicolas Heess, Joel Veness, Alex Pritzel, Pablo Sprechmann, et al · 2019
Closest in time.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Deirdre Quillen, Chelsea Finn, and Sergey Levine · 2019
Closest in time.
Promp: Proximal meta-policy search
Jonas Rothfuss, Dennis Lee, Ignasi Clavera, Tamim Asfour, and Pieter Abbeel · 2019
Closest in time.
Fast context adaptation via meta-learning
Luisa M Zintgraf, Kyriacos Shiarlis, Vitaly Kurin, Katja Hofmann, and Shimon Whiteson · 2019
Closest in time.