Fetching the paper…
Reading the bibliography…
We consider a context-dependent Reinforcement Learning (RL) setting, which is characterized by: a) an unknown finite number of not directly observable contexts; b) abrupt (discontinuous) context changes occurring during an episode; and c) Markovian context evolution.
Reinforcement learning in non-stationary environments
Sindhu Padakandla, Prabuchandran K. J., and Shalabh Bhatnagar · 1905
Earlier work this paper cites.
Optimal control of Markov processes with incomplete state information
Karl J Astrom · 1965
Earlier work this paper cites.
Nonnegative matrices in the mathematical sciences
Abraham Berman and Robert J Plemmons · 1994
Earlier work this paper cites.
A constructive definition of Dirichlet priors
Jayaram Sethuraman · 1994
Earlier work this paper cites.
Sensor/actuator failure detection in the Vista F-16 by multiple model adaptive estimation
Timothy E Menke and Peter S Maybeck · 1995
Earlier work this paper cites.
Hidden-mode Markov decision processes for nonstationary sequential decision making
Samuel PM Choi, Dit-Yan Yeung, and Nevin L Zhang · 2000
Earlier work this paper cites.
Value-function approximations for partially observable Markov decision processes
Milos Hauskrecht · 2000
Earlier work this paper cites.
Variational inference for Dirichlet process mixtures
David M Blei, Michael I Jordan, et al · 2006
Earlier work this paper cites.
Dealing with non-stationary environments using context detection
Bruno Castro da Silva, Eduardo W. Basso, Ana L. C. Bazzan, and Paulo Martins Engel · 2006
Earlier work this paper cites.
On perron–frobenius property of matrices having some negative entries
Dimitrios Noutsos · 2006
Earlier work this paper cites.
Point-based value iteration for continuous pomdps
Josep M Porta, Nikos Vlassis, Matthijs TJ Spaan, and Pascal Poupart · 2006
Earlier work this paper cites.
Hierarchical Dirichlet processes
Yee Whye Teh, Michael I Jordan, Matthew J Beal, and David M Blei · 2006
Earlier work this paper cites.
System design for uncertainty
Franz S Hover and Michael S Triantafyllou · 2009
Earlier work this paper cites.
Planning under uncertainty for robotic tasks with mixed observability
Sylvie CW Ong, Shao Wei Png, David Hsu, and Wee Sun Lee · 2010
Earlier work this paper cites.
Bayesian nonparametric inference of switching dynamic linear models
Emily Fox, Erik B Sudderth, Michael I Jordan, and Alan S Willsky · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Truly nonparametric online variational inference for hierarchical Dirichlet processes
Michael Bryant and Erik Sudderth · 2012
Earlier work this paper cites.
Towards continual reinforcement learning: A review and perspectives
Khimya Khetarpal, Matthew Riemer, Irina Rish, and Doina Precup · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Online multi-task learning for policy gradient methods
Haitham Bou-Ammar, Eric Eaton, Paul Ruvolo, and Matthew E. Taylor · 2014
Earlier work this paper cites.
Sequential decision-making under non-stationary environments via sequential change-point detection
Emmanuel Hadoux, Aurélie Beynier, and Paul Weng · 2014
Earlier work this paper cites.
Weight uncertainty in neural network
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Earlier work this paper cites.
Contextual markov decision processes
Assaf Hallak, Dotan Di Castro, and Shie Mannor · 2015
Earlier work this paper cites.
Deep recurrent Q-learning for partially observable MDPs
Matthew Hausknecht and Peter Stone · 2015
Earlier work this paper cites.
Reliable and scalable variational inference for the hierarchical Dirichlet process
Michael Hughes, Dae Il Kim, and Erik Sudderth · 2015
Cited alongside, same era.
Variational dropout and the local reparameterization trick
Durk P Kingma, Tim Salimans, and Max Welling · 2015
Cited alongside, same era.
Hidden parameter Markov decision processes: A semiparametric regression approach for discovering latent task parametrizations
Finale Doshi-Velez and George Dimitri Konidaris · 2016
Cited alongside, same era.
RL 2 : Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
Switching linear dynamics for variational Bayes filtering
Philip Becker-Ehmck, Jan Peters, and Patrick Van Der Smagt · 2019
Later among the works it cites.
Learning action representations for reinforcement learning
Yash Chandak, Georgios Theocharous, James Kostas, Scott M. Jordan, and Philip S. Thomas · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Later among the works it cites.
Hyunjik Kim, Andriy Mnih, Jonathan Schwarz, Marta Garnelo, Ali Eslami, Dan Rosenbaum, Oriol Vinyals, and Yee Whye Teh · 2019
Later among the works it cites.
Bayesian policy optimization for model uncertainty
Gilwoo Lee, Brian Hou, Aditya Mandalika, Jeongseok Lee, Sanjiban Choudhury, and Siddhartha S. Srinivasa · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Quickest change detection approach to optimal control in Markov decision processes with model changes
Taposh Banerjee, Miao Liu, and Jonathan P. How · 2017
Cited alongside, same era.
Variational inference: A review for statisticians
David M Blei, Alp Kucukelbir, and Jon D McAuliffe · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
On improving deep reinforcement learning for POMDPs
Pengfei Zhu, Xin Li, Pascal Poupart, and Guanghui Miao · 2017
Cited alongside, same era.
Openai spinning up documentation, 2018
Josh Achiam · 2018
Cited alongside, same era.
Context-aware policy reuse
Siyuan Li, Fangda Gu, Guangxiang Zhu, and Chongjie Zhang · 2019
Later among the works it cites.
gym-cartpole-swingup. a simple, continuous-control environment for openai gym
Angelo Lovatto · 2019
Later among the works it cites.
Recurrent attentive neural process for sequential data
Shenghao Qin, Jiacheng Zhu, Jimmy Qin, Wenshuo Wang, and Ding Zhao · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Chelsea Finn, Sergey Levine, and Deirdre Quillen · 2019
Later among the works it cites.
Learning to learn without forgetting by maximizing transfer and minimizing interference
Matthew Riemer, Ignacio Cases, Robert Ajemian, Miao Liu, Irina Rish, Yuhai Tu, and Gerald Tesauro · 2019
Later among the works it cites.
Experience replay for continual learning
David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy P. Lillicrap, and Gregory Wayne · 2019
Later among the works it cites.
ProMP: Proximal meta-policy search
Jonas Rothfuss, Dennis Lee, Ignasi Clavera, Tamim Asfour, and Pieter Abbeel · 2019
Later among the works it cites.
Optimizing for the future in non-stationary MDPs
Yash Chandak, Georgios Theocharous, Shiv Shankar, Martha White, Sridhar Mahadevan, and Philip Thomas · 2020
Later among the works it cites.
Collapsed amortized variational inference for switching nonlinear dynamical systems
Zhe Dong, Bryan Seybold, Kevin Murphy, and Hung Bui · 2020
Later among the works it cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Ilya Kostrikov, Denis Yarats, and Rob Fergus · 2020
Later among the works it cites.
Deep reinforcement learning amidst lifelong non-stationarity
Annie Xie, James Harrison, and Chelsea Finn · 2020
Later among the works it cites.
Task-agnostic online reinforcement learning with an infinite mixture of Gaussian processes
Mengdi Xu, Wenhao Ding, Jiacheng Zhu, Zuxin Liu, Baiming Chen, and Ding Zhao · 2020
Later among the works it cites.
Soft actor-critic (sac) implementation in pytorch
Denis Yarats and Ilya Kostrikov · 2020
Later among the works it cites.
Invariant causal prediction for block MDPs
Amy Zhang, Clare Lyle, Shagun Sodhani, Angelos Filos, Marta Kwiatkowska, Joelle Pineau, Yarin Gal, and Doina Precup · 2020
Later among the works it cites.
Varibad: A very good method for bayes-adaptive deep RL via meta-learning
Luisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, and Shimon Whiteson · 2020
Later among the works it cites.
Carl: A benchmark for contextual and adaptive reinforcement learning
Carolin Benjamins, Theresa Eimer, Frederik Schubert, André Biedenkapp, Bodo Rosenhahn, Frank Hutter, and Marius Lindauer · 2021
Later among the works it cites.
A continual learning survey: Defying forgetting in classification tasks
Matthias Delange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Ales Leonardis, Greg Slabaugh, and Tinne Tuytelaars · 2021
Later among the works it cites.
Generalization in reinforcement learning by soft data augmentation
Nicklas Hansen and Xiaolong Wang · 2021
Later among the works it cites.
Learning to fly—a gym environment with pybullet physics for reinforcement learning of multi-agent quadcopter control
Jacopo Panerati, Hehui Zheng, SiQi Zhou, James Xu, Amanda Prorok, and Angela P. Schoellig · 2021
Later among the works it cites.
Mbrl-lib: A modular library for model-based reinforcement learning
Luis Pineda, Brandon Amos, Amy Zhang, Nathan O. Lambert, and Roberto Calandra · 2021
Later among the works it cites.