Fetching the paper…
Reading the bibliography…
In reinforcement learning (RL), when defining a Markov Decision Process (MDP), the environment dynamics is implicitly assumed to be stationary.
Stochastic processes
J. L. Doob · 1953
Earlier work this paper cites.
On the identification problem
L Zadeh · 1956
Earlier work this paper cites.
Dynamic Programming
Richard Bellman · 1957
Earlier work this paper cites.
Numerical identification of linear dynamic systems from normal operating records
Karl Johan Åström and Torsten Bohlin · 1965
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
Applied Nonlinear Control
J.J.E. Slotine and W. Li · 1991
Earlier work this paper cites.
Markov Chains and Stochastic Stability
S.P. Meyn and R.L. Tweedie · 1993
Earlier work this paper cites.
Continual learning in reinforcement environments
Mark Bishop Ring et al · 1994
Earlier work this paper cites.
Markov decision processes: Discrete stochastic dynamic programming
Martin L Puterman · 1995
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Optimal robot excitation and identification
Jan Swevers, Chris Ganseman, D Bilgin Tukel, Joris De Schutter, and Hendrik Van Brussel · 1997
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Leslie Pack Kaelbling, Michael L Littman, and Anthony R Cassandra · 1998
Earlier work this paper cites.
Lifelong learning algorithms
Sebastian Thrun · 1998
Earlier work this paper cites.
Computing the physical parameters of rigid-body motion from video
Kiran S Bhat, Steven M Seitz, Jovan Popović, and Pradeep K Khosla · 2002
Earlier work this paper cites.
Multiple model-based reinforcement learning
Kenji Doya, Kazuyuki Samejima, Ken-ichi Katagiri, and Mitsuo Kawato · 2002
Earlier work this paper cites.
Metrics for finite markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2004
Earlier work this paper cites.
Probabilistic policy reuse in a reinforcement learning agent
Fernando Fernández and Manuela Veloso · 2006
Earlier work this paper cites.
System identification without lennart ljung: what would have been different?
Michel Gevers et al · 2006
Earlier work this paper cites.
Perspectives on system identification
Lennart Ljung · 2009
Earlier work this paper cites.
Learning complex motions by sequencing simpler motion templates
Gerhard Neumann, Wolfgang Maass, and Jan Peters · 2009
Earlier work this paper cites.
Bisimulation metrics for continuous markov decision processes
Norm Ferns, Prakash Panangaden, and Doina Precup · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Subspace identification for linear systems: Theory—Implementation—Applications
Peter Van Overschee and BL De Moor · 2012
Earlier work this paper cites.
Robustness and generalization
Huan Xu and Shie Mannor · 2012
Earlier work this paper cites.
Finale Doshi-Velez and George Konidaris · 2013
Earlier work this paper cites.
A survey on concept drift adaptation
João Gama, Indrundefined Žliobaitundefined, Albert Bifet, Mykola Pechenizkiy, and Abdelhamid Bouchachia · 2014
Earlier work this paper cites.
Facial landmark detection by deep multi-task learning
Zhanpeng Zhang, Ping Luo, Chen Change Loy, and Xiaoou Tang · 2014
Earlier work this paper cites.
Learning visual predictive models of physics for playing billiards
Katerina Fragkiadaki, Pulkit Agrawal, Sergey Levine, and Jitendra Malik · 2015
Earlier work this paper cites.
Contextual markov decision processes, 2015
Assaf Hallak, Dotan Di Castro, and Shie Mannor · 2015
Earlier work this paper cites.
Abstraction selection in model-based reinforcement learning
Nan Jiang, Alex Kulesza, and Satinder Singh · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Andrei A Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Tobias Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Cited alongside, same era.
Successor features for transfer in reinforcement learning
André Barreto, Will Dabney, Rémi Munos, Jonathan J Hunt, Tom Schaul, Hado Van Hasselt, and David Silver · 2016
Cited alongside, same era.
David Ha, Andrew Dai, and Quoc V Le · 2016
Cited alongside, same era.
Terrain-adaptive locomotion skills using deep reinforcement learning
Xue Bin Peng, Glen Berseth, and Michiel Van de Panne · 2016
Cited alongside, same era.
Epopt: Learning robust neural network policies using model ensembles
On tiny episodic memories in continual learning
Arslan Chaudhry, Marcus Rohrbach, Mohamed Elhoseiny, Thalaiyasingam Ajanthan, Puneet K Dokania, Philip HS Torr, and Marc’Aurelio Ranzato · 2019
Later among the works it cites.
System identification: A machine learning perspective
Alessandro Chiuso and Gianluigi Pillonetto · 2019
Later among the works it cites.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
Ignasi Clavera, Anusha Nagabandi, Simin Liu, Ronald S. Fearing, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2019
Later among the works it cites.
Provably efficient RL with rich observations via latent state decoding
Simon Du, Akshay Krishnamurthy, Nan Jiang, Alekh Agarwal, Miroslav Dudik, and John Langford · 2019
Later among the works it cites.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aravind Rajeswaran, Sarvjeet Ghotra, Balaraman Ravindran, and Sergey Levine · 2016
Cited alongside, same era.
Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell · 2016
Cited alongside, same era.
Modular multitask reinforcement learning with policy sketches
Jacob Andreas, Dan Klein, and Sergey Levine · 2017
Cited alongside, same era.
Learning modular neural network policies for multi-task and multi-robot transfer
Coline Devin, Abhishek Gupta, Trevor Darrell, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Ubernet: Training a universal convolutional neural network for low-, mid-, and high-level vision using diverse datasets and limited memory
Iasonas Kokkinos · 2017
Cited alongside, same era.
Gradient episodic memory for continual learning
David Lopez-Paz and Marc’Aurelio Ranzato · 2017
Cited alongside, same era.
A simple neural attentive meta-learner
Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel · 2017
Cited alongside, same era.
Variational recurrent models for solving partially observable control tasks
Dongqi Han, Kenji Doya, and Jun Tani · 2019
Later among the works it cites.
Meta reinforcement learning as task inference
Jan Humplik, Alexandre Galashov, Leonard Hasenclever, Pedro A Ortega, Yee Whye Teh, and Nicolas Heess · 2019
Later among the works it cites.
Meta-learning representations for continual learning
Khurram Javed and Martha White · 2019
Later among the works it cites.
The marabou framework for verification and analysis of deep neural networks
Guy Katz, Derek A Huang, Duligur Ibeling, Kyle Julian, Christopher Lazarus, Rachel Lim, Parth Shah, Shantanu Thakoor, Haoze Wu, Aleksandar Zeljić, et al · 2019
Later among the works it cites.
Hypernetwork functional image representation
Sylwester Klocek, Łukasz Maziarka, Maciej Wołczyk, Jacek Tabor, Jakub Nowak, and Marek Śmieja · 2019
Later among the works it cites.
Modular universal reparameterization: Deep multi-task learning across diverse domains
Elliot Meyerson and Risto Miikkulainen · 2019
Later among the works it cites.
Deep online learning via meta-learning: Continual adaptation for model-based RL
Anusha Nagabandi, Chelsea Finn, and Sergey Levine · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Chelsea Finn, Sergey Levine, and Deirdre Quillen · 2019
Later among the works it cites.
Densephysnet: Learning dense physical object representations via multi-step dynamic interactions
Zhenjia Xu, Jiajun Wu, Andy Zeng, Joshua B Tenenbaum, and Shuran Song · 2019
Later among the works it cites.
Improving sample efficiency in model-free reinforcement learning from images
Denis Yarats, Amy Zhang, Ilya Kostrikov, Brandon Amos, Joelle Pineau, and Rob Fergus · 2019
Later among the works it cites.
Learning causal state representations of partially observable environments
Amy Zhang, Zachary C. Lipton, Luis Pineda, Kamyar Azizzadenesheli, Anima Anandkumar, Laurent Itti, Joelle Pineau, and Tommaso Furlanello · 2019
Later among the works it cites.
Environment probing interaction policies
Wenxuan Zhou, Lerrel Pinto, and Abhinav Gupta · 2019
Later among the works it cites.
Fast context adaptation via meta-learning
Luisa Zintgraf, Kyriacos Shiarli, Vitaly Kurin, Katja Hofmann, and Shimon Whiteson · 2019
Later among the works it cites.
Embracing change: Continual learning in deep neural networks
Raia Hadsell, Dushyant Rao, Andrei A Rusu, and Razvan Pascanu · 2020
Later among the works it cites.
Multitask soft option learning, 2020
Maximilian Igl, Andrew Gambardella, Jinke He, Nantas Nardelli, N. Siddharth, Wendelin Böhmer, and Shimon Whiteson · 2020
Later among the works it cites.
Pierre-Alexandre Kamienny, Matteo Pirotta, Alessandro Lazaric, Thibault Lavril, Nicolas Usunier, and Ludovic Denoyer · 2020
Later among the works it cites.
Towards continual reinforcement learning: A review and perspectives
Khimya Khetarpal, Matthew Riemer, Irina Rish, and Doina Precup · 2020
Later among the works it cites.
Context-aware dynamics model for generalization in model-based reinforcement learning
Kimin Lee, Younggyo Seo, Seunghyun Lee, Honglak Lee, and Jinwoo Shin · 2020
Later among the works it cites.
No-regret exploration in contextual reinforcement learning
Aditya Modi and Ambuj Tewari · 2020
Later among the works it cites.
Generalized Hidden Parameter MDPs:Transferable Model-Based RL in a Handful of Trials
Christian Perez, Felipe Petroski Such, and Theofanis Karaletsos · 2020
Later among the works it cites.
Toward Training Recurrent Neural Networks for Lifelong Learning
Shagun Sodhani, Sarath Chandar, and Yoshua Bengio · 2020
Later among the works it cites.
Deep reinforcement learning amidst lifelong non-stationarity, 2020
Annie Xie, James Harrison, and Chelsea Finn · 2020
Later among the works it cites.
Latent state models for meta-reinforcement learning from images
Tony Z. Zhao, Anusha Nagabandi, Kate Rakelly, Chelsea Finn, and Sergey Levine · 2020
Later among the works it cites.
Transfer learning in deep reinforcement learning: A survey
Zhuangdi Zhu, Kaixiang Lin, and Jiayu Zhou · 2020
Later among the works it cites.
Hyperdynamics: Generating expert dynamics models by observation
Zhou Xian, Shamit Lal, Hsiao-Yu Tung, Emmanouil Antonios Platanios, and Katerina Fragkiadaki · 2021
Closest in time.