Fetching the paper…
Reading the bibliography…
We propose a simple, practical, and intuitive approach for domain adaptation in reinforcement learning.
Dual control theory. i
AA Feldbaum · 1960
Earlier work this paper cites.
Training and tracking in robotics
Oliver G Selfridge, Richard S Sutton, and Andrew G Barto · 1985
Earlier work this paper cites.
Adaptive control of linearizable systems
Sosale Shankara Sastry and Alberto Isidori · 1989
Earlier work this paper cites.
Neural networks for control and system identification
Paul J Werbos · 1989
Earlier work this paper cites.
Adaptive dual control methods: An overview
Björn Wittenmark · 1995
Earlier work this paper cites.
System identification
Lennart Ljung · 1999
Earlier work this paper cites.
Using options for knowledge transfer in reinforcement learning
Theodore J Perkins, Doina Precup, et al · 1999
Earlier work this paper cites.
Risk-sensitive reinforcement learning
Oliver Mihatsch and Ralph Neuneier · 2002
Earlier work this paper cites.
Multitask reinforcement learning on the distribution of mdps
Fumihide Tanaka and Masayuki Yamamura · 2003
Earlier work this paper cites.
Transfer of experience between reinforcement learning environments with progressive difficulty
Michael G Madden and Tom Howley · 2004
Earlier work this paper cites.
An algebraic approach to abstraction in reinforcement learning
Balaraman Ravindran and Andrew G Barto · 2004
Earlier work this paper cites.
Learning and evaluating classifiers under sample selection bias
Bianca Zadrozny · 2004
Earlier work this paper cites.
Path integrals and symmetry breaking for optimal control theory
Hilbert J Kappen · 2005
Earlier work this paper cites.
Improving action selection in mdp’s via knowledge transfer
Alexander A Sherstov and Peter Stone · 2005
Earlier work this paper cites.
Model transfer for markov decision tasks via parameter matching
Funlade T Sunmola and Jeremy L Wyatt · 2006
Earlier work this paper cites.
Discriminative learning for differing training and test distributions
Steffen Bickel, Michael Brückner, and Tobias Scheffer · 2007
Earlier work this paper cites.
Linearly-solvable markov decision problems
Emanuel Todorov · 2007
Earlier work this paper cites.
Knowledge transfer in reinforcement learning
Alessandro Lazaric · 2008
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Daphne Koller and Nir Friedman · 2009
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Matthew E Taylor and Peter Stone · 2009
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Marc Toussaint · 2009
Earlier work this paper cites.
Modeling Purposeful Adaptive Behavior with the Principle of Maximum Causal Entropy
Brian D. Ziebart · 2010
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
The transferability approach: Crossing the reality gap in evolutionary robotics
Sylvain Koos, Jean-Baptiste Mouret, and Stéphane Doncieux · 2012
Earlier work this paper cites.
Agnostic system identification for model-based reinforcement learning
Stephane Ross and J Andrew Bagnell · 2012
Earlier work this paper cites.
Humanoid robots learning to walk faster: From the real world to simulation and back
Alon Farchy, Samuel Barrett, Patrick MacAlpine, and Peter Stone · 2013
Earlier work this paper cites.
Unsupervised visual domain adaptation using subspace alignment
Basura Fernando, Amaury Habrard, Marc Sebban, and Tinne Tuytelaars · 2013
Earlier work this paper cites.
On stochastic optimal control and reinforcement learning by approximate inference
Konrad Rawlik, Marc Toussaint, and Sethu Vijayakumar · 2013
Cited alongside, same era.
Scaling up robust mdps by reinforcement learning
Aviv Tamar, Huan Xu, and Shie Mannor · 2013
Cited alongside, same era.
Adaptive model predictive control for constrained linear systems
Marko Tanaskovic, Lorenzo Fagiano, Roy Smith, Paul Goulart, and Manfred Morari · 2013
Cited alongside, same era.
Domain adaptation and sample bias correction theory and algorithm for regression
Corinna Cortes and Mehryar Mohri · 2014
Cited alongside, same era.
Reinforcement learning with multi-fidelity simulators
Mark Cutler, Thomas J Walsh, and Jonathan P How · 2014
Cited alongside, same era.
Policy evaluation with temporal differences: A survey and comparison
Robust and efficient transfer learning with hidden parameter markov decision processes
Taylor W Killian, Samuel Daulton, George Konidaris, and Finale Doshi-Velez · 2017
Later among the works it cites.
A simple neural attentive meta-learner
Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel · 2017
Later among the works it cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Later among the works it cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel · 2017
Later among the works it cites.
Unifying task specification in reinforcement learning
Martha White · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christoph Dann, Gerhard Neumann, Jan Peters, et al · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Model predictive path integral control using covariance variable importance sampling
Grady Williams, Andrew Aldrich, and Evangelos Theodorou · 2015
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Cited alongside, same era.
Rlˆ2: Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Cited alongside, same era.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, François Laviolette, Mario Marchand, and Victor Lempitsky · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Addressing appearance change in outdoor robotics with adversarial domain adaptation
Markus Wulfmeier, Alex Bewley, and Ingmar Posner · 2017
Later among the works it cites.
Preparing for the unknown: Learning a universal policy with online system identification
Wenhao Yu, Jie Tan, C Karen Liu, and Greg Turk · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Later among the works it cites.
Using simulation and domain adaptation to improve efficiency of deep robotic grasping
Konstantinos Bousmalis, Alex Irpan, Paul Wohlhart, Yunfei Bai, Matthew Kelcey, Mrinal Kalakrishnan, Laura Downs, Julian Ibarz, Peter Pastor, Kurt Konolige, et al · 2018
Later among the works it cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Later among the works it cites.
Learning to adapt: Meta-learning for model-based control
Ignasi Clavera, Anusha Nagabandi, Ronald S Fearing, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2018
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2018
Later among the works it cites.
Transfer learning for related reinforcement learning tasks via image-to-image translation
Shani Gamrian and Yoav Goldberg · 2018
Later among the works it cites.
Tf-agents: A library for reinforcement learning in tensorflow, 2018
Sergio Guadarrama, Anoop Korattikara, Oscar Ramirez, Pablo Castro, Ethan Holly, Sam Fishman, Ke Wang, Ekaterina Gonina, Chris Harris, Vincent Vanhoucke, et al · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Learning latent dynamics for planning from pixels
Danijar Hafner, Timothy Lillicrap, Ian Fischer, Ruben Villegas, David Ha, Honglak Lee, and James Davidson · 2018
Later among the works it cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Later among the works it cites.
Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection
Sergey Levine, Peter Pastor, Alex Krizhevsky, Julian Ibarz, and Deirdre Quillen · 2018
Later among the works it cites.
Detecting and correcting for label shift with black box predictors
Zachary C Lipton, Yu-Xiang Wang, and Alex Smola · 2018
Later among the works it cites.
Sim-to-real transfer of robotic control with dynamics randomization
Xue Bin Peng, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Later among the works it cites.
Cycle-consistent adversarial learning as approximate bayesian inference
Louis C Tiao, Edwin V Bonilla, and Fabio Ramos · 2018
Later among the works it cites.
Closing the sim-to-real loop: Adapting simulation randomization with real world experience
Yevgen Chebotar, Ankur Handa, Viktor Makoviychuk, Miles Macklin, Jan Issac, Nathan Ratliff, and Dieter Fox · 2019
Later among the works it cites.
When to trust your model: Model-based policy optimization
Michael Janner, Justin Fu, Marvin Zhang, and Sergey Levine · 2019
Later among the works it cites.
A review of domain adaptation without target labels
Wouter Marco Kouw and Marco Loog · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Chelsea Finn, Sergey Levine, and Deirdre Quillen · 2019
Later among the works it cites.
V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control
H Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark, Hubert Soyer, Jack W Rae, Seb Noury, Arun Ahuja, Siqi Liu, Dhruva Tirumala, et al · 2019
Later among the works it cites.
Planning and execution using inaccurate models with provable guarantees
Anirudh Vemula, Yash Oza, J Andrew Bagnell, and Maxim Likhachev · 2020
Closest in time.