Fetching the paper…
Reading the bibliography…
One of the key reasons for the high sample complexity in reinforcement learning (RL) is the inability to transfer knowledge from one task to another.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B · 1992
Earlier work this paper cites.
Reinforcement learning: A survey
Kaelbling, L. P., Littman, M. L., and Moore, A. W · 1996
Earlier work this paper cites.
Multitask learning
Caruana, R · 1997
Earlier work this paper cites.
The acquisition of skilled motor performance: fast and slow experience-driven changes in primary motor cortex
Karni, A., Meyer, G., Rey-Hipolito, C., Jezzard, P., Adams, M. M., Turner, R., and Ungerleider, L. G · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton, R. S., Precup, D., and Singh, S · 1999
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y., Russell, S. J., et al · 2000
Earlier work this paper cites.
Acquisition of stand-up behavior by a real robot using hierarchical reinforcement learning
Morimoto, J. and Doya, K · 2001
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
Precup, D., Sutton, R. S., and Dasgupta, S · 2001
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Intrinsically motivated learning of hierarchical collections of skills
Barto, A. G., Singh, S., and Chentanez, N · 2004
Earlier work this paper cites.
Maximum margin planning
Ratliff, N. D., Bagnell, J. A., and Zinkevich, M. A · 2006
Earlier work this paper cites.
Adapting svm classifiers to data with shifted distributions
Yang, J., Yan, R., and Hauptmann, A. G · 2007
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
A survey on transfer learning
Pan, S. J. and Yang, Q · 2009
Earlier work this paper cites.
Transfer learning for reinforcement learning domains: A survey
Taylor, M. E. and Stone, P · 2009
Earlier work this paper cites.
Learning-based control strategy for safe human-robot interaction exploiting task and robot redundancies
Calinon, S., Sardellitti, I., and Caldwell, D. G · 2010
Earlier work this paper cites.
Adapting visual category models to new domains
Saenko, K., Kulis, B., Fritz, M., and Darrell, T · 2010
Earlier work this paper cites.
Transfer learning
Torrey, L. and Shavlik, J · 2010
Earlier work this paper cites.
Tabula rasa: Model transfer for object category detection
Aytar, Y. and Zisserman, A · 2011
Earlier work this paper cites.
Domain adaptation for object recognition: An unsupervised approach
Gopalan, R., Li, R., and Chellappa, R · 2011
Earlier work this paper cites.
What you saw is not what you get: Domain adaptation using asymmetric kernel transforms
Kulis, B., Saenko, K., and Darrell, T · 2011
Cited alongside, same era.
Energy efficient use of robotics in the automobile industry
Meike, D. and Ribickis, L · 2011
Cited alongside, same era.
Robust visual domain adaptation with low-rank reconstruction
Jhuo, I.-H., Liu, D., Lee, D., and Chang, S.-F · 2012
Cited alongside, same era.
Unsupervised visual domain adaptation using subspace alignment
Fernando, B., Habrard, A., Sebban, M., and Tuytelaars, T · 2013
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R., Donahue, J., Darrell, T., and Malik, J · 2014
Cited alongside, same era.
Continuous manifold based adaptation for evolving visual domains
Hoffman, J., Darrell, T., and Saenko, K · 2014
Learning modular neural network policies for multi-task and multi-robot transfer
Devin, C., Gupta, A., Darrell, T., Abbeel, P., and Levine, S · 2017
Later among the works it cites.
Multi-task self-supervised visual learning
Doersch, C. and Zisserman, A · 2017
Later among the works it cites.
Meta learning shared hierarchies
Frans, K., Ho, J., Chen, X., Abbeel, P., and Schulman, J · 2017
Later among the works it cites.
Schema networks: Zero-shot transfer with a generative causal model of intuitive physics
Kansky, K., Silver, T., Mély, D. A., Eldawy, M., Lázaro-Gredilla, M., Lou, X., Dorfman, N., Sidor, S., Phoenix, S., and George, D · 2017
Later among the works it cites.
Ubernet: Training a universal convolutional neural network for low-, mid-, and high-level vision using diverse datasets and limited memory
Kokkinos, I · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M · 2014
Cited alongside, same era.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Cited alongside, same era.
Actor-mimic: Deep multitask and transfer reinforcement learning
Parisotto, E., Ba, J. L., and Salakhutdinov, R · 2015
Cited alongside, same era.
Universal value function approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D · 2015
Cited alongside, same era.
Levy, A., Platt, R., and Saenko, K · 2017
Later among the works it cites.
Infogail: Interpretable imitation learning from visual demonstrations
Li, Y., Song, J., and Ermon, S · 2017
Later among the works it cites.
Deep decentralized multi-task multi-agent reinforcement learning under partial observability
Omidshafiei, S., Pazis, J., Amato, C., How, J. P., and Vian, J · 2017
Later among the works it cites.
Learning to push by grasping: Using multiple tasks for effective learning
Pinto, L. and Gupta, A · 2017
Later among the works it cites.
Asymmetric actor critic for image-based robot learning
Pinto, L., Andrychowicz, M., Welinder, P., Zaremba, W., and Abbeel, P · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al · 2017
Later among the works it cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P · 2017
Later among the works it cites.
Deep drone racing: Learning agile flight in dynamic environments
Kaufmann, E., Loquercio, A., Ranftl, R., Dosovitskiy, A., Koltun, V., and Scaramuzza, D · 2018
Later among the works it cites.
Data-efficient hierarchical reinforcement learning
Nachum, O., Gu, S. S., Lee, H., and Levine, S · 2018
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Nagabandi, A., Kahn, G., Fearing, R. S., and Levine, S · 2018
Later among the works it cites.
Visual reinforcement learning with imagined goals
Nair, A. V., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S · 2018
Later among the works it cites.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., Welinder, P., et al · 2018
Later among the works it cites.
Goal-conditioned imitation learning
Ding, Y., Florensa, C., Abbeel, P., and Phielipp, M · 2019
Later among the works it cites.
Sub-policy adaptation for hierarchical reinforcement learning
Li, A. C., Florensa, C., Clavera, I., and Abbeel, P · 2019
Later among the works it cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2019
Later among the works it cites.