Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) is an important machine learning paradigm for solving sequential decision-making problems.
M. L. Puterman, “Markov decision processes,” in Stochastic Models , 1990, vol. 2, pp. 331–434
1990
Earlier work this paper cites.
M. B. Ring, “Child: A first step towards continual learning,” Mach. Learn. , vol. 28, no. 1, pp. 77–104, 1997
1997
Earlier work this paper cites.
A. G. Barto and S. Mahadevan, “Recent advances in hierarchical reinforcement learning,” Discrete Event Dyn. Syst. , vol. 13, pp. 341–379, 2003
2003
Earlier work this paper cites.
A. Stocco, C. Lebiere, and J. R. Anderson, “Conditional routing of information to the cortex: A model of the basal ganglia’s role in cognitive coordination.” Psychological Review , vol. 117, no. 2, p. 541, 2010
2010
Earlier work this paper cites.
Y. Bengio, “Deep learning of representations for unsupervised and transfer learning,” in ICML , vol. 27, 2012, pp. 17–36
2012
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in IROS , 2012, pp. 5026–5033
2012
Earlier work this paper cites.
M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, “The arcade learning environment: An evaluation platform for general agents,” Artif. Intell. Res. , vol. 47, no. 1, pp. 253–279, 2013
2013
Earlier work this paper cites.
H. Bou-Ammar, E. Eaton, P. Ruvolo, and M. E. Taylor, “Online multi-task learning for policy gradient methods,” in ICML , vol. 32, 2014, pp. 1206–1214
2014
Earlier work this paper cites.
H. Bou-Ammar, E. Eaton, J. Luna, and P. Ruvolo, “Autonomous cross-domain knowledge transfer in lifelong policy gradientreinforcement learning,” in IJCAI , 2015, pp. 3345–3351
2015
Earlier work this paper cites.
H. Bou-Ammar, R. Tutunov, and E. Eaton, “Safe policy search for lifelong reinforcement learning with sublinearregret,” in ICML , vol. 37, 2015, pp. 2361–2369
2015
Earlier work this paper cites.
V. Mnih et al. , “Human-level control through deep reinforcement learning,” Nature , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
D. Isele, M. Rostami, and E. Eaton, “Using task features for zero-shot knowledge transfer in lifelong learning,” in IJCAI , 2016, pp. 1620–1626
2016
Earlier work this paper cites.
A. A. Rusu et al. , “Progressive neural networks,” ArXiv preprint , vol. abs/1606.04671, 2016
2016
Earlier work this paper cites.
C. Beattie et al. , “Deepmind lab,” ArXiv preprint , vol. abs/1612.03801, 2016
2016
Earlier work this paper cites.
R. J. Boucherie and N. M. Van Dijk, Markov Decision Processes in Practice . Springer, 2017
2017
Earlier work this paper cites.
J. Kirkpatrick et al. , “Overcoming catastrophic forgetting in neural networks,” PNAS , vol. 114, no. 13, pp. 3521–3526, 2017
2017
Earlier work this paper cites.
D. Lopez-Paz and M. Ranzato, “Gradient episodic memory for continual learning,” in NeurIPS , vol. 30, 2017, pp. 6467–6476
2017
Earlier work this paper cites.
C. Tessler, S. Givony, T. Zahavy, D. J. Mankowitz, and S. Mannor, “A deep hierarchical approach to lifelong learning in minecraft,” in AAAI , vol. 31, no. 1, 2017, pp. 1553–1561
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Y. Zhan, H. B. Ammar, and M. E. Taylor, “Scalable lifelong reinforcement learning,” Pattern Recognit. , vol. 72, pp. 407–418, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction . MIT Press, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
C. Kaplanis, M. Shanahan, and C. Clopath, “Continual reinforcement learning with complex synapses,” in ICML , vol. 80, 2018, pp. 2502–2511
2018
Earlier work this paper cites.
Z. Li and D. Hoiem, “Learning without forgetting,” Trans. Pattern Anal. Mach. Intell. , vol. 40, no. 12, pp. 2935–2947, 2018
2018
Earlier work this paper cites.
D. Isele and A. Cosgun, “Selective experience replay for lifelong learning,” in AAAI , vol. 32, no. 1, 2018, pp. 3302–3309
2018
Earlier work this paper cites.
D. Abel, D. Arumugam, L. Lehnert, and M. L. Littman, “State abstractions for lifelong reinforcement learning,” in ICML , vol. 80, 2018, pp. 10–19
2018
Earlier work this paper cites.
D. Abel, Y. Jinnai, S. Y. Guo, G. D. Konidaris, and M. L. Littman, “Policy and value transfer in lifelong reinforcement learning,” in ICML , vol. 80, 2018, pp. 20–29
2018
Earlier work this paper cites.
J. Schwarz et al. , “Progress & compress: A scalable framework for continual learning,” in ICML , vol. 80, 2018, pp. 4535–4544
2018
Earlier work this paper cites.
J. A. Mendez, S. Shivkumar, and E. Eaton, “Lifelong inverse reinforcement learning,” in NeurIPS , vol. 31, 2018, pp. 4507–4518
2018
Earlier work this paper cites.
E. Meyerson and R. Miikkulainen, “Beyond shared hierarchies: Deep multitask learning through soft layerordering,” in ICLR , 2018, pp. 1–14
2018
Earlier work this paper cites.
A. J. Kell, D. L. Yamins, E. N. Shook, S. V. Norman-Haignere, and J. H. McDermott, “A task-optimized neural network replicates human auditory behavior, predicts brain responses, and reveals a cortical processing hierarchy,” Neuron , vol. 98, no. 3, pp. 630–644, 2018
2018
Earlier work this paper cites.
A. Mallya and S. Lazebnik, “Packnet: Adding multiple tasks to a single network by iterative pruning,” in CVPR , 2018, pp. 7765–7773
2018
Earlier work this paper cites.
L. Espeholt et al. , “IMPALA: scalable distributed deep-rl with importance weighted actor-learnerarchitectures,” in ICML , vol. 80, 2018, pp. 1406–1415
2018
Earlier work this paper cites.
C. Sherstan, M. C. Machado, and P. M. Pilarski, “Accelerating learning in constructive predictive frameworks with the successor representation,” in IROS , 2018, pp. 2997–3003
2018
Earlier work this paper cites.
G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter, “Continual lifelong learning with neural networks: A review,” Neural Networks , vol. 113, pp. 54–71, 2019
2019
Earlier work this paper cites.
B. van Niekerk, S. James, A. C. Earle, and B. Rosman, “Composing value functions in reinforcement learning,” in ICML , vol. 97, 2019, pp. 6401–6409
2019
Earlier work this paper cites.
R. Traoré et al. , “Discorl: Continual reinforcement learning via policy distillation,” in NeurIPS DRL Workshop , 2019
2019
Earlier work this paper cites.
B. Liu, L. Wang, and M. Liu, “Lifelong federated reinforcement learning: A learning architecture for navigation in cloud robotic systems,” IEEE Rob. Autom. Lett. , vol. 4, no. 4, pp. 4555–4562, 2019
2019
Earlier work this paper cites.
D. Rolnick, A. Ahuja, J. Schwarz, T. P. Lillicrap, and G. Wayne, “Experience replay for continual learning,” in NeurIPS , vol. 32, 2019, pp. 348–358
2019
Earlier work this paper cites.
F. M. Garcia and P. S. Thomas, “A meta-mdp approach to exploration for lifelong reinforcement learning,” in NeurIPS , vol. 32, 2019, pp. 5692–5701
2019
Earlier work this paper cites.
A. Nagabandi, C. Finn, and S. Levine, “Deep online learning via meta-learning: Continual adaptation for model-basedRL,” in ICLR , 2019, pp. 1–15
2019
Earlier work this paper cites.
D. Grbic and S. Risi, “Towards continual reinforcement learning through evolutionary meta-learning,” in GECCO , 2019, p. 119–120
2019
Earlier work this paper cites.
C. Kaplanis, M. Shanahan, and C. Clopath, “Policy consolidation for continual reinforcement learning,” in ICML , vol. 97, 2019, pp. 3242–3251
2019
Earlier work this paper cites.
R. Traoré, H. Caselles-Dupré, T. Lesort, T. Sun, N. Díaz-Rodríguez, and D. Filliat, “Continual reinforcement learning deployed in real-life using policy distillation and Sim2Real transfer,” in ICML MLRL Workshop , 2019
2019
Earlier work this paper cites.
C. Doyle, M. Guériau, and I. Dusparic, “Variational policy chaining for lifelong reinforcement learning,” in ICTAI , 2019, pp. 1546–1550
2019
Earlier work this paper cites.
M. Chang, A. Gupta, S. Levine, and T. L. Griffiths, “Automatically composing representation transformations as a meansfor generalization,” in ICLR , 2019, pp. 1–23
2019
Earlier work this paper cites.
A. Nagabandi et al. , “Learning to adapt in dynamic, real-world environments through meta-reinforcementlearning,” in ICLR , 2019
2019
Earlier work this paper cites.
Z. Ding and H. Dong, Challenges of Reinforcement Learning . Springer Singapore, 2020
2020
Earlier work this paper cites.
T. Lesort, V. Lomonaco, A. Stoian, D. Maltoni, D. Filliat, and N. Díaz-Rodríguez, “Continual learning for robotics: Definition, framework, learning strategies, opportunities and challenges,” Inf. Fusion , vol. 58, pp. 52–68, 2020
2020
Earlier work this paper cites.
N. Vithayathil Varghese and Q. H. Mahmoud, “A survey of multi-task deep reinforcement learning,” Electron. , vol. 9, no. 9, 2020
2020
Earlier work this paper cites.
V. Lomonaco, K. Desai, E. Culurciello, and D. Maltoni, “Continual reinforcement learning in 3d non-stationary environments,” in CVPR , 2020
2020
Earlier work this paper cites.
N. Bard et al. , “The Hanabi challenge: A new frontier for AI research,” Artif. Intell. , vol. 280, p. 103216, 2020
2020
Earlier work this paper cites.
T. Yu et al. , “Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,” in CoRL , vol. 100, 2020, pp. 1094–1100
2020
Earlier work this paper cites.
B. Wu, J. K. Gupta, and M. Kochenderfer, “Model primitives for hierarchical lifelong reinforcement learning,” in AAMAS , vol. 34, no. 1, 2020, pp. 1–38
2020
Earlier work this paper cites.
A. Xie, J. Harrison, and C. Finn, “Deep reinforcement learning amidst lifelong non-stationarity,” in ICML LML Workshop , 2020
2020
Earlier work this paper cites.
G. N. Tasse, S. James, and B. Rosman, “A boolean task algebra for reinforcement learning,” in NeurIPS , vol. 33, 2020, pp. 9497–9507
2020
Earlier work this paper cites.
J. A. Mendez, B. Wang, and E. Eaton, “Lifelong policy gradient learning of factored policies for fastertraining without forgetting,” in NeurIPS , vol. 33, 2020, pp. 14 398–14 409
2020
Earlier work this paper cites.
Y. Chandak, G. Theocharous, C. Nota, and P. S. Thomas, “Lifelong learning with a changing action set,” in AAAI , vol. 34, no. 04, 2020, pp. 3373–3380
2020
Earlier work this paper cites.
M. Xu, W. Ding, J. Zhu, Z. Liu, B. Chen, and D. Zhao, “Task-agnostic online reinforcement learning with an infinite mixtureof gaussian processes,” in NeurIPS , vol. 33, 2020, pp. 6429–6440
2020
Earlier work this paper cites.
G. N. Tasse, S. James, and B. Rosman, “Logical composition in lifelong reinforcement learning,” in ICML LML Workshop , 2020
2020
Earlier work this paper cites.
J. von Oswald, C. Henning, J. Sacramento, and B. F. Grewe, “Continual learning with hypernetworks,” in ICLR , 2020, pp. 1–28
2020
Earlier work this paper cites.
M. Wortsman et al. , “Supermasks in superposition,” in NeurIPS , vol. 33, 2020, pp. 15 173–15 184
2020
Earlier work this paper cites.
N. Wang, D. Zhang, and Y. Wang, “Learning to navigate for mobile robot with continual reinforcement learning,” in CCC , 2020, pp. 3701–3706
2020
Earlier work this paper cites.
N. Stiennon et al. , “Learning to summarize with human feedback,” in NeurIPS , vol. 33, 2020, pp. 3008–3021
2020
Earlier work this paper cites.
N. Justesen, P. Bontrager, J. Togelius, and S. Risi, “Deep learning for video game playing,” IEEE Trans. Games , vol. 12, no. 1, pp. 1–20, 2020
2020
Earlier work this paper cites.
T. Kobayashi and T. Sugino, “Reinforcement learning for quadrupedal locomotion with design of continual–hierarchical curriculum,” Eng. Appl. Artif. Intell. , vol. 95, p. 103869, 2020
2020
Earlier work this paper cites.
S. C. Raparthy, E. Hambro, R. Kirk, M. Henaff, and R. Raileanu, “Generalization to new sequential decision making tasks with in-contextlearning,” in ICML , 2024
2020
Earlier work this paper cites.
G. Dulac-Arnold et al. , “Challenges of real-world reinforcement learning: definitions, benchmarks and analysis,” Mach. Learn. , vol. 110, no. 9, pp. 2419–2468, 2021
2021
Earlier work this paper cites.
M. Wolczyk, M. Zajac, R. Pascanu, L. Kucinski, and P. Milos, “Continual world: A robotic benchmark for continual reinforcementlearning,” in NeurIPS , vol. 34, 2021, pp. 28 496–28 510
2021
Earlier work this paper cites.
H. Nekoei, A. Badrinaaraayanan, A. C. Courville, and S. Chandar, “Continuous coordination as a realistic scenario for lifelong learning,” in ICML , vol. 139, 2021, pp. 8016–8024
2021
Earlier work this paper cites.
K. Lu, A. Grover, P. Abbeel, and I. Mordatch, “Reset-free lifelong learning with skill-space planning,” in ICLR , 2021, pp. 1–20
2021
Earlier work this paper cites.
C. Atkinson, B. McCane, L. Szymanski, and A. Robins, “Pseudo-rehearsal: Achieving deep reinforcement learning without catastrophic forgetting,” Neurocomputing , vol. 428, pp. 291–307, 2021
2021
Earlier work this paper cites.
Y. Jiang, S. Bharadwaj, B. Wu, R. Shah, U. Topcu, and P. Stone, “Temporal-logic-based reward shaping for continuing reinforcement learningtasks,” in AAAI , vol. 35, no. 9, 2021, pp. 7995–8003
2021
Cited alongside, same era.
E. Lecarpentier, D. Abel, K. Asadi, Y. Jinnai, E. Rachelson, and M. L. Littman, “Lipschitz lifelong reinforcement learning,” in AAAI , vol. 35, no. 9, 2021, pp. 8270–8278
2021
Cited alongside, same era.
Y. Huang, K. Xie, H. Bharadhwaj, and F. Shkurti, “Continual model-based reinforcement learning with hypernetworks,” in ICRA , 2021, pp. 799–805
2021
Cited alongside, same era.
Y.-M. Qian, F.-Z. Xiong, and Z.-Y. Liu, “Zero-shot policy generation in lifelong reinforcement learning,” Neurocomputing , vol. 446, pp. 65–73, 2021
2021
Cited alongside, same era.
C. Li, Y. Li, Y. Zhao, P. Peng, and X. Geng, “SLER: Self-generated long-term experience replay for continual reinforcement learning,” Appl. Intell. , vol. 51, no. 1, pp. 185–201, 2021
J. Mendez-Mendez, L. P. Kaelbling, and T. Lozano-Pérez, “Embodied lifelong learning for task and motion planning,” in CoRL , vol. 229, 2023, pp. 2134–2150
2023
Later among the works it cites.
S. Rajeswar et al. , “Mastering the unsupervised reinforcement learning benchmark from pixels,” in ICML , vol. 202, 2023, pp. 28 598–28 617
2023
Later among the works it cites.
Y. Fang et al. , “Scrl: Self-supervised continual reinforcement learning for domain adaptation,” in AIoTSys , 2023, pp. 55–63
2023
Later among the works it cites.
L. Wang, X. Zhang, H. Su, and J. Zhu, “A comprehensive survey of continual learning: Theory, method and application,” Trans. Pattern Anal. Mach. Intell. , vol. 46, no. 8, pp. 5362–5383, 2024
2024
Later among the works it cites.
A. Muppidi, Z. Zhang, and H. Yang, “Fast TRAC: A parameter-free optimizer for lifelong reinforcementlearning,” in NeurIPS , 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
N. Bougie and R. Ichise, “Intrinsically motivated lifelong exploration in reinforcement learning,” in AISC , 2021, pp. 109–120
2021
Cited alongside, same era.
R. Mowakeaa, S.-J. Kim, and D. K. Emge, “Kernel-based lifelong policy gradient reinforcement learning,” in ICASSP , 2021, pp. 3500–3504
2021
Cited alongside, same era.
Y. Shi, L. Yuan, Y. Chen, and J. Feng, “Continual learning via bit-level information preserving,” in CVPR , 2021, pp. 16 674–16 683
2021
Cited alongside, same era.
H. Caselles-Dupré, M. Garcia-Ortiz, and D. Filliat, “S-trigger: Continual state representation learning via self-triggered generative replay,” in IJCNN , 2021, pp. 1–7
2021
Cited alongside, same era.
K. Chu, X. Zhu, and W. Zhu, “Accelerating lifelong reinforcement learning via reshaping rewards,” in SMC , 2021, pp. 619–624
2021
Cited alongside, same era.
J. A. Mendez and E. Eaton, “Lifelong learning of compositional structures,” in ICLR , 2021, pp. 1–25
2021
Cited alongside, same era.
S. Pateria, B. Subagdja, A.-h. Tan, and C. Quek, “Hierarchical reinforcement learning: A comprehensive survey,” ACM Comput. Surv. , vol. 54, no. 5, pp. 1–35, 2021
2021
Cited alongside, same era.
2024
Later among the works it cites.
S. Dohare, J. F. Hernandez-Garcia, Q. Lan, P. Rahman, A. R. Mahmood, and R. S. Sutton, “Loss of plasticity in deep continual learning,” Nature , vol. 632, no. 8026, pp. 768–774, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
A. Juliani and J. T. Ash, “A study of plasticity loss in on-policy deep reinforcement learning,” in NeurIPS , vol. 37, 2024, pp. 113 884–113 910
2024
Later among the works it cites.
M. Malagon, J. Ceberio, and J. A. Lozano, “Self-composing policies for scalable continual reinforcement learning,” in ICML , vol. 235, 2024, pp. 34 432–34 460
2024
Later among the works it cites.
Z. Liu, C. Du, W. S. Lee, and M. Lin, “Locality sensitive sparse encoding for learning world models online,” in ICLR , 2024, pp. 1–18
2024
Later among the works it cites.
H. Zhang et al. , “CPPO: continual learning for reinforcement learning with human feedback,” in ICLR , 2024, pp. 1–24
2024
Later among the works it cites.
T. Zhang et al. , “Dynamics-adaptive continual reinforcement learning via progressive contextualization,” IEEE Trans. Neural Networks Learn. Syst. , vol. 35, no. 10, pp. 14 588–14 602, 2024
2024
Later among the works it cites.
J. Dick, S. Nath, C. Peridis, E. Benjamin, S. Kolouri, and A. Soltoggio, “Statistical context detection for deep lifelong reinforcement learning,” in CoLLAs , 2024
2024
Later among the works it cites.
C. Zhao, J. Xu, R. Peng, X. Chen, K. Mei, and X. Lan, “Experience consistency distillation continual reinforcement learning for robotic manipulation tasks,” in ICRA , vol. 33, 2024, pp. 501–507
2024
Later among the works it cites.
M. Xu, X. Chen, and J. Wang, “Policy Correction and State-Conditioned Action Evaluation for Few-Shot Lifelong Deep Reinforcement Learning,” IEEE Trans. Neural Networks Learn. Syst. , pp. 1–15, 2024
2024
Later among the works it cites.
V. K. Chauhan, J. Zhou, P. Lu, S. Molaei, and D. A. Clifton, “A brief review of hypernetworks in deep learning,” Artif. Intell. , vol. 57, no. 9, p. 250, 2024
2024
Later among the works it cites.
M. Xu, X. Chen, and J. Wang, “Policy correction and state-conditioned action evaluation for few-shot lifelong deep reinforcement learning,” IEEE Trans. Neural Networks Learn. Syst. , pp. 1–15, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
K. Chen et al. , “Llm-assisted multi-teacher continual learning for visual question answering in robotic surgery,” in ICRA , 2024, pp. 10 772–10 778
2024
Later among the works it cites.
H. Bai et al. , “Digirl: Training in-the-wild device-control agents with autonomousreinforcement learning,” in NeurIPS , 2024
2024
Later among the works it cites.
K. Frans, S. Park, P. Abbeel, and S. Levine, “Unsupervised zero-shot reinforcement learning via functional rewardencodings,” in ICML , vol. 235, 2024, pp. 13 927–13 942
2024
Later among the works it cites.
C. Ying et al. , “PEAC: unsupervised pre-training for cross-embodiment reinforcementlearning,” in NeurIPS , vol. 37, 2024, pp. 54 632–54 669
2024
Later among the works it cites.
T. Xu, Z. Li, and Q. Ren, “Meta-reinforcement learning robust to distributional shift via performinglifelong in-context learning,” in ICML , vol. 235, 2024, pp. 55 112–55 125
2024
Later among the works it cites.
E. Meyer, A. White, and M. C. Machado, “Harnessing discrete representations for continual reinforcement learning,” in RCL , 2024
2024
Later among the works it cites.
R. Chua, A. Ghosh, C. Kaplanis, B. A. Richards, and D. Precup, “Learning successor features the simple way,” in NeurIPS , vol. 37, 2024
2024
Later among the works it cites.
J. Josifovski, S. Auddy, M. Malmir, J. Piater, A. Knoll, and N. Navarro-Guerrero, “Continual domain randomization,” in IROS , 2024, pp. 4965–4972
2024
Later among the works it cites.
T. Schmied, F. Paischer, V. Patil, M. Hofmarcher, R. Pascanu, and S. Hochreiter, “Retrieval-augmented decision transformer: External memory for in-context RL,” in CoLLAs , 2024
2024
Later among the works it cites.
J. Liu, J. Hao, Y. Ma, and S. Xia, “Unlock the cognitive generalization of deep reinforcement learningvia granular ball representation,” in ICML , vol. 235, 2024, pp. 31 062–31 079
2024
Later among the works it cites.
2024
Later among the works it cites.
Q. Delfosse, J. Blüml, B. Gregori, and K. Kersting, “Hackatari: Atari learning environments for robust and continual reinforcement learning,” in RLC IPRL Workshop , 2024
2024
Later among the works it cites.
V. Shulev and K. Sima’an, “Continual reinforcement learning for controlled text generation,” in COLING , 2024, pp. 3881–3889
2024
Later among the works it cites.
G. Zheng, S. Zhou, V. Braverman, M. A. Jacobs, and V. S. Parekh, “Selective experience replay compression using coresets for lifelong deep reinforcement learning in medical imaging,” in MIDL , 2024, pp. 1751–1764
2024
Later among the works it cites.
A. Q. Md et al. , “A novel approach for self-driving car in partially observable environment using life long reinforcement learning,” Sustainable Energy, Grids and Networks , vol. 38, p. 101356, 2024
2024
Later among the works it cites.
G. Tziafas and H. Kasaei, “Lifelong robot library learning: Bootstrapping composable and generalizable skills for embodied control with language models,” in ICRA , 2024, pp. 515–522
2024
Later among the works it cites.
G. Wang et al. , “Voyager: An open-ended embodied agent with large language models,” Trans. Mach. Learn. Res. , 2024
2024
Later among the works it cites.
B. Kim, M. Seo, and J. Choi, “Online continual learning for interactive instruction following agents,” in ICLR , 2024, pp. 1–18
2024
Later among the works it cites.
S. E. Ada and E. Ugur, “Unsupervised meta-testing with conditional neural processes for hybrid meta-reinforcement learning,” IEEE Rob. Autom. Lett. , vol. 9, no. 10, pp. 8427–8434, 2024
2024
Later among the works it cites.
Z. Wang, E. Yang, L. Shen, and H. Huang, “A comprehensive survey of forgetting in deep learning beyond continual learning,” Trans. Pattern Anal. Mach. Intell. , vol. 47, no. 3, pp. 1464–1483, 2025
2025
Closest in time.
C. Lyle, G. Sokar, R. Pascanu, and A. Gyorgy, “What can grokking teach us about learning under nonstationarity?” in CoLLAs , 2025
2025
Closest in time.
R. Rubavicius, P. D. Fagan, A. Lascarides, and S. Ramamoorthy, “Secure: Semantics-aware embodied conversation under unawareness for lifelong robot learning,” in CoLLAs , 2025
2025
Closest in time.
2025
Closest in time.
E. Elelimy, D. Szepesvari, M. White, and M. Bowling, “Rethinking the foundations for continual reinforcement learning,” in RCL , 2025
2025
Closest in time.
G. Mesbahi, P. M. Panahi, O. Mastikhina, S. Tang, M. White, and A. White, “Position: Lifetime tuning is incompatible with continual reinforcement learning,” in ICML Position Paper Track , 2025
2025
Closest in time.
A. Kobanda, O.-A. Maillard, and R. Portelas, “A continual offline reinforcement learning benchmark for navigation tasks,” in CoG , 2025, pp. 1–8
2025
Closest in time.
C. Yao et al. , “From general relation patterns to task-specific decision-making in continual multi-agent coordination,” in IJCAI , vol. 1, 2025, pp. 6821–6829
2025
Closest in time.
C. Pan et al. , “Multi-granularity knowledge transfer for continual reinforcement learning,” in IJCAI , 2025
2025
Closest in time.
H. Tang, J. Obando-Ceron, P. S. Castro, A. Courville, and G. Berseth, “Mitigating plasticity loss in continual reinforcement learning by reducing churn,” in ICML , 2025
2025
Closest in time.
Y. Meng et al. , “Preserving and combining knowledge in robotic lifelong reinforcement learning,” Mach. Intell. , vol. 7, no. 2, pp. 256–269, 2025
2025
Closest in time.
N. D. Palo, L. Hasenclever, J. Humplik, and A. Byravan, “Diffusion augmented agents: A framework for efficient exploration and transfer learning,” in CoLLAs , vol. 274, 2025, pp. 268–284
2025
Closest in time.
L. Yuan, L. Li, Z. Zhang, F. Zhang, C. Guan, and Y. Yu, “Multiagent continual coordination via progressive task contextualization,” IEEE Trans. Neural Networks Learn. Syst. , vol. 36, no. 4, pp. 6326–6340, 2025
2025
Closest in time.
H. Fu, Y. Sun, M. Littman, and G. Konidaris, “Knowledge retention in continual model-based reinforcement learning,” in ICML , 2025
2025
Closest in time.
E. Piccoli, M. Li, G. Carfı, V. Lomonaco, and D. Bacciu, “Combining pre-trained models for enhanced feature representation in reinforcement learning,” in CoLLAs , 2025
2025
Closest in time.
R. Surdej, M. Bortkiewicz, A. Lewandowski, M. Ostaszewski, and C. Lyle, “Balancing expressivity and robustness: Constrained rational activations for reinforcement learning,” in CoLLAs , 2025
2025
Closest in time.
J. Hu et al. , “Tackling continual offline RL through selective weights activation on aligned spaces,” in NeurIPS , 2025
2025
Closest in time.
——, “Continual knowledge adaptation for reinforcement learning,” in NeurIPS , 2025
2025
Closest in time.
A. Jaziri, “Mitigating the stability-plasticity dilemma in adaptive train scheduling with curriculum-driven continual dqn expansion,” in CoLLAs , 2025
2025
Closest in time.
W. Yue, B. Liu, and P. Stone, “T-DGR: A trajectory-based deep generative replay method for continual learning in decision making,” in CoLLAs , 2025, pp. 481–497
2025
Closest in time.
D. Guo et al. , “DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning,” Nature , vol. 645, no. 8081, pp. 633–638, 2025
2025
Closest in time.
J. H. et al
2025
Closest in time.
A. Zhao, E. Zhu, R. Lu, M. Lin, Y.-J. Liu, and G. Huang, “Self-referencing agents for unsupervised reinforcement learning,” Neural Networks , vol. 188, p. 107448, 2025
2025
Closest in time.
K. Garces, J. Xuan, and H. Zuo, “Adaptive successor features composition for transfer reinforcement learning,” IEEE Transactions on Artificial Intelligence , pp. 1–12, 2025
2025
Closest in time.
H. Ahn, J. Hyeon, Y. Oh, B. Hwang, and T. Moon, “Prevalence of negative transfer in continual reinforcement learning: Analyses and a simple baseline,” in ICLR , 2025
2025
Closest in time.
N. Botteghi, M. Poel, and C. Brune, “Unsupervised representation learning in deep reinforcement learning: A review,” IEEE Control Systems , vol. 45, no. 2, pp. 26–68, 2025
2025
Closest in time.
2025
Closest in time.
J. Wang, R. Chandra, and S. Zhang, “Towards provable emergence of in-context reinforcement learning,” in NeurIPS , 2025
2025
Closest in time.
X. Liu, Y. Bai, Y. Lu, A. Soltoggio, and S. Kolouri, “Wasserstein task embedding for measuring task similarities,” Neural Networks , vol. 181, p. 106796, 2025
2025
Closest in time.
2025
Closest in time.
X. Zeng, H. Luo, Z. Wang, S. Li, Z. Shen, and T. Li, “A continual learning approach for embodied question answering with generative adversarial imitation learning,” in ICASSP , 2025, pp. 1–5
2025
Closest in time.
J. Beck et al. , “A tutorial on meta-reinforcement learning,” Foundations and Trends in Machine Learning , vol. 18, no. 2-3, pp. 224–384, 2025
2025
Closest in time.
K. Sun et al. , “Principled fast and meta knowledge learners for continual reinforcement learning,” in ICLR , 2026
2026
Closest in time.
K. Cobbe, C. Hesse, J. Hilton, and J. Schulman, “Leveraging procedural generation to benchmark reinforcement learning,” in ICML , vol. 119, 2020, pp. 2048–2056
2056
Closest in time.
L. Chen, S. Jayanthi, R. R. Paleja, D. Martin, V. Zakharov, and M. Gombolay, “Fast lifelong adaptive inverse reinforcement learning from demonstrations,” in CoRL , vol. 205, 2023, pp. 2083–2094
2094
Closest in time.