Fetching the paper…
Reading the bibliography…
Fine-tuning is a widespread technique that allows practitioners to transfer pre-trained capabilities, as recently showcased by the successful applications of foundation models.
Reinforcement learning for robots using neural networks
Lin, L.-J · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
A framework for behavioural cloning
Bain, M. and Sammut, C · 1995
Earlier work this paper cites.
Measuring statistical dependence with hilbert-schmidt norms
Gretton, A., Bousquet, O., Smola, A., and Schölkopf, B · 2005
Earlier work this paper cites.
Efficient reductions for imitation learning
Ross, S. and Bagnell, D · 2010
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M · 2013
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Girshick, R., Donahue, J., Darrell, T., and Malik, J · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Yosinski, J., Clune, J., Bengio, Y., and Lipson, H · 2014
Earlier work this paper cites.
Actor-mimic: Deep multitask and transfer reinforcement learning
Parisotto, E., Ba, J. L., and Salakhutdinov, R · 2015
Earlier work this paper cites.
Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., and Hadsell, R · 2016
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al · 2017
Earlier work this paper cites.
Time limits in reinforcement learning
Pardo, F., Tavakoli, A., Levdik, V., and Kormushev, P · 2017
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
Rebuffi, S.-A., Kolesnikov, A., Sperl, G., and Lampert, C. H · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Memory aware synapses: Learning what (not) to forget
Aljundi, R., Babiloni, F., Elhoseiny, M., Rohrbach, M., and Tuytelaars, T · 2018
Earlier work this paper cites.
Exploration by random network distillation
Burda, Y., Edwards, H., Storkey, A., and Klimov, O · 2018
Earlier work this paper cites.
Measuring catastrophic forgetting in neural networks
Kemker, R., McClure, M., Abitino, A., Hayes, T., and Kanan, C · 2018
Earlier work this paper cites.
Packnet: Adding multiple tasks to a single network by iterative pruning
Mallya, A. and Lazebnik, S · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Kickstarting deep reinforcement learning
Schmitt, S., Hudson, J. J., Zidek, A., Osindero, S., Doersch, C., Czarnecki, W. M., Leibo, J. Z., Kuttler, H., Zisserman, A., Simonyan, K., et al · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
On tiny episodic memories in continual learning
Chaudhry, A., Rohrbach, M., Elhoseiny, M., Ajanthan, T., Dokania, P. K., Torr, P. H., and Ranzato, M · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Similarity of neural network representations revisited
Kornblith, S., Norouzi, M., Lee, H., and Hinton, G · 2019
Earlier work this paper cites.
Experience replay for continual learning
Rolnick, D., Ahuja, A., Schwarz, J., Lillicrap, T., and Wayne, G · 2019
Earlier work this paper cites.
Ray interference: a source of plateaus in deep reinforcement learning, 2019
Schaul, T., Borsa, D., Modayil, J., and Pascanu, R · 2019
Earlier work this paper cites.
Discorl: Continual reinforcement learning via policy distillation
Traoré, R., Caselles-Dupré, H., Lesort, T., Sun, T., Cai, G., Díaz-Rodríguez, N., and Filliat, D · 2019
Earlier work this paper cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Earlier work this paper cites.
Chemberta: Large-scale self-supervised pretraining for molecular property prediction
Chithrananda, S., Grand, G., and Ramsundar, B · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
The nethack learning environment
Küttler, H., Nardelli, N., Miller, A., Raileanu, R., Selvatici, M., Grefenstette, E., and Rocktäschel, T · 2020
Cited alongside, same era.
Awac: Accelerating online reinforcement learning with offline datasets
Nair, A., Gupta, A., Dalal, M., and Levine, S · 2020
Cited alongside, same era.
What is being transferred in transfer learning?
Neyshabur, B., Sedghi, H., and Zhang, C · 2020
Cited alongside, same era.
Petrenko, A., Huang, Z., Kumar, T., Sukhatme, G. S., and Koltun, V · 2020
Building a subspace of policies for scalable continual learning
Gaya, J.-B., Doan, T., Caccia, L., Soulier, L., Denoyer, L., and Raileanu, R · 2022
Later among the works it cites.
Towards continual reinforcement learning: A review and perspectives
Khetarpal, K., Riemer, M., Rish, I., and Precup, D · 2022
Later among the works it cites.
Pre-training for robots: Offline rl enables learning new tasks from a handful of trials
Kumar, A., Singh, A., Ebert, F., Yang, Y., Finn, C., and Levine, S · 2022
Later among the works it cites.
How to spend your robot time: Bridging kickstarting and offline reinforcement learning for vision-based robotic manipulation
Lee, A. X., Devin, C., Springenberg, J. T., Zhou, Y., Lampe, T., Abdolmaleki, A., and Bousmalis, K · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Anatomy of catastrophic forgetting: Hidden representations and task semantics
Ramasesh, V. V., Dyer, E., and Raghu, M · 2020
Cited alongside, same era.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S · 2020
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Cited alongside, same era.
Rethinking experience replay: a bag of tricks for continual learning
Buzzega, P., Boschini, M., Porrello, A., and Calderara, S · 2021
Cited alongside, same era.
Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research
Ceron, J. S. O. and Castro, P. S · 2021
Cited alongside, same era.
A continual learning survey: Defying forgetting in classification tasks
De Lange, M., Aljundi, R., Masana, M., Parisot, S., Jia, X., Leonardis, A., Slabaugh, G., and Tuytelaars, T · 2021
Cited alongside, same era.
Insights from the neurips 2021 nethack challenge
Hambro, E., Mohanty, S., Babaev, D., Byeon, M., Chakraborty, D., Grefenstette, E., Jiang, M., Daejin, J., Kanervisto, A., Kim, J., et al · 2021
Cited alongside, same era.
Lesort, T., Ostapenko, O., Misra, D., Arefin, M. R., Rodríguez, P., Charlin, L., and Rish, I · 2022
Later among the works it cites.
On the effectiveness of fine-tuning versus meta-reinforcement learning
Mandi, Z., Abbeel, P., and James, S · 2022
Later among the works it cites.
Modular lifelong reinforcement learning via neural composition
Mendez, J. A., van Seijen, H., and Eaton, E · 2022
Later among the works it cites.
Architecture matters in continual learning
Mirzadeh, S. I., Chaudhry, A., Yin, D., Nguyen, T., Pascanu, R., Gorur, D., and Farajtabar, M · 2022
Later among the works it cites.
Improving intrinsic exploration with language abstractions
Mu, J., Zhong, V., Raileanu, R., Jiang, M., Goodman, N., Rocktäschel, T., and Grefenstette, E · 2022
Later among the works it cites.
Cora: Benchmarks, baselines, and metrics as a platform for continual reinforcement learning agents
Powers, S., Xing, E., Kolve, E., Mottaghi, R., and Gupta, A · 2022
Later among the works it cites.
Effect of scale on catastrophic forgetting in neural networks
Ramasesh, V. V., Lewkowycz, A., and Dyer, E · 2022
Later among the works it cites.
Probing transfer in deep reinforcement learning without task engineering
Rusu, A. A., Flennerhag, S., Rao, D., Pascanu, R., and Hadsell, R · 2022
Later among the works it cites.
Fine-tuning image transformers using learnable memory
Sandler, M., Zhmoginov, A., Vladymyrov, M., and Jackson, A · 2022
Later among the works it cites.
Reinforcement learning with action-free pre-training from videos
Seo, Y., Lee, K., James, S. L., and Abbeel, P · 2022
Later among the works it cites.
Disentangling transfer in continual reinforcement learning
Wolczyk, M., Zając, M., Pascanu, R., Kuciński, Ł., and Miłoś, P · 2022
Later among the works it cites.
Bigssl: Exploring the frontier of large-scale semi-supervised learning for automatic speech recognition
Zhang, Y., Park, D. S., Han, W., Qin, J., Gulati, A., Shor, J., Jansen, A., Xu, Y., Huang, Y., Wang, S., et al · 2022
Later among the works it cites.
Efficient online reinforcement learning with offline data
Ball, P. J., Smith, L., Kostrikov, I., and Levine, S · 2023
Later among the works it cites.
Language acquisition: do children and language models follow similar learning stages?
Evanson, L., Lakretz, Y., and King, J.-R · 2023
Later among the works it cites.
The effectiveness of world models for continual reinforcement learning, 2023
Kessler, S., Ostaszewski, M., Bortkiewicz, M., Żarski, M., Wołczyk, M., Parker-Holder, J., Roberts, S. J., and Miłoś, P · 2023
Later among the works it cites.
Motif: Intrinsic motivation from artificial intelligence feedback
Klissarov, M., D’Oro, P., Sodhani, S., Raileanu, R., Bacon, P.-L., Vincent, P., Zhang, A., and Henaff, M · 2023
Later among the works it cites.
An empirical study of catastrophic forgetting in large language models during continual fine-tuning
Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., and Zhang, Y · 2023
Later among the works it cites.
NetHack Home Page
NetHack DevTeam · 2023
Later among the works it cites.
Piterbarg, U., Pinto, L., and Fergus, R · 2023
Later among the works it cites.
Scaling laws for imitation learning in nethack
Tuyls, J., Madeka, D., Torkkola, K., Foster, D., Narasimhan, K., and Kakade, S · 2023
Later among the works it cites.
Foundations for transfer in reinforcement learning: A taxonomy of knowledge modalities, 2023
Wulfmeier, M., Byravan, A., Bechtle, S., Hausman, K., and Heess, N · 2023
Later among the works it cites.
Xu, L., Xie, H., Qin, S.-Z. J., Tao, X., and Wang, F. L · 2023
Later among the works it cites.
What is essential for unseen goal generalization of offline goal-conditioned rl?
Yang, R., Yong, L., Ma, X., Hu, H., Zhang, C., and Zhang, T · 2023
Later among the works it cites.
Adaptive policy learning for offline-to-online reinforcement learning
Zheng, H., Luo, X., Wei, P., Song, X., Li, D., and Jiang, J · 2023
Later among the works it cites.
Evron, I., Goldfarb, D., Weinberger, N., Soudry, D., and Hand, P · 2024
Closest in time.
Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning
Nakamoto, M., Zhai, S., Singh, A., Sobol Mark, M., Ma, Y., Finn, C., Kumar, A., and Levine, S · 2024
Closest in time.