Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) has been successful in various domains like robotics, game playing, and simulation.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Simple, scalable adaptation for neural machine translation
Bapna, A., Arivazhagan, N., and Firat, O. (2019) · 1909
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
Berner, C., Brockman, G., Chan, B., Cheung, V., Dkebiak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al. (2019) · 1912
Earlier work this paper cites.
Reinforcement learning upside down: Don’t predict rewards–just map them to actions
Schmidhuber, J. (2019) · 1912
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
McCloskey, M. and Cohen, N. (1989) · 1989
Earlier work this paper cites.
Multitask learning
Caruana, R. (1997) · 1997
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Multitask reinforcement learning on the distribution of mdps
Tanaka, F. and Yamamura, M. (2003) · 2003
Earlier work this paper cites.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S. (2020) · 2004
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G. (2008) · 2008
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y. (2012) · 2012
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Maas, A. L., Hannun, A. Y., Ng, A. Y., et al. (2013) · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2015) · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016) · 2016
Earlier work this paper cites.
Learning shared representations in multi-task reinforcement learning
Borsa, D., Graepel, T., and Shawe-Taylor, J. (2016) · 2016
Earlier work this paper cites.
Policy distillation
Rusu, A. A., Colmenarejo, S. G., Gülçehre, Ç., Desjardins, G., Kirkpatrick, J., Pascanu, R., Mnih, V., Kavukcuoglu, K., and Hadsell, R. (2016a) · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., Dieleman, S., Grewe, D., Nham, J., Kalchbrenner, N., Sutskever, I., Lillicrap, T. P., Leach, M., Kavukcuoglu, K., Graepel, T., and Hassabis, D. (2016) · 2016
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
Andreas, J., Klein, D., and Levine, S. (2017) · 2017
Earlier work this paper cites.
Scalable multitask policy gradient reinforcement learning
El Bsat, S., Bou Ammar, H., and Taylor, M. (2017) · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., et al. (2017) · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
Lopez-Paz, D. and Ranzato, M. (2017) · 2017
Earlier work this paper cites.
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., et al. (2017) · 2017
Earlier work this paper cites.
EPOpt: Learning robust neural network policies using model ensembles
Rajeswaran, A., Ghotra, S., Ravindran, B., and Levine, S. (2017) · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
Rebuffi, S.-A., Bilen, H., and Vedaldi, A. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, l., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Continual learning through synaptic intelligence
Zenke, F., Poole, B., and Ganguli, S. (2017) · 2017
Earlier work this paper cites.
Memory aware synapses: Learning what (not) to forget
Aljundi, R., Babiloni, F., Elhoseiny, M., Rohrbach, M., and Tuytelaars, T. (2018) · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2018) · 2018
Earlier work this paper cites.
Packnet: Adding multiple tasks to a single network by iterative pruning
Mallya, A. and Lazebnik, S. (2018) · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al. (2018) · 2018
Earlier work this paper cites.
Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T. P., and Riedmiller, M. A. (2018) · 2018
Earlier work this paper cites.
Rudder: Return decomposition for delayed rewards
Arjona-Medina, J. A., Gillhofer, M., Widrich, M., Unterthiner, T., Brandstetter, J., and Hochreiter, S. (2019) · 2019
Earlier work this paper cites.
Efficient lifelong learning with A-GEM
Chaudhry, A., Ranzato, M., Rohrbach, M., and Elhoseiny, M. (2019) · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K. (2019) · 2019
Cited alongside, same era.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J. (2019) · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. (2019) · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. (2019) · 2019
Cited alongside, same era.
Grandmaster level in starcraft II using multi-agent reinforcement learning
Uni [mask]: Unified inference in sequential decision problems
Carroll, M., Paradise, O., Lin, J., Georgescu, R., Sun, M., Bignell, D., Milani, S., Hofmann, K., Hausknecht, M., Dragan, A., et al. (2022) · 2022
Later among the works it cites.
Magnetic control of tokamak plasmas through deep reinforcement learning
Degrave, J., Felici, F., Buchli, J., Neunert, M., Tracey, B., Carpanese, F., Ewalds, T., Hafner, R., Abdolmaleki, A., de Las Casas, D., et al. (2022) · 2022
Later among the works it cites.
Cloob: Modern hopfield networks with infoloob outperform clip
Fürst, A., Rumetshofer, E., Lehner, J., Tran, V., Tang, F., Ramsauer, H., Kreil, D., Kopp, M., Klambauer, G., Bitto-Nemling, A., and Hochreiter, S. (2022) · 2022
Later among the works it cites.
Metamorph: Learning universal controllers with transformers
Gupta, A., Fan, L., Ganguli, S., and Fei-Fei, L. (2022) · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R. B. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., Oh, J., Horgan, D., Kroiss, M., Danihelka, I., Huang, A., Sifre, L., Cai, T., Agapiou, J. P., Jaderberg, M., Vezhnevets, A. S., Leblond, R., Pohlen, T., Dalibard, V., Budden, D., Sulsky, Y., Molloy, J., Paine, T. L., Gülçehre, Ç., Wang, Z., Pfaff, T., Wu, Y., Ring, R., Yogatama, D., Wünsch, D., McKinney, K., Smith, O., Schaul, T., Lillicrap, T. P., Kavukcuoglu, K., Hassabis, D., Apps, C., and Silver, D. (2019) · 2019
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, Y., Mohamed, A., and Auli, M. (2020) · 2020
Cited alongside, same era.
Autonomous navigation of stratospheric balloons using reinforcement learning
Bellemare, M. G., Candido, S., Castro, P. S., Gong, J., Machado, M. C., Moitra, S., Ponda, S. S., and Wang, Z. (2020) · 2020
Cited alongside, same era.
Sharing knowledge in multi-task deep reinforcement learning
D’Eramo, C., Tateo, D., Bonarini, A., Restelli, M., and Peters, J. (2020) · 2020
Cited alongside, same era.
Embracing change: Continual learning in deep neural networks
Hadsell, R., Rao, D., Rusu, A. A., and Pascanu, R. (2020) · 2020
Cited alongside, same era.
Multitask soft option learning
Igl, M., Gambardella, A., He, J., Nardelli, N., Siddharth, N., Boehmer, W., and Whiteson, S. (2020) · 2020
Cited alongside, same era.
Continual learning for robotics: Definition, framework, learning strategies, opportunities and challenges
Lesort, T., Lomonaco, V., Stoian, A., Maltoni, D., Filliat, D., and Díaz-Rodríguez, N. (2020) · 2020
Cited alongside, same era.
Ho, J., Chan, W., Saharia, C., Whang, J., Gao, R., Gritsenko, A. A., Kingma, D. P., Poole, B., Norouzi, M., Fleet, D. J., and Salimans, T. (2022) · 2022
Later among the works it cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2022) · 2022
Later among the works it cites.
Vima: General robot manipulation with multimodal prompts
Jiang, Y., Gupta, A., Zhang, Z., Wang, G., Dou, Y., Chen, Y., Fei-Fei, L., Anandkumar, A., Zhu, Y., and Fan, L. (2022) · 2022
Later among the works it cites.
Towards continual reinforcement learning: A review and perspectives
Khetarpal, K., Riemer, M., Rish, I., and Precup, D. (2022) · 2022
Later among the works it cites.
In-context reinforcement learning with algorithm distillation
Laskin, M., Wang, L., Oh, J., Parisotto, E., Spencer, S., Steigerwald, R., Strouse, D., Hansen, S., Filos, A., Brooks, E., et al. (2022) · 2022
Later among the works it cites.
Multi-game decision transformers
Lee, K.-H., Nachum, O., Yang, M., Lee, L., Freeman, D., Xu, W., Guadarrama, S., Fischer, I., Jang, E., Michalewski, H., et al. (2022) · 2022
Later among the works it cites.
On the effectiveness of fine-tuning versus meta-reinforcement learning
Mandi, Z., Abbeel, P., and James, S. (2022) · 2022
Later among the works it cites.
Unipelt: A unified framework for parameter-efficient language model tuning
Mao, Y., Mathias, L., Hou, R., Almahairi, A., Ma, H., Han, J., Yih, S., and Khabsa, M. (2022) · 2022
Later among the works it cites.
History compression via language models in reinforcement learning
Paischer, F., Adler, T., Patil, V., Bitto-Nemling, A., Holzleitner, M., Lehner, S., Eghbal-Zadeh, H., and Hochreiter, S. (2022) · 2022
Later among the works it cites.
Robust speech recognition via large-scale weak supervision
Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I. (2022) · 2022
Later among the works it cites.
Reed, S. E., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., Eccles, T., Bruce, J., Razavi, A., Edwards, A., Heess, N., Chen, Y., Hadsell, R., Vinyals, O., Bordbar, M., and de Freitas, N. (2022) · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022) · 2022
Later among the works it cites.
A dataset perspective on offline reinforcement learning
Schweighofer, K., Dinu, M.-c., Radler, A., Hofmarcher, M., Patil, V. P., Bitto-Nemling, A., Eghbal-zadeh, H., and Hochreiter, S. (2022) · 2022
Later among the works it cites.
Starformer: Transformer with state-action-reward representations for visual reinforcement learning
Shang, J., Kahatapitiya, K., Li, X., and Ryoo, M. S. (2022) · 2022
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D. (2022) · 2022
Later among the works it cites.
How crucial is transformer in decision transformer?
Siebenborn, M., Belousov, B., Huang, J., and Peters, J. (2022) · 2022
Later among the works it cites.
Make-a-video: Text-to-video generation without text-video data
Singer, U., Polyak, A., Hayes, T., Yin, X., An, J., Zhang, S., Hu, Q., Yang, H., Ashual, O., Gafni, O., et al. (2022) · 2022
Later among the works it cites.
Coda-prompt: Continual decomposed attention-based prompting for rehearsal-free continual learning
Smith, J. S., Karlinsky, L., Gutta, V., Cascante-Bonilla, P., Kim, D., Arbelle, A., Panda, R., Feris, R., and Kira, Z. (2022) · 2022
Later among the works it cites.
Reactive exploration to cope with non-stationarity in lifelong reinforcement learning
Steinparz, C. A., Schmied, T., Paischer, F., Dinu, M.-C., Patil, V. P., Bitto-Nemling, A., Eghbal-zadeh, H., and Hochreiter, S. (2022) · 2022
Later among the works it cites.
Smart: Self-supervised multi-task pretraining with control transformers
Sun, Y., Ma, S., Madaan, R., Bonatti, R., Huang, F., and Kapoor, A. (2022) · 2022
Later among the works it cites.
Dualprompt: Complementary prompting for rehearsal-free continual learning
Wang, Z., Zhang, Z., Ebrahimi, S., Sun, R., Zhang, H., Lee, C., Ren, X., Su, G., Perot, V., Dy, J. G., and Pfister, T. (2022b) · 2022
Later among the works it cites.
Disentangling transfer in continual reinforcement learning
Wolczyk, M., Zajkac, M., Pascanu, R., Kuciński, L., and Miloś, P. (2022) · 2022
Later among the works it cites.
Online decision transformer
Zheng, Q., Zhang, A., and Grover, A. (2022) · 2022
Later among the works it cites.
A domain-agnostic approach for characterization of lifelong learning systems
Baker, M. M., New, A., Aguilar-Simon, M., Al-Halah, Z., Arnold, S. M., Ben-Iwhiwhu, E., Brna, A. P., Brooks, E., Brown, R. C., Daniels, Z., et al. (2023) · 2023
Closest in time.
A survey on transformers in reinforcement learning
Li, W., Luo, H., Lin, Z., Zhang, C., Lu, Z., and Ye, D. (2023) · 2023
Closest in time.
Emergent agentic transformer from chain of hindsight experience
Liu, H. and Abbeel, P. (2023) · 2023
Closest in time.
Semantic HELM: an interpretable memory for reinforcement learning
Paischer, F., Adler, T., Hofmarcher, M., and Hochreiter, S. (2023) · 2023
Closest in time.
Progressive prompts: Continual learning for language models
Razdaibiedina, A., Mao, Y., Hou, R., Khabsa, M., Lewis, M., and Almahairi, A. (2023) · 2023
Closest in time.
Foundation models for decision making: Problems, methods, and opportunities
Yang, S., Nachum, O., Du, Y., Wei, J., Abbeel, P., and Schuurmans, D. (2023) · 2023
Closest in time.