Fetching the paper…
Reading the bibliography…
Recently, it has been shown that transformers pre-trained on diverse datasets with multi-episode contexts can generalize to new reinforcement learning tasks in-context.
Distance metric learning for large margin nearest neighbor classification
Weinberger, K. Q. and Saul, L. K · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E · 2010
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2013
Earlier work this paper cites.
Deep reinforcement learning in large discrete action spaces
Dulac-Arnold, G., Evans, R., van Hasselt, H., Sunehag, P., Lillicrap, T., Hunt, J., Mann, T., Weber, T., Degris, T., and Coppin, B · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Schroff, F., Kalenichenko, D., and Philbin, J · 2015
Earlier work this paper cites.
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Towards generalization and simplicity in continuous control
Rajeswaran, A., Lowrey, K., Todorov, E. V., and Kakade, S. M · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Minimalistic gridworld environment for openai gym
Chevalier-Boisvert, M., Willems, L., and Pal, S · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Earlier work this paper cites.
A dissection of overfitting and generalization in continuous reinforcement learning
Zhang, A., Ballas, N., and Pineau, J · 2018
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Earlier work this paper cites.
Experiment tracking with weights and biases, 2020
Biewald, L · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Lifelong learning with a changing action set
Chandak, Y., Theocharous, G., Nota, C., and Thomas, P · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
Generalization to new actions in reinforcement learning
Jain, A., Szot, A., and Lim, J. J · 2020
Cited alongside, same era.
A survey on contrastive self-supervised learning
Jaiswal, A., Babu, A. R., Zadeh, M. Z., Banerjee, D., and Makedon, F · 2020
Cited alongside, same era.
Multi-game decision transformers
Lee, K.-H., Nachum, O., Yang, M. S., Lee, L., Freeman, D., Guadarrama, S., Fischer, I., Xu, W., Jang, E., Michalewski, H., et al · 2022
Later among the works it cites.
Lin, Q., Liu, H., and Sengupta, B · 2022
Later among the works it cites.
Prompting decision transformer for few-shot policy generalization
Xu, M., Shen, Y., Zhang, S., Lu, Y., Zhao, D., Tenenbaum, J., and Gan, C · 2022
Later among the works it cites.
Headless language models: Learning without predicting with contrastive weight tying
Godey, N., de la Clergerie, É., and Sagot, B · 2023
Closest in time.
Towards general-purpose in-context learning agents
Kirsch, L., Harrison, J., Freeman, C. D., Sohl-Dickstein, J., and Schmidhuber, J · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li, L., Yang, R., and Luo, D · 2020
Cited alongside, same era.
Offline policy evaluation with new arms
London, B. and Joachims, T · 2020
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Cited alongside, same era.
Offline meta reinforcement learning–identifiability challenges and effective data collection strategies
Dorfman, R., Shenfeld, I., and Tamar, A · 2021
Cited alongside, same era.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
Gu, A., Goel, K., and Ré, C · 2021
Cited alongside, same era.
Know your action set: Learning action relations for reinforcement learning
Jain, A., Kosaka, N., Kim, K.-M., and Lim, J. J · 2021
Cited alongside, same era.
Lee, J. N., Xie, A., Pacchiano, A., Chandak, Y., Finn, C., Nachum, O., and Brunskill, E · 2023
Closest in time.
A survey on transformers in reinforcement learning
Li, W., Luo, H., Lin, Z., Zhang, C., Lu, Z., and Ye, D · 2023
Closest in time.
Lin, L., Bai, Y., and Mei, S · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G · 2023
Closest in time.
Structured state space models for in-context reinforcement learning
Lu, C., Schroecker, Y., Gu, A., Parisotto, E., Foerster, J., Singh, S., and Behbahani, F · 2023
Closest in time.
Large language models as general pattern machines
Mirchandani, S., Xia, F., Florence, P., Ichter, B., Driess, D., Arenas, M. G., Rao, K., Sadigh, D., and Zeng, A · 2023
Closest in time.
XLand-minigrid: Scalable meta-reinforcement learning environments in JAX
Nikulin, A., Kurenkov, V., Zisman, I., Sinii, V., Agarkov, A., and Kolesnikov, S · 2023
Closest in time.
Generalization to new sequential decision making tasks with in-context learning
Raparthy, S. C., Hambro, E., Kirk, R., Henaff, M., and Raileanu, R · 2023
Closest in time.
Images speak in images: A generalist painter for in-context visual learning
Wang, X., Wang, W., Cao, Y., Shen, C., and Huang, T · 2023
Closest in time.
Action pick-up in dynamic action space reinforcement learning
Ye, J., Li, X., Wu, P., and Wang, F · 2023
Closest in time.
Emergence of in-context reinforcement learning from noise distillation
Zisman, I., Kurenkov, V., Nikulin, A., Sinii, V., and Kolesnikov, S · 2023
Closest in time.
Transformers learn temporal difference methods for in-context reinforcement learning, 2024
Wang, J., Blaser, E., Daneshmand, H., and Zhang, S · 2024
Closest in time.
Tinyllama: An open-source small language model, 2024
Zhang, P., Zeng, G., Wang, T., and Lu, W · 2024
Closest in time.