Fetching the paper…
Reading the bibliography…
Recently, extensive studies in Reinforcement Learning have been carried out on the ability of transformers to adapt in-context to various environments and tasks.
Spatial localization does not require the presence of local cues
Morris, R. G · 1981
Earlier work this paper cites.
Learning to learn using gradient descent
Hochreiter, S., Younger, A. S., and Conwell, P. R · 2001
Earlier work this paper cites.
(more) efficient reinforcement learning via posterior sampling
Osband, I., Russo, D., and Van Roy, B · 2013
Earlier work this paper cites.
Deepmind lab
Beattie, C., Leibo, J. Z., Teplyashin, D., Ward, T., Wainwright, M., Küttler, H., Lefrancq, A., Green, S., Valdés, V., Sadik, A., Schrittwieser, J., Anderson, K., York, S., Cant, M., Cain, A., Bolton, A., Gaffney, S., King, H., Hassabis, D., Legg, S., and Petersen, S · 2016
Earlier work this paper cites.
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Duan, Y., Schulman, J., Chen, X., Bartlett, P. L., Sutskever, I., and Abbeel, P · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Wang, J. X., Kurth-Nelson, Z., Tirumala, D., Soyer, H., Leibo, J. Z., Munos, R., Blundell, C., Kumaran, D., and Botvinick, M · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Vinyals, O., Babuschkin, I., Czarnecki, W. M., Mathieu, M., Dudzik, A., Chung, J., Choi, D. H., Powell, R., Ewalds, T., Georgiev, P., et al · 2019
Earlier work this paper cites.
Better-than-demonstrator imitation learning via automatically-ranked demonstrations
Brown, D. S., Goo, W., and Niekum, S · 2020
Earlier work this paper cites.
The nethack learning environment
Küttler, H., Nardelli, N., Miller, A. H., Raileanu, R., Selvatici, M., Grefenstette, E., and Rocktäschel, T · 2020
Earlier work this paper cites.
Mastering atari, go, chess and shogi by planning with a learned model
Schrittwieser, J., Antonoglou, I., Hubert, T., Simonyan, K., Sifre, L., Schmitt, S., Guez, A., Lockhart, E., Hassabis, D., Graepel, T., et al · 2020
Earlier work this paper cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Earlier work this paper cites.
Super-human performance in gran turismo sport using deep reinforcement learning
Fuchs, F., Song, Y., Kaufmann, E., Scaramuzza, D., and Dürr, P · 2021
Earlier work this paper cites.
Benchmarking the spectrum of agent capabilities
Hafner, D · 2021
Earlier work this paper cites.
Offline reinforcement learning as one big sequence modeling problem
Janner, M., Li, Q., and Levine, S · 2021
Cited alongside, same era.
Transformers can do bayesian inference
Muller, S., Hollmann, N., Arango, S. P., Grabocka, J., and Hutter, F · 2021
Cited alongside, same era.
Stable-baselines3: Reliable reinforcement learning implementations
Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., and Dormann, N · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2021
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., Ibarz, J., Ichter, B., Irpan, A., Jackson, T., Jesmonth, S., Joshi, N. J., Julian, R. C., Kalashnikov, D., Kuang, Y., Leal, I., Lee, K.-H., Levine, S., Lu, Y., Malla, U., Manjunath, D., Mordatch, I., Nachum, O., Parada, C., Peralta, J., Perez, E., Pertsch, K., Quiambao, J., Rao, K., Ryoo, M., Salazar, G., Sanketi, P., Sayed, K., Singh, J., Sontakke, S., Stone, A., Tan, C., Tran, H., Vanhoucke, V., Vega, S., Vuong, Q., Xia, F., Xiao, T., Xu, P., Xu, S., Yu, T., and Zitkovich, B · 2022
suessmann/agentic-transformer-pytorch: v1.0, 2024
Zisman, I · 2022
Later among the works it cites.
Transformers in reinforcement learning: A survey
Agarwal, P., Rahman, A. A., St-Charles, P.-L., Prince, S. J. D., and Kahou, S. E · 2023
Closest in time.
Deep rl at scale: Sorting waste in office buildings with a fleet of mobile manipulators
Herzog, A., Rao, K., Hausman, K., Lu, Y., Wohlhart, P., Yan, M., Lin, J., Arenas, M. G., Xiao, T., Kappler, D., Ho, D., Rettinghouse, J., Chebotar, Y., Lee, K.-H., Gopalakrishnan, K., Julian, R., Li, A., Fu, C. K., Wei, B., Ramesh, S., Holden, K., Kleiven, K., Rendleman, D., Kirmani, S., Bingham, J., Weisz, J., Xu, Y., Lu, W., Bennice, M., Fong, C., Do, D., Lam, J., Bai, Y., Holson, B., Quinlan, M., Brown, N., Kalakrishnan, M., Ibarz, J., Pastor, P., and Levine, S · 2023
Closest in time.
Towards general-purpose in-context learning agents
Kirsch, L., Harrison, J., Freeman, C. D., Sohl-Dickstein, J., and Schmidhuber, J · 2023
Closest in time.
Katakomba: Tools and benchmarks for data-driven nethack
Kurenkov, V., Nikulin, A., Tarasov, D., and Kolesnikov, S · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Why can gpt learn in-context? language models secretly perform gradient descent as meta optimizers
Dai, D., Sun, Y., Dong, L., Hao, Y., Sui, Z., and Wei, F · 2022
Cited alongside, same era.
Dungeons and data: A large-scale nethack dataset
Hambro, E., Raileanu, R., Rothermel, D., Mella, V., Rocktäschel, T., Küttler, H., and Murray, N · 2022
Cited alongside, same era.
General-purpose in-context learning by meta-learning transformers
Kirsch, L., Harrison, J., Sohl-Dickstein, J., and Metz, L · 2022
Cited alongside, same era.
In-context reinforcement learning with algorithm distillation
Laskin, M., Wang, L., Oh, J., Parisotto, E., Spencer, S., Steigerwald, R., Strouse, D., Hansen, S., Filos, A., Brooks, E., et al · 2022
Cited alongside, same era.
Multi-game decision transformers
Lee, K.-H., Nachum, O., Yang, M. S., Lee, L., Freeman, D., Guadarrama, S., Fischer, I., Xu, W., Jang, E., Michalewski, H., et al · 2022
Cited alongside, same era.
Switch trajectory transformer with distributional value approximation for multi-task reinforcement learning
Lin, Q., Liu, H., and Sengupta, B · 2022
Cited alongside, same era.
A generalist agent
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., Eccles, T., Bruce, J., Razavi, A., Edwards, A., Heess, N., Chen, Y., Hadsell, R., Vinyals, O., Bordbar, M., and de Freitas, N · 2022
Cited alongside, same era.
Lee, J. N., Xie, A., Pacchiano, A., Chandak, Y., Finn, C., Nachum, O., and Brunskill, E · 2023
Closest in time.
Emergent agentic transformer from chain of hindsight experience
Liu, H. and Abbeel, P · 2023
Closest in time.
Transformers are sample-efficient world models
Micheli, V., Alonso, E., and Fleuret, F · 2023
Closest in time.
Xland-minigrid: Scalable meta-reinforcement learning environments in jax
Nikulin, A., Kurenkov, V., Zisman, I., Agarkov, A., Sinii, V., and Kolesnikov, S · 2023
Closest in time.
Cross-episodic curriculum for transformer agents
Shi, L. X., Jiang, Y., Grigsby, J., Fan, L. J., and Zhu, Y · 2023
Closest in time.
In-context reinforcement learning for variable action spaces
Sinii, V., Nikulin, A., Kurenkov, V., Zisman, I., and Kolesnikov, S · 2023
Closest in time.
Shimmy: Gymnasium and PettingZoo Wrappers for Commonly Used Environments, June 2023
Tai, J. J., Towers, M., and Tower, E · 2023
Closest in time.
Transformers learn in-context by gradient descent
Von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Closest in time.
An information-theoretic analysis of in-context learning
Jeon, H. J., Lee, J. D., Lei, Q., and Roy, B. V · 2024
Closest in time.
Transformers learn temporal difference methods for in-context reinforcement learning, 2024
Wang, J., Blaser, E., Daneshmand, H., and Zhang, S · 2024
Closest in time.