Fetching the paper…
Reading the bibliography…
Foundation models have shown impressive adaptation and scalability in supervised and self-supervised learning problems, but so far these successes have not fully translated to reinforcement learning (RL).
Interaction between learning and development
L. Vygotsky · 1978
Earlier work this paper cites.
Curious model-building control systems
J. Schmidhuber · 1991
Earlier work this paper cites.
Learning complex, extended sequences using the principle of history compression
J. Schmidhuber · 1992
Earlier work this paper cites.
Roma: Multi-agent reinforcement learning with emergent roles
T. Wang, H. Dong, V. Lesser, and C. Zhang · 2003
Earlier work this paper cites.
Learning to incentivize other learning agents
J. Yang, A. Li, M. Farajtabar, P. Sunehag, E. Hughes, and H. Zha · 2006
Earlier work this paper cites.
An experiment in automatic game design
J. Togelius and J. Schmidhuber · 2008
Earlier work this paper cites.
Ad hoc autonomous agent teams: Collaboration without pre-coordination
P. Stone, G. A. Kaminka, S. Kraus, and J. S. Rosenschein · 2010
Earlier work this paper cites.
On the properties of neural machine translation: Encoder-decoder approaches
K. Cho, B. Van Merriënboer, D. Bahdanau, and Y. Bengio · 2014
Earlier work this paper cites.
Robots that can adapt like animals
A. Cully, J. Clune, D. Tarapore, and J.-B. Mouret · 2015
Earlier work this paper cites.
Fictitious self-play in extensive-form games
J. Heinrich, M. Lanctot, and D. Silver · 2015
Earlier work this paper cites.
Prioritized experience replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
M. Andrychowicz, M. Denil, S. Gomez, M. W. Hoffman, D. Pfau, T. Schaul, B. Shillingford, and N. De Freitas · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
The Malmo platform for artificial intelligence experimentation
M. Johnson, K. Hofmann, T. Hutton, and D. Bignell · 2016
Earlier work this paper cites.
Safe and efficient off-policy reinforcement learning
R. Munos, T. Stepleton, A. Harutyunyan, and M. Bellemare · 2016
Earlier work this paper cites.
Learning to reinforcement learn
J. X. Wang, Z. Kurth-Nelson, D. Tirumala, H. Soyer, J. Z. Leibo, R. Munos, C. Blundell, D. Kumaran, and M. Botvinick · 2016
Earlier work this paper cites.
RL 2 : Fast reinforcement learning via slow reinforcement learning, 2017
Y. Duan, J. Schulman, X. Chen, P. L. Bartlett, I. Sutskever, and P. Abbeel · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Population based training of neural networks
M. Jaderberg, V. Dalibard, S. Osindero, W. M. Czarnecki, J. Donahue, A. Razavi, O. Vinyals, T. Green, I. Dunning, K. Simonyan, et al · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
C. Sun, A. Shrivastava, S. Singh, and A. Gupta · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Natural value approximators: Learning when to trust past estimates
Z. Xu, J. Modayil, H. P. van Hasselt, A. Barreto, D. Silver, and T. Schaul · 2017
Earlier work this paper cites.
Re-evaluating evaluation
D. Balduzzi, K. Tuyls, J. Perolat, and T. Graepel · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Earlier work this paper cites.
Minimalistic gridworld environment for OpenAI Gym
M. Chevalier-Boisvert, L. Willems, and S. Pal · 2018
Earlier work this paper cites.
Quantifying generalization in reinforcement learning
K. Cobbe, O. Klimov, C. Hesse, T. Kim, and J. Schulman · 2018
Earlier work this paper cites.
Probabilistic model-agnostic meta-learning
C. Finn, K. Xu, and S. Levine · 2018
Earlier work this paper cites.
Procedural level generation improves generality of deep reinforcement learning
N. Justesen, R. R. Torrado, P. Bontrager, A. Khalifa, J. Togelius, and S. Risi · 2018
Earlier work this paper cites.
Exploring the limits of weakly supervised pretraining
D. Mahajan, R. Girshick, V. Ramanathan, K. He, M. Paluri, Y. Li, A. Bharambe, and L. van der Maaten · 2018
Earlier work this paper cites.
Kickstarting deep reinforcement learning
S. Schmitt, J. J. Hudson, A. Zidek, S. Osindero, C. Doersch, W. M. Czarnecki, J. Z. Leibo, H. Kuttler, A. Zisserman, K. Simonyan, et al · 2018
Earlier work this paper cites.
Some considerations on learning to explore via meta-reinforcement learning
B. C. Stadie, G. Yang, R. Houthooft, X. Chen, Y. Duan, Y. Wu, P. Abbeel, and I. Sutskever · 2018
Earlier work this paper cites.
Intrinsic motivation and automatic curricula via asymmetric self-play
S. Sukhbaatar, Z. Lin, I. Kostrikov, G. Synnaeve, A. Szlam, and R. Fergus · 2018
Earlier work this paper cites.
Prefrontal cortex as a meta-reinforcement learning system
J. X. Wang, Z. Kurth-Nelson, D. Kumaran, D. Tirumala, H. Soyer, J. Z. Leibo, D. Hassabis, and M. Botvinick · 2018
Earlier work this paper cites.
Meta-gradient reinforcement learning
Z. Xu, H. P. van Hasselt, and D. Silver · 2018
Earlier work this paper cites.
Solving rubik’s cube with a robot hand
I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, et al · 2019
Earlier work this paper cites.
Dota 2 with large scale deep reinforcement learning
C. Berner, G. Brockman, B. Chan, V. Cheung, P. Debiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Józefowicz, S. Gray, C. Olsson, J. Pachocki, M. Petrov, H. P. de Oliveira Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Tang, F. Wolski, and S. Zhang · 2019
Earlier work this paper cites.
On the utility of learning about humans for human-ai coordination, 2019
M. Carroll, R. Shah, M. K. Ho, T. L. Griffiths, S. A. Seshia, P. Abbeel, and A. Dragan · 2019
Earlier work this paper cites.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
I. Clavera, A. Nagabandi, S. Liu, R. S. Fearing, P. Abbeel, S. Levine, and C. Finn · 2019
Earlier work this paper cites.
J. Clune · 2019
Earlier work this paper cites.
Distilling policy distillation
W. M. Czarnecki, R. Pascanu, S. Osindero, S. Jayakumar, G. Swirszcz, and M. Jaderberg · 2019
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. Le, and R. Salakhutdinov · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
A meta-mdp approach to exploration for lifelong reinforcement learning
F. Garcia and P. S. Thomas · 2019
Cited alongside, same era.
Meta reinforcement learning as task inference
J. Humplik, A. Galashov, L. Hasenclever, P. A. Ortega, Y. W. Teh, and N. Heess · 2019
Cited alongside, same era.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
M. Jaderberg, W. M. Czarnecki, I. Dunning, L. Marris, G. Lever, A. G. Castañeda, C. Beattie, N. C. Rabinowitz, A. S. Morcos, A. Ruderman, N. Sonnerat, T. Green, L. Deason, J. Z. Leibo, D. Silver, D. Hassabis, K. Kavukcuoglu, and T. Graepel · 2019
Cited alongside, same era.
Obstacle Tower: A Generalization Challenge in Vision, Control, and Planning
A. Juliani, A. Khalifa, V. Berges, J. Harper, E. Teng, H. Henry, A. Crespi, J. Togelius, and D. Lange · 2019
RMA: Rapid motor adaptation for legged robots
A. Kumar, Z. Fu, D. Pathak, and J. Malik · 2021
Later among the works it cites.
Gradients are not all you need
L. Metz, C. D. Freeman, S. S. Schoenholz, and T. Kachman · 2021
Later among the works it cites.
Open-ended learning leads to generally capable agents
OEL Team, A. Stooke, A. Mahajan, C. Barros, C. Deck, J. Bauer, J. Sygnowski, M. Trebacz, M. Jaderberg, M. Mathieu, N. McAleese, N. Bradley-Schmieg, N. Wong, N. Porcel, R. Raileanu, S. Hughes-Fitt, V. Dalibard, and W. M. Czarnecki · 2021
Later among the works it cites.
Asymmetric self-play for automatic goal discovery in robotic manipulation, 2021
OpenAI, M. Plappert, R. Sampedro, T. Xu, I. Akkaya, V. Kosaraju, P. Welinder, R. D’Sa, A. Petron, H. P. de Oliveira Pinto, A. Paino, H. Noh, L. Weng, Q. Yuan, C. Chu, and W. Zaremba · 2021
Later among the works it cites.
Training larger networks for deep reinforcement learning, 2021
K. Ota, D. K. Jha, and A. Kanezaki · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Emergent coordination through competition
S. Liu, G. Lever, N. Heess, J. Merel, S. Tunyasuvunakool, and T. Graepel · 2019
Cited alongside, same era.
Meta-learning of sequential strategies
P. A. Ortega, J. X. Wang, M. Rowland, T. Genewein, Z. Kurth-Nelson, R. Pascanu, N. Heess, J. Veness, A. Pritzel, P. Sprechmann, et al · 2019
Cited alongside, same era.
Teacher algorithms for curriculum learning of deep RL in continuously parameterized environments
R. Portelas, C. Colas, K. Hofmann, and P. Oudeyer · 2019
Cited alongside, same era.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
K. Rakelly, A. Zhou, C. Finn, S. Levine, and D. Quillen · 2019
Cited alongside, same era.
Grandmaster level in starcraft II using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, J. Oh, . D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. P. Agapiou, M. Jaderberg, A. S. Vezhnevets, R. Leblond, T. Pohlen, V. Dalibard, D. Budden, Y. Sulsky, J. Molloy, T. L. Paine, Ç. Gülçehre, Z. Wang, T. Pfaff, Y. Wu, R. Ring, D. Yogatama, D. Wünsch, K. McKinney, O. Smith, T. Schaul, T. P. Lillicrap, K. Kavukcuoglu, D. Hassabis, C. Apps, and D. Silver · 2019
Cited alongside, same era.
R. Wang, J. Lehman, J. Clune, and K. O. Stanley · 2019
Cited alongside, same era.
Single episode policy transfer in reinforcement learning
J. Yang, B. Petersen, H. Zha, and D. Faissol · 2019
Cited alongside, same era.
Later among the works it cites.
Meta Reinforcement Learning through Memory
E. Parisotto · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, et al · 2021
Later among the works it cites.
Minihack the planet: A sandbox for open-ended reinforcement learning research
M. Samvelyan, R. Kirk, V. Kurin, J. Parker-Holder, M. Jiang, E. Hambro, F. Petroni, H. Kuttler, E. Grefenstette, and T. Rocktäschel · 2021
Later among the works it cites.
Collaborating with humans without human data
D. Strouse, K. McKee, M. Botvinick, E. Hughes, and R. Everett · 2021
Later among the works it cites.
No dice: An investigation of the bias-variance tradeoff in meta-gradients
R. Vuorio, J. A. Beck, G. Farquhar, J. N. Foerster, and S. Whiteson · 2021
Later among the works it cites.
Alchemy: A structured task distribution for meta-reinforcement learning
J. X. Wang, M. King, N. Porcel, Z. Kurth-Nelson, T. Zhu, C. Deck, P. Choy, M. Cassin, M. Reynolds, F. Song, et al · 2021
Later among the works it cites.
Avalon: A benchmark for RL generalization using procedurally generated worlds
J. Albrecht, A. J. Fetterman, B. Fogelman, E. Kitanidis, B. Wróblewski, N. Seo, M. Rosenthal, M. Knutins, Z. Polizzi, J. B. Simon, and K. Qiu · 2022
Later among the works it cites.
A practical guide for studying human behavior in the lab
J. Barbosa, H. Stein, S. Zorowitz, Y. Niv, C. Summerfield, S. Soto-Faraco, and A. Hyafil · 2022
Later among the works it cites.
Deep surrogate assisted generation of environments
V. Bhatt, B. Tjanaka, M. C. Fontaine, and S. Nikolaidis · 2022
Later among the works it cites.
Stabilizing off-policy deep reinforcement learning from pixels
E. Cetin, P. J. Ball, S. Roberts, and O. Celiktutan · 2022
Later among the works it cites.
Towards learning universal hyperparameter optimizers with transformers
Y. Chen, X. Song, C. Lee, Z. Wang, Q. Zhang, D. Dohan, K. Kawakami, G. Kochanski, A. Doucet, M. Ranzato, S. Perel, and N. de Freitas · 2022
Later among the works it cites.
Pareto actor-critic for equilibrium selection in multi-agent reinforcement learning
F. Christianos, G. Papoudakis, and S. V. Albrecht · 2022
Later among the works it cites.
Learning robust real-time cultural transmission without human data, 2022
Cultural General Intelligence Team, A. Bhoopchand, B. Brownfield, A. Collister, A. D. Lago, A. Edwards, R. Everett, A. Frechette, Y. G. Oliveira, E. Hughes, K. W. Mathewson, P. Mendolicchio, J. Pawar, M. Pislar, A. Platonov, E. Senter, S. Singh, A. Zacherl, and L. M. Zhang · 2022
Later among the works it cites.
ProcTHOR: Large-scale embodied AI using procedural generation
M. Deitke, E. VanderBilt, A. Herrasti, L. Weihs, J. Salvador, K. Ehsani, W. Han, E. Kolve, A. Farhadi, A. Kembhavi, and R. Mottaghi · 2022
Later among the works it cites.
Minedojo: Building open-ended embodied agents with internet-scale knowledge
L. Fan, G. Wang, Y. Jiang, A. Mandlekar, Y. Yang, H. Zhu, A. Tang, D.-A. Huang, Y. Zhu, and A. Anandkumar · 2022
Later among the works it cites.
Model-value inconsistency as a signal for epistemic uncertainty
A. Filos, E. Vértes, Z. Marinho, G. Farquhar, D. Borsa, A. L. Friesen, F. M. P. Behbahani, T. Schaul, A. Barreto, and S. Osindero · 2022
Later among the works it cites.
Bootstrapped meta-learning
S. Flennerhag, Y. Schroecker, T. Zahavy, H. van Hasselt, D. Silver, and S. Singh · 2022
Later among the works it cites.
Multi-agent deep reinforcement learning: a survey
S. Gronauer and K. Diepold · 2022
Later among the works it cites.
Benchmarking the spectrum of agent capabilities
D. Hafner · 2022
Later among the works it cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Later among the works it cites.
Minerl diamond 2021 competition: Overview, results, and lessons learned, 2022
A. Kanervisto, S. Milani, K. Ramanauskas, N. Topin, Z. Lin, J. Li, J. Shi, D. Ye, Q. Fu, W. Yang, W. Hong, Z. Huang, H. Chen, G. Zeng, Y. Lin, V. Micheli, E. Alonso, F. Fleuret, A. Nikulin, Y. Belousov, O. Svidchenko, and A. Shpilman · 2022
Later among the works it cites.
General-purpose in-context learning by meta-learning transformers
L. Kirsch, J. Harrison, J. Sohl-Dickstein, and L. Metz · 2022
Later among the works it cites.
In-context reinforcement learning with algorithm distillation, 2022
M. Laskin, L. Wang, J. Oh, E. Parisotto, S. Spencer, R. Steigerwald, D. Strouse, S. Hansen, A. Filos, E. Brooks, M. Gazeau, H. Sahni, S. Singh, and V. Mnih · 2022
Later among the works it cites.
Multi-game decision transformers
K.-H. Lee, O. Nachum, S. Yang, L. Lee, C. D. Freeman, S. Guadarrama, I. Fischer, W. Xu, E. Jang, H. Michalewski, and I. Mordatch · 2022
Later among the works it cites.
Transformers are meta-reinforcement learners
L. C. Melo · 2022
Later among the works it cites.
Improving intrinsic exploration with language abstractions
J. Mu, V. Zhong, R. Raileanu, M. Jiang, N. Goodman, T. Rocktäschel, and E. Grefenstette · 2022
Later among the works it cites.
The primacy bias in deep reinforcement learning
E. Nikishin, M. Schwarzer, P. D’Oro, P.-L. Bacon, and A. Courville · 2022
Later among the works it cites.
Evolving curricula with regret-based environment design
J. Parker-Holder, M. Jiang, M. Dennis, M. Samvelyan, J. Foerster, E. Grefenstette, and T. Rocktäschel · 2022
Later among the works it cites.
When should agents explore?
M. Pislar, D. Szepesvari, G. Ostrovski, D. L. Borsa, and T. Schaul · 2022
Later among the works it cites.
A generalist agent
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-maron, M. Giménez, Y. Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. Edwards, N. Heess, Y. Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas · 2022
Later among the works it cites.
Can wikipedia help offline reinforcement learning?
M. Reid, Y. Yamada, and S. S. Gu · 2022
Later among the works it cites.
MAESTRO: Open-ended environment design for multi-agent reinforcement learning
M. Samvelyan, A. Khan, M. D. Dennis, M. Jiang, J. Parker-Holder, J. N. Foerster, R. Raileanu, and T. Rocktäschel · 2022
Later among the works it cites.
LAION-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. W. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. R. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev · 2022
Later among the works it cites.
Scaling laws vs model architectures: How does inductive bias influence scaling?, 2022
Y. Tay, M. Dehghani, S. Abnar, H. W. Chung, W. Fedus, J. Rao, S. Narang, V. Q. Tran, D. Yogatama, and D. Metzler · 2022
Later among the works it cites.
Scaling vision transformers
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2022
Later among the works it cites.
Fast adaptation via meta reinforcement learning
L. Zintgraf · 2022
Later among the works it cites.