Fetching the paper…
Reading the bibliography…
We introduce AMAGO, an in-context Reinforcement Learning (RL) agent that uses sequence models to tackle the challenges of generalization, long-term memory, and meta-learning.
“Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook”, 1987
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
“Learning to achieve goals”
Leslie Kaelbling · 1993
Earlier work this paper cites.
“Acting optimally in partially observable stochastic domains”
Anthony Cassandra, Leslie Kaelbling and Michael Littman · 1994
Earlier work this paper cites.
“Learning phrase representations using RNN encoder-decoder for statistical machine translation”
Kyunghyun Cho et al · 2014
Earlier work this paper cites.
“Human-level control through deep reinforcement learning”
Volodymyr Mnih et al · 2015
Earlier work this paper cites.
“Contextual markov decision processes”
Assaf Hallak, Dotan Di and Shie Mannor · 2015
Earlier work this paper cites.
“Deep recurrent q-learning for partially observable mdps”
Matthew Hausknecht and Peter Stone · 2015
Earlier work this paper cites.
“Memory-based control with recurrent neural networks”
Nicolas Heess, Jonathan Hunt, Timothy Lillicrap and David Silver · 2015
Earlier work this paper cites.
“Continuous control with deep reinforcement learning”
Timothy Lillicrap et al · 2015
Earlier work this paper cites.
“Learning to reinforcement learn”
Jane Wang et al · 2016
Earlier work this paper cites.
“RL 2 : Fast reinforcement learning via slow reinforcement learning”
Yan Duan et al · 2016
Earlier work this paper cites.
“Learning values across many orders of magnitude”
Hado van Hasselt et al · 2016
Earlier work this paper cites.
Greg Brockman et al · 2016
Earlier work this paper cites.
Jimmy Ba, Jamie Kiros and Geoffrey Hinton · 2016
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“A simple neural attentive meta-learner”
Nikhil Mishra, Mostafa Rohaninejad, Xi Chen and Pieter Abbeel · 2017
Earlier work this paper cites.
“Hindsight experience replay”
Marcin Andrychowicz et al · 2017
Earlier work this paper cites.
“Model-agnostic meta-learning for fast adaptation of deep networks”
Chelsea Finn, Pieter Abbeel and Sergey Levine · 2017
Earlier work this paper cites.
“A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play”
David Silver et al · 2018
Earlier work this paper cites.
“Some considerations on learning to explore via meta-reinforcement learning”
Bradly Stadie et al · 2018
Earlier work this paper cites.
“Assessing generalization in deep reinforcement learning”
Charles Packer et al · 2018
Earlier work this paper cites.
“A dissection of overfitting and generalization in continuous reinforcement learning”
Amy Zhang, Nicolas Ballas and Joelle Pineau · 2018
Earlier work this paper cites.
“Visual reinforcement learning with imagined goals”
Ashvin Nair et al · 2018
Earlier work this paper cites.
“Deep reinforcement learning that matters”
Peter Henderson et al · 2018
Earlier work this paper cites.
“Recurrent experience replay in distributed reinforcement learning”
Steven Kapturowski et al · 2018
Earlier work this paper cites.
“Addressing function approximation error in actor-critic methods”
Scott Fujimoto, Herke Hoof and David Meger · 2018
Earlier work this paper cites.
“Time limits in reinforcement learning”
Fabio Pardo, Arash Tavakoli, Vitaly Levdik and Petar Kormushev · 2018
Earlier work this paper cites.
“Overcoming exploration in reinforcement learning with demonstrations”
Ashvin Nair et al · 2018
Earlier work this paper cites.
“Rainbow: Combining improvements in deep reinforcement learning”
Matteo Hessel et al · 2018
Earlier work this paper cites.
“Soft actor-critic algorithms and applications”
Tuomas Haarnoja et al · 2018
Earlier work this paper cites.
“Dota 2 with large scale deep reinforcement learning”
Christopher Berner et al · 2019
Earlier work this paper cites.
“Grandmaster level in StarCraft II using multi-agent reinforcement learning”
Oriol Vinyals et al · 2019
Earlier work this paper cites.
“Meta-learning of sequential strategies”
Pedro Ortega et al · 2019
Earlier work this paper cites.
“Reinforcement learning, fast and slow”
Matthew Botvinick et al · 2019
Earlier work this paper cites.
Rasool Fakoor, Pratik Chaudhari, Stefano Soatto and Alexander Smola · 2019
Earlier work this paper cites.
“Observational overfitting in reinforcement learning”
Xingyou Song et al · 2019
Earlier work this paper cites.
“Meta reinforcement learning as task inference”
Jan Humplik et al · 2019
Earlier work this paper cites.
“Efficient off-policy meta-reinforcement learning via probabilistic context variables”
Kate Rakelly et al · 2019
Earlier work this paper cites.
“Learning to reach goals via iterated supervised learning”
Dibya Ghosh et al · 2019
Cited alongside, same era.
Aviral Kumar, Xue Peng and Sergey Levine · 2019
Cited alongside, same era.
“Reinforcement Learning Upside Down: Don’t Predict Rewards–Just Map Them to Actions”
Juergen Schmidhuber · 2019
Cited alongside, same era.
“Training agents using upside-down reinforcement learning”
Rupesh Srivastava et al · 2019
Cited alongside, same era.
“Soft actor-critic for discrete action settings”
Petros Christodoulou · 2019
“Normformer: Improved transformer pretraining with extra normalization”
Sam Shleifer, Jason Weston and Myle Ott · 2021
Later among the works it cites.
“A Closer Look at Advantage-Filtered Behavioral Cloning in High-Noise Datasets”
Jake Grigsby and Yanjun Qi · 2021
Later among the works it cites.
“Exploration in approximate hyper-state space for meta reinforcement learning”
Luisa Zintgraf et al · 2021
Later among the works it cites.
“A Survey of Generalisation in Deep Reinforcement Learning”, 2022
Robert Kirk, Amy Zhang, Edward Grefenstette and Tim Rocktäschel · 2022
Later among the works it cites.
“Minedojo: Building open-ended embodied agents with internet-scale knowledge”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Transformers without tears: Improving the normalization of self-attention”
Toan Nguyen and Julian Salazar · 2019
Cited alongside, same era.
“Hyperbolic discounting and learning over multiple horizons”
William Fedus et al · 2019
Cited alongside, same era.
“Agent57: Outperforming the atari human benchmark”
Adriàènech Badia et al · 2020
Cited alongside, same era.
“Do transformers need deep long-range memory”
Jack Rae and Ali Razavi · 2020
Cited alongside, same era.
Pierre-Alexandre Kamienny et al · 2020
Cited alongside, same era.
“OCEAN: Online Task Inference for Compositional Tasks with Context Adaptation”
Hongyu Ren et al · 2020
Cited alongside, same era.
“Rewriting history with inverse rl: Hindsight inference for policy improvement”
Ben Eysenbach, Xinyang Geng, Sergey Levine and Russ Salakhutdinov · 2020
Cited alongside, same era.
Linxi Fan et al · 2022
Later among the works it cites.
Scott Reed et al · 2022
Later among the works it cites.
“Fast adaptation via meta reinforcement learning”, 2022
Luisa Zintgraf · 2022
Later among the works it cites.
“Transformers are meta-reinforcement learners”
Luckeciano Melo · 2022
Later among the works it cites.
“Recurrent Model-Free RL can be a Strong Baseline for Many POMDPs”, 2022
Tianwei Ni, Benjamin Eysenbach and Ruslan Salakhutdinov · 2022
Later among the works it cites.
“Multi-game decision transformers”
Kuang-Huei Lee et al · 2022
Later among the works it cites.
Qinqing Zheng, Amy Zhang and Aditya Grover · 2022
Later among the works it cites.
“In-context Reinforcement Learning with Algorithm Distillation”
Michael Laskin et al · 2022
Later among the works it cites.
“From play to policy: Conditional behavior generation from uncurated robot data”
Zichen Cui, Yibin Wang, Nur Muhammad and Lerrel Pinto · 2022
Later among the works it cites.
“Prompting decision transformer for few-shot policy generalization”
Mengdi Xu et al · 2022
Later among the works it cites.
“Imitating Past Successes can be Very Suboptimal”
Benjamin Eysenbach, Soumith Udatha, Sergey Levine and Ruslan Salakhutdinov · 2022
Later among the works it cites.
“You Can’t Count on Luck: Why Decision Transformers Fail in Stochastic Environments”
Keiran Paster, Sheila McIlraith and Jimmy Ba · 2022
Later among the works it cites.
“When does return-conditioned supervised learning work for offline reinforcement learning?”
David Brandfonbrener et al · 2022
Later among the works it cites.
“Rethinking goal-conditioned supervised learning and its connection to offline rl”
Rui Yang et al · 2022
Later among the works it cites.
“Generalization, Mayhems and Limits in Recurrent Proximal Policy Optimization”
Marco Pleines, Matthias Pallasch, Frank Zimmer and Mike Preuss · 2022
Later among the works it cites.
“Sample-efficient reinforcement learning by breaking the replay ratio barrier”
Pierluca D’Oro et al · 2022
Later among the works it cites.
“The primacy bias in deep reinforcement learning”
Evgenii Nikishin et al · 2022
Later among the works it cites.
“Evaluating Long-Term Memory in 3D Mazes”
Jurgis Pasukonis, Timothy Lillicrap and Danijar Hafner · 2022
Later among the works it cites.
“Efficient Multi-Horizon Learning for Off-Policy Reinforcement Learning”
Raja Ali, Nasik Nafi, Kevin Duong and William Hsu · 2022
Later among the works it cites.
“FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness”
Tri Dao et al · 2022
Later among the works it cites.
“In defense of the unitary scalarization for deep multi-task learning”
Vitaly Kurin et al · 2022
Later among the works it cites.
“Human-Timescale Adaptation in an Open-Ended Task Space”
Adaptive Team et al · 2023
Closest in time.
“A Survey of Meta-Reinforcement Learning”
Jacob Beck et al · 2023
Closest in time.
“Cross-Episodic Curriculum for Transformer Agents”
Lucy Shi et al · 2023
Closest in time.
“POPGym: Benchmarking Partially Observable Reinforcement Learning”
Steven Morad et al · 2023
Closest in time.
“Structured state space models for in-context reinforcement learning”
Chris Lu et al · 2023
Closest in time.
“Llama 2: Open foundation and fine-tuned chat models”
Hugo Touvron et al · 2023
Closest in time.
“Loss of plasticity in continual deep reinforcement learning”
Zaheer Abbas et al · 2023
Closest in time.
“Stabilizing Transformer Training by Preventing Attention Entropy Collapse”
Shuangfei Zhai et al · 2023
Closest in time.
“Efficient Deep Reinforcement Learning Requires Regulating Overfitting”
Qiyang Li, Aviral Kumar, Ilya Kostrikov and Sergey Levine · 2023
Closest in time.
“When Do Transformers Shine in RL? Decoupling Memory from Credit Assignment”
Tianwei Ni, Michel Ma, Benjamin Eysenbach and Pierre-Luc Bacon · 2023
Closest in time.
“Bigger, Better, Faster: Human-level Atari with human-level efficiency”
Max Schwarzer et al · 2023
Closest in time.
“Leveraging procedural generation to benchmark reinforcement learning”
Karl Cobbe, Chris Hesse, Jacob Hilton and John Schulman · 2056
Closest in time.