Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) agents typically optimize their policies by performing expensive backward passes to update their network parameters.
Learning to predict by the methods of temporal differences
Richard S. Sutton · 1988
Earlier work this paper cites.
Average cost temporal-difference learning
John N. Tsitsiklis and Benjamin Van Roy · 1999
Earlier work this paper cites.
Exploration and apprenticeship learning in reinforcement learning
Pieter Abbeel and Andrew Y. Ng · 2005
Earlier work this paper cites.
A primal-dual perspective of online learning algorithms
Shai Shalev-Shwartz and Yoram Singer · 2007
Earlier work this paper cites.
Contextual bandits with linear payoff functions
Wei Chu, Lihong Li, Lev Reyzin, and Robert Schapire · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning
Yan Duan, John Schulman, Xi Chen, Peter L. Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Learning to reinforcement learn
Jane X. Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z. Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Earlier work this paper cites.
Minimax regret bounds for reinforcement learning
Mohammad Gheshlaghi Azar, Ian Osband, and Rémi Munos · 2017
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A simple neural attentive meta-learner
Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel · 2018
Earlier work this paper cites.
Been there, done that: Meta-learning with episodic recall
Samuel Ritter, Jane X. Wang, Zeb Kurth-Nelson, Siddhant M. Jayakumar, Charles Blundell, Razvan Pascanu, and Matthew Botvinick · 2018
Earlier work this paper cites.
A tutorial on thompson sampling
Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, Zheng Wen, et al · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction (2nd Edition)
Richard S Sutton and Andrew G Barto · 2018
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2019
Earlier work this paper cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Chelsea Finn, Sergey Levine, and Deirdre Quillen · 2019
Earlier work this paper cites.
Some Considerations on Learning to Explore via Meta-Reinforcement Learning
Bradly C. Stadie, Ge Yang, Rein Houthooft, Xi Chen, Yan Duan, Yuhuai Wu, Pieter Abbeel, and Ilya Sutskever · 2019
Earlier work this paper cites.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
Karl Cobbe, Christopher Hesse, Jacob Hilton, and John Schulman · 2020
Earlier work this paper cites.
Varibad: A very good method for bayes-adaptive deep RL via meta-learning
Luisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, and Shimon Whiteson · 2020
Earlier work this paper cites.
Learning to cooperate with unseen agent via meta-reinforcement learning
Rujikorn Charakorn, Poramate Manoonpong, and Nat Dilokthanakul · 2021
Earlier work this paper cites.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Earlier work this paper cites.
Muesli: Combining improvements in policy optimization
Matteo Hessel, Ivo Danihelka, Fabio Viola, Arthur Guez, Simon Schmitt, Laurent Sifre, Theophane Weber, David Silver, and Hado Van Hasselt · 2021
Cited alongside, same era.
A Survey of Zero-shot Generalisation in Deep Reinforcement Learning
Robert Kirk, Amy Zhang, Edward Grefenstette, and Tim Rocktäschel · 2021
Cited alongside, same era.
Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Avnish Narayan, Hayden Shively, Adithya Bellathur, Karol Hausman, Chelsea Finn, and Sergey Levine · 2021
Cited alongside, same era.
When does return-conditioned supervised learning work for offline reinforcement learning?
David Brandfonbrener, Alberto Bietti, Jacob Buckman, Romain Laroche, and Joan Bruna · 2022
Cited alongside, same era.
Rvs: What is essential for offline RL via supervised learning?
Scott Emmons, Benjamin Eysenbach, Ilya Kostrikov, and Sergey Levine · 2022
Cited alongside, same era.
Cross-episodic curriculum for transformer agents
Lucy Xiaoyang Shi, Yunfan Jiang, Jake Grigsby, Linxi Fan, and Yuke Zhu · 2023
Later among the works it cites.
In-Context Reinforcement Learning for Variable Action Spaces
Viacheslav Sinii, Alexander Nikulin, Vladislav Kurenkov, Ilya Zisman, and Sergey Kolesnikov · 2023
Later among the works it cites.
Simplified state space layers for sequence modeling
Jimmy T. H. Smith, Andrew Warrington, and Scott W. Linderman · 2023
Later among the works it cites.
Large Sequence Models for Sequential Decision-Making: A Survey
Muning Wen, Runji Lin, Hanjing Wang, Yaodong Yang, Ying Wen, Luo Mai, Jun Wang, Haifeng Zhang, and Weinan Zhang · 2023
Later among the works it cites.
Emergence of In-Context Reinforcement Learning from Noise Distillation
Ilya Zisman, Vladislav Kurenkov, Alexander Nikulin, Viacheslav Sinii, and Sergey Kolesnikov · 2023
Later among the works it cites.
In-Context Language Learning: Architectures and Algorithms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Generalized decision transformer for offline hindsight information matching
Hiroki Furuta, Yutaka Matsuo, and Shixiang Shane Gu · 2022
Cited alongside, same era.
Meta-rl for multi-agent rl: Learning to adapt to evolving agents
Matthias Gerstgrasser and David C Parkes · 2022
Cited alongside, same era.
Goal-conditioned reinforcement learning: Problems and solutions
Minghuan Liu, Menghui Zhu, and Weinan Zhang · 2022
Cited alongside, same era.
Transformers are meta-reinforcement learners
Luckeciano C. Melo · 2022
Cited alongside, same era.
Prompting decision transformer for few-shot policy generalization
Mengdi Xu, Yikang Shen, Shun Zhang, Yuchen Lu, Ding Zhao, Joshua B. Tenenbaum, and Chuang Gan · 2022
Cited alongside, same era.
Human-timescale adaptation in an open-ended task space
Jakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand, Nathalie Bradley-Schmieg, Michael Chang, Natalie Clay, Adrian Collister, Vibhavari Dasagi, Lucy Gonzalez, Karol Gregor, Edward Hughes, Sheleem Kashem, Maria Loks-Thompson, Hannah Openshaw, Jack Parker-Holder, Shreya Pathak, Nicolas Perez Nieves, Nemanja Rakicevic, Tim Rocktäschel, Yannick Schroecker, Satinder Singh, Jakub Sygnowski, Karl Tuyls, Sarah York, Alexander Zacherl, and Lei M. Zhang · 2023
Cited alongside, same era.
A Survey of Meta-Reinforcement Learning
Jacob Beck, Risto Vuorio, Evan Zheran Liu, Zheng Xiong, Luisa Zintgraf, Chelsea Finn, and Shimon Whiteson · 2023
Cited alongside, same era.
Ekin Akyürek, Bailin Wang, Yoon Kim, and Jacob Andreas · 2024
Later among the works it cites.
xLSTM: Extended Long Short-Term Memory
Maximilian Beck, Korbinian Pöppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, Günter Klambauer, Johannes Brandstetter, and Sepp Hochreiter · 2024
Later among the works it cites.
Artificial Generational Intelligence: Cultural Accumulation in Reinforcement Learning
Jonathan Cook, Chris Lu, Edward Hughes, Joel Z. Leibo, and Jakob Nicolaus Foerster · 2024
Later among the works it cites.
In-context Exploration-Exploitation for Reinforcement Learning
Zhenwen Dai, Federico Tomasi, and Sina Ghiassian · 2024
Later among the works it cites.
In-Context Reinforcement Learning Without Optimal Action Labels
Juncheng Dong, Moyang Guo, Ethan X. Fang, Zhuoran Yang, and Vahid Tarokh · 2024
Later among the works it cites.
ReLIC: A Recipe for 64k Steps of In-Context Reinforcement Learning for Embodied AI
Ahmad Elawady, Gunjan Chhablani, Ram Ramrakhya, Karmesh Yadav, Dhruv Batra, Zsolt Kira, and Andrew Szot · 2024
Later among the works it cites.
AMAGO-2: Breaking the Multi-Task Barrier in Meta-Reinforcement Learning with Transformers
Jake Grigsby, Justin Sasek, Samyak Parajuli, Daniel Adebi, Amy Zhang, and Yuke Zhu · 2024
Later among the works it cites.
In-Context Decision Transformer: Reinforcement Learning via Hierarchical Chain-of-Thought
Sili Huang, Jifeng Hu, Hechang Chen, Lichao Sun, and Bo Yang · 2024
Later among the works it cites.
Decision Mamba: Reinforcement Learning via Hybrid Selective Sequence Modeling
Sili Huang, Jifeng Hu, Zhejian Yang, Liwei Yang, Tao Luo, Hechang Chen, Lichao Sun, and Bo Yang · 2024
Later among the works it cites.
V-learning—a simple, efficient, decentralized algorithm for multiagent reinforcement learning
Chi Jin, Qinghua Liu, Yuanhao Wang, and Tiancheng Yu · 2024
Later among the works it cites.
Can large language models explore in-context?
Akshay Krishnamurthy, Keegan Harris, Dylan J. Foster, Cyril Zhang, and Aleksandrs Slivkins · 2024
Later among the works it cites.
Do LLM Agents Have Regret? A Case Study in Online Learning and Games
Chanwoo Park, Xiangyu Liu, Asuman Ozdaglar, and Kaiqing Zhang · 2024
Later among the works it cites.
A Large Recurrent Action Model: xLSTM enables Fast Inference for Robotics Tasks
Thomas Schmied, Thomas Adler, Vihang Patil, Maximilian Beck, Korbinian Pöppel, Johannes Brandstetter, Günter Klambauer, Razvan Pascanu, and Sepp Hochreiter · 2024
Later among the works it cites.
Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models
Chengshuai Shi, Kun Yang, Jing Yang, and Cong Shen · 2024
Later among the works it cites.
Transformers Learn Temporal Difference Methods for In-Context Reinforcement Learning
Jiuqi Wang, Ethan Blaser, Hadi Daneshmand, and Shangtong Zhang · 2024
Later among the works it cites.
Hierarchical Prompt Decision Transformer: Improving Few-Shot Policy Generalization with Global and Adaptive Guidance
Zhe Wang, Haozhu Wang, and Yanjun Qi · 2024
Later among the works it cites.
Meta-Reinforcement Learning Robust to Distributional Shift Via Performing Lifelong In-Context Learning
Tengye Xu, Zihao Li, and Qinyuan Ren · 2024
Later among the works it cites.
A survey on model compression for large language models
Xunyu Zhu, Jian Li, Yong Liu, Can Ma, and Weiping Wang · 2024
Later among the works it cites.
N-Gram Induction Heads for In-Context RL: Improving Stability and Reducing Data Needs
Ilya Zisman, Alexander Nikulin, Andrei Polubarov, Nikita Lyubaykin, and Vladislav Kurenkov · 2024
Later among the works it cites.