Fetching the paper…
Reading the bibliography…
Integrating large language models (LLMs) as priors in reinforcement learning (RL) offers significant advantages but comes with substantial computational costs.
Mujoco: A physics engine for model-based control
Emanuel Todorov et al · 2012
Earlier work this paper cites.
Deep reinforcement learning in large discrete action spaces
Gabriel Dulac-Arnold, Richard Evans, Hado van Hasselt, Peter Sunehag, Timothy Lillicrap, Jonathan Hunt, Timothy Mann, Theophane Weber, Thomas Degris, and Ben Coppin · 2015
Earlier work this paper cites.
Bayesian policy gradient algorithms
Mohammad Ghavamzadeh et al · 2015
Earlier work this paper cites.
Bayesian reinforcement learning: A survey
Daniel Russo et al · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn et al · 2017
Earlier work this paper cites.
Survey of model-based reinforcement learning: Applications on robotics
Athanasios S Polydoros and Lazaros Nalpantidis · 2017
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning
Tuomas Haarnoja et al · 2018
Earlier work this paper cites.
Reinforcement learning and control as probabilistic inference
Sergey Levine · 2018
Earlier work this paper cites.
Reptile: a scalable metalearning algorithm
Alex Nichol and John Schulman · 2018
Earlier work this paper cites.
Meta-gradient reinforcement learning
Zeyuan Xu, Hado van Hasselt, and David Silver · 2018
Earlier work this paper cites.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2019
Earlier work this paper cites.
Textworld: A learning environment for text-based games
Marc-Alexandre Côté et al · 2019
Cited alongside, same era.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Chelsea Finn, Sergey Levine, and Deirdre Quillen · 2019
Cited alongside, same era.
Habitat: A platform for embodied ai research
Manolis Savva et al · 2019
Cited alongside, same era.
AlphaStar: Mastering the real-time strategy game StarCraft II
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michael Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Grounding large language models in interactive environments with online reinforcement learning
Thomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier, Olivier Sigaud, and Pierre-Yves Oudeyer · 2023
Later among the works it cites.
Guiding pretraining in reinforcement learning with large language models
Yuqing Du, Olivia Watkins, Zihan Wang, Cédric Colas, Trevor Darrell, Pieter Abbeel, Abhishek Gupta, and Jacob Andreas · 2023
Later among the works it cites.
Mastering diverse domains through world models
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap · 2023
Later among the works it cites.
Temporal difference learning for model predictive control
Nicklas Hansen, Haochen Wang, and Xiaoxing Su · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Cote, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht · 2020
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Cited alongside, same era.
Maniskill: Generalizable manipulation skill benchmark with large-scale demonstrations
Tongzhou Mu, Zhan Ling, Fanbo Xiang, Derek Yang, Xuanlin Li, Stone Gray, Andrew Viswanath, Pete Roberson, Shixiang Zhao, Juan Li, Hao Gao, Jiahong Zhang, Zilin Wang, Peter Liang, and Sergey Levine · 2021
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Edward Hu et al · 2022
Cited alongside, same era.
Planning with diffusion for flexible behavior synthesis
Michael Janner, Qiyang Li, and Sergey Levine · 2022
Cited alongside, same era.
Retro-rl: Reinforcing nominal controller with deep reinforcement learning for tilting-rotor drones
I Made Aswin Nahrendra, Christian Tirtawardhana, Byeongho Yu, Eungchang Mason Lee, and Hyun Myung · 2022
Cited alongside, same era.
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu · 2023
Later among the works it cites.
Reward design with language models
Minae Kwon, Sang Michael Xie, Kalesha Bullard, and Dorsa Sadigh · 2023
Later among the works it cites.
Eureka: Human-level reward design via coding large language models
Yecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang, Osbert Bastani, Dinesh Jayaraman, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2023
Later among the works it cites.
Voyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2023
Later among the works it cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao et al · 2023
Later among the works it cites.
Building cooperative embodied agents modularly with large language models
Hongxin Zhang, Weihua Tian, Jiaxu Du, Yucheng Zhao, Rui Xia, Rui Ding, Xin Ma, Tao Min, Wenqiang Xue, Yihuai Zhu, et al · 2023
Later among the works it cites.
Efficient reinforcement learning with large language model priors
Xue Yan, Yan Song, Xidong Feng, Mengyue Yang, Haifeng Zhang, Haitham Bou Ammar, and Jun Wang · 2024
Later among the works it cites.
Mulberry: A multimodal reasoning model with structured representations
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Later among the works it cites.