Fetching the paper…
Reading the bibliography…
In sequential decision-making (SDM) tasks, methods like reinforcement learning (RL) and heuristic search have made notable advances in specific cases.
Model predictive control: Theory and practice—a survey
Carlos E Garcia, David M Prett, and Manfred Morari · 1989
Earlier work this paper cites.
Power-steering control architecture for automatic driving
José Eugenio Naranjo, Carlos González, Ricardo García, Teresa de Pedro, and Rodolfo E Haber · 2005
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih · 2013
Earlier work this paper cites.
Deep reinforcement learning in large discrete action spaces
Gabriel Dulac-Arnold, Richard Evans, Hado van Hasselt, Peter Sunehag, Timothy Lillicrap, Jonathan Hunt, Timothy Mann, Theophane Weber, Thomas Degris, and Ben Coppin · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Christopher J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Earlier work this paper cites.
Alphago, deep learning, and the future of the human microscopist
Scott R Granter, Andrew H Beck, and David J Papke Jr · 2017
Earlier work this paper cites.
Survey of model-based reinforcement learning: Applications on robotics
Athanasios S Polydoros and Lazaros Nalpantidis · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Reinforcement learning and control as probabilistic inference: Tutorial and review
Sergey Levine · 2018
Earlier work this paper cites.
Quantifying generalization in reinforcement learning
Karl Cobbe, Oleg Klimov, Chris Hesse, Taehoon Kim, and John Schulman · 2019
Earlier work this paper cites.
Textworld: A learning environment for text-based games
Marc-Alexandre Côté, Akos Kádár, Xingdi Yuan, Ben Kybartas, Tavian Barnes, Emery Fine, James Moore, Matthew Hausknecht, Layla El Asri, Mahmoud Adada, et al · 2019
Earlier work this paper cites.
Dialogue generation: From imitation learning to inverse reinforcement learning
Ziming Li, Julia Kiseleva, and Maarten De Rijke · 2019
Earlier work this paper cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Audrey Dudzik, Junyoung Chung, Arthur J Tashman, David Horgan, Peter Kroiss, Max Powell, et al · 2019
Earlier work this paper cites.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht · 2020
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Earlier work this paper cites.
Fudge: Controlled text generation with future discriminators
Kevin Yang and Dan Klein · 2021
Earlier work this paper cites.
Safe learning in robotics: From learning-based control to safe reinforcement learning
Lukas Brunke, Melissa Greeff, Adam W Hall, Zhaocong Yuan, Siqi Zhou, Jacopo Panerati, and Angela P Schoellig · 2022
Cited alongside, same era.
Rl with kl penalties is better viewed as bayesian inference
Tomasz Korbak, Ethan Perez, and Christopher L Buckley · 2022
Cited alongside, same era.
Conversational ai: Dialogue systems, conversational agents, and chatbots
Michael McTear · 2022
Cited alongside, same era.
Quantum decision making in automatic driving
Qingyuan Song, Weiping Fu, Wen Wang, Yuan Sun, Denggui Wang, and Jincao Zhou · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2022
Cited alongside, same era.
Monte carlo tree search: A review of recent modifications and applications
Maciej Świechowski, Konrad Godlewski, Bartosz Sawicki, and Jacek Mańdziuk · 2023
Later among the works it cites.
Increasing brain-llm alignment via information-theoretic compression
Mycal Tucker and Greta Tuckute · 2023
Later among the works it cites.
Voyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2023
Later among the works it cites.
Xue Yan, Yan Song, Xinyu Cui, Filippos Christianos, Haifeng Zhang, David Henry Mguni, and Jun Wang · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al · 2023
Cited alongside, same era.
Large language models can implement policy iteration
Ethan Brooks, Logan A Walls, Richard Lewis, and Satinder Singh · 2023
Cited alongside, same era.
Grounding large language models in interactive environments with online reinforcement learning
Thomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier, Olivier Sigaud, and Pierre-Yves Oudeyer · 2023
Cited alongside, same era.
Pangu-agent: A fine-tunable generalist agent with structured reasoning, 2023
Filippos Christianos, Georgios Papoudakis, Matthieu Zimmer, Thomas Coste, Zhihao Wu, Jingxuan Chen, Khyati Khandelwal, James Doran, Xidong Feng, Jiacheng Liu, Zheng Xiong, Yicheng Luo, Jianye Hao, Kun Shao, Haitham Bou-Ammar, and Jun Wang · 2023
Cited alongside, same era.
Alphazero-like tree-search can guide large language model decoding and training
Xidong Feng, Ziyu Wan, Muning Wen, Ying Wen, Weinan Zhang, and Jun Wang · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu · 2023
Cited alongside, same era.
Amortizing intractable inference in large language models
Edward J Hu, Moksh Jain, Eric Elmoznino, Younesse Kaddar, Guillaume Lajoie, Yoshua Bengio, and Nikolay Malkin · 2023
Cited alongside, same era.
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan · 2023
Later among the works it cites.
Proagent: Building proactive cooperative ai with large language models
Ceyao Zhang, Kaijie Yang, Siyi Hu, Zihao Wang, Guanghe Li, Yihang Sun, Cheng Zhang, Zhaowei Zhang, Anji Liu, Song-Chun Zhu, et al · 2023
Later among the works it cites.
Large language models can implement policy iteration
Ethan Brooks, Logan Walls, Richard L Lewis, and Satinder Singh · 2024
Closest in time.
Large language monkeys: Scaling inference compute with repeated sampling
Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le, Christopher Ré, and Azalia Mirhoseini · 2024
Closest in time.
Stochastic q-learning for large discrete action spaces
Fares Fourati, Vaneet Aggarwal, and Mohamed-Slim Alouini · 2024
Closest in time.
Distilled self-critique of llms with synthetic data: a bayesian perspective
Victor Gallego · 2024
Closest in time.
Q-probe: A lightweight approach to reward maximization for language models
Kenneth Li, Samy Jelassi, Hugh Zhang, Sham Kakade, Martin Wattenberg, and David Brandfonbrener · 2024
Closest in time.
Llm critics help catch llm bugs
Nat McAleese, Rai Michael Pokorny, Juan Felipe Ceron Uribe, Evgenia Nitishinskaya, Maja Trebacz, and Jan Leike · 2024
Closest in time.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Closest in time.
Entropy-regularized token-level policy optimization for large language models
Muning Wen, Cheng Deng, Jun Wang, Weinan Zhang, and Ying Wen · 2024
Closest in time.
An efficient end-to-end training approach for zero-shot human-ai coordination
Xue Yan, Jiaxian Guo, Xingzhou Lou, Jun Wang, Haifeng Zhang, and Yali Du · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Closest in time.
Pre-training and in-context learning is bayesian inference a la de finetti
Naimeng Ye, Hanming Yang, Andrew Siah, and Hongseok Namkoong · 2024
Closest in time.
Reflect-rl: Two-player online rl fine-tuning for lms
Runlong Zhou, Simon S Du, and Beibin Li · 2024
Closest in time.