Fetching the paper…
Reading the bibliography…
This survey organizes the intricate literature on the design and optimization of emerging structures around post-trained LMs.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift, 2019
Yaniv Ovadia, Emily Fertig, Jie Ren, et al · 1906
Earlier work this paper cites.
The early phase of neural network training, 2020
Jonathan Frankle, David J Schwab, Ari S Morcos · 2002
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, et al · 2010
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, Geoffrey Hinton · 2015
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, et al · 2016
Earlier work this paper cites.
Reinforcement Learning: An Introduction
Richard S Sutton, Andrew G Barto · 2018
Earlier work this paper cites.
A survey on deep learning: Algorithms, techniques, and applications
Samira Pouyanfar, Saad Sadiq, Yilin Yan, et al · 2018
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, et al · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, et al · 2019
Earlier work this paper cites.
Critical learning periods in deep neural networks, 2019
Alessandro Achille, Matteo Rovere, Stefano Soatto · 2019
Earlier work this paper cites.
Embracing change: Continual learning in deep neural networks
Raia Hadsell, Dushyant Rao, Andrei A Rusu, et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, et al · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, et al · 2020
Earlier work this paper cites.
Hover: A dataset for many-hop fact extraction and claim verification
Yichen Jiang, Shikha Bordia, Zheng Zhong, et al · 2020
Earlier work this paper cites.
Reinforcement learning is supervised learning on optimized data, 2020
Ben Eysenbach, Aviral Kumar, Abhishek Gupta · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the MATH dataset
Dan Hendrycks, Collin Burns, Saurav Kadavath, et al · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, et al · 2021
Earlier work this paper cites.
Leveraging passage retrieval with generative models for open domain question answering
Gautier Izacard, Edouard Grave · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, et al · 2022
Earlier work this paper cites.
Rlprompt: Optimizing discrete text prompts with reinforcement learning, 2022
Mingkai Deng, Jianyu Wang, Cheng-Ping Hsieh, et al · 2022
Earlier work this paper cites.
TEMPERA: test-time prompting via reinforcement learning
Tianjun Zhang, Xuezhi Wang, Denny Zhou, et al · 2022
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback, 2022
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, et al · 2022
Earlier work this paper cites.
Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive NLP
Omar Khattab, Keshav Santhanam, Xiang Lisa Li, et al · 2022
Earlier work this paper cites.
Language models as agent models
Jacob Andreas · 2022
Earlier work this paper cites.
A mixture of surprises for unsupervised reinforcement learning
Andrew Zhao, Matthieu Lin, Yangguang Li, et al · 2022
Earlier work this paper cites.
The primacy bias in deep reinforcement learning, 2022
Evgenii Nikishin, Max Schwarzer, Pierluca D’Oro, et al · 2022
Earlier work this paper cites.
No, to the right: Online language corrections for robotic manipulation via shared autonomy
Yuchen Cui, Siddharth Karamcheti, Raj Palleti, et al · 2023
Earlier work this paper cites.
Llf-bench: Benchmark for interactive learning from language feedback, 2023
Ching-An Cheng, Andrey Kolobov, Dipendra Misra, et al · 2023
Earlier work this paper cites.
Dspy: Compiling declarative language model calls into self-improving pipelines
Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, et al · 2023
Earlier work this paper cites.
Large language models are human-level prompt engineers
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, et al · 2023
Earlier work this paper cites.
Automatic prompt optimization with "gradient descent" and beam search
Reid Pryzant, Dan Iter, Jerry Li, et al · 2023
Earlier work this paper cites.
Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation
Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, et al · 2023
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries, 2023
Patrick Chao, Alexander Robey, Edgar Dobriban, et al · 2023
Earlier work this paper cites.
R Thomas McCoy, Shunyu Yao, Dan Friedman, et al · 2023
Earlier work this paper cites.
Studying large language model generalization with influence functions, 2023
Roger Grosse, Juhan Bae, Cem Anil, et al · 2023
Earlier work this paper cites.
In-context retrieval-augmented language models
Ori Ram, Yoav Levine, Itay Dalmedigos, et al · 2023
Earlier work this paper cites.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph C O’Brien, Carrie Jun Cai, et al · 2023
Earlier work this paper cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, et al · 2023
Earlier work this paper cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, et al · 2023
Earlier work this paper cites.
Transformers learn in-context by gradient descent
Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, et al · 2023
Earlier work this paper cites.
Reflexion: language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, et al · 2023
Earlier work this paper cites.
Train once, get a family: State-adaptive balances for offline-to-online reinforcement learning
Shenzhi Wang, Qisen Yang, Jiawei Gao, et al · 2023
Earlier work this paper cites.
Theory of mind in large language models: Examining performance of 11 state-of-the-art models vs. children aged 7–10 on advanced tests
Max van Duijn, Bram van Dijk, Tom Kouwenhoven, et al · 2023
Earlier work this paper cites.
Vima: General robot manipulation with multimodal prompts
Yunfan Jiang, Agrim Gupta, Zichen Zhang, et al · 2023
Earlier work this paper cites.
OpenHands: An Open Platform for AI Software Developers as Generalist Agents, 2024
Xingyao Wang, Boxuan Li, Yufan Song, et al · 2024
Earlier work this paper cites.
MLE-bench: Evaluating machine learning agents on machine learning engineering
Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, et al · 2024
Earlier work this paper cites.
What are tools anyway? a survey from the language model perspective
Zhiruo Wang, Zhoujun Cheng, Hao Zhu, et al · 2024
Earlier work this paper cites.
Expel: Llm agents are experiential learners
Andrew Zhao, Daniel Huang, Quentin Xu, et al · 2024
Earlier work this paper cites.
Claude, 2024
Anthropic · 2024
Earlier work this paper cites.
Chatgpt, 2024
OpenAI · 2024
Earlier work this paper cites.
The shift from models to compound ai systems, 2024
Matei Zaharia, Omar Khattab, Lingjiao Chen, et al · 2024
Earlier work this paper cites.
Language agents: From next-token prediction to digital automation
Shunyu Yao · 2024
Cited alongside, same era.
The landscape of emerging AI agent architectures for reasoning, planning, and tool calling: A survey
Tula Masterman, Sandi Besen, Mason Sawtell, et al · 2024
Cited alongside, same era.
Trace is the new autodiff - unlocking efficient optimization of computational workflows
Ching-An Cheng, Allen Nie, Adith Swaminathan · 2024
Cited alongside, same era.
AdalFlow: The Library for Large Language Model (LLM) Applications, 2024
Li Yin · 2024
Cited alongside, same era.
Llm-based optimization of compound ai systems: A survey, 2024
Matthieu Lin, Jenny Sheng, Andrew Zhao, et al · 2024
Cited alongside, same era.
Function calling, 2025
OpenAI · 2025
Closest in time.
Optimizing generative ai by backpropagating language model feedback
Mert Yuksekgonul, Federico Bianchi, Joseph Boen, et al · 2025
Closest in time.
Inducing programmatic skills for agentic tasks, 2025
Zora Zhiruo Wang, Apurva Gandhi, Graham Neubig, et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI, Daya Guo, Dejian Yang, et al · 2025
Closest in time.
Are my optimized prompts compromised? exploring vulnerabilities of llm-based optimizers, 2025
Andrew Zhao, Reshmi Ghosh, Vitor Carvalho, et al · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, et al · 2024
Cited alongside, same era.
Promptagent: Strategic planning with language models enables expert-level prompt optimization
Xinyuan Wang, Chenxi Li, Zhen Wang, et al · 2024
Cited alongside, same era.
Andrew Zhao, Quentin Xu, Matthieu Lin, et al · 2024
Cited alongside, same era.
Prewrite: Prompt rewriting with reinforcement learning, 2024
Weize Kong, Spurthi Amba Hombaiah, Mingyang Zhang, et al · 2024
Cited alongside, same era.
Stableprompt: Automatic prompt tuning using reinforcement learning for large language models, 2024
Minchan Kwon, Gaeun Kim, Jongsuk Kim, et al · 2024
Cited alongside, same era.
Agentless: Demystifying llm-based software engineering agents
Chunqiu Steven Xia, Yinlin Deng, Soren Dunn, et al · 2024
Cited alongside, same era.
SWE-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, et al · 2024
Cited alongside, same era.
New tools for building agents, 2025
OpenAI · 2025
Closest in time.
The rise and potential of large language model based agents: a survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, et al · 2025
Closest in time.
The second half, 2025
Shunyu Yao · 2025
Closest in time.
A survey on the feedback mechanism of LLM-based AI agents
Zhipeng Liu, Xuefeng Bai, Kehai Chen, et al · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI · 2025
Closest in time.
A survey on post-training of large language models, 2025
Guiyao Tie, Zeli Zhao, Dingjie Song, et al · 2025
Closest in time.
A survey of post-training scaling in large language models
Kenny Shijian Lai, Jasur Mirzakhalov, Karan Singla, et al · 2025
Closest in time.
A survey on the optimization of large language model-based agents, 2025
Shangheng Du, Jiabao Zhao, Jinxin Shi, et al · 2025
Closest in time.
Absolute zero: Reinforced self-play reasoning with zero data, 2025
Andrew Zhao, Yiran Wu, Yang Yue, et al · 2025
Closest in time.
Reinforcement learning for reasoning in large language models with one training example, 2025
Yiping Wang, Qing Yang, Zhiyuan Zeng, et al · 2025
Closest in time.
Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution
Yuxiang Wei, Olivier Duchenne, Jade Copet, et al · 2025
Closest in time.
Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions
Terry Yue Zhuo, Vu Minh Chien, Jenny Chim, et al · 2025
Closest in time.
Asymmetry of verification and verifier’s rule, 2025
Jason Wei · 2025
Closest in time.
AFlow: Automating agentic workflow generation
Jiayi Zhang, Jinyu Xiang, Zhaoyang Yu, et al · 2025
Closest in time.
Openai harmony response format
Dominik Kundel · 2025
Closest in time.
Codex cloud: Internet access, 2025
OpenAI · 2025
Closest in time.
Claude’s system prompt: Chatbots are more than just models, 2025
Drew Breunig · 2025
Closest in time.
Taming knowledge conflicts in language models, 2025
Gaotang Li, Yuzhong Chen, Hanghang Tong · 2025
Closest in time.
Claude system prompt leak, 2025
Ásgeir Thor Johnson · 2025
Closest in time.
The "think" tool: Enabling claude to stop and think in complex tool use situations, 2025
Anthropic · 2025
Closest in time.
Human as a tool, 2025
LangChain · 2025
Closest in time.
Openai agents sdk, 2025
OpenAI · 2025
Closest in time.
gpt-oss-120b & gpt-oss-20b model card, 2025
OpenAI · 2025
Closest in time.
Operator system card
OpenAI · 2025
Closest in time.
Pou: Proof-of-use to counter tool-call hacking in deepresearch agents, 2025
Shengjie Ma, Chenlong Deng, Jiaxin Mao, et al · 2025
Closest in time.
How to think about agent frameworks, 2025
LangChain · 2025
Closest in time.
The ai scientist-v2: Workshop-level automated scientific discovery via agentic tree search
Yutaro Yamada, Robert Tjarko Lange, Cong Lu, et al · 2025
Closest in time.
Windsurf editor
Windsurf · 2025
Closest in time.
Cursor: The ai-powered code editor, 2023
Cursor · 2025
Closest in time.
Copilot (gpt-4) [large language model], 2025
Microsoft · 2025
Closest in time.
Swe-lancer: Can frontier llms earn $1 million from real-world freelance software engineering?, 2025
Samuel Miserendino, Michele Wang, Tejal Patwardhan, et al · 2025
Closest in time.
Sweet-rl: Training multi-turn llm agents on collaborative reasoning tasks, 2025
Yifei Zhou, Song Jiang, Yuandong Tian, et al · 2025
Closest in time.
Browsecomp: A simple yet challenging benchmark for browsing agents, 2025
Jason Wei, Zhiqing Sun, Spencer Papay, et al · 2025
Closest in time.
Paperbench: Evaluating ai’s ability to replicate ai research, 2025
Giulio Starace, Oliver Jaffe, Dane Sherburn, et al · 2025
Closest in time.
Hcast: Human-calibrated autonomy software tasks, 2025
David Rein, Joel Becker, Amy Deng, et al · 2025
Closest in time.
A survey of automatic prompt engineering: An optimization perspective, 2025
Wenwu Li, Xiangfeng Wang, Wenhao Li, et al · 2025
Closest in time.
Contextual experience replay for continual learning of language agents, 2025
Yitao Liu, Chenglei Si, Karthik R Narasimhan, et al · 2025
Closest in time.
Skillweaver: Web agents can self-improve by discovering and honing skills, 2025
Boyuan Zheng, Michael Y Fatemi, Xiaolong Jin, et al · 2025
Closest in time.
How to correctly do semantic backpropagation on language-based agentic systems, 2025
Wenyi Wang, Hisham Abdullah Alyahya, Dylan R Ashley, et al · 2025
Closest in time.
Measuring ai ability to complete long tasks, 2025
Thomas Kwa, Ben West, Joel Becker, et al · 2025
Closest in time.
Bang Liu, Xinfeng Li, Jiayi Zhang, et al · 2025
Closest in time.
Openai memory announcement, 2025
OpenAI · 2025
Closest in time.
Dynamic cheatsheet: Test-time learning with adaptive memory, 2025
Mirac Suzgun, Mert Yuksekgonul, Federico Bianchi, et al · 2025
Closest in time.
The bitter lesson, 2019
Richard S Sutton · 2025
Closest in time.
Sleep-time compute: Beyond inference scaling at test-time, 2025
Kevin Lin, Charlie Snell, Yu Wang, et al · 2025
Closest in time.
Lattice: Learning to efficiently compress the memory, 2025
Mahdi Karami, Vahab Mirrokni · 2025
Closest in time.
Langprobe: a language programs benchmark, 2025
Shangyin Tan, Lakshya A Agrawal, Arnav Singhvi, et al · 2025
Closest in time.