Fetching the paper…
Reading the bibliography…
The rapid evolution of agentic AI marks a new phase in artificial intelligence, where Large Language Models (LLMs) no longer merely respond but act, reason, and adapt.
Strips: A new approach to the application of theorem proving to problem solving
Richard E. Fikes and Nils J. Nilsson. 1971 · 1971
Earlier work this paper cites.
Support-Vector Networks
Corinna Cortes and Vladimir Vapnik. 1995 · 1995
Earlier work this paper cites.
PDDL - The Planning Domain Definition Language
Malik Ghallab, Craig Knoblock, David Wilkins, Anthony Barrett, and Daniel Weld. 1998 · 1998
Earlier work this paper cites.
Longformer: The Long-Document Transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Dense Passage Retrieval for Open-Domain Question Answering
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen tau Yih. 2020 · 2004
Earlier work this paper cites.
Automation and customization of rendered web pages. In Proceedings of the 18th annual ACM symposium on User interface software and technology . 163–172
Michael Bolin, Matthew Webber, Philip Rha, Tom Wilson, and Robert C Miller. 2005 · 2005
Earlier work this paper cites.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, et al · 2005
Earlier work this paper cites.
CoScripter: Sharing ‘How-to’Knowledge in the Enterprise
Gilly Leshed, Eben Haber, Tessa Lau, and Allen Cypher. 2007 · 2007
Earlier work this paper cites.
Rethinking Attention with Performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, David Belanger, Lucy Colwell, and Adrian Weller. 2022 · 2009
Earlier work this paper cites.
Sikuli: using GUI screenshots for search and automation. In Proceedings of the 22nd annual ACM symposium on User interface software and technology . 183–192
Tom Yeh, Tsung-Hsiang Chang, and Robert C Miller. 2009 · 2009
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks. In Proceedings of the 25th International Conference on Neural Information Processing Systems (NeurIPS) . 1097–1105
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012 · 2012
Earlier work this paper cites.
Reran: Timing-and touch-sensitive record and replay for android. In 2013 35th international conference on software engineering (ICSE) . IEEE, 72–81
Lorenzo Gomez, Iulian Neamtiu, Tanzirul Azim, and Todd Millstein. 2013 · 2013
Earlier work this paper cites.
Ringer: web automation by demonstration. In Proceedings of the 2016 ACM SIGPLAN international conference on object-oriented programming, systems, languages, and applications . 748–764
Shaon Barman, Sarah Chasins, Rastislav Bodik, and Sumit Gulwani. 2016 · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, et al · 2016
Earlier work this paper cites.
SUGILITE: creating multimodal smartphone automation by demonstration. In Proceedings of the 2017 CHI conference on human factors in computing systems . 6038–6049
Toby Jia-Jun Li, Amos Azaria, and Brad A Myers. 2017 · 2017
Earlier work this paper cites.
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell. 2017 · 2017
Earlier work this paper cites.
Proximal Policy Optimization Algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Attention Is All You Need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS) . Curran Associates, Inc., 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Improving Language Understanding by Generative PreTraining
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Earlier work this paper cites.
Sara: self-replay augmented record and replay for android in industrial cases. In Proceedings of the 28th acm sigsoft international symposium on software testing and analysis . 90–100
Jiaqi Guo, Shuyue Li, Jian-Guang Lou, Zijiang Yang, and Ting Liu. 2019 · 2019
Earlier work this paper cites.
Big bird: transformers for longer sequences. In Proceedings of the 34th International Conference on Neural Information Processing Systems (Vancouver, BC, Canada) (NIPS ’20) . Curran Associates Inc., Red Hook, NY, USA, Article 1450, 15 pages
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020 · 2020
Earlier work this paper cites.
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, et al · 2021
Earlier work this paper cites.
WebGPT: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2021 · 2021
Earlier work this paper cites.
RoFormer: Enhanced Transformer with Rotary Position Embedding
Jianlin Su et al · 2021
Earlier work this paper cites.
K-Adapter: Infusing Knowledge into Pre-Trained Models with Adapters. In Findings of ACL-IJCNLP
Ruize Wang, Duyu Tang, Nan Duan, Zhongyu Wei, Xuanjing Huang, Jianshu Ji, Guihong Cao, Daxin Jiang, and Ming Zhou. 2021 · 2021
Earlier work this paper cites.
Reactive synthesis of software robots in RPA from user interface logs
Simone Agostinelli, Marco Lupia, Andrea Marrella, and Massimo Mecella. 2022 · 2022
Earlier work this paper cites.
Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, et al · 2022
Earlier work this paper cites.
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. In NeurIPS
Tri Dao et al · 2022
Earlier work this paper cites.
Langchain
Chase Harrison. 2022-10 · 2022
Earlier work this paper cites.
Atlas: Few-shot Learning with Retrieval Augmented Language Models
Gautier Izacard, Patrick Lewis, Maria Lomeli, et al · 2022
Earlier work this paper cites.
Internet-Augmented Dialogue Generation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Ireland, 8460–8478
Mojtaba Komeili, Kurt Shuster, and Jason Weston. 2022 · 2022
Earlier work this paper cites.
Locating and Editing Factual Associations in GPT. In NeurIPS
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022 · 2022
Earlier work this paper cites.
Mass-Editing Memory in a Transformer
Kevin Meng and Arnold P. et al. 2022 · 2022
Earlier work this paper cites.
Memory-Based Model Editing at Scale. In ICML
Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D. Manning, and Chelsea Finn. 2022 · 2022
Earlier work this paper cites.
ChatGPT: Optimizing Language Models for Dialogue
OpenAI. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, et al · 2022
Earlier work this paper cites.
Measuring and Narrowing the Compositionality Gap in Language Models
Ofir Press et al · 2022
Earlier work this paper cites.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V. Le, and Denny Zhou. 2022 · 2022
Earlier work this paper cites.
Memorizing Transformers. In ICLR
Yuhuai Wu, Markus N. Rabe, DeLesley Hutchins, and Christian Szegedy. 2022 · 2022
Earlier work this paper cites.
Beyond Goldfish Memory: Long-Term Open-Domain Conversation. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, Dublin, Ireland, 5180–5197
Jing Xu, Arthur Szlam, and Jason Weston. 2022 · 2022
Earlier work this paper cites.
WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents. In Advances in Neural Information Processing Systems 35 (NeurIPS 2022)
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. 2022 · 2022
Earlier work this paper cites.
Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2023 · 2023
Earlier work this paper cites.
Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
Wenhu Chen, Xueguang Ma, Xinyi Wang, and William W. Cohen. 2023 · 2023
Earlier work this paper cites.
COOPER: Coordinating Specialized Agents towards a Complex Dialogue Goal
Yi Cheng, Wenge Liu, Jian Wang, Chak Tou Leong, Yi Ouyang, Wenjie Li, Xian Wu, and Yefeng Zheng. 2023 · 2023
Earlier work this paper cites.
PAL: Program-aided Language Models. In Proceedings of the 40th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 202) , Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). PMLR, 10764–10799
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023 · 2023
Earlier work this paper cites.
Critic: Large language models can self-correct with tool-interactive critiquing
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2023 · 2023
Earlier work this paper cites.
Leveraging pre-trained large language models to construct and utilize world models for model-based task planning. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS ’23) . Curran Associates Inc., Red Hook, NY, USA, Article 3459, 14 pages
Lin Guan, Karthik Valmeekam, Sarath Sreedharan, and Subbarao Kambhampati. 2023 · 2023
Earlier work this paper cites.
A real-world webagent with planning, long context understanding, and program synthesis
Izzeddin Gur, Hiroki Furuta, Austin Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust. 2023 · 2023
Earlier work this paper cites.
Reasoning with Language Model is Planning with World Model. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 8154–8173
Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Wang, and Zhiting Hu. 2023 · 2023
Earlier work this paper cites.
MetaGPT: Meta programming for a multi-agent collaborative framework. In The Twelfth International Conference on Learning Representations
Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, et al · 2023
Earlier work this paper cites.
LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 13358–13376
Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023a · 2023
Earlier work this paper cites.
Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression
Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu. 2023b · 2023
Earlier work this paper cites.
Language models can solve computer tasks
Geunwoo Kim, Pierre Baldi, and Stephen McAleer. 2023 · 2023
Earlier work this paper cites.
API Reference for langchain.memory.summary.ConversationSummaryMemory
Langchain Team. 2023 · 2023
Earlier work this paper cites.
CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society
Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. 2023b · 2023
Earlier work this paper cites.
Code as Policies: Language Model Programs for Embodied Control. In 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 9493–9500
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. 2023b · 2023
Earlier work this paper cites.
Encouraging divergent thinking in large language models through multi-agent debate
Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Shuming Shi, and Zhaopeng Tu. 2023a · 2023
Earlier work this paper cites.
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2023b · 2023
Earlier work this paper cites.
Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models. In Advances in Neural Information Processing Systems , Vol. 36
Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, and Jianfeng Gao. 2023b · 2023
Earlier work this paper cites.
Query rewriting in retrieval-augmented large language models. In The 2023 Conference on Empirical Methods in Natural Language Processing
Xinbei Ma, Yeyun Gong, Pengcheng He, Nan Duan, et al · 2023
Earlier work this paper cites.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, et al · 2023
Earlier work this paper cites.
Function calling and other API updates
OpenAI. 2023 · 2023
Earlier work this paper cites.
Generative Agents: Interactive Simulacra of Human Behavior
Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023 · 2023
Earlier work this paper cites.
Measuring and Narrowing the Compositionality Gap in Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapore, 5687–5711
Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah Smith, and Mike Lewis. 2023 · 2023
Earlier work this paper cites.
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. 2023 · 2023
Earlier work this paper cites.
HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face. In Advances in Neural Information Processing Systems , Vol. 36
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2023 · 2023
Earlier work this paper cites.
S-LoRA: Serving Thousands of Concurrent LoRA Adapters
Yifan Sheng et al · 2023
Earlier work this paper cites.
Reflexion: language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curran Associates, Inc., 8634–8652
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023 · 2023
Earlier work this paper cites.
Suno: AI Music Generation Platform
Suno. 2023 · 2023
Earlier work this paper cites.
ViperGPT: Visual Inference via Python Execution for Reasoning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
G. Surís, S. Adeli, D. Bashkirova, J. Carreira, and C. Vondrick. 2023 · 2023
Earlier work this paper cites.
Freshllms: Refreshing large language models with search engine augmentation
Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, et al · 2023
Earlier work this paper cites.
Query2doc: Query expansion with large language models
Liang Wang, Nan Yang, and Furu Wei. 2023b · 2023
Earlier work this paper cites.
Generative query reformulation for effective adhoc search
Xiao Wang, Sean MacAvaney, Craig Macdonald, and Iadh Ounis. 2023a · 2023
Earlier work this paper cites.
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Qingyun Wu et al · 2023
Earlier work this paper cites.
Recomp: Improving retrieval-augmented lms with compression and selective augmentation
Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2023 · 2023
Earlier work this paper cites.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In Advances in Neural Information Processing Systems (NeurIPS) 2023 . 11809–11822
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023a · 2023
Earlier work this paper cites.
AppAgent: Multimodal Agents as Smartphone Users
Chi Zhang, Zhao Yang, Jiaxuan Liu, Yucheng Han, Xin Chen, Zebiao Huang, Bin Fu, and Gang Yu. 2023b · 2023
Earlier work this paper cites.
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, and Beidi Chen. 2023a · 2023
Earlier work this paper cites.
MemoryBank: Enhancing Large Language Models with Long-Term Memory
Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang. 2023 · 2023
Earlier work this paper cites.
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, et al · 2024
Earlier work this paper cites.
Web agents with world models: Learning and leveraging environment dynamics in web navigation
Hyungjoo Chae, Namyoung Kim, Kai Tzu-iunn Ong, Minju Gwak, Gwanwoo Song, Jihoon Kim, Sunghwan Kim, Dongha Lee, and Jinyoung Yeo. 2024 · 2024
Earlier work this paper cites.
GUIDE: Graphical User Interface Data for Execution
Rajat Chawla, Adarsh Jha, Muskaan Kumar, Mukunda NS, and Ishaan Bhola. 2024 · 2024
Cited alongside, same era.
Justin Chih-Yao Chen, Swarnadeep Saha, Elias Stengel-Eskin, and Mohit Bansal. 2024b · 2024
Cited alongside, same era.
SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Kanzhi Cheng, Qiushi Sun, Yougang Chu, Fangzhi Xu, Yantao Li, Jianbing Zhang, and Zhiyong Wu. 2024 · 2024
Cited alongside, same era.
Don’t teach. Incentivize
Hyung Won Chung. 2024 · 2024
Cited alongside, same era.
SKVQ: Sliding-window Key and Value Cache Quantization for Large Language Models
Llms get lost in multi-turn conversation
Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, and Jennifer Neville. 2025 · 2025
Closest in time.
Bespoke-Stratos: The unreasonable effectiveness of reasoning distillation
Bespoke Labs. 2025 · 2025
Closest in time.
WebSailor: Navigating Super-human Reasoning for Web Agent
Kuan Li, Zhongwang Zhang, Huifeng Yin, Rui Ye, Yida Zhao, Liwen Zhang, Litu Ou, Dingchu Zhang, et al · 2025
Closest in time.
Shiyuan Li, Yixin Liu, Qingsong Wen, Chengqi Zhang, and Shirui Pan. 2025g · 2025
Closest in time.
Chain-of-agents: End-to-end agent foundation models via multi-agent distillation and agentic rl
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Haojie Duanmu, Zhihang Yuan, Xiuhong Li, Jiangfei Duan, Xingcheng Zhang, and Dahua Lin. 2024 · 2024
Cited alongside, same era.
From local to global: A graph rag approach to query-focused summarization
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2024 · 2024
Cited alongside, same era.
AI Agents are about 90% software engineering and 10% AI. Do you agree?
Rakesh Gohel. 2024 · 2024
Cited alongside, same era.
CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing. In Proceedings of the International Conference on Learning Representations (ICLR)
Zhibin Gou, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. 2024 · 2024
Cited alongside, same era.
A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis. In The Twelfth International Conference on Learning Representations
Izzeddin Gur, Hiroki Furuta, Austin V Huang, Mustafa Safdari, Yutaka Matsuo, Douglas Eck, and Aleksandra Faust. 2024 · 2024
Cited alongside, same era.
Bolei He, Nuo Chen, Xinran He, Lingyong Yan, Zhenkai Wei, Jinchang Luo, and Zhen-Hua Ling. 2024a · 2024
Cited alongside, same era.
Webvoyager: Building an end-to-end web agent with large multimodal models
Hongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Yong Dai, Hongming Zhang, Zhenzhong Lan, and Dong Yu. 2024b · 2024
Cited alongside, same era.
CogAgent: A Visual Language Model for GUI Agents. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 14281–14290
Wenyi Hong, Weihan Wang, Qingsong Lv, Jiazheng Xu, Wenmeng Yu, Junhui Ji, Yan Wang, Zihan Wang, et al · 2024
Cited alongside, same era.
Weizhen Li, Jianbo Lin, Zhuosong Jiang, Jingyi Cao, Xinpeng Liu, Jiayu Zhang, Zhenqiang Huang, Qianben Chen, et al · 2025
Closest in time.
Search-o1: Agentic search-enhanced large reasoning models
Xiaoxi Li, Guanting Dong, Jiajie Jin, Yuyao Zhang, Yujia Zhou, Yutao Zhu, Peitian Zhang, and Zhicheng Dou. 2025a · 2025
Closest in time.
Webthinker: Empowering large reasoning models with deep research capability
Xiaoxi Li, Jiajie Jin, Guanting Dong, Hongjin Qian, Yutao Zhu, Yongkang Wu, Ji-Rong Wen, and Zhicheng Dou. 2025e · 2025
Closest in time.
ToRL: Tool-integrated Reinforcement Learning for LLM Agents
Xuefeng Li, Haoyang Zou, Pengfei Liu, et al · 2025
Closest in time.
Marft: Multi-agent reinforcement fine-tuning
Junwei Liao, Muning Wen, Jun Wang, and Weinan Zhang. 2025 · 2025
Closest in time.
Mind with Eyes: from Language Reasoning to Multimodal Reasoning
Zhiyu Lin, Yifei Gao, Xian Zhao, Yunfan Yang, and Jitao Sang. 2025 · 2025
Closest in time.
Llm collaboration with multi-agent reinforcement learning
Shuo Liu, Zeyu Liang, Xueguang Lyu, and Christopher Amato. 2025c · 2025
Closest in time.
STEVE: A Step Verification Pipeline for Computer-use Agent Training
Fanbin Lu, Zhisheng Zhong, Ziqin Wei, Shu Liu, Chi-Wing Fu, and Jiaya Jia. 2025f · 2025
Closest in time.
Deepdive: Advancing deep search agents with knowledge graphs and multi-turn rl
Rui Lu, Zhenyu Hou, Zihan Wang, Hanchen Zhang, Xiao Liu, Yujiang Li, Shi Feng, Jie Tang, and Yuxiao Dong. 2025c · 2025
Closest in time.
UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning
Zhengxi Lu, Yuxiang Chai, Yaxuan Guo, Xi Yin, Liang Liu, Hao Wang, Han Xiao, Shuai Ren, Guanjing Xiong, and Hongsheng Li. 2025a · 2025
Closest in time.
Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Xufang Luo, Yuge Zhang, Zhiyuan He, Zilong Wang, Siyun Zhao, Dongsheng Li, Luna K. Qiu, and Yuqing Yang. 2025e · 2025
Closest in time.
Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving
Xinji Mai, Haotian Xu, Xing W, Weinong Wang, Yingying Zhang, Wenqiang Zhang, Jian Hu, et al · 2025
Closest in time.
SYNTHETIC-1: Two Million Collaboratively Generated Reasoning Traces from Deepseek-R1
Justus Mattern, Sami Jaghouar, Manveer Basra, Jannik Straube, Matthew Di Ferrante, Felix Gabriel, Jack Min Ong, Vincent Weisser, and Johannes Hagemann. 2025 · 2025
Closest in time.
O 2 -Searcher: A Searching-based Agent Model for Open-domain QA via Reinforcement Learning
Jiaqi Mei et al · 2025
Closest in time.
Multi-Agent Tool-Integrated Policy Optimization
Zhanfeng Mo, Xingxuan Li, Yuntao Chen, and Lidong Bing. 2025 · 2025
Closest in time.
Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, and Tatsunori Hashimoto. 2025 · 2025
Closest in time.
Sfr-deepresearch: Towards effective reinforcement learning for autonomously reasoning single agents
Xuan-Phi Nguyen, Shrey Pandit, Revanth Gangi Reddy, Austin Xu, Silvio Savarese, Caiming Xiong, and Shafiq Joty. 2025 · 2025
Closest in time.
Introducing OpenAI o3 and o4-mini
OpenAI. 2025 · 2025
Closest in time.
LieRE: Lie Rotational Positional Encodings
Sophie Ostmeier, Brian Axelrod, Maya Varma, Michael E. Moseley, Akshay Chaudhari, and Curtis Langlotz. 2025 · 2025
Closest in time.
Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents. In Findings of the Association for Computational Linguistics: ACL 2025 , Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 6300–6323
Vardaan Pahuja, Yadong Lu, Corby Rosset, Boyu Gou, Arindam Mitra, Spencer Whitehead, Yu Su, and Ahmed Hassan Awadallah. 2025 · 2025
Closest in time.
SeCom: On Memory Construction and Retrieval for Personalized Conversational Agents. In ICLR 2025
Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang, Xufang Luo, Hao Cheng, Dongsheng Li, Yuqing Yang, Chin-Yew Lin, H. Vicky Zhao, Lili Qiu, and Jianfeng Gao. 2025 · 2025
Closest in time.
BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation
Bo Pang, Hanze Dong, Jiacheng Xu, Silvio Savarese, Yingbo Zhou, and Caiming Xiong. 2025 · 2025
Closest in time.
ToolRL: Reinforcement Learning for Tool-Use in Large Reasoning Models
Cheng Qian, Emre Can Acar, Qi He, Hongru Wang, Xiusi Chen, Dilek Hakkani-Tür, Gokhan Tur, and Heng Ji. 2025a · 2025
Closest in time.
Agentic knowledgeable self-awareness
Shuofei Qiao, Zhisong Qiu, Baochang Ren, Xiaobin Wang, Xiangyuan Ru, Ningyu Zhang, Xiang Chen, Yong Jiang, et al · 2025
Closest in time.
WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
Zile Qiao, Guoxin Chen, Xuanzhong Chen, Donglei Yu, Wenbiao Yin, Xinyu Wang, Zhen Zhang, Baixuan Li, et al · 2025
Closest in time.
UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Yujia Qin, Shiguang Liu, Qianli Ma, Jingyu Li, Xiaojun Xiao, Kai Cai, Chuang Li, et al · 2025
Closest in time.
Zep: A Temporal Knowledge Graph Architecture for Agent Memory
Preston Rasmussen, Pavlo Paliychuk, Travis Beauvais, Jack Ryan, and Daniel Chalef. 2025 · 2025
Closest in time.
MemInsight: Autonomous Memory Augmentation for LLM Agents
Rana Salama, Jason Cai, Michelle Yuan, Anna Currey, Monica Sunkara, Yi Zhang, and Yassine Benajiba. 2025 · 2025
Closest in time.
rStar2-Agent: Agentic Reasoning Technical Report
Ning Shang, Yifei Liu, Yi Zhu, Li Lyna Zhang, Weijiang Xu, Xinyu Guan, Buze Zhang, Bingcheng Dong, et al · 2025
Closest in time.
DAST: Difficulty-Adaptive Slow-Thinking for Large Reasoning Models
Yi Shen, Jian Zhang, Jieyun Huang, Shuming Shi, Wenjing Zhang, Jiangze Yan, Ning Wang, Kai Wang, Zhaoxiang Liu, and Shiguo Lian. 2025 · 2025
Closest in time.
Welcome to the Era of Experience
David Silver and Richard S. Sutton. 2025 · 2025
Closest in time.
Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
Joykirat Singh, Raghav Magazine, Yash Pandya, and Akshay Nambi. 2025 · 2025
Closest in time.
R1-searcher: Incentivizing the search capability in llms via reinforcement learning
Huatong Song, Jinhao Jiang, Yingqian Min, Jie Chen, Zhipeng Chen, Wayne Xin Zhao, Lei Fang, and Ji-Rong Wen. 2025a · 2025
Closest in time.
R1-Searcher++: Incentivizing the Dynamic Knowledge Acquisition of LLMs via Reinforcement Learning
Huatong Song, Jinhao Jiang, Wenqing Tian, Zhipeng Chen, Yuhuan Wu, Jiahao Zhao, Yingqian Min, Wayne Xin Zhao, Lei Fang, and Ji-Rong Wen. 2025b · 2025
Closest in time.
Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Yi Su, Dian Yu, Linfeng Song, Juntao Li, Haitao Mi, Zhaopeng Tu, Min Zhang, and Dong Yu. 2025 · 2025
Closest in time.
H-MEM: Hierarchical Memory for High-Efficiency Long-Term Reasoning in LLM Agents
Haoran Sun and Shaoning Zeng. 2025 · 2025
Closest in time.
OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 5555–5579
Qiushi Sun, Kanzhi Cheng, Zichen Ding, Chuanyang Jin, Yian Wang, Fangzhi Xu, Zhenyu Wu, Chengyou Jia, Liheng Chen, Zhoumianze Liu, Ben Kao, Guohao Li, Junxian He, Yu Qiao, and Zhiyong Wu. 2025a · 2025
Closest in time.
GUI-Xplore: Empowering Generalizable GUI Agents with One Exploration. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 19477–19486
Yuchen Sun, Shanhui Zhao, Tao Yu, Hao Wen, Samith Va, Mengwei Xu, Yuanchun Li, and Chongyang Zhang. 2025d · 2025
Closest in time.
MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
Liujian Tang, Shaokang Dong, Yijia Huang, Minqi Xiang, Hongtao Ruan, Bin Wang, Shuo Li, Zhiheng Xi, et al · 2025
Closest in time.
Kimi k1.5: Scaling Reinforcement Learning with LLMs
Kimi Team, Angang Du, Bofei Gao, Bowei Xing, Changjiu Jiang, Cheng Chen, Cheng Li, Chenjun Xiao, et al · 2025
Closest in time.
QwQ-32B: Embracing the Power of Reinforcement Learning
Qwen Team. 2025 · 2025
Closest in time.
Nemori: Self-Organizing Agent Memory Inspired by Cognitive Science
Unknown. 2025 · 2025
Closest in time.
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
Hanlin Wang, Chak Tou Leong, Jiashuo Wang, Jian Wang, and Wenjie Li. 2025c · 2025
Closest in time.
Acting Less is Reasoning More: Teaching Models Optimal Tool Calls via Reinforcement Learning
Han Wang, Ziwei Wang, Xinyi Yang, Keming Lu, Qi Liu, Shumin Deng, Wenxuan Zhang, and Minlie Huang. 2025h · 2025
Closest in time.
Junyang Wang, Haiyang Xu, Xi Zhang, Ming Yan, Ji Zhang, Fei Huang, and Jitao Sang. 2025i · 2025
Closest in time.
The Illusion of Role Separation: Hidden Shortcuts in LLM Role Learning (and How to Fix Them)
Zihao Wang, Yibo Jiang, Jiahao Yu, and Heqing Huang. 2025a · 2025
Closest in time.
AutoTIR: Autonomous Tools Integrated Reasoning via Reinforcement Learning
Yifan Wei, Xiaoyan Yu, Yixuan Weng, Tengfei Pan, Angsheng Li, and Li Du. 2025c · 2025
Closest in time.
MMSearch-R1: Incentivizing LMMs to Search
Jiong Wu et al · 2025
Closest in time.
ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
Xixi Wu, Kuan Li, Yida Zhao, Liwen Zhang, Litu Ou, Huifeng Yin, Zhongwang Zhang, Yong Jiang, et al · 2025
Closest in time.
On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
Yongliang Wu, Yizhou Zhou, Ziheng Zhou, Yingzhe Peng, Xinyu Ye, Xinting Hu, Wenbo Zhu, Lu Qi, Mingzhuo Yang, and Xu Yang. 2025g · 2025
Closest in time.
UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
Han Xiao, Guozhi Wang, Yuxiang Chai, Zimu Lu, Weifeng Lin, Hao He, Lue Fan, Liuyang Bian, et al · 2025
Closest in time.
Zhiyou Xiao, Qinhan Yu, Binghui Li, Geng Chen, Chong Chen, and Wentao Zhang. 2025b · 2025
Closest in time.
GUI-PRA: Process Reward Agent for GUI Tasks
Tao Xiong, Xavier Hu, Yurun Chen, Yuhang Liu, Changqiao Wu, Pengzhi Gao, Wei Liu, Jian Luan, and Shengyu Zhang. 2025 · 2025
Closest in time.
A-MEM: Agentic Memory for LLM Agents
Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang. 2025a · 2025
Closest in time.
Shuo Yan et al · 2025
Closest in time.
An Yang, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Jingren Zhou, et al · 2025
Closest in time.
MasHost Builds It All: Autonomous Multi-Agent System Directed by Reinforcement Learning
Kuo Yang, Xingjie Yang, Linhui Yu, Qing Xu, Yan Fang, Xu Wang, Zhengyang Zhou, and Yang Wang. 2025d · 2025
Closest in time.
MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
Xiao Yang, Jiawei Chen, Jun Luo, Zhengwei Fang, Yinpeng Dong, Hang Su, and Jun Zhu. 2025a · 2025
Closest in time.
Aria-UI: Visual Grounding for GUI Instructions. In Findings of the Association for Computational Linguistics: ACL 2025 , Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar (Eds.). Association for Computational Linguistics, Vienna, Austria, 22418–22433
Yuhao Yang, Yue Wang, Dongxu Li, Ziyang Luo, Bei Chen, Chao Huang, and Junnan Li. 2025c · 2025
Closest in time.
The Second Half of AI
Shunyu Yao. 2025 · 2025
Closest in time.
Mobile-Agent-v3: Foundamental Agents for GUI Automation
Jiabo Ye, Xi Zhang, Haiyang Xu, Haowei Liu, Junyang Wang, Zhaoqing Zhu, Ziwei Zheng, Feiyu Gao, et al · 2025
Closest in time.
Demystifying Long Chain-of-Thought Reasoning in LLMs
Edward Yeo, Yuxuan Tong, Morry Niu, Graham Neubig, and Xiang Yue. 2025 · 2025
Closest in time.
DAPO: An Open-Source LLM Reinforcement Learning Framework
Q. Yu et al · 2025
Closest in time.
Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training
Siyu Yuan, Zehui Chen, Zhiheng Xi, Junjie Ye, Zhengyin Du, and Jiecao Chen. 2025a · 2025
Closest in time.
Wisdom of the Crowd: Reinforcement Learning from Coevolutionary Collective Feedback
Wenzhen Yuan, Shengji Tang, Weihao Lin, Jiacheng Ruan, Ganqu Cui, Bo Zhang, Tao Chen, Ting Liu, et al · 2025
Closest in time.
Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
Xinbin Yuan, Jian Zhang, Kaixin Li, Zhuoxuan Cai, Lujian Yao, Jie Chen, Enguang Wang, Qibin Hou, et al · 2025
Closest in time.
Reinforce LLM Reasoning through Multi-Agent Reflection
Yurun Yuan and Tengyang Xie. 2025 · 2025
Closest in time.
Intrinsic Memory Agents: Heterogeneous Multi-Agent LLM Systems through Structured Contextual Memory
Sizhe Yuen, Francisco Gomez Medina, Ting Su, Yali Du, and Adam J. Sobey. 2025 · 2025
Closest in time.
Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering
Guangtao Zeng, Maohao Shen, Delin Chen, Zhenting Qi, Subhro Das, Dan Gutfreund, David Cox, Gregory Wornell, et al · 2025
Closest in time.
FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction
Zhiyuan Zeng, Jiashuo Liu, Siyuan Chen, Tianci He, Yali Liao, Jinpeng Wang, Zaiyuan Wang, Yang Yang, et al · 2025
Closest in time.
Tool-N1: Training General Tool-Use in Large Reasoning Models
Kaiyan Zhang et al · 2025
Closest in time.
Agent-as-Tool: A Study on the Hierarchical Decision Making with Reinforcement Learning
Yanfei Zhang. 2025 · 2025
Closest in time.
Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework
Zeyu Zhang, Quanyu Dai, Rui Li, Xiaohe Bo, Xu Chen, and Zhenhua Dong. 2025a · 2025
Closest in time.
Pre-training Large Memory Language Models with Internal and External Knowledge
Linxi Zhao, Sofian Zalouk, Christian K. Belardi, Justin Lovelace, Jin Peng Zhou, Ryan T. Noonan, Dongyoung Go, Kilian Q. Weinberger, Yoav Artzi, and Jennifer J. Sun. 2025c · 2025
Closest in time.
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
Qingfei Zhao, Ruobing Wang, Dingling Xu, Daren Zha, and Limin Liu. 2025b · 2025
Closest in time.
MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents
Zijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim, Alok Prakash, Daniela Rus, Jinhua Zhao, Bryan Kian Hsiang Low, and Paul Pu Liang. 2025 · 2025
Closest in time.