Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are increasingly integrated into autonomous systems, giving rise to a new class of software known as Agentware, where LLM-powered agents perform complex, open-ended tasks in domains such as software engineering, customer service, and data analysis.
“Dapper, a Large-Scale Distributed Systems Tracing Infrastructure”, 2010
Benjamin. Sigelman et al · 2010
Earlier work this paper cites.
“A Survey of Data-Driven and Knowledge-Aware eXplainable AI”
Xiao-Hui Li et al · 2020
Earlier work this paper cites.
“Measuring Massive Multitask Language Understanding”
Dan Hendrycks et al · 2021
Earlier work this paper cites.
“Efficient Training of Language Models to Fill in the Middle”, 2022
Mohammad Bavarian et al · 2022
Earlier work this paper cites.
“Towards training reproducible deep learning models”
Boyuan Chen et al · 2022
Earlier work this paper cites.
“Enjoy your observability: an industrial survey of microservice tracing and analysis”
B. Li et al · 2022
Earlier work this paper cites.
“Chain-of-thought prompting elicits reasoning in large language models”
Jason Wei et al · 2022
Earlier work this paper cites.
“InCoder: A Generative Model for Code Infilling and Synthesis”
Daniel Fried et al · 2023
Earlier work this paper cites.
Lei Huang et al · 2023
Earlier work this paper cites.
“AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap”, 2023
Q. Liao and Jennifer Vaughan · 2023
Earlier work this paper cites.
“Reflexion: language agents with verbal reinforcement learning”
Noah Shinn et al · 2023
Earlier work this paper cites.
“Self-Consistency Improves Chain of Thought Reasoning in Language Models”, 2023
Xuezhi Wang et al · 2023
Earlier work this paper cites.
“Phoenix by Arize” last accessed: 2024-10-08, 2024
Arize · 2024
Earlier work this paper cites.
“SWE-bench Lite: A Canonical Subset for Efficient Evaluation of Language Models as Software Engineers” last accessed: 2024-10-08, 2024
Jiayi Carlos E. John · 2024
Earlier work this paper cites.
“A Survey on Evaluation of Large Language Models”
Yupeng Chang et al · 2024
Earlier work this paper cites.
“django Web Framework” last accessed: 2024-10-08, 2024
Django · 2024
Earlier work this paper cites.
“AgentOps: Enabling Observability of LLM Agents”
Liming Dong, Qinghua Lu and Liming Zhu · 2024
Earlier work this paper cites.
“PromptExp: Multi-granularity Prompt Explanation of Large Language Models” under review, preprint available on arXiv, 2024
Ximing Dong et al · 2024
Cited alongside, same era.
“Dynatrace” last accessed: 2024-10-08, 2024
Dynatrace · 2024
Cited alongside, same era.
“GitHub Copilot” last accessed: 2024-10-09, 2024
GitHub · 2024
Cited alongside, same era.
“DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning”
Siyuan Guo et al · 2024
Cited alongside, same era.
URL: https://www.haptik.ai/
“Haptik, Drive Business Efficiency at Scale with Generative AI” last accessed: 2024-10-08, 2024 · 2024
Cited alongside, same era.
“Rethinking Software Engineering in the Era of Foundation Models: A Curated Catalogue of Challenges in the Development of Trustworthy FMware”
Tula Masterman, Sandi Besen, Mason Sawtell and Alex Chao · 2024
Closest in time.
“Nebuly - Explicit and Implicit LLM User Feedback Quickguide” last accessed: 2024-10-08, 2024
Nebuly · 2024
Closest in time.
“Introducing OpenAI o1” last accessed: 2024-10-08, 2024
OpenAI · 2024
Closest in time.
“Qwak - LLMops” last accessed: 2024-10-08, 2024
Qwak · 2024
Closest in time.
“Traceloop” last accessed: 2024-10-08, 2024
Traceloop · 2024
Closest in time.
“Weights and Biases - Weave” last accessed: 2024-10-08, 2024
Weights and Biases · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ahmed. Hassan et al · 2024
Cited alongside, same era.
“Towards AI-Native Software Engineering (SE 3.0): A Vision and a Challenge Roadmap”, 2024
Ahmed. Hassan et al · 2024
Cited alongside, same era.
“Helicone” last accessed: 2024-10-08, 2024
Helicone · 2024
Cited alongside, same era.
“Automated Design of Agentic Systems”, 2024
Shengran Hu, Cong Lu and Jeff Clune · 2024
Cited alongside, same era.
“Humanloop” last accessed: 2024-10-08, 2024
Humanloop · 2024
Cited alongside, same era.
URL: https://www.cognition.ai/blog/introducing-devin
“Introducing Devin, the first AI software engineer” last accessed: 2024-10-08, 2024 · 2024
Cited alongside, same era.
“Language Models for Code Completion: A Practical Evaluation”
M. Izadi et al · 2024
Cited alongside, same era.
“WhyLabs” last accessed: 2024-10-08, 2024
WhyLabs · 2024
Closest in time.
Fangzhi Xu et al · 2024
Closest in time.
“Tree of thoughts: deliberate problem solving with large language models”
Shunyu Yao et al · 2024
Closest in time.
“AutoCodeRover: Autonomous Program Improvement”
Yuntong Zhang, Haifeng Ruan, Zhiyu Fan and Abhik Roychoudhury · 2024
Closest in time.
“Reasoning models don’t always say what they think” Available at https://www.anthropic.com/research/reasoning-models-dont-say-think , Anthropic Technical Report, 2025
Anthropic Safety Team · 2025
Closest in time.
“Chain-of-Thought Is Not Explainability”, 2025
Fazl Barez et al · 2025
Closest in time.
“Reasoning Models Don’t Always Say What They Think”, 2025
Yanda Chen et al · 2025
Closest in time.
“DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning”
DeepSeek-AI et al · 2025
Closest in time.
“GDB: The GNU Project Debugger” Accessed: 2025-05-26, https://sourceware.org/gdb/
2025
Closest in time.
“Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces”, 2025
DiJia Su et al · 2025
Closest in time.
“OpenHands: An Open Platform for AI Software Developers as Generalist Agents”, 2025
Xingyao Wang et al · 2025
Closest in time.