Fetching the paper…
Reading the bibliography…
Agents built on LLMs are increasingly deployed across diverse domains, automating complex decision-making and task execution.
FixDrive: Automatically repairing autonomous vehicle driving behaviour for $0.08 per violation
Sun, Y., Poskitt, C. M., Wang, K., and Sun, J · 1933
Earlier work this paper cites.
Principles of model checking
Baier, C., and Katoen, J · 2008
Earlier work this paper cites.
What can you verify and enforce at runtime?
Falcone, Y., Fernandez, J., and Mounier, L · 2012
Earlier work this paper cites.
The Definitive ANTLR 4 Reference
Parr, T · 2013
Earlier work this paper cites.
Detecting standard violation errors in smart contracts
Li, A., and Long, F · 2018
Earlier work this paper cites.
A survey of challenges for runtime verification from advanced application domains (beyond software)
Sánchez, C., Schneider, G., Ahrendt, W., Bartocci, E., Bianculli, D., Colombo, C., Falcone, Y., Francalanza, A., Krstic, S., Lourenço, J. M., Nickovic, D., Pace, G. J., Rufino, J., Signoles, J., Traytel, D., and Weiss, A · 2019
Earlier work this paper cites.
LawBreaker: An approach for specifying traffic laws and fuzzing autonomous vehicles
Sun, Y., Poskitt, C. M., Sun, J., Chen, Y., and Yang, Z · 2022
Earlier work this paper cites.
Runtime verification for trustworthy computing
Abela, R., Colombo, C., Curmi, A., Fenech, M., Vella, M., and Ferrando, A · 2023
Earlier work this paper cites.
CAMEL: communicative agents for "mind" exploration of large language model society
Li, G., Hammoud, H., Itani, H., Khizbullin, D., and Ghanem, B · 2023
Earlier work this paper cites.
A language agent for autonomous driving
Mao, J., Ye, J., Qian, Y., Pavone, M., and Wang, Y · 2023
Earlier work this paper cites.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S · 2023
Earlier work this paper cites.
Pedro, R., Castro, D., Carreira, P., and Santos, N · 2023
Earlier work this paper cites.
Reflexion: language agents with verbal reinforcement learning
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S · 2023
Earlier work this paper cites.
ReAct: Synergizing reasoning and acting in language models
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K. R., and Cao, Y · 2023
Earlier work this paper cites.
LLMArena: Assessing capabilities of large language models in dynamic multi-agent environments
Chen, J., Hu, X., Liu, S., Huang, S., Tu, W., He, Z., and Wen, L · 2024
Earlier work this paper cites.
AgentPoison: Red-teaming LLM agents via poisoning memory or knowledge bases
Chen, Z., Xiang, Z., Xiao, C., Song, D., and Li, B · 2024
Earlier work this paper cites.
A survey on in-context learning
Dong, Q., Li, L., Dai, D., Zheng, C., Ma, J., Li, R., Xia, H., Xu, J., Wu, Z., Chang, B., Sun, X., Li, L., and Sui, Z · 2024
Earlier work this paper cites.
Safeguarding large language models: A survey
Dong, Y., Mu, R., Zhang, Y., Sun, S., Zhang, T., Wu, C., Jin, G., Qi, Y., Hu, J., Meng, J., Bensalem, S., and Huang, X · 2024
Earlier work this paper cites.
RedCode: Risky code execution and generation benchmark for code agents
Guo, C., Liu, X., Xie, C., Zhou, A., Zeng, Y., Lin, Z., Song, D., and Li, B · 2024
Earlier work this paper cites.
Large language model based multi-agents: A survey of progress and challenges
Guo, T., Chen, X., Wang, Y., Chang, R., Pei, S., Chawla, N. V., Wiest, O., and Zhang, X · 2024
Cited alongside, same era.
LLM multi-agent systems: Challenges and open problems
Han, S., Zhang, Q., Yao, Y., Jin, W., Xu, Z., and He, C · 2024
Cited alongside, same era.
Efficient detection of toxic prompts in large language models
Liu, Y., Yu, J., Sun, H., Shi, L., Deng, G., Chen, Y., and Liu, Y · 2024
Cited alongside, same era.
Identifying the risks of LM agents with an LM-emulated sandbox
Ruan, Y., Dong, H., Wang, A., Pitis, S., Zhou, Y., Ba, J., Dubois, Y., Maddison, C. J., and Hashimoto, T · 2024
Cited alongside, same era.
EHRAgent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records
Shi, W., Xu, R., Zhuang, Y., Yu, Y., Zhang, J., Wu, H., Zhu, Y., Ho, J. C., Yang, C., and Wang, M. D · 2024
Cited alongside, same era.
GPT-4V(ision) is a generalist web agent, if grounded
Zheng, B., Gou, B., Kil, J., Sun, H., and Su, Y · 2024
Later among the works it cites.
ALI-Agent: Assessing LLMs’ alignment with human values via agent-based evaluation
Zheng, J., Wang, H., Zhang, A., Nguyen, T. D., Sun, J., and Chua, T · 2024
Later among the works it cites.
https://github.com/haoyuwang99/AgentSpec , 2025
AgentSpec · 2025
Closest in time.
Apollo Self-Driving
Baidu Apollo · 2025
Closest in time.
When AI thinks it will lose, it sometimes cheats, study finds
Booth, H · 2025
Closest in time.
AI agents under threat: A survey of key security challenges and future pathways
Deng, Z., Guo, Y., Han, C., Ma, W., Xiong, J., Wen, S., and Xiang, Y · 2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Prioritizing safeguarding over autonomy: Risks of LLM agents for science
Tang, X., Jin, Q., Zhu, K., Yuan, T., Zhang, Y., Zhou, W., Qu, M., Zhao, Y., Tang, J., Zhang, Z., Cohan, A., Lu, Z., and Gerstein, M · 2024
Cited alongside, same era.
Voyager: An open-ended embodied agent with large language models
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A · 2024
Cited alongside, same era.
μ \mu Drive: User-controlled autonomous driving
Wang, K., Poskitt, C. M., Sun, Y., Sun, J., Wang, J., Cheng, P., and Chen, J · 2024
Cited alongside, same era.
A survey on large language model based autonomous agents
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., Zhao, W. X., Wei, Z., and Wen, J · 2024
Cited alongside, same era.
Executable code actions elicit better LLM agents
Wang, X., Chen, Y., Yuan, L., Zhang, Y., Li, Y., Peng, H., and Ji, H · 2024
Cited alongside, same era.
GuardAgent: Safeguard LLM agents by a guard agent via knowledge-enabled reasoning
Xiang, Z., Zheng, L., Li, Y., Hong, J., Li, Q., Xie, H., Zhang, J., Xiong, Z., Xie, C., Yang, C., Song, D., and Li, B · 2024
Cited alongside, same era.
Understanding the weakness of large language model agents within a complex android environment
Xing, M., Zhang, R., Xue, H., Chen, Q., Yang, F., and Xiao, Z · 2024
Cited alongside, same era.
llama.cpp: LLM inference in C/C++
Gerganov, G., and ggml-org Community · 2025
Closest in time.
LangChain
LangChain Contributors · 2025
Closest in time.
LangChain Expression Language (LCEL)
LangChain Contributors · 2025
Closest in time.
Eia: Environmental injection attack on generalist web agents for privacy leakage
Liao, Z., Mo, L., Xu, C., Kang, M., Zhang, J., Xiao, C., Tian, Y., Li, B., and Sun, H · 2025
Closest in time.
What are AI guardrails?
McKinsey & Company · 2025
Closest in time.
Real estate listing gaffe exposes widespread use of AI in Australian industry – and potential risks
McLeod, C · 2025
Closest in time.
AutoGen: A framework for building AI agents and applications
Microsoft · 2025
Closest in time.
CROW: eliminating backdoors from large language models via internal consistency regularization
Min, N. M., Pham, L. H., Li, Y., and Sun, J · 2025
Closest in time.
NeMo: A scalable generative AI framework
NVIDIA · 2025
Closest in time.
The rise and potential of large language model based agents: A survey
Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., Wang, J., Jin, S., Zhou, E., Zheng, R., Fan, X., Wang, X., Xiong, L., Zhou, Y., Wang, W., Jiang, C., Zou, Y., Liu, X., Yin, Z., Dou, S., Weng, R., Qin, W., Zheng, Y., Qiu, X., Huang, X., Zhang, Q., and Gui, T · 2025
Closest in time.
LLMScan: Causal scan for LLM misbehavior detection
Zhang, M., Goh, K. K., Zhang, P., and Sun, J · 2025
Closest in time.
Position: Trustworthy AI agents require the integration of large language models and formal methods
Zhang, Y., Cai, Y., Zuo, X., Luan, X., Wang, K., Hou, Z., Zhang, Y., Wei, Z., Sun, M., Sun, J., Sun, J., and Dong, J. S · 2025
Closest in time.