Fetching the paper…
Reading the bibliography…
The strong planning and reasoning capabilities of Large Language Models (LLMs) have fostered the development of agent-based systems capable of leveraging external tools and interacting with increasingly complex environments.
An empirical study of the reliability of unix utilities
Miller, B. P., Fredriksen, L., and So, B · 1990
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P · 2002
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A · 2018
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. D. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
CodeRL: Mastering code generation through pretrained models and deep reinforcement learning
Le, H., Wang, Y., Gotmare, A. D., Savarese, S., and Hoi, S · 2022
Earlier work this paper cites.
Competition-level code generation with alphacode
Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Dal Lago, A., et al · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Perez, F. and Ribeiro, I · 2022
Earlier work this paper cites.
Prompt injection attacks against GPT-3
Willison, S · 2022
Earlier work this paper cites.
https://learnprompting.org/docs/prompt_hacking/defensive_measures/sandwich_defense , 2023
Sandwitch defense · 2023
Earlier work this paper cites.
Claude family, 2023
Anthropic · 2023
Earlier work this paper cites.
Pal: Program-aided language models
Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G · 2023
Earlier work this paper cites.
Gemini family, 2023
Google · 2023
Earlier work this paper cites.
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M · 2023
Earlier work this paper cites.
A real-world webagent with planning, long context understanding, and program synthesis
Gur, I., Furuta, H., Huang, A., Safdari, M., Matsuo, Y., Eck, D., and Faust, A · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., et al · 2023
Earlier work this paper cites.
Baseline defenses for adversarial attacks against aligned language models
Jain, N., Schwarzschild, A., Wen, Y., Somepalli, G., Kirchenbauer, J., Chiang, P.-y., Goldblum, M., Saha, A., Geiping, J., and Goldstein, T · 2023
Cited alongside, same era.
Prompt injection attack against llm-integrated applications
Liu, Y., Deng, G., Li, Y., Wang, K., Wang, Z., Wang, X., Zhang, T., Liu, Y., Wang, H., Zheng, Y., et al · 2023
Cited alongside, same era.
Ultimate ChatGPT prompt engineering guide for general users and developers
Mendes, A · 2023
Cited alongside, same era.
Chatgpt plugins, 2023b
OpenAI · 2023
Cited alongside, same era.
Gorilla: Large language model connected with massive apis
Patil, S. G., Zhang, T., Wang, X., and Gonzalez, J. E · 2023
Cited alongside, same era.
Openai o1, 2024
OpenAI · 2024
Later among the works it cites.
Goex: Perspectives and designs towards a runtime for autonomous llm applications
Patil, S. G., Zhang, T., Fang, V., Huang, R., Hao, A., Casado, M., Gonzalez, J. E., Popa, R. A., Stoica, I., et al · 2024
Later among the works it cites.
Fine-tuned deberta-v3-base for prompt injection detection, 2024
ProtectAI · 2024
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2024
Later among the works it cites.
The instruction hierarchy: Training llms to prioritize privileged instructions
Wallace, E., Xiao, K., Leike, R., Weng, L., Heidecke, J., and Beutel, A · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Toolllm: Facilitating large language models to master 16000+ real-world apis
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., et al · 2023
Cited alongside, same era.
Ignore this title and HackAPrompt: Exposing systemic vulnerabilities of LLMs through a global prompt hacking competition
Schulhoff, S., Pinto, J., Khan, A., Bouchard, L.-F., Si, C., Anati, S., Tagliabue, V., Kost, A., Carnahan, C., and Boyd-Graber, J · 2023
Cited alongside, same era.
The dual llm pattern for building ai assistants that can resist prompt injection
Willison, S · 2023
Cited alongside, same era.
Delimiters won’t save you from prompt injection
Willison, S · 2023
Cited alongside, same era.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Yu, J., Lin, X., and Xing, X · 2023
Cited alongside, same era.
Agentdojo: A dynamic environment to evaluate attacks and defenses for llm agents
Debenedetti, E., Zhang, J., Balunović, M., Beurer-Kellner, L., Fischer, M., and Tramèr, F · 2024
Cited alongside, same era.
Mind2web: Towards a generalist agent for the web
Deng, X., Gu, Y., Zheng, B., Chen, S., Stevens, S., Wang, B., Sun, H., and Su, Y · 2024
Cited alongside, same era.
Dissecting adversarial robustness of multimodal lm agents
Wu, C. H., Shah, R. R., Koh, J. Y., Salakhutdinov, R., Fried, D., and Raghunathan, A · 2024
Later among the works it cites.
Advweb: Controllable black-box attacks on vlm-powered web agents
Xu, C., Kang, M., Zhang, J., Liao, Z., Mo, L., Yuan, M., Sun, H., and Li, B · 2024
Later among the works it cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang, K., Lu, K., Bao, K., Yang, K., Yu, L., Li, M., Xue, M., Zhang, P., Zhu, Q., Men, R., Lin, R., Li, T., Xia, T., Ren, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Wan, Y., Liu, Y., Cui, Z., Zhang, Z., and Qiu, Z · 2024
Later among the works it cites.
Promptfuzz: Harnessing fuzzing techniques for robust testing of prompt injection in llms
Yu, J., Shao, Y., Miao, H., Shi, J., and Xing, X · 2024
Later among the works it cites.
Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents
Zhan, Q., Liang, Z., Ying, Z., and Kang, D · 2024
Later among the works it cites.
Attacking vision-language computer agents via pop-ups, 2024
Zhang, Y., Yu, T., and Yang, D · 2024
Later among the works it cites.
Webarena: A realistic web environment for building autonomous agents
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al · 2025
Closest in time.
Operator – an agent that can use its own browser to perform tasks for you., 2025
OpenAI · 2025
Closest in time.
Qwq-32b: Embracing the power of reinforcement learning, March 2025
Qwen Team · 2025
Closest in time.
IsolateGPT: An Execution Isolation Architecture for LLM-Based Systems
Wu, Y., Roesner, F., Kohno, T., Zhang, N., and Iqbal, U · 2025
Closest in time.
Ai domination: Remote controlling chatgpt zombai instances, January 2025
wunderwuzzi · 2025
Closest in time.