Fetching the paper…
Reading the bibliography…
AI agents aim to solve complex tasks by combining text-based reasoning with external tool calls.
“Language models are few-shot learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry and Amanda Askell · 1901
Earlier work this paper cites.
“Intelligent agents: Theory and practice”
Michael Wooldridge and Nicholas Jennings · 1995
Earlier work this paper cites.
“A critique of the deepsec platform for security analysis of deep learning models”
Nicholas Carlini · 2019
Earlier work this paper cites.
“Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks”
Francesco Croce and Matthias Hein · 2020
Earlier work this paper cites.
“On Adaptive Attacks to Adversarial Example Defenses”
Florian Tramèr, Nicholas Carlini, Wieland Brendel and Aleksander Madry · 2020
Earlier work this paper cites.
“RobustBench: a standardized adversarial robustness benchmark”
Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal and Matthias Hein · 2021
Earlier work this paper cites.
“WebGPT: Browser-assisted question-answering with human feedback”
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju and William Saunders · 2021
Earlier work this paper cites.
“Training a helpful and harmless assistant with reinforcement learning from human feedback”
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli and Tom Henighan · 2022
Earlier work this paper cites.
“Exploiting GPT-3 prompts with malicious inputs that order the model to ignore its previous directions”, https://x.com/goodside/status/1569128808308957185, 2022
Riley Goodside · 2022
Earlier work this paper cites.
“Language models as zero-shot planners: Extracting actionable knowledge for embodied agents”
Wenlong Huang, Pieter Abbeel, Deepak Pathak and Igor Mordatch · 2022
Earlier work this paper cites.
“Large language models are zero-shot reasoners”
Takeshi Kojima, Shixiang Gu, Machel Reid, Yutaka Matsuo and Yusuke Iwasawa · 2022
Earlier work this paper cites.
“Training language models to follow instructions with human feedback”
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama and Alex Ray · 2022
Earlier work this paper cites.
“Ignore previous prompt: Attack techniques for language models”
Fábio Perez and Ian Ribeiro · 2022
Earlier work this paper cites.
“Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI”
Mahima Pushkarna, Andrew Zaldivar and Oddur Kjartansson · 2022
Earlier work this paper cites.
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay and Jost Springenberg · 2022
Earlier work this paper cites.
“LaMDA: Language models for dialog applications”
Romal Thoppilan, Daniel De, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker and Yu Du · 2022
Earlier work this paper cites.
“Chain-of-thought prompting elicits reasoning in large language models”
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc Le and Denny Zhou · 2022
Earlier work this paper cites.
“Prompt injection attacks against GPT-3”, https://simonwillison.net/2022/Sep/12/prompt-injection/ , 2022
Simon Willison · 2022
Earlier work this paper cites.
“You can’t solve AI security problems with more AI”, https://simonwillison.net/2022/Sep/17/prompt-injection-more-ai/ , 2022
Simon Willison · 2022
Earlier work this paper cites.
“WebShop: Towards scalable real-world web interaction with grounded language agents”
Shunyu Yao, Howard Chen, John Yang and Karthik Narasimhan · 2022
Earlier work this paper cites.
“ReAct: Synergizing reasoning and acting in language models”
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan and Yuan Cao · 2022
Earlier work this paper cites.
“Introducing Command R+: Our new, most powerful model in the Command R family”, https://cohere.com/command, 2023
Cohere · 2023
Earlier work this paper cites.
“Misusing Tools in Large Language Models With Visual Adversarial Examples”
Xiaohan Fu, Zihan Wang, Shuheng Li, Rajesh Gupta, Niloofar Mireshghallah, Taylor Berg-Kirkpatrick and Earlence Fernandes · 2023
Earlier work this paper cites.
“PAL: Program-aided language models”
Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan and Graham Neubig · 2023
Earlier work this paper cites.
“Gemini: a family of highly capable multimodal models”
Gemini Team · 2023
Earlier work this paper cites.
“Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection”
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz and Mario Fritz · 2023
Cited alongside, same era.
“Function calling”, https://cookbook.openai.com/examples/how_to_call_functions_with_chat_models , 2023
Colin Jarvis and Joe Palermo · 2023
Cited alongside, same era.
“Intro to Large Language Models”, https://www.youtube.com/watch?v=zjkBMFhNj_g , 2023
Andrej Karpathy · 2023
Cited alongside, same era.
“Language models can solve computer tasks”
Geunwoo Kim, Pierre Baldi and Stephen McAleer · 2023
Cited alongside, same era.
“Evaluating Language-Model Agents on Realistic Autonomous Tasks”
Megan Kinniment, Lucas Sato, Haoxing Du, Brian Goodrich, Max Hasin, Lawrence Chan, Luke Miles, Tao. Lin, Hjalmar Wijk, Joel Burget, Aaron Ho, Elizabeth Barnes and Paul Christiano · 2023
Cited alongside, same era.
“The Claude 3 Model Family: Opus, Sonnet, Haiku”, https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bbc618857627/Model_Card_Claude_3.pdf , 2024
Anthropic · 2024
Closest in time.
“Tool use (function calling)”, https://docs.anthropic.com/en/docs/tool-use , 2024
Anthropic · 2024
Closest in time.
“JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models”, 2024
Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George. Pappas, Florian Tramèr, Hamed Hassani and Eric Wong · 2024
Closest in time.
“StruQ: Defending Against Prompt Injection with Structured Queries”
Sizhe Chen, Julien Piet, Chawin Sitawarin and David Wagner · 2024
Closest in time.
“Dataset and Lessons Learned from the 2024 SaTML LLM Capture-the-Flag Competition”, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“AgentSims: An Open-Source Sandbox for Large Language Model Evaluation”, 2023
Jiaju Lin, Haoran Zhao, Aochi Zhang, Yiting Wu, Huqiuyue Ping and Qin Chen · 2023
Cited alongside, same era.
“AgentBench: Evaluating LLMs as Agents”, 2023
Xiao Liu et al · 2023
Cited alongside, same era.
“Prompt Injection attack against LLM-integrated Applications”
Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng and Yang Liu · 2023
Cited alongside, same era.
“Formalizing and Benchmarking Prompt Injection Attacks and Defenses”, 2023
Yupei Liu, Yuqi Jia, Runpeng Geng, Jinyuan Jia and Neil Gong · 2023
Cited alongside, same era.
“Inverse Scaling Prize: Second Round Winners”, 2023
Ian McKenzie, Alexander Lyzhov, Alicia Parrish, Ameya Prabhu, Aaron Mueller, Najoung Kim, Sam Bowman and Ethan Perez · 2023
Cited alongside, same era.
“Inverse Scaling: When Bigger Isn’t Better”
Ian McKenzie, Alexander Lyzhov, Michael Pieler, Alicia Parrish, Aaron Mueller, Ameya Prabhu, Euan McLean, Aaron Kirtland, Alexis Ross and Alisa Liu · 2023
Cited alongside, same era.
“Gorilla: Large Language Model Connected with Massive APIs”, 2023
Shishir. Patil, Tianjun Zhang, Xin Wang and Joseph. Gonzalez · 2023
Cited alongside, same era.
Edoardo Debenedetti et al · 2024
Closest in time.
“Coercing LLMs to do and reveal (almost) anything”
Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah, Yuxin Wen and Tom Goldstein · 2024
Closest in time.
“Defending Against Indirect Prompt Injection Attacks With Spotlighting”, 2024
Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger and Emre Kiciman · 2024
Closest in time.
“Llama-3 Function Calling Demo”, https://nbsanity.com/static/d06085f1dacae8c9de9402f2d7428de2/demo.html , 2024
Hamel Husain · 2024
Closest in time.
“Exploiting programmatic behavior of llms: Dual-use through standard security attacks”
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia and Tatsunori Hashimoto · 2024
Closest in time.
“ChainGuard”, https://lakeraai.github.io/chainguard/ , 2024
Lakera · 2024
Closest in time.
“Hugging Face prompt injection identification”, https://python.langchain.com/v0.1/docs/guides/productionization/safety/hugging_face_prompt_injection/ , 2024
LangChain · 2024
Closest in time.
“Sandwich Defense”, https://learnprompting.org/docs/prompt_hacking/defensive_measures/sandwich_defense , 2024
Learn Prompting · 2024
Closest in time.
“Chameleon: Plug-and-play compositional reasoning with large language models”
Pan Lu, Baolin Peng, Hao Cheng, Michel Galley, Kai-Wei Chang, Ying Wu, Song-Chun Zhu and Jianfeng Gao · 2024
Closest in time.
“HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal”, 2024
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth and Dan Hendrycks · 2024
Closest in time.
“Can LLMs Follow Simple Rules?”, 2024
Norman Mu, Sarah Chen, Zifan Wang, Sizhe Chen, David Karamardian, Lulwa Aljeraisy, Basel Alomair, Dan Hendrycks and David Wagner · 2024
Closest in time.
“Neural Exec: Learning (and Learning from) Execution Triggers for Prompt Injection Attacks”, 2024
Dario Pasquini, Martin Strohmeier and Carmela Troncoso · 2024
Closest in time.
“Fine-Tuned DeBERTa-v3-base for Prompt Injection Detection”
ProtectAI · 2024
Closest in time.
“Identifying the Risks of LM Agents with an LM-Emulated Sandbox”
Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris. Maddison and Tatsunori Hashimoto · 2024
Closest in time.
“HuggingGPT: Solving AI tasks with ChatGPT and its friends in Hugging Face”
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu and Yueting Zhuang · 2024
Closest in time.
“The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions”, 2024
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke and Alex Beutel · 2024
Closest in time.
“SecGPT: An execution isolation architecture for LLM-based systems”
Yuhao Wu, Franziska Roesner, Tadayoshi Kohno, Ning Zhang and Umar Iqbal · 2024
Closest in time.
“Berkeley Function Calling Leaderboard”, https://gorilla.cs.berkeley.edu/blogs/8_berkeley_function_calling_leaderboard.html , 2024
Fanjia Yan, Huanzhi Mao, Charlie-Jie Ji, Tianjun Zhang, Shishir. Patil, Ion Stoica and Joseph. Gonzalez · 2024
Closest in time.
Qiusi Zhan, Zhixiang Liang, Zifan Ying and Daniel Kang · 2024
Closest in time.
“Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?”
Egor Zverev, Sahar Abdelnabi, Mario Fritz and Christoph Lampert · 2024
Closest in time.