Fetching the paper…
Reading the bibliography…
Large language model (LLM)-based agents combine LLMs with external tools to automate tasks such as scheduling meetings, managing documents, or booking travel.
S. Gu and L. Rigazio, “Towards Deep Neural Network Architectures Robust to Adversarial Examples,” in International Conference on Learning Representations (ICLR) , 2015
2015
Earlier work this paper cites.
D. Meng and H. Chen, “MagNet: A Two-Pronged Defense against Adversarial Examples,” in Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, CCS 2017, Dallas, TX, USA, October 30 - November 03, 2017 . ACM, 2017, pp. 135–147
2017
Earlier work this paper cites.
J. H. Metzen, T. Genewein, V. Fischer, and B. Bischoff, “On Detecting Adversarial Perturbations,” in International Conference on Learning Representations (ICLR) , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. N. Bhagoji, D. Cullina, C. Sitawarin, and P. Mittal, “Enhancing robustness of machine learning systems via data transformations,” in 52nd Annual Conference on Information Sciences and Systems, CISS 2018, Princeton, NJ, USA, March 21-23, 2018 . IEEE, 2018, pp. 1–5
2018
Earlier work this paper cites.
P. Samangouei, M. Kabkab, and R. Chellappa, “Defense-GAN: Protecting Classifiers Against Adversarial Attacks Using Generative Models,” in International Conference on Learning Representations (ICLR) , 2018
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
C. Zhu, Y. Cheng, Z. Gan, S. Sun, and T. G. et al., “FreeLB: Enhanced Adversarial Training for Natural Language Understanding,” in International Conference on Learning Representations (ICLR) , 2020
2020
Earlier work this paper cites.
N. Carlini, F. Tramèr, E. Wallace, M. Jagielski, and A. H. et al., “Extracting Training Data from Large Language Models,” in USENIX Security Symposium , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
2022
Earlier work this paper cites.
H. Chase. (2022) LangChain. Langchain. Original-date: 2022-10-17T02:58:36Z. [Online]. Available: https://github.com/langchain-ai/langchain
2022
Earlier work this paper cites.
W. Nie, B. Guo, Y. Huang, C. Xiao, and A. V. et al., “Diffusion Models for Adversarial Purification,” in International Conference on Machine Learning (ICML) , ser. Proceedings of Machine Learning Research, vol. 162. PMLR, 2022, pp. 16 805–16 827
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
E. Crothers, N. Japkowicz, and H. L. Viktor, “Machine-generated text: A comprehensive survey of threat models and detection methods,” IEEE Access , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection,” in Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, AISec 2023, Copenhagen, Denmark, 30 November 2023 . ACM, 2023, pp. 79–90. [Online]. Available: https://doi.org/10.1145/3605764.3623985
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
M. Shergadwala. (2023) Prompt injection attacks in various LLMs. [Online]. Available: https://medium.com/@murtuza.shergadwala/prompt-injection-attacks-in-various-llms-206f56cd6ee9
2023
Cited alongside, same era.
A. Wei, N. Haghtalab, and J. Steinhardt, “Jailbroken: How Does LLM Safety Training Fail?” in Annual Conference on Neural Information Processing Systems (NeurIPS) , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
Notion. (2024) Notion AI | now with q&a. [Online]. Available: https://www.notion.so/product/ai
2024
Closest in time.
OpenAI. (2024) Integrate the OpenAI (ChatGPT) API with the gmail API. OpenAI. [Online]. Available: https://pipedream.com/apps/openai/integrations/gmail
2024
Closest in time.
Clockwise. (2024) AI calendar | AI scheduling assistant | clockwise. Clockwise. [Online]. Available: https://www.getclockwise.com/ai
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Lee, “0xk1h0 - github: Jailbreak prompts collection,” 2023. [Online]. Available: https://github.com/0xk1h0/ChatGPT_DAN
2023
Cited alongside, same era.
W. Zhang. (2023) Prompt injection attack on GPT-4 — robust intelligence. [Online]. Available: https://www.robustintelligence.com/blog-posts/prompt-injection-attack-on-gpt-4
2023
Cited alongside, same era.
LaurieWired [@lauriewired]. (2023) Novel jailbreak technique via typoglycemia. [Online]. Available: https://twitter.com/lauriewired/status/1682825249203662848
2023
Cited alongside, same era.
2023
Cited alongside, same era.
S. Armstrong and R. Gorman, “Using GPT-eliezer against ChatGPT jailbreaking,” 2023. [Online]. Available: https://www.alignmentforum.org/posts/pNcFYZnPdXyL2RfgA/using-gpt-eliezer-against-chatgpt-jailbreaking
2023
Cited alongside, same era.
2023
Cited alongside, same era.
LearnPrompting. (2023) Learn prompting. [Online]. Available: https://learnprompting.org/docs/category/-defensive-measures
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
Microsoft. (2024) Personal AI assistant | microsoft copilot. [Online]. Available: https://www.microsoft.com/en-us/microsoft-copilot/personal-ai-assistant
2024
Closest in time.
J. Rehberger. (2024) Microsoft copilot: From prompt injection to exfiltration of personal information · embrace the red. [Online]. Available: https://embracethered.com/blog/posts/2024/m365-copilot-prompt-injection-tool-invocation-and-data-exfil-using-ascii-smuggling/
2024
Closest in time.
E. Debenedetti, J. Rando, D. Paleka, S. F. Florin, D. Albastroiu, Y. L. Niv Cohen, A. S. Reshmi Ghosh, Rui Wen, S. Z.-B. Giovanni Cherubin, R. Schmid, V. Klem, S. K. Takahiro Miki, Chenhao Li, M. Fritz, F. Tramèr, S. Abdelnabi, and L. Schönherr, “Dataset and lessons learned from the 2024 satml llm capture-the-flag competition,” in NeurIPS Dataset and Benchmark Track , 2024
2024
Closest in time.
2024
Closest in time.
Mitre. (2024) LLM jailbreak | MITRE ATLAS™. [Online]. Available: https://atlas.mitre.org/techniques/AML.T0054
2024
Closest in time.
2024
Closest in time.
J. Rehberger. (2024) Google AI studio: LLM-powered data exfiltration hits again! quickly fixed. · embrace the red. [Online]. Available: https://embracethered.com/blog/posts/2024/google-ai-studio-data-exfiltration-now-fixed/
2024
Closest in time.
(2024) Llama 3.2: Revolutionizing edge AI and vision with open, customizable models. [Online]. Available: https://ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Q. Team. (2024) Qwen2.5: A party of foundation models! Section: blog. [Online]. Available: http://qwenlm.github.io/blog/qwen2.5/
2024
Closest in time.
2024
Closest in time.
(2024) Llama 3.1 | Model Cards and Prompt formats. [Online]. Available: https://www.llama.com/docs/model-cards-and-prompt-formats/llama3_1/#-tool-calling-(8b/70b/405b)-
2024
Closest in time.
OWASP. (2024) OWASP top 10 for LLM applications. OWASP. [Online]. Available: https://www.llmtop10.com
2024
Closest in time.