Fetching the paper…
Reading the bibliography…
Autonomous AI agents that can follow instructions and perform complex multi-step tasks have tremendous potential to boost human productivity.
Differentially private empirical risk minimization
K. Chaudhuri, C. Monteleoni, and A. D. Sarwate · 2011
Earlier work this paper cites.
Deep learning with differential privacy
M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang · 2016
Earlier work this paper cites.
Membership inference attacks against machine learning models
R. Shokri, M. Stronati, C. Song, and V. Shmatikov · 2017
Earlier work this paper cites.
Extracting Training Data from Large Language Models
N. Carlini, F. Tramèr, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. B. Brown, D. X. Song, Ú. Erlingsson, A. Oprea, and C. Raffel · 2020
Earlier work this paper cites.
What does it mean for a language model to preserve privacy?
H. Brown, K. Lee, F. Mireshghallah, R. Shokri, and F. Tramèr · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. H. Chi, Q. Le, and D. Zhou · 2022
Earlier work this paper cites.
Mind2web: Towards a generalist agent for the web, 2023
X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica · 2023
Earlier work this paper cites.
BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Earlier work this paper cites.
N. Mireshghallah, H. Kim, X. Zhou, Y. Tsvetkov, M. Sap, R. Shokri, and Y. Choi · 2023
Earlier work this paper cites.
Gorilla: Large language model connected with massive apis
S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez · 2023
Earlier work this paper cites.
Beyond memorization: Violating privacy via inference with large language models
R. Staab, M. Vero, M. Balunovi’c, and M. T. Vechev · 2023
Earlier work this paper cites.
Set-of-Mark prompting unleashes extraordinary visual grounding in gpt-4v
J. Yang, H. Zhang, F. Li, X. Zou, C. Li, and J. Gao · 2023
Cited alongside, same era.
Webarena: A realistic web environment for building autonomous agents
S. Zhou, F. F. Xu, H. Zhu, X. Zhou, R. Lo, A. Sridhar, X. Cheng, Y. Bisk, D. Fried, U. Alon, et al · 2023
Cited alongside, same era.
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
M. Andriushchenko, A. Souly, M. Dziemian, D. Duenas, M. Lin, J. Wang, D. Hendrycks, A. Zou, Z. Kolter, M. Fredrikson, E. Winsor, J. Wynne, Y. Gal, and X. Davies · 2024
Cited alongside, same era.
Airgapagent: Protecting privacy-conscious conversational agents
E. Bagdasarian, R. Yi, S. Ghalebikesabi, P. Kairouz, M. Gruteser, S. Oh, B. Balle, and D. Ramage · 2024
Cited alongside, same era.
Ci-bench: Benchmarking contextual integrity of ai assistants on synthetic data
Eia: Environmental injection attack on generalist web agents for privacy leakage, 2024
Z. Liao, L. Mo, C. Xu, M. Kang, J. Zhang, C. Xiao, Y. Tian, B. Li, and H. Sun · 2024
Later among the works it cites.
Meta Llama · 2024
Later among the works it cites.
PrivAgent: Agentic-based Red-teaming for LLM Privacy Leakage, 2024
Y. Nie, Z. Wang, Y. Yu, X. Wu, X. Zhao, W. Guo, and D. Song · 2024
Later among the works it cites.
tinybenchmarks: evaluating llms with fewer examples, 2024
F. M. Polo, L. Weber, L. Choshen, Y. Sun, G. Xu, and M. Yurochkin · 2024
Later among the works it cites.
Identifying the risks of lm agents with an lm-emulated sandbox
Y. Ruan, H. Dong, A. Wang, S. Pitis, Y. Zhou, J. Ba, Y. Dubois, C. J. Maddison, and T. Hashimoto · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Cheng, D. Wan, M. Abueg, S. Ghalebikesabi, R. Yi, E. Bagdasarian, B. Balle, S. Mellem, and S. O’Banion · 2024
Cited alongside, same era.
Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents
E. Debenedetti, J. Zhang, M. Balunovic, L. Beurer-Kellner, M. Fischer, and F. Tramèr · 2024
Cited alongside, same era.
Do membership inference attacks work on large language models?
M. Duan, A. Suri, N. Mireshghallah, S. Min, W. Shi, L. S. Zettlemoyer, Y. Tsvetkov, Y. Choi, D. Evans, and H. Hajishirzi · 2024
Cited alongside, same era.
Operationalizing contextual integrity in privacy-conscious assistants
S. Ghalebikesabi, E. Bagdasaryan, R. Yi, I. Yona, I. Shumailov, A. Pappu, C. Shi, L. Weidinger, R. Stanforth, L. Berrada, et al · 2024
Cited alongside, same era.
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024
E. Glazer, E. Erdil, T. Besiroglu, D. Chicharro, E. Chen, A. Gunning, C. F. Olsson, J.-S. Denain, A. Ho, E. de Oliveira Santos, O. Järviniemi, M. Barnett, R. Sandler, M. Vrzala, J. Sevilla, Q. Ren, E. Pratt, L. Levine, G. Barkley, N. Stewart, B. Grechuk, T. Grechuk, S. V. Enugandla, and M. Wildon · 2024
Cited alongside, same era.
Y. He, E. Wang, Y. Rong, Z. Cheng, and H. Chen · 2024
Cited alongside, same era.
Déjà vu memorization in vision-language models, 2024
B. Jayaraman, C. Guo, and K. Chaudhuri · 2024
Cited alongside, same era.
Visualwebarena: Evaluating multimodal agents on realistic visual web tasks
J. Y. Koh, R. Lo, L. Jang, V. Duvvur, M. C. Lim, P.-Y. Huang, G. Neubig, S. Zhou, R. Salakhutdinov, and D. Fried · 2024
Cited alongside, same era.
Y. Shao, T. Li, W. Shi, Y. Liu, and D. Yang · 2024
Later among the works it cites.
A new era in llm security: Exploring security concerns in real-world llm-based systems
F. Wu, N. Zhang, S. Jha, P. D. McDaniel, and C. Xiao · 2024
Later among the works it cites.
A survey on large language model (llm) security and privacy: The good, the bad, and the ugly
Y. Yao, J. Duan, K. Xu, Y. Cai, Z. Sun, and Y. Zhang · 2024
Later among the works it cites.
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents, 2024
H. Zhang, J. Huang, K. Mei, Y. Yao, Z. Wang, C. Zhan, H. Wang, and Y. Zhang · 2024
Later among the works it cites.
Gpt-4v(ision) is a generalist web agent, if grounded, 2024
B. Zheng, B. Gou, J. Kil, H. Sun, and Y. Su · 2024
Later among the works it cites.
Claude 3.5 Sonnet with Computer Use, 2024
Anthropic · 2025
Closest in time.
WASP: Benchmarking web agent security against prompt injection attacks
I. Evtimov, A. Zharmagambetov, A. Grattafiori, C. Guo, and K. Chaudhuri · 2025
Closest in time.