Fetching the paper…
Reading the bibliography…
Inspired by the rapid development of Large Language Models (LLMs), LLM agents have evolved to perform complex tasks.
Mitigating Poisoning Attacks on Machine Learning Models: A Data Provenance Based Approach. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security (AISec ’17) . Association for Computing Machinery, New York, NY, USA, 103–110
Nathalie Baracaldo, Bryant Chen, Heiko Ludwig, and Jaehoon Amir Safavi. 2017 · 2017
Earlier work this paper cites.
Ethical Challenges in Data-Driven Dialogue Systems
Peter Henderson, Koustuv Sinha, Nicolas Angelard-Gontier, Nan Rosemary Ke, Genevieve Fried, Ryan Lowe, and Joelle Pineau. 2017 · 2017
Earlier work this paper cites.
Universal Language Model Fine-tuning for Text Classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (Melbourne, Australia), Iryna Gurevych and Yusuke Miyao (Eds.). Association for Computational Linguistics, 328–339
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Earlier work this paper cites.
Privacy and Security of Big Data in AI Systems: A Research and Standards Perspective. In 2019 IEEE International Conference on Big Data (Big Data) . 5737–5743
Saharnaz Dilmaghani, Matthias R. Brust, Grégoire Danoy, Natalia Cassagnes, Johnatan Pecero, and Pascal Bouvry. 2019 · 2019
Earlier work this paper cites.
Machine Learning Security: Threats, Countermeasures, and Evaluations
Mingfu Xue, Chengxiang Yuan, Heyi Wu, Yushu Zhang, and Weiqiang Liu. 2020 · 2020
Earlier work this paper cites.
Extracting Training Data from Large Language Models. In 30th USENIX Security Symposium (USENIX Security 21) . 2633–2650
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, Alina Oprea, and Colin Raffel. 2021 · 2021
Earlier work this paper cites.
BERTective: Language Models and Contextual Information for Deception Detection. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (Online). Association for Computational Linguistics, 2699–2708
Tommaso Fornaciari, Federico Bianchi, Massimo Poesio, and Dirk Hovy. 2021 · 2021
Earlier work this paper cites.
Membership Privacy for Machine Learning Models Through Knowledge Transfer
Virat Shejwalkar and Amir Houmansadr. 2021 · 2021
Earlier work this paper cites.
Data-Free Model Extraction. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 4769–4778
Jean-Baptiste Truong, Pratyush Maini, Robert J. Walls, and Nicolas Papernot. 2021 · 2021
Earlier work this paper cites.
Differentially Private Fine-tuning of Language Models. In International Conference on Learning Representations
Da Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi, Huseyin A. Inan, Gautam Kamath, Janardhan Kulkarni, Yin Tat Lee, Andre Manoel, Lukas Wutschitz, Sergey Yekhanin, and Huishuai Zhang. 2021 · 2021
Earlier work this paper cites.
ChatGPT
2022 · 2022
Earlier work this paper cites.
LangChain
Harrison Chase. 2022 · 2022
Earlier work this paper cites.
Deduplicating Training Data Makes Language Models Better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (Dublin, Ireland), Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (Eds.). Association for Computational Linguistics, 8424–8445
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2022 · 2022
Earlier work this paper cites.
Multi-Objective Learning to Overcome Catastrophic Forgetting in Time-series Applications
Reem A. Mahmoud and Hazem Hajj. 2022 · 2022
Earlier work this paper cites.
Downstream Task Performance of BERT Models Pre-Trained Using Automatically De-Identified Clinical Data. In Proceedings of the Thirteenth Language Resources and Evaluation Conference (Marseille, France), Nicoletta Calzolari, Frédéric Béchet, Philippe Blache, Khalid Choukri, Christopher Cieri, Thierry Declerck, Sara Goggi, Hitoshi Isahara, Bente Maegaard, Joseph Mariani, Hélène Mazo, Jan Odijk, and Stelios Piperidis (Eds.). European Language Resources Association, 4245–4252
Thomas Vakili, Anastasios Lamproudis, Aron Henriksson, and Hercules Dalianis. 2022 · 2022
Earlier work this paper cites.
Gemini - Chat to Supercharge Your Ideas
2023 · 2023
Earlier work this paper cites.
Malicious ChatGPT Agents: How GPTs Can Quietly Grab Your Data (Demo) · Embrace The Red
Embrace The Red 2023 · 2023
Earlier work this paper cites.
What Are Large Language Model (LLM) Agents and Autonomous Agents
Prompt Engineering 2023 · 2023
Earlier work this paper cites.
Characterizing Attribution and Fluency Tradeoffs for Retrieval-Augmented Large Language Models
Renat Aksitov, Chung-Ching Chang, David Reitter, Siamak Shakeri, and Yunhsuan Sung. 2023 · 2023
Earlier work this paper cites.
Frontier AI Regulation: Managing Emerging Risks to Public Safety
Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, Cullen O’Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, Ben Chang, Tantum Collins, Tim Fist, Gillian Hadfield, Alan Hayes, Lewis Ho, Sara Hooker, Eric Horvitz, Noam Kolt, Jonas Schuett, Yonadav Shavit, Divya Siddarth, Robert Trager, and Kevin Wolf. 2023 · 2023
Earlier work this paper cites.
ChemCrow: Augmenting Large-Language Models with Chemistry Tools
Andres M. Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D. White, and Philippe Schwaller. 2023 · 2023
Earlier work this paper cites.
P. V. Sai Charan, Hrushikesh Chunduri, P. Mohan Anand, and Sandeep K. Shukla. 2023 · 2023
Earlier work this paper cites.
MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang, Ying Zhang, Zefeng Li, Haoyu Wang, Tianwei Zhang, and Yang Liu. 2023 · 2023
Earlier work this paper cites.
Toxicity in ChatGPT: Analyzing Persona-assigned Language Models
Ameet Deshpande, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, and Karthik Narasimhan. 2023 · 2023
Earlier work this paper cites.
Chain-of-Verification Reduces Hallucination in Large Language Models
Shehzaad Dhuliawala, Mojtaba Komeili, Jing Xu, Roberta Raileanu, Xian Li, Asli Celikyilmaz, and Jason Weston. 2023 · 2023
Earlier work this paper cites.
Decoding the Threat Landscape : ChatGPT, FraudGPT, and WormGPT in Social Engineering Attacks
Polra Victor Falade. 2023 · 2023
Earlier work this paper cites.
Wenjie Fu, Huandong Wang, Chen Gao, Guanghua Liu, Yong Li, and Tao Jiang. 2023 · 2023
Earlier work this paper cites.
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023 · 2023
Earlier work this paper cites.
MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Ceyao Zhang, Jinlin Wang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber. 2023 · 2023
Earlier work this paper cites.
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023 · 2023
Earlier work this paper cites.
Training Data Extraction From Pre-trained Language Models: A Survey. In Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023) (Toronto, Canada), Anaelia Ovalle, Kai-Wei Chang, Ninareh Mehrabi, Yada Pruksachatkun, Aram Galystan, Jwala Dhamala, Apurv Verma, Trista Cao, Anoop Kumar, and Rahul Gupta (Eds.). Association for Computational Linguistics, 260–275
Shotaro Ishihara. 2023 · 2023
Earlier work this paper cites.
Combing for Credentials: Active Pattern Extraction from Smart Reply
Bargav Jayaraman, Esha Ghosh, Melissa Chase, Sambuddha Roy, Wei Dai, and David Evans. 2023 · 2023
Cited alongside, same era.
Towards Mitigating Hallucination in Large Language Models via Self-Reflection
Ziwei Ji, Tiezheng Yu, Yan Xu, Nayeon Lee, Etsuko Ishii, and Pascale Fung. 2023 · 2023
Cited alongside, same era.
ProPILE: Probing Privacy Leakage in Large Language Models
Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and Seong Joon Oh. 2023 · 2023
Cited alongside, same era.
Multi-Step Jailbreaking Privacy Attacks on ChatGPT
Haoran Li, Dadi Guo, Wei Fan, Mingshi Xu, Jie Huang, Fanpu Meng, and Yangqiu Song. 2023 · 2023
Cited alongside, same era.
Noise Imitation Based Adversarial Training for Robust Multimodal Sentiment Analysis
Ziqi Yuan, Yihe Liu, Hua Xu, and Kai Gao. 2024 · 2023
Later among the works it cites.
Investigating the Catastrophic Forgetting in Multimodal Large Language Models
Yuexiang Zhai, Shengbang Tong, Xiao Li, Mu Cai, Qing Qu, Yong Jae Lee, and Yi Ma. 2023 · 2023
Later among the works it cites.
Synergizing Human-AI Agency: A Guide of 23 Heuristics for Service Co-Creation with LLM-Based Agents
Qingxiao Zheng, Zhongwei Xu, Abhinav Choudhary, Yuting Chen, Yongming Li, and Yun Huang. 2023 · 2023
Later among the works it cites.
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiaju Lin, Haoran Zhao, Aochi Zhang, Yiting Wu, Huqiuyue Ping, and Qin Chen. 2023 · 2023
Cited alongside, same era.
Zero-Resource Hallucination Prevention for Large Language Models
Junyu Luo, Cao Xiao, and Fenglong Ma. 2023a · 2023
Cited alongside, same era.
RoCo: Dialectic Multi-Robot Collaboration with Large Language Models
Zhao Mandi, Shreeya Jain, and Shuran Song. 2023 · 2023
Cited alongside, same era.
Mitigating Catastrophic Forgetting with Complementary Layered Learning
Sean Mondesire and R. Paul Wiegand. 2023 · 2023
Cited alongside, same era.
Med-Flamingo: A Multimodal Medical Few-shot Learner. In Proceedings of the 3rd Machine Learning for Health Symposium (Proceedings of Machine Learning Research, Vol. 225) . PMLR, 353–367
Michael Moor, Qian Huang, Shirley Wu, Michihiro Yasunaga, Yash Dalmia, Jure Leskovec, Cyril Zakka, Eduardo Pontes Reis, and Pranav Rajpurkar. 2023 · 2023
Cited alongside, same era.
Babyagi, 2023
Yohei Nakajima. [n. d.] · 2023
Cited alongside, same era.
Controlling the Extraction of Memorized Data from Large Language Models via Prompt-Tuning
Mustafa Safa Ozdayi, Charith Peris, Jack FitzGerald, Christophe Dupuy, Jimit Majmudar, Haidar Khan, Rahil Parikh, and Rahul Gupta. 2023 · 2023
Cited alongside, same era.
Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (2023-10-29) (UIST ’23) . Association for Computing Machinery, 1–22
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. [n. d.] · 2023
Cited alongside, same era.
AirGapAgent: Protecting Privacy-Conscious Conversational Agents. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security (CCS ’24) . Association for Computing Machinery, 3868–3882
Eugene Bagdasarian, Ren Yi, Sahra Ghalebikesabi, Peter Kairouz, Marco Gruteser, Sewoong Oh, Borja Balle, and Daniel Ramage. 2024 · 2024
Closest in time.
Security and Privacy Challenges of Large Language Models: A Survey
Badhan Chandra Das, M. Hadi Amini, and Yanzhao Wu. 2024 · 2024
Closest in time.
Multi-Agent Software Development through Cross-Team Collaboration
Zhuoyun Du, Chen Qian, Wei Liu, Zihao Xie, Yifei Wang, Yufan Dang, Weize Chen, and Cheng Yang. 2024 · 2024
Closest in time.
LLM Agents Can Autonomously Exploit One-day Vulnerabilities
Richard Fang, Rohan Bindu, Akul Gupta, and Daniel Kang. 2024 · 2024
Closest in time.
Imprompter: Tricking LLM Agents into Improper Tool Use
Xiaohan Fu, Shuheng Li, Zihan Wang, Yihao Liu, Rajesh K. Gupta, Taylor Berg-Kirkpatrick, and Earlence Fernandes. 2024 · 2024
Closest in time.
CoCA: Regaining Safety-awareness of Multimodal Large Language Models with Constitutional Calibration
Jiahui Gao, Renjie Pi, Tianyang Han, Han Wu, Lanqing Hong, Lingpeng Kong, Xin Jiang, and Zhenguo Li. 2024 · 2024
Closest in time.
Large Language Model Based Multi-Agents: A Survey of Progress and Challenges
Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V. Chawla, Olaf Wiest, and Xiangliang Zhang. 2024 · 2024
Closest in time.
Defending Against Indirect Prompt Injection Attacks With Spotlighting
Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman. 2024 · 2024
Closest in time.
Fine-Tuning Large Language Models with Sequential Instructions
Hanxu Hu, Pinzhen Chen, and Edoardo M. Ponti. 2024 · 2024
Closest in time.
Sleeper Agents: Training Deceptive LLMs That Persist Through Safety Training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez. 2024 · 2024
Closest in time.
Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities
Tianjie Ju, Yiting Wang, Xinbei Ma, Pengzhou Cheng, Haodong Zhao, Yulong Wang, Lifeng Liu, Jian Xie, Zhuosheng Zhang, and Gongshen Liu. 2024 · 2024
Closest in time.
User Inference Attacks on Large Language Models
Nikhil Kandpal, Krishna Pillutla, Alina Oprea, Peter Kairouz, Christopher A. Choquette-Choo, and Zheng Xu. 2024 · 2024
Closest in time.
Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision
Seongyun Lee, Sue Hyun Park, Yongrae Jo, and Minjoon Seo. 2024 · 2024
Closest in time.
Mitigating Object Hallucinations in Large Vision-Language Models through Visual Contrastive Decoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 13872–13882
Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing. 2024 · 2024
Closest in time.
PrivAgent: Agentic-based Red-teaming for LLM Privacy Leakage
Yuzhou Nie, Zhun Wang, Ye Yu, Xian Wu, Xuandong Zhao, Wenbo Guo, and Dawn Song. 2024 · 2024
Closest in time.
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, and Balcom. 2024 · 2024
Closest in time.
Empowering Language Models with Active Inquiry for Deeper Understanding
Jing-Cheng Pang, Heng-Bo Fan, Pengyuan Wang, Jia-Hao Xiao, Nan Tang, Si-Hang Yang, Chengxing Jia, Sheng-Jun Huang, and Yang Yu. 2024 · 2024
Closest in time.
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J. Maddison, and Tatsunori Hashimoto. 2024 · 2024
Closest in time.
InferDPT: Privacy-Preserving Inference for Black-box Large Language Model
Meng Tong, Kejiang Chen, Jie Zhang, Yuang Qi, Weiming Zhang, Nenghai Yu, Tianwei Zhang, and Zhikun Zhang. 2024 · 2024
Closest in time.
Large Multimodal Agents: A Survey
Junlin Xie, Zhihong Chen, Ruifei Zhang, Xiang Wan, and Guanbin Li. 2024 · 2024
Closest in time.
VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models
Ziyi Yin, Muchao Ye, Tianrong Zhang, Tianyu Du, Jinguo Zhu, Han Liu, Jinghui Chen, Ting Wang, and Fenglong Ma. 2024 · 2024
Closest in time.
HallE-Control: Controlling Object Hallucination in Large Multimodal Models
Bohan Zhai, Shijia Yang, Chenfeng Xu, Sheng Shen, Kurt Keutzer, Chunyuan Li, and Manling Li. 2024 · 2024
Closest in time.
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang. 2024 · 2024
Closest in time.
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics . 8865–8887
Zhexin Zhang, Junxiao Yang, Pei Ke, Fei Mi, Hongning Wang, and Minlie Huang. 2024d · 2024
Closest in time.
SAUP: Situation Awareness Uncertainty Propagation on LLM Agent
Qiwei Zhao, Xujiang Zhao, Yanchi Liu, Wei Cheng, Yiyou Sun, Mika Oishi, Takao Osaki, Katsushi Matsuda, Huaxiu Yao, and Haifeng Chen. 2024 · 2024
Closest in time.
Memorybank: Enhancing large language models with long-term memory. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 38. 19724–19731
Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang. 2024 · 2024
Closest in time.
PoisonedRAG: Knowledge Poisoning Attacks to Retrieval-Augmented Generation of Large Language Models. In 34th USENIX Security Symposium (USENIX Security 25) . 3827–3844
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2025 · 2025
Closest in time.