Fetching the paper…
Reading the bibliography…
With the continuous development of large language models (LLMs), transformer-based models have made groundbreaking advances in numerous natural language processing (NLP) tasks, leading to the emergence of a series of agents that use LLMs as their control hub.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
Membership inference attacks from first principles. In 2022 IEEE Symposium on Security and Privacy (SP) . IEEE, 1897–1914
Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, and Florian Tramer. 2022 · 1914
Earlier work this paper cites.
CrowS-Pairs: A challenge dataset for measuring social biases in masked language models. In 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020 . Association for Computational Linguistics (ACL), 1953–1967
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R Bowman. 2020 · 1967
Earlier work this paper cites.
Privacy as contextual integrity
Helen Nissenbaum. 2004 · 2004
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2013 · 2013
Earlier work this paper cites.
The algorithmic foundations of differential privacy
Cynthia Dwork, Aaron Roth, et al · 2014
Earlier work this paper cites.
Practical planning: extending the classical AI planning paradigm
David E Wilkins. 2014 · 2014
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015 · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Distillation as a defense to adversarial perturbations against deep neural networks. In 2016 IEEE symposium on security and privacy (SP) . IEEE, 582–597
Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami. 2016 · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Regulation (EU) 2016/679 of the European Parliament and of the Council
Protection Regulation. 2016 · 2016
Earlier work this paper cites.
A causal framework for discovering and removing direct and indirect discrimination
Lu Zhang, Yongkai Wu, and Xintao Wu. 2016 · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan. 2017 · 2017
Earlier work this paper cites.
Towards evaluating the robustness of neural networks. In 2017 ieee symposium on security and privacy (sp) . Ieee, 39–57
Nicholas Carlini and David Wagner. 2017 · 2017
Earlier work this paper cites.
Conscientious classification: A data scientist’s guide to discrimination-aware classification
Brian d’Alessandro, Cathy O’Neil, and Tom LaGatta. 2017 · 2017
Earlier work this paper cites.
Algorithmic Bias in Autonomous Systems.. In Ijcai , Vol. 17. 4691–4697
David Danks and Alex John London. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models. In 2017 IEEE symposium on security and privacy (SP) . IEEE, 3–18
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. In International conference on machine learning . PMLR, 274–283
Anish Athalye, Nicholas Carlini, and David Wagner. 2018 · 2018
Earlier work this paper cites.
Bias on the web
Ricardo Baeza-Yates. 2018 · 2018
Earlier work this paper cites.
Thermometer encoding: One hot way to resist adversarial examples. In International conference on learning representations
Jacob Buckman, Aurko Roy, Colin Raffel, and Ian Goodfellow. 2018 · 2018
Earlier work this paper cites.
Gender shades: Intersectional accuracy disparities in commercial gender classification. In Conference on fairness, accountability and transparency . PMLR, 77–91
Joy Buolamwini and Timnit Gebru. 2018 · 2018
Earlier work this paper cites.
Learning and memorization. In International conference on machine learning . PMLR, 755–763
Satrajit Chatterjee. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin. 2018 · 2018
Earlier work this paper cites.
Robust physical-world attacks on deep learning visual classification. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1625–1634
Kevin Eykholt, Ivan Evtimov, Earlence Fernandes, Bo Li, Amir Rahmati, Chaowei Xiao, Atul Prakash, Tadayoshi Kohno, and Dawn Song. 2018 · 2018
Earlier work this paper cites.
Hallucinations in neural machine translation
Katherine Lee, Orhan Firat, Ashish Agarwal, Clara Fannjiang, and David Sussillo. 2018 · 2018
Earlier work this paper cites.
Object hallucination in image captioning
Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. 2018 · 2018
Earlier work this paper cites.
Thieves on sesame street! model extraction of bert-based apis
Kalpesh Krishna, Gaurav Singh Tomar, Ankur P Parikh, Nicolas Papernot, and Mohit Iyyer. 2019 · 2019
Earlier work this paper cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 2019
Earlier work this paper cites.
On measuring social biases in sentence encoders
Chandler May, Alex Wang, Shikha Bordia, Samuel R Bowman, and Rachel Rudinger. 2019 · 2019
Earlier work this paper cites.
Model reconstruction from model explanations. In Proceedings of the Conference on Fairness, Accountability, and Transparency . 1–9
Smitha Milli, Ludwig Schmidt, Anca D Dragan, and Moritz Hardt. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
A framework for understanding unintended consequences of machine learning
Harini Suresh and John V Guttag. 2019 · 2019
Earlier work this paper cites.
Fairness without harm: Decoupled classifiers with preference guarantees. In International Conference on Machine Learning . PMLR, 6373–6382
Berk Ustun, Yang Liu, and David Parkes. 2019 · 2019
Earlier work this paper cites.
Towards causal vqa: Revealing and reducing spurious correlations by invariant and covariant semantic editing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 9690–9698
Vedika Agarwal, Rakshith Shetty, and Mario Fritz. 2020 · 2020
Earlier work this paper cites.
Cryptanalytic extraction of neural network models. In Annual international cryptology conference . Springer, 189–218
Nicholas Carlini, Matthew Jagielski, and Ilya Mironov. 2020 · 2020
Earlier work this paper cites.
Hopskipjumpattack: A query-efficient decision-based attack. In 2020 ieee symposium on security and privacy (sp) . IEEE, 1277–1294
Jianbo Chen, Michael I Jordan, and Martin J Wainwright. 2020 · 2020
Earlier work this paper cites.
Does learning require memorization? a short tale about a long tail. In Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing . 954–959
Vitaly Feldman. 2020 · 2020
Earlier work this paper cites.
Codebert: A pre-trained model for programming and natural languages
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, et al · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020 · 2020
Earlier work this paper cites.
Dataset Security for Machine Learning: Data Poisoning, Backdoor Attacks, and Defenses
Micah Goldblum, Dimitris Tsipras, Chulin Xie, Xinyun Chen, Avi Schwarzschild, Dawn Song, Aleksander Madry, Bo Li, and Tom Goldstein. 2020 · 2020
Earlier work this paper cites.
High accuracy and high fidelity extraction of neural networks. In 29th USENIX security symposium (USENIX Security 20) . 1345–1362
Matthew Jagielski, Nicholas Carlini, David Berthelot, Alex Kurakin, and Nicolas Papernot. 2020 · 2020
Earlier work this paper cites.
Thai Le, Noseong Park, and Dongwon Lee. 2020 · 2020
Earlier work this paper cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2020 · 2020
Earlier work this paper cites.
Infobert: Improving robustness of language models from an information theoretic perspective
Boxin Wang, Shuohang Wang, Yu Cheng, Zhe Gan, Ruoxi Jia, Bo Li, and Jingjing Liu. 2020 · 2020
Earlier work this paper cites.
Measuring and reducing gendered correlations in pre-trained models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, Ed Chi, and Slav Petrov. 2020 · 2020
Earlier work this paper cites.
Detecting hallucinated content in conditional neural sequence generation
Chunting Zhou, Graham Neubig, Jiatao Gu, Mona Diab, Paco Guzman, Luke Zettlemoyer, and Marjan Ghazvininejad. 2020 · 2020
Earlier work this paper cites.
E2B Sandbox
2021 · 2021
Earlier work this paper cites.
Rongzhou Bao, Jiayi Wang, and Hai Zhao. 2021 · 2021
Earlier work this paper cites.
Soumya Barikeri, Anne Lauscher, Ivan Vulić, and Goran Glavaš. 2021 · 2021
Earlier work this paper cites.
Machine unlearning. In 2021 IEEE Symposium on Security and Privacy (SP) . IEEE, 141–159
Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot. 2021 · 2021
Earlier work this paper cites.
Extracting training data from large language models. In 30th USENIX Security Symposium (USENIX Security 21) . 2633–2650
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
How should pre-trained language models be fine-tuned towards adversarial robustness?
Xinshuai Dong, Anh Tuan Luu, Min Lin, Shuicheng Yan, and Hanwang Zhang. 2021 · 2021
Earlier work this paper cites.
Cert-RNN: Towards Certifying the Robustness of Recurrent Neural Networks
Tianyu Du, Shouling Ji, Lujia Shen, Yao Zhang, Jinfeng Li, Jie Shi, Chengfang Fang, Jianwei Yin, Raheem Beyah, and Ting Wang. 2021 · 2021
Earlier work this paper cites.
Neural path hunter: Reducing hallucination in dialogue systems via path grounding
Nouha Dziri, Andrea Madotto, Osmar Zaïane, and Avishek Joey Bose. 2021 · 2021
Earlier work this paper cites.
Amnesiac machine learning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 11516–11524
Laura Graves, Vineel Nagisetty, and Vijay Ganesh. 2021 · 2021
Earlier work this paper cites.
Detecting emergent intersectional biases: Contextualized word embeddings contain a distribution of human-like biases. In Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society . 122–133
Wei Guo and Aylin Caliskan. 2021 · 2021
Earlier work this paper cites.
Model extraction and adversarial transferability, your BERT is vulnerable!
Xuanli He, Lingjuan Lyu, Qiongkai Xu, and Lichao Sun. 2021 · 2021
Earlier work this paper cites.
Learning and evaluating a differentially private pre-trained language model. In Findings of the Association for Computational Linguistics: EMNLP 2021 . 1178–1189
Shlomo Hoory, Amir Feder, Avichai Tendler, Sofia Erell, Alon Peled-Cohen, Itay Laish, Hootan Nakhost, Uri Stemmer, Ayelet Benjamini, Avinatan Hassidim, et al · 2021
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021 · 2021
Earlier work this paper cites.
Yannik Keller, Jan Mackensen, and Steffen Eger. 2021 · 2021
Earlier work this paper cites.
Sustainable Modular Debiasing of Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2021 . 4782–4797
Anne Lauscher, Tobias Lueken, and Goran Glavaš. 2021 · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2021 · 2021
Earlier work this paper cites.
Anonymisation models for text data: State of the art, challenges and future directions. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . 4188–4203
Pierre Lison, Ildikó Pilán, David Sánchez, Montserrat Batet, and Lilja Øvrelid. 2021 · 2021
Earlier work this paper cites.
A survey on bias and fairness in machine learning
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021 · 2021
Earlier work this paper cites.
StereoSet: Measuring stereotypical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . 5356–5371
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021 · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Earlier work this paper cites.
Descent-to-delete: Gradient-based methods for machine unlearning. In Algorithmic Learning Theory . PMLR, 931–962
Seth Neel, Aaron Roth, and Saeed Sharifi-Malvajerdi. 2021 · 2021
Earlier work this paper cites.
Sashank Santhanam, Behnam Hedayatnia, Spandana Gella, Aishwarya Padmakumar, Seokhwan Kim, Yang Liu, and Dilek Hakkani-Tur. 2021 · 2021
Earlier work this paper cites.
Backdoor Pre-trained Models Can Transfer to All
Lujia Shen, Shouling Ji, Xuhong Zhang, Jinfeng Li, Jing Chen, Jie Shi, Chengfang Fang, Jianwei Yin, and Ting Wang. 2021 · 2021
Earlier work this paper cites.
On memorization in probabilistic deep generative models
Gerrit van den Burg and Chris Williams. 2021 · 2021
Earlier work this paper cites.
Yue Wang, Weishi Wang, Shafiq Joty, and Steven CH Hoi. 2021 · 2021
Earlier work this paper cites.
Student surpasses teacher: Imitation attack for black-box NLP APIs
Qiongkai Xu, Xuanli He, Lingjuan Lyu, Lizhen Qu, and Gholamreza Haffari. 2021 · 2021
Earlier work this paper cites.
Differential Privacy for Text Analytics via Natural Text Sanitization. In Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 . 3853–3866
Xiang Yue, Minxin Du, Tianhao Wang, Yaliang Li, Huan Sun, and Sherman SM Chow. 2021 · 2021
Earlier work this paper cites.
Grey-box extraction of natural language models. In International Conference on Machine Learning . PMLR, 12278–12286
Santiago Zanella-Beguelin, Shruti Tople, Andrew Paverd, and Boris Köpf. 2021 · 2021
Earlier work this paper cites.
Defense against synonym substitution-based adversarial attacks via Dirichlet neighborhood ensemble. In Association for Computational Linguistics (ACL)
Yi Zhou, Xiaoqing Zheng, Cho-Jui Hsieh, Kai-Wei Chang, and Xuanjing Huan. 2021 · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al · 2022
Earlier work this paper cites.
Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures
Eugene Bagdasaryan and Vitaly Shmatikov. 2022 · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
Let there be a clock on the beach: Reducing object hallucination in image captioning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 1381–1390
Ali Furkan Biten, Lluís Gómez, and Dimosthenis Karatzas. 2022 · 2022
Earlier work this paper cites.
LangChain
Harrison Chase. 2022 · 2022
Earlier work this paper cites.
Measuring fairness with biased rulers: A comparative study on bias metrics for pre-trained language models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics . Association for Computational Linguistics, 1693–1706
Pieter Delobelle, Ewoenam Kwaku Tokpo, Toon Calders, and Bettina Berendt. 2022 · 2022
Earlier work this paper cites.
Demographic-aware language model fine-tuning as a bias mitigation technique. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing (Volume 2: Short Papers) . 311–319
Aparna Garimella, Rada Mihalcea, and Akhash Amarnath. 2022 · 2022
Earlier work this paper cites.
Balancing out Bias: Achieving Fairness Through Balanced Training. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 11335–11350
Xudong Han, Timothy Baldwin, and Trevor Cohn. 2022a · 2022
Earlier work this paper cites.
Towards equal opportunity fairness through adversarial learning
Xudong Han, Timothy Baldwin, and Trevor Cohn. 2022b · 2022
Earlier work this paper cites.
Cater: Intellectual property protection on text generation apis via conditional watermarks
Xuanli He, Qiongkai Xu, Yi Zeng, Lingjuan Lyu, Fangzhao Wu, Jiwei Li, and Ruoxi Jia. 2022 · 2022
Earlier work this paper cites.
Are Large Pre-Trained Language Models Leaking Your Personal Information?. In 2022 Findings of the Association for Computational Linguistics: EMNLP 2022
Jie Huang, Hanyin Shao, and Kevin Chen Chuan Chang. 2022b · 2022
Earlier work this paper cites.
Unmasking the mask–evaluating social biases in masked language models. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 36. 11954–11962
Masahiro Kaneko and Danushka Bollegala. 2022 · 2022
Earlier work this paper cites.
Deduplicating Training Data Makes Language Models Better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 8424–8445
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2022 · 2022
Earlier work this paper cites.
Text adversarial purification as defense against adversarial attacks
Linyang Li, Demin Song, and Xipeng Qiu. 2022a · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Earlier work this paper cites.
LlamaIndex
Jerry Liu. 2022 · 2022
Earlier work this paper cites.
" That Is a Suspicious Reaction!": Interpreting Logits Variation to Detect NLP Adversarial Attacks
Edoardo Mosca, Shreyash Agarwal, Javier Rando, and Georg Groh. 2022 · 2022
Earlier work this paper cites.
A survey of machine unlearning
Thanh Tam Nguyen, Thanh Trung Huynh, Phi Le Nguyen, Alan Wee-Chung Liew, Hongzhi Yin, and Quoc Viet Hung Nguyen. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
BBQ: A hand-built bias benchmark for question answering
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022 · 2022
Earlier work this paper cites.
Asleep at the keyboard? assessing the security of github copilot’s code contributions. In 2022 IEEE Symposium on Security and Privacy (SP) . IEEE, 754–768
Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri. 2022 · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro. 2022 · 2022
Earlier work this paper cites.
First the Worst: Finding Better Gender Translations During Beam Search. In Findings of the Association for Computational Linguistics: ACL 2022 . 3814–3823
Danielle Saunders, Rosie Sallis, and Bill Byrne. 2022 · 2022
Earlier work this paper cites.
" I’m sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset
Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022 · 2022
Earlier work this paper cites.
Mutual information alleviates hallucinations in abstractive summarization
Liam Van der Poel, Ryan Cotterell, and Clara Meister. 2022 · 2022
Cited alongside, same era.
On the Importance of Difficulty Calibration in Membership Inference Attacks. In International Conference on Learning Representations
Lauren Watson, Chuan Guo, Graham Cormode, and Alexandre Sablayrolles. 2022 · 2022
Cited alongside, same era.
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
Differentially Private Fine-tuning of Language Models. In International Conference on Learning Representations
Da Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi, Huseyin A Inan, Gautam Kamath, Janardhan Kulkarni, Yin Tat Lee, Andre Manoel, Lukas Wutschitz, Sergey Yekhanin, and Huishuai Zhang. 2022 · 2022
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023 · 2023
Later among the works it cites.
Who we are
2024 · 2024
Closest in time.
Are you still on track!? Catching LLM Task Drift with Activations
Sahar Abdelnabi, Aideen Fay, Giovanni Cherubin, Ahmed Salem, Mario Fritz, and Andrew Paverd. 2024 · 2024
Closest in time.
Investigating the prompt leakage effect and black-box defenses for multi-turn LLM interactions
Divyansh Agarwal, Alexander R Fabbri, Philippe Laban, Shafiq Joty, Caiming Xiong, and Chien-Sheng Wu. 2024 · 2024
Closest in time.
Many-shot jailbreaking
Cem Anil, Esin Durmus, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Nina Rimsky, Meg Tong, Jesse Mu, Daniel Ford, et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Cited alongside, same era.
Vlstereoset: A study of stereotypical bias in pre-trained vision-language models. Association for Computational Linguistics
Kankan Zhou, Yibin LAI, and Jing Jiang. 2022 · 2022
Cited alongside, same era.
TrojanPuzzle: Covertly Poisoning Code-Suggestion Models
Hojjat Aghakhani, Wei Dai, Andre Manoel, Xavier Fernandes, Anant Kharkar, Christopher Kruegel, Giovanni Vigna, David Evans, Ben Zorn, and Robert Sim. 2023 · 2023
Cited alongside, same era.
Detecting language model attacks with perplexity
Gabriel Alon and Michael Kamfonas. 2023 · 2023
Cited alongside, same era.
Exploiting Biased Models to De-bias Text: A Gender-Fair Rewriting Model. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 4486–4506
Chantal Amrhein, Florian Schottmann, Rico Sennrich, and Samuel Läubli. 2023 · 2023
Cited alongside, same era.
Red-teaming large language models using chain of utterances for safety-alignment
Rishabh Bhardwaj and Soujanya Poria. 2023 · 2023
Cited alongside, same era.
On the independence of association bias and empirical fairness in language models. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency . 370–378
Laura Cabello, Anna Katrine Jørgensen, and Anders Søgaard. 2023 · 2023
Cited alongside, same era.
Gamegpt: Multi-agent collaborative framework for game development
Dake Chen, Hanbin Wang, Yunhao Huo, Yuzhao Li, and Haoyang Zhang. 2023 · 2023
Cited alongside, same era.
Air Gap: Protecting Privacy-Conscious Conversational Agents
Eugene Bagdasaryan, Ren Yi, Sahra Ghalebikesabi, Peter Kairouz, Marco Gruteser, Sewoong Oh, Borja Balle, and Daniel Ramage. 2024 · 2024
Closest in time.
Emergent and predictable memorization in large language models
Stella Biderman, Usvsn Prashanth, Lintang Sutawika, Hailey Schoelkopf, Quentin Anthony, Shivanshu Purohit, and Edward Raff. 2024 · 2024
Closest in time.
VisDiaHalBench: A Visual Dialogue Benchmark For Diagnosing Hallucination in Large Vision-Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) . 12161–12176
Qingxing Cao, Junhao Cheng, Xiaodan Liang, and Liang Lin. 2024 · 2024
Closest in time.
Are aligned neural networks adversarially aligned?
Nicholas Carlini, Milad Nasr, Christopher A Choquette-Choo, Matthew Jagielski, Irena Gao, Pang Wei W Koh, Daphne Ippolito, Florian Tramer, and Ludwig Schmidt. 2024a · 2024
Closest in time.
Stealing part of a production language model
Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke, Jonathan Hayase, A Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, et al · 2024
Closest in time.
Jailbreakbench: An open robustness benchmark for jailbreaking large language models
Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al · 2024
Closest in time.
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models
Kedi Chen, Qin Chen, Jie Zhou, Yishen He, and Liang He. 2024a · 2024
Closest in time.
StruQ: Defending against prompt injection with structured queries
Sizhe Chen, Julien Piet, Chawin Sitawarin, and David Wagner. 2024c · 2024
Closest in time.
AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
Zhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song, and Bo Li. 2024d · 2024
Closest in time.
Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
Yixin Cheng, Markos Georgopoulos, Volkan Cevher, and Grigorios G Chrysos. 2024 · 2024
Closest in time.
Comprehensive assessment of jailbreak attacks against llms
Junjie Chu, Yugeng Liu, Ziqing Yang, Xinyue Shen, Michael Backes, and Yang Zhang. 2024 · 2024
Closest in time.
Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications
Stav Cohen, Ron Bitton, and Ben Nassi. 2024 · 2024
Closest in time.
Risk taxonomy, mitigation, and assessment benchmarks of large language model systems
Tianyu Cui, Yanling Wang, Chuanpu Fu, Yong Xiao, Sijia Li, Xinhao Deng, Yunpeng Liu, Qinglin Zhang, Ziyi Qiu, Peiyang Li, et al · 2024
Closest in time.
AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways
Zehang Deng, Yongjian Guo, Changzhou Han, Wanlun Ma, Junwu Xiong, Sheng Wen, and Yang Xiang. 2024a · 2024
Closest in time.
OpenBias: Open-set Bias Detection in Text-to-Image Generative Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 12225–12235
Moreno D’Incà, Elia Peruzzo, Massimiliano Mancini, Dejia Xu, Vidit Goel, Xingqian Xu, Zhangyang Wang, Humphrey Shi, and Nicu Sebe. 2024 · 2024
Closest in time.
Detection, Diagnosis, and Explanation: A Benchmark for Chinese Medial Hallucination Evaluation. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) . 4784–4794
Chengfeng Dou, Ying Zhang, Yanyuan Chen, Zhi Jin, Wenpin Jiao, Haiyan Zhao, and Yu Huang. 2024 · 2024
Closest in time.
h4rm3l: A Dynamic Benchmark of Composable Jailbreak Attacks for LLM Safety Assessment
Moussa Koulako Bala Doumbouya, Ananjan Nandi, Gabriel Poesia, Davide Ghilardi, Anna Goldie, Federico Bianchi, Dan Jurafsky, and Christopher D Manning. 2024 · 2024
Closest in time.
Multi-modal hallucination control by visual information grounding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 14303–14312
Alessandro Favero, Luca Zancato, Matthew Trager, Siddharth Choudhary, Pramuditha Perera, Alessandro Achille, Ashwin Swaminathan, and Stefano Soatto. 2024 · 2024
Closest in time.
Bias and fairness in large language models: A survey
Isabel O Gallegos, Ryan A Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K Ahmed. 2024 · 2024
Closest in time.
Agent smith: A single image can jailbreak one million multimodal llm agents exponentially fast
Xiangming Gu, Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Ye Wang, Jing Jiang, and Min Lin. 2024 · 2024
Closest in time.
Visogender: A dataset for benchmarking gender bias in image-text pronoun resolution
Siobhan Mackenzie Hall, Fernanda Gonçalves Abrantes, Hanwen Zhu, Grace Sodunke, Aleksandar Shtedritski, and Hannah Rose Kirk. 2024 · 2024
Closest in time.
Defending Against Indirect Prompt Injection Attacks With Spotlighting
Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati, Yonatan Zunger, and Emre Kiciman. 2024 · 2024
Closest in time.
SocialCounterfactuals: Probing and Mitigating Intersectional Social Biases in Vision-Language Models with Counterfactual Examples. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 11975–11985
Phillip Howard, Avinash Madasu, Tiep Le, Gustavo Lujan Moreno, Anahita Bhiwandiwalla, and Vasudev Lal. 2024 · 2024
Closest in time.
Semantic-guided Prompt Organization for Universal Goal Hijacking against LLMs
Yihao Huang, Chong Wang, Xiaojun Jia, Qing Guo, Felix Juefei-Xu, Jian Zhang, Geguang Pu, and Yang Liu. 2024b · 2024
Closest in time.
Evan Hubinger, Carson Denison, and Jesse Mu et al. 2024 · 2024
Closest in time.
PLeak: Prompt Leaking Attacks against Large Language Model Applications
Bo Hui, Haolin Yuan, Neil Gong, Philippe Burlina, and Yinzhi Cao. 2024 · 2024
Closest in time.
Defending large language models against jailbreak attacks via semantic smoothing
Jiabao Ji, Bairu Hou, Alexander Robey, George J Pappas, Hamed Hassani, Yang Zhang, Eric Wong, and Shiyu Chang. 2024 · 2024
Closest in time.
Hallucination augmented contrastive learning for multimodal large language model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 27036–27046
Chaoya Jiang, Haiyang Xu, Mengfan Dong, Jiaxing Chen, Wei Ye, Ming Yan, Qinghao Ye, Ji Zhang, Fei Huang, and Shikun Zhang. 2024 · 2024
Closest in time.
Exploring Backdoor Attacks against Large Language Model-based Decision Making
Ruochen Jiao, Shaoyuan Xie, Justin Yue, Takami Sato, Lixu Wang, Yixuan Wang, Qi Alfred Chen, and Qi Zhu. 2024 · 2024
Closest in time.
Attackeval: How to evaluate the effectiveness of jailbreak attacking on large language models
Mingyu Jin, Suiyuan Zhu, Beichen Wang, Zihao Zhou, Chong Zhang, Yongfeng Zhang, et al · 2024
Closest in time.
THRONE: An object-based hallucination benchmark for the free-form generations of large vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 27228–27238
Prannay Kaul, Zhizhong Li, Hao Yang, Yonatan Dukler, Ashwin Swaminathan, CJ Taylor, and Stefano Soatto. 2024 · 2024
Closest in time.
Propile: Probing privacy leakage in large language models
Siwon Kim, Sangdoo Yun, Hwaran Lee, Martin Gubri, Sungroh Yoon, and Seong Joon Oh. 2024 · 2024
Closest in time.
Subaru Kimura, Ryota Tanaka, Shumpei Miyawaki, Jun Suzuki, and Keisuke Sakaguchi. 2024 · 2024
Closest in time.
InjectBench: An Indirect Prompt Injection Benchmarking Framework
Nicholas Ka-Shing Kong. 2024 · 2024
Closest in time.
Mitigating object hallucinations in large vision-language models through visual contrastive decoding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 13872–13882
Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing. 2024 · 2024
Closest in time.
Vocabulary Attack to Hijack Large Language Model Applications
Patrick Levi and Christoph P Neumann. 2024 · 2024
Closest in time.
Yifan Li, Hangyu Guo, Kun Zhou, Wayne Xin Zhao, and Ji-Rong Wen. 2024b · 2024
Closest in time.
BackdoorLLM: A Comprehensive Benchmark for Backdoor Attacks on Large Language Models
Yige Li, Hanxun Huang, Yunhan Zhao, Xingjun Ma, and Jun Sun. 2024c · 2024
Closest in time.
Siyuan Liang, Kuanrong Liu, Jiajun Gong, Jiawei Liang, Yuan Xun, Ee-Chien Chang, and Xiaochun Cao. 2024b · 2024
Closest in time.
Why Are My Prompts Leaked? Unraveling Prompt Extraction Threats in Customized Large Language Models
Zi Liang, Haibo Hu, Qingqing Ye, Yaxin Xiao, and Haoyang Li. 2024a · 2024
Closest in time.
Compromising Embodied Agents with Contextual Backdoor Attacks
Aishan Liu, Yuguang Zhou, and Xianglong Liu et al. 2024e · 2024
Closest in time.
Odyssey: Empowering Agents with Open-World Skills
Shunyu Liu, Yaoru Li, Kongcheng Zhang, Zhenyu Cui, Wenkai Fang, Yuxuan Zheng, Tongya Zheng, and Mingli Song. 2024b · 2024
Closest in time.
Tong Liu, Yingjie Zhang, Zhe Zhao, Yinpeng Dong, Guozhu Meng, and Kai Chen. 2024d · 2024
Closest in time.
Automatic and universal prompt injection attacks against large language models
Xiaogeng Liu, Zhiyuan Yu, Yizhe Zhang, Ning Zhang, and Chaowei Xiao. 2024c · 2024
Closest in time.
Jiarui Lu, Thomas Holleis, and Yizhe Zhang et al. 2024a · 2024
Closest in time.
Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao. 2024a · 2024
Closest in time.
HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation
Wen Luo, Tianshu Shen, Wei Li, Guangyue Peng, Richeng Xuan, Houfeng Wang, and Xi Yang. 2024b · 2024
Closest in time.
Connecting certified and adversarial training
Yuhao Mao, Mark Müller, Marc Fischer, and Martin Vechev. 2024 · 2024
Closest in time.
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory. In The Twelfth International Conference on Learning Representations
Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi. 2024 · 2024
Closest in time.
Jailbreaking attack against multimodal large language model
Zhenxing Niu, Haodong Ren, Xinbo Gao, Gang Hua, and Rong Jin. 2024 · 2024
Closest in time.
Neural Exec: Learning (and Learning from) Execution Triggers for Prompt Injection Attacks
Dario Pasquini, Martin Strohmeier, and Carmela Troncoso. 2024 · 2024
Closest in time.
Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks. In The Twelfth International Conference on Learning Representations
Vaidehi Patil, Peter Hase, and Mohit Bansal. 2024 · 2024
Closest in time.
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models
Delong Ran, Jinyuan Liu, Yichen Gong, Jingyi Zheng, Xinlei He, Tianshuo Cong, and Anyu Wang. 2024 · 2024
Closest in time.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2024 · 2024
Closest in time.
Prompt stealing attacks against large language models
Zeyang Sha and Yang Zhang. 2024 · 2024
Closest in time.
SPML: A DSL for Defending Language Models Against Prompt Attacks
Reshabh K Sharma, Vinayak Gupta, and Dan Grossman. 2024 · 2024
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face
Yongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li, Weiming Lu, and Yueting Zhuang. 2024c · 2024
Closest in time.
Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles
Xiongtao Sun, Deyue Zhang, Dongdong Yang, Quanchen Zou, and Hui Li. 2024b · 2024
Closest in time.
Benchmarking Hallucination in Large Language Models based on Unanswerable Math Word Problem
Yuhong Sun, Zhangyue Yin, Qipeng Guo, Jiawen Wu, Xipeng Qiu, and Hui Zhao. 2024a · 2024
Closest in time.
Discover the best AI tools with Supertools
Supertools. [n. d.] · 2024
Closest in time.
GenderCARE: A Comprehensive Framework for Assessing and Reducing Gender Bias in Large Language Models. In Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security
Kunsheng Tang, Wenbo Zhou, Jie Zhang, Aishan Liu, Gelei Deng, Shuai Li, Peigui Qi, Weiming Zhang, Tianwei Zhang, and Nenghai Yu. 2024c · 2024
Closest in time.
HypoTermQA: Hypothetical Terms Dataset for Benchmarking Hallucination Tendency of LLMs
Cem Uluoglakci and Tugba Taskaya Temizel. 2024 · 2024
Closest in time.
The EU Artificial Intelligence Act
European Union. 2024 · 2024
Closest in time.
The instruction hierarchy: Training llms to prioritize privileged instructions
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel. 2024 · 2024
Closest in time.
Transferable multimodal attack on vision-language pre-training models. In 2024 IEEE Symposium on Security and Privacy (SP) . IEEE Computer Society, 102–102
Haodi Wang, Kai Dong, Zhilei Zhu, Haotong Qin, Aishan Liu, Xiaolin Fang, Jiakai Wang, and Xianglong Liu. 2024a · 2024
Closest in time.
Model Surgery: Modulating LLM’s Behavior Via Simple Parameter Editing
Huanqian Wang, Yang Yue, Rui Lu, Jingxin Shi, Andrew Zhao, Shenzhi Wang, Shiji Song, and Gao Huang. 2024d · 2024
Closest in time.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al · 2024
Closest in time.
Detoxifying Large Language Models via Knowledge Editing
Mengru Wang, Ningyu Zhang, Ziwen Xu, Zekun Xi, Shumin Deng, Yunzhi Yao, Qishen Zhang, Linyi Yang, Jindong Wang, and Huajun Chen. 2024e · 2024
Closest in time.
Context Injection Attacks on Large Language Models
Cheng’an Wei, Kai Chen, Yue Zhao, Yujia Gong, Lu Xiang, and Shenchen Zhu. 2024 · 2024
Closest in time.
Detecting, explaining, and mitigating memorization in diffusion models. In The Twelfth International Conference on Learning Representations
Yuxin Wen, Yuchen Liu, Chen Chen, and Lingjuan Lyu. 2024 · 2024
Closest in time.
Hallucination benchmark in medical visual question answering
Jinge Wu, Yunsoo Kim, and Honghan Wu. 2024b · 2024
Closest in time.
AUTOHALLUSION: Automatic Generation of Hallucination Benchmarks for Vision-Language Models
Xiyang Wu, Tianrui Guan, Dianqi Li, Shuaiyi Huang, Xiaoyu Liu, Xijun Wang, Ruiqi Xian, Abhinav Shrivastava, Furong Huang, Jordan Lee Boyd-Graber, et al · 2024
Closest in time.
Certifiably Robust RAG against Retrieval Corruption
Chong Xiang, Tong Wu, Zexuan Zhong, David Wagner, Danqi Chen, and Prateek Mittal. 2024 · 2024
Closest in time.
Safedecoding: Defending against jailbreak attacks via safety-aware decoding
Zhangchen Xu, Fengqing Jiang, Luyao Niu, Jinyuan Jia, Bill Yuchen Lin, and Radha Poovendran. 2024a · 2024
Closest in time.
Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
Zhao Xu, Fan Liu, and Hao Liu. 2024b · 2024
Closest in time.
Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents
Wenkai Yang, Xiaohan Bi, Yankai Lin, Sishuo Chen, Jie Zhou, and Xu Sun. 2024a · 2024
Closest in time.
GuardT2I: Defending Text-to-Image Models from Adversarial Prompts
Yijun Yang, Ruiyuan Gao, Xiao Yang, Jianyuan Zhong, and Qiang Xu. 2024b · 2024
Closest in time.
PRSA: Prompt reverse stealing attacks against large language models
Yong Yang, Xuhong Zhang, Yi Jiang, Xi Chen, Haoyu Wang, Shouling Ji, and Zonghui Wang. 2024c · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024 · 2024
Closest in time.
Vlattack: Multimodal adversarial attacks on vision-language tasks via pre-trained models
Ziyi Yin, Muchao Ye, Tianrong Zhang, Tianyu Du, Jinguo Zhu, Han Liu, Jinghui Chen, Ting Wang, and Fenglong Ma. 2024 · 2024
Closest in time.
Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 12944–12953
Qifan Yu, Juncheng Li, Longhui Wei, Liang Pang, Wentao Ye, Bosheng Qin, Siliang Tang, Qi Tian, and Yueting Zhuang. 2024 · 2024
Closest in time.
Qingbin Zeng, Qinglong Yang, Shunan Dong, Heming Du, Liang Zheng, Fengli Xu, and Yong Li. 2024a · 2024
Closest in time.
Mitigating the Privacy Issues in Retrieval-Augmented Generation (RAG) via Pure Synthetic Data
Shenglai Zeng, Jiankun Zhang, Pengfei He, Jie Ren, Tianqi Zheng, Hanqing Lu, Han Xu, Hui Liu, Yue Xing, and Jiliang Tang. 2024b · 2024
Closest in time.
The good and the bad: Exploring privacy issues in retrieval-augmented generation (rag)
Shenglai Zeng, Jiankun Zhang, Pengfei He, Yue Xing, Yiding Liu, Han Xu, Jie Ren, Shuaiqiang Wang, Dawei Yin, Yi Chang, et al · 2024
Closest in time.
Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents
Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang. 2024 · 2024
Closest in time.
Extracting Prompts by Inverting LLM Outputs
Collin Zhang, John X Morris, and Vitaly Shmatikov. 2024g · 2024
Closest in time.
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
Hanrong Zhang, Jingyuan Huang, Kai Mei, Yifei Yao, Zhenting Wang, Chenlu Zhan, Hongwei Wang, and Yongfeng Zhang. 2024d · 2024
Closest in time.
Kechi Zhang, Jia Li, Ge Li, Xianjie Shi, and Zhi Jin. 2024e · 2024
Closest in time.
Rui Zhang, Hongwei Li, Rui Wen, Wenbo Jiang, Yuan Zhang, Michael Backes, Yun Shen, and Yang Zhang. 2024f · 2024
Closest in time.
Privacyasst: Safeguarding user privacy in tool-using large language model agents
Xinyu Zhang, Huiyu Xu, Zhongjie Ba, Zhibo Wang, Yuan Hong, Jian Liu, Zhan Qin, and Kui Ren. 2024i · 2024
Closest in time.
Effective prompt extraction from language models
Yiming Zhang, Nicholas Carlini, and Daphne Ippolito. 2024b · 2024
Closest in time.
Yuxiang Zhang, Jing Chen, Junjie Wang, Yaxin Liu, Cheng Yang, Chufan Shi, Xinyu Zhu, Zihao Lin, Hanwen Wan, Yujiu Yang, et al · 2024
Closest in time.
A Survey on the Memory Mechanism of Large Language Model based Agents
Zeyu Zhang, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Quanyu Dai, Jieming Zhu, Zhenhua Dong, and Ji-Rong Wen. 2024a · 2024
Closest in time.
On large language models’ resilience to coercive interrogation. In 2024 IEEE Symposium on Security and Privacy (SP) . IEEE Computer Society, 252–252
Zhuo Zhang, Guangyu Shen, Guanhong Tao, Siyuan Cheng, and Xiangyu Zhang. 2024h · 2024
Closest in time.
Kening Zheng, Junkai Chen, Yibo Yan, Xin Zou, and Xuming Hu. 2024 · 2024
Closest in time.
PoisonedRAG: Knowledge Poisoning Attacks to Retrieval-Augmented Generation of Large Language Models
Wei Zou, Runpeng Geng, Binghui Wang, and Jinyuan Jia. 2024 · 2024
Closest in time.
Order-disorder: Imitation adversarial attacks for black-box neural ranking models. In Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security . 2025–2039
Jiawei Liu, Yangyang Kang, Di Tang, Kaisong Song, Changlong Sun, Xiaofeng Wang, Wei Lu, and Xiaozhong Liu. 2022a · 2039
Closest in time.