Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have significantly transformed the landscape of Natural Language Processing (NLP).
Real or fake? learning to discriminate machine from human generated text
Anton Bakhtin, Sam Gross, Myle Ott, Yuntian Deng, Marc’Aurelio Ranzato, and Arthur D. Szlam. 2019 · 1906
Earlier work this paper cites.
Release strategies and the social impacts of language models
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, and Jasmine Wang. 2019 · 1908
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M. Ziegler, Nisan Stiennon, Jeff Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 1909
Earlier work this paper cites.
Attacking neural text detectors
Max Wolff. 2020 · 2002
Earlier work this paper cites.
Just how toxic is data poisoning? a unified benchmark for backdoor and data poisoning attacks
Avi Schwarzschild, Micah Goldblum, Arjun Gupta, John P. Dickerson, and Tom Goldstein. 2020 · 2006
Earlier work this paper cites.
Learning from examples to improve code completion systems
Marcel Bruch, Martin Monperrus, and Mira Mezini. 2009 · 2009
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020 · 2009
Earlier work this paper cites.
Onion: A simple and effective defense against textual backdoor attacks
Fanchao Qi, Yangyi Chen, Mukai Li, Yuan Yao, Zhiyuan Liu, and Maosong Sun. 2020 · 2011
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Yinzhi Cao and Junfeng Yang. 2015 · 2015
Earlier work this paper cites.
" why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Automatically learning semantic features for defect prediction
Song Wang, Taiyue Liu, and Lin Tan. 2016 · 2016
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Xiaodong Song. 2017 · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
code2seq: Generating sequences from structured representations of code
Uri Alon, Shaked Brody, Omer Levy, and Eran Yahav. 2018 · 2018
Earlier work this paper cites.
Fine-pruning: Defending against backdooring attacks on deep neural networks
Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg. 2018 · 2018
Earlier work this paper cites.
Gltr: Statistical detection and visualization of generated text
Sebastian Gehrmann, Hendrik Strobelt, and Alexander M. Rush. 2019 · 2019
Earlier work this paper cites.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Earlier work this paper cites.
Program Synthesis and Semantic Parsing with Learned Code Idioms . Curran Associates Inc., Red Hook, NY, USA
Richard Shin, Miltiadis Allamanis, Marc Brockschmidt, and Oleksandr Polozov. 2019 · 2019
Earlier work this paper cites.
Realm: Retrieval-augmented language model pre-training
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2020
Earlier work this paper cites.
Bert-based sentiment analysis: A software engineering perspective
Himanshu Batra, Narinder Singh Punn, Sanjay Kumar Sonbhadra, and Sonali Agarwal. 2021 · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021 · 2021
Earlier work this paper cites.
Badnl: Backdoor attacks against nlp models with semantic-preserving improvements
Xiaoyi Chen, Ahmed Salem, Dingfan Chen, Michael Backes, Shiqing Ma, Qingni Shen, Zhonghai Wu, and Yang Zhang. 2021 · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Neural attention distillation: Erasing backdoor triggers from deep neural networks
Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma. 2021 · 2021
Earlier work this paper cites.
A survey of transformers
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu. 2021 · 2021
Earlier work this paper cites.
P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks
Xiao Liu, Kaixuan Ji, Yicheng Fu, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2021 · 2021
Earlier work this paper cites.
Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D Manning. 2021 · 2021
Earlier work this paper cites.
Asleep at the keyboard? assessing the security of github copilot’s code contributions
Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri. 2021 · 2021
Earlier work this paper cites.
Hidden killer: Invisible textual backdoor attacks with syntactic trigger
Fanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang, Zhiyuan Liu, Yasheng Wang, and Maosong Sun. 2021 · 2021
Earlier work this paper cites.
Concealed data poisoning attacks on NLP models
Eric Wallace, Tony Zhao, Shi Feng, and Sameer Singh. 2021 · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Laura Weidinger, John F. J. Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, Zachary Kenton, Sande Minnich Brown, William T. Hawkins, Tom Stepleton, Courtney Biles, Abeba Birhane, Julia Haas, Laura Rimell, Lisa Anne Hendricks, William S. Isaac, Sean Legassick, Geoffrey Irving, and Iason Gabriel. 2021 · 2021
Earlier work this paper cites.
Performance evaluation of adversarial attacks: Discrepancies and solutions
Jing Wu, Mingyi Zhou, Ce Zhu, Yipeng Liu, Mehrtash Harandi, and Li Li. 2021 · 2021
Earlier work this paper cites.
Wenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu Sun, and Bin He. 2021 · 2021
Earlier work this paper cites.
Bert-coqac: Bert-based conversational question answering in context
Munazza Zaib, Dai Hoang Tran, Subhash Sagar, Adnan Mahmood, Wei Emma Zhang, and Quan Z. Sheng. 2021 · 2021
Earlier work this paper cites.
Differentiable prompt makes pre-trained language models better few-shot learners
Ningyu Zhang, Luoqiu Li, Xiang Chen, Shumin Deng, Zhen Bi, Chuanqi Tan, Fei Huang, and Huajun Chen. 2021 · 2021
Earlier work this paper cites.
Badprompt: Backdoor attacks on continuous prompts
Xiangrui Cai, Haidong Xu, Sihan Xu, Ying Zhang, and Xiaojie Yuan. 2022 · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, John Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Benjamin Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zachary Dodds, T. J. Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom B. Brown, Nicholas Joseph, Sam McCandlish, Christopher Olah, Jared Kaplan, and Jack Clark. 2022 · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield-Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan. 2022 · 2022
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, Yuhuai Wu, Behnam Neyshabur, Guy Gur-Ari, and Vedant Misra. 2022 · 2022
Earlier work this paper cites.
Memory-based model editing at scale
Eric Mitchell, Charles Lin, Antoine Bosselut, Christopher D Manning, and Chelsea Finn. 2022 · 2022
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe. 2022 · 2022
Earlier work this paper cites.
Examining zero-shot vulnerability repair with large language models
Hammond Pearce, Benjamin Tan, Baleegh Ahmad, Ramesh Karri, and Brendan Dolan-Gavitt. 2022 · 2022
Earlier work this paper cites.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro. 2022 · 2022
Earlier work this paper cites.
Lost at c: A user study on the security implications of large language model code assistants
Gustavo Sandoval, Hammond A. Pearce, Teo Nys, Ramesh Karri, Siddharth Garg, and Brendan Dolan-Gavitt. 2022 · 2022
Earlier work this paper cites.
Damith Chamalke Senadeera and Julia Ive. 2022 · 2022
Earlier work this paper cites.
An empirical study of code smells in transformer-based code generation techniques
Mohammed Latif Siddiq, Shafayat H. Majumder, Maisha R. Mim, Sourov Jajodia, and Joanna C. S. Santos. 2022 · 2022
Earlier work this paper cites.
Ai bot chatgpt writes smart essays-should academics worry?
Chris Stokel-Walker. 2022 · 2022
Earlier work this paper cites.
Evaluating the factual consistency of large language models through summarization
Derek Tam, Anisha Mascarenhas, Shiyue Zhang, Sarah Kwan, Mohit Bansal, and Colin Raffel. 2022 · 2022
Earlier work this paper cites.
Memorization without overfitting: Analyzing the training dynamics of large language models
Kushal Tirumala, Aram Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. 2022 · 2022
Earlier work this paper cites.
Chain of thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed H. Chi, Quoc V Le, and Denny Zhou. 2022 · 2022
Earlier work this paper cites.
STar: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah Goodman. 2022 · 2022
Cited alongside, same era.
Moderate-fitting as a natural backdoor defender for pre-trained language models
Biru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen, Weilin Zhao, Chong Fu, Yangdong Deng, Zhiyuan Liu, Jingang Wang, Wei Wu, Maosong Sun, and Ming Gu. 2022 · 2022
Cited alongside, same era.
Openai. ai text classifier
2023 · 2023
Cited alongside, same era.
Zerogpt: Ai text detector
2023 · 2023
Cited alongside, same era.
How to catch an AI liar: Lie detection in black-box LLMs by asking unrelated questions
Anonymous. 2023 · 2023
Cited alongside, same era.
Towards a robust detection of language model generated text: Is chatgpt that easy to detect?
Wissam Antoun, Virginie Mouilleron, Benoît Sagot, and Djamé Seddah. 2023 · 2023
Controlling the extraction of memorized data from large language models via prompt-tuning
Mustafa Safa Ozdayi, Charith Peris, Jack FitzGerald, Christophe Dupuy, Jimit Majmudar, Haidar Khan, Rahil Parikh, and Rahul Gupta. 2023 · 2023
Later among the works it cites.
On the risk of misinformation pollution with large language models
Yikang Pan, Liangming Pan, Wenhu Chen, Preslav Nakov, Min-Yen Kan, and William Wang. 2023b · 2023
Later among the works it cites.
To chatgpt, or not to chatgpt: That is the question!
Alessandro Pegoraro, Kavita Kumari, Hossein Fereidooni, and Ahmad-Reza Sadeghi. 2023 · 2023
Later among the works it cites.
Are you copying my model? protecting the copyright of large language models for eaas via backdoor watermark
Wenjun Peng, Jingwei Yi, Fangzhao Wu, Shangxi Wu, Bin Benjamin Zhu, Lingjuan Lyu, Binxing Jiao, Tong Xu, Guangzhong Sun, and Xing Xie. 2023 · 2023
Later among the works it cites.
Adding instructions during pretraining: Effective way of controlling toxicity in language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Is github’s copilot as bad as humans at introducing vulnerabilities in code?
Owura Asare, Meiyappan Nagappan, and N. Asokan. 2023 · 2023
Cited alongside, same era.
Red-teaming large language models using chain of utterances for safety-alignment
Rishabh Bhardwaj and Soujanya Poria. 2023 · 2023
Cited alongside, same era.
Investigating answerability of llms for long-form question answering
Meghana Moorthy Bhat, Rui Meng, Ye Liu, Yingbo Zhou, and Semih Yavuz. 2023 · 2023
Cited alongside, same era.
Purple Llama CyberSecEval: A secure coding benchmark for language models
Manish Bhatt, Sahana Chennabasappa, Cyrus Nikolaidis, Shengye Wan, Ivan Evtimov, Dominik Gabi, Daniel Song, Faizan Ahmad, Cornelius Aschermann, Lorenzo Fontana, Sasha Frolov, Ravi Prakash Giri, Dhaval Kapil, Yiannis Kozyrakis, David LeBlanc, James Milazzo, Aleksandar Straumann, Gabriel Synnaeve, Varun Vontimitta, Spencer Whitman, and Joshua Saxe. 2023 · 2023
Cited alongside, same era.
What can we learn from data leakage and unlearning for law?
Jaydeep Borkar. 2023 · 2023
Cited alongside, same era.
Explore, establish, exploit: Red teaming language models from scratch
Stephen Casper, Jason Lin, Joe Kwon, Gatlen Culp, and Dylan Hadfield-Menell. 2023 · 2023
Cited alongside, same era.
Shrimai Prabhumoye, Mostofa Patwary, Mohammad Shoeybi, and Bryan Catanzaro. 2023 · 2023
Later among the works it cites.
Beyond black box ai-generated plagiarism detection: From sentence to document level
Mujahid Ali Quidwai, Chun Xing Li, and Parijat Dube. 2023 · 2023
Later among the works it cites.
Tricking llms into disobedience: Understanding, analyzing, and preventing jailbreaks
Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya, and Monojit Choudhury. 2023 · 2023
Later among the works it cites.
Generating phishing attacks using chatgpt
Sayak Saha Roy, Krishna Vamsi Naragam, and Shirin Nilizadeh. 2023 · 2023
Later among the works it cites.
Can ai-generated text be reliably detected?
Vinu Sankar Sadasivan, Aounon Kumar, S. Balasubramanian, Wenxiao Wang, and Soheil Feizi. 2023 · 2023
Later among the works it cites.
Analysis of chatgpt on source code
Ahmed R. Sadik, Antonello Ceravola, Frank Joublin, and Jibesh Patra. 2023 · 2023
Later among the works it cites.
Hanyin Shao, Jie Huang, Shen Zheng, and Kevin Chen-Chuan Chang. 2023 · 2023
Later among the works it cites.
Survey of vulnerabilities in large language models revealed by adversarial attacks
Erfan Shayegani, Md. Abdullah Al Mamun, Yu Fu, Pedram Zaree, Yue Dong, and Nael B. Abu-Ghazaleh. 2023 · 2023
Later among the works it cites.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. 2023 · 2023
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023 · 2023
Later among the works it cites.
Mondrian: Prompt abstraction attack against large language models for cheaper api pricing
Wai Man Si, Michael Backes, and Yang Zhang. 2023 · 2023
Later among the works it cites.
Seeing seeds beyond weeds: Green teaming generative ai for beneficial uses
Logan Stapleton, Jordan Taylor, Sarah Fox, Tongshuang Wu, and Haiyi Zhu. 2023 · 2023
Later among the works it cites.
Detectllm: Leveraging log rank information for zero-shot detection of machine-generated text
Jinyan Su, Terry Yue Zhuo, Di Wang, and Preslav Nakov. 2023 · 2023
Later among the works it cites.
Safety assessment of chinese large language models
Hao Sun, Zhexin Zhang, Jiawen Deng, Jiale Cheng, and Minlie Huang. 2023 · 2023
Later among the works it cites.
Evaluating the factual consistency of large language models through news summarization
Derek Tam, Anisha Mascarenhas, Shiyue Zhang, Sarah Kwan, Mohit Bansal, and Colin Raffel. 2023 · 2023
Later among the works it cites.
Stanford alpaca: an instruction-following llama model (2023)
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023 · 2023
Later among the works it cites.
Zeroleak: Using llms for scalable and cost effective side-channel patching
M. Caner Tol and Berk Sunar. 2023 · 2023
Later among the works it cites.
Christoforos Vasilatos, Manaar Alam, Talal Rahwan, Yasir Zaki, and Michail Maniatakos. 2023 · 2023
Later among the works it cites.
Disinformation capabilities of large language models
Ivan Vykopal, Mat’uvs Pikuliak, Ivan Srba, Róbert Móro, Dominik Macko, and Mária Bieliková. 2023 · 2023
Later among the works it cites.
Poisoning language models during instruction tuning
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023 · 2023
Later among the works it cites.
Testing of detection tools for ai-generated text
Debora Weber-Wulff, Alla Anohina-Naumeca, Sonja Bjelobaba, Tom’aš Foltýnek, Jean Gabriel Guerrero-Dib, Olumide Popoola, Petr Sigut, and Lorna Waddington. 2023 · 2023
Later among the works it cites.
Unveiling the implicit toxicity in large language models
Jiaxin Wen, Pei Ke, Hao Sun, Zhexin Zhang, Chengfei Li, Jinfeng Bai, and Minlie Huang. 2023 · 2023
Later among the works it cites.
An evaluation on large language model outputs: Discourse and memorization
Adrian de Wynter, Xun Wang, Alex Sokolov, Qilong Gu, and Si-Qing Chen. 2023 · 2023
Later among the works it cites.
Gpt paternity test: Gpt generated text detection with gpt genetic inheritance
Xiao Yu, Yuang Qi, Kejiang Chen, Guoqiang Chen, Xi Yang, Pengyuan Zhu, Weiming Zhang, and Neng H. Yu. 2023 · 2023
Later among the works it cites.
G3detector: General gpt-generated text detector
Haolan Zhan, Xuanli He, Qiongkai Xu, Yuxiang Wu, and Pontus Stenetorp. 2023 · 2023
Later among the works it cites.
Memorybank: Enhancing large language models with long-term memory
Wanjun Zhong, Lianghong Guo, Qi-Fei Gao, He Ye, and Yanlin Wang. 2023 · 2023
Later among the works it cites.
Red teaming chatgpt via jailbreaking: Bias, robustness, reliability and toxicity
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing. 2023 · 2023
Later among the works it cites.
To code, or not to code? exploring impact of code in pre-training
Viraat Aryabumi, Yixuan Su, Raymond Ma, Adrien Morisot, Ivan Zhang, Acyr F. Locatelli, Marzieh Fadaee, A. Ustun, and Sara Hooker. 2024 · 2024
Closest in time.
Unfair alignment: Examining safety alignment across vision encoder layers in vision-language models
Saketh Bachu, Erfan Shayegani, Trishna Chakraborty, Rohit Lal, Arindam Dutta, Chengyu Song, Yue Dong, Nael Abu-Ghazaleh, and Amit K. Roy-Chowdhury. 2024 · 2024
Closest in time.
Can textual unlearning solve cross-modality safety alignment?
Trishna Chakraborty, Erfan Shayegani, Zikui Cai, Nael B. Abu-Ghazaleh, M. Salman Asif, Yue Dong, Amit Roy-Chowdhury, and Chengyu Song. 2024 · 2024
Closest in time.
Agentdojo: A dynamic environment to evaluate attacks and defenses for llm agents
Edoardo Debenedetti, Jie Zhang, Mislav Balunovi’c, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr. 2024 · 2024
Closest in time.
Refusal-trained llms are easily jailbroken as browser agents
Priyanshu Kumar, Elaine Lau, Saranya Vijayakumar, Tu Trinh, Scale Red Team, Elaine Chang, Vaughn Robinson, Sean Hendryx, Shuyan Zhou, Matt Fredrikson, Summer Yue, and Zifan Wang. 2024 · 2024
Closest in time.
Eia: Environmental injection attack on generalist web agents for privacy leakage
Zeyi Liao, Lingbo Mo, Chejian Xu, Mintong Kang, Jiawei Zhang, Chaowei Xiao, Yuan Tian, Bo Li, and Huan Sun. 2024 · 2024
Closest in time.
Great, now write an article about that: The crescendo multi-turn llm jailbreak attack
Mark Russinovich, Ahmed Salem, and Ronen Eldan. 2024 · 2024
Closest in time.
Rainbow teaming: Open-ended generation of diverse adversarial prompts
Mikayel Samvelyan, Sharath Chandra Raparthy, Andrei Lupu, Eric Hambro, Aram H. Markosyan, Manish Bhatt, Yuning Mao, Minqi Jiang, Jack Parker-Holder, Jakob Foerster, Tim Rocktäschel, and Roberta Raileanu. 2024 · 2024
Closest in time.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Erfan Shayegani, Yue Dong, and Nael Abu-Ghazaleh. 2024 · 2024
Closest in time.
Calibration and correctness of language models for code
Claudio Spiess, David Gros, Kunal Suresh Pai, Michael Pradel, Md Rafiqul Islam Rabin, Susmit Jha, Prem Devanbu, and Toufique Ahmed. 2024 · 2024
Closest in time.
Localizing paragraph memorization in language models
Niklas Stoehr, Mitchell Gordon, Chiyuan Zhang, and Owen Lewis. 2024 · 2024
Closest in time.
Ke Yang, Jiateng Liu, John Wu, Chaoqi Yang, Y. Fung, Sha Li, Zixuan Huang, Xu Cao, Xingyao Wang, Yiquan Wang, Heng Ji, and Chengxiang Zhai. 2024 · 2024
Closest in time.
Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents
Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang. 2024 · 2024
Closest in time.
Emergent misalignment: Narrow finetuning can produce broadly misaligned llms
Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martín Soto, Nathan Labenz, and Owain Evans. 2025 · 2025
Closest in time.
Llm-powered gui agents in phone automation: Surveying progress and prospects
Guangyi Liu, Pengxiang Zhao, Liang Liu, Yaxuan Guo, Han Xiao, Weifeng Lin, Yuxiang Chai, Yue Han, Shuai Ren, Hao Wang, Xiaoyu Liang, Wenhao Wang, Tianze Wu, Linghao Li, Hao Wang, Guanjing Xiong, Yong Liu, and Hongsheng Li. 2025 · 2025
Closest in time.
Gui-r1 : A generalist r1-style vision-language action model for gui agents
Run Luo, Lu Wang, Wanwei He, and Xiaobo Xia. 2025 · 2025
Closest in time.
Erfan Shayegani, G M Shahariar, Sara Abdali, Lei Yu, Nael Abu-Ghazaleh, and Yue Dong. 2025 · 2025
Closest in time.
Zhaofeng Wu, Xinyan Velocity Yu, Dani Yogatama, Jiasen Lu, and Yoon Kim. 2025 · 2025
Closest in time.
BlueSuffix: Reinforced blue teaming for vision-language models against jailbreak attacks
Yunhan Zhao, Xiang Zheng, Lin Luo, Yige Li, Xingjun Ma, and Yu-Gang Jiang. 2025 · 2025
Closest in time.
Are large pre-trained language models leaking your personal information?
Jie Huang, Hanyin Shao, and Kevin Chen-Chuan Chang. 2022 · 2047
Closest in time.