Fetching the paper…
Reading the bibliography…
In recent years, generative AI (GenAI), like large language models and text-to-image models, has received significant attention across various domains.
Intellectual property protection of dnn models
Sen Peng, Yufei Chen, Jie Xu, Zizhuo Chen, Cong Wang, and Xiaohua Jia · 1911
Earlier work this paper cites.
Membership inference attacks from first principles
Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, and Florian Tramer · 1914
Earlier work this paper cites.
Ethical issues in status planning
Juan Cobarrubias et al · 1983
Earlier work this paper cites.
Secure spread spectrum watermarking for images, audio and video
Ingemar J Cox, Joe Kilian, Tom Leighton, and Talal Shamoon · 1996
Earlier work this paper cites.
Watermarking digital images for copyright protection
JJK ó Ruanaidh, WJ Dowling, and FM Boland · 1996
Earlier work this paper cites.
Natural language processing for information assurance and security: an overview and implementations
Mikhail J Atallah, Craig J McDonough, Victor Raskin, and Sergei Nirenburg · 2001
Earlier work this paper cites.
Natural language watermarking and tamperproofing
Mikhail J Atallah, Victor Raskin, Christian F Hempelmann, Mercan Karahan, Radu Sion, Umut Topkara, and Katrina E Triezenberg · 2002
Earlier work this paper cites.
Noun-verb based technique of text watermarking using recursive decent semantic net parsers
Xingming Sun and Alex Jessey Asiimwe · 2005
Earlier work this paper cites.
Digital camera identification from sensor pattern noise
Jan Lukas, Jessica Fridrich, and Miroslav Goljan · 2006
Earlier work this paper cites.
The hiding virtues of ambiguity: quantifiably resilient watermarking of natural language text through synonym substitutions
Umut Topkara, Mercan Topkara, and Mikhail J Atallah · 2006
Earlier work this paper cites.
Human-level artificial general intelligence and the possibility of a technological singularity: A reaction to ray kurzweil’s the singularity is near, and mcdermott’s critique of kurzweil
Ben Goertzel · 2007
Earlier work this paper cites.
Natural language watermarking via morphosyntactic alterations
Hasan Mesut Meral, Bülent Sankur, A Sumru Özsoy, Tunga Güngör, and Emre Sevinç · 2009
Earlier work this paper cites.
A new approach of the cryptographic attacks
Otilia Cangea and Gabriela Moise · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2013
Earlier work this paper cites.
Privacy in pharmacogenetics: An { \{ End-to-End } \} case study of personalized warfarin dosing
Matthew Fredrikson, Eric Lantz, Somesh Jha, Simon Lin, David Page, and Thomas Ristenpart · 2014
Earlier work this paper cites.
Model inversion attacks that exploit confidence information and basic countermeasures
Matt Fredrikson, Somesh Jha, and Thomas Ristenpart · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli · 2015
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Earlier work this paper cites.
Deepfool: a simple and accurate method to fool deep neural networks
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard · 2016
Earlier work this paper cites.
Transferability in machine learning: from phenomena to black-box attacks using adversarial samples
Nicolas Papernot, Patrick McDaniel, and Ian Goodfellow · 2016
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Digital watermarking techniques for security applications
Sonam Tyagi, Harsh Vikram Singh, Raghav Agarwal, and Sandeep Kumar Gangwar · 2016
Earlier work this paper cites.
Replacement attack: A new zero text watermarking attack
Morteza Bashardoost, Mohd Shafry Mohd Rahim, Tanzila Saba, and Amjad Rehman · 2017
Earlier work this paper cites.
Tom B Brown, Dandelion Mané, Aurko Roy, Martín Abadi, and Justin Gilmer · 2017
Earlier work this paper cites.
Towards evaluating the robustness of neural networks
Nicholas Carlini and David Wagner · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Generating steganographic text with lstms
Tina Fang, Martin Jaggi, and Katerina Argyraki · 2017
Earlier work this paper cites.
On the (statistical) detection of adversarial examples
Kathrin Grosse, Praveen Manoharan, Nicolas Papernot, Michael Backes, and Patrick McDaniel · 2017
Earlier work this paper cites.
Generating steganographic images via adversarial training
Jamie Hayes and George Danezis · 2017
Earlier work this paper cites.
Logan: Membership inference attacks against generative models
Jamie Hayes, Luca Melis, George Danezis, and Emiliano De Cristofaro · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Earlier work this paper cites.
Neural trojans
Yuntao Liu, Yang Xie, and Ankur Srivastava · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov · 2017
Earlier work this paper cites.
Embedding watermarks into deep neural networks
Yusuke Uchida, Yuki Nagai, Shigeyuki Sakazawa, and Shin’ichi Satoh · 2017
Earlier work this paper cites.
Attacks on digital watermarks: classification, implications, benchmarks
Yukti Varshney · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A zero-watermarking scheme for prose writings
Meng Yingjie, Liu Huiran, Shang Tong, and Teng Xiaoyu · 2017
Earlier work this paper cites.
Men also like shopping: Reducing gender bias amplification using corpus-level constraints
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2017
Earlier work this paper cites.
Two-stream neural networks for tampered face detection
Peng Zhou, Xintong Han, Vlad I Morariu, and Larry S Davis · 2017
Earlier work this paper cites.
Turning your weakness into a strength: Watermarking deep neural networks by backdooring
Yossi Adi, Carsten Baum, Moustapha Cisse, Benny Pinkas, and Joseph Keshet · 2018
Earlier work this paper cites.
Mesonet: a compact facial video forgery detection network
Darius Afchar, Vincent Nozick, Junichi Yamagishi, and Isao Echizen · 2018
Earlier work this paper cites.
Adversarial attacks and defences: A survey
Anirban Chakraborty, Manaar Alam, Vishal Dey, Anupam Chattopadhyay, and Debdeep Mukhopadhyay · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Boosting adversarial attacks with momentum
Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li · 2018
Earlier work this paper cites.
Robust physical-world attacks on deep learning visual classification
Kevin Eykholt, Ivan Evtimov, Earlence Fernandes, Bo Li, Amir Rahmati, Chaowei Xiao, Atul Prakash, Tadayoshi Kohno, and Dawn Song · 2018
Earlier work this paper cites.
Deepfake video detection using recurrent neural networks
David Güera and Edward J Delp · 2018
Earlier work this paper cites.
Exposing deepfake videos by detecting face warping artifacts
Yuezun Li and Siwei Lyu · 2018
Earlier work this paper cites.
In ictu oculi: Exposing ai created fake videos by detecting eye blinking
Yuezun Li, Ming-Ching Chang, and Siwei Lyu · 2018
Earlier work this paper cites.
Detection of adversarial training examples in poisoning attacks through anomaly detection
Andrea Paudice, Luis Muñoz-González, Andras Gyorgy, and Emil C Lupu · 2018
Earlier work this paper cites.
Moral reasoning, 2018
H. S. Richardson · 2018
Earlier work this paper cites.
Ahmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang, Mario Fritz, and Michael Backes · 2018
Earlier work this paper cites.
Clean-label backdoor attacks
Alexander Turner, Dimitris Tsipras, and Aleksander Madry · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
Samuel Yeom, Irene Giacomelli, Matt Fredrikson, and Somesh Jha · 2018
Earlier work this paper cites.
Protecting intellectual property of deep neural networks with watermarking
Jialong Zhang, Zhongshu Gu, Jiyong Jang, Hui Wu, Marc Ph Stoecklin, Heqing Huang, and Ian Molloy · 2018
Earlier work this paper cites.
Hidden: Hiding data with deep networks
Jiren Zhu, Russell Kaplan, Justin Johnson, and Li Fei-Fei · 2018
Earlier work this paper cites.
Deepfake video detection through optical flow based cnn
Irene Amerini, Leonardo Galteri, Roberto Caldelli, and Alberto Del Bimbo · 2019
Earlier work this paper cites.
Evaluating the underlying gender bias in contextualized word embeddings
Christine Basta, Marta R Costa-Jussà, and Noe Casas · 2019
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song · 2019
Earlier work this paper cites.
Deepinspect: A black-box trojan detection and mitigation framework for deep neural networks
Huili Chen, Cheng Fu, Jishen Zhao, and Farinaz Koushanfar · 2019
Earlier work this paper cites.
Noiseprint: A cnn-based camera model fingerprint
Davide Cozzolino and Luisa Verdoliva · 2019
Earlier work this paper cites.
Sparse and imperceivable adversarial attacks
Francesco Croce and Matthias Hein · 2019
Earlier work this paper cites.
Towards near-imperceptible steganographic text
Falcon Z Dai and Zheng Cai · 2019
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu · 2019
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston · 2019
Earlier work this paper cites.
Gltr: Statistical detection and visualization of generated text
Sebastian Gehrmann, Hendrik Strobelt, and Alexander M Rush · 2019
Earlier work this paper cites.
Assessing the factual accuracy of generated text
Ben Goodrich, Vinay Rao, Peter J Liu, and Mohammad Saleh · 2019
Earlier work this paper cites.
Badnets: Evaluating backdooring attacks on deep neural networks
Tianyu Gu, Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg · 2019
Earlier work this paper cites.
Tabor: A highly accurate approach to inspecting and restoring trojan backdoors in ai systems
Wenbo Guo, Lun Wang, Xinyu Xing, Min Du, and Dawn Song · 2019
Earlier work this paper cites.
Monte carlo and reconstruction membership inference attacks against generative models
Benjamin Hilprecht, Martin Härterich, and Daniel Bernau · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2019
Earlier work this paper cites.
Automatic detection of generated text is easiest when humans are fooled
Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck · 2019
Earlier work this paper cites.
Semantic adversarial attacks: Parametric transformations that fool deep classifiers
Ameya Joshi, Amitangshu Mukherjee, Soumik Sarkar, and Chinmay Hegde · 2019
Earlier work this paper cites.
How to prove your model belongs to you: A blind-watermark based framework to protect intellectual property of dnn
Zheng Li, Chengyu Hu, Yang Zhang, and Shanqing Guo · 2019
Earlier work this paper cites.
Do gans leave artificial fingerprints?
Francesco Marra, Diego Gragnaniello, Luisa Verdoliva, and Giovanni Poggi · 2019
Earlier work this paper cites.
Sparsefool: a few pixels make a big difference
Apostolos Modas, Seyed-Mohsen Moosavi-Dezfooli, and Pascal Frossard · 2019
Earlier work this paper cites.
Comprehensive privacy analysis of deep learning: Passive and active white-box inference attacks against centralized and federated learning
Milad Nasr, Reza Shokri, and Amir Houmansadr · 2019
Earlier work this paper cites.
Detecting gan generated fake images using co-occurrence matrices
Lakshmanan Nataraj, Tajuddin Manhar Mohammed, Shivkumar Chandrasekaran, Arjuna Flenner, Jawadul H Bappy, Amit K Roy-Chowdhury, and BS Manjunath · 2019
Earlier work this paper cites.
Capsule-forensics: Using capsule networks to detect forged images and videos
Huy H Nguyen, Junichi Yamagishi, and Isao Echizen · 2019
Earlier work this paper cites.
Label sanitization against label flipping poisoning attacks
Andrea Paudice, Luis Muñoz-González, and Emil C Lupu · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher · 2019
Earlier work this paper cites.
Faceforensics++: Learning to detect manipulated facial images
Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner · 2019
Earlier work this paper cites.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A Smith, and Yejin Choi · 2019
Earlier work this paper cites.
Generalization in generation: A closer look at exposure bias
Florian Schmidt · 2019
Earlier work this paper cites.
Release strategies and the social impacts of language models
Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al · 2019
Earlier work this paper cites.
On nmt search errors and model errors: Cat got your tongue?
Felix Stahlberg and Bill Byrne · 2019
Earlier work this paper cites.
Neural cleanse: Identifying and mitigating backdoor attacks in neural networks
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y Zhao · 2019
Earlier work this paper cites.
Deepfake bot submissions to federal public comment websites cannot be distinguished from human submissions
Max Weiss · 2019
Earlier work this paper cites.
Attributing fake images to gans: Learning and analyzing gan fingerprints
Ning Yu, Larry S Davis, and Mario Fritz · 2019
Earlier work this paper cites.
Generating synthetic data in finance: opportunities, challenges and pitfalls
Samuel A Assefa, Danial Dervovic, Mahmoud Mahfouz, Robert E Tillman, Prashant Reddy, and Manuela Veloso · 2020
Earlier work this paper cites.
Generating fact checking explanations
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein · 2020
Earlier work this paper cites.
Influence functions in deep learning are fragile
Samyadeep Basu, Philip Pope, and Soheil Feizi · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Gan-leaks: A taxonomy of membership inference attacks against generative models
Dingfan Chen, Ning Yu, Yang Zhang, and Mario Fritz · 2020
Earlier work this paper cites.
Fakecatcher: Detection of synthetic portrait videos using biological signals
Umur Aybars Ciftci, Ilke Demir, and Lijun Yin · 2020
Earlier work this paper cites.
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
Francesco Croce and Matthias Hein · 2020
Earlier work this paper cites.
Backdooring convolutional neural networks via targeted weight perturbations
Jacob Dumford and Walter Scheirer · 2020
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
Vitaly Feldman and Chiyuan Zhang · 2020
Earlier work this paper cites.
Leveraging frequency analysis for deep fake image recognition
Joel Frank, Thorsten Eisenhofer, Lea Schönherr, Asja Fischer, Dorothea Kolossa, and Thorsten Holz · 2020
Earlier work this paper cites.
Scalable detection of offensive and non-compliant content/logo in product images
Shreyansh Gandhi, Samrat Kokkula, Abon Chaudhuri, Alessandro Magnani, Theban Stanley, Behzad Ahmadi, Venkatesh Kandaswamy, Omer Ovenc, and Shie Mannor · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith · 2020
Earlier work this paper cites.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2020
Earlier work this paper cites.
Semantic object accuracy for generative text-to-image synthesis
Tobias Hinz, Stefan Heinrich, and Stefan Wermter · 2020
Earlier work this paper cites.
Membership inference attacks on sequence-to-sequence models: Is my data in your machine translation system?
Sorami Hisamoto, Matt Post, and Kevin Duh · 2020
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
On the effectiveness of mitigating data poisoning attacks with gradient shaping
Sanghyun Hong, Varun Chandrasekaran, Yiğitcan Kaya, Tudor Dumitraş, and Nicolas Papernot · 2020
Earlier work this paper cites.
Universal physical camouflage attacks on object detectors
Lifeng Huang, Chengying Gao, Yuyin Zhou, Cihang Xie, Alan L Yuille, Changqing Zou, and Ning Liu · 2020
Earlier work this paper cites.
Social biases in nlp models as barriers for persons with disabilities
Ben Hutchinson, Vinodkumar Prabhakaran, Emily Denton, Kellie Webster, Yu Zhong, and Stephen Denuyl · 2020
Earlier work this paper cites.
Decentralized attribution of generative models
Changhoon Kim, Yi Ren, and Yezhou Yang · 2020
Earlier work this paper cites.
Deep partition aggregation: Provable defense against general poisoning attacks
Alexander Levine and Soheil Feizi · 2020
Earlier work this paper cites.
Face x-ray for more general face forgery detection
Lingzhi Li, Jianmin Bao, Ting Zhang, Hao Yang, Dong Chen, Fang Wen, and Baining Guo · 2020
Earlier work this paper cites.
Transformers are better than humans at identifying generated text
Antonis Maronikolakis, Mark Stevenson, and Hinrich Schütze · 2020
Earlier work this paper cites.
Two-branch recurrent network for isolating deepfakes in videos
Iacopo Masi, Aditya Killekar, Royston Marian Mascarenhas, Shenoy Pratik Gurudatt, and Wael AbdAlmageed · 2020
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald · 2020
Earlier work this paper cites.
The radicalization risks of gpt-3 and advanced neural language models
Kris McGuffie and Alex Newhouse · 2020
Earlier work this paper cites.
Deepfakes detection with automatic face weighting
Daniel Mas Montserrat, Hanxiang Hao, Sri K Yarlagadda, Sriram Baireddy, Ruiting Shao, János Horváth, Emily Bartusiak, Justin Yang, David Guera, Fengqing Zhu, et al · 2020
Earlier work this paper cites.
Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp
John X Morris, Eli Lifland, Jin Yong Yoo, Jake Grigsby, Di Jin, and Yanjun Qi · 2020
Earlier work this paper cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2020
Earlier work this paper cites.
Data augmentation of high frequency financial data using generative adversarial network
Yusuke Naritomi and Takanori Adachi · 2020
Earlier work this paper cites.
Input-aware dynamic backdoor attack
Tuan Anh Nguyen and Anh Tran · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Tbt: Targeted neural network attack with bit trojan
Adnan Siraj Rakin, Zhezhi He, and Deliang Fan · 2020
Earlier work this paper cites.
Disrupting deepfakes: Adversarial attacks against conditional image translation networks and facial manipulation systems
Nataniel Ruiz, Sarah Adel Bargal, and Stan Sclaroff · 2020
Earlier work this paper cites.
Hidden trigger backdoor attacks
Aniruddha Saha, Akshayvarun Subramanya, and Hamed Pirsiavash · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh · 2020
Earlier work this paper cites.
Information leakage in embedding models
Congzheng Song and Ananth Raghunathan · 2020
Earlier work this paper cites.
Systematic evaluation of backdoor data poisoning attacks on image classifiers
Loc Truong, Chace Jones, Brian Hutchinson, Andrew August, Brenda Praggastis, Robert Jasper, Nicole Nichols, and Aaron Tuor · 2020
Earlier work this paper cites.
Authorship attribution for neural text generation
Adaku Uchendu, Thai Le, Kai Shu, and Dongwon Lee · 2020
Earlier work this paper cites.
Concealed data poisoning attacks on nlp models
Eric Wallace, Tony Z Zhao, Shi Feng, and Sameer Singh · 2020
Earlier work this paper cites.
On exposure bias, hallucination and domain shift in neural machine translation
Chaojun Wang and Rico Sennrich · 2020
Earlier work this paper cites.
Practical detection of trojan neural networks: Data-limited and data-free cases
Ren Wang, Gaoyuan Zhang, Sijia Liu, Pin-Yu Chen, Jinjun Xiong, and Meng Wang · 2020
Earlier work this paper cites.
Disrupting image-translation-based deepfake algorithms with adversarial attacks
Chin-Yuan Yeh, Hsi-Wen Chen, Shang-Lun Tsai, and Sheng-De Wang · 2020
Earlier work this paper cites.
Responsible disclosure of generative models using scalable fingerprinting
Ning Yu, Vladislav Skripniuk, Dingfan Chen, Larry Davis, and Mario Fritz · 2020
Earlier work this paper cites.
Analyzing information leakage of updates to natural language models
Santiago Zanella-Béguelin, Lukas Wutschitz, Shruti Tople, Victor Rühle, Andrew Paverd, Olga Ohrimenko, Boris Köpf, and Marc Brockschmidt · 2020
Earlier work this paper cites.
Adversarial attacks on deep-learning models in natural language processing: A survey
Wei Emma Zhang, Quan Z Sheng, Ahoud Alhazmi, and Chenliang Li · 2020
Earlier work this paper cites.
Clean-label backdoor attacks on video recognition models
Shihao Zhao, Xingjun Ma, Xiang Zheng, James Bailey, Jingjing Chen, and Yu-Gang Jiang · 2020
Earlier work this paper cites.
Detecting hallucinated content in conditional neural sequence generation
Chunting Zhou, Graham Neubig, Jiatao Gu, Mona Diab, Paco Guzman, Luke Zettlemoyer, and Marjan Ghazvininejad · 2020
Earlier work this paper cites.
Adversarial watermarking transformer: Towards tracing text provenance with data hiding
Sahar Abdelnabi and Mario Fritz · 2021
Earlier work this paper cites.
Large language models associate muslims with violence
Abubakar Abid, Maheen Farooqi, and James Zou · 2021
Earlier work this paper cites.
Molgpt: molecular generation using a transformer-decoder model
Viraj Bagal, Rishal Aggarwal, PK Vinod, and U Deva Priyakumar · 2021
Earlier work this paper cites.
On training sample memorization: Lessons from benchmarking generative modeling with a large-scale competition
Ching-Yuan Bai, Hsuan-Tien Lin, Colin Raffel, and Wendy Chi-wen Kan · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
Large image datasets: A pyrrhic win for computer vision?
Abeba Birhane and Vinay Uday Prabhu · 2021
Earlier work this paper cites.
Truth, lies, and automation
Ben Buchanan, Andrew Lohn, Micah Musser, and Katerina Sedova · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
Refit: a unified watermark removal framework for deep learning systems with limited data
Xinyun Chen, Wenxiao Wang, Chris Bender, Yiming Ding, Ruoxi Jia, Bo Li, and Dawn Song · 2021
Earlier work this paper cites.
Label-only membership inference attacks
Christopher A Choquette-Choo, Florian Tramer, Nicholas Carlini, and Nicolas Papernot · 2021
Earlier work this paper cites.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov · 2021
Earlier work this paper cites.
Lira: Learnable, imperceptible and robust backdoor attacks
Khoa Doan, Yingjie Lao, Weijie Zhao, and Ping Li · 2021
Earlier work this paper cites.
Generative adversarial networks in finance: an overview
Florian Eckerli and Joerg Osterrieder · 2021
Earlier work this paper cites.
When do gans replicate? on the choice of dataset size
Qianli Feng, Chenqi Guo, Fabian Benitez-Quiroz, and Aleix M Martinez · 2021
Cited alongside, same era.
Unsupervised and distributional detection of machine-generated text
Matthias Gallé, Jos Rozen, Germán Kruszewski, and Hady Elsahar · 2021
Cited alongside, same era.
Effective and efficient vote attack on capsule networks
Jindong Gu, Baoyuan Wu, and Volker Tresp · 2021
Cited alongside, same era.
Gradient-based adversarial attacks against text transformers
Chuan Guo, Alexandre Sablayrolles, Hervé Jégou, and Douwe Kiela · 2021
Cited alongside, same era.
Clipscore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi · 2021
Cited alongside, same era.
Membership inference of diffusion models
Hailong Hu and Jun Pang · 2023
Later among the works it cites.
Yue Huang and Lichao Sun · 2023
Later among the works it cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al · 2023
Later among the works it cites.
Preventing generation of verbatim memorization in language models gives a false sense of privacy
Daphne Ippolito, Florian Tramèr, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher A Choquette-Choo, and Nicholas Carlini · 2023
Later among the works it cites.
Llm platform security: Applying a systematic evaluation framework to openai’s chatgpt plugins
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exposing gan-generated faces using inconsistent corneal specular highlights
Shu Hu, Yuezun Li, and Siwei Lyu · 2021
Cited alongside, same era.
Initiative defense against facial manipulation
Qidong Huang, Jie Zhang, Wenbo Zhou, Weiming Zhang, and Nenghai Yu · 2021
Cited alongside, same era.
Membership inference attack susceptibility of clinical language models
Abhyuday Jagannatha, Bhanu Pratap Singh Rawat, and Hong Yu · 2021
Cited alongside, same era.
Intrinsic certified robustness of bagging against data poisoning attacks
Jinyuan Jia, Xiaoyu Cao, and Neil Zhenqiang Gong · 2021
Cited alongside, same era.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2021
Cited alongside, same era.
Does bert pretrained on clinical notes reveal sensitive data?
Eric Lehman, Sarthak Jain, Karl Pichotta, Yoav Goldberg, and Byron C Wallace · 2021
Cited alongside, same era.
Invisible backdoor attack with sample-specific triggers
Yuezun Li, Yiming Li, Baoyuan Wu, Longkang Li, Ran He, and Siwei Lyu · 2021
Cited alongside, same era.
Umar Iqbal, Tadayoshi Kohno, and Franziska Roesner · 2023
Later among the works it cites.
Leveraging generative ai models for synthetic data generation in healthcare: Balancing research and privacy
Aryan Jadon and Shashank Kumar · 2023
Later among the works it cites.
A systematic review of hate speech automatic detection using natural language processing
Md Saroar Jahan and Mourad Oussalah · 2023
Later among the works it cites.
Baseline defenses for adversarial attacks against aligned language models
Neel Jain, Avi Schwarzschild, Yuxin Wen, Gowthami Somepalli, John Kirchenbauer, Ping-yeh Chiang, Micah Goldblum, Aniruddha Saha, Jonas Geiping, and Tom Goldstein · 2023
Later among the works it cites.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Later among the works it cites.
Xiaojun Jia, Yuefeng Chen, Xiaofeng Mao, Ranjie Duan, Jindong Gu, Rong Zhang, Hui Xue, and Xiaochun Cao · 2023
Later among the works it cites.
Active retrieval augmented generation
Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig · 2023
Later among the works it cites.
Automatically auditing large language models via discrete optimization
Erik Jones, Anca Dragan, Aditi Raghunathan, and Jacob Steinhardt · 2023
Later among the works it cites.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
Daniel Kang, Xuechen Li, Ion Stoica, Carlos Guestrin, Matei Zaharia, and Tatsunori Hashimoto · 2023
Later among the works it cites.
Mitigating approximate memorization in language models via dissimilarity learned policy
Aly M Kassem · 2023
Later among the works it cites.
Roast: Robustifying language models via adversarial perturbation with selective training
Jaehyung Kim, Yuning Mao, Rui Hou, Hanchao Yu, Davis Liang, Pascale Fung, Qifan Wang, Fuli Feng, Lifu Huang, and Madian Khabsa · 2023
Later among the works it cites.
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein · 2023
Later among the works it cites.
On robustness-accuracy characterization of large language models using synthetic datasets
Ching-Yun Ko, Pin-Yu Chen, Payel Das, Yung-Sung Chuang, and Luca Daniel · 2023
Later among the works it cites.
An efficient membership inference attack for the diffusion model by proximal initialization
Fei Kong, Jinhao Duan, RuiPeng Ma, Hengtao Shen, Xiaofeng Zhu, Xiaoshuang Shi, and Kaidi Xu · 2023
Later among the works it cites.
Character as pixels: A controllable prompt adversarial attacking framework for black-box text guided image generation models
Ziyi Kou, Shichao Pei, Yijun Tian, and Xiangliang Zhang · 2023
Later among the works it cites.
Large language models and generative ai in finance: An analysis of chatgpt, bard, and bing ai
David Krause · 2023
Later among the works it cites.
Ablating concepts in text-to-image diffusion models
Nupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman, Richard Zhang, and Jun-Yan Zhu · 2023
Later among the works it cites.
The rise of generative artificial intelligence in healthcare
Murat Kuzlu, Zhenxin Xiao, Salih Sarp, Ferhat Ozgur Catak, Necip Gurler, and Ozgur Guler · 2023
Later among the works it cites.
Preethi Lahoti, Nicholas Blumm, Xiao Ma, Raghavendra Kotikalapudi, Sahitya Potluri, Qijun Tan, Hansa Srinivasan, Ben Packer, Ahmad Beirami, Alex Beutel, et al · 2023
Later among the works it cites.
Influencer backdoor attack on semantic segmentation
Haoheng Lan, Jindong Gu, Philip Torr, and Hengshuang Zhao · 2023
Later among the works it cites.
Open sesame! universal black box jailbreaking of large language models
Raz Lapid, Ron Langberg, and Moshe Sipper · 2023
Later among the works it cites.
Single-model attribution via final-layer inversion
Mike Laszkiewicz, Jonas Ricker, Johannes Lederer, and Asja Fischer · 2023
Later among the works it cites.
Aligning text-to-image models using human feedback
Kimin Lee, Hao Liu, Moonkyung Ryu, Olivia Watkins, Yuqing Du, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, and Shixiang Shane Gu · 2023
Later among the works it cites.
Dall· e 2 fails to reliably capture common syntactic processes
Evelina Leivada, Elliot Murphy, and Gary Marcus · 2023
Later among the works it cites.
Not with my name! inferring artists’ names of input strings employed by diffusion models
Roberto Leotta, Oliver Giudice, Luca Guarnera, and Sebastiano Battiato · 2023
Later among the works it cites.
Halueval: A large-scale hallucination evaluation benchmark for large language models
Junyi Li, Xiaoxue Cheng, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen · 2023
Later among the works it cites.
Mist: Towards improved adversarial examples for diffusion models
Chumeng Liang and Xiaoyu Wu · 2023
Later among the works it cites.
Coco: Coherence-enhanced machine-generated text detection under low resource with contrastive learning
Xiaoming Liu, Zhaohan Zhang, Yichen Wang, Hang Pu, Yu Lan, and Chao Shen · 2023
Later among the works it cites.
Explainable clinical coding with in-domain adapted transformers
Guillermo López-García, José M Jerez, Nuria Ribelles, Emilio Alba, and Francisco J Veredas · 2023
Later among the works it cites.
Analyzing leakage of personally identifiable information in language models
Nils Lukas, Ahmed Salem, Robert Sim, Shruti Tople, Lukas Wutschitz, and Santiago Zanella-Béguelin · 2023
Later among the works it cites.
Differentially private latent diffusion models
Saiyue Lyu, Margarita Vinaroz, Michael F Liu, and Mijung Park · 2023
Later among the works it cites.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Potsawee Manakul, Adian Liusie, and Mark JF Gales · 2023
Later among the works it cites.
A holistic approach to undesired content detection in the real world
Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng · 2023
Later among the works it cites.
Membership inference attacks against diffusion models
Tomoya Matsumoto, Takayuki Miura, and Naoto Yanai · 2023
Later among the works it cites.
Membership inference attacks against language models via neighbourhood comparison
Justus Mattern, Fatemehsadat Mireshghallah, Zhijing Jin, Bernhard Schölkopf, Mrinmaya Sachan, and Taylor Berg-Kirkpatrick · 2023
Later among the works it cites.
Did the neurons read your book? document-level membership inference for large language models
Matthieu Meeus, Shubham Jain, Marek Rei, and Yves-Alexandre de Montjoye · 2023
Later among the works it cites.
Assert: Automated safety scenario red teaming for evaluating the robustness of large language models
Alex Mei, Sharon Levy, and William Yang Wang · 2023
Later among the works it cites.
Selfcheck: Using llms to zero-shot check their own step-by-step reasoning
Ning Miao, Yee Whye Teh, and Tom Rainforth · 2023
Later among the works it cites.
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Wei Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Detectgpt: Zero-shot machine-generated text detection using probability curvature
Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D Manning, and Chelsea Finn · 2023
Later among the works it cites.
Sandra Mitrović, Davide Andreoletti, and Omran Ayoub · 2023
Later among the works it cites.
Use of llms for illicit purposes: Threats, prevention measures, and vulnerabilities
Maximilian Mozes, Xuanli He, Bennett Kleinberg, and Lewis D Griffin · 2023
Later among the works it cites.
Controlled decoding from language models
Sidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li, Tao Wang, Yanping Huang, Zhifeng Chen, Heng-Tze Cheng, Michael Collins, Trevor Strohman, et al · 2023
Later among the works it cites.
A multimodal comparison of latent denoising diffusion probabilistic models and generative adversarial networks for medical image synthesis
Gustav Müller-Franzes, Jan Moritz Niehues, Firas Khader, Soroosh Tayebi Arasteh, Christoph Haarburger, Christiane Kuhl, Tianci Wang, Tianyu Han, Teresa Nolte, Sven Nebelung, et al · 2023
Later among the works it cites.
Social biases through the text-to-image generation lens
Ranjita Naik and Besmira Nushi · 2023
Later among the works it cites.
Modelling temporal document sequences for clinical icd coding
Clarence Boon Liang Ng, Diogo Santos, and Marek Rei · 2023
Later among the works it cites.
Ores: Open-vocabulary responsible visual synthesis
Minheng Ni, Chenfei Wu, Xiaodong Wang, Shengming Yin, Lijuan Wang, Zicheng Liu, and Nan Duan · 2023
Later among the works it cites.
Attributing image generative models using latent fingerprints
Guangyu Nie, Changhoon Kim, Yezhou Yang, and Yi Ren · 2023
Later among the works it cites.
Automated radiology report generation using transformers
Wimukthi Nimalsiri, Mahela Hennayake, Kasun Rathnayake, Thanuja D Ambegoda, and Dulani Meedeniya · 2023
Later among the works it cites.
Diffmix: Diffusion model-based data synthesis for nuclei segmentation and classification in imbalanced pathology image datasets
Hyun-Jic Oh and Won-Ki Jeong · 2023
Later among the works it cites.
Membership inference attacks with token-level deduplication on korean language models
Myung Gyo Oh, Leo Hyun Park, Jaeuk Kim, Jaewoo Park, and Taekyoung Kwon · 2023
Later among the works it cites.
Towards universal fake image detectors that generalize across generative models
Utkarsh Ojha, Yuheng Li, and Yong Jae Lee · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
A taxonomy of prompt modifiers for text-to-image generation
Jonas Oppenlaender · 2023
Later among the works it cites.
Unsupervised medical image translation with adversarial diffusion models
Muzaffer Özbey, Onat Dalmaz, Salman UH Dar, Hasan A Bedel, Şaban Özturk, Alper Güngör, and Tolga Çukur · 2023
Later among the works it cites.
Controlling the extraction of memorized data from large language models via prompt-tuning
Mustafa Safa Ozdayi, Charith Peris, Jack Fitzgerald, Christophe Dupuy, Jimit Majmudar, Haidar Khan, Rahil Parikh, and Rahul Gupta · 2023
Later among the works it cites.
How to catch an ai liar: Lie detection in black-box llms by asking unrelated questions
Lorenzo Pacchiardi, Alex J Chan, Sören Mindermann, Ilan Moscovitz, Alexa Y Pan, Yarin Gal, Owain Evans, and Jan Brauner · 2023
Later among the works it cites.
2d medical image synthesis using transformer-based denoising diffusion probabilistic model
Shaoyan Pan, Tonghe Wang, Richard LJ Qiu, Marian Axente, Chih-Wei Chang, Junbo Peng, Ashish B Patel, Joseph Shelton, Sagar A Patel, Justin Roper, et al · 2023
Later among the works it cites.
White-box membership inference attacks against diffusion models
Yan Pang, Tianhao Wang, Xuhui Kang, Mengdi Huai, and Yang Zhang · 2023
Later among the works it cites.
Trak: Attributing model behavior at scale
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry · 2023
Later among the works it cites.
Analyzing bias in diffusion-based face generation models
Malsha V Perera and Vishal M Patel · 2023
Later among the works it cites.
Circumventing concept erasure methods for text-to-image generative models
Minh Pham, Kelly O Marshall, Niv Cohen, Govind Mittal, and Chinmay Hegde · 2023
Later among the works it cites.
Huachuan Qiu, Shuai Zhang, Anqi Li, Hongliang He, and Zhenzhong Lan · 2023
Later among the works it cites.
Unsafe diffusion: On the generation of unsafe images and hateful memes from text-to-image models
Yiting Qu, Xinyue Shen, Xinlei He, Michael Backes, Savvas Zannettou, and Yang Zhang · 2023
Later among the works it cites.
Aart: Ai-assisted red-teaming with diverse data generation for new llm-powered applications
Bhaktipriya Radharapu, Kevin Robinson, Lora Aroyo, and Preethi Lahoti · 2023
Later among the works it cites.
In-context retrieval-augmented language models
Ori Ram, Yoav Levine, Itay Dalmedigos, Dor Muhlgay, Amnon Shashua, Kevin Leyton-Brown, and Yoav Shoham · 2023
Later among the works it cites.
Role and challenges of chatgpt and similar generative artificial intelligence in finance and accounting
Nitin Rane · 2023
Later among the works it cites.
Tricking llms into disobedience: Understanding, analyzing, and preventing jailbreaks
Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya, and Monojit Choudhury · 2023
Later among the works it cites.
Generative ai in finance: Risks and potential solutions
Nydia Remolina · 2023
Later among the works it cites.
My art my choice: Adversarial protection against unruly ai
Anthony Rhodes, Ram Bhagat, Umur Aybars Ciftci, and Ilke Demir · 2023
Later among the works it cites.
Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Nataniel Ruiz, Yuanzhen Li, Varun Jampani, Yael Pritch, Michael Rubinstein, and Kfir Aberman · 2023
Later among the works it cites.
Can ai-generated text be reliably detected?
Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, and Soheil Feizi · 2023
Later among the works it cites.
Raising the cost of malicious ai-powered image editing
Hadi Salman, Alaa Khaddaj, Guillaume Leclerc, Andrew Ilyas, and Aleksander Madry · 2023
Later among the works it cites.
Autextification: automatic text identification
Areg Mikael Sarvazyan, José Ángel González, M Franco Salvador, Francisco Rangel, Berta Chulvi, and Paolo Rosso · 2023
Later among the works it cites.
Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models
Patrick Schramowski, Manuel Brack, Björn Deiseroth, and Kristian Kersting · 2023
Later among the works it cites.
De-fake: Detection and attribution of fake images generated by text-to-image generation models
Zeyang Sha, Zheng Li, Ning Yu, and Yang Zhang · 2023
Later among the works it cites.
Generative artificial intelligence in finance
Ghiath Shabsigh and El Bachir Boukherouaa · 2023
Later among the works it cites.
Scalable and transferable black-box jailbreaks for language models via persona modulation
Rusheb Shah, Soroush Pour, Arush Tagade, Stephen Casper, Javier Rando, et al · 2023
Later among the works it cites.
Glaze: Protecting artists from style mimicry by { \{ Text-to-Image } \} models
Shawn Shan, Jenna Cryan, Emily Wenger, Haitao Zheng, Rana Hanocka, and Ben Y Zhao · 2023
Later among the works it cites.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R Johnston, et al · 2023
Later among the works it cites.
Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models
Erfan Shayegani, Yue Dong, and Nael B. Abu-Ghazaleh · 2023
Later among the works it cites.
Prompt stealing attacks against text-to-image generation models
Xinyue Shen, Yiting Qu, Michael Backes, and Yang Zhang · 2023
Later among the works it cites.
In-context pretraining: Language modeling beyond document boundaries
Weijia Shi, Sewon Min, Maria Lomeli, Chunting Zhou, Margaret Li, Victoria Lin, Noah A Smith, Luke Zettlemoyer, Scott Yih, and Mike Lewis · 2023
Later among the works it cites.
A comprehensive review of generative ai in healthcare
Yasin Shokrollahi, Sahar Yarmohammadtoosky, Matthew M Nikahd, Pengfei Dong, Xianqi Li, and Linxia Gu · 2023
Later among the works it cites.
On the exploitability of instruction tuning
Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, and Tom Goldstein · 2023
Later among the works it cites.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al · 2023
Later among the works it cites.
Break it, imitate it, fix it: Robustness by generating human-like attacks
Aradhana Sinha, Ananth Balashankar, Ahmad Beirami, Thi Avrahami, Jilin Chen, and Alex Beutel · 2023
Later among the works it cites.
Diffusion art or digital forgery? investigating data replication in diffusion models
Gowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping, and Tom Goldstein · 2023
Later among the works it cites.
Beyond memorization: Violating privacy via inference with large language models
Robin Staab, Mark Vero, Mislav Balunović, and Martin Vechev · 2023
Later among the works it cites.
Rickrolling the artist: Injecting backdoors into text encoders for text-to-image synthesis
Lukas Struppek, Dominik Hintersdorf, and Kristian Kersting · 2023
Later among the works it cites.
Vipergpt: Visual inference via python execution for reasoning
Dídac Surís, Sachit Menon, and Carl Vondrick · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models, 2023
Google Gemini Team · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo et al. Touvron · 2023
Later among the works it cites.
Ring-a-bell! how reliable are concept removal methods for diffusion models?
Yu-Lin Tsai, Chia-Yi Hsu, Chulin Xie, Chih-Hsun Lin, Jia-You Chen, Bo Li, Pin-Yu Chen, Chia-Mu Yu, and Chun-Ying Huang · 2023
Later among the works it cites.
Med-halt: Medical domain hallucination test for large language models
Logesh Kumar Umapathi, Ankit Pal, and Malaikannan Sankarasubbu · 2023
Later among the works it cites.
Anti-dreambooth: Protecting users from personalized text-to-image synthesis
Thanh Van Le, Hao Phung, Thuan Hoang Nguyen, Quan Dao, Ngoc N Tran, and Anh Tran · 2023
Later among the works it cites.
Chatgpt: The transformative influence of generative ai on science and healthcare
Julian Varghese and Julius Chapiro · 2023
Later among the works it cites.
Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu · 2023
Later among the works it cites.
Using artificial intelligence in craft education: crafting with text-to-image generative models
Henriikka Vartiainen and Matti Tedre · 2023
Later among the works it cites.
Bagm: A backdoor attack for manipulating text-to-image generative models
Jordan Vice, Naveed Akhtar, Richard Hartley, and Ajmal Mian · 2023
Later among the works it cites.
Fairpy: A toolkit for evaluation of social biases and their mitigation in large language models
Hrishikesh Viswanath and Tianyi Zhang · 2023
Later among the works it cites.
Freshllms: Refreshing large language models with search engine augmentation
Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, et al · 2023
Later among the works it cites.
Poisoning language models during instruction tuning
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein · 2023
Later among the works it cites.
Generative ai in operational risk management: Harnessing the future of finance
Yanqing Wang · 2023
Later among the works it cites.
A reproducible extraction of training images from diffusion models
Ryan Webster · 2023
Later among the works it cites.
Llmdet: A third party large language models generated text detection tool
Kangxi Wu, Liang Pang, Huawei Shen, Xueqi Cheng, and Tat-Seng Chua · 2023
Later among the works it cites.
Toward effective protection against diffusion-based mimicry through score distillation
Haotian Xue, Chumeng Liang, Xiaoyu Wu, and Yongxin Chen · 2023
Later among the works it cites.
Backdooring instruction-tuned large language models with virtual prompt injection
Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin · 2023
Later among the works it cites.
Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Pinjia He, Shuming Shi, and Zhaopeng Tu · 2023
Later among the works it cites.
Narcissus: A practical clean-label backdoor attack with limited information
Yi Zeng, Minzhou Pan, Hoang Anh Just, Lingjuan Lyu, Meikang Qiu, and Ruoxi Jia · 2023
Later among the works it cites.
Text-to-image diffusion models can be easily backdoored through multimodal data poisoning
Shengfang Zhai, Yinpeng Dong, Qingni Shen, Shi Pu, Yuejian Fang, and Hang Su · 2023
Later among the works it cites.
Intriguing properties of data attribution on diffusion models
Xiaosen Zheng, Tianyu Pang, Chao Du, Jing Jiang, and Min Lin · 2023
Later among the works it cites.
A transformer-based representation-learning model with unified processing of multimodal input for clinical diagnostics
Hong-Yu Zhou, Yizhou Yu, Chengdi Wang, Shu Zhang, Yuanxu Gao, Jia Pan, Jun Shao, Guangming Lu, Kang Zhang, and Weimin Li · 2023
Later among the works it cites.
Promptbench: Towards evaluating the robustness of large language models on adversarial prompts
Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, et al · 2023
Later among the works it cites.
A pilot study of query-free adversarial attack against stable diffusion
Haomin Zhuang, Yihua Zhang, and Sijia Liu · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
URL https://github.com/pharmapsychotic/clip-interrogator
Clip interrogator, 2024 · 2024
Closest in time.
Detectors for safe and reliable llms: Implementations, uses, and limitations
Swapnaja Achintalwar, Adriana Alvarado Garcia, Ateret Anaby-Tavor, Ioana Baldini, Sara E Berger, Bishwaranjan Bhattacharjee, Djallel Bouneffouf, Subhajit Chaudhury, Pin-Yu Chen, Lamogha Chiazor, et al · 2024
Closest in time.
Elijah: Eliminating backdoors injected in diffusion models via distribution shift
Shengwei An, Sheng-Yen Chou, Kaiyuan Zhang, Qiuling Xu, Guanhong Tao, Guangyu Shen, Siyuan Cheng, Shiqing Ma, Pin-Yu Chen, Tsung-Yi Ho, et al · 2024
Closest in time.
Special characters attack: Toward scalable training data extraction from large language models
Yang Bai, Ge Pei, Jindong Gu, Yong Yang, and Xingjun Ma · 2024
Closest in time.
Theoretical guarantees on the best-of-n alignment policy
Ahmad Beirami, Alekh Agarwal, Jonathan Berant, Alexander D’Amour, Jacob Eisenstein, Chirag Nagpal, and Ananda Theertha Suresh · 2024
Closest in time.
Red teaming gpt-4v: Are gpt-4v safe against uni/multi-modal jailbreak attacks?
Shuo Chen, Zhen Han, Bailan He, Zifeng Ding, Wenqian Yu, Philip Torr, Volker Tresp, and Jindong Gu · 2024
Closest in time.
Generative ai in medical practice: In-depth exploration of privacy and security challenges
Yan Chen and Pouyan Esmaeilzadeh · 2024
Closest in time.
Villandiffusion: A unified backdoor attack framework for diffusion models
Sheng-Yen Chou, Pin-Yu Chen, and Tsung-Yi Ho · 2024
Closest in time.
Stable diffusion is unstable
Chengbin Du, Yanxi Li, Zhongwei Qiu, and Chang Xu · 2024
Closest in time.
Reinforcement learning for fine-tuning text-to-image diffusion models
Ying Fan, Olivia Watkins, Yuqing Du, Hao Liu, Moonkyung Ryu, Craig Boutilier, Pieter Abbeel, Mohammad Ghavamzadeh, Kangwook Lee, and Kimin Lee · 2024
Closest in time.
Mathematical capabilities of chatgpt
Simon Frieder, Luca Pinchetti, Ryan-Rhys Griffiths, Tommaso Salvatori, Thomas Lukasiewicz, Philipp Petersen, and Julius Berner · 2024
Closest in time.
Unified concept editing in diffusion models
Rohit Gandikota, Hadas Orgad, Yonatan Belinkov, Joanna Materzyńska, and David Bau · 2024
Closest in time.
Inducing high energy-latency of large vision-language models with verbose images
Kuofeng Gao, Yang Bai, Jindong Gu, Shu-Tao Xia, Philip Torr, Zhifeng Li, and Wei Liu · 2024
Closest in time.
Spotting llms with binoculars: Zero-shot detection of machine-generated text
Abhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Goldstein · 2024
Closest in time.
Optimizing prompts for text-to-image generation
Yaru Hao, Zewen Chi, Li Dong, and Furu Wei · 2024
Closest in time.
Xiaomeng Hu, Pin-Yu Chen, and Tsung-Yi Ho · 2024
Closest in time.
Unfamiliar finetuning examples control how language models hallucinate
Katie Kang, Eric Wallace, Claire Tomlin, Aviral Kumar, and Sergey Levine · 2024
Closest in time.
Stable bias: Evaluating societal representations in diffusion models
Sasha Luccioni, Christopher Akiki, Margaret Mitchell, and Yacine Jernite · 2024
Closest in time.
An image is worth 1000 lies: Transferability of adversarial images across prompts on vision-language models
Haochen Luo, Jindong Gu, Fengyuan Liu, and Philip Torr · 2024
Closest in time.
Dall·e 2 pre-training mitigations, 2024a
OpenAI · 2024
Closest in time.
Sora, 2024b
OpenAI · 2024
Closest in time.
Language model tokenizers introduce unfairness between languages
Aleksandar Petrov, Emanuele La Malfa, Philip Torr, and Adel Bibi · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Closest in time.
Gemini or chatgpt? efficiency, performance, and adaptability of cutting-edge generative artificial intelligence (ai) in finance and accounting
Nitin Rane, Saurabh Choudhary, and Jayesh Rane · 2024
Closest in time.
Journey of hallucination-minimized generative ai solutions for financial decision makers
Sohini Roychowdhury · 2024
Closest in time.
Generative ai for transformative healthcare: A comprehensive study of emerging models, applications, case studies and limitations
Siva Sai, Aanchal Gaur, Revant Sai, Vinay Chamola, Mohsen Guizani, and Joel JPC Rodrigues · 2024
Closest in time.
Tokenization counts: the impact of tokenization on arithmetic in frontier llms
Aaditya K Singh and DJ Strouse · 2024
Closest in time.
Hansa Srinivasan, Candice Schumann, Aradhana Sinha, David Madras, Gbolahan Oluwafemi Olanubi, Alex Beutel, Susanna Ricco, and Jilin Chen · 2024
Closest in time.
Gradient-based language model red teaming
Nevan Wichers, Carson Denison, and Ahmad Beirami · 2024
Closest in time.
Imagereward: Learning and evaluating human preferences for text-to-image generation
Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong · 2024
Closest in time.
Minimalism is king! high-frequency energy-based screening for data-efficient backdoor attacks
Yuan Xun, Xiaojun Jia, Jindong Gu, Xinwei Liu, Qing Guo, and Xiaochun Cao · 2024
Closest in time.
Sneakyprompt: Jailbreaking text-to-image generative models
Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong, and Yinzhi Cao · 2024
Closest in time.
Rigorllm: Resilient guardrails for large language models against undesired content
Zhuowen Yuan, Zidi Xiong, Yi Zeng, Ning Yu, Ruoxi Jia, Dawn Song, and Bo Li · 2024
Closest in time.
Autodefense: Multi-agent llm defense against jailbreak attacks
Yifan Zeng, Yiran Wu, Xiao Zhang, Huazheng Wang, and Qingyun Wu · 2024
Closest in time.
Explainability for large language models: A survey
Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du · 2024
Closest in time.
Duwak: Dual watermarks in large language models
Chaoyi Zhu, Jeroen Galjaard, Pin-Yu Chen, and Lydia Y Chen · 2024
Closest in time.
Human preference score: Better aligning text-to-image models with human preference
Xiaoshi Wu, Keqiang Sun, Feng Zhu, Rui Zhao, and Hongsheng Li · 2096
Closest in time.