Fetching the paper…
Reading the bibliography…
Deep learning models are often trained on distributed, web-scale datasets crawled from the internet.
Building a large annotated corpus of English: The Penn Treebank
Mary Ann Marcinkiewicz · 1994
Earlier work this paper cites.
Information quality work organization in Wikipedia
Besiki Stvilia, Michael B Twidale, Linda C Smith, and Les Gasser · 2008
Earlier work this paper cites.
80 million tiny images: A large data set for nonparametric object and scene recognition
Antonio Torralba, Rob Fergus, and William T Freeman · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images, 2009
Alex Krizhevsky · 2009
Earlier work this paper cites.
Attribute and simile classifiers for face verification
Neeraj Kumar, Alexander C Berg, Peter N Belhumeur, and Shree K Nayar · 2009
Earlier work this paper cites.
Beyond vandalism: Wikipedia trolls
Pnina Shachaf and Noriko Hara · 2010
Earlier work this paper cites.
Bagging classifiers for fighting poisoning attacks in adversarial classification tasks
Battista Biggio, Igino Corona, Giorgio Fumera, Giorgio Giacinto, and Fabio Roli · 2011
Earlier work this paper cites.
Support vector machines under adversarial label noise
Battista Biggio, Blaine Nelson, and Pavel Laskov · 2011
Earlier work this paper cites.
Poisoning attacks against support vector machines
Battista Biggio, Blaine Nelson, and Pavel Laskov · 2012
Earlier work this paper cites.
You are what you include: large-scale evaluation of remote Javascript inclusions
Nick Nikiforakis, Luca Invernizzi, Alexandros Kapravelos, Steven Van Acker, Wouter Joosen, Christopher Kruegel, Frank Piessens, and Giovanni Vigna · 2012
Earlier work this paper cites.
pHash: The open source perceptual hash library
Evan Klinger and David Starkweather · 2013
Earlier work this paper cites.
Certificate transparency
Ben Laurie · 2014
Earlier work this paper cites.
The ghosts of banking past: Empirical analysis of closed bank websites
Tyler Moore and Richard Clayton · 2014
Earlier work this paper cites.
A data-driven approach to cleaning large face datasets
Hong-Wei Ng and Stefan Winkler · 2014
Earlier work this paper cites.
Deep face recognition
Omkar M Parkhi, Andrea Vedaldi, and Andrew Zisserman · 2015
Earlier work this paper cites.
The abandoned side of the Internet: Hijacking Internet resources when domain names expire
Johann Schlamp, Josef Gustafsson, Matthias Wählisch, Thomas C Schmidt, and Georg Carle · 2015
Earlier work this paper cites.
Is feature selection secure against training data poisoning?
Huang Xiao, Battista Biggio, Gavin Brown, Giorgio Fumera, Claudia Eckert, and Fabio Roli · 2015
Earlier work this paper cites.
WikiQA: A challenge dataset for open-domain question answering
Yi Yang, Wen-tau Yih, and Christopher Meek · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Earlier work this paper cites.
The MegaFace benchmark: 1 million faces for recognition at scale
Ira Kemelmacher-Shlizerman, Steven M Seitz, Daniel Miller, and Evan Brossard · 2016
Earlier work this paper cites.
Generating text from structured data with application to the biography domain
Rémi Lebret, David Grangier, and Michael Auli · 2016
Earlier work this paper cites.
Domain-Z: 28 registrations later measuring the exploitation of residual trust in domains
Chaz Lever, Robert Walls, Yacin Nadji, David Dagon, Patrick McDaniel, and Manos Antonakakis · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Bitcoin and cryptocurrency technologies: a comprehensive introduction
Arvind Narayanan, Joseph Bonneau, Edward Felten, Andrew Miller, and Steven Goldfeder · 2016
Earlier work this paper cites.
YFCC100M: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Earlier work this paper cites.
The do’s and don’ts for CNN-based face verification
Ankan Bansal, Carlos Castillo, Rajeev Ranjan, and Rama Chellappa · 2017
Earlier work this paper cites.
UMDfaces: An annotated face dataset for training deep networks
Ankan Bansal, Anirudh Nanduri, Carlos D Castillo, Rajeev Ranjan, and Rama Chellappa · 2017
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song · 2017
Earlier work this paper cites.
Badnets: Identifying vulnerabilities in the machine learning model supply chain
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg · 2017
Earlier work this paper cites.
Game of registrars: An empirical analysis of post-expiration domain name takeovers
Tobias Lauinger, Abdelberi Chaabane, Ahmet Salih Buyukkayhan, Kaan Onarlioglu, and William Robertson · 2017
Earlier work this paper cites.
Towards poisoning of deep learning algorithms with back-gradient optimization
Luis Muñoz-González, Battista Biggio, Ambra Demontis, Andrea Paudice, Vasin Wongrassamee, Emil C Lupu, and Fabio Roli · 2017
Earlier work this paper cites.
Deep learning is robust to massive label noise
David Rolnick, Andreas Veit, Serge Belongie, and Nir Shavit · 2017
Earlier work this paper cites.
Contour: A practical system for binary transparency
Mustafa Al-Bassam and Sarah Meiklejohn · 2018
Earlier work this paper cites.
VGGFace2: A dataset for recognising faces across pose and age
Qiong Cao, Li Shen, Weidi Xie, Omkar M Parkhi, and Andrew Zisserman · 2018
Earlier work this paper cites.
Detecting backdoor attacks on deep neural networks by activation clustering
Bryant Chen, Wilka Carvalho, Nathalie Baracaldo, Heiko Ludwig, Benjamin Edwards, Taesung Lee, Ian Molloy, and Biplav Srivastava · 2018
Earlier work this paper cites.
Manipulating machine learning: Poisoning attacks and countermeasures for regression learning
Matthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu, Cristina Nita-Rotaru, and Bo Li · 2018
Earlier work this paper cites.
From deletion to re-registration in zero seconds: Domain registrar behaviour during the drop
Tobias Lauinger, Ahmet S Buyukkayhan, Abdelberi Chaabane, William Robertson, and Engin Kirda · 2018
Earlier work this paper cites.
Poison frogs! targeted clean-label poisoning attacks on neural networks
Ali Shafahi, W Ronny Huang, Mahyar Najibi, Octavian Suciu, Christoph Studer, Tudor Dumitras, and Tom Goldstein · 2018
Earlier work this paper cites.
Conceptual Captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Cited alongside, same era.
Spectral signatures in backdoor attacks
Brandon Tran, Jerry Li, and Aleksander Madry · 2018
Cited alongside, same era.
A new backdoor attack in cnns by training set corruption without label poisoning
Mauro Barni, Kassem Kallas, and Benedetta Tondi · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Strip: A defence against trojan attacks on deep neural networks
Yansong Gao, Change Xu, Derui Wang, Shiping Chen, Damith C Ranasinghe, and Surya Nepal · 2019
Cited alongside, same era.
Badnets: Evaluating backdooring attacks on deep neural networks
Anti-backdoor learning: Training clean models on poisoned data
Yige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu, Bo Li, and Xingjun Ma · 2021
Later among the works it cites.
Backdoor attack in the physical world
Yiming Li, Tongqing Zhai, Yong Jiang, Zhifeng Li, and Shu-Tao Xia · 2021
Later among the works it cites.
Invisible backdoor attack with sample-specific triggers
Yuezun Li, Yiming Li, Baoyuan Wu, Longkang Li, Ran He, and Siwei Lyu · 2021
Later among the works it cites.
What’s in the box? an analysis of undesirable content in the Common Crawl corpus
Alexandra Luccioni and Joseph Viviano · 2021
Later among the works it cites.
Wanet–imperceptible warping-based backdoor attack
Anh Nguyen and Anh Tran · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tianyu Gu, Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg · 2019
Cited alongside, same era.
Neuroninspect: Detecting backdoors in neural networks via output explanations
Xijie Huang, Moustafa Alzantot, and Mani Srivastava · 2019
Cited alongside, same era.
Abs: Scanning neural networks for back-doors by artificial brain stimulation
Yingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma, Yousra Aafer, and Xiangyu Zhang · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Cited alongside, same era.
Clean-label backdoor attacks, 2019
Alexander Turner, Dimitris Tsipras, and Aleksander Madry · 2019
Cited alongside, same era.
Label-consistent backdoor attacks
Alexander Turner, Dimitris Tsipras, and Aleksander Madry · 2019
Cited alongside, same era.
Latent backdoor attacks on deep neural networks
Yuanshun Yao, Huiying Li, Haitao Zheng, and Ben Y Zhao · 2019
Cited alongside, same era.
Fanchao Qi, Mukai Li, Yangyi Chen, Zhengyan Zhang, Zhiyuan Liu, Yasheng Wang, and Maosong Sun · 2021
Later among the works it cites.
Deepsweep: An evaluation framework for mitigating dnn backdoor attacks using data augmentation
Han Qiu, Yi Zeng, Shangwei Guo, Tianwei Zhang, Meikang Qiu, and Bhavani Thuraisingham · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
LAION-400M: Open dataset of clip-filtered 400 million image-text pairs
Christoph Schuhmann, Richard Vencu, Romain Beaumont, Robert Kaczmarczyk, Clayton Mullis, Aarush Katta, Theo Coombes, Jenia Jitsev, and Aran Komatsuzaki · 2021
Later among the works it cites.
You autocomplete me: Poisoning vulnerabilities in neural code completion
Roei Schuster, Congzheng Song, Eran Tromer, and Vitaly Shmatikov · 2021
Later among the works it cites.
Backdoor scanning for deep neural networks through k-arm optimization
Guangyu Shen, Yingqi Liu, Guanhong Tao, Shengwei An, Qiuling Xu, Siyuan Cheng, Shiqing Ma, and Xiangyu Zhang · 2021
Later among the works it cites.
Adversarial neuron pruning purifies backdoored deep models
Dongxian Wu and Yisen Wang · 2021
Later among the works it cites.
Detecting ai trojans using meta neural analysis
Xiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov, Carl A Gunter, and Bo Li · 2021
Later among the works it cites.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Later among the works it cites.
COYO-700M: Image-text pair dataset
Minwoo Byeon, Beomhee Park, Haecheon Kim, Sungjun Lee, Woonhyuk Baek, and Saehoon Kim · 2022
Later among the works it cites.
PaLM: Scaling language modeling with Pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Later among the works it cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Later among the works it cites.
Adversarial detection avoidance attacks: Evaluating the robustness of perceptual hashing-based client-side scanning
Shubham Jain, Ana-Maria Crețu, and Yves-Alexandre de Montjoye · 2022
Later among the works it cites.
Introducing Whisper
OpenAI · 2022
Later among the works it cites.
Red-teaming the Stable Diffusion safety filter
Javier Rando, Daniel Paleka, David Lindner, Lennart Heim, and Florian Tramèr · 2022
Later among the works it cites.
Dynamic backdoor attacks against machine learning models
Ahmed Salem, Rui Wen, Michael Backes, Shiqing Ma, and Yang Zhang · 2022
Later among the works it cites.
LAION-5B: An open large-scale dataset for training next generation image-text models
Christoph Schuhmann, Romain Beaumont, Richard Vencu, Cade Gordon, Ross Wightman, Mehdi Cherti, Theo Coombes, Aarush Katta, Clayton Mullis, Mitchell Wortsman, et al · 2022
Later among the works it cites.
LAION-Aesthetics
Christoph Schumann and Romain Beaumont · 2022
Later among the works it cites.
Poison forensics: Traceback of data poisoning attacks in neural networks
Shawn Shan, Arjun Nitin Bhagoji, Haitao Zheng, and Ben Y Zhao · 2022
Later among the works it cites.
Flava: A foundational language and vision alignment model
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela · 2022
Later among the works it cites.
Domains do change their spots: Quantifying potential abuse of residual trust
Johnny So, Najmeh Miramirkhani, Michael Ferdman, and Nick Nikiforakis · 2022
Later among the works it cites.
Sleeper agent: Scalable hidden trigger backdoors for neural networks trained from scratch
Hossein Souri, Liam Fowl, Rama Chellappa, Micah Goldblum, and Tom Goldstein · 2022
Later among the works it cites.
Learning to break deep perceptual hashing: The use case NeuralHash
Lukas Struppek, Dominik Hintersdorf, Daniel Neider, and Kristian Kersting · 2022
Later among the works it cites.
Truth serum: Poisoning machine learning models to reveal their secrets
Florian Tramèr, Reza Shokri, Ayrton San Joaquin, Hoang Le, Matthew Jagielski, Sanghyun Hong, and Nicholas Carlini · 2022
Later among the works it cites.
REVISE: A tool for measuring and mitigating bias in visual datasets
Angelina Wang, Alexander Liu, Ryan Zhang, Anat Kleiman, Leslie Kim, Dora Zhao, Iroha Shirai, Arvind Narayanan, and Olga Russakovsky · 2022
Later among the works it cites.
Reliability of wikipedia — Wikipedia, the free encyclopedia, 2022
Wikipedia contributors · 2022
Later among the works it cites.
Wikipedia:Go ahead, vandalize, 2022
Wikipedia contributors · 2022
Later among the works it cites.
Venomave: Targeted poisoning against speech recognition
Hojjat Aghakhani, Lea Schönherr, Thorsten Eisenhofer, Dorothea Kolossa, Thorsten Holz, Christopher Kruegel, and Giovanni Vigna · 2023
Closest in time.
Obelisc: An open web-scale filtered dataset of interleaved image-text documents, 2023
Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman, Amanpreet Singh, Anton Lozhkov, Thomas Wang, Siddharth Karamcheti, Alexander M. Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh · 2023
Closest in time.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Closest in time.
Robust speech recognition via large-scale weak supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever · 2023
Closest in time.
Diffusers: State-of-the-art diffusion models
Patrick von Platen, Suraj Patil, Anton Lozhkov, Pedro Cuenca, Nathan Lambert, Kashif Rasul, Mishig Davaadorj, and Thomas Wolf · 2023
Closest in time.
Multimodal C4: An open, billion-scale corpus of images interleaved with text
Wanrong Zhu, Jack Hessel, Anas Awadalla, Samir Yitzhak Gadre, Jesse Dodge, Alex Fang, Youngjae Yu, Ludwig Schmidt, William Yang Wang, and Yejin Choi · 2023
Closest in time.