Fetching the paper…
Reading the bibliography…
Learning reward models from pairwise comparisons is a fundamental component in a number of domains, including autonomous control, conversational agents, and recommendation systems, as part of a broad goal of aligning automated decisions with user preferences.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Mathematics for economists
Carl P Simon, Lawrence Blume, et al · 1994
Earlier work this paper cites.
Neural network learning: Theoretical foundations
Martin Anthony and Peter L Bartlett · 1999
Earlier work this paper cites.
Backdoor embedding in convolutional neural network models via invisible perturbation
Haoti Zhong, Cong Liao, Anna Cinzia Squicciarini, Sencun Zhu, and David Miller · 2000
Earlier work this paper cites.
Support vector machines under adversarial label noise
Battista Biggio, Blaine Nelson, and Pavel Laskov · 2011
Earlier work this paper cites.
Poisoning attacks against support vector machines
Battista Biggio, Blaine Nelson, and Pavel Laskov · 2012
Earlier work this paper cites.
Adversarial label flips attack on support vector machines
Han Xiao, Huang Xiao, and Claudia Eckert · 2012
Earlier work this paper cites.
Learning trajectory preferences for manipulators via iterative improvement
Ashesh Jain, Brian Wojcik, Thorsten Joachims, and Ashutosh Saxena · 2013
Earlier work this paper cites.
Using machine teaching to identify optimal training-set attacks on machine learners
Shike Mei and Xiaojin Zhu · 2015
Earlier work this paper cites.
Debugging machine learning tasks
Aleksandar Chakarov, Aditya V. Nori, Sriram K. Rajamani, Shayak Sen, and Deepak Vijaykeerthy · 2016
Earlier work this paper cites.
Data poisoning attacks on factorization-based collaborative filtering
Bo Li, Yining Wang, Aarti Singh, and Yevgeniy Vorobeychik · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Earlier work this paper cites.
Robust linear regression against training data poisoning
Chang Liu, Bo Li, Yevgeniy Vorobeychik, and Alina Oprea · 2017
Earlier work this paper cites.
Neural trojans
Yuntao Liu, Yang Xie, and Ankur Srivastava · 2017
Earlier work this paper cites.
Towards poisoning of deep learning algorithms with back-gradient optimization
Luis Muñoz-González, Battista Biggio, Ambra Demontis, Andrea Paudice, Vasin Wongrassamee, Emil C Lupu, and Fabio Roli · 2017
Earlier work this paper cites.
Certified defenses for data poisoning attacks
Jacob Steinhardt, Pang Wei W Koh, and Percy S Liang · 2017
Earlier work this paper cites.
Generative poisoning attack method against neural networks
Chaofei Yang, Qing Wu, Hai Li, and Yiran Chen · 2017
Earlier work this paper cites.
Sentinet: Detecting localized universal attacks against deep learning systems
Edward Chou, Florian Tramèr, and Giancarlo Pellegrino · 2018
Earlier work this paper cites.
Poisoning attacks to graph-based recommender systems
Minghong Fang, Guolei Yang, Neil Zhenqiang Gong, and Jia Liu · 2018
Earlier work this paper cites.
Manipulating machine learning: Poisoning attacks and countermeasures for regression learning
Matthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu, Cristina Nita-Rotaru, and Bo Li · 2018
Earlier work this paper cites.
Eliciting pairwise preferences in recommender systems
Saikishore Kalloori, Francesco Ricci, and Rosella Gennari · 2018
Earlier work this paper cites.
Fine-pruning: Defending against backdooring attacks on deep neural networks
Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg · 2018
Earlier work this paper cites.
A voting-based system for ethical decision making
Ritesh Noothigattu, Snehalkumar Gaikwad, Edmond Awad, Sohan Dsouza, Iyad Rahwan, Pradeep Ravikumar, and Ariel Procaccia · 2018
Earlier work this paper cites.
Pairwise preferences learning for recommender systems
Nunung Nurul Qomariyah · 2018
Earlier work this paper cites.
When does machine learning { \{ FAIL } \} ? generalized transferability for evasion and poisoning attacks
Octavian Suciu, Radu Marginean, Yigitcan Kaya, Hal Daume III, and Tudor Dumitras · 2018
Earlier work this paper cites.
Spectral signatures in backdoor attacks
Brandon Tran, Jerry Li, and Aleksander Madry · 2018
Earlier work this paper cites.
Adversarial machine learning
Yevgeniy Vorobeychik and Kantarcioglu · 2018
Earlier work this paper cites.
Deepinspect: A black-box trojan detection and mitigation framework for deep neural networks
Huili Chen, Cheng Fu, Jishen Zhao, and Farinaz Koushanfar · 2019
Earlier work this paper cites.
Robust anomaly detection and backdoor attack detection via differential privacy
Min Du, R. Jia, and Dawn Xiaodong Song · 2019
Earlier work this paper cites.
Robust anomaly detection and backdoor attack detection via differential privacy
Min Du, Ruoxi Jia, and Dawn Song · 2019
Earlier work this paper cites.
Badnets: Evaluating backdooring attacks on deep neural networks
Tianyu Gu, Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg · 2019
Earlier work this paper cites.
Sensitive-sample fingerprinting of deep neural networks
Zecheng He, Tianwei Zhang, and Ruby Lee · 2019
Earlier work this paper cites.
Neuroninspect: Detecting backdoors in neural networks via output explanations
Xijie Huang, Moustafa Farid Alzantot, and Mani B. Srivastava · 2019
Earlier work this paper cites.
Universal litmus patterns: Revealing backdoor attacks in cnns
Soheil Kolouri, Aniruddha Saha, Hamed Pirsiavash, and Heiko Hoffmann · 2019
Earlier work this paper cites.
Abs: Scanning neural networks for back-doors by artificial brain stimulation
Yingqi Liu, Wen-Chuan Lee, Guanhong Tao, Shiqing Ma, Yousra Aafer, and X. Zhang · 2019
Earlier work this paper cites.
Label sanitization against label flipping poisoning attacks
Andrea Paudice, Luis Muñoz-González, and Emil C. Lupu · 2019
Earlier work this paper cites.
Demon in the variant: Statistical analysis of dnns for robust backdoor contamination detection
Di Tang, Xiaofeng Wang, Haixu Tang, and Kehuan Zhang · 2019
Earlier work this paper cites.
Neural cleanse: Identifying and mitigating backdoor attacks in neural networks
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y. Zhao · 2019
Cited alongside, same era.
Learning and decision-making from rank data
Lirong Xia · 2019
Cited alongside, same era.
Influence function based data poisoning attacks to top-n recommender systems
Minghong Fang, Neil Zhenqiang Gong, and Jia Liu · 2020
Cited alongside, same era.
On the effectiveness of mitigating data poisoning attacks with gradient shaping
Sanghyun Hong, Varun Chandrasekaran, Yigitcan Kaya, Tudor Dumitras, and Nicolas Papernot · 2020
Cited alongside, same era.
Metapoison: Practical general-purpose clean-label data poisoning
W Ronny Huang, Jonas Geiping, Liam Fowl, Gavin Taylor, and Tom Goldstein · 2020
Cited alongside, same era.
Backdoor scanning for deep neural networks through k-arm optimization
Guangyu Shen, Yingqi Liu, Guanhong Tao, Shengwei An, Qiuling Xu, Siyuan Cheng, Shiqing Ma, and X. Zhang · 2021
Later among the works it cites.
Robust learning for data poisoning attacks
Yunjuan Wang, Poorya Mianjy, and Raman Arora · 2021
Later among the works it cites.
Adversarial neuron pruning purifies backdoored deep models
Dongxian Wu and Yisen Wang · 2021
Later among the works it cites.
Detecting ai trojans using meta neural analysis
Xiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov, Carl A Gunter, and Bo Li · 2021
Later among the works it cites.
Topological detection of trojaned neural networks
Songzhu Zheng, Yikai Zhang, Hubert Wagner, Mayank Goswami, and Chao Chen · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mojan Javaheripi, Mohammad Samragh, Gregory Fields, Tara Javidi, and Farinaz Koushanfar · 2020
Cited alongside, same era.
Weight poisoning attacks on pre-trained models
Keita Kurita, Paul Michel, and Graham Neubig · 2020
Cited alongside, same era.
Backdooring and poisoning neural networks with image-scaling attacks
Erwin Quiring and Konrad Rieck · 2020
Cited alongside, same era.
Certified robustness to label-flipping attacks via randomized smoothing
Elan Rosenfeld, Ezra Winston, Pradeep Ravikumar, and Zico Kolter · 2020
Cited alongside, same era.
Hidden trigger backdoor attacks
Aniruddha Saha, Akshayvarun Subramanya, and Hamed Pirsiavash · 2020
Cited alongside, same era.
Exposing backdoors in robust machine learning models
Ezekiel Olamide Soremekun, Sakshi Udeshi, Sudipta Chattopadhyay, and Andreas Zeller · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
Clean-image backdoor: Attacking multi-label models with poisoned labels only
Kangjie Chen, Xiaoxuan Lou, Guowen Xu, Jiwei Li, and Tianwei Zhang · 2022
Later among the works it cites.
Towards effective and robust neural trojan defenses via input filtering
Kien Do, Haripriya Harikumar, Hung Le, Dung Nguyen, Truyen Tran, Santu Rana, Dang Nguyen, Willy Susilo, and Svetha Venkatesh · 2022
Later among the works it cites.
A feature-based on-line detector to remove adversarial-backdoors by iterative demarcation
Hao Fu, Akshaj Kumar Veldanda, Prashanth Krishnamurthy, Siddharth Garg, and Farshad Khorrami · 2022
Later among the works it cites.
Defending against the label-flipping attack in federated learning
Najeeb Moharram Jebreel, Josep Domingo-Ferrer, David Sánchez, and Alberto Blanco-Justicia · 2022
Later among the works it cites.
An adaptive black-box defense against trojan attacks
Guanxiong Liu, Abdallah Khreishah, Fatima Sharadgah, and Issa Khalil · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
Dynamic backdoor attacks against machine learning models
Ahmed Salem, Rui Wen, Michael Backes, Shiqing Ma, and Yang Zhang · 2022
Later among the works it cites.
A survey of neural trojan attacks and defenses in deep learning
Jie Wang, Ghulam Mubashar Hassan, and Naveed Akhtar · 2022
Later among the works it cites.
Training with more confidence: Mitigating injected and natural backdoors during training
Zhenting Wang, Hailun Ding, Juan Zhai, and Shiqing Ma · 2022
Later among the works it cites.
Data poisoning attacks against machine learning algorithms
Fahri Anıl Yerlikaya and Şerif Bahtiyar · 2022
Later among the works it cites.
Towards class-oriented poisoning attacks against neural networks
Bingyin Zhao and Yingjie Lao · 2022
Later among the works it cites.
Trojdiff: Trojan attacks on diffusion models with diverse targets
Weixin Chen, Dawn Song, and Bo Li · 2023
Later among the works it cites.
Wild patterns reloaded: A survey of machine learning security against training data poisoning
Antonio Emanuele Cinà, Kathrin Grosse, Ambra Demontis, Sebastiano Vascon, Werner Zellinger, Bernhard A Moser, Alina Oprea, Battista Biggio, Marcello Pelillo, and Fabio Roli · 2023
Later among the works it cites.
Label poisoning is all you need
Rishi D Jha, Jonathan Hayase, and Sewoong Oh · 2023
Later among the works it cites.
Pore: Provably robust recommender systems against data poisoning attacks
Jinyuan Jia, Yupei Liu, Yuepeng Hu, and Neil Zhenqiang Gong · 2023
Later among the works it cites.
Data quality detection mechanism against label flipping attacks in federated learning
Yifeng Jiang, Weiwen Zhang, and Yanxi Chen · 2023
Later among the works it cites.
Exclusive: Openai used kenyan workers on less than $2 per hour to make chatgpt less toxic
Billy Perrigo · 2023
Later among the works it cites.
Universal jailbreak backdoors from poisoned human feedback
Javier Rando and Florian Tramèr · 2023
Later among the works it cites.
Django: Detecting trojans in object detection models via gaussian focus calibration
Guangyu Shen, Siyuan Cheng, Guanhong Tao, Kaiyuan Zhang, Yingqi Liu, Shengwei An, Shiqing Ma, and Xiangyu Zhang · 2023
Later among the works it cites.
Badgpt: Exploring security vulnerabilities of chatgpt via backdoor attacks to instructgpt
Jiawen Shi, Yixin Liu, Pan Zhou, and Lichao Sun · 2023
Later among the works it cites.
On the exploitability of instruction tuning
Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, and Tom Goldstein · 2023
Later among the works it cites.
Aligning large multimodal models with factually augmented rlhf
Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liang-Yan Gui, Yu-Xiong Wang, Yiming Yang, et al · 2023
Later among the works it cites.
What distributions are robust to indiscriminate poisoning attacks for linear learners?
Fnu Suya, Xiao Zhang, Yuan Tian, and David Evans · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Poisoning language models during instruction tuning
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein · 2023
Later among the works it cites.
On the exploitability of reinforcement learning with human feedback for large language models
Jiongxiao Wang, Junlin Wu, Muhao Chen, Yevgeniy Vorobeychik, and Chaowei Xiao · 2023
Later among the works it cites.
A brief overview of chatgpt: The history, status quo and potential future development
Tianyu Wu, Shizhu He, Jingping Liu, Siqi Sun, Kang Liu, Qing-Long Han, and Yang Tang · 2023
Later among the works it cites.
Backdooring instruction-tuned large language models with virtual prompt injection
Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin · 2023
Later among the works it cites.
Meta-Sift: How to sift out a clean subset in the presence of data poisoning?
Yi Zeng, Minzhou Pan, Himanshu Jahagirdar, Ming Jin, Lingjuan Lyu, and Ruoxi Jia · 2023
Later among the works it cites.
Poisoning web-scale training datasets is practical
Nicholas Carlini, Matthew Jagielski, Christopher A Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr · 2024
Closest in time.
Lfighter: Defending against the label-flipping attack in federated learning
Najeeb Moharram Jebreel, Josep Domingo-Ferrer, David Sánchez, and Alberto Blanco-Justicia · 2024
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2024
Closest in time.