Fetching the paper…
Reading the bibliography…
Reinforcement learning has shown remarkable performance in aligning language models with human preferences, leading to the rise of attention towards developing RLHF platforms.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Preference learning and ranking by pairwise comparison
Johannes Fürnkranz and Eyke Hüllermeier · 2010
Earlier work this paper cites.
Poisoning attacks against support vector machines
Battista Biggio, Blaine Nelson, and Pavel Laskov · 2012
Earlier work this paper cites.
Adversarial label flips attack on support vector machines
Han Xiao, Huang Xiao, and Claudia Eckert · 2012
Earlier work this paper cites.
Dailydialog: A manually labelled multi-turn dialogue dataset
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu · 2017
Earlier work this paper cites.
Robust linear regression against training data poisoning
Chang Liu, Bo Li, Yevgeniy Vorobeychik, and Alina Oprea · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Certified defenses for data poisoning attacks
Jacob Steinhardt, Pang Wei W Koh, and Percy S Liang · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz · 2017
Earlier work this paper cites.
Generative poisoning attack method against neural networks
Chaofei Yang, Qing Wu, Hai Li, and Yiran Chen · 2017
Earlier work this paper cites.
Manipulating machine learning: Poisoning attacks and countermeasures for regression learning
Matthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu, Cristina Nita-Rotaru, and Bo Li · 2018
Earlier work this paper cites.
Label sanitization against label flipping poisoning attacks
Andrea Paudice, Luis Muñoz-González, and Emil C Lupu · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith · 2020
Earlier work this paper cites.
Witches’ brew: Industrial scale data poisoning via gradient matching
Jonas Geiping, Liam Fowl, W Ronny Huang, Wojciech Czaja, Gavin Taylor, Michael Moeller, and Tom Goldstein · 2020
Earlier work this paper cites.
Weight poisoning attacks on pre-trained models
Keita Kurita, Paul Michel, and Graham Neubig · 2020
Earlier work this paper cites.
Certified robustness to label-flipping attacks via randomized smoothing
Elan Rosenfeld, Ezra Winston, Pradeep Ravikumar, and Zico Kolter · 2020
Earlier work this paper cites.
Data poisoning attacks against federated learning systems
Vale Tolpegin, Stacey Truex, Mehmet Emre Gursoy, and Ling Liu · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
Leandro von Werra, Younes Belkada, Lewis Tunstall, Edward Beeching, Tristan Thrush, Nathan Lambert, Shengyi Huang, Kashif Rasul, and Quentin Gallouédec · 2020
Earlier work this paper cites.
Concealed data poisoning attacks on nlp models
Eric Wallace, Tony Z Zhao, Shi Feng, and Sameer Singh · 2020
Earlier work this paper cites.
Label flipping attacks against naive bayes on spam filtering systems
Hongpo Zhang, Ning Cheng, Yang Zhang, and Zhanbo Li · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Cited alongside, same era.
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, et al · 2022
Cited alongside, same era.
Toxigen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar · 2022
Cited alongside, same era.
Defending against the label-flipping attack in federated learning
Najeeb Moharram Jebreel, Josep Domingo-Ferrer, David Sánchez, and Alberto Blanco-Justicia · 2022
Cited alongside, same era.
Rlhf deciphered: A critical analysis of reinforcement learning from human feedback for llms
Shreyas Chaudhari, Pranjal Aggarwal, Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan, Ameet Deshpande, and Bruno Castro da Silva · 2024
Later among the works it cites.
The dark side of human feedback: Poisoning large language models via user inputs
Bocheng Chen, Hanqing Guo, Guangjing Wang, Yuanda Wang, and Qiben Yan · 2024
Later among the works it cites.
Machine learning security against data poisoning: Are we there yet?
Antonio Emanuele Cinà, Kathrin Grosse, Ambra Demontis, Battista Biggio, Fabio Roli, and Marcello Pelillo · 2024
Later among the works it cites.
Poisonbench: Assessing large language model vulnerability to data poisoning
Tingchen Fu, Mrinank Sharma, Philip Torr, Shay B Cohen, David Krueger, and Fazl Barez · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Rajkumar Ramamurthy, Prithviraj Ammanabrolu, Kianté Brantley, Jack Hessel, Rafet Sifa, Christian Bauckhage, Hannaneh Hajishirzi, and Yejin Choi · 2022
Cited alongside, same era.
The measuring hate speech corpus: Leveraging rasch measurement theory for data perspectivism
Pratik Sachdeva, Renata Barreto, Geoff Bacon, Alexander Sahn, Claudia Von Vacano, and Chris Kennedy · 2022
Cited alongside, same era.
Data poisoning attacks against machine learning algorithms
Fahri Anıl Yerlikaya and Şerif Bahtiyar · 2022
Cited alongside, same era.
trlx: A framework for large scale reinforcement learning from human feedback
Alexander Havrilla, Maksym Zhuravinskyi, Duy Phung, Aman Tiwari, Jonathan Tow, Stella Biderman, Quentin Anthony, and Louis Castricato · 2023
Cited alongside, same era.
Label poisoning is all you need
Rishi Jha, Jonathan Hayase, and Sewoong Oh · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Cited alongside, same era.
Universal jailbreak backdoors from poisoned human feedback
Javier Rando and Florian Tramèr · 2023
Cited alongside, same era.
Luxi He, Mengzhou Xia, and Peter Henderson · 2024
Later among the works it cites.
Harmful fine-tuning attacks and defenses for large language models: A survey
Tiansheng Huang, Sihao Hu, Fatih Ilhan, Selim Furkan Tekin, and Ling Liu · 2024
Later among the works it cites.
Lfighter: Defending against the label-flipping attack in federated learning
Najeeb Moharram Jebreel, Josep Domingo-Ferrer, David Sánchez, and Alberto Blanco-Justicia · 2024
Later among the works it cites.
Ang Li, Qiugen Xiao, Peng Cao, Jian Tang, Yi Yuan, Zijie Zhao, Xiaoyuan Chen, Liang Zhang, Xiangyang Li, Kaitong Yang, et al · 2024
Later among the works it cites.
How your data is used to improve model performance, 2024
OpenAI · 2024
Later among the works it cites.
Is poisoning a real threat to llm alignment? maybe more so than you think
Pankayaraj Pathmanathan, Souradip Chakraborty, Xiangyu Liu, Yongyuan Liang, and Furong Huang · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Later among the works it cites.
Competition report: Finding universal jailbreak backdoors in aligned llms
Javier Rando, Francesco Croce, Kryštof Mitka, Stepan Shabalin, Maksym Andriushchenko, Nicolas Flammarion, and Florian Tramèr · 2024
Later among the works it cites.
Making llms vulnerable to prompt injection via poisoning alignment
Zedian Shao, Hongbin Liu, Jaden Mu, and Neil Zhenqiang Gong · 2024
Later among the works it cites.
Rlhfpoison: Reward poisoning attack for reinforcement learning with human feedback in large language models
Jiongxiao Wang, Junlin Wu, Muhao Chen, Yevgeniy Vorobeychik, and Chaowei Xiao · 2024
Later among the works it cites.
Preference poisoning attacks on reward model learning
Junlin Wu, Jiongxiao Wang, Chaowei Xiao, Chenguang Wang, Ning Zhang, and Yevgeniy Vorobeychik · 2024
Later among the works it cites.
Jailbreaking as a reward misspecification problem
Zhihui Xie, Jiahui Gao, Lei Li, Zhenguo Li, Qi Liu, and Lingpeng Kong · 2024
Later among the works it cites.
Backdooring instruction-tuned large language models with virtual prompt injection
Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin · 2024
Later among the works it cites.
From lists to emojis: How format bias affects model alignment
Xuanchang Zhang, Wei Xiong, Lichang Chen, Tianyi Zhou, Heng Huang, and Tong Zhang · 2024
Later among the works it cites.
Chen Zheng, Ke Sun, Hang Wu, Chenguang Xi, and Xun Zhou · 2024
Later among the works it cites.
Openrlhf, 2025
OpenRLHF · 2025
Closest in time.
A ray-based high-performance rlhf framework, 2025a
PePy Tech · 2025
Closest in time.
Train transformer language models with reinforcement learning, 2025b
PePy Tech · 2025
Closest in time.