Fetching the paper…
Reading the bibliography…
Reinforcement Learning with Human Feedback (RLHF) is a methodology designed to align Large Language Models (LLMs) with human preferences, playing an important role in LLMs alignment.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 1909
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry. 1952 · 1952
Earlier work this paper cites.
Strategy-proofness and pivotal voters: a direct proof of the gibbard-satterthwaite theorem
Salvador Barberá. 1983 · 1983
Earlier work this paper cites.
Bad characters: Imperceptible nlp attacks
Nicholas Boucher, Ilia Shumailov, Ross Anderson, and Nicolas Papernot. 2022 · 2004
Earlier work this paper cites.
Poisoning attacks against support vector machines
Battista Biggio, Blaine Nelson, and Pavel Laskov. 2012 · 2012
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song. 2017 · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Generative poisoning attack method against neural networks
Chaofei Yang, Qing Wu, Hai Li, and Yiran Chen. 2017 · 2017
Earlier work this paper cites.
Reliability and learnability of human bandit feedback for sequence-to-sequence reinforcement learning
Julia Kreutzer, Joshua Uyheng, and Stefan Riezler. 2018 · 2018
Earlier work this paper cites.
Data poisoning attacks in contextual bandits
Yuzhe Ma, Kwang-Sung Jun, Lihong Li, and Xiaojin Zhu. 2018 · 2018
Earlier work this paper cites.
Poison frogs! targeted clean-label poisoning attacks on neural networks
Ali Shafahi, W Ronny Huang, Mahyar Najibi, Octavian Suciu, Christoph Studer, Tudor Dumitras, and Tom Goldstein. 2018 · 2018
Earlier work this paper cites.
Invisible poisoning: Highly stealthy targeted poisoning attack
Jinyin Chen, Haibin Zheng, Mengmeng Su, Tianyu Du, Changting Lin, and Shouling Ji. 2020 · 2019
Earlier work this paper cites.
A backdoor attack against lstm-based text classification systems
Jiazhu Dai, Chuanshuai Chen, and Yufeng Li. 2019 · 2019
Cited alongside, same era.
Badnets: Evaluating backdooring attacks on deep neural networks
Tianyu Gu, Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg. 2019 · 2019
Cited alongside, same era.
Policy poisoning in batch reinforcement learning and control
Yuzhe Ma, Xuezhou Zhang, Wen Sun, and Jerry Zhu. 2019 · 2019
Cited alongside, same era.
Weight poisoning attacks on pretrained models
Keita Kurita, Paul Michel, and Graham Neubig. 2020 · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020 · 2020
Cited alongside, same era.
Adaptive reward-poisoning attacks against reinforcement learning
Xuezhou Zhang, Yuzhe Ma, Adish Singla, and Xiaojin Zhu. 2020 · 2020
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023 · 2023
Closest in time.
Safe rlhf: Safe reinforcement learning from human feedback
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang. 2023 · 2023
Closest in time.
Alpacaeval: An automatic evaluator of instruction-following models
Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Universal jailbreak backdoors from poisoned human feedback
Javier Rando and Florian Tramèr. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Exploring data and model poisoning attacks to deep learning-based nlp systems
Fiammetta Marulli, Laura Verde, and Lelio Campanile. 2021 · 2021
Cited alongside, same era.
Webgpt: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al. 2021 · 2021
Cited alongside, same era.
Concealed data poisoning attacks on nlp models
Eric Wallace, Tony Zhao, Shi Feng, and Sameer Singh. 2021 · 2021
Cited alongside, same era.
Putting words into the system’s mouth: A targeted attack on neural machine translation using monolingual data poisoning
Jun Wang, Chang Xu, Francisco Guzmán, Ahmed El-Kishky, Yuqing Tang, Benjamin Rubinstein, and Trevor Cohn. 2021 · 2021
Cited alongside, same era.
A targeted attack on black-box neural machine translation with parallel data poisoning
Chang Xu, Jun Wang, Yuqing Tang, Francisco Guzmán, Benjamin IP Rubinstein, and Trevor Cohn. 2021 · 2021
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
Closest in time.
Poster: Badgpt: Exploring security vulnerabilities of chatgpt via backdoor attacks to instructgpt
Jiawen Shi, Yixin Liu, Pan Zhou, and Lichao Sun. 2023 · 2023
Closest in time.
On the exploitability of instruction tuning
Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, and Tom Goldstein. 2023 · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Closest in time.
Poisoning language models during instruction tuning
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. 2023 · 2023
Closest in time.
Backdooring instruction-tuned large language models with virtual prompt injection
Jun Yan, Vikas Yadav, Shiyang Li, Lichang Chen, Zheng Tang, Hai Wang, Vijay Srinivasan, Xiang Ren, and Hongxia Jin. 2023 · 2023
Closest in time.
Secrets of rlhf in large language models part i: Ppo
Rui Zheng, Shihan Dou, Songyang Gao, Wei Shen, Binghai Wang, Yan Liu, Senjie Jin, Qin Liu, Limao Xiong, Lu Chen, et al. 2023 · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2024 · 2024
Closest in time.