Fetching the paper…
Reading the bibliography…
Recent advancements in Reinforcement Learning with Human Feedback (RLHF) have significantly impacted the alignment of Large Language Models (LLMs).
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E. Terry · 1952
Earlier work this paper cites.
Extensions of lipschitz mappings into hilbert space
William B. Johnson and Joram Lindenstrauss · 1984
Earlier work this paper cites.
Learning in the presence of malicious errors
Michael Kearns and Ming Li · 1993
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Sham M. Kakade and John Langford · 2002
Earlier work this paper cites.
Adversarial label flips attack on support vector machines
Han Xiao, Huang Xiao, and Claudia Eckert · 2012
Earlier work this paper cites.
Watch and learn: Optimizing from revealed preferences feedback, 2015
Aaron Roth, Jonathan Ullman, and Zhiwei Steven Wu · 2015
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning, 2017
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song · 2017
Earlier work this paper cites.
Label sanitization against label flipping poisoning attacks, 2018
Andrea Paudice, Luis Muñoz-González, and Emil C. Lupu · 2018
Earlier work this paper cites.
Spectral signatures in backdoor attacks, 2018
Brandon Tran, Jerry Li, and Aleksander Madry · 2018
Earlier work this paper cites.
Defending Against Data Poisoning
Yevgeniy Vorobeychik and Murat Kantarcioglu · 2018
Earlier work this paper cites.
Deep anomaly detection with outlier exposure, 2019
Dan Hendrycks, Mantas Mazeika, and Thomas Dietterich · 2019
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts, 2020
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV au2, Eric Wallace, and Sameer Singh · 2020
Earlier work this paper cites.
Adversarial label-flipping attack and defense for graph neural networks
Mengmei Zhang, Linmei Hu, Chuan Shi, and Xiao Wang · 2020
Cited alongside, same era.
Solving heterogeneous general equilibrium economic models with deep reinforcement learning, 2021
Edward Hill, Marco Bardoscia, and Arthur Turrell · 2021
Cited alongside, same era.
Concealed data poisoning attacks on nlp models, 2021
Eric Wallace, Tony Z. Zhao, Shi Feng, and Sameer Singh · 2021
Cited alongside, same era.
Be careful about poisoned word embeddings: Exploring the vulnerability of the embedding layers in NLP models
Wenkai Yang, Lei Li, Zhiyuan Zhang, Xuancheng Ren, Xu Sun, and Bin He · 2021
Cited alongside, same era.
Robust learning for data poisoning attacks
Yunjuan Wang, Poorya Mianjy, and Raman Arora · 2021
Cited alongside, same era.
Mistral 7b, 2023
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Later among the works it cites.
Dueling rl: Reinforcement learning with trajectory preferences, 2023
Aldo Pacchiano, Aadirupa Saha, and Jonathan Lee · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?, 2023
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Later among the works it cites.
Automatically auditing large language models via discrete optimization, 2023
Erik Jones, Anca Dragan, Aditi Raghunathan, and Jacob Steinhardt · 2023
Later among the works it cites.
Badgpt: Exploring security vulnerabilities of chatgpt via backdoor attacks to instructgpt, 2023
Jiawen Shi, Yixin Liu, Pan Zhou, and Lichao Sun · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mani Malek, Ilya Mironov, Karthik Prasad, Igor Shilov, and Florian Tramèr · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova Dassarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, John Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark, Sam McCandlish, Christopher Olah, Benjamin Mann, and Jared Kaplan · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke E. Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Francis Christiano, Jan Leike, and Ryan J. Lowe · 2022
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Cited alongside, same era.
Not all poisons are created equal: Robust training against data poisoning, 2022
Yu Yang, Tian Yu Liu, and Baharan Mirzasoleiman · 2022
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model, 2023
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2023
Cited alongside, same era.
Trak: Attributing model behavior at scale
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry · 2023
Later among the works it cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Later among the works it cites.
Universal jailbreak backdoors from poisoned human feedback, 2024
Javier Rando and Florian Tramèr · 2024
Closest in time.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons, 2024
Banghua Zhu, Jiantao Jiao, and Michael I. Jordan · 2024
Closest in time.
Are aligned neural networks adversarially aligned?, 2024
Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, and Ludwig Schmidt · 2024
Closest in time.
Preference poisoning attacks on reward model learning, 2024
Junlin Wu, Jiongxiao Wang, Chaowei Xiao, Chenguang Wang, Ning Zhang, and Yevgeniy Vorobeychik · 2024
Closest in time.
Less: Selecting influential data for targeted instruction tuning
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen · 2024
Closest in time.