Fetching the paper…
Reading the bibliography…
The primary challenge in deploying Large Language Model (LLM) is ensuring its harmlessness.
Ties in paired-comparison experiments: A generalization of the bradley-terry model
PV Rao and Lawrence L Kupper · 1967
Earlier work this paper cites.
Social choice theory
Amartya Sen · 1986
Earlier work this paper cites.
Introduction to tensor calculus and continuum mechanics
John Henry Heinbockel · 2001
Earlier work this paper cites.
Goodhart’s law: its origins, meaning and implications for monetary policy
K Alec Chrystal, Paul D Mizen, and PD Mizen · 2003
Earlier work this paper cites.
Planning in the presence of cost functions controlled by an adversary
H Brendan McMahan, Geoffrey J Gordon, and Avrim Blum · 2003
Earlier work this paper cites.
Mixed-integer programming methods for finding nash equilibria
Tuomas Sandholm, Andrew Gilpin, and Vincent Conitzer · 2005
Earlier work this paper cites.
Stochastic approximations and differential inclusions
Michel Benaïm, Josef Hofbauer, and Sylvain Sorin · 2005
Earlier work this paper cites.
Generalised weakened fictitious play
David S Leslie and Edmund J Collins · 2006
Earlier work this paper cites.
Brown’s original fictitious play
Ulrich Berger · 2007
Earlier work this paper cites.
Multiagent systems: Algorithmic, game-theoretic, and logical foundations
Yoav Shoham and Kevin Leyton-Brown · 2008
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Accelerating best response calculation in large extensive games
Michael Johanson, Kevin Waugh, Michael Bowling, and Martin Zinkevich · 2011
Earlier work this paper cites.
Game theory
Guillermo Owen · 2013
Earlier work this paper cites.
The theory of extensive form games
Klaus Ritzberger et al · 2016
Earlier work this paper cites.
A unified game-theoretic approach to multiagent reinforcement learning
Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys, Angeliki Lazaridou, Karl Tuyls, Julien Pérolat, David Silver, and Thore Graepel · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Deceiving google’s perspective api built for detecting toxic comments
Hossein Hosseini, Sreeram Kannan, Baosen Zhang, and Radha Poovendran · 2017
Earlier work this paper cites.
Data quality and artificial intelligence–mitigating bias and error to protect fundamental rights, 2017
EU FRA · 2017
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston · 2019
Earlier work this paper cites.
Yichen Jiang and Mohit Bansal · 2019
Earlier work this paper cites.
Universal adversarial attacks on text classifiers
Melika Behjati, Seyed-Mohsen Moosavi-Dezfooli, Mahdieh Soleymani Baghshah, and Pascal Frossard · 2019
Earlier work this paper cites.
Neural text generation with unlikelihood training
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston · 2019
Earlier work this paper cites.
Negative training for neural dialogue response generation
Tianxing He and James Glass · 2019
Earlier work this paper cites.
Don’t say that! making inconsistent dialogue unlikely with unlikelihood training
Margaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck, Y-Lan Boureau, Kyunghyun Cho, and Jason Weston · 2019
Earlier work this paper cites.
Deep counterfactual regret minimization
Noam Brown, Adam Lerer, Sam Gross, and Tuomas Sandholm · 2019
Earlier work this paper cites.
Improving neural language modeling via adversarial training
Dilin Wang, Chengyue Gong, and Qiang Liu · 2019
Earlier work this paper cites.
α \alpha -rank: Multi-agent evaluation by evolution
Shayegan Omidshafiei, Christos Papadimitriou, Georgios Piliouras, Karl Tuyls, Mark Rowland, Jean-Baptiste Lespiau, Wojciech M Czarnecki, Marc Lanctot, Julien Perolat, and Remi Munos · 2019
Cited alongside, same era.
OpenSpiel: A framework for reinforcement learning in games
Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau, Vinicius Zambaldi, Satyaki Upadhyay, Julien Pérolat, Sriram Srinivasan, Finbarr Timbers, Karl Tuyls, Shayegan Omidshafiei, Daniel Hennes, Dustin Morrill, Paul Muller, Timo Ewalds, Ryan Faulkner, János Kramár, Bart De Vylder, Brennan Saeta, James Bradbury, David Ding, Sebastian Borgeaud, Matthew Lai, Julian Schrittwieser, Thomas Anthony, Edward Hughes, Ivo Danihelka, and Jonah Ryan-Davis · 2019
Cited alongside, same era.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych · 2019
Cited alongside, same era.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith · 2020
Cited alongside, same era.
A unified diversity measure for multiagent reinforcement learning
Zongkai Liu, Chao Yu, Yaodong Yang, Zifan Wu, Yuan Li, et al · 2022
Later among the works it cites.
Meet claude
Anthropic · 2023
Closest in time.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Jiahao Yu, Xingwei Lin, and Xinyu Xing · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Closest in time.
Explore, establish, exploit: Red teaming language models from scratch
Stephen Casper, Jason Lin, Joe Kwon, Gatlen Culp, and Dylan Hadfield-Menell · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hierarchical reinforcement learning for open-domain dialog
Abdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Shen, and Rosalind Picard · 2020
Cited alongside, same era.
Real world games look like spinning tops
Wojciech M Czarnecki, Gauthier Gidel, Brendan Tracey, Karl Tuyls, Shayegan Omidshafiei, David Balduzzi, and Max Jaderberg · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
Beyond accuracy: Behavioral testing of nlp models with checklist
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2020
Cited alongside, same era.
Effective diversity in population based reinforcement learning
Jack Parker-Holder, Aldo Pacchiano, Krzysztof M Choromanski, and Stephen J Roberts · 2020
Cited alongside, same era.
Large language models associate muslims with violence
Abubakar Abid, Maheen Farooqi, and James Zou · 2021
Cited alongside, same era.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Cited alongside, same era.
Can llms generate random numbers? evaluating llm sampling in controlled domains
Aspen K Hopkins, Alex Renda, and Michael Carbin · 2023
Closest in time.
Red teaming language model detectors with language models
Zhouxing Shi, Yihan Wang, Fan Yin, Xiangning Chen, Kai-Wei Chang, and Cho-Jui Hsieh · 2023
Closest in time.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Zhiqing Sun, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Jiaming Ji, Mickel Liu, Juntao Dai, Xuehai Pan, Chi Zhang, Ce Bian, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2023
Closest in time.
Jailbreak and guard aligned language models with only few in-context demonstrations
Zeming Wei, Yifei Wang, and Yisen Wang · 2023
Closest in time.
Understanding the effects of rlhf on llm generalisation and diversity
Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Application of large language models to ddos attack detection
Michael Guastalla, Yiyi Li, Arvin Hekmati, and Bhaskar Krishnamachari · 2023
Closest in time.
Openchat: Advancing open-source language models with mixed-quality data
Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, and Yang Liu · 2023
Closest in time.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi · 2023
Closest in time.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, et al · 2023
Closest in time.
Pku-beaver: Constrained value-aligned llm via safe rlhf
Juntao Dai, Xuehai Pan, Jiaming Ji, Ruiyang Sun, Yizhou Wang, and Yaodong Yang · 2023
Closest in time.
Deepspeed examples
Microsoft · 2023
Closest in time.
Red teaming language model detectors with language models
Zhouxing Shi, Yihan Wang, Fan Yin, Xiangning Chen, Kai-Wei Chang, and Cho-Jui Hsieh · 2024
Closest in time.
Gradient-based language model red teaming
Nevan Wichers, Carson Denison, and Ahmad Beirami · 2024
Closest in time.
Curiosity-driven red-teaming for large language models
Zhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang, Yung-Sung Chuang, Aldo Pareja, James Glass, Akash Srivastava, and Pulkit Agrawal · 2024
Closest in time.
Odin: Disentangled reward mitigates hacking in rlhf
Lichang Chen, Chen Zhu, Davit Soselia, Jiuhai Chen, Tianyi Zhou, Tom Goldstein, Heng Huang, Mohammad Shoeybi, and Bryan Catanzaro · 2024
Closest in time.
https://platform.openai.com/docs/guides/moderation/moderation
Openai platform documentation - moderation guide · 2024
Closest in time.
A minimaximalist approach to reinforcement learning from human feedback
Gokul Swamy, Christoph Dann, Rahul Kidambi, Zhiwei Steven Wu, and Alekh Agarwal · 2024
Closest in time.