Fetching the paper…
Reading the bibliography…
Reinforcement learning from human feedback (RLHF) is a technique for training AI systems to align with human goals.
The Assistive Multi-Armed Bandit
Lawrence Chan, Dylan Hadfield-Menell, Siddhartha Srinivasa, and Anca Dragan · 1901
Earlier work this paper cites.
Tom Everitt, Marcus Hutter, Ramana Kumar, and Victoria Krakovna · 1908
Earlier work this paper cites.
The Belmont report: ethical principles and guidelines for the protection of human subjects of research , volume 1
United States National Commission for the Protection of Human Subjects · 1978
Earlier work this paper cites.
Social choice theory
Amartya Sen · 1986
Earlier work this paper cites.
The ‘awful idea of accountability’: inscribing people into the measurement of objects
Keith Hoskin · 1996
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Andrew Y Ng, Stuart Russell, et al · 2000
Earlier work this paper cites.
Quantifying differences in reward functions
Adam Gleave, Michael Dennis, Shane Legg, Stuart Russell, and Jan Leike · 2006
Earlier work this paper cites.
The netflix prize
James Bennett, Stan Lanning, et al · 2007
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Deepak Ramachandran and Eyal Amir · 2007
Earlier work this paper cites.
Fda and clinical drug trials: a short history
Suzanne Junod · 2008
Earlier work this paper cites.
Tamer: Training an agent manually via evaluative reinforcement
W Bradley Knox and Peter Stone · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Brian D Ziebart, Andrew L Maas, J Andrew Bagnell, Anind K Dey, et al · 2008
Earlier work this paper cites.
Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in ai
Alon Jacovi, Ana Marasović, Tim Miller, and Yoav Goldberg · 2010
Earlier work this paper cites.
Leviathan and the air-pump: Hobbes, Boyle, and the experimental life
Steven Shapin and Simon Schaffer · 2011
Earlier work this paper cites.
Ranking vs. preference: a comparative study of self-reporting
Georgios N Yannakakis and John Hallam · 2011
Earlier work this paper cites.
Bpr: Bayesian personalized ranking from implicit feedback
Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme · 2012
Earlier work this paper cites.
Machine learning techniques for anomaly detection: an overview
Salima Omar, Asri Ngadi, and Hamid H Jebur · 2013
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Revealed preference theory , volume 56
Christopher P Chambers and Federico Echenique · 2016
Earlier work this paper cites.
Cooperative inverse reinforcement learning
Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan · 2016
Earlier work this paper cites.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Learning mixtures of plackett-luce models
Zhibing Zhao, Peter Piech, and Lirong Xia · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan · 2017
Earlier work this paper cites.
Delving into adversarial attacks on deep policies
Jernej Kos and Dawn Song · 2017
Earlier work this paper cites.
Interactive learning from policy-dependent human feedback
James MacGlashan, Mark K Ho, Robert Loftin, Bei Peng, Guan Wang, David L Roberts, Matthew E Taylor, and Michael L Littman · 2017
Earlier work this paper cites.
Reinforcement learning for bandit neural machine translation with simulated human feedback
Khanh Nguyen, Hal Daumé III, and Jordan Boyd-Graber · 2017
Earlier work this paper cites.
Active preference-based learning of reward functions
Dorsa Sadigh, Anca D Dragan, Shankar Sastry, and Sanjit A Seshia · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, Johannes Fürnkranz, et al · 2017
Earlier work this paper cites.
Learning from physical human corrections, one feature at a time
Andrea Bajcsy, Dylan P Losey, Marcia K O’Malley, and Anca D Dragan · 2018
Earlier work this paper cites.
Batch active preference-based learning of reward functions
Erdem Biyik and Dorsa Sadigh · 2018
Earlier work this paper cites.
Ai governance: a research agenda
Allan Dafoe · 2018
Earlier work this paper cites.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Earlier work this paper cites.
Reward learning from human preferences and demonstrations in atari
Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei · 2018
Earlier work this paper cites.
Deep reinforcement learning doesn’t work yet
Alex Irpan · 2018
Earlier work this paper cites.
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Safe reinforcement learning via probabilistic shields
Nils Jansen, Bettina Könighofer, Sebastian Junges, Alexandru C Serban, and Roderick Bloem · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Earlier work this paper cites.
Categorizing variants of goodhart’s law
David Manheim and Scott Garrabrant · 2018
Earlier work this paper cites.
Occam’s razor is insufficient to infer the preferences of irrational agents
Soren Mindermann and Stuart Armstrong · 2018
Earlier work this paper cites.
Improving stability in deep reinforcement learning with weight averaging
Evgenii Nikishin, Pavel Izmailov, Ben Athiwaratkun, Dmitrii Podoprikhin, Timur Garipov, Pavel Shvechikov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Earlier work this paper cites.
A voting-based system for ethical decision making
Ritesh Noothigattu, Snehalkumar Gaikwad, Edmond Awad, Sohan Dsouza, Iyad Rahwan, Pradeep Ravikumar, and Ariel Procaccia · 2018
Earlier work this paper cites.
Deep reinforcement learning from policy-dependent human feedback
Dilip Arumugam, Jun Ki Lee, Sophie Saskin, and Michael L Littman · 2019
Earlier work this paper cites.
Asking easy questions: A user-friendly approach to active reward learning
Erdem Biyik, Malayandi Palan, Nicholas C. Landolfi, Dylan P. Losey, and Dorsa Sadigh · 2019
Earlier work this paper cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
Daniel Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum · 2019
Earlier work this paper cites.
Scaling data-driven robotics with reward sketching and batch reinforcement learning
Serkan Cabi, Sergio Gómez Colmenarejo, Alexander Novikov, Ksenia Konyushkova, Scott Reed, Rae Jeong, Konrad Zolna, Yusuf Aytar, David Budden, Mel Vecerik, et al · 2019
Earlier work this paper cites.
Worst-case guarantees
Paul Christiano · 2019
Earlier work this paper cites.
Standards for ai governance: international standards to enable global coordination in ai research & development
Peter Cihon · 2019
Earlier work this paper cites.
The mandela effect and new memory
Aaron French · 2019
Earlier work this paper cites.
From language to goals: Inverse reinforcement learning for vision-based instruction following
Justin Fu, Anoop Korattikara, Sergey Levine, and Sergio Guadarrama · 2019
Earlier work this paper cites.
Using natural language for reward shaping in reinforcement learning
Prasoon Goyal, Scott Niekum, and Raymond J Mooney · 2019
Earlier work this paper cites.
A survey of reinforcement learning informed by natural language
Jelena Luketina, Nantas Nardelli, Gregory Farquhar, Jakob Foerster, Jacob Andreas, Edward Grefenstette, Shimon Whiteson, and Tim Rocktäschel · 2019
Earlier work this paper cites.
What you see is what you get? the impact of representation criteria on human bias in hiring
Andi Peng, Besmira Nushi, Emre Kıcıman, Kori Inkpen, Siddharth Suri, and Ece Kamar · 2019
Earlier work this paper cites.
Ai governance and the policymaking process: key considerations for reducing ai risk
Brandon Perry and Risto Uuk · 2019
Earlier work this paper cites.
Where Do You Think You’re Going?: Inferring Beliefs about Dynamics from Behavior
Siddharth Reddy, Anca D. Dragan, and Sergey Levine · 2019
Earlier work this paper cites.
On the feasibility of learning, rather than assuming, human biases for reward inference
Rohin Shah, Noah Gundotra, Pieter Abbeel, and Anca Dragan · 2019
Earlier work this paper cites.
Optimal policies tend to seek power
Alexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch, and Prasad Tadepalli · 2019
Earlier work this paper cites.
Adversarial examples: Opportunities and challenges
Jiliang Zhang and Chen Li · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Preference learning along multiple criteria: A game-theoretic perspective
Kush Bhatia, Ashwin Pananjady, Peter Bartlett, Anca Dragan, and Martin J Wainwright · 2020
Earlier work this paper cites.
Active preference-based gaussian process regression for reward learning
Erdem Biyik, Nicolas Huynh, Mykel J. Kochenderfer, and Dorsa Sadigh · 2020
Earlier work this paper cites.
Quantifying hypothesis space misspecification in learning from human–robot demonstrations and physical corrections
Andreea Bobu, Andrea Bajcsy, Jaime F Fisac, Sampada Deglurkar, and Anca D Dragan · 2020
Earlier work this paper cites.
Why ai chips matter
CSET Policy Brief · 2020
Earlier work this paper cites.
Safe imitation learning via fast bayesian reward inference from preferences
Daniel Brown, Russell Coleman, Ravi Srinivasan, and Scott Niekum · 2020
Earlier work this paper cites.
Achilles heels for agi/asi via decision theoretic adversaries
Stephen Casper · 2020
Earlier work this paper cites.
An mturk crisis? shifts in data quality and the impact on study results
Michael Chmielewski and Sarah C Kucker · 2020
Earlier work this paper cites.
Curiosity killed the cat and the asymptotically optimal agent
Michael K Cohen and Marcus Hutter · 2020
Earlier work this paper cites.
Ai research considerations for human existential safety (arches)
Andrew Critch and David Krueger · 2020
Earlier work this paper cites.
Challenges of reinforcement learning
Zihan Ding and Hao Dong · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
Understanding rl vision
Jacob Hilton, Nick Cammarata, Shan Carter, Gabriel Goh, and Chris Olah · 2020
Earlier work this paper cites.
An overview of 11 proposals for building safe advanced ai
Evan Hubinger · 2020
Earlier work this paper cites.
Reward-rational (implicit) choice: A unifying formalism for reward learning
Hong Jun Jeon, Smitha Milli, and Anca Dragan · 2020
Earlier work this paper cites.
Specification gaming: the flip side of ai ingenuity
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Earlier work this paper cites.
Hidden incentives for auto-induced distributional shift, 2020
David Krueger, Tegan Maharaj, and Jan Leike · 2020
Earlier work this paper cites.
Transparency in artificial intelligence
Stefan Larsson and Fredrik Heintz · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu · 2020
Earlier work this paper cites.
Understanding learned reward functions
Eric J Michaud, Adam Gleave, and Stuart Russell · 2020
Earlier work this paper cites.
Literal or pedagogic human? analyzing human model misspecification in objective learning
Smitha Milli and Anca D Dragan · 2020
Earlier work this paper cites.
The windfall clause: Distributing the benefits of ai for the common good
Cullen O’Keefe, Peter Cihon, Ben Garfinkel, Carrick Flynn, Jade Leung, and Allan Dafoe · 2020
Earlier work this paper cites.
Assisted Perception: Optimizing Observations to Communicate State
Siddharth Reddy, Sergey Levine, and Anca D Dragan · 2020
Earlier work this paper cites.
Learning to be safe: Deep rl with a safety critic
Krishnan Srinivasan, Benjamin Eysenbach, Sehoon Ha, Jie Tan, and Chelsea Finn · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
Consequences of misaligned ai
Simon Zhuang and Dylan Hadfield-Menell · 2020
Cited alongside, same era.
A survey of exploration methods in reinforcement learning
Susan Amin, Maziar Gomrokchi, Harsh Satija, Herke van Hoof, and Doina Precup · 2021
Cited alongside, same era.
Why ai alignment could be hard with modern deep learning
Ajeya Cotra · 2021
Cited alongside, same era.
Hard choices in artificial intelligence
Roel Dobbe, Thomas Krendl Gilbert, and Yonatan Mintz · 2021
Cited alongside, same era.
Governing ai safety through independent audits
Gregory Falco, Ben Shneiderman, Julia Badger, Ryan Carrier, Anton Dahbura, David Danks, Martin Eling, Alwyn Goodloe, Jerry Gupta, Christopher Hart, et al · 2021
Adversarial training for high-stakes reliability
Daniel Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman, Peter Schmidt-Nielsen, Tao Lin, Adam Scherlis, Noa Nabeshima, Benjamin Weinstein-Raun, Daniel de Haas, et al · 2022
Later among the works it cites.
Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs
Afra Feyza Akyürek, Ekin Akyürek, Aman Madaan, Ashwin Kalyan, Peter Clark, Derry Wijaya, and Niket Tandon · 2023
Closest in time.
Jailbreak chat
Alex Albert · 2023
Closest in time.
Frontier ai regulation: Managing emerging risks to public safety, 2023
Markus Anderljung, Joslyn Barnhart, Jade Leung, Anton Korinek, Cullen O’Keefe, Jess Whittlestone, Shahar Avin, Miles Brundage, Justin Bullock, Duncan Cass-Beggs, Ben Chang, Tantum Collins, Tim Fist, Gillian Hadfield, Alan Hayes, Lewis Ho, Sara Hooker, Eric Horvitz, Noam Kolt, Jonas Schuett, Yonadav Shavit, Divya Siddarth, Robert Trager, and Kevin Wolf · 2023
Closest in time.
Introducing claude, 2023
Anthropic · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Choice set misspecification in reward inference
Rachel Freedman, Rohin Shah, and Anca Dragan · 2021
Cited alongside, same era.
Subjectifying objectivity: Delineating tastes in theoretical quantum gravity research
Thomas Krendl Gilbert and Andrew Loveridge · 2021
Cited alongside, same era.
The disagreement deconvolution: Bringing machine learning performance metrics in line with reality
Mitchell L Gordon, Kaitlyn Zhou, Kayur Patel, Tatsunori Hashimoto, and Michael S Bernstein · 2021
Cited alongside, same era.
Unsolved problems in ml safety
Dan Hendrycks, Nicholas Carlini, John Schulman, and Jacob Steinhardt · 2021
Cited alongside, same era.
Alignment of Language Agents, March 2021
Zachary Kenton, Tom Everitt, Laura Weidinger, Iason Gabriel, Vladimir Mikulik, and Geoffrey Irving · 2021
Cited alongside, same era.
A distributional approach to controlled text generation
Muhammad Khalifa, Hady Elsahar, and Marc Dymetman · 2021
Cited alongside, same era.
Artificial Intelligence Can Persuade Humans
Hui Bai · 2023
Closest in time.
Active reward learning from multiple teachers
Peter Barnett, Rachel Freedman, Justin Svegliato, and Stuart Russell · 2023
Closest in time.
Which examples should be multiply annotated? active learning when annotators may disagree
Connor Baumler, Anna Sotnikova, and Hal Daumé III · 2023
Closest in time.
Aligning robot and human representations
Andreea Bobu, Andi Peng, Pulkit Agrawal, Julie Shah, and Anca D Dragan · 2023
Closest in time.
Settling the reward hypothesis
Michael Bowling, John D Martin, David Abel, and Will Dabney · 2023
Closest in time.
Characterizing Manipulation from AI Systems, March 2023
Micah Carroll, Alan Chan, Henry Ashton, and David Krueger · 2023
Closest in time.
Improving code generation by training with natural language feedback
Angelica Chen, Jérémy Scheurer, Tomasz Korbak, Jon Ander Campos, Jun Shern Chan, Samuel R Bowman, Kyunghyun Cho, and Ethan Perez · 2023
Closest in time.
Thoughts on the impact of rlhf research, Jan 2023
Paul Christiano · 2023
Closest in time.
Raft: Reward ranked finetuning for generative foundation model alignment
Hanze Dong, Wei Xiong, Deepanshu Goyal, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang · 2023
Closest in time.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch · 2023
Closest in time.
Gpts are gpts: An early look at the labor market impact potential of large language models, 2023
Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock · 2023
Closest in time.
Moral machine or tyranny of the majority?
Michael Feffer, Hoda Heidari, and Zachary C Lipton · 2023
Closest in time.
"federal trade commission civil investigative demand schedule ftc file no. 232-3044", July 2023
FTC · 2023
Closest in time.
Causal abstraction for faithful model interpretation
Atticus Geiger, Chris Potts, and Thomas Icard · 2023
Closest in time.
Chatgpt outperforms crowd-workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli · 2023
Closest in time.
Aligning language models with preferences through f-divergence minimization, 2023
Dongyoung Go, Tomasz Korbak, Germán Kruszewski, Jos Rozen, Nahyeon Ryu, and Marc Dymetman · 2023
Closest in time.
Bard, 2023
Google · 2023
Closest in time.
Susceptibility to Influence of Large Language Models, March 2023
Lewis D. Griffin, Bennett Kleinberg, Maximilian Mozes, Kimberly T. Mai, Maria Vau, Matthew Caldwell, and Augustine Marvor-Parker · 2023
Closest in time.
Ground (less) truth: A causal framework for proxy labels in human-algorithm decision-making
Luke Guerdan, Amanda Coston, Zhiwei Steven Wu, and Kenneth Holstein · 2023
Closest in time.
Regulatory markets: The future of ai governance
Gillian K Hadfield and Jack Clark · 2023
Closest in time.
The hidden workforce that helped filter violence and abuse out of chatgpt, 2023
Karen Hao · 2023
Closest in time.
Jochen Hartmann, Jasper Schwenzow, and Maximilian Witte · 2023
Closest in time.
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun · 2023
Closest in time.
Measuring and manipulating knowledge representations in language models
Evan Hernandez, Belinda Z Li, and Jacob Andreas · 2023
Closest in time.
Aligning language models with offline reinforcement learning from human feedback
Jian Hu, Li Tao, June Yang, and Chandler Zhou · 2023
Closest in time.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Closest in time.
Toward comprehensive risk assessments and assurance of ai-based systems
Heidy Khlaaf · 2023
Closest in time.
Aligning large language models through synthetic feedback
Sungdong Kim, Sanghwan Bae, Jamin Shin, Soyoung Kang, Donghyun Kwak, Kang Min Yoo, and Minjoon Seo · 2023
Closest in time.
Hannah Rose Kirk, Bertie Vidgen, Paul Röttger, and Scott A Hale · 2023
Closest in time.
Pretraining language models with human preferences, 2023
Tomasz Korbak, Kejian Shi, Angelica Chen, Rasika Bhalerao, Christopher L. Buckley, Jason Phang, Samuel R. Bowman, and Ethan Perez · 2023
Closest in time.
Power-seeking can be probable and predictive for trained agents
Victoria Krakovna and Janos Kramar · 2023
Closest in time.
Rlaif: Scaling reinforcement learning from human feedback with ai feedback, 2023
Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Lu, Thomas Mesnard, Colton Bishop, Victor Carbune, and Abhinav Rastogi · 2023
Closest in time.
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Closest in time.
Learning safety constraints from demonstrations with unknown rewards
David Lindner, Xin Chen, Sebastian Tschiatschek, Katja Hofmann, and Andreas Krause · 2023
Closest in time.
Jailbreaking chatgpt via prompt engineering: An empirical study, 2023
Yi Liu, Gelei Deng, Zhengzi Xu, Yuekang Li, Yaowen Zheng, Ying Zhang, Lida Zhao, Tianwei Zhang, and Yang Liu · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2023
Closest in time.
On the fragility of learned reward functions
Lev McKinney, Yawen Duan, David Krueger, and Adam Gleave · 2023
Closest in time.
Locating and editing factual associations in gpt, 2023
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2023
Closest in time.
Auditing large language models: a three-layered approach
Jakob Mökander, Jonas Schuett, Hannah Rose Kirk, and Luciano Floridi · 2023
Closest in time.
Chat gpt "dan" (and other "jailbreaks")
A.J. Oneal · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Diagnosis, feedback, adaptation: A human-in-the-loop framework for test-time policy adaptation
Andi Peng, Aviv Netanyahu, Mark K Ho, Tianmin Shu, Andreea Bobu, Julie Shah, and Pulkit Agrawal · 2023
Closest in time.
Exclusive: The $2 per hour workers who made chatgpt safer, 2023
Billy Perrigo · 2023
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D Manning, and Chelsea Finn · 2023
Closest in time.
Alexandre Rame, Guillaume Couairon, Mustafa Shukor, Corentin Dancette, Jean-Baptiste Gaya, Laure Soulier, and Matthieu Cord · 2023
Closest in time.
Tricking llms into disobedience: Understanding, analyzing, and preventing jailbreaks, 2023
Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya, and Monojit Choudhury · 2023
Closest in time.
Supporting human-ai collaboration in auditing llms with llms
Charvi Rastogi, Marco Tulio Ribeiro, Nicholas King, and Saleema Amershi · 2023
Closest in time.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell · 2023
Closest in time.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto · 2023
Closest in time.
Training language models with language feedback at scale
Jérémy Scheurer, Jon Ander Campos, Tomasz Korbak, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, and Ethan Perez · 2023
Closest in time.
What does it take to catch a chinchilla? verifying rules on large-scale neural network training via compute monitoring, 2023
Yonadav Shavit · 2023
Closest in time.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2023
Closest in time.
Model evaluation for extreme risks
Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, et al · 2023
Closest in time.
Fairness in preference-based reinforcement learning, 2023
Umer Siddique, Abhinav Sinha, and Yongcan Cao · 2023
Closest in time.
Invariance in policy optimisation and partial identifiability in reward learning
Joar Max Viktor Skalse, Matthew Farrugia-Roberts, Stuart Russell, Alessandro Abate, and Adam Gleave · 2023
Closest in time.
Reward collapse in aligning large language models
Ziang Song, Tianle Cai, Jason D Lee, and Weijie J Su · 2023
Closest in time.
Emergent Deception and Emergent Optimization, February 2023
Jacob Steinhardt · 2023
Closest in time.
Towards Modeling and Influencing the Dynamics of Human Learning, January 2023
Ran Tian, Masayoshi Tomizuka, Anca Dragan, and Andrea Bajcsy · 2023
Closest in time.
Causal confusion and reward misidentification in preference-based reward learning
Jeremy Tien, Jerry Zhi-Yang He, Zackory Erickson, Anca Dragan, and Daniel S Brown · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Closest in time.
Survey on reinforcement learning for language processing
Victor Uc-Cetina, Nicolas Navarro-Guerrero, Anabel Martin-Gonzalez, Cornelius Weber, and Stefan Wermter · 2023
Closest in time.
Veniamin Veselovsky, Manoel Horta Ribeiro, and Robert West · 2023
Closest in time.
Microsoft’s Bing is an emotionally manipulative liar, and people love it, February 2023
James Vincent · 2023
Closest in time.
Poisoning language models during instruction tuning
Alex Wan, Eric Wallace, Sheng Shen, and Dan Klein · 2023
Closest in time.
Aligning large language models with human: A survey, 2023
Yufei Wang, Wanjun Zhong, Liangyou Li, Fei Mi, Xingshan Zeng, Wenyong Huang, Lifeng Shang, Xin Jiang, and Qun Liu · 2023
Closest in time.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Closest in time.
Prompt injection
Simon Willison · 2023
Closest in time.
Fundamental limitations of alignment in large language models
Yotam Wolf, Noam Wies, Yoav Levine, and Amnon Shashua · 2023
Closest in time.
Fine-grained human feedback gives better rewards for language model training, 2023
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi · 2023
Closest in time.
Selfee: Iterative self-revising llm empowered by self-feedback generation, 2023
Seonghyeon Ye, Yongrae Jo, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, and Minjoon Seo · 2023
Closest in time.
Language to rewards for robotic skill synthesis
Wenhao Yu, Nimrod Gileadi, Chuyuan Fu, Sean Kirmani, Kuang-Huei Lee, Montse Gonzalez Arenas, Hao-Tien Lewis Chiang, Tom Erez, Leonard Hasenclever, Jan Humplik, Brian Ichter, Ted Xiao, Peng Xu, Andy Zeng, Tingnan Zhang, Nicolas Heess, Dorsa Sadigh, Jie Tan, Yuval Tassa, and Fei Xia · 2023
Closest in time.
Rrhf: Rank responses to align language models with human feedback without tears, 2023
Zheng Yuan, Hongyi Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang · 2023
Closest in time.
Clare: Conservative model-based reward learning for offline inverse reinforcement learning
Sheng Yue, Guanbo Wang, Wei Shao, Zhaofeng Zhang, Sen Lin, Ju Ren, and Junshan Zhang · 2023
Closest in time.
How language model hallucinations can snowball
Muru Zhang, Ofir Press, William Merrill, Alisa Liu, and Noah A Smith · 2023
Closest in time.
Lima: Less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, et al · 2023
Closest in time.
Principled reinforcement learning with human feedback from pairwise or k k -wise comparisons
Banghua Zhu, Jiantao Jiao, and Michael I Jordan · 2023
Closest in time.