Fetching the paper…
Reading the bibliography…
Human feedback is commonly utilized to finetune AI assistants.
Assessing sensitivity to an unobserved binary covariate in an observational study with binary outcome
Paul R Rosenbaum and Donald B Rubin · 1983
Earlier work this paper cites.
Sensitivity analysis for selection bias and unmeasured confounding in missing data and causal inference models
James M Robins, Andrea Rotnitzky, and Daniel O Scharfstein · 2000
Earlier work this paper cites.
MCMC using Hamiltonian dynamics
Radford M Neal et al · 2011
Earlier work this paper cites.
The No-U-Turn sampler: Adaptively setting path lengths in Hamiltonian Monte Carlo
Matthew D Hoffman, Andrew Gelman, et al · 2014
Earlier work this paper cites.
Learning mixtures of Plackett-Luce models
Zhibing Zhao, Peter Piech, and Lirong Xia · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Program induction by rationale generation: Learning to solve and explain algebraic word problems
Wang Ling, Dani Yogatama, Chris Dyer, and Phil Blunsom · 2017
Earlier work this paper cites.
AI safety via debate, 2018
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: A research direction, 2018
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Earlier work this paper cites.
Occam’s Razor is insufficient to infer the preferences of irrational agents
Soren Mindermann and Stuart Armstrong · 2018
Earlier work this paper cites.
Composable effects for flexible and accelerated probabilistic programming in NumPyro
Du Phan, Neeraj Pradhan, and Martin Jankowiak · 2019
Earlier work this paper cites.
On the feasibility of learning, rather than assuming, human biases for reward inference
Rohin Shah, Noah Gundotra, Pieter Abbeel, and Anca Dragan · 2019
Earlier work this paper cites.
An MTurk crisis? Shifts in data quality and the impact on study results
Michael Chmielewski and Sarah C Kucker · 2020
Cited alongside, same era.
WebGPT: Browser-assisted question-answering with human feedback
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, et al · 2021
Cited alongside, same era.
Fine-tuning language models to find agreement among humans with diverse preferences
Michiel Bakker, Martin Chadwick, Hannah Sheahan, Michael Tessler, Lucy Campbell-Gillingham, Jan Balaguer, Nat McAleese, Amelia Glaese, John Aslanides, Matt Botvinick, et al · 2022
Cited alongside, same era.
Measuring progress on scalable oversight for large language models
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Christopher Olah, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Jackson Kernion, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Liane Lovitt, Nelson Elhage, Nicholas Schiefer, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Robin Larson, Sam McCandlish, Sandipan Kundu, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, and Jared Kaplan · 2022
Self-critiquing models for assisting human evaluators
William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike · 2022
Later among the works it cites.
Claude 2, 2023
Anthropic · 2023
Closest in time.
Open problems and fundamental limitations of reinforcement learning from human feedback, 2023
Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, Tony Wang, Samuel Marks, Charbel-Raphaël Segerie, Micah Carroll, Andi Peng, Phillip Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, Anand Siththaranjan, Max Nadeau, Eric J. Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Bıyık, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell · 2023
Closest in time.
Why AI alignment could be hard with modern deep learning
Ajeya Cotra · 2023
Closest in time.
Maintenance Phase: Debunking the junk science behind health fads, wellness scams and nonsensical nutrition advice., October 2020
Aubrey Gordon and Michael Hobbes · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Scaling laws for reward model overoptimization
Leo Gao, John Schulman, and Jacob Hilton · 2022
Cited alongside, same era.
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, et al · 2022
Cited alongside, same era.
On the sensitivity of reward inference to misspecified human models
Joey Hong, Kush Bhatia, and Anca Dragan · 2022
Cited alongside, same era.
TruthfulQA: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans · 2022
Cited alongside, same era.
David Lindner and Mennatallah El-Assady · 2022
Cited alongside, same era.
Introducing chatgpt, 2022
OpenAI · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Modeling and mitigating human annotation errors to design efficient stream processing systems with human-in-the-loop machine learning
Rahul Pandey, Hemant Purohit, Carlos Castillo, and Valerie L Shalin · 2022
Cited alongside, same era.
Closest in time.
The false promise of imitating proprietary LLMs
Arnav Gudibande, Eric Wallace, Charlie Snell, Xinyang Geng, Hao Liu, Pieter Abbeel, Sergey Levine, and Dawn Song · 2023
Closest in time.
Understanding the effects of rlhf on llm generalisation and diversity
Robert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina, Eric Hambro, Edward Grefenstette, and Roberta Raileanu · 2023
Closest in time.
GPT-4 technical report, 2023
OpenAI · 2023
Closest in time.
Question decomposition improves the faithfulness of model-generated reasoning
Ansh Radhakrishnan, Karina Nguyen, Anna Chen, Carol Chen, Carson Denison, Danny Hernandez, Esin Durmus, Evan Hubinger, Jackson Kernion, Kamilė Lukošiūtė, et al · 2023
Closest in time.
Blog post on the AI Alignment Forum, Jul 2023
Nina Rimsky · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Miles Turpin, Julian Michael, Ethan Perez, and Samuel R. Bowman · 2023
Closest in time.