Fetching the paper…
Reading the bibliography…
As AI agents surpass human capabilities, scalable oversight -- the problem of effectively supplying human feedback to potentially superhuman AI models -- becomes increasingly critical to ensure alignment.
Coherent Extrapolated Volition, 2004
Eliezer Yudkowsky · 2004
Earlier work this paper cites.
Deep Reinforcement Learning from Human Preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts, 2018
Paul Christiano, Buck Shlegeris, and Dario Amodei · 2018
Earlier work this paper cites.
AI safety via debate, October 2018
Geoffrey Irving, Paul Christiano, and Dario Amodei · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction, 2018
Jan Leike, David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and Shane Legg · 2018
Earlier work this paper cites.
AI safety via market making — LessWrong, June 2020
Evan Hubinger · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Earlier work this paper cites.
Measuring Progress on Scalable Oversight for Large Language Models, November 2022
Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Christopher Olah, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Jackson Kernion, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Liane Lovitt, Nelson Elhage, Nicholas Schiefer, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Robin Larson, Sam McCandlish, Sandipan Kundu, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, and Jared Kaplan · 2022
Cited alongside, same era.
Self-critiquing models for assisting human evaluators, 2022
William Saunders, Catherine Yeh, Jeff Wu, Steven Bills, Long Ouyang, Jonathan Ward, and Jan Leike · 2022
Cited alongside, same era.
Debate helps supervise unreliable experts, 2023
Julian Michael, Salsabila Mahdi, David Rein, Jackson Petty, Julien Dirani, Vishakh Padmakumar, and Samuel R. Bowman · 2023
Cited alongside, same era.
Anthropic Fall 2023 Debate Progress Update
Ansh Radhakrishnan · 2023
Cited alongside, same era.
Reinforcement learning from multi-role debates as feedback for bias mitigation in llms, 2024
Ruoxi Cheng, Haoxuan Ma, Shuirong Cao, Jiaqi Li, Aihua Pei, Zhiqiang Wang, Pengliang Ji, Haoyu Wang, and Jiaqi Huo · 2024
Later among the works it cites.
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2024
Elliot Glazer, Ege Erdil, Tamay Besiroglu, Diego Chicharro, Evan Chen, Alex Gunning, Caroline Falkman Olsson, Jean-Stanislas Denain, Anson Ho, Emily de Oliveira Santos, Olli Järviniemi, Matthew Barnett, Robert Sandler, Matej Vrzala, Jaime Sevilla, Qiuyu Ren, Elizabeth Pratt, Lionel Levine, Grant Barkley, Natalie Stewart, Bogdan Grechuk, Tetiana Grechuk, Shreepranav Varma Enugandla, and Mark Wildon · 2024
Later among the works it cites.
On scalable oversight with weak LLMs judging strong LLMs, July 2024
Zachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen, Samuel Albanie, Jannis Bulian, Rishabh Agarwal, David Lindner, Yunhao Tang, Noah D. Goodman, and Rohin Shah · 2024
Later among the works it cites.
Debating with More Persuasive LLMs Leads to More Truthful Answers, July 2024
Akbir Khan, John Hughes, Dan Valentine, Laura Ruis, Kshitij Sachan, Ansh Radhakrishnan, Edward Grefenstette, Samuel R. Bowman, Tim Rocktäschel, and Ethan Perez · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Training Language Models to Win Debates with Self-Play Improves Judge Accuracy
Samuel Arnesen, David Rein, and Julian Michael · 2024
Cited alongside, same era.
Scalable ai safety via doubly-efficient debate
Jonah Brown-Cohen, Geoffrey Irving, and Georgios Piliouras · 2024
Cited alongside, same era.
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeffrey Wu · 2024
Cited alongside, same era.
Later among the works it cites.
Improving weak-to-strong generalization with scalable oversight and ensemble learning, 2024
Jitao Sang, Yuhang Wang, Jing Zhang, Yanxu Zhu, Chao Kong, Junhong Ye, Shuyu Wei, and Jinlin Xiao · 2024
Later among the works it cites.
Debate helps weak-to-strong generalization, 2025
Hao Lang, Fei Huang, and Yongbin Li · 2025
Closest in time.