Fetching the paper…
Reading the bibliography…
Scalable oversight, the process by which weaker AI systems supervise stronger ones, has been proposed as a key strategy to control future superintelligent systems.
Nim, a game with a complete mathematical theory
Bouton, C. L · 1901
Earlier work this paper cites.
The proposed uscf rating system, its development, theory, and applications
Elo, A. E · 1967
Earlier work this paper cites.
Modified reactor safety goal policy statement
US Nuclear Regulatory Commission · 2001
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
Christiano, P., Shlegeris, B., and Amodei, D · 2018
Earlier work this paper cites.
Irving, G., Christiano, P., and Amodei, D · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Earlier work this paper cites.
The akaike information criterion: Background, derivation, properties, application, interpretation, and refinements
Cavanaugh, J. E. and Neath, A. A · 2019
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Ai safety via market making
Hubinger, E · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Measuring coding challenge competence with APPS
Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2021
Earlier work this paper cites.
Revisiting neural scaling laws in language and vision
Alabdulmohsin, I. M., Neyshabur, B., and Zhai, X · 2022
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models
Bowman, S. R., Hyun, J., Perez, E., Chen, E., Pettit, C., Heiner, S., Lukošiūtė, K., Askell, A., Jones, A., Chen, A., et al · 2022
Earlier work this paper cites.
QuALITY: Question answering with long input texts, yes!
Pang, R. Y., Parrish, A., Joshi, N., Nangia, N., Phang, J., Chen, A., Padmakumar, V., Ma, J., Thompson, J., He, H., and Bowman, S · 2022
Earlier work this paper cites.
Self-critiquing models for assisting human evaluators
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J · 2022
Cited alongside, same era.
Scaling laws for generative mixed-modal language models
Aghajanyan, A., Yu, L., Conneau, A., Hsu, W.-N., Hambardzumyan, K., Zhang, S., Roller, S., Goyal, N., Levy, O., and Zettlemoyer, L · 2023
Cited alongside, same era.
Scalable ai safety via doubly-efficient debate
Brown-Cohen, J., Irving, G., and Piliouras, G · 2023
Cited alongside, same era.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al · 2023
Cited alongside, same era.
Reproducible scaling laws for contrastive language-image learning
Cherti, M., Beaumont, R., Wightman, R., Wortsman, M., Ilharco, G., Gordon, C., Schuhmann, C., Schmidt, L., and Jitsev, J · 2023
Debating with more persuasive llms leads to more truthful answers
Khan, A., Hughes, J., Valentine, D., Ruis, L., Sachan, K., Radhakrishnan, A., Grefenstette, E., Bowman, S. R., Rocktäschel, T., and Perez, E · 2024
Later among the works it cites.
Theoretical analysis of weak-to-strong generalization
Lang, H., Sontag, D., and Vijayaraghavan, A · 2024
Later among the works it cites.
Touvron, H. et al · 2024
Later among the works it cites.
Yang, A. et al · 2024
Later among the works it cites.
Finding deceivers in social context with large language models and how to find them: the case of the mafia game
Yoo, B. and Kim, K.-J · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Fy2023 q2 aviation safety progress report
Federal Aviation Administration · 2023
Cited alongside, same era.
Gemini: A family of highly capable multimodal models
Google DeepMind · 2023
Cited alongside, same era.
Ai control: Improving safety despite intentional subversion
Greenblatt, R., Shlegeris, B., Sachan, K., and Roger, F · 2023
Cited alongside, same era.
Debate helps supervise unreliable experts
Michael, J., Mahdi, S., Rein, D., Petty, J., Dirani, J., Padmakumar, V., and Bowman, S. R · 2023
Cited alongside, same era.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Saparov, A. and He, H · 2023
Cited alongside, same era.
Introducing claude 3.5 haiku
Anthropic · 2024
Cited alongside, same era.
Vision superalignment: Weak-to-strong generalization for vision foundation models
Guo, J., Chen, H., Wang, C., Han, K., Xu, C., and Wang, Y · 2024
Cited alongside, same era.
Recommendations for technical ai safety research directions
Anthropic Alignment Science Team · 2025
Closest in time.
How ai takeover might happen in 2 years, February 2025
Clymer, J · 2025
Closest in time.
Among us: A sandbox for agentic deception
Golechha, S. and Garriga-Alonso, A · 2025
Closest in time.
Llm mafia game, 2025
Guzus · 2025
Closest in time.
Ai 2027: A scenario forecast
Kokotajlo, D., Lifland, E., Larsen, T., Dean, R., and Vollmer, J · 2025
Closest in time.
Introducing superalignment
OpenAI · 2025
Closest in time.
An approach to technical agi safety and security
Shah, R., Irpan, A., Turner, A. M., Wang, A., Conmy, A., Lindner, D., Brown-Cohen, J., Ho, L., Nanda, N., Popa, R. A., et al · 2025
Closest in time.
The leaderboard illusion, 2025
Singh, S., Nan, Y., Wang, A., D’Souza, D., Kapoor, S., Üstün, A., Koyejo, S., Deng, Y., Longpre, S., Smith, N., Ermis, B., Fadaee, M., and Hooker, S · 2025
Closest in time.
A benchmark for scalable oversight protocols
Sudhir, A. P., Kaunismaa, J., and Panickssery, A · 2025
Closest in time.