Fetching the paper…
Reading the bibliography…
Scalable oversight protocols aim to enable humans to accurately supervise superhuman AI.
Concrete problems in AI safety
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané · 2016
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
P. Christiano, B. Shlegeris, and D. Amodei · 2018
Earlier work this paper cites.
G. Irving, P. Christiano, and D. Amodei · 2018
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction
J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg · 2018
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
C. Clark, K. Lee, M.-W. Chang, T. Kwiatkowski, M. Collins, and K. Toutanova · 2019
Earlier work this paper cites.
Writeup: Progress on AI Safety via Debate, 2020
B. Barnes and P. Christiano · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
AI safety via market making
E. Hubinger · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
The case for aligning narrowly superhuman models
A. Cotra · 2021
Earlier work this paper cites.
Risks from Learned Optimization in Advanced Machine Learning Systems
E. Hubinger, C. van Merwijk, V. Mikulik, J. Skalse, and S. Garrabrant · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2021
Earlier work this paper cites.
QuALITY: Question answering with long input texts, yes!
R. Y. Pang, A. Parrish, N. Joshi, N. Nangia, J. Phang, A. Chen, V. Padmakumar, J. Ma, J. Thompson, H. He, et al · 2021
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models
S. R. Bowman, J. Hyun, E. Perez, E. Chen, C. Pettit, S. Heiner, K. Lukošiūtė, A. Askell, A. Jones, A. Chen, et al · 2022
Earlier work this paper cites.
Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover, July 2022
A. Cotra · 2022
Earlier work this paper cites.
The alignment problem from a deep learning perspective
R. Ngo, L. Chan, and S. Mindermann · 2022
Cited alongside, same era.
Self-critiquing models for assisting human evaluators
W. Saunders, C. Yeh, J. Wu, S. Bills, L. Ouyang, J. Ward, and J. Leike · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
Scalable AI safety via doubly-efficient debate
J. Brown-Cohen, G. Irving, and G. Piliouras · 2023
Cited alongside, same era.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Can chatgpt defend its belief in truth? Evaluating LLM reasoning via debate
B. Wang, X. Yue, and H. Sun · 2023
Later among the works it cites.
MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al · 2023
Later among the works it cites.
On large language models’ selection bias in multi-choice questions
C. Zheng, H. Zhou, F. Meng, J. Zhou, and M. Huang · 2023
Later among the works it cites.
R. Agarwal, A. Singh, L. M. Zhang, B. Bohnet, S. Chan, A. Anand, Z. Abbas, A. Nova, J. D. Co-Reyes, E. Chu, et al · 2024
Closest in time.
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Burns, P. Izmailov, J. H. Kirchner, B. Baker, L. Gao, L. Aschenbrenner, Y. Chen, A. Ecoffet, M. Joglekar, J. Leike, et al · 2023
Cited alongside, same era.
Scheming AIs: Will AIs fake alignment during training in order to get power?
J. Carlsmith · 2023
Cited alongside, same era.
Chateval: Towards better LLM-based evaluators through multi-agent debate
C.-M. Chan, W. Chen, Y. Su, J. Yu, W. Xue, S. Zhang, J. Fu, and Z. Liu · 2023
Cited alongside, same era.
Improving factuality and reasoning in language models through multiagent debate
Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch · 2023
Cited alongside, same era.
Pal: Program-aided language models
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Gemini Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Cited alongside, same era.
Large language models cannot self-correct reasoning yet
J. Huang, X. Chen, S. Mishra, H. S. Zheng, A. W. Yu, X. Song, and D. Zhou · 2023
Cited alongside, same era.
Encouraging divergent thinking in large language models through multi-agent debate
T. Liang, Z. He, W. Jiao, X. Wang, Y. Wang, R. Wang, Y. Yang, Z. Tu, and S. Shi · 2023
Cited alongside, same era.
C. Denison, M. MacDiarmid, F. Barez, D. Duvenaud, S. Kravec, S. Marks, N. Schiefer, R. Soklaski, A. Tamkin, J. Kaplan, et al · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, et al · 2024
Closest in time.
Debating with more persuasive LLMs leads to more truthful answers
A. Khan, J. Hughes, D. Valentine, L. Ruis, K. Sachan, A. Radhakrishnan, E. Grefenstette, S. R. Bowman, T. Rocktäschel, and E. Perez · 2024
Closest in time.
PRD: Peer rank and discussion improve large language model based evaluations
R. Li, T. Patel, and X. Du · 2024
Closest in time.
LLM evaluators recognize and favor their own generations
A. Panickssery, S. R. Bowman, and S. Feng · 2024
Closest in time.
Let models speak ciphers: Multiagent debate through embeddings
C. Pham, B. Liu, Y. Yang, Z. Chen, T. Liu, J. Yuan, B. A. Plummer, Z. Wang, and H. Yang · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
Open consultancy: Letting untrusted AIs choose what answer to argue for, 2024
F. Roger · 2024
Closest in time.
Large language models are inconsistent and biased evaluators
R. Stureborg, D. Alikaniotis, and Y. Suhara · 2024
Closest in time.
Replacing judges with juries: Evaluating LLM generations with a panel of diverse models
P. Verga, S. Hofstatter, S. Althammer, Y. Su, A. Piktus, A. Arkhangorodsky, M. Xu, N. White, and P. Lewis · 2024
Closest in time.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2024
Closest in time.