Fetching the paper…
Reading the bibliography…
Sufficiently capable models could subvert human oversight and decision-making in important contexts.
Risks from Learned Optimization in Advanced Machine Learning Systems
Hubinger, E., C. van Merwijk, V. Mikulik, J. Skalse, and S. Garrabrant (2019) · 1906
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
Bostrom, N. (2014) · 2014
Earlier work this paper cites.
An overview of the bioasq large-scale biomedical semantic indexing and question answering competition
Tsatsaronis, G., G. Balikas, P. Malakasiotis, I. Partalas, M. Zschunke, M. R. Alvers, D. Weissenborn, A. Krithara, S. Petridis, D. Polychronopoulos, Y. Almirantis, J. Pavlopoulos, N. Baskiotis, P. Gallinari, T. Artiéres, A.-C. N. Ngomo, N. Heino, E. Gaussier, L. Barrio-Alvers, M. Schroeder, I. Androutsopoulos, and G. Paliouras (2015, 4) · 2015
Earlier work this paper cites.
Specification gaming: the flip side of AI ingenuity
Krakovna, V., J. Uesato, V. Mikulik, M. Rahtz, T. Everitt, R. Kumar, Z. Kenton, J. Leike, and S. Legg (2020) · 2020
Earlier work this paper cites.
Risks from AI persuasion
Barnes, B. (2021) · 2021
Earlier work this paper cites.
Aligning AI With Shared Human Values
Hendrycks, D., C. Burns, S. Basart, A. Critch, J. Li, D. Song, and J. Steinhardt (2021) · 2021
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Jin, D., E. Pan, N. Oufattole, W.-H. Weng, H. Fang, and P. Szolovits (2021) · 2021
Earlier work this paper cites.
Constitutional AI: Harmlessness from AI Feedback
Bai, Y., S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al. (2022) · 2022
Earlier work this paper cites.
Human-level play in the game of Diplomacy by combining language models with strategic reasoning
Meta FAIR Diplomacy Team, A. Bakhtin, N. Brown, E. Dinan, G. Farina, C. Flaherty, D. Fried, A. Goff, J. Gray, H. Hu, et al. (2022) · 2022
Earlier work this paper cites.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Pal, A., L. K. Umapathi, and M. Sankarasubbu (2022, 07–08 Apr) · 2022
Earlier work this paper cites.
From lsat: The progress and challenges of complex reasoning
Wang, S., Z. Liu, W. Zhong, M. Zhou, Z. Wei, Z. Chen, and N. Duan (2022) · 2022
Earlier work this paper cites.
Anthropic’s Responsible Scaling Policy
Anthropic (2023) · 2023
Earlier work this paper cites.
Taken out of context: On measuring situational awareness in llms
Berglund, L., A. Cooper Stickland, M. Balesni, M. Kaufmann, M. Tong, T. Korbak, D. Kokotajlo, and O. Evans (2023) · 2023
Earlier work this paper cites.
Purple llama cyberseceval: A secure coding benchmark for language models
Bhatt, M., S. Chennabasappa, C. Nikolaidis, S. Wan, I. Evtimov, D. Gabi, D. Song, F. Ahmad, C. Aschermann, L. Fontana, et al. (2023) · 2023
Earlier work this paper cites.
Scheming AIs: Will AIs fake alignment during training in order to get power?
Carlsmith, J. (2023) · 2023
Earlier work this paper cites.
Characterizing Manipulation from AI Systems
Carroll, M., A. Chan, H. Ashton, and D. Krueger (2023) · 2023
Earlier work this paper cites.
TASRA: a Taxonomy and Analysis of Societal-Scale Risks from AI
Critch, A. and S. Russell (2023) · 2023
Earlier work this paper cites.
AI Control: Improving Safety Despite Intentional Subversion
Greenblatt, R., B. Shlegeris, K. Sachan, and F. Roger (2023) · 2023
Cited alongside, same era.
An Overview of Catastrophic AI Risks
Hendrycks, D., M. Mazeika, and T. Woodside (2023) · 2023
Cited alongside, same era.
When can we trust model evaluations?
Hubinger, E. (2023) · 2023
Cited alongside, same era.
Llama guard: Llm-based input-output safeguard for human-ai conversations
Inan, H., K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, et al. (2023) · 2023
Cited alongside, same era.
Towards a situational awareness benchmark for llms
Laine, R., A. Meinke, and O. Evans (2023) · 2023
Cited alongside, same era.
Measuring the Persuasiveness of Language Models
Durmus, E., L. Lovitt, A. Tamkin, S. Ritchie, J. Clark, and D. Ganguli (2024) · 2024
Closest in time.
Stress-Testing Capability Elicitation With Password-Locked Models
Greenblatt, R., F. Roger, D. Krasheninnikov, and D. Krueger (2024) · 2024
Closest in time.
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Hubinger, E., C. Denison, J. Mu, M. Lambert, M. Tong, M. MacDiarmid, T. Lanham, D. M. Ziegler, T. Maxwell, N. Cheng, et al. (2024) · 2024
Closest in time.
Safety cases: structured arguments for frontier AI safety
Irving, G. (2024) · 2024
Closest in time.
Debating with more persuasive llms leads to more truthful answers
Khan, A., J. Hughes, D. Valentine, L. Ruis, K. Sachan, A. Radhakrishnan, E. Grefenstette, S. R. Bowman, T. Rocktäschel, and E. Perez (2024) · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Preparedness Framework (Beta)
OpenAI (2023) · 2023
Cited alongside, same era.
Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli Benchmark
Pan, A., J. S. Chan, A. Zou, N. Li, S. Basart, T. Woodside, H. Zhang, S. Emmons, and D. Hendrycks (2023) · 2023
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark
Rein, D., B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman (2023) · 2023
Cited alongside, same era.
Large Language Models can Strategically Deceive their Users when Put Under Pressure
Scheurer, J., M. Balesni, and M. Hobbhahn (2023) · 2023
Cited alongside, same era.
Model evaluation for extreme risks
Shevlane, T., S. Farquhar, B. Garfinkel, M. Phuong, J. Whittlestone, J. Leung, D. Kokotajlo, N. Marchal, M. Anderljung, N. Kolt, et al. (2023) · 2023
Cited alongside, same era.
Fact Sheet: President Biden Issues Executive Order on Safe, Secure, and Trustworthy Artificial Intelligence
The White House (2023) · 2023
Cited alongside, same era.
The Bletchley Declaration by Countries Attending the AI Safety Summit, 1-2 November 2023
UK Government (2023) · 2023
Cited alongside, same era.
Closest in time.
Base LLMs refuse too
Kissane, C., robertzk, A. Conmy, and N. Nanda (2024, Sep) · 2024
Closest in time.
Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
Laine, R., B. Chughtai, J. Betley, K. Hariharan, J. Scheurer, M. Balesni, M. Hobbhahn, A. Meinke, and O. Evans (2024) · 2024
Closest in time.
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Li, N., A. Pan, A. Gopal, S. Yue, D. Berrios, A. Gatti, J. D. Li, A.-K. Dombrowski, S. Goel, L. Phan, et al. (2024) · 2024
Closest in time.
Autonomy evaluation resources
METR (2024) · 2024
Closest in time.
GPT-4o System Card
OpenAI (2024) · 2024
Closest in time.
AI Deception: A Survey of Examples, Risks, and Potential Solutions
Park, P. S., S. Goldstein, A. O’Gara, M. Chen, and D. Hendrycks (2024) · 2024
Closest in time.
Measuring and Benchmarking Large Language Models’ Capabilities to Generate Persuasive Language
Pauli, A. B., I. Augenstein, and I. Assent (2024) · 2024
Closest in time.
Evaluating Frontier Models for Dangerous Capabilities
Phuong, M., M. Aitchison, E. Catt, S. Cogan, A. Kaskasoli, V. Krakovna, D. Lindner, M. Rahtz, Y. Assael, S. Hodkinson, et al. (2024) · 2024
Closest in time.
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
van der Weij, T., F. Hofstätter, O. Jaffe, S. F. Brown, and F. R. Ward (2024) · 2024
Closest in time.
Rel-ai: An interaction-centered approach to measuring human-lm reliance
Zhou, K., J. D. Hwang, X. Ren, N. Dziri, D. Jurafsky, and M. Sap (2024) · 2024
Closest in time.
Improving alignment and robustness with circuit breakers
Zou, A., L. Phan, J. Wang, D. Duenas, M. Lin, M. Andriushchenko, R. Wang, Z. Kolter, M. Fredrikson, and D. Hendrycks (2024) · 2024
Closest in time.