Fetching the paper…
Reading the bibliography…
Trustworthy capability evaluations are crucial for ensuring the safety of AI systems, and are becoming a key component of AI regulation.
The problem with metrics is a fundamental problem for AI
Thomas, R. and Uminsky, D. (2020) · 2002
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2020) · 2009
Earlier work this paper cites.
Towards formal definitions of blameworthiness, intention, and moral responsibility
Halpern, J. and Kleiman-Weiner, M. (2018) · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Talmor, A., Herzig, J., Lourie, N., and Berant, J. (2018) · 2018
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 2020
Earlier work this paper cites.
LoRA: Low-Rank Adaptation of Large Language Models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2021) · 2021
Earlier work this paper cites.
Risks from Learned Optimization in Advanced Machine Learning Systems
Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., and Garrabrant, S. (2021) · 2021
Earlier work this paper cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. (2022) · 2022
Earlier work this paper cites.
The alignment problem from a deep learning perspective
Ngo, R., Chan, L., and Mindermann, S. (2022) · 2022
Earlier work this paper cites.
Discovering Language Model Behaviors with Model-Written Evaluations
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., Jones, A., Chen, A., Mann, B., Israel, B., Seethor, B., McKinnon, C., Olah, C., Yan, D., Amodei, D., Amodei, D., Drain, D., Li, D., Tran-Johnson, E., Khundadze, G., Kernion, J., Landis, J., Kerr, J., Mueller, J., Hyun, J., Landau, J., Ndousse, K., Goldberg, L., Lovitt, L., Lucas, M., Sellitto, M., Zhang, M., Kingsland, N., Elhage, N., Joseph, N., Mercado, N., DasSarma, N., Rausch, O., Larson, R., McCandlish, S., Johnston, S., Kravec, S., Showk, S. E., Lanham, T., Telleen-Lawton, T., Brown, T., Henighan, T., Hume, T., Bai, Y., Hatfield-Dodds, Z., Clark, J., Bowman, S. R., Askell, A., Grosse, R., Hernandez, D., Ganguli, D., Hubinger, E., Schiefer, N., and Kaplan, J. (2022) · 2022
Earlier work this paper cites.
Anthropic’s Responsible Scaling Policy
Anthropic (2023) · 2023
Earlier work this paper cites.
Definitions of intent suitable for algorithms
Ashton, H. (2023) · 2023
Earlier work this paper cites.
Eight Things to Know about Large Language Models
Bowman, S. R. (2023) · 2023
Earlier work this paper cites.
Poisoning Web-Scale Training Datasets is Practical
Carlini, N., Jagielski, M., Choquette-Choo, C. A., Paleka, D., Pearce, W., Anderson, H., Terzis, A., Thomas, K., and Tramèr, F. (2023) · 2023
Earlier work this paper cites.
Scheming AIs: Will AIs fake alignment during training in order to get power?
Carlsmith, J. (2023) · 2023
Earlier work this paper cites.
AI control: Improving safety despite intentional subversion
Greenblatt, R., Shlegeris, B., Sachan, K., and Roger, F. (2023) · 2023
Earlier work this paper cites.
Language models represent space and time
Gurnee, W. and Tegmark, M. (2023) · 2023
Earlier work this paper cites.
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
Jain, S., Kirk, R., Lubana, E. S., Dick, R. P., Tanaka, H., Grefenstette, E., Rockt"aschel, T., and Krueger, D. S. (2023) · 2023
Earlier work this paper cites.
Mistral 7B
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. (2023) · 2023
Earlier work this paper cites.
Governing General Purpose AI
Maham, P. and Küspert, S. (2023) · 2023
Earlier work this paper cites.
Preparedness Framework (Beta)
OpenAI (2023) · 2023
Earlier work this paper cites.
How to catch an AI liar: Lie detection in black-box LLMs by asking unrelated questions
Pacchiardi, L., Chan, A. J., Mindermann, S., Moscovitz, I., Pan, A. Y., Gal, Y., Evans, O., and Brauner, J. (2023) · 2023
Earlier work this paper cites.
Pretraining on the test set is all you need
Schaeffer, R. (2023) · 2023
Earlier work this paper cites.
Practices for Governing Agentic AI Systems
Shavit, Y., Agarwal, S., Brundage, M., Adler, S., O’Keefe, C., Campbell, R., Lee, T., Mishkin, P., Eloundou, T., Hickey, A., et al. (2023) · 2023
Earlier work this paper cites.
Model evaluation for extreme risks
Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., et al. (2023) · 2023
Earlier work this paper cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T. (2023) · 2023
Cited alongside, same era.
Poisoning Language Models During Instruction Tuning
Wan, A., Wallace, E., Shen, S., and Klein, D. (2023) · 2023
Cited alongside, same era.
Voyager: An open-ended embodied agent with large language models
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A. (2023) · 2023
Cited alongside, same era.
Inspect An open-source framework for large language model evaluations
AI Safety Institute (2024) · 2024
Cited alongside, same era.
The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Li, N., Pan, A., Gopal, A., Yue, S., Berrios, D., Gatti, A., Li, J. D., Dombrowski, A.-K., Goel, S., Phan, L., et al. (2024) · 2024
Closest in time.
Eight Methods to Evaluate Robust Unlearning in LLMs
Lynch, A., Guo, P., Ewart, A., Casper, S., and Hadfield-Menell, D. (2024) · 2024
Closest in time.
Eliciting Latent Knowledge from Quirky Language Models
Mallen, A., Brumley, M., Kharchenko, J., and Belrose, N. (2024) · 2024
Closest in time.
Trojan Detection in Large Language Models: Insights from The Trojan Detection Challenge
Maloyan, N., Verma, E., Nutfullin, B., and Ashinov, B. (2024) · 2024
Closest in time.
Introducing Meta Llama 3: The most capable openly available LLM to date
Meta (2024) · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover
Ajeya Cotra (2022) · 2024
Cited alongside, same era.
Fun story from our internal testing on Claude 3 Opus
Alex Albert (2024) · 2024
Cited alongside, same era.
Introducing the next generation of Claude
Anthropic (2024a) · 2024
Cited alongside, same era.
Simple probes can catch sleeper agents
Anthropic (2024b) · 2024
Cited alongside, same era.
The Claude 3 Model Family: Opus, Sonnet, Haiku
Anthropic (2024) · 2024
Cited alongside, same era.
Challenges in evaluating AI systems
Antrhopic (2024) · 2024
Cited alongside, same era.
Understanding strategic deception and deceptive alignment
Apollo Research (2023) · 2024
Cited alongside, same era.
Black-Box Access is Insufficient for Rigorous AI Audits
Casper, S., Ezell, C., Siegmann, C., Kolt, N., Curtis, T. L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., et al. (2024) · 2024
Cited alongside, same era.
Closest in time.
Details about metr’s preliminary evaluation of openai o1-preview
METR (2024) · 2024
Closest in time.
U.S. Artificial Intelligence Safety Institute | | NIST
NIST (2024) · 2024
Closest in time.
Developing beneficial AGI safely and responsibly
Open AI (2024) · 2024
Closest in time.
GPT-4 Technical Report
OpenAI (2024) · 2024
Closest in time.
Models documentation
OpenAI (2024a) · 2024
Closest in time.
OpenAI o1 System Card
OpenAI (2024c) · 2024
Closest in time.
Evaluating Frontier Models for Dangerous Capabilities
Phuong, M., Aitchison, M., Catt, E., Cogan, S., Kaskasoli, A., Krakovna, V., Lindner, D., Rahtz, M., Assael, Y., Hodkinson, S., et al. (2024) · 2024
Closest in time.
Future events as backdoor triggers: investigating temporal vulnerabilities in LLMs
Price, S., Strickland, A. C., and Bowman, S. (2024) · 2024
Closest in time.
Password-locked models: a stress case for capabilities evaluation
Roger, F. (2023) · 2024
Closest in time.
Biden-Harris Administration Announces Key AI Actions 180 Days Following President Biden’s Landmark Executive Order
The White House (2024) · 2024
Closest in time.
AI Safety Institute approach to evaluations
UK AISI (2024) · 2024
Closest in time.
Capabilities and risks from frontier AI
UK DSIT (2023a) · 2024
Closest in time.
Emerging processes for frontier AI safety
UK DSIT (2023b) · 2024
Closest in time.
Implementing the UK’s AI Regulatory Principles
UK DSIT (2023c) · 2024
Closest in time.
AI Safety Institute approach to evaluations
UK DSIT (2024) · 2024
Closest in time.
Trading Off Compute in Training and Inference
Villalobos, P. and Atkinson, D. (2023) · 2024
Closest in time.
Rethinking Generative Large Language Model Evaluation for Semantic Comprehension
Wei, F., Chen, X., and Luo, L. (2024) · 2024
Closest in time.
Overfitting
Wikipedia (2024a) · 2024
Closest in time.
Sandbagging
Wikipedia (2024b) · 2024
Closest in time.
Volkswagen emissions scandal
Wikipedia (2024c) · 2024
Closest in time.