Fetching the paper…
Reading the bibliography…
Capability evaluations are required to understand and regulate AI systems that may be deployed or further developed.
Backdoor Learning: A Survey, 2022
Li, Y., Jiang, Y., Li, Z., and Xia, S.-T · 2007
Earlier work this paper cites.
Scalable agent alignment via reward modeling: a research direction, 2018
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Measuring coding challenge competence with apps
Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., et al · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J · 2022
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Anthropic’s Responsible Scaling Policy, 2023
Anthropic · 2023
Earlier work this paper cites.
Structured access for third-party research on frontier AI models: investigating researchers’ model access requirements, 2023
Bucknall, B. S. and Trager, R. F · 2023
Earlier work this paper cites.
Poisoning Web-Scale Training Datasets is Practical, 2023
Carlini, N., Jagielski, M., Choquette-Choo, C. A., Paleka, D., Pearce, W., Anderson, H., Terzis, A., Thomas, K., and Tramèr, F · 2023
Earlier work this paper cites.
Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research
Hubinger, E., Schiefer, N., Denison, C., and Perez, E · 2023
Earlier work this paper cites.
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
Jain, S., Kirk, R., Lubana, E. S., Dick, R. P., Tanaka, H., Grefenstette, E., Rockt"aschel, T., and Krueger, D. S · 2023
Earlier work this paper cites.
Mistral 7B, 2023
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Earlier work this paper cites.
Can generalist foundation models outcompete special-purpose tuning? case study in medicine
Nori, H., Lee, Y. T., Zhang, S., Carignan, D., Edgar, R., Fusi, N., King, N., Larson, J., Li, Y., Liu, W., et al · 2023
Earlier work this paper cites.
Preparedness Framework (Beta), 2023
OpenAI · 2023
Earlier work this paper cites.
Steering llama 2 via contrastive activation addition
Rimsky, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., and Turner, A. M · 2023
Earlier work this paper cites.
Survey of vulnerabilities in large language models revealed by adversarial attacks
Shayegani, E., Mamun, M. A. A., Fu, Y., Zaree, P., Dong, Y., and Abu-Ghazaleh, N · 2023
Cited alongside, same era.
Model evaluation for extreme risks
Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., et al · 2023
Cited alongside, same era.
Llama 2: Open Foundation and Fine-Tuned Chat Models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Cited alongside, same era.
Activation addition: Steering language models without optimization
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training, 2024
Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., Jermyn, A., Askell, A., Radhakrishnan, A., Anil, C., Duvenaud, D., Ganguli, D., Barez, F., Clark, J., Ndousse, K., Sachan, K., Sellitto, M., Sharma, M., DasSarma, N., Grosse, R., Kravec, S., Bai, Y., Witten, Z., Favaro, M., Brauner, J., Karnofsky, H., Christiano, P., Bowman, S. R., Graham, L., Kaplan, J., Mindermann, S., Greenblatt, R., Shlegeris, B., Schiefer, N., and Perez, E · 2024
Later among the works it cites.
Mechanistically Eliciting Latent Behaviors in Language Models
Mack, A. and Turner, A · 2024
Later among the works it cites.
Balancing label quantity and quality for scalable elicitation
Mallen, A. and Belrose, N · 2024
Later among the works it cites.
Trojan Detection in Large Language Models: Insights from The Trojan Detection Challenge
Maloyan, N., Verma, E., Nutfullin, B., and Ashinov, B · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Turner, A. M., Thiergart, L., Leech, G., Udell, D., Vazquez, J. J., Mini, U., and MacDiarmid, M · 2023
Cited alongside, same era.
Poisoning Language Models During Instruction Tuning, 2023
Wan, A., Wallace, E., Shen, S., and Klein, D · 2023
Cited alongside, same era.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models, 2023
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Cited alongside, same era.
Representation engineering: A top-down approach to ai transparency
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al · 2023
Cited alongside, same era.
Agarwal, R., Singh, A., Zhang, L. M., Bohnet, B., Chan, S., Anand, A., Abbas, Z., Nova, A., Co-Reyes, J. D., Chu, E., et al · 2024
Cited alongside, same era.
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
Andriushchenko, M., Croce, F., and Flammarion, N · 2024
Cited alongside, same era.
Many-shot jailbreaking
Anil, C., Durmus, E., Sharma, M., Benton, J., Kundu, S., Batson, J., Rimsky, N., Tong, M., Mu, J., Ford, D., et al · 2024
Cited alongside, same era.
Sabotage evaluations for frontier models
Benton, J., Wagner, M., Christiansen, E., Anil, C., Perez, E., Srivastav, J., Durmus, E., Ganguli, D., Kravec, S., Shlegeris, B., Kaplan, J., Karnofsky, H., Hubinger, E., Grosse, R., Bowman, S. R., and Duvenaud, D · 2024
Cited alongside, same era.
Playing chess with large language models
Carlini, N · 2024
Cited alongside, same era.
Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D · 2024
Later among the works it cites.
Mistral-7B-Instruct-v0.2
Mistral AI Team · 2024
Later among the works it cites.
Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models, 2024
Nevo, S., Lahav, D., Karpur, A., Bar-On, Y., and Bradley, H. A · 2024
Later among the works it cites.
U.S. Artificial Intelligence Safety Institute | | NIST
NIST · 2024
Later among the works it cites.
GPT-4 Technical Report, 2024
OpenAI · 2024
Later among the works it cites.
OpenAI 01 System Card
OpenAI · 2024
Later among the works it cites.
Latent adversarial training improves robustness to persistent harmful behaviors in llms, 2024
Sheshadri, A., Ewart, A., Guo, P., Lynch, A., Wu, C., Hebbar, V., Sleight, H., Stickland, A. C., Perez, E., Hadfield-Menell, D., and Casper, S · 2024
Later among the works it cites.
Analyzing the generalization and reliability of steering vectors, 2024
Tan, D., Chanin, D., Lynch, A., Kanoulas, D., Paige, B., Garriga-Alonso, A., and Kirk, R · 2024
Later among the works it cites.
AI Safety Institute approach to evaluations
UK AISI · 2024
Later among the works it cites.
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
van der Weij, T., Hofst"atter, F., Jaffe, O., Brown, S. F., and Ward, F. R · 2024
Later among the works it cites.
Trading Off Compute in Training and Inference
Villalobos, P. and Atkinson, D · 2024
Later among the works it cites.
repeng, 2024
Vogel, T · 2024
Later among the works it cites.
Robust prompt optimization for defending language models against jailbreaking attacks
Zhou, A., Li, B., and Wang, H · 2024
Later among the works it cites.
Bigcodebench: Benchmarking code generation with diverse function calls and complex instructions
Zhuo, T. Y., Vu, M. C., Chim, J., Hu, H., Yu, W., Widyasari, R., Yusuf, I. N. B., Zhan, H., He, J., Paul, I., et al · 2024
Later among the works it cites.
Improving Alignment and Robustness with Short Circuiting
Zou, A., Phan, L., Wang, J., Duenas, D., Lin, M., Andriushchenko, M., Wang, R., Kolter, Z., Fredrikson, M., and Hendrycks, D · 2024
Later among the works it cites.