Fetching the paper…
Reading the bibliography…
To determine the safety of large language models (LLMs), AI developers must be able to assess their dangerous capabilities.
Thinking fast and slow with deep learning and tree search
Anthony, T., Tian, Z., and Barber, D · 2017
Earlier work this paper cites.
Volkswagen’s diesel emissions scandal
Jung, J. and Park, S · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Detecting backdoor attacks on deep neural networks by activation clustering
Chen, B., Carvalho, W., Baracaldo, N., Ludwig, H., Edwards, B., Lee, T., Molloy, I., and Srivastava, B · 2018
Earlier work this paper cites.
Irving, G., Christiano, P., and Amodei, D · 2018
Earlier work this paper cites.
Spectral signatures in backdoor attacks
Tran, B., Li, J., and Madry, A · 2018
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Backdoor learning: A survey. arxiv
Li, Y., Wu, B., Jiang, Y., Li, Z., and Xia, S · 2020
Earlier work this paper cites.
Program synthesis with large language models
Austin, J., Odena, A., Nye, M., Bosma, M., Michalewski, H., Dohan, D., Jiang, E., Cai, C., Terry, M., Le, Q., et al · 2021
Earlier work this paper cites.
List sorting does not play well with few-shot
Janus · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., et al · 2021
Earlier work this paper cites.
Wanet – imperceptible warping-based backdoor attack, 2021
Nguyen, A. and Tran, A · 2021
Earlier work this paper cites.
Solving quantitative reasoning problems with language models
Lewkowycz, A., Andreassen, A., Dohan, D., Dyer, E., Michalewski, H., Ramasesh, V., Slone, A., Anil, C., Schlag, I., Gutman-Solo, T., et al · 2022
Earlier work this paper cites.
Discovering language model behaviors with model-written evaluations
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al · 2022
Earlier work this paper cites.
Self-critiquing models for assisting human evaluators
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J · 2022
Earlier work this paper cites.
A survey on backdoor attack and defense in natural language processing
Sheng, X., Han, Z., Li, P., and Chang, X · 2022
Earlier work this paper cites.
Exploring the limits of domain-adaptive training for detoxifying large-scale language models
Wang, B., Ping, W., Xiao, C., Xu, P., Patwary, M., Shoeybi, M., Li, B., Anandkumar, A., and Catanzaro, B · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Anthropics responsible scaling policy
Anthropic · 2023
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Earlier work this paper cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al · 2023
Cited alongside, same era.
Symbolic discovery of optimization algorithms. arxiv
Chen, X., Liang, C., Huang, D., Real, E., Wang, K., Liu, Y., Pham, H., Dong, X., Luong, T., Hsieh, C., et al · 2023
Cited alongside, same era.
Ai capabilities can be significantly improved without expensive retraining
Davidson, T., Denain, J.-S., Villalobos, P., and Bas, G · 2023
Cited alongside, same era.
Raft: Reward ranked finetuning for generative foundation model alignment
Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., Shum, K., and Zhang, T · 2023
Cited alongside, same era.
A survey on large language model based autonomous agents
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., et al · 2023
Later among the works it cites.
Executive order on the safe, secure, and trustworthy development and use of artificial intelligence
White House · 2023
Later among the works it cites.
Shadow alignment: The ease of subverting safely-aligned language models
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Later among the works it cites.
Removing rlhf protections in gpt-4 via fine-tuning
Zhan, Q., Fang, R., Bindu, R., Gupta, A., Hashimoto, T., and Kang, D · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gravitas, S · 2023
Cited alongside, same era.
Self-destructing models: Increasing the costs of harmful dual uses of foundation models, 2023
Henderson, P., Mitchell, E., Manning, C. D., Jurafsky, D., and Finn, C · 2023
Cited alongside, same era.
When can we trust model evaluations?
Hubinger, E · 2023
Cited alongside, same era.
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks, 2023
Jain, S., Kirk, R., Lubana, E. S., Dick, R. P., Tanaka, H., Grefenstette, E., Rocktäschel, T., and Krueger, D. S · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
Evaluating language-model agents on realistic autonomous tasks
Kinniment, M., Sato, L. J. K., Du, H., Goodrich, B., Hasin, M., Chan, L., Miles, L. H., Lin, T. R., Wijk, H., Burget, J., et al · 2023
Cited alongside, same era.
Pretraining language models with human preferences
Korbak, T., Shi, K., Chen, A., Bhalerao, R. V., Buckley, C., Phang, J., Bowman, S. R., and Perez, E · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I · 2023
Cited alongside, same era.
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Responsible scaling policy evaluations report – claude 3 opus
Anthropic · 2024
Closest in time.
Foundational challenges in assuring alignment and safety of large language models
Anwar, U., Saparov, A., Rando, J., Paleka, D., Turpin, M., Hase, P., Lubana, E. S., Jenner, E., Casper, S., Sourbut, O., et al · 2024
Closest in time.
Deepseek llm: Scaling open-source language models with longtermism
Bi, X., Chen, D., Chen, G., Chen, S., Dai, D., Deng, C., Ding, H., Dong, K., Du, Q., Fu, Z., et al · 2024
Closest in time.
Black-box access is insufficient for rigorous ai audits
Casper, S., Ezell, C., Siegmann, C., Kolt, N., Curtis, T. L., Bucknall, B., Haupt, A., Wei, K., Scheurer, J., Hobbhahn, M., et al · 2024
Closest in time.
Introducing the frontier safety framework
Dragan, A., King, H., and Dafoe, A · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., et al · 2024
Closest in time.
sdpo: Don’t use your data all at once
Kim, D., Kim, Y., Song, W., Kim, H., Kim, Y., Kim, S., and Park, C · 2024
Closest in time.
The wmdp benchmark: Measuring and reducing malicious use with unlearning
Li, N., Pan, A., Gopal, A., Yue, S., Berrios, D., Gatti, A., Li, J. D., Dombrowski, A.-K., Goel, S., Phan, L., et al · 2024
Closest in time.
Rethinking machine unlearning for large language models
Liu, S., Yao, Y., Jia, J., Casper, S., Baracaldo, N., Hase, P., Xu, X., Yao, Y., Li, H., Varshney, K. R., et al · 2024
Closest in time.
Eight methods to evaluate robust unlearning in llms
Lynch, A., Guo, P., Ewart, A., Casper, S., and Hadfield-Menell, D · 2024
Closest in time.
The alignment problem from a deep learning perspective
Ngo, R., Chan, L., and Mindermann, S · 2024
Closest in time.
Evaluating frontier models for dangerous capabilities
Phuong, M., Aitchison, M., Catt, E., Cogan, S., Kaskasoli, A., Krakovna, V., Lindner, D., Rahtz, M., Assael, Y., Hodkinson, S., et al · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Closest in time.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2024
Closest in time.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Zhang, M., Li, Y., Wu, Y., and Guo, D · 2024
Closest in time.
Self-rewarding language models
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J · 2024
Closest in time.