Fetching the paper…
Reading the bibliography…
As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes.
A research agenda: Dynamic models to defend against correlated attacks
Goodfellow, I. (2019) · 1903
Earlier work this paper cites.
Fundamentals of data mining in genomics and proteomics
Dubitzky, W., Granzow, M., and Berrar, D. P. (2007) · 2007
Earlier work this paper cites.
Red Team: How to succeed by thinking like the enemy
Zenko, M. (2015) · 2015
Earlier work this paper cites.
Red teams
Christiano, P. (2016) · 2016
Earlier work this paper cites.
Irving, G., Christiano, P., and Amodei, D. (2018) · 2018
Earlier work this paper cites.
Professional Red Teaming: Conducting Successful Cybersecurity Engagements
Oakley, J. G. (2019) · 2019
Earlier work this paper cites.
Measuring coding challenge competence with apps
Hendrycks, D., Basart, S., Kadavath, S., Mazeika, M., Arora, A., Guo, E., Burns, C., Puranik, S., He, H., Song, D., et al. (2021) · 2021
Earlier work this paper cites.
A review on text steganography techniques
Majeed, M. A., Sulaiman, R., Shukur, Z., and Hasan, M. K. (2021) · 2021
Earlier work this paper cites.
Measuring progress on scalable oversight for large language models
Bowman, S. R., Hyun, J., Perez, E., Chen, E., Pettit, C., Heiner, S., Lukošiūtė, K., Askell, A., Jones, A., Chen, A., et al. (2022) · 2022
Earlier work this paper cites.
The alignment problem from a deep learning perspective
Ngo, R., Chan, L., and Mindermann, S. (2022) · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Earlier work this paper cites.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G. (2022) · 2022
Cited alongside, same era.
Steganography in chain of thought reasoning
Ray, A. (2022) · 2022
Cited alongside, same era.
Self-critiquing models for assisting human evaluators
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J. (2022) · 2022
Cited alongside, same era.
Scheming ais: Will ais fake alignment during training in order to get power?
Carlsmith, J. (2023) · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., et al. (2023) · 2023
The new bing: Our approach to responsible ai
Microsoft (2023) · 2023
Closest in time.
Testing language model agents safely in the wild
Naihin, S., Atkinson, D., Green, M., Hamadi, M., Swift, C., Schonholtz, D., Kalai, A. T., and Bau, D. (2023) · 2023
Closest in time.
Preventing language models from hiding their reasoning
Roger, F. and Greenblatt, R. (2023) · 2023
Closest in time.
Model evaluation for extreme risks
Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., et al. (2023) · 2023
Closest in time.
Untrusted smart models and trusted dumb models
Shlegeris, B. (2023) · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The case for ensuring that powerful ais are controlled
Greenblatt, R. and Shlegeris, B. (2023) · 2023
Cited alongside, same era.
When can we trust model evaluations?
Hubinger, E. (2023) · 2023
Cited alongside, same era.
Model organisms of misalignment: The case for a new pillar of alignment research
Hubinger, E., Schiefer, N., Denison, C., and Perez, E. (2023) · 2023
Cited alongside, same era.
Ai alignment: A comprehensive survey
Ji, J., Qiu, T., Chen, B., Zhang, B., Lou, H., Wang, K., Duan, Y., He, Z., Zhou, J., Zhang, Z., et al. (2023) · 2023
Cited alongside, same era.
Introducing superalignment
Leike, J. and Sutskever, I. (2023) · 2023
Cited alongside, same era.
Debate helps supervise unreliable experts
Michael, J., Mahdi, S., Rein, D., Petty, J., Dirani, J., Padmakumar, V., and Bowman, S. R. (2023) · 2023
Cited alongside, same era.
A watermark for large language models
Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., and Goldstein, T. (2023a)
Cited in the paper.
Meta-level adversarial evaluation of oversight techniques might allow robust measurement of their adequacy
Shlegeris, B. and Greenblatt, R. (2023) · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023) · 2023
Closest in time.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J. (2023) · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M. (2023) · 2023
Closest in time.
Notes on control evaluations for safety cases
Greenblatt, R., Shlegeris, B., and Roger, F. (2024) · 2024
Closest in time.