Fetching the paper…
Reading the bibliography…
As the capabilities of large machine learning models continue to grow, and as the autonomy afforded to such models continues to expand, the spectre of a new adversary looms: the models themselves.
Unsolved problems in ml safety
Hendrycks, D., Carlini, N., Schulman, J., and Steinhardt, J · 2021
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems, 2021
Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., and Garrabrant, S · 2021
Earlier work this paper cites.
Aging with grace: Lifelong model editing with discrete key-value adaptors
Hartvigsen, T., Sankaranarayanan, S., Palangi, H., Kim, Y., and Ghassemi, M · 2022
Earlier work this paper cites.
A modern self-referential weight matrix that learns to modify itself
Irie, K., Schlag, I., Csordás, R., and Schmidhuber, J · 2022
Earlier work this paper cites.
Memory-based model editing at scale
Mitchell, E., Lin, C., Bosselut, A., Manning, C. D., and Finn, C · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Earlier work this paper cites.
Linear adversarial concept erasure, 2022
Ravfogel, S., Twiton, M., Goldberg, Y., and Cotterell, R · 2022
Earlier work this paper cites.
Self-instruct: Aligning language model with self generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2022
Earlier work this paper cites.
Semdedup: Data-efficient learning at web-scale through semantic deduplication
Abbas, A., Tirumala, K., Simig, D., Ganguli, S., and Morcos, A. S · 2023
Earlier work this paper cites.
Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs
Akyürek, A. F., Akyürek, E., Madaan, A., Kalyan, A., Clark, P., Wijaya, D., and Tandon, N · 2023
Cited alongside, same era.
Poisoning web-scale training datasets is practical
Carlini, N., Jagielski, M., Choquette-Choo, C. A., Paleka, D., Pearce, W., Anderson, H., Terzis, A., Thomas, K., and Tramèr, F · 2023
Cited alongside, same era.
Facade: A framework for adversarial circuit anomaly detection and evaluation, 2023
Carranza, A., Pai, D., Tandon, A., Schaeffer, R., and Koyejo, S · 2023
Cited alongside, same era.
Mechanistic anomaly detection and elk, 2022a
Christiano, P · 2023
Cited alongside, same era.
Can we efficiently distinguish different mechanisms?, 2022b
Christiano, P · 2023
Cited alongside, same era.
The alignment problem from a deep learning perspective, 2023
Ngo, R., Chan, L., and Mindermann, S · 2023
Closest in time.
Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark, 2023
Pan, A., Shern, C. J., Zou, A., Li, N., Basart, S., Woodside, T., Ng, J., Zhang, H., Emmons, S., and Hendrycks, D · 2023
Closest in time.
Peng, B., Li, C., He, P., Galley, M., and Gao, J · 2023
Closest in time.
Are emergent abilities of large language models a mirage?, 2023
Schaeffer, R., Miranda, B., and Koyejo, S · 2023
Closest in time.
Model evaluation for extreme risks, 2023
Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., Ho, L., Siddarth, D., Avin, S., Hawkins, W., Kim, B., Gabriel, I., Bolina, V., Clark, J., Bengio, Y., Christiano, P., and Dafoe, A · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Detecting edit failures in large language models: An improved specificity benchmark
Hoelscher-Obermaier, J., Persson, J., Kran, E., Konstas, I., and Barez, F · 2023
Cited alongside, same era.
Aligning text-to-image models using human feedback, 2023
Lee, K., Liu, H., Ryu, M., Watkins, O., Du, Y., Boutilier, C., Abbeel, P., Ghavamzadeh, M., and Gu, S. S · 2023
Cited alongside, same era.
Internet explorer: Targeted representation learning on the open web
Li, A. C., Brown, E., Efros, A. A., and Pathak, D · 2023
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022a
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J
Cited in the paper.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al
Cited in the paper.
Eliminating meta optimization through self-referential meta learning
Kirsch, L. and Schmidhuber, J
Cited in the paper.
Self-referential meta learning
Kirsch, L. and Schmidhuber, J
Cited in the paper.
Principle-driven self-alignment of language models from scratch with minimal human supervision
Sun, Z., Shen, Y., Zhou, Q., Zhang, H., Chen, Z., Cox, D., Yang, Y., and Gan, C · 2023
Closest in time.
Doremi: Optimizing data mixtures speeds up language model pretraining
Xie, S. M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P., Le, Q. V., Ma, T., and Yu, A. W · 2023
Closest in time.
Wizardlm: Empowering large language models to follow complex instructions
Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., and Jiang, D · 2023
Closest in time.