Fetching the paper…
Reading the bibliography…
Current large language models have dangerous capabilities, which are likely to become more problematic in the future.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017) · 2017
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. (2020) · 2020
Earlier work this paper cites.
Distinguishing three alignment taxes
Leike, J. (2022) · 2022
Earlier work this paper cites.
Discovering language model behaviors with model-written evaluations
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al. (2022) · 2022
Earlier work this paper cites.
Language models are better than humans at next-token prediction
Shlegeris, B., Roger, F., Chan, L., and McLean, E. (2022) · 2022
Earlier work this paper cites.
Natural language processing with transformers
Tunstall, L., Von Werra, L., and Wolf, T. (2022) · 2022
Earlier work this paper cites.
Open problems and fundamental limitations of reinforcement learning from human feedback
Casper, S., Davies, X., Shi, C., Gilbert, T. K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., et al. (2023) · 2023
Earlier work this paper cites.
Fast machine unlearning without retraining through selective synaptic dampening
Foster, J., Schoepf, S., and Brintrup, A. (2023) · 2023
Cited alongside, same era.
Improving activation steering in language models with mean-centring
Jorgensen, O., Cope, D., Schoots, N., and Shanahan, M. (2023) · 2023
Cited alongside, same era.
Inference-time intervention: Eliciting truthful answers from a language model
Li, K., Patel, O., Viégas, F., Pfister, H., and Wattenberg, M. (2023) · 2023
Cited alongside, same era.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G. (2023) · 2023
Cited alongside, same era.
Copy suppression: Comprehensively understanding an attention head
Steering llama 2 via contrastive activation addition
Rimsky, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., and Turner, A. M. (2023) · 2023
Later among the works it cites.
Exploring the landscape of machine unlearning: A survey and taxonomy
Shaik, T., Tao, X., Xie, H., Li, L., Zhu, X., and Li, Q. (2023) · 2023
Later among the works it cites.
Model evaluation for extreme risks
Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., Ho, L., Siddarth, D., Avin, S., Hawkins, W., Kim, B., Gabriel, I., Bolina, V., Clark, J., Bengio, Y., Christiano, P., and Dafoe, A. (2023) · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al. (2023) · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
McDougall, C., Conmy, A., Rushing, C., McGrath, T., and Nanda, N. (2023) · 2023
Cited alongside, same era.
Dissecting large language models
Pochinkov, N. and Schoots, N. (2023) · 2023
Cited alongside, same era.
Turner, A., Thiergart, L., Udell, D., Leech, G., Mini, U., and MacDiarmid, M. (2023) · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al. (2023) · 2023
Later among the works it cites.