Fetching the paper…
Reading the bibliography…
AI assistants such as ChatGPT are trained to respond to users by saying, "I am a large language model".
NLTK: The natural language toolkit
Bird, S. and Loper, E · 2004
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Ai safety via debate, 2018
Irving, G., Christiano, P., and Amodei, D · 2018
Earlier work this paper cites.
Long-term planning and situational awareness in openai five, 2019
Raiman, J., Zhang, S., and Wolski, F · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Sensors and AI techniques for situational awareness in autonomous ships: A review
Thombre, S., Zhao, Z., Ramm-Schmidt, H., Vallet García, J. M., Malkamäki, T., Nikolskiy, S., Hammarberg, T., Nuortie, H., H. Bhuiyan, M. Z., Särkkä, S., and Lehtola, V. V · 2020
Earlier work this paper cites.
Fine-tuning language models from human preferences, 2020
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al · 2021
Earlier work this paper cites.
Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover
Cotra, A · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Earlier work this paper cites.
The pile v2, 2022
CarperAI · 2022
Earlier work this paper cites.
Self-Locating Beliefs
Egan, A. and Titelbaum, M. G · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Earlier work this paper cites.
Discovering language model behaviors with model-written evaluations, 2022
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., Jones, A., Chen, A., Mann, B., Israel, B., Seethor, B., McKinnon, C., Olah, C., Yan, D., Amodei, D., Amodei, D., Drain, D., Li, D., Tran-Johnson, E., Khundadze, G., Kernion, J., Landis, J., Kerr, J., Mueller, J., Hyun, J., Landau, J., Ndousse, K., Goldberg, L., Lovitt, L., Lucas, M., Sellitto, M., Zhang, M., Kingsland, N., Elhage, N., Joseph, N., Mercado, N., DasSarma, N., Rausch, O., Larson, R., McCandlish, S., Johnston, S., Kravec, S., El Showk, S., Lanham, T., Telleen-Lawton, T., Brown, T., Henighan, T., Hume, T., Bai, Y., Hatfield-Dodds, Z., Clark, J., Bowman, S. R., Askell, A., Grosse, R., Hernandez, D., Ganguli, D., Hubinger, E., Schiefer, N., and Kaplan, J · 2022
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Earlier work this paper cites.
Self-consistency of large language models under ambiguity
Bartsch, H., Jorgensen, O., Rosati, D., Hoelscher-Obermaier, J., and Pfau, J · 2023
Earlier work this paper cites.
Taken out of context: On measuring situational awareness in LLMs, 2023
Berglund, L., Stickland, A. C., Balesni, M., Kaufmann, M., Tong, M., Korbak, T., Kokotajlo, D., and Evans, O · 2023
Earlier work this paper cites.
Two failures of self-consistency in the multi-step reasoning of llms, 2023
Chen, A., Phang, J., Parrish, A., Padmakumar, V., Zhao, C., Bowman, S. R., and Cho, K · 2023
Earlier work this paper cites.
Mamba: Linear-time sequence modeling with selective state spaces
Gu, A. and Dao, T · 2023
Earlier work this paper cites.
A survey of reinforcement learning from human feedback, 2023
Kaufmann, T., Weng, P., Bengs, V., and Hüllermeier, E · 2023
Earlier work this paper cites.
Fantom: A benchmark for stress-testing machine theory of mind in interactions, 2023
Kim, H., Sclar, M., Zhou, X., Bras, R. L., Kim, G., Choi, Y., and Sap, M · 2023
Cited alongside, same era.
Evaluating language-model agents on realistic autonomous tasks
Kinniment, M., Jun, L., Sato, K., Du, H., Goodrich, B., Hasin, M., Chan, L., Miles, L. H., Lin, T. R., Wijk, H., Burget, J., Ho, A., Barnes, E., and Christiano, P. F · 2023
Cited alongside, same era.
Understanding the effects of rlhf on llm generalisation and diversity
Kirk, R., Mediratta, I., Nalmpantis, C., Luketina, J., Hambro, E., Grefenstette, E., and Raileanu, R · 2023
Cited alongside, same era.
Out-of-context meta-learning in large language models
Krasheninnikov, D., Krasheninnikov, E., and Krueger, D · 2023
Cited alongside, same era.
Holistic evaluation of language models, 2023
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., Ré, C., Acosta-Navas, D., Hudson, D. A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F., Ren, H., Yao, H., Wang, J., Santhanam, K., Orr, L., Zheng, L., Yuksekgonul, M., Suzgun, M., Kim, N., Guha, N., Chatterji, N., Khattab, O., Henderson, P., Huang, Q., Chi, R., Xie, S. M., Santurkar, S., Ganguli, S., Hashimoto, T., Icard, T., Zhang, T., Chaudhary, V., Wang, W., Li, X., Mai, Y., Zhang, Y., and Koreeda, Y · 2023
Anthropic console, 2024b
Anthropic · 2024
Closest in time.
Claude, 2024c
Anthropic · 2024
Closest in time.
Foundational challenges in assuring alignment and safety of large language models, 2024
Anwar, U., Saparov, A., Rando, J., Paleka, D., Turpin, M., Hase, P., Lubana, E. S., Jenner, E., Casper, S., Sourbut, O., Edelman, B. L., Zhang, Z., Günther, M., Korinek, A., Hernandez-Orallo, J., Hammond, L., Bigelow, E., Pan, A., Langosco, L., Korbak, T., Zhang, H., Zhong, R., hÉigeartaigh, S. Ó., Recchia, G., Corsi, G., Chan, A., Anderljung, M., Edwards, L., Bengio, Y., Chen, D., Albanie, S., Maharaj, T., Foerster, J., Tramer, F., He, H., Kasirzadeh, A., Choi, Y., and Krueger, D · 2024
Closest in time.
arxiv dataset, 2024
arXiv.org submitters · 2024
Closest in time.
Here is claude 3’s system prompt …
Askell, A · 2024
Closest in time.
Chatbot arena: An open platform for evaluating llms by human preference, 2024
Chiang, W.-L., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhang, H., Zhu, B., Jordan, M., Gonzalez, J. E., and Stoica, I · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bolaa: Benchmarking and orchestrating llm-augmented autonomous agents, 2023
Liu, Z., Yao, W., Zhang, J., Xue, L., Heinecke, S., Murthy, R., Feng, Y., Chen, Z., Niebles, J. C., Arpit, D., Xu, R., Mui, P., Wang, H., Xiong, C., and Savarese, S · 2023
Cited alongside, same era.
Introspective capabilities in large language models
Long, R · 2023
Cited alongside, same era.
Tell, don’t show: Declarative facts influence how llms generalize
Meinke, A. and Evans, O · 2023
Cited alongside, same era.
Gaia: a benchmark for general ai assistants, 2023
Mialon, G., Fourrier, C., Swift, C., Wolf, T., LeCun, Y., and Scialom, T · 2023
Cited alongside, same era.
The alignment problem from a deep learning perspective, 2023
Ngo, R., Chan, L., and Mindermann, S · 2023
Cited alongside, same era.
Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark
Pan, A., Chan, J. S., Zou, A., Li, N., Basart, S., Woodside, T., Zhang, H., Emmons, S., and Hendrycks, D · 2023
Cited alongside, same era.
Towards evaluating ai systems for moral status using self-reports
Perez, E. and Long, R · 2023
Cited alongside, same era.
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., et al · 2024
Closest in time.
Towards a situational awareness benchmark for LLMs
Laine, R., Meinke, A., and Evans, O · 2024
Closest in time.
The wmdp benchmark: Measuring and reducing malicious use with unlearning, 2024
Li, N., Pan, A., Gopal, A., Yue, S., Berrios, D., Gatti, A., Li, J. D., Dombrowski, A.-K., Goel, S., Phan, L., Mukobi, G., Helm-Burger, N., Lababidi, R., Justen, L., Liu, A. B., Chen, M., Barrass, I., Zhang, O., Zhu, X., Tamirisa, R., Bharathi, B., Khoja, A., Zhao, Z., Herbert-Voss, A., Breuer, C. B., Zou, A., Mazeika, M., Wang, Z., Oswal, P., Liu, W., Hunt, A. A., Tienken-Harder, J., Shih, K. Y., Talley, K., Guan, J., Kaplan, R., Steneker, I., Campbell, D., Jokubaitis, B., Levinson, A., Wang, J., Qian, W., Karmakar, K. K., Basart, S., Fitz, S., Levine, M., Kumaraguru, P., Tupakula, U., Varadharajan, V., Shoshitaishvili, Y., Ba, J., Esvelt, K. M., Wang, A., and Hendrycks, D · 2024
Closest in time.
Inverse scaling: When bigger isn’t better, 2024
McKenzie, I. R., Lyzhov, A., Pieler, M., Parrish, A., Mueller, A., Prabhu, A., McLean, E., Kirtland, A., Ross, A., Liu, A., Gritsevskiy, A., Wurgaft, D., Kauffman, D., Recchia, G., Liu, J., Cavanagh, J., Weiss, M., Huang, S., Droid, T. F., Tseng, T., Korbak, T., Shen, X., Zhang, Y., Zhou, Z., Kim, N., Bowman, S. R., and Perez, E · 2024
Closest in time.
Secret collusion among generative ai agents, 2024
Motwani, S. R., Baranchuk, M., Strohmeier, M., Bolina, V., Torr, P. H. S., Hammond, L., and de Witt, C. S · 2024
Closest in time.
GitHub - openai/evals: Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks. — github.com
OpenAI · 2024
Closest in time.
Chatgpt, 2024a
OpenAI · 2024
Closest in time.
Platform, 2024b
OpenAI · 2024
Closest in time.
LM Situational Awareness, Evaluation Proposal: Violating Imitation, April 2023
Pfau, J · 2024
Closest in time.
Replicate - run ai with an api
Replicate · 2024
Closest in time.
Can language models explain their own classification behavior?, 2024
Sherburn, D., Chughtai, B., and Evans, O · 2024
Closest in time.
Treutlein, J., Choi, D., Betley, J., Anil, C., Marks, S., Grosse, R. B., and Evans, O · 2024
Closest in time.
Ai sandbagging: Language models can strategically underperform on evaluations
van der Weij, T., Hofstätter, F., Jaffe, O., Brown, S., and Ward, F. R · 2024
Closest in time.
Wildchat: 1m chatGPT interaction logs in the wild
Zhao, W., Ren, X., Hessel, J., Cardie, C., Choi, Y., and Deng, Y · 2024
Closest in time.