Fetching the paper…
Reading the bibliography…
Language models exhibit complex, diverse behaviors when prompted with free-form text, making it difficult to characterize the space of possible outputs.
An algorithm for quadratic programming
Frank, M. and Wolfe, P · 1956
Earlier work this paper cites.
An introduction to variational methods for graphical models
Jordan, M. I., Ghahramani, Z., Jaakkola, T. S., and Saul, L. K · 1999
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
Duality between subgradient and conditional gradient methods
Bach, F. R · 2012
Earlier work this paper cites.
Diagnostic and Statistical Manual of Mental Disorders (DSM-5®)
APA · 2013
Earlier work this paper cites.
Variational inference: A review for statisticians
Blei, D. M., Kucukelbir, A., and McAuliffe, J. D · 2017
Earlier work this paper cites.
Enabling language models to fill in the blanks
Donahue, C., Lee, M., and Liang, P · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
von Werra, L., Belkada, Y., Tunstall, L., Beeching, E., Thrush, T., Lambert, N., Huang, S., Rasul, K., and Gallouédec, Q · 2020
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Earlier work this paper cites.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Earlier work this paper cites.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Cited alongside, same era.
Enhancing chat language models by scaling high-quality instructional conversations
Ding, N., Chen, Y., Xu, B., Qin, Y., Zheng, Z., Hu, S., Liu, Z., Sun, M., and Zhou, B · 2023
Cited alongside, same era.
Morris, J. X., Zhao, W., Chiu, J. T., Shmatikov, V., and Rush, A. M · 2023
Cited alongside, same era.
Chenghaomou/text-dedup: Reference snapshot, September 2023
Mou, C., Ha, C., Enevoldsen, K., and Liu, P · 2023
Cited alongside, same era.
Eliciting language model behaviors using reverse language models
Pfau, J., Infanger, A., Sheshadri, A., Panda, A., Michael, J., and Huebner, C · 2023
Jailbreaking leading safety-aligned llms with simple adaptive attacks
Andriushchenko, M., Croce, F., and Flammarion, N · 2024
Later among the works it cites.
Curiosity-driven red-teaming for large language models
Hong, Z.-W., Shenfeld, I., Wang, T.-H., Chuang, Y.-S., Pareja, A., Glass, J. R., Srivastava, A., and Agrawal, P · 2024
Later among the works it cites.
Llm defenses are not robust to multi-turn human jailbreaks yet
Li, N., Han, Z., Steneker, I., Primack, W., Goodside, R., Zhang, H., Wang, Z., Menghini, C., and Yue, S · 2024
Later among the works it cites.
AutoDAN: Generating stealthy jailbreak prompts on aligned large language models
Liu, X., Xu, N., Chen, M., and Xiao, C · 2024
Later among the works it cites.
Openai model specification
OpenAI · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., et al · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Cited alongside, same era.
A conversation with bing’s chatbot left me deeply unsettled
Roose, K · 2023
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Cited alongside, same era.
Reward collapse in aligning large language models, 2023
Song, Z., Cai, T., Lee, J. D., and Su, W. J · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2023
Cited alongside, same era.
Penedo, G., Kydlíček, H., Ben Allal, L., Lozhkov, A., Mitchell, M., Raffel, C., Von Werra, L., and Wolf, T · 2024
Later among the works it cites.
A multimodal automated interpretability agent
Shaham, T. R., Schwettmann, S., Wang, F., Rajaram, A., Hernandez, E., Andreas, J., and Torralba, A · 2024
Later among the works it cites.
Factuality of large language models: A survey, 2024
Wang, Y., Wang, M., Manzoor, M. A., Liu, F., Georgiev, G., Das, R. J., and Nakov, P · 2024
Later among the works it cites.
Wildchat: 1m chatGPT interaction logs in the wild
Zhao, W., Ren, X., Hessel, J., Cardie, C., Choi, Y., and Deng, Y · 2024
Later among the works it cites.
Infinite backrooms: Dreams of an electric mind, 2024
Ayrey, A · 2025
Closest in time.