Fetching the paper…
Reading the bibliography…
As large language models (LLMs) are becoming more capable and widespread, the study of their failure cases is becoming increasingly important.
"go with the winners" algorithms
Aldous, D. and Vazirani, U · 1994
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts, 2020
Shin, T., Razeghi, Y., IV, R. L. L., Wallace, E., and Singh, S · 2010
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset, 2021
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models, 2021
Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A · 2021
Earlier work this paper cites.
Exploring length generalization in large language models, 2022
Anil, C., Wu, Y., Andreassen, A., Lewkowycz, A., Misra, V., Ramasesh, V., Slone, A., Gur-Ari, G., Dyer, E., and Neyshabur, B · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Earlier work this paper cites.
Red teaming language models with language models, 2022
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Earlier work this paper cites.
Solving math word problems with process- and outcome-based feedback, 2022
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Earlier work this paper cites.
Detecting language model attacks with perplexity, 2023
Alon, G. and Kamfonas, M · 2023
Earlier work this paper cites.
Mart: Improving llm safety with multi-round automatic red-teaming, 2023
Ge, S., Zhou, C., Hou, R., Khabsa, M., Wang, Y.-C., Wang, Q., Han, J., and Mao, Y · 2023
Earlier work this paper cites.
Llama guard: Llm-based input-output safeguard for human-ai conversations, 2023
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., and Khabsa, M · 2023
Earlier work this paper cites.
Language-driven representation learning for robotics, 2023
Karamcheti, S., Nair, S., Chen, A. S., Kollar, T., Finn, C., Sadigh, D., and Liang, P · 2023
Earlier work this paper cites.
Code as policies: Language model programs for embodied control, 2023
Liang, J., Huang, W., Xia, F., Xu, P., Hausman, K., Ichter, B., Florence, P., and Zeng, A · 2023
Earlier work this paper cites.
Let’s verify step by step, 2023
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Earlier work this paper cites.
Lost in the middle: How language models use long contexts, 2023
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Meta AI · 2023
Earlier work this paper cites.
Nemo guardrails: A toolkit for controllable and safe llm applications with programmable rails, 2023
Rebedea, T., Dinu, R., Sreedhar, M., Parisien, C., and Cohen, J · 2023
Earlier work this paper cites.
Gpqa: A graduate-level google-proof q&a benchmark, 2023
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2023
Earlier work this paper cites.
Mathematical discoveries from program search with large language models
Romera-Paredes, B., Barekatain, M., Novikov, A., Balog, M., Kumar, M. P., Dupont, E., Ruiz, F. J. R., Ellenberg, J. S., Wang, P., Fawzi, O., Kohli, P., Fawzi, A., Grochow, J., Lodi, A., Mouret, J.-B., Ringer, T., and Yu, T · 2023
Cited alongside, same era.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery, 2023
Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., and Goldstein, T · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models, 2023
Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M · 2023
Cited alongside, same era.
Improving alignment and robustness with circuit breakers, 2024
Zou, A., Phan, L., Wang, J., Duenas, D., Lin, M., Andriushchenko, M., Wang, R., Kolter, Z., Fredrikson, M., and Hendrycks, D · 2023
Cited alongside, same era.
Fast adversarial attacks on language models in one gpu minute, 2024
Sadasivan, V. S., Saha, S., Sriramanan, G., Kattakinda, P., Chegini, A., and Feizi, S · 2024
Later among the works it cites.
Rainbow teaming: Open-ended generation of diverse adversarial prompts, 2024
Samvelyan, M., Raparthy, S. C., Lupu, A., Hambro, E., Markosyan, A. H., Bhatt, M., Mao, Y., Jiang, M., Parker-Holder, J., Foerster, J., Rocktäschel, T., and Raileanu, R · 2024
Later among the works it cites.
Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024
Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y. K., Wu, Y., and Guo, D · 2024
Later among the works it cites.
Latent adversarial training improves robustness to persistent harmful behaviors in llms, 2024
Sheshadri, A., Ewart, A., Guo, P., Lynch, A., Wu, C., Hebbar, V., Sleight, H., Stickland, A. C., Perez, E., Hadfield-Menell, D., and Casper, S · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ahn, J., Verma, R., Lou, R., Liu, D., Zhang, R., and Yin, W · 2024
Cited alongside, same era.
Liar: Leveraging alignment (best-of-n) to jailbreak llms in seconds, 2024
Beetham, J., Chakraborty, S., Wang, M., Huang, F., Bedi, A. S., and Shah, M · 2024
Cited alongside, same era.
A survey on in-context learning, 2024
Dong, Q., Li, L., Dai, D., Zheng, C., Ma, J., Li, R., Xia, H., Xu, J., Wu, Z., Liu, T., Chang, B., Sun, X., Li, L., and Sui, Z · 2024
Cited alongside, same era.
Stream of search (sos): Learning to search in language, 2024
Gandhi, K., Lee, D., Grand, G., Liu, M., Cheng, W., Sharma, A., and Goodman, N. D · 2024
Cited alongside, same era.
Query-based adversarial prompt generation, 2024
Hayase, J., Borevkovic, E., Carlini, N., Tramèr, F., and Nasr, M · 2024
Cited alongside, same era.
Hughes, J., Price, S., Lynch, A., Schaeffer, R., Barez, F., Koyejo, S., Sleight, H., Jones, E., Perez, E., and Sharma, M · 2024
Cited alongside, same era.
Improved techniques for optimization-based jailbreaking on large language models, 2024
Jia, X., Pang, T., Du, C., Huang, Y., Gu, J., Liu, Y., Cao, X., and Lin, M · 2024
Cited alongside, same era.
Open sesame! universal black box jailbreaking of large language models, 2024
Lapid, R., Langberg, R., and Sipper, M · 2024
Cited alongside, same era.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
A strongreject for empty jailbreaks, 2024
Souly, A., Lu, Q., Bowen, D., Trinh, T., Hsieh, E., Pandey, S., Abbeel, P., Svegliato, J., Emmons, S., Watkins, O., and Toyer, S · 2024
Later among the works it cites.
On the self-verification limitations of large language models on reasoning and planning tasks, 2024
Stechly, K., Valmeekam, K., and Kambhampati, S · 2024
Later among the works it cites.
Chatgpt for robotics: Design principles and model abilities
Vemprala, S. H., Bonatti, R., Bucker, A., and Kapoor, A · 2024
Later among the works it cites.
Dissecting adversarial robustness of multimodal lm agents, 2024
Wu, C. H., Shah, R., Koh, J. Y., Salakhutdinov, R., Fried, D., and Raghunathan, A · 2024
Later among the works it cites.
Efficient adversarial training in llms with continuous attacks, 2024
Xhonneux, S., Sordoni, A., Günnemann, S., Gidel, G., and Schwinn, L · 2024
Later among the works it cites.
Monte carlo tree search boosts reasoning via iterative preference learning, 2024
Xie, Y., Goyal, A., Zheng, W., Kan, M.-Y., Lillicrap, T. P., Kawaguchi, K., and Shieh, M · 2024
Later among the works it cites.
Ovm, outcome-supervised value models for planning in mathematical reasoning, 2024
Yu, F., Gao, A., and Wang, B · 2024
Later among the works it cites.
Textgrad: Automatic "differentiation" via text, 2024
Yuksekgonul, M., Bianchi, F., Boen, J., Liu, S., Huang, Z., Guestrin, C., and Zou, J · 2024
Later among the works it cites.
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., and Shi, W · 2024
Later among the works it cites.
Zheng, X., Lou, J., Cao, B., Wen, X., Ji, Y., Lin, H., Lu, Y., Han, X., Zhang, D., and Sun, L · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI · 2025
Closest in time.
Deliberative alignment: Reasoning enables safer language models, 2025
Guan, M. Y., Joglekar, M., Wallace, E., Jain, S., Barak, B., Helyar, A., Dias, R., Vallone, A., Ren, H., Wei, J., Chung, H. W., Toyer, S., Heidecke, J., Beutel, A., and Glaese, A · 2025
Closest in time.
Jailbreaking to jailbreak, 2025
Kritz, J., Robinson, V., Vacareanu, R., Varjavand, B., Choi, M., Gogov, B., Team, S. R., Yue, S., Primack, W. E., and Wang, Z · 2025
Closest in time.
Towards system 2 reasoning in llms: Learning how to think with meta chain-of-thought, 2025
Xiang, V., Snell, C., Gandhi, K., Albalak, A., Singh, A., Blagden, C., Phung, D., Rafailov, R., Lile, N., Mahan, D., Castricato, L., Franken, J.-P., Haber, N., and Finn, C · 2025
Closest in time.