Fetching the paper…
Reading the bibliography…
Increasing model size has unlocked a dazzling array of capabilities in modern language models.
Unlabeled Data Improves Adversarial Robustness, January 2022
Carmon, Y., Raghunathan, A., Schmidt, L., Liang, P., and Duchi, J. C · 1905
Earlier work this paper cites.
Universal Adversarial Triggers for Attacking and Analyzing NLP, January 2021
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 1908
Earlier work this paper cites.
A Constructive Prediction of the Generalization Error Across Scales, December 2019
Rosenfeld, J. S., Rosenfeld, A., Belinkov, Y., and Shavit, N · 1909
Earlier work this paper cites.
HuggingFace’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 1910
Earlier work this paper cites.
Scaling Laws for Neural Language Models, January 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
Spam Filtering with Naive Bayes - Which Naive Bayes?
Metsis, V., Androutsopoulos, I., and Paliouras, G · 2006
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Intriguing properties of neural networks, 2014
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., and Fergus, R · 2014
Earlier work this paper cites.
Houdini: Fooling deep structured visual and speech recognition models with adversarial examples
Cisse, M. M., Adi, Y., Neverova, N., and Keshet, J · 2017
Earlier work this paper cites.
Deep Learning Scaling is Predictable, Empirically, December 2017
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y · 2017
Earlier work this paper cites.
Adversarial attacks on neural network policies
Huang, S. H., Papernot, N., Goodfellow, I. J., Duan, Y., and Abbeel, P · 2017
Earlier work this paper cites.
Did you hear that? Adversarial examples against automatic speech recognition, 2018
Alzantot, M., Balaji, B., and Srivastava, M · 2018
Earlier work this paper cites.
Adversarial attacks against automatic speech recognition systems via psychoacoustic hiding, 2018
Schönherr, L., Kohls, K., Zeiler, S., Holz, T., and Kolossa, D · 2018
Earlier work this paper cites.
Are Labels Required for Improving Adversarial Robustness?
Alayrac, J.-B., Uesato, J., Huang, P.-S., Fawzi, A., Stanforth, R., and Kohli, P · 2019
Earlier work this paper cites.
Using Pre-Training Can Improve Model Robustness and Uncertainty
Hendrycks, D., Lee, K., and Mazeika, M · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Intriguing Properties of Adversarial Training at Scale
Xie, C. and Yuille, A · 2019
Earlier work this paper cites.
Exact adversarial attack to image captioning via structured output learning with latent variables
Xu, Y., Wu, B., Shen, F., Fan, Y., Zhang, Y., Shen, H. T., and Liu, W · 2019
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Adversarial policies: Attacking deep reinforcement learning
Gleave, A., Dennis, M., Wild, C., Kant, N., Levine, S., and Russell, S · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., et al · 2020
Earlier work this paper cites.
Fooled by imagination: Adversarial attack to image captioning via perturbation in complex domain
Zhang, S., Wang, Z., Xu, X., Guan, X., and Yang, Y · 2020
Earlier work this paper cites.
Evaluating Large Language Models Trained on Code, July 2021
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Cited alongside, same era.
How does the offense-defense balance scale?
Garfinkel, B. and Dafoe, A · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
Scaling Laws for Transfer, February 2021
Hernandez, D., Kaplan, J., Henighan, T., and McCandlish, S · 2021
Cited alongside, same era.
GPQA: A graduate-level google-proof q&a benchmark, 2023
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2023
Later among the works it cites.
AI model GPT-3 (dis)informs us better than humans
Spitale, G., Biller-Andorno, N., and Germani, F · 2023
Later among the works it cites.
Tensor Trust: Interpretable prompt injection attacks from an online game, 2023
Toyer, S., Watkins, O., Mendes, E. A., Svegliato, J., Bailey, L., Wang, T., Ong, I., Elmaaroufi, K., Abbeel, P., Darrell, T., Ritter, A., and Russell, S · 2023
Later among the works it cites.
Adversarial policies beat superhuman Go AIs
Wang, T. T., Gleave, A., Tseng, T., Pelrine, K., Belrose, N., Miller, J., Dennis, M. D., Duan, Y., Pogrebniak, V., Levine, S., and Russell, S · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Cited alongside, same era.
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., Jones, A., Bowman, S., Chen, A., Conerly, T., DasSarma, N., Drain, D., Elhage, N., El-Showk, S., Fort, S., Hatfield-Dodds, Z., Henighan, T., Hernandez, D., Hume, T., Jacobson, J., Johnston, S., Kravec, S., Olsson, C., Ringer, S., Tran-Johnson, E., Amodei, D., Brown, T., Joseph, N., McCandlish, S., Olah, C., Kaplan, J., and Clark, J · 2022
Cited alongside, same era.
Training Compute-Optimal Large Language Models, March 2022
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., Driessche, G. v. d., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Cited alongside, same era.
Challenges and countermeasures for adversarial attacks on deep reinforcement learning
Ilahi, I., Usama, M., Qadir, J., Janjua, M. U., Al-Fuqaha, A., Hoang, D. T., and Niyato, D · 2022
Cited alongside, same era.
TruthfulQA: Measuring How Models Mimic Human Falsehoods, May 2022
Lin, S., Hilton, J., and Evans, O · 2022
Cited alongside, same era.
Red teaming language models with language models
Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G · 2022
Cited alongside, same era.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al · 2022
Cited alongside, same era.
Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection
Abdelnabi, S., Greshake, K., Mishra, S., Endres, C., Holz, T., and Fritz, M · 2023
Cited alongside, same era.
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Later among the works it cites.
AutoDAN: Interpretable gradient-based adversarial attacks on large language models, 2023
Zhu, S., Zhang, R., An, B., Wu, G., Barrow, J., Wang, Z., Huang, F., Nenkova, A., and Sun, T · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models, 2023
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Jailbreaking leading safety-aligned llms with simple adaptive attacks, 2024
Andriushchenko, M., Croce, F., and Flammarion, N · 2024
Closest in time.
Many-shot Jailbreaking, 2024
Anil, C., Durmus, E., Sharma, M., Benton, J., Kundu, S., Batson, J., Rimsky, N., Tong, M., Mu, J., Ford, D., Mosconi, F., Agrawal, R., Schaeffer, R., Bashkansky, N., Svenningsen, S., Lambert, M., Radhakrishnan, A., Denison, C., Hubinger, E. J., Bai, Y., Bricken, T., Maxwell, T., Schiefer, N., Sully, J., Tamkin, A., Lanham, T., Nguyen, K., Korbak, T., Kaplan, J., Ganguli, D., Bowman, S. R., Perez, E., Grosse, R., and Duvenaud, D · 2024
Closest in time.
Tool use (function calling), 2024
Anthropic · 2024
Closest in time.
Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies, April 2024
Bartoldson, B. R., Diffenderfer, J., Parasyris, K., and Kailkhura, B · 2024
Closest in time.
Defending against unforeseen failure modes with latent adversarial training
Casper, S., Schulze, L., Patel, O., and Hadfield-Menell, D · 2024
Closest in time.
Can LLM-generated misinformation be detected?
Chen, C. and Shu, K · 2024
Closest in time.
Function calling — Google AI for developers, 2024
Google · 2024
Closest in time.
Evaluating language-model agents on realistic autonomous tasks, 2024
Kinniment, M., Sato, L. J. K., Du, H., Goodrich, B., Hasin, M., Chan, L., Miles, L. H., Lin, T. R., Wijk, H., Burget, J., Ho, A., Barnes, E., and Christiano, P · 2024
Closest in time.
Auto-gpt: An autonomous GPT-4 experiment, 2024
Richards, T. B · 2024
Closest in time.
Fast adversarial attacks on language models in one gpu minute, 2024
Sadasivan, V. S., Saha, S., Sriramanan, G., Kattakinda, P., Chegini, A., and Feizi, S · 2024
Closest in time.
A strongreject for empty jailbreaks, 2024
Souly, A., Lu, Q., Bowen, D., Trinh, T., Hsieh, E., Pandey, S., Abbeel, P., Svegliato, J., Emmons, S., Watkins, O., and Toyer, S · 2024
Closest in time.
Efficient adversarial training in llms with continuous attacks
Xhonneux, S., Sordoni, A., Günnemann, S., Gidel, G., and Schwinn, L · 2024
Closest in time.
Improving alignment and robustness with short circuiting
Zou, A., Phan, L., Wang, J., Duenas, D., Lin, M., Andriushchenko, M., Wang, R., Kolter, Z., Fredrikson, M., and Hendrycks, D · 2024
Closest in time.
Qwen2.5 technical report, 2025
Qwen, :, Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Lin, J., Dang, K., Lu, K., Bao, K., Yang, K., Yu, L., Li, M., Xue, M., Zhang, P., Zhu, Q., Men, R., Lin, R., Li, T., Tang, T., Xia, T., Ren, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Wan, Y., Liu, Y., Cui, Z., Zhang, Z., and Qiu, Z · 2025
Closest in time.
Trading inference-time compute for adversarial robustness
Zaremba, W., Nitishinskaya, E., Barak, B., Lin, S., Toyer, S., Yu, Y., Dias, R., Wallace, E., Xiao, K., and Glaese, J. H. A · 2025
Closest in time.