Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have been observed to process non-human-readable text sequences, such as jailbreak prompts, often viewed as a bug for aligned LLMs.
Efficient saliency maps for explainable ai, 2020
Mundhenk, T. N., Chen, B. Y., and Friedland, G · 1911
Earlier work this paper cites.
A gold standard dependency corpus for English
Silveira, N., Dozat, T., de Marneffe, M.-C., Bowman, S., Connor, M., Bauer, J., and Manning, C. D · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I. J., and Fergus, R · 2014
Earlier work this paper cites.
Deep neural networks are easily fooled: High confidence predictions for unrecognizable images
Nguyen, A., Yosinski, J., and Clune, J · 2015
Earlier work this paper cites.
Transferability in machine learning: from phenomena to black-box attacks using adversarial samples
Papernot, N., McDaniel, P., and Goodfellow, I · 2016
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J · 2018
Earlier work this paper cites.
A structural probe for finding syntax in word representations
Hewitt, J. and Manning, C. D · 2019
Earlier work this paper cites.
Adversarial examples are not bugs, they are features
Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., and Madry, A · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I · 2019
Earlier work this paper cites.
Universal adversarial triggers for attacking and analyzing nlp
Wallace, E., Feng, S., Kandpal, N., Gardner, M., and Singh, S · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S · 2020
Earlier work this paper cites.
Sinha, K., Parthasarathi, P., Pineau, J., and Williams, A · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Cited alongside, same era.
Adversarial attacks on gpt-4 via simple random search, 2022
Andriushchenko, M · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., et al · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
Meet claude
anthropic · 2024
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Omni-math: A universal olympiad level mathematic benchmark for large language models
Gao, B., Song, F., Yang, Z., Cai, Z., Miao, Y., Dong, Q., Li, L., Ma, C., Chen, L., Xu, R., et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
Automatically auditing large language models via discrete optimization
Jones, E., Dragan, A., Raghunathan, A., and Steinhardt, J · 2023
Cited alongside, same era.
Unnatural language processing: How do language models handle machine-generated prompts?
Kervadec, C., Franzon, F., and Baroni, M · 2023
Cited alongside, same era.
Alpacaeval: An automatic evaluator of instruction-following models
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Autodan: Generating stealthy jailbreak prompts on aligned large language models
Liu, X., Xu, N., Chen, M., and Xiao, C · 2023
Cited alongside, same era.
Later among the works it cites.
Instruction following without instruction tuning, 2024
Hewitt, J., Liu, N. F., Liang, P., and Manning, C. D · 2024
Later among the works it cites.
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al · 2024
Later among the works it cites.
Mission: Impossible language models
Kallini, J., Papadimitriou, I., Futrell, R., Mahowald, K., and Potts, C · 2024
Later among the works it cites.
Mixeval: Deriving wisdom of the crowd from llm benchmark mixtures
Ni, J., Xue, F., Yue, X., Deng, Y., Shah, M., Jain, K., Neubig, G., and You, Y · 2024
Later among the works it cites.
Introducing chatgpt
OpenAI · 2024
Later among the works it cites.
Let’s think dot by dot: Hidden computation in transformer language models
Pfau, J., Merrill, W., and Bowman, S. R · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Team, G., Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., et al · 2024
Later among the works it cites.
Accelerating greedy coordinate gradient via probe sampling
Zhao, Y., Zheng, W., Cai, T., Do, X. L., Kawaguchi, K., Goyal, A., and Shieh, M · 2024
Later among the works it cites.