Fetching the paper…
Reading the bibliography…
The proliferation of large language models (LLMs) has underscored concerns regarding their security vulnerabilities, notably against jailbreak attacks, where adversaries design jailbreak prompts to circumvent safety mechanisms for potential misuse.
K. Sparck Jones, “A statistical interpretation of term specificity and its application in retrieval,” Journal of documentation , vol. 60, no. 5, pp. 493–502, 2004
2004
Earlier work this paper cites.
M. M. Malik, C. Heinzl, and M. E. Groeller, “Comparative visualization for parameter studies of dataset series,” IEEE TVCG , vol. 16, no. 5, pp. 829–840, 2010
2010
Earlier work this paper cites.
A. Karpathy, J. Johnson, and L. Fei-Fei, “Visualizing and understanding recurrent networks,” arXiv , 2015
2015
Earlier work this paper cites.
J. Li, X. Chen, E. Hovy, and D. Jurafsky, “Visualizing and understanding neural models in nlp,” arXiv , 2016
2016
Earlier work this paper cites.
Y. Ming, S. Cao, R. Zhang, Z. Li, Y. Chen, Y. Song, and H. Qu, “Understanding hidden memories of recurrent neural networks,” in IEEE VAST . IEEE, 2017, pp. 13–24
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” NeurIPS , vol. 30, 2017
2017
Earlier work this paper cites.
R. Speer, J. Chin, and C. Havasi, “Conceptnet 5.5: An open multilingual graph of general knowledge,” in Proceedings of AAAI , vol. 31, no. 1, 2017
2017
Earlier work this paper cites.
A. Shafahi, W. R. Huang, M. Najibi, O. Suciu, C. Studer, T. Dumitras, and T. Goldstein, “Poison frogs! targeted clean-label poisoning attacks on neural networks,” NeurIPS , vol. 31, 2018
2018
Earlier work this paper cites.
M. Liu, S. Liu, H. Su, K. Cao, and J. Zhu, “Analyzing the noise robustness of deep neural networks,” in 2018 IEEE Conference on Visual Analytics Science and Technology (VAST) . IEEE, 2018, pp. 60–71
2018
Earlier work this paper cites.
H. Strobelt, S. Gehrmann, H. Pfister, and A. M. Rush, “Lstmvis: A tool for visual analysis of hidden state dynamics in recurrent neural networks,” IEEE TVCG , vol. 24, no. 1, pp. 667–676, 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” arXiv , 2018
2018
Earlier work this paper cites.
S. Liu, Z. Li, T. Li, V. Srikumar, V. Pascucci, and P.-T. Bremer, “Nlize: A perturbation-driven visual interrogation tool for analyzing and interpreting natural language inference models,” IEEE TVCG , vol. 25, no. 1, pp. 651–660, 2019
2019
Earlier work this paper cites.
N. Das, H. Park, Z. J. Wang, F. Hohman, R. Firstman, E. Rogers, and D. H. P. Chau, “Bluff: Interactively deciphering adversarial attacks on deep neural networks,” in IEEE VIS . IEEE, 2020, pp. 271–275
2020
Earlier work this paper cites.
Y. Ma, T. Xie, J. Li, and R. Maciejewski, “Explaining vulnerabilities to adversarial machine learning through visual analytics,” IEEE TVCG , vol. 26, no. 1, pp. 1075–1085, 2020
2020
Earlier work this paper cites.
J. Wexler, M. Pushkarna, T. Bolukbasi, M. Wattenberg, F. Viégas, and J. Wilson, “The what-if tool: Interactive probing of machine learning models,” IEEE TVCG , vol. 26, no. 1, pp. 56–65, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
W. He, J. Wang, H. Guo, H.-W. Shen, and T. Peterka, “Cecav-dnn: Collective ensemble comparison and visualization using deep neural networks,” Visual Informatics , vol. 4, no. 2, pp. 109–121, 2020
2020
Earlier work this paper cites.
F. Cheng, Y. Ming, and H. Qu, “Dece: Decision explorer with counterfactual explanations for machine learning models,” IEEE TVCG , vol. 27, no. 2, pp. 1438–1447, 2021
2021
Earlier work this paper cites.
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang, “Glm: General language model pretraining with autoregressive blank infilling,” arXiv , 2021
2021
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” NeurIPS , vol. 35, pp. 27 730–27 744, 2022
2022
Earlier work this paper cites.
F. Perez and I. Ribeiro, “Ignore previous prompt: Attack techniques for language models,” arXiv , 2022
2022
Earlier work this paper cites.
D. Ziegler, S. Nix, L. Chan, T. Bauman, P. Schmidt-Nielsen, T. Lin, A. Scherlis, N. Nabeshima, B. Weinstein-Raun, D. de Haas et al. , “Adversarial training for high-stakes reliability,” NeurIPS , vol. 35, pp. 9274–9286, 2022
2022
Earlier work this paper cites.
Y. Feng, J. Chen, K. Huang, J. K. Wong, H. Ye, W. Zhang, R. Zhu, X. Luo, and W. Chen, “ipoet: interactive painting poetry creation with visual multimodal analysis,” Journal of Visualization , pp. 1–15, 2022
2022
Earlier work this paper cites.
A. Boggust, B. Hoover, A. Satyanarayan, and H. Strobelt, “Shared interest: Measuring human-ai alignment to identify recurring patterns in model behavior,” in Proceedings of CHI , 2022, pp. 1–17
2022
Earlier work this paper cites.
N. Feldhus, A. M. Ravichandran, and S. Möller, “Mediators: Conversational agents explaining nlp model behavior,” arXiv , 2022
2022
Earlier work this paper cites.
P. P. Liang, Y. Lyu, G. Chhablani, N. Jain, Z. Deng, X. Wang, L.-P. Morency, and R. Salakhutdinov, “Multiviz: Towards visualizing and understanding multimodal models,” arXiv , 2022
2022
Earlier work this paper cites.
Z. Li, X. Wang, W. Yang, J. Wu, Z. Zhang, Z. Liu, M. Sun, H. Zhang, and S. Liu, “A unified understanding of deep nlp models for text classification,” IEEE TVCG , vol. 28, no. 12, pp. 4980–4994, 2022
2022
Earlier work this paper cites.
2022
Cited alongside, same era.
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar et al. , “Llama: Open and efficient foundation language models,” arXiv , 2023
2023
Cited alongside, same era.
S. Brade, B. Wang, M. Sousa, S. Oore, and T. Grossman, “Promptify: Text-to-image generation through interactive prompt exploration with large language models,” in Proceedings of ACM UIST , 2023, pp. 1–14
2023
Cited alongside, same era.
T. Angert, M. Suzara, J. Han, C. Pondoc, and H. Subramonyam, “Spellburst: A node-based interface for exploratory creative coding with natural language prompts,” in Proceedings of ACM UIST , 2023, pp. 1–22
2023
Cited alongside, same era.
“OpenAI’s Usage Policies,” https://openai.com/policies/usage-policies , accessed on October 1st, 2023
2023
Later among the works it cites.
H. Huang, Z. Zhao, M. Backes, Y. Shen, and Y. Zhang, “Composite backdoor attacks against large language models,” arXiv , 2023
2023
Later among the works it cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in NeurIPS , vol. 36. Curran Associates, Inc., 2023, pp. 34 892–34 916
2023
Later among the works it cites.
N. Carlini, M. Nasr, C. A. Choquette-Choo, M. Jagielski, I. Gao, P. W. W. Koh, D. Ippolito, F. Tramer, and L. Schmidt, “Are aligned neural networks adversarially aligned?” in NeurIPS , vol. 36. Curran Associates, Inc., 2023, pp. 61 478–61 500
2023
Later among the works it cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “Gpt-4 technical report,” arXiv , 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Z. Peng, X. Wang, Q. Han, J. Zhu, X. Ma, and H. Qu, “Storyfier: Exploring vocabulary learning support with text generation models,” in Proceedings of ACM UIST , 2023, pp. 1–16
2023
Cited alongside, same era.
Y. Liu, Z. Wen, L. Weng, O. Woodman, Y. Yang, and W. Chen, “Sprout: Authoring programming tutorials with interactive visualization of large language model generation process,” arXiv , 2023
2023
Cited alongside, same era.
M. X. Liu, A. Sarkar, C. Negreanu, B. Zorn, J. Williams, N. Toronto, and A. D. Gordon, ““what it wants me to say”: Bridging the abstraction gap between end-user programmers and code-generating large language models,” in Proceedings of CHI , 2023, pp. 1–31
2023
Cited alongside, same era.
T. Xie, F. Zhou, Z. Cheng, P. Shi, L. Weng, Y. Liu, T. J. Hua, J. Zhao, Q. Liu, C. Liu et al. , “Openagents: An open platform for language agents in the wild,” arXiv , 2023
2023
Cited alongside, same era.
E. Shayegani, M. A. A. Mamun, Y. Fu, P. Zaree, Y. Dong, and N. Abu-Ghazaleh, “Survey of vulnerabilities in large language models revealed by adversarial attacks,” arXiv , 2023
2023
Cited alongside, same era.
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong, “Jailbreaking black box large language models in twenty queries,” arXiv , 2023
2023
Cited alongside, same era.
X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang, “”do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models,” arXiv , 2023
2023
Cited alongside, same era.
A. Zou, Z. Wang, J. Z. Kolter, and M. Fredrikson, “Universal and transferable adversarial attacks on aligned language models,” arXiv , 2023
2023
Cited alongside, same era.
2024
Closest in time.
L. Weng, X. Wang, J. Lu, Y. Feng, Y. Liu, and W. Chen, “Insightlens: Discovering and exploring insights from conversational contexts in large-language-model-powered data analysis,” arXiv , 2024
2024
Closest in time.
Y. Liu, G. Deng, Z. Xu, Y. Li, Y. Zheng, Y. Zhang, L. Zhao, T. Zhang, and Y. Liu, “Jailbreaking chatgpt via prompt engineering: An empirical study,” arXiv , 2024
2024
Closest in time.
A. Wei, N. Haghtalab, and J. Steinhardt, “Jailbroken: How does llm safety training fail?” NeurIPS , vol. 36, 2024
2024
Closest in time.
G. Deng, Y. Liu, Y. Li, K. Wang, Y. Zhang, Z. Li, H. Wang, T. Zhang, and Y. Liu, “Masterkey: Automated jailbreaking of large language model chatbots,” arXiv , 2024
2024
Closest in time.
J. Chu, Y. Liu, Z. Yang, X. Shen, M. Backes, and Y. Zhang, “Comprehensive assessment of jailbreak attacks against llms,” arXiv , 2024
2024
Closest in time.
P. Ding, J. Kuang, D. Ma, X. Cao, Y. Xian, J. Chen, and S. Huang, “A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily,” arXiv , 2024
2024
Closest in time.
X. Liu, N. Xu, M. Chen, and C. Xiao, “Autodan: Generating stealthy jailbreak prompts on aligned large language models,” arXiv , 2024
2024
Closest in time.
M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li et al. , “Harmbench: A standardized evaluation framework for automated red teaming and robust refusal,” arXiv , 2024
2024
Closest in time.
Y. Yuan, W. Jiao, W. Wang, J.-t. Huang, P. He, S. Shi, and Z. Tu, “Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher,” arXiv , 2024
2024
Closest in time.
X. Chen, X. Zhang, Z. Wang, K. Yu, W. Kam-Kwai, H. Guo, and S. Chen, “Visual analytics for security threats detection in ethereum consensus layer,” Journal of Visualization , vol. 27, no. 3, pp. 469–483, 2024
2024
Closest in time.
Z. Jin, S. Liu, H. Li, X. Zhao, and H. Qu, “Jailbreakhunter: A visual analytics approach for jailbreak prompts discovery from large-scale human-llm conversational datasets,” arXiv , 2024
2024
Closest in time.
D. Deng, C. Zhang, H. Zheng, Y. Pu, S. Ji, and Y. Wu, “Adversaflow: Visual red teaming for large language models with multi-level adversarial flow,” IEEE TVCG , 2024
2024
Closest in time.
J. Lu, B. Pan, J. Chen, Y. Feng, J. Hu, Y. Peng, and W. Chen, “Agentlens: Visual analysis for agent behaviors in llm-based autonomous systems,” arXiv , 2024
2024
Closest in time.
S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, and K. Narasimhan, “Tree of thoughts: Deliberate problem solving with large language models,” NeurIPS , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
“OpenAI’s Embedding Models,” https://platform.openai.com/docs/guides/embeddings , accessed on March 1st, 2024
2024
Closest in time.
C. Shi, X. Wang, Q. Ge, S. Gao, X. Yang, T. Gui, Q. Zhang, X. Huang, X. Zhao, and D. Lin, “Navigating the overkill in large language models,” arXiv , 2024
2024
Closest in time.
W. Yang, X. Bi, Y. Lin, S. Chen, J. Zhou, and X. Sun, “Watch out for your agents! investigating backdoor threats to llm-based agents,” arXiv , 2024
2024
Closest in time.
Z. Zhou, H. Yu, X. Zhang, R. Xu, F. Huang, and Y. Li, “How alignment and jailbreak work: Explain llm safety through intermediate hidden states,” arXiv , 2024
2024
Closest in time.
“GPT-4V(ision) System Card,” https://openai.com/research/gpt-4v-system-card , accessed on March 1st, 2024
2024
Closest in time.
E. Shayegani, Y. Dong, and N. Abu-Ghazaleh, “Jailbreak in pieces: Compositional adversarial attacks on multi-modal language models,” in ICLR , 2024
2024
Closest in time.
X. Qi, K. Huang, A. Panda, P. Henderson, M. Wang, and P. Mittal, “Visual adversarial examples jailbreak aligned large language models,” Proceedings of AAAI , vol. 38, no. 19, pp. 21 527–21 536, Mar. 2024
2024
Closest in time.