Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have gained significant attention but also raised concerns due to the risk of misuse.
1956
Earlier work this paper cites.
E. Hoque and G. Carenini, “Convis: A visual text analytic system for exploring blog conversations,” in Comput. Graphics Forum , vol. 33, no. 3. Wiley Online Library, 2014, pp. 221–230
2014
Earlier work this paper cites.
E. Hoque and G. Carenini, “Convisit: Interactive topic modeling for exploring asynchronous online conversations,” in Proc. IUI . New York: ACM, 2015, pp. 169–180
2015
Earlier work this paper cites.
M. El-Assady, V. Gold, C. Acevedo, C. Collins, and D. Keim, “Contovi: Multi-party conversation exploration using topic-space views,” in Comput. Graphics Forum , vol. 35, no. 3. Wiley Online Library, 2016, pp. 431–440
2016
Earlier work this paper cites.
S. Fu, J. Zhao, W. Cui, and H. Qu, “Visual analysis of mooc forums with iforum,” IEEE Trans. Visual. Comput. Graphics , vol. 23, no. 1, pp. 201–210, 2016
2016
Earlier work this paper cites.
E. Hoque and G. Carenini, “Multiconvis: A visual text analytics system for exploring a collection of online conversations,” in Proc. IUI . New York: ACM, 2016, pp. 96–107
2016
Earlier work this paper cites.
Y. Ming, S. Cao, R. Zhang, Z. Li, Y. Chen, Y. Song, and H. Qu, “Understanding hidden memories of recurrent neural networks,” in VAST . IEEE, 2017, pp. 13–24
2017
Earlier work this paper cites.
H. Strobelt, S. Gehrmann, H. Pfister, and A. M. Rush, “Lstmvis: A tool for visual analysis of hidden state dynamics in recurrent neural networks,” IEEE Trans. Visual. Comput. Graphics , vol. 24, no. 1, pp. 667–676, 2018
2018
Earlier work this paper cites.
S. Fu, J. Zhao, H. F. Cheng, H. Zhu, and J. Marlow, “T-cal: Understanding team conversational data with calendar-based visualization,” in Proc. CHI . New York: ACM, 2018, pp. 1–13
2018
Earlier work this paper cites.
S. Fu, Y. Wang, Y. Yang, Q. Bi, F. Guo, and H. Qu, “Visforum: A visual analysis system for exploring user groups in online forums,” ACM Transactions on Interactive Intelligent Systems , vol. 8, no. 1, pp. 3:1–3:21, 2018
2018
Earlier work this paper cites.
M. El-Assady, R. Sevastjanova, D. Keim, and C. Collins, “Threadreconstructor: Modeling reply-chains to untangle conversational text through visual analytics,” in Comput. Graphics Forum , vol. 37, no. 3. Wiley Online Library, 2018, pp. 351–365
2018
Earlier work this paper cites.
J.-S. Wong et al. , “Messagelens: A visual analytics system to support multifaceted exploration of mooc forum discussions,” Vis. Informatics , vol. 2, no. 1, pp. 37–49, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
H. Strobelt, S. Gehrmann, M. Behrisch, A. Perer, H. Pfister, and A. M. Rush, “Seq2seq-vis: A visual debugging tool for sequence-to-sequence models,” IEEE Trans. Visual. Comput. Graphics , vol. 25, no. 1, pp. 353–363, 2019
2019
Earlier work this paper cites.
J. Vig, “A multiscale visualization of attention in the transformer model,” in Proc. ACL . Florence, Italy: ACL, 2019, pp. 37–42
2019
Earlier work this paper cites.
C. Park, I. Na, Y. Jo, S. Shin, J. Yoo, B. C. Kwon, J. Zhao, H. Noh, Y. Lee, and J. Choo, “Sanvis: Visual analytics for understanding self-attention networks,” in VIS . IEEE, 2019, pp. 146–150
2019
Earlier work this paper cites.
N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in Proc. EMNLP . Hong Kong, China: ACL, 2019, pp. 3982–3992
2019
Earlier work this paper cites.
B. Hoover, H. Strobelt, and S. Gehrmann, “exBERT: A Visual Analysis Tool to Explore Learned Representations in Transformer Models,” in Proc. ACL . Online: ACL, 2020, pp. 187–196
2020
Earlier work this paper cites.
T. Wu, K. Wongsuphasawat, D. Ren, K. Patel, and C. DuBois, “Tempura: Query analysis with structural templates,” in Proc. CHI . New York: ACM, 2020, pp. 1–12
2020
Earlier work this paper cites.
P. Röttger, B. Vidgen, D. Nguyen, Z. Waseem, H. Margetts, and J. B. Pierrehumbert, “Hatecheck: Functional tests for hate speech detection models,” in Proc. ACL . ACL, 2021, pp. 41–58
2021
Earlier work this paper cites.
J. F. DeRose, J. Wang, and M. Berger, “Attention flows: Analyzing and comparing attention mechanisms in language models,” IEEE Trans. Visual. Comput. Graphics , vol. 27, no. 2, pp. 1160–1170, 2021
2021
Earlier work this paper cites.
R. Li, E. Hoque, G. Carenini, R. Lester, and R. Chau, “Conviscope: Visual analytics for exploring patient conversations,” in VIS . IEEE, 2021, pp. 151–155
2021
Cited alongside, same era.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” in NeurIPS , vol. 35, 2022, pp. 27 730–27 744
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Later among the works it cites.
A. Wei, N. Haghtalab, and J. Steinhardt, “Jailbroken: How does llm safety training fail?” in NeurIPS , vol. 36, 2023, pp. 80 079–80 110
2023
Later among the works it cites.
D. Kang, X. Li, I. Stoica, C. Guestrin, M. Zaharia, and T. Hashimoto, “Exploiting programmatic behavior of LLMs: Dual-use through standard security attacks,” in Proc. ICML , 2023
2023
Later among the works it cites.
B. Deng, W. Wang, F. Feng, Y. Deng, Q. Wang, and X. He, “Attack prompt generation for red teaming and defending large language models,” in Findings of ACL: EMNLP . Singapore: ACL, 2023, pp. 2176–2189
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Hartvigsen, S. Gabriel, H. Palangi, M. Sap, D. Ray, and E. Kamar, “ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection,” in Proc. ACL . Dublin, Ireland: ACL, 2022, pp. 3309–3326
2022
Cited alongside, same era.
P. Röttger, H. Seelawi, D. Nozza, Z. Talat, and B. Vidgen, “Multilingual HateCheck: Functional tests for multilingual hate speech detection models,” in WOAH . Seattle, Washington (Hybrid): ACL, 2022, pp. 154–169
2022
Cited alongside, same era.
2022
Cited alongside, same era.
F. Perez and I. Ribeiro, “Ignore previous prompt: Attack techniques for language models,” in NeurIPS Workshop MLSW , 2022
2022
Cited alongside, same era.
R. Sevastjanova, A. Kalouli, C. Beck, H. Hauptmann, and M. El-Assady, “Lmfingerprints: Visual explanations of language model embedding spaces through layerwise contextualization scores,” in Comput. Graphics Forum , vol. 41, no. 3. Wiley Online Library, 2022, pp. 295–307
2022
Cited alongside, same era.
Z. Li, X. Wang, W. Yang, J. Wu, Z. Zhang, Z. Liu, M. Sun, H. Zhang, and S. Liu, “A unified understanding of deep nlp models for text classification,” IEEE Trans. Visual. Comput. Graphics , vol. 28, no. 12, pp. 4980–4994, 2022
2022
Cited alongside, same era.
H. Strobelt, A. Webson, V. Sanh, B. Hoover, J. Beyer, H. Pfister, and A. M. Rush, “Interactive and visual prompt engineering for ad-hoc task adaptation with large language models,” IEEE Trans. Visual. Comput. Graphics , vol. 29, no. 1, pp. 1146–1156, 2022
2022
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, and E. Wong, “Jailbreaking black box large language models in twenty queries,” in NeurIPS Workshop R0-FoMo , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
S. Schulhoff, J. Pinto, A. Khan, L.-F. Bouchard, C. Si, S. Anati, V. Tagliabue, A. Kost, C. Carnahan, and J. Boyd-Graber, “Ignore this title and HackAPrompt: Exposing systemic vulnerabilities of LLMs through a global prompt hacking competition,” in Proc. EMNLP . Singapore: ACL, 2023, pp. 4945–4977
2023
Later among the works it cites.
Z. J. Wang, F. Hohman, and D. H. Chau, “WizMap: Scalable interactive visualization for exploring large machine learning embeddings,” in Proc. ACL . Toronto, Canada: ACL, 2023, pp. 516–523
2023
Later among the works it cites.
Z. Jin, X. Wang, F. Cheng, C. Sun, Q. Liu, and H. Qu, “Shortcutlens: A visual analytics approach for exploring shortcuts in natural language understanding dataset,” IEEE Trans. Visual. Comput. Graphics , pp. 1–15, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing et al. , “Judging llm-as-a-judge with mt-bench and chatbot arena,” in NeurIPS , vol. 36, 2023, pp. 46 595–46 623
2023
Later among the works it cites.
W. Zhao, X. Ren, J. Hessel, C. Cardie, Y. Choi, and Y. Deng, “(InThe)WildChat: 570k chatgpt interaction logs in the wild,” in ICLR , 2023
2023
Later among the works it cites.
2024
Closest in time.
OpenAI, “Moderation - openai api,” 2024, https://platform.openai.com/docs/guides/moderation
2024
Closest in time.
Y. Huang, S. Gupta, M. Xia, K. Li, and D. Chen, “Catastrophic jailbreak of open-source LLMs via exploiting generation,” in ICLR , 2024
2024
Closest in time.
Y. Yuan, W. Jiao, W. Wang, J. tse Huang, P. He, S. Shi, and Z. Tu, “GPT-4 is too smart to be safe: Stealthy chat with LLMs via cipher,” in ICLR , 2024
2024
Closest in time.
C. Yeh, Y. Chen, A. Wu, C. Chen, F. Viégas, and M. Wattenberg, “Attentionviz: A global view of transformer attention,” IEEE Trans. Visual. Comput. Graphics , vol. 30, no. 1, pp. 262–272, 2024
2024
Closest in time.