Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have increasingly become pivotal in content generation with notable societal impact.
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: Bert: Pre-training of deep bidirectional transformers for language understanding (2019)
2019
Earlier work this paper cites.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I.: Language models are unsupervised multitask learners (2019), https://api.semanticscholar.org/CorpusID:160025533
2019
Earlier work this paper cites.
Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A.: Language models are few-shot learners (2020)
2020
Earlier work this paper cites.
Paluri, S.: Human computer interaction (12 2020)
2020
Earlier work this paper cites.
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Kernion, J., Ndousse, K., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Kaplan, J.: A general language assistant as a laboratory for alignment (2021)
2021
Earlier work this paper cites.
Dasigi, P., Lo, K., Beltagy, I., Cohan, A., Smith, N.A., Gardner, M.: A dataset of information-seeking questions and answers anchored in research papers (2021)
2021
Earlier work this paper cites.
Apruzzese, G., Anderson, H.S., Dambra, S., Freeman, D., Pierazzi, F., Roundy, K.A.: "real attackers don’t compute gradients": Bridging the gap between adversarial ml research and practice (2022)
2022
Earlier work this paper cites.
Lees, A., Tran, V.Q., Tay, Y., Sorensen, J., Gupta, J., Metzler, D., Vasserman, L.: A new generation of perspective api: Efficient multilingual character-level transformers (2022)
2022
Earlier work this paper cites.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al.: Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems 35
2022
Earlier work this paper cites.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q.V., Zhou, D., et al.: Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35
2022
Earlier work this paper cites.
Openai says a bug leaked sensitive chatgpt user data. https://www.engadget.com/chatgpt-briefly-went-offline-after-a-bug-revealed-user-chat-histories-115632504.html , engadget. Accessed:2023-8-20
2023
Earlier work this paper cites.
Chatgpt: Us lawyer admits using ai for case research (2023), https://www.bbc.co.uk/news/world-us-canada-65735769 , accessed: 2023-08-23
2023
Earlier work this paper cites.
’he would still be here’: Man dies by suicide after talking with ai chatbot, widow says (2023), https://www.vice.com/en/article/pkadgm/man-dies-by-suicide-after-talking-with-ai-chatbot-widow-says , accessed: 2023-08-23
2023
Earlier work this paper cites.
Bano, M., Zowghi, D., Shea, P., Ibarra, G.: Investigating responsible ai for scientific research: An empirical study (2023)
2023
Earlier work this paper cites.
Bhardwaj, R., Poria, S.: Red-teaming large language models using chain of utterances for safety-alignment (2023)
2023
Earlier work this paper cites.
Bhatt, M., Chennabasappa, S., Nikolaidis, C., Wan, S., Evtimov, I., Gabi, D., Song, D., Ahmad, F., Aschermann, C., Fontana, L., Frolov, S., Giri, R.P., Kapil, D., Kozyrakis, Y., LeBlanc, D., Milazzo, J., Straumann, A., Synnaeve, G., Vontimitta, V., Whitman, S., Saxe, J.: Purple llama cyberseceval: A secure coding benchmark for language models (2023)
2023
Earlier work this paper cites.
Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G., Wong, E.: Jailbreaking black box large language models in twenty queries (Oct 2023)
2023
Earlier work this paper cites.
Cinà, A.E., Grosse, K., Demontis, A., Vascon, S., Zellinger, W., Moser, B.A., Oprea, A., Biggio, B., Pelillo, M., Roli, F.: Wild patterns reloaded: A survey of machine learning security against training data poisoning. ACM Computing Surveys 55
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Ding, P., Kuang, J., Ma, D., Cao, X., Xian, Y., Chen, J., Huang, S.: A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily (Nov 2023)
2023
Earlier work this paper cites.
Dong, Q., Li, L., Dai, D., Zheng, C., Wu, Z., Chang, B., Sun, X., Xu, J., Li, L., Sui, Z.: A survey on in-context learning (2023)
2023
Earlier work this paper cites.
Frieder, S., Pinchetti, L., Chevalier, A., Griffiths, R.R., Salvatori, T., Lukasiewicz, T., Petersen, P.C., Berner, J.: Mathematical capabilities of chatgpt (2023)
2023
Earlier work this paper cites.
Gopal, A., Helm-Burger, N., Justen, L., Soice, E.H., Tzeng, T., Jeyapragasan, G., Grimm, S., Mueller, B., Esvelt, K.M.: Will releasing the weights of future large language models grant widespread access to pandemic agents? (2023)
2023
Earlier work this paper cites.
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., Fritz, M.: Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection (2023)
2023
Earlier work this paper cites.
Huang, X., Ruan, W., Huang, W., Jin, G., Dong, Y., Wu, C., Bensalem, S., Mu, R., Qi, Y., Zhao, X., Cai, K., Zhang, Y., Wu, S., Xu, P., Wu, D., Freitas, A., Mustafa, M.A.: A survey of safety and trustworthiness of large language models through the lens of verification and validation (2023)
2023
Earlier work this paper cites.
Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., Khabsa, M.: Llama guard: Llm-based input-output safeguard for human-ai conversations (2023)
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
Lee, P.: Learning from tay’s introduction. https://blogs.microsoft.com/blog/2016/03/25/learning-tays-introduction/ , accessed: 2023-08-20
2023
Cited alongside, same era.
Li, H., Guo, D., Fan, W., Xu, M., Huang, J., Meng, F., Song, Y.: Multi-step jailbreaking privacy attacks on chatgpt (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al.: Gpt-4 technical report (2024)
2024
Closest in time.
Anthropic: Introducing the next generation of claude. https://www.anthropic.com/news/claude-3-family (2023), accessed: 2024-06-05
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
Liu, H., Ning, R., Teng, Z., Liu, J., Zhou, Q., Zhang, Y.: Evaluating the logical reasoning ability of chatgpt and gpt-4 (2023)
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
OWASP: OWASP Top 10 for LLM Applications (2023), https://owasp.org/www-project-top-10-for-large-language-model-applications/
2023
Cited alongside, same era.
Röttger, P., Kirk, H.R., Vidgen, B., Attanasio, G., Bianchi, F., Hovy, D.: Xstest: A test suite for identifying exaggerated safety behaviours in large language models (2023)
2023
Cited alongside, same era.
2024
Closest in time.
Chao, P., Debenedetti, E., Robey, A., Andriushchenko, M., Croce, F., Sehwag, V., Dobriban, E., Flammarion, N., Pappas, G.J., Tramer, F., Hassani, H., Wong, E.: Jailbreakbench: An open robustness benchmark for jailbreaking large language models (2024)
2024
Closest in time.
Chiang, W.L., Zheng, L., Sheng, Y., Angelopoulos, A.N., Li, T., Li, D., Zhang, H., Zhu, B., Jordan, M., Gonzalez, J.E., Stoica, I.: Chatbot arena: An open platform for evaluating llms by human preference (2024)
2024
Closest in time.
Face, H.: Llm safety leaderboard. https://huggingface.co/spaces/AI-Secure/llm-trustworthy-leaderboard , accessed: 2024-06-05
2024
Closest in time.
Face, H.: Open llm leaderboard. https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard , accessed: 2024-06-05
2024
Closest in time.
Face, H.: Perplexity of fixed-length models. https://huggingface.co/docs/transformers/perplexity , accessed: 2024-06-05
2024
Closest in time.
Face, H.: Tokenizers. https://huggingface.co/docs/tokenizers/index , accessed: 2024-06-05
2024
Closest in time.
GitHub: Github copilot: Your ai pair programmer. https://github.com/features/copilot , accessed: 2024-06-05
2024
Closest in time.
Google: Google gemini advanced. https://gemini.google.com/advanced , accessed: 2024-06-05
2024
Closest in time.
Guo, T., Chen, X., Wang, Y., Chang, R., Pei, S., Chawla, N.V., Wiest, O., Zhang, X.: Large language model based multi-agents: A survey of progress and challenges (2024)
2024
Closest in time.
Hamiel, N.: Reducing the impact of prompt injection attacks through design. https://research.kudelskisecurity.com/2023/05/25/reducing-the-impact-of-prompt-injection-attacks-through-design/ , accessed: 2024-06-05
2024
Closest in time.
Huang, Y., Gupta, S., Xia, M., Li, K., Chen, D.: Catastrophic jailbreak of open-source LLMs via exploiting generation. In: The Twelfth International Conference on Learning Representations (2024), https://openreview.net/forum?id=r42tSSCHPh
2024
Closest in time.
Hugging Face: Meta llama. https://huggingface.co/meta-llama (2023), accessed: 2024-02-14
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
Meta: Introducing meta llama 3: The most capable openly available llm to date. https://ai.meta.com/blog/meta-llama-3/ , accessed: 2024-06-05
2024
Closest in time.
OpenAI: Research overview. https://openai.com/research/overview (2023), accessed: 2024-02-14
2024
Closest in time.
Rachel Draelos, MD, P.: Bias, toxicity, and jailbreaking large language models (llms). https://glassboxmedicine.com/2023/11/28/bias-toxicity-and-jailbreaking-large-language-models-llms/ , accessed: 2024-06-05
2024
Closest in time.
Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M.S., Love, J., Tafti, P.: Gemma: Open models based on gemini research and technology (2024)
2024
Closest in time.
University, S.: Cs324: Data. https://stanford-cs324.github.io/winter2022/lectures/data/ , accessed: 2024-06-05
2024
Closest in time.
Xu, H., Wang, S., Li, N., Wang, K., Zhao, Y., Chen, K., Yu, T., Liu, Y., Wang, H.: Large language models for cyber security: A systematic literature review (2024)
2024
Closest in time.
Xu, Z., Liu, Y., Deng, G., Li, Y., Picek, S.: A comprehensive study of jailbreak attack versus defense for large language models (2024)
2024
Closest in time.
Zeng, Y., Lin, H., Zhang, J., Yang, D., Jia, R., Shi, W.: How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms (2024)
2024
Closest in time.