Fetching the paper…
Reading the bibliography…
Recent advancements in generative AI have enabled ubiquitous access to large language models (LLMs).
Using thematic analysis in psychology
Virginia Braun and Victoria Clarke · 2006
Earlier work this paper cites.
Studies in machiavellianism
Richard Christie and Florence L Geis · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy · 2015
Earlier work this paper cites.
False sense of security: A study on the effectivity of jailbreak detection in banking apps
Ansgar Kellner, Micha Horlboge, Konrad Rieck, and Christian Wressnegger · 2019
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Earlier work this paper cites.
Weight poisoning attacks on pretrained models
Keita Kurita, Paul Michel, and Graham Neubig · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2021
Earlier work this paper cites.
You autocomplete me: Poisoning vulnerabilities in neural code completion
Roei Schuster, Congzheng Song, Eran Tromer, and Vitaly Shmatikov · 2021
Earlier work this paper cites.
Story centaur: Large language model few shot learning as a creative writing tool
Ben Swanson, Kory Mathewson, Ben Pietrzak, Sherol Chen, and Monica Dinalescu · 2021
Earlier work this paper cites.
Ethical and social risks of harm from language models
Laura Weidinger, John Mellor, Maribeth Rauh, Conor Griffin, Jonathan Uesato, Po-Sen Huang, Myra Cheng, Mia Glaese, Borja Balle, Atoosa Kasirzadeh, et al · 2021
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
Membership inference attacks from first principles
Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, and Florian Tramer · 2022
Earlier work this paper cites.
Deduplicating training data mitigates privacy risks in language models
Nikhil Kandpal, Eric Wallace, and Colin Raffel · 2022
Earlier work this paper cites.
Design guidelines for prompt engineering text-to-image generative models
Vivian Liu and Lydia B Chilton · 2022
Earlier work this paper cites.
Chatgpt has a devastating sense of humor
Farhad Manjoo · 2022
Earlier work this paper cites.
Quantifying privacy risks of masked language models using membership inference attacks
Fatemehsadat Mireshghallah, Kartik Goyal, Archit Uniyal, Taylor Berg-Kirkpatrick, and Reza Shokri · 2022
Earlier work this paper cites.
Introducing chatgpt
OpenAI · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Ignore previous prompt: Attack techniques for language models
Fábio Perez and Ian Ribeiro · 2022
Cited alongside, same era.
Truth serum: Poisoning machine learning models to reveal their secrets
Florian Tramèr, Reza Shokri, Ayrton San Joaquin, Hoang Le, Matthew Jagielski, Sanghyun Hong, and Nicholas Carlini · 2022
Cited alongside, same era.
Jailbreak chat
Alex Albert · 2023
Cited alongside, same era.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
How to jailbreak chatgpt: Best prompts & more
Joel Loynds · 2023
Later among the works it cites.
Membership inference attacks against language models via neighbourhood comparison
Justus Mattern, Fatemehsadat Mireshghallah, Zhijing Jin, Bernhard Schölkopf, Mrinmaya Sachan, and Taylor Berg-Kirkpatrick · 2023
Later among the works it cites.
Chatgpt “dan" (and other “jailbreaks")
AJ ONeal · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Openai usage policies
OpenAI · 2023
Later among the works it cites.
Usage policies
OpenAI · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ask me anything: A simple strategy for prompting language models
Simran Arora, Avanika Narayan, Mayee F. Chen, Laurel J. Orr, Neel Guha, Kush Bhatia, Ines Chami, and Christopher Ré · 2023
Cited alongside, same era.
Number of chatgpt users (2023)
Fabio Duarte · 2023
Cited alongside, same era.
Flowgpt: Fast & free ai & gpts bots store
FlowGPT · 2023
Cited alongside, same era.
It’s scary easy to use chatgpt to write phishing emails
Bree Fowler · 2023
Cited alongside, same era.
Chatgpt jailbreak prompts
Jonathan Gan · 2023
Cited alongside, same era.
Introducing palm 2
Zoubin Ghahramani · 2023
Cited alongside, same era.
PRAW · 2023
Later among the works it cites.
Tricking llms into disobedience: Understanding, analyzing, and preventing jailbreaks
Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya, and Monojit Choudhury · 2023
Later among the works it cites.
Exploring chatgpt vs open-source models on slightly harder tasks
Marco Tulio Ribeiro · 2023
Later among the works it cites.
Selenium automates browsers
Selenium · 2023
Later among the works it cites.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2023
Later among the works it cites.
The artificially intelligent entrepreneur: Chatgpt, prompt engineering, and entrepreneurial rhetoric creation
Cole E Short and Jeremy C Short · 2023
Later among the works it cites.
Hackers with ai are harder to stop, microsoft says
Catherine Stupp · 2023
Later among the works it cites.
Poisoning language models during instruction tuning
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein · 2023
Later among the works it cites.
Fundamental limitations of alignment in large language models
Yotam Wolf, Noam Wies, Yoav Levine, and Amnon Shashua · 2023
Later among the works it cites.
Tidybot: Personalized robot assistance with large language models
Jimmy Wu, Rika Antonova, Adam Kan, Marion Lepert, Andy Zeng, Shuran Song, Jeannette Bohg, Szymon Rusinkiewicz, and Thomas Funkhouser · 2023
Later among the works it cites.
Codeipprompt: Intellectual property infringement assessment of code language models
Zhiyuan Yu, Yuhao Wu, Ning Zhang, Chenguang Wang, Yevgeniy Vorobeychik, and Chaowei Xiao · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2024
Closest in time.