Fetching the paper…
Reading the bibliography…
Generative Artificial Intelligence (GenAI) is becoming ubiquitous in our daily lives.
Fairness and abstraction in sociotechnical systems
Andrew D Selbst, Danah Boyd, Sorelle A Friedler, Suresh Venkatasubramanian, and Janet Vertesi · 2019
Earlier work this paper cites.
Vulnerability disclosure mechanisms: A synthesis and framework for market-based and non-market-based disclosures
Ali Ahmed, Amit Deokar, and Ho Cheung Brian Lee · 2021
Earlier work this paper cites.
LangChain, October 2022
Harrison Chase · 2022
Earlier work this paper cites.
Measuring and narrowing the compositionality gap in language models, 2023
Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A. Smith, and Mike Lewis · 2023
Earlier work this paper cites.
Generative ai applications in healthcare
J. Smith and R. Johnson · 2023
Earlier work this paper cites.
Generative ai as a force multiplier in defense
J. Smith and R. Lee · 2023
Earlier work this paper cites.
Fairlearn: Assessing and improving fairness of ai systems
Hilde Weerts, Miroslav Dudik, Richard Edgar, Adrin Jalali, Roman Lutz, and Michael Madaio · 2023
Earlier work this paper cites.
Autogen: Enabling next-gen llm applications via multi-agent conversation framework
Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, and Chi Wang · 2023
Cited alongside, same era.
Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts
Jiahao Yu, Xingwei Lin, and Xinyu Xing · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al · 2024
Cited alongside, same era.
The wmdp benchmark: Measuring and reducing malicious use with unlearning
Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al · 2024
Closest in time.
Datasets for large language models: A comprehensive survey
Yang Liu, Jiahuan Cao, Chongyu Liu, Kai Ding, and Lianwen Jin · 2024
Closest in time.
Codechameleon: Personalized encryption framework for jailbreaking large language models
Huijie Lv, Xiao Wang, Yuansen Zhang, Caishuang Huang, Shihan Dou, Junjie Ye, Tao Gui, Qi Zhang, and Xuanjing Huang · 2024
Closest in time.
Announcing microsoft’s open automation framework to red team generative ai systems, 2024
Microsoft · 2024
Closest in time.
Semantic Kernel, June 2024
Microsoft · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Many-shot jailbreaking
Cem Anil, Esin Durmus, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Nina Rimsky, Meg Tong, Jesse Mu, Daniel Ford, et al · 2024
Cited alongside, same era.
Generative ai for financial forecasting
A. Brown and C. Lee · 2024
Cited alongside, same era.
garak: A framework for security probing large language models
Leon Derczynski, Erick Galinkin, Jeffrey Martin, Subho Majumdar, and Nanna Inie · 2024
Cited alongside, same era.
Jailbreaking black box large language models in twenty queries
Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong
Cited in the paper.
AI security risk assessment using counterfit
Ram Shankar Siva Kumar
Cited in the paper.
You shall not pass: the spells behind gandalf extbar lakera - protecting AI teams that disrupt the world
Max Mathys
Cited in the paper.
Tree of attacks: Jailbreaking black-box LLMs automatically
Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi
Cited in the paper.
InterpretML: A unified framework for machine learning interpretability
Harsha Nori, Samuel Jenkins, Paul Koch, and Rich Caruana
Cited in the paper.
Closest in time.
Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi · 2024
Closest in time.