Fetching the paper…
Reading the bibliography…
The growing use of large language model (LLM)-based conversational agents to manage sensitive user data raises significant privacy concerns.
Randomized response: A survey technique for eliminating evasive answer bias
Stanley L Warner · 1965
Earlier work this paper cites.
Computer security technology planning study
James P Anderson et al · 1972
Earlier work this paper cites.
Role-based access control
Ravi S Sandhu · 1998
Earlier work this paper cites.
A language-based approach to security
Fred B Schneider, Greg Morrisett, and Robert Harper · 2001
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Limiting privacy breaches in privacy preserving data mining
Alexandre Evfimievski, Johannes Gehrke, and Ramakrishnan Srikant · 2003
Earlier work this paper cites.
Privacy as contextual integrity
Helen Nissenbaum · 2004
Earlier work this paper cites.
A reference monitor for workflow systems with constrained task execution
Jason Crampton · 2005
Earlier work this paper cites.
Privacy and contextual integrity: framework and applications
A Barth, A Datta, J C Mitchell, and H Nissenbaum · 2006
Earlier work this paper cites.
Rule-based access control for social networks
Barbara Carminati, Elena Ferrari, and Andrea Perego · 2006
Earlier work this paper cites.
Differential privacy
Cynthia Dwork · 2006
Earlier work this paper cites.
Lessons learned from the deployment of a smartphone-based access-control system
Lujo Bauer, Lorrie Faith Cranor, Michael K Reiter, and Kami Vaniea · 2007
Earlier work this paper cites.
Privacy in context: Technology, policy, and the integrity of social life
Helen Nissenbaum · 2009
Earlier work this paper cites.
Privacy in "the cloud" applying Nissenbaum’s theory of contextual integrity
Frances S Grodzinsky and Herman T Tavani · 2011
Earlier work this paper cites.
What can we learn privately?
Shiva Prasad Kasiviswanathan, Homin K Lee, Kobbi Nissim, Sofya Raskhodnikova, and Adam Smith · 2011
Earlier work this paper cites.
The air gap: SCADA’s enduring security myth
Eric Byres · 2013
Earlier work this paper cites.
Local privacy and statistical minimax rates
John C Duchi, Michael I Jordan, and Martin J Wainwright · 2013
Earlier work this paper cites.
Extremal mechanisms for local differential privacy
Peter Kairouz, Sewoong Oh, and Pramod Viswanath · 2014
Earlier work this paper cites.
Implicit contextual integrity in online social networks
Natalia Criado and Jose M Such · 2015
Earlier work this paper cites.
Android permissions remystified: A field study on contextual integrity
Primal Wijesekera, Arjun Baokar, Ashkan Hosseini, Serge Egelman, David Wagner, and Konstantin Beznosov · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Contextual integrity through the lens of computer science
Sebastian Benthall, Seda Gürses, and Helen Nissenbaum · 2017
Earlier work this paper cites.
Privacy at scale: Local differential privacy in practice
Graham Cormode, Somesh Jha, Tejas Kulkarni, Ninghui Li, Divesh Srivastava, and Tianhao Wang · 2018
Earlier work this paper cites.
VACCINE: Using contextual integrity for data leakage detection
Yan Shvartzshnaider, Zvonimir Pavlinovic, Ananth Balashankar, Thomas Wies, Lakshminarayanan Subramanian, Helen Nissenbaum, and Prateek Mittal · 2019
Cited alongside, same era.
Universal adversarial triggers for attacking and analyzing nlp
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh · 2019
Cited alongside, same era.
Context aware local differential privacy
Jayadev Acharya, Kallista Bonawitz, Peter Kairouz, Daniel Ramage, and Ziteng Sun · 2020
Cited alongside, same era.
Aquilis: Using contextual integrity for privacy protection on mobile devices
Abhishek Kumar, Tristan Braud, Young D Kwon, and Pan Hui · 2020
Cited alongside, same era.
A smart-contract-based access control framework for cloud smart healthcare system
Akanksha Saini, Qingyi Zhu, Navneet Singh, Yong Xiang, Longxiang Gao, and Yushu Zhang · 2020
Cited alongside, same era.
A survey on air-gap attacks: Fundamentals, transport means, attack scenarios and challenges
Jangyong Park, Jaehoon Yoo, Jaehyun Yu, Jiho Lee, and JaeSeung Song · 2023
Later among the works it cites.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2023
Later among the works it cites.
“What can ChatGPT do?” analyzing early reactions to the innovative AI chatbot on Twitter
Viriya Taecharungroj · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models, 2023
Gemini Team · 2023
Later among the works it cites.
DecodingTrust: A comprehensive assessment of trustworthiness in GPT models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P. Parikh · 2020
Cited alongside, same era.
Phishing attacks: A recent comprehensive study and a new anatomy
Zainab Alkhalil, Chaminda Hewage, Liqaa Nawaf, and Imtiaz Khan · 2021
Cited alongside, same era.
Tag: Gradient attack on transformer-based language models
Jieren Deng, Yijue Wang, Ji Li, Chao Shang, Hang Liu, Sanguthevar Rajasekaran, and Caiwen Ding · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Cited alongside, same era.
Towards user profile meta-ontology
Ankica Barisic and Marco Winckler · 2022
Cited alongside, same era.
What does it mean for a language model to preserve privacy?
Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian Tramèr · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, et al · 2022
Cited alongside, same era.
A survey on large language model based autonomous agents
Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Later among the works it cites.
Fundamental limitations of alignment in large language models
Yotam Wolf, Noam Wies, Oshri Avnery, Yoav Levine, and Amnon Shashua · 2023
Later among the works it cites.
The rise and potential of large language model based agents: A survey
Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, et al · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Are you still on track!? Catching LLM task drift with activations
Sahar Abdelnabi, Aideen Fay, Giovanni Cherubin, Ahmed Salem, Mario Fritz, and Andrew Paverd · 2024
Closest in time.
Do membership inference attacks work on large language models?
Michael Duan, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettlemoyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi · 2024
Closest in time.
The ethics of advanced AI assistants
Iason Gabriel, Arianna Manzini, Geoff Keeling, Lisa Anne Hendricks, Verena Rieser, Hasan Iqbal, Nenad Tomašev, et al · 2024
Closest in time.
Operationalizing contextual integrity in privacy-conscious assistants
Sahra Ghalebikesabi, Eugene Bagdasaryan, Ren Yi, Itay Yona, Ilia Shumailov, Aneesh Pappu, Chongyang Shi, Laura Weidinger, Robert Stanforth, Leonard Berrada, et al · 2024
Closest in time.
Gemini for Google Workspace, 2024
Google · 2024
Closest in time.
Can LLMs keep a secret? Testing privacy implications of language models via contextual integrity theory
Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi · 2024
Closest in time.
Teach LLMs to phish: Stealing private information from language models
Ashwinee Panda, Christopher A. Choquette-Choo, Zhengming Zhang, Yaoqing Yang, and Prateek Mittal · 2024
Closest in time.
Toolformer: Language models can teach themselves to use tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessi, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2024
Closest in time.
Rethinking interpretability in the era of large language models
Chandan Singh, Jeevana Priya Inala, Michel Galley, Rich Caruana, and Jianfeng Gao · 2024
Closest in time.
TrustLLM: Trustworthiness in large language models
Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, Xiner Li, et al · 2024
Closest in time.
The instruction hierarchy: Training llms to prioritize privileged instructions
Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel · 2024
Closest in time.
A new era in LLM security: Exploring security concerns in real-world LLM-based systems
Fangzhou Wu, Ning Zhang, Somesh Jha, Patrick McDaniel, and Chaowei Xiao · 2024
Closest in time.
Can LLMs separate instructions from data? and what do we even mean by that?
Egor Zverev, Sahar Abdelnabi, Mario Fritz, and Christoph H Lampert · 2024
Closest in time.