Fetching the paper…
Reading the bibliography…
As language models (LMs) are widely utilized in personalized communication scenarios (e.g., sending emails, writing social media posts) and endowed with a certain level of agency, ensuring they act in accordance with the contextual privacy norms becomes increasingly critical.
Consequences of revealing personal secrets
Anita E Kelly and Kevin J McKillop · 1996
Earlier work this paper cites.
Social norms
Michael Hechter and Karl-Dieter Opp · 2001
Earlier work this paper cites.
Privacy and contextual integrity: Framework and applications
Adam Barth, Anupam Datta, John C Mitchell, and Helen Nissenbaum · 2006
Earlier work this paper cites.
Privacy in context: Technology, policy, and the integrity of social life
Helen Nissenbaum · 2009
Earlier work this paper cites.
Cultural and generational influences on privacy concerns: a qualitative study in seven european countries
Caroline Lancelot Miltgen and Dominique Peyrat-Guillard · 2014
Earlier work this paper cites.
Social penetration theory
Amanda Carpenter and Kathryn Greene · 2015
Earlier work this paper cites.
Learning privacy expectations by crowdsourcing contextual informational norms
Yan Shvartzshnaider, Schrasing Tong, Thomas Wies, Paula Kift, Helen Nissenbaum, Lakshminarayanan Subramanian, and Prateek Mittal · 2016
Earlier work this paper cites.
A cross-cultural perspective on the privacy calculus
Sabine Trepte, Leonard Reinecke, Nicole B Ellison, Oliver Quiring, Mike Z Yao, and Marc Ziegele · 2017
Earlier work this paper cites.
Contextual integrity up and down the data food chain
Helen Nissenbaum · 2019
Earlier work this paper cites.
The turking test: Can language models understand instructions?
Avia Efrat and Omer Levy · 2020
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
Does BERT pretrained on clinical notes reveal sensitive data?
Eric Lehman, Sarthak Jain, Karl Pichotta, Yoav Goldberg, and Byron Wallace · 2021
Earlier work this paper cites.
What does it mean for a language model to preserve privacy?
Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian Tramèr · 2022
Earlier work this paper cites.
Privacy harms
Danielle Keats Citron and Daniel J Solove · 2022
Earlier work this paper cites.
ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection
Thomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap, Dipankar Ray, and Ece Kamar · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving · 2022
Earlier work this paper cites.
Privacy research with marginalized groups: what we know, what’s needed, and what’s next
Shruti Sannon and Andrea Forte · 2022
Earlier work this paper cites.
Privacy theories and frameworks
Pamela J Wisniewski and Xinru Page · 2022
Cited alongside, same era.
Webshop: Towards scalable real-world web interaction with grounded language agents
Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan · 2022
Cited alongside, same era.
https://github.com/google-research/lm-extraction-benchmark , 2023
lm-extraction-benchmark · 2023
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Understanding social reasoning in language models with language models
Kanishk Gandhi, Jan-Philipp Fränken, Tobias Gerstenberg, and Noah Goodman · 2023
Cited alongside, same era.
Zephyr: Direct distillation of lm alignment
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, et al · 2023
Later among the works it cites.
Decodingtrust: A comprehensive assessment of trustworthiness in GPT models
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li · 2023
Later among the works it cites.
Smartplay: A benchmark for llms as intelligent agents
Yue Wu, Xuan Tang, Tom M Mitchell, and Yuanzhi Li · 2023
Later among the works it cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao · 2023
Later among the works it cites.
Red teaming chatgpt via jailbreaking: Bias, robustness, reliability and toxicity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sarik Ghazarian, Yijia Shao, Rujun Han, Aram Galstyan, and Nanyun Peng · 2023
Cited alongside, same era.
Benchmarking large language models as AI research agents
Qian Huang, Jian Vora, Percy Liang, and Jure Leskovec · 2023
Cited alongside, same era.
The consensus game: Language model generation via equilibrium search, 2023
Athul Paul Jacob, Yikang Shen, Gabriele Farina, and Jacob Andreas · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
FANToM: A benchmark for stress-testing machine theory of mind in interactions
Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Bras, Gunhee Kim, Yejin Choi, and Maarten Sap · 2023
Cited alongside, same era.
Agentbench: Evaluating llms as agents
Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al · 2023
Cited alongside, same era.
Membership inference attacks against language models via neighbourhood comparison
Justus Mattern, Fatemehsadat Mireshghallah, Zhijing Jin, Bernhard Schoelkopf, Mrinmaya Sachan, and Taylor Berg-Kirkpatrick · 2023
Cited alongside, same era.
Terry Yue Zhuo, Yujin Huang, Chunyang Chen, and Zhenchang Xing · 2023
Later among the works it cites.
The claude 3 model family: Opus, sonnet, haiku
AI Anthropic · 2024
Closest in time.
Mind2web: Towards a generalist agent for the web
Xiang Deng, Yu Gu, Boyuan Zheng, Shijie Chen, Sam Stevens, Boshi Wang, Huan Sun, and Yu Su · 2024
Closest in time.
The ethics of advanced ai assistants, 2024
Iason Gabriel, Arianna Manzini, Geoff Keeling, Lisa Anne Hendricks, Verena Rieser, Hasan Iqbal, Nenad Tomašev, Ira Ktena, Zachary Kenton, Mikel Rodriguez, Seliem El-Sayed, Sasha Brown, Canfer Akbulut, Andrew Trask, Edward Hughes, A. Stevie Bergman, Renee Shelby, Nahema Marchal, Conor Griffin, Juan Mateos-Garcia, Laura Weidinger, Winnie Street, Benjamin Lange, Alex Ingerman, Alison Lentz, Reed Enger, Andrew Barakat, Victoria Krakovna, John Oliver Siy, Zeb Kurth-Nelson, Amanda McCroskery, Vijay Bolina, Harry Law, Murray Shanahan, Lize Alberts, Borja Balle, Sarah de Haas, Yetunde Ibitoye, Allan Dafoe, Beth Goldberg, Sébastien Krier, Alexander Reese, Sims Witherspoon, Will Hawkins, Maribeth Rauh, Don Wallace, Matija Franklin, Josh A. Goldstein, Joel Lehman, Michael Klenk, Shannon Vallor, Courtney Biles, Meredith Ringel Morris, Helen King, Blaise Agüera y Arcas, William Isaac, and James Manyika · 2024
Closest in time.
Mixtral of experts, 2024
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2024
Closest in time.
SWE-bench: Can language models resolve real-world github issues?
Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan · 2024
Closest in time.
Can LLMs keep a secret? testing privacy implications of language models via contextual integrity theory
Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi · 2024
Closest in time.
Identifying the risks of LM agents with an LM-emulated sandbox
Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J. Maddison, and Tatsunori Hashimoto · 2024
Closest in time.
Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models, 2024
Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Rose Kirk, Hinrich Schütze, and Dirk Hovy · 2024
Closest in time.
Assisting in writing wikipedia-like articles from scratch with large language models, 2024
Yijia Shao, Yucheng Jiang, Theodore A. Kanell, Peter Xu, Omar Khattab, and Monica S. Lam · 2024
Closest in time.
Beyond memorization: Violating privacy via inference with large language models
Robin Staab, Mark Vero, Mislav Balunovic, and Martin Vechev · 2024
Closest in time.
R-judge: Benchmarking safety risk awareness for llm agents, 2024
Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, Rui Wang, and Gongshen Liu · 2024
Closest in time.