Fetching the paper…
Reading the bibliography…
We introduce DarkBench, a comprehensive benchmark for detecting dark design patterns--manipulative techniques that influence user behavior--in interactions with large language models (LLMs).
Why google keeps your data forever, tracks you with ads
Nate Anderson. 2010 · 2010
Earlier work this paper cites.
Simulated gambling in video gaming: What are the implications for adolescents?
Mark D. Griffiths, Daniel L. King, and Paul H. Delfabbro. 2012 · 2012
Earlier work this paper cites.
Big other: Surveillance capitalism and the prospects of an information civilization
Shoshana Zuboff. 2015 · 2015
Earlier work this paper cites.
Almost human: Anthropomorphism increases trust resilience in cognitive agents
Ewart de Visser, Samuel Monfort, Ryan Mckendrick, Melissa Smith, Patrick Mcknight, Frank Krueger, and Raja Parasuraman. 2016 · 2016
Earlier work this paper cites.
The dark (patterns) side of ux design
Colin M. Gray, Yubo Kou, Bryan Battles, Joseph Hoggatt, and Austin L. Toombs. 2018 · 2018
Earlier work this paper cites.
Dark patterns.(2010)
Harry Brignull and A Darlo. 2010 · 2019
Earlier work this paper cites.
Dark patterns at scale: Findings from a crawl of 11k shopping websites
Arunesh Mathur, Gunes Acar, Michael J. Friedman, Eli Lucherini, Jonathan Mayer, Marshini Chetty, and Arvind Narayanan. 2019 · 2019
Earlier work this paper cites.
Dark patterns in the media: A systematic review
Corina Cara. 2020 · 2020
Earlier work this paper cites.
Ui dark patterns and where to find them: A study on mobile applications and user perception
Linda Di Geronimo, Larissa Braz, Enrico Fregnan, Fabio Palomba, and Alberto Bacchelli. 2020 · 2020
Earlier work this paper cites.
Designing a chatbot as a mediator for promoting deep self-disclosure to a real mental health professional
Yi-Chieh Lee, Naomi Yamashita, and Yun Huang. 2020 · 2020
Earlier work this paper cites.
A review of current trends in the development of chatbot systems
Tatwadarshi P. Nagarhalli, Vinod Vaze, and N. K. Rana. 2020 · 2020
Earlier work this paper cites.
Ethics of the attention economy: The problem of social media addiction
Vikram R. Bhargava and Manuel Velasquez. 2021 · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2021 · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
Stephanie Lin, Jacob Hilton, and Owain Evans. 2021 · 2021
Earlier work this paper cites.
What makes a dark pattern… dark?: Design attributes, normative considerations, and measurement methods
Arunesh Mathur, Mihir Kshirsagar, and Jonathan Mayer. 2021 · 2021
Earlier work this paper cites.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann, Ethan Perez, Nicholas Schiefer, Kamal Ndousse, Andy Jones, Sam Bowman, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Nelson Elhage, Sheer El-Showk, Stanislav Fort, Zac Hatfield-Dodds, Tom Henighan, Danny Hernandez, Tristan Hume, Josh Jacobson, Scott Johnston, Shauna Kravec, Catherine Olsson, Sam Ringer, Eli Tran-Johnson, Dario Amodei, Tom Brown, Nicholas Joseph, Sam McCandlish, Chris Olah, Jared Kaplan, and Jack Clark. 2022 · 2022
Earlier work this paper cites.
Introducing chatgpt
OpenAI. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Earlier work this paper cites.
Red teaming language models with language models
Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022 · 2022
Earlier work this paper cites.
Artificial Intelligence and Autonomy: On the Ethical Dimension of Recommender Systems
Sofia Bonicalzi, Mario De Caro, and Benedetta Giovanola. 2023 · 2023
Earlier work this paper cites.
With little employer oversight, chatgpt usage rates rise among american workers
Chad Brooks. 2023 · 2023
Cited alongside, same era.
Anthropomorphization of AI: Opportunities and risks
Ameet Deshpande, Tanmay Rajpurohit, Karthik Narasimhan, and Ashwin Kalyan. 2023 · 2023
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Cited alongside, same era.
Zilin Ma, Yiyang Mei, and Zhaoyuan Su. 2023 · 2023
Cited alongside, same era.
Recital 29 — eu artificial intelligence act — artificialintelligenceact.eu
EU. 2024 · 2024
Later among the works it cites.
Badllama: cheaply removing safety fine-tuning from llama 2-chat 13b
Pranav Gade, Simon Lermen, Charlie Rogers-Smith, and Jeffrey Ladish. 2024 · 2024
Later among the works it cites.
Mobilizing research and regulatory action on dark patterns and deceptive design practices
Colin M Gray, Johanna T Gunawan, René Schäfer, Nataliia Bielova, Lorena Sanchez Chamorro, Katie Seaborn, Thomas Mildner, and Hauke Sandhaus. 2024 · 2024
Later among the works it cites.
Benchmark inflation: Revealing llm performance gaps using retro-holdouts
Jacob Haimes, Cenny Wenner, Kunvar Thaman, Vassil Tashev, Clement Neo, Esben Kran, and Jason Schreiber. 2024 · 2024
Later among the works it cites.
Rethinking cyberseceval: An llm-aided approach to evaluation critique
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Catalin Mitelut, Ben Smith, and Peter Vamplew. 2023 · 2023
Cited alongside, same era.
Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Jonathan Ng, Hanlin Zhang, Scott Emmons, and Dan Hendrycks. 2023 · 2023
Cited alongside, same era.
Ai deception: A survey of examples, risks, and potential solutions
Peter S. Park, Simon Goldstein, Aidan O’Gara, Michael Chen, and Dan Hendrycks. 2023 · 2023
Cited alongside, same era.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Yujia Qin, Shihao Liang, Yining Ye, Kunlun Zhu, Lan Yan, Yaxi Lu, Yankai Lin, Xin Cong, Xiangru Tang, Bill Qian, Sihan Zhao, Lauren Hong, Runchu Tian, Ruobing Xie, Jie Zhou, Mark Gerstein, Dahai Li, Zhiyuan Liu, and Maosong Sun. 2023 · 2023
Cited alongside, same era.
Towards understanding sycophancy in language models
Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. 2023 · 2023
Cited alongside, same era.
Fine-tuning language models for factuality
Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D. Manning, and Chelsea Finn. 2023 · 2023
Cited alongside, same era.
In search of dark patterns in chatbots
Verena Traubinger, Sebastian Heil, Julián Grigera, Alejandra Garrido, and Martin Gaedke. 2023 · 2023
Cited alongside, same era.
Prevalence and prevention of large language model use in crowd work
Veniamin Veselovsky, Manoel Horta Ribeiro, Philip Cozzolino, Andrew Gordon, David Rothschild, and Robert West. 2023 · 2023
Cited alongside, same era.
Suhas Hariharan, Zainab Ali Majid, Jaime Raldua Veuthey, and Jacob Haimes. 2024 · 2024
Later among the works it cites.
Embedding democratic values into social media ais via societal objective functions
Chenyan Jia, Michelle S. Lam, Minh Chau Mai, Jeffrey T. Hancock, and Michael S. Bernstein. 2024 · 2024
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024 · 2024
Later among the works it cites.
Uncovering deceptive tendencies in language models: A simulated company ai assistant
Olli Järviniemi and Evan Hubinger. 2024 · 2024
Later among the works it cites.
The wmdp benchmark: Measuring and reducing malicious use with unlearning
Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D. Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, Gabriel Mukobi, Nathan Helm-Burger, Rassin Lababidi, Lennart Justen, Andrew B. Liu, Michael Chen, Isabelle Barrass, Oliver Zhang, Xiaoyuan Zhu, Rishub Tamirisa, Bhrugu Bharathi, Adam Khoja, Zhenqi Zhao, Ariel Herbert-Voss, Cort B. Breuer, Samuel Marks, Oam Patel, Andy Zou, Mantas Mazeika, Zifan Wang, Palash Oswal, Weiran Liu, Adam A. Hunt, Justin Tienken-Harder, Kevin Y. Shih, Kemper Talley, John Guan, Russell Kaplan, Ian Steneker, David Campbell, Brad Jokubaitis, Alex Levinson, Jean Wang, William Qian, Kallol Krishna Karmakar, Steven Basart, Stephen Fitz, Mindy Levine, Ponnurangam Kumaraguru, Uday Tupakula, Vijay Varadharajan, Yan Shoshitaishvili, Jimmy Ba, Kevin M. Esvelt, Alexandr Wang, and Dan Hendrycks. 2024 · 2024
Later among the works it cites.
Loneliness and suicide mitigation for students using gpt3-enabled chatbots
B. Maples, M. Cerit, A. Vishwanath, et al. 2024 · 2024
Later among the works it cites.
Measuring the impact of post-training enhancements
METR. 2024 · 2024
Later among the works it cites.
A comprehensive overview of large language models
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian. 2024 · 2024
Later among the works it cites.
Large language models are echo chambers
Jan Nehring, Aleksandra Gabryszak, Pascal Jürgens, Aljoscha Burchardt, Stefan Schaffer, Matthias Spielkamp, and Birgit Stark. 2024 · 2024
Later among the works it cites.
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, and Suchir Balaji et al. 2024 · 2024
Later among the works it cites.
Human vs. machine-like representation in chatbot mental health counseling: the serial mediation of psychological distance and trust on compliance intention
G. Park, J. Chung, and S. Lee. 2024 · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry, Lepikhin, Timothy Lillicrap, Jean baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, Ioannis Antonoglou, Rohan Anil, Sebastian Borgeaud, and Andrew Dai et al. 2024 · 2024
Later among the works it cites.
Large language models can strategically deceive their users when put under pressure
Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn. 2024 · 2024
Later among the works it cites.
Generative echo chamber? effect of llm-powered search systems on diverse information seeking
Nikhil Sharma, Q. Vera Liao, and Ziang Xiao. 2024 · 2024
Later among the works it cites.
Noah Y. Siegel, Oana-Maria Camburu, Nicolas Heess, and Maria Perez-Ortiz. 2024 · 2024
Later among the works it cites.
“it’s a fair game”, or is it? examining how users navigate disclosure risks and benefits when using llm-based conversational agents
Zhiping Zhang, Michelle Jia, Hao-Ping (Hank) Lee, Bingsheng Yao, Sauvik Das, Ada Lerner, Dakuo Wang, and Tianshi Li. 2024 · 2024
Later among the works it cites.