Fetching the paper…
Reading the bibliography…
Conversational LLMs function as black box systems, leaving users guessing about why they see the output they do.
Eliza—a computer program for the study of natural language communication between man and machine
Joseph Weizenbaum · 1966
Earlier work this paper cites.
Automobile design liability
Richard M Goodman · 1970
Earlier work this paper cites.
Human factors in aviation
Earl L Wiener and David C Nagel · 1988
Earlier work this paper cites.
Design: Cultural probes
William Gaver, Anthony Dunne, and Elena Pacenti · 1999
Earlier work this paper cites.
Technology probes: inspiring design for and with families
Hilary Hutchinson, Wendy Mackay, Bo Westerlund, Benjamin B Bederson, Allison Druin, Catherine Plaisant, Michel Beaudouin-Lafon, Stéphane Conversy, Helen Evans, Heiko Hansen, et al · 2003
Earlier work this paper cites.
Cold war hothouses: inventing postwar culture, from cockpit to playboy
Beatriz Colomina, Annmarie Brennan, and Jeannie Kim · 2004
Earlier work this paper cites.
The problem of education-based discrimination
Stuart Tannock · 2008
Earlier work this paper cites.
The balanced accuracy and its posterior distribution
Kay Henning Brodersen, Cheng Soon Ong, Klaas Enno Stephan, and Joachim M Buhmann · 2010
Earlier work this paper cites.
Differentiation and discrimination: Understanding social class and social exclusion in leading law firms
Louise Ashley and Laura Empson · 2013
Earlier work this paper cites.
Age discrimination in the evaluation of job applicants
Ben Richardson, Janine Webb, Lynne Webber, and Kaye Smith · 2013
Earlier work this paper cites.
Gender bias in academic recruitment
Giovanni Abramo, Ciriaco Andrea D’Angelo, and Francesco Rosati · 2016
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
First i" like" it, then i hide it: Folk theories of social feeds
Motahhare Eslami, Karrie Karahalios, Christian Sandvig, Kristen Vaccaro, Aimee Rickman, Kevin Hamilton, and Alex Kirlik · 2016
Earlier work this paper cites.
Discovery of grounded theory: Strategies for qualitative research
Barney Glaser and Anselm Strauss · 2017
Earlier work this paper cites.
What are we talking about when we talk about holistic review? selective college admissions and its effects on low-ses students
Michael N Bastedo, Nicholas A Bowman, Kristen M Glasener, and Jandi L Kelly · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Earlier work this paper cites.
Older workers need not apply? ageist language in job ads and age discrimination in hiring
Ian Burn, Patrick Button, Luis Felipe Munguia Corella, and David Neumark · 2019
Earlier work this paper cites.
Human-centered tools for coping with imperfect algorithms during medical decision-making
Carrie J Cai, Emily Reif, Narayan Hegde, Jason Hipp, Been Kim, Daniel Smilkov, Martin Wattenberg, Fernanda Viegas, Greg S Corrado, Martin C Stumpe, et al · 2019
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Earlier work this paper cites.
Pandora talks: Personality and demographics on reddit
Matej Gjurković, Mladen Karan, Iva Vukojević, Mihaela Bošnjak, and Jan Šnajder · 2020
Cited alongside, same era.
Transparency and trust in artificial intelligence systems
Felix Biessmann Philipp Schmidt and Timm Teubner · 2020
Cited alongside, same era.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov · 2022
Cited alongside, same era.
” because ai is 100% right and safe”: User attitudes and sources of ai authority in india
Shivani Kapania, Oliver Siy, Gabe Clapper, Azhagu Meena SP, and Nithya Sambasivan · 2022
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations
Ethan Perez, Sam Ringer, Kamilė Lukošiūtė, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, et al · 2022
Cited alongside, same era.
The system model and the user model: Exploring ai dashboard design
Fernanda Viégas and Martin Wattenberg · 2023
Later among the works it cites.
Xintao Wang, Yaying Fei, Ziang Leng, and Cheng Li · 2023
Later among the works it cites.
Bias and fairness in chatbots: An overview
Jintang Xue, Yun-Cheng Wang, Chengwei Wei, Xiaofeng Liu, Jonghye Woo, and C-C Jay Kuo · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
L Zheng, WL Chiang, Y Sheng, S Zhuang, Z Wu, Y Zhuang, Z Lin, Z Li, D Li, and E Xing · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Meysam Alizadeh, Maël Kubli, Zeynab Samei, Shirin Dehghani, Juan Diego Bermeo, Maria Korobeynikova, and Fabrizio Gilardi · 2023
Cited alongside, same era.
Places: Prompting language models for social conversation synthesis
Maximillian Chen, Alexandros Papangelis, Chenyang Tao, Seokhwan Kim, Andy Rosenbaum, Yang Liu, Zhou Yu, and Dilek Hakkani-Tur · 2023
Cited alongside, same era.
Do models explain themselves? counterfactual simulatability of natural language explanations
Yanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao, He He, Jacob Steinhardt, Zhou Yu, and Kathleen McKeown · 2023
Cited alongside, same era.
Chatgpt outperforms crowd workers for text-annotation tasks
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli · 2023
Cited alongside, same era.
Topical-chat: Towards knowledge-grounded open-domain conversations
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tur · 2023
Cited alongside, same era.
Measuring and manipulating knowledge representations in language models
Evan Hernandez, Belinda Z Li, and Jacob Andreas · 2023
Cited alongside, same era.
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al · 2023
Later among the works it cites.
Chainforge: A visual toolkit for prompt engineering and llm hypothesis testing
Ian Arawjo, Chelse Swoopes, Priyan Vaithilingam, Martin Wattenberg, and Elena L Glassman · 2024
Closest in time.
Chatbot arena: An open platform for evaluating llms by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E Gonzalez, et al · 2024
Closest in time.
Folk psychological attributions of consciousness to large language models
Clara Colombatto and Stephen M Fleming · 2024
Closest in time.
Llm comparator: Visual analytics for side-by-side evaluation of large language models
Minsuk Kahng, Ian Tenney, Mahima Pushkarna, Michael Xieyang Liu, James Wexler, Emily Reif, Krystal Kallarackal, Minsuk Chang, Michael Terry, and Lucas Dixon · 2024
Closest in time.
Measuring and controlling persona drift in language model dialogs
Kenneth Li, Tianle Liu, Naomi Bashkansky, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2024
Closest in time.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2024
Closest in time.
React v18.2: The library for web and native user interfaces
Meta Open Source · 2024
Closest in time.
Interpreting GPT: the Logit Lens, 2020
nostalgebraist · 2024
Closest in time.
How your data is used to improve model performances, 2024
OpenAI · 2024
Closest in time.
Flask v3.0.x
Pallets · 2024
Closest in time.
ChatGPT continues to be one of the fastest-growing services ever, 2023
Jon Porter · 2024
Closest in time.
chat.openai.com Traffic & Engagement Analysis, 2024
similarweb · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman · 2024
Closest in time.