Fetching the paper…
Reading the bibliography…
Recent advances in general-purpose AI underscore the urgent need to align AI systems with human goals and values.
Some moral and technical consequences of automation: As machines learn they may develop unforeseen strategies at rates that baffle their programmers
Norbert Wiener · 1960
Earlier work this paper cites.
Self-transcendence as a human phenomenon
Viktor E Frankl · 1966
Earlier work this paper cites.
Are there universal aspects in the structure and contents of human values?
Shalom H Schwartz · 1994
Earlier work this paper cites.
The multiply motivated self
Constantine Sedikides and Michael J Strube · 1995
Earlier work this paper cites.
Intelligent agents
Michael Wooldridge · 1999
Earlier work this paper cites.
Living with socially intelligent agents
Kerstin Dautenhahn, Chrystopher L Nehaniv, and K Dautenhahn · 2000
Earlier work this paper cites.
Software law: intelligent agents: a curse or a blessing? a survey of the legal aspects of the application of intelligent software systems
Kees Stuurman and Hugo Wijnands · 2001
Earlier work this paper cites.
The truth about lie detectors (aka polygraph tests)
American Psychological Association et al · 2004
Earlier work this paper cites.
Soylent: a word processor with a crowd inside
Michael S Bernstein, Greg Little, Robert C Miller, Björn Hartmann, Mark S Ackerman, David R Karger, David Crowell, and Katrina Panovich · 2010
Earlier work this paper cites.
Social choice and individual values
Kenneth J Arrow · 2012
Earlier work this paper cites.
An overview of the schwartz theory of basic values
Shalom H Schwartz · 2012
Earlier work this paper cites.
Moral foundations theory: The pragmatic validity of moral pluralism
Jesse Graham, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P Wojcik, and Peter H Ditto · 2013
Earlier work this paper cites.
Values levers: Building ethics into design
Katie Shilton · 2013
Earlier work this paper cites.
White paper: Value alignment in autonomous systems
Stuart Russell · 2014
Earlier work this paper cites.
Algorithmic bias: From discrimination discovery to fairness-aware data mining
Sara Hajian, Francesco Bonchi, and Carlos Castillo · 2016
Earlier work this paper cites.
The social dilemma of autonomous vehicles
Jean-François Bonnefon, Azim Shariff, and Iyad Rahwan · 2016
Earlier work this paper cites.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Artificial intelligence: a modern approach
Stuart J Russell and Peter Norvig · 2016
Earlier work this paper cites.
Yes, we are worried about the existential risk of artificial intelligence
Allan Dafoe and Stuart Russell · 2016
Earlier work this paper cites.
Privacy, algorithms, and artificial intelligence
Catherine Tucker, A Agrawal, J Gans, and A Goldfarb · 2018
Earlier work this paper cites.
Introducing artificial intelligence training in medical education
Ketan Paranjape, Michiel Schinkel, Rishi Nannan Panday, Josip Car, Prabath Nanayakkara, et al · 2019
Earlier work this paper cites.
Understanding the effect of accuracy on trust in machine learning models
Ming Yin, Jennifer Wortman Vaughan, and Hanna Wallach · 2019
Earlier work this paper cites.
Improving fairness in machine learning systems: What do industry practitioners need?
Kenneth Holstein, Jennifer Wortman Vaughan, Hal Daumé III, Miro Dudik, and Hanna Wallach · 2019
Earlier work this paper cites.
Will you accept an imperfect ai? exploring designs for adjusting end-user expectations of ai systems
Rafal Kocielnik, Saleema Amershi, and Paul N Bennett · 2019
Earlier work this paper cites.
Designing theory-driven user-centric explainable ai
Danding Wang, Qian Yang, Ashraf Abdul, and Brian Y Lim · 2019
Earlier work this paper cites.
" hello ai": uncovering the onboarding needs of medical practitioners for human-ai collaborative decision-making
Carrie J Cai, Samantha Winter, David Steiner, Lauren Wilcox, and Michael Terry · 2019
Earlier work this paper cites.
Human-ai collaboration in data science: Exploring data scientists’ perceptions of automated ai
Dakuo Wang, Justin D Weisz, Michael Muller, Parikshit Ram, Werner Geyer, Casey Dugan, Yla Tausczik, Horst Samulowitz, and Alexander Gray · 2019
Earlier work this paper cites.
Aila: Attentive interactive labeling assistant for document classification through attention-based deep neural networks
Minsuk Choi, Cheonbok Park, Soyoung Yang, Yonggyu Kim, Jaegul Choo, and Sungsoo Ray Hong · 2019
Earlier work this paper cites.
Atmseer: Increasing transparency and controllability in automated machine learning
Qianwen Wang, Yao Ming, Zhihua Jin, Qiaomu Shen, Dongyu Liu, Micah J Smith, Kalyan Veeramachaneni, and Huamin Qu · 2019
Earlier work this paper cites.
What can ai do for me? evaluating machine learning interpretations in cooperative play
Shi Feng and Jordan Boyd-Graber · 2019
Earlier work this paper cites.
Guidelines for human-ai interaction
Saleema Amershi, Dan Weld, Mihaela Vorvoreanu, Adam Fourney, Besmira Nushi, Penny Collisson, Jina Suh, Shamsi Iqbal, Paul N Bennett, Kori Inkpen, et al · 2019
Earlier work this paper cites.
Allennlp interpret: A framework for explaining predictions of nlp models
Eric Wallace, Jens Tuyls, Junlin Wang, Sanjay Subramanian, Matt Gardner, and Sameer Singh · 2019
Earlier work this paper cites.
Plan, write, and revise: an interactive system for open-domain story generation
Seraphina Goldfarb-Tarrant, Haining Feng, and Nanyun Peng · 2019
Earlier work this paper cites.
On the utility of learning about humans for human-ai coordination
Micah Carroll, Rohin Shah, Mark K Ho, Tom Griffiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan · 2019
Earlier work this paper cites.
Heidl: Learning linguistic expressions with deep learning and human-in-the-loop
Prithviraj Sen, Yunyao Li, Eser Kandogan, Yiwei Yang, and Walter Lasecki · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving · 2019
Earlier work this paper cites.
Human compatible: Artificial intelligence and the problem of control, 2020
Thomas A Hemphill · 2020
Earlier work this paper cites.
Specification gaming: the flip side of ai ingenuity
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Earlier work this paper cites.
How useful are the machine-generated interpretations to general users? a human evaluation on guessing the incorrectly predicted labels
Hua Shen and Ting-Hao Huang · 2020
Earlier work this paper cites.
Artificial intelligence, values, and alignment
Iason Gabriel · 2020
Earlier work this paper cites.
Social chemistry 101: Learning to reason about social and moral norms
Maxwell Forbes, Jena D Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi · 2020
Earlier work this paper cites.
What is ai literacy? competencies and design considerations
Duri Long and Brian Magerko · 2020
Earlier work this paper cites.
Intersectional ai: A study of how information science students think about ethics and their impact
Nora McDonald and Shimei Pan · 2020
Earlier work this paper cites.
The language interpretability tool: Extensible, interactive visualizations and analysis for nlp models
Ian Tenney, James Wexler, Jasmijn Bastings, Tolga Bolukbasi, Andy Coenen, Sebastian Gehrmann, Ellen Jiang, Mahima Pushkarna, Carey Radebaugh, Emily Reif, et al · 2020
Earlier work this paper cites.
Conservatism
Andy Hamilton · 2020
Earlier work this paper cites.
Collecting the public perception of ai and robot rights
Gabriel Lima, Changyeon Kim, Seungho Ryu, Chihyung Jeon, and Meeyoung Cha · 2020
Earlier work this paper cites.
Human-ai collaboration in a cooperative game setting: Measuring social perception and outcomes
Zahra Ashktorab, Q Vera Liao, Casey Dugan, James Johnson, Qian Pan, Wei Zhang, Sadhana Kumaravel, and Murray Campbell · 2020
Earlier work this paper cites.
No explainability without accountability: An empirical study of explanations and feedback in interactive ml
Alison Smith-Renner, Ron Fan, Melissa Birchfield, Tongshuang Wu, Jordan Boyd-Graber, Daniel S Weld, and Leah Findlater · 2020
Earlier work this paper cites.
Digging into user control: perceptions of adherence and instability in transparent models
Alison Smith-Renner, Varun Kumar, Jordan Boyd-Graber, Kevin Seppi, and Leah Findlater · 2020
Earlier work this paper cites.
Questioning the ai: informing design practices for explainable ai user experiences
Q Vera Liao, Daniel Gruen, and Sarah Miller · 2020
Earlier work this paper cites.
" why is’ chicago’deceptive?" towards building model-driven tutorials for humans
Vivian Lai, Han Liu, and Chenhao Tan · 2020
Earlier work this paper cites.
Human-in-the-loop ai in government: A case study
Lanthao Benedikt, Chaitanya Joshi, Louisa Nolan, Ruben Henstra-Hill, Luke Shaw, and Sharon Hook · 2020
Earlier work this paper cites.
Conceptual metaphors impact perceptions of human-ai collaboration
Pranav Khadpe, Ranjay Krishna, Li Fei-Fei, Jeffrey T Hancock, and Michael S Bernstein · 2020
Earlier work this paper cites.
Co-designing checklists to understand organizational challenges and opportunities around fairness in ai
Michael A Madaio, Luke Stark, Jennifer Wortman Vaughan, and Hanna Wallach · 2020
Earlier work this paper cites.
Re-examining whether, why, and how human-ai interaction is uniquely difficult to design
Qian Yang, Aaron Steinfeld, Carolyn Rosé, and John Zimmerman · 2020
Earlier work this paper cites.
Multi-modal interactive task learning from demonstrations and natural language instructions
Toby Jia-Jun Li · 2020
Earlier work this paper cites.
Joint policy search for multi-agent collaboration with imperfect information
Yuandong Tian, Qucheng Gong, and Yu Jiang · 2020
Earlier work this paper cites.
Evaluating explainable ai: Which algorithmic explanations help users predict model behavior?
Peter Hase and Mohit Bansal · 2020
Earlier work this paper cites.
Interactive weak supervision: Learning useful heuristics for data labeling
Benedikt Boecking, Willie Neiswanger, Eric Xing, and Artur Dubrawski · 2020
Earlier work this paper cites.
Would you rather? a new benchmark for learning machine alignment with cultural values and social preferences
Yi Tay, Donovan Ong, Jie Fu, Alvin Chan, Nancy Chen, Anh Tuan Luu, and Christopher Pal · 2020
Earlier work this paper cites.
Find: Human-in-the-loop debugging deep text classifiers
Piyawat Lertvittayakumjorn, Lucia Specia, and Francesca Toni · 2020
Earlier work this paper cites.
Dialogue response ranking training with large-scale human feedback data
Xiang Gao, Yizhe Zhang, Michel Galley, Chris Brockett, and William B Dolan · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Earlier work this paper cites.
Human-ai interaction in human resource management: Understanding why employees resist algorithmic evaluation at workplaces and how to mitigate burdens
Hyanghee Park, Daehwan Ahn, Kartik Hosanagar, and Joonhwan Lee · 2021
Earlier work this paper cites.
A scalable framework for learning from implicit user feedback to improve natural language understanding in large-scale conversational ai systems
Sunghyun Park, Han Li, Ameen Patel, Sidharth Mudgal, Sungjin Lee, Young-Bum Kim, Spyros Matsoukas, and Ruhi Sarikaya · 2021
Earlier work this paper cites.
“was it “stated” or was it “claimed”?: How linguistic bias affects generative language models
Roma Patel and Ellie Pavlick · 2021
Earlier work this paper cites.
Crowdplay: Crowdsourcing human demonstrations for offline learning
Matthias Gerstgrasser, Rakshit Trivedi, and David C Parkes · 2021
Earlier work this paper cites.
Problematic machine behavior: A systematic literature review of algorithm audits
Jack Bandy · 2021
Earlier work this paper cites.
The design of reciprocal learning between human and artificial intelligence
Alexey Zagalsky, Dov Te’eni, Inbal Yahav, David G Schwartz, Gahl Silverman, Daniel Cohen, Yossi Mann, and Dafna Lewinsky · 2021
Earlier work this paper cites.
Say it all: Feedback for improving non-visual presentation accessibility
Yi-Hao Peng, JiWoong Jang, Jeffrey P Bigham, and Amy Pavel · 2021
Earlier work this paper cites.
The importance of modeling social factors of language: Theory and practice
Dirk Hovy and Diyi Yang · 2021
Earlier work this paper cites.
The prisma 2020 statement: an updated guideline for reporting systematic reviews
Matthew J Page, Joanne E McKenzie, Patrick M Bossuyt, Isabelle Boutron, Tammy C Hoffmann, Cynthia D Mulrow, Larissa Shamseer, Jennifer M Tetzlaff, Elie A Akl, Sue E Brennan, et al · 2021
Earlier work this paper cites.
Who is included in human perceptions of ai?: Trust and perceived fairness around healthcare ai and cultural mistrust
Min Kyung Lee and Katherine Rich · 2021
Earlier work this paper cites.
How should ai systems talk to users when collecting their personal information? effects of role framing and self-referencing on human-ai interaction
Mengqi Liao and S Shyam Sundar · 2021
Earlier work this paper cites.
Human perceptions on moral responsibility of ai: A case study in ai-assisted bail decision-making
Gabriel Lima, Nina Grgić-Hlača, and Meeyoung Cha · 2021
Earlier work this paper cites.
Towards fairness in practice: A practitioner-oriented rubric for evaluating fair ml toolkits
Brianna Richardson, Jean Garcia-Gathright, Samuel F Way, Jennifer Thom, and Henriette Cramer · 2021
Earlier work this paper cites.
Effect of information presentation on fairness perceptions of machine learning predictors
Niels Van Berkel, Jorge Goncalves, Daniel Russo, Simo Hosio, and Mikael B Skov · 2021
Earlier work this paper cites.
Towards mutual theory of mind in human-ai interaction: How language reflects what students perceive about a virtual teaching assistant
Qiaosi Wang, Koustuv Saha, Eric Gregori, David Joyner, and Ashok Goel · 2021
Earlier work this paper cites.
The landscape and gaps in open source fairness toolkits
Michelle Seng Ah Lee and Jat Singh · 2021
Earlier work this paper cites.
Where responsible ai meets reality: Practitioner perspectives on enablers for shifting organizational practices
Bogdana Rakova, Jingying Yang, Henriette Cramer, and Rumman Chowdhury · 2021
Earlier work this paper cites.
From philosophy to interfaces: an explanatory method and a tool inspired by achinstein’s theory of explanation
Francesco Sovrano and Fabio Vitali · 2021
Earlier work this paper cites.
Visual, textual or hybrid: the effect of user expertise on different explanations
Maxwell Szymanski, Martijn Millecamp, and Katrien Verbert · 2021
Earlier work this paper cites.
How to support users in understanding intelligent systems? structuring the discussion
Malin Eiband, Daniel Buschek, and Heinrich Hussmann · 2021
Earlier work this paper cites.
I think i get your point, ai! the illusion of explanatory depth in explainable ai
Michael Chromik, Malin Eiband, Felicitas Buchner, Adrian Krüger, and Andreas Butz · 2021
Earlier work this paper cites.
Are explanations helpful? a comparative study of the effects of explanations in ai-assisted decision-making
Xinru Wang and Ming Yin · 2021
Earlier work this paper cites.
Planning for natural language failures with the ai playbook
Matthew K Hong, Adam Fourney, Derek DeBellis, and Saleema Amershi · 2021
Earlier work this paper cites.
The effects of warmth and competence perceptions on users’ choice of an ai system
Zohar Gilad, Ofra Amir, and Liat Levontin · 2021
Earlier work this paper cites.
To trust or to think: cognitive forcing functions can reduce overreliance on ai in ai-assisted decision-making
Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z Gajos · 2021
Earlier work this paper cites.
How to evaluate trust in ai-assisted decision making? a survey of empirical methodologies
Oleksandra Vereschak, Gilles Bailly, and Baptiste Caramiaux · 2021
Earlier work this paper cites.
Perfection not required? human-ai partnerships in code translation
Justin D Weisz, Michael Muller, Stephanie Houde, John Richards, Steven I Ross, Fernando Martinez, Mayank Agarwal, and Kartik Talamadupula · 2021
Earlier work this paper cites.
Datasheets for datasets help ml engineers notice and understand ethical issues in training data
Karen L Boyd · 2021
Earlier work this paper cites.
Do datasets have politics? disciplinary values in computer vision dataset development
Morgan Klaus Scheuerman, Alex Hanna, and Emily Denton · 2021
Earlier work this paper cites.
Ai in global health: the view from the front lines
Azra Ismail and Neha Kumar · 2021
Earlier work this paper cites.
Avoiding the turing tarpit: Learning conversational programming by starting from code’s purpose
Kathryn Cunningham, Barbara J Ericson, Rahul Agrawal Bejarano, and Mark Guzdial · 2021
Earlier work this paper cites.
Cody: An ai-based system to semi-automate coding for qualitative research
Tim Rietz and Alexander Maedche · 2021
Earlier work this paper cites.
Does the whole exceed its parts? the effect of ai explanations on complementary team performance
Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel Weld · 2021
Earlier work this paper cites.
Designing ai for trust and collaboration in time-constrained medical decisions: a sociotechnical lens
Maia Jacobs, Jeffrey He, Melanie F. Pradier, Barbara Lam, Andrew C Ahn, Thomas H McCoy, Roy H Perlis, Finale Doshi-Velez, and Krzysztof Z Gajos · 2021
Earlier work this paper cites.
Supporting serendipity: Opportunities and challenges for human-ai collaboration in qualitative analysis
Jialun Aaron Jiang, Kandrea Wade, Casey Fiesler, and Jed R Brubaker · 2021
Earlier work this paper cites.
Understanding the effect of out-of-distribution examples and interactive explanations on human-ai decision making
Han Liu, Vivian Lai, and Chenhao Tan · 2021
Earlier work this paper cites.
" an ideal human" expectations of ai teammates in human-ai teaming
Rui Zhang, Nathan J McNeese, Guo Freeman, and Geoff Musick · 2021
Earlier work this paper cites.
Marcelle: composing interactive machine learning workflows and interfaces
Jules Françoise, Baptiste Caramiaux, and Téo Sanchez · 2021
Earlier work this paper cites.
Discovering and validating ai errors with crowdsourced failure reports
Ángel Alexander Cabrera, Abraham J Druck, Jason I Hong, and Adam Perer · 2021
Earlier work this paper cites.
Recast: Enabling user recourse and interpretability of toxicity detection models with interactive visualization
Austin P Wright, Omar Shaikh, Haekyu Park, Will Epperson, Muhammed Ahmed, Stephane Pinel, Duen Horng Chau, and Diyi Yang · 2021
Earlier work this paper cites.
Increasing the speed and accuracy of data labeling through an ai assisted interface
Michael Desmond, Michael Muller, Zahra Ashktorab, Casey Dugan, Evelyn Duesterwald, Kristina Brimijoin, Catherine Finegan-Dollak, Michelle Brachman, Aabhas Sharma, Narendra Nath Joshi, et al · 2021
Earlier work this paper cites.
Transferable dialogue systems and user simulators
Bo-Hsiang Tseng, Yinpei Dai, Florian Kreyssig, and Bill Byrne · 2021
Earlier work this paper cites.
Case study: Deontological ethics in nlp
Shrimai Prabhumoye, Brendon Boldt, Ruslan Salakhutdinov, and Alan W Black · 2021
Earlier work this paper cites.
The utility of explainable ai in ad hoc human-machine teaming
Rohan Paleja, Muyleng Ghuy, Nadun Ranawaka Arachchige, Reed Jensen, and Matthew Gombolay · 2021
Earlier work this paper cites.
A human-machine collaborative framework for evaluating malevolence in dialogues
Yangjun Zhang, Pengjie Ren, and Maarten de Rijke · 2021
Earlier work this paper cites.
Summvis: Interactive visual analysis of models, data, and evaluation for text summarization
Jesse Vig, Wojciech Kryściński, Karan Goel, and Nazneen Rajani · 2021
Earlier work this paper cites.
Societal biases in language generation: Progress and challenges
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng · 2021
Earlier work this paper cites.
Widening the pipeline in human-guided reinforcement learning with explanation and context-aware data augmentation
Lin Guan, Mudit Verma, Suna Sihang Guo, Ruohan Zhang, and Subbarao Kambhampati · 2021
Earlier work this paper cites.
Textual time travel: A temporally informed approach to theory of mind
Akshatha Arodi and Jackie Chi Kit Cheung · 2021
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, et al · 2021
Earlier work this paper cites.
Surveying wonderland for many more literature visualization techniques
Richard Brath · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al · 2022
Earlier work this paper cites.
Shared interest: Measuring human-ai alignment to identify recurring patterns in model behavior
Angie Boggust, Benjamin Hoover, Arvind Satyanarayan, and Hendrik Strobelt · 2022
Earlier work this paper cites.
Responsible ai systems: who are the stakeholders?
Advait Deshpande and Helen Sharp · 2022
Earlier work this paper cites.
Identifying the human values behind arguments
Johannes Kiesel, Milad Alshomary, Nicolas Handke, Xiaoni Cai, Henning Wachsmuth, and Benno Stein · 2022
Earlier work this paper cites.
When to make exceptions: Exploring language models as accounts of human moral judgment
Zhijing Jin, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, and Bernhard Schölkopf · 2022
Earlier work this paper cites.
Aligning to social norms and values in interactive narratives
Prithviraj Ammanabrolu, Liwei Jiang, Maarten Sap, Hannaneh Hajishirzi, and Yejin Choi · 2022
Earlier work this paper cites.
Moral: Aligning ai with human norms through multi-objective reinforced active learning
Markus Peschl, Arkady Zgonnikov, Frans A Oliehoek, and Luciano C Siebert · 2022
Earlier work this paper cites.
Bertscore is unfair: On social bias in language model-based metrics for text generation
Tianxiang Sun, Junliang He, Xipeng Qiu, and Xuan-Jing Huang · 2022
Earlier work this paper cites.
Potato: The portable text annotation tool
Jiaxin Pei, Aparna Ananthasubramaniam, Xingyao Wang, Naitian Zhou, Apostolos Dedeloudis, Jackson Sargent, and David Jurgens · 2022
Earlier work this paper cites.
Sensible ai: Re-imagining interpretability and explainability using sensemaking theory
Harmanpreet Kaur, Eytan Adar, Eric Gilbert, and Cliff Lampe · 2022
Earlier work this paper cites.
Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts
Tongshuang Wu, Michael Terry, and Carrie Jun Cai · 2022
Earlier work this paper cites.
Fine-tuning language models to find agreement among humans with diverse preferences
Michiel Bakker, Martin Chadwick, Hannah Sheahan, Michael Tessler, Lucy Campbell-Gillingham, Jan Balaguer, Nat McAleese, Amelia Glaese, John Aslanides, Matt Botvinick, et al · 2022
Earlier work this paper cites.
Human vs. ai: Understanding the impact of anthropomorphism on consumer response to chatbots from the perspective of trust and relationship norms
Xusen Cheng, Xiaoping Zhang, Jason Cohen, and Jian Mou · 2022
Earlier work this paper cites.
Ai is mastering language. should we trust what it says?
Steven Johnson and Nikita Iziev · 2022
Earlier work this paper cites.
Owning mistakes sincerely: Strategies for mitigating ai errors
Amama Mahmood, Jeanie W Fung, Isabel Won, and Chien-Ming Huang · 2022
Earlier work this paper cites.
Building trust in interactive machine learning via user contributed interpretable rules
Lijie Guo, Elizabeth M Daly, Oznur Alkan, Massimiliano Mattetti, Owen Cornec, and Bart Knijnenburg · 2022
Earlier work this paper cites.
Examining ai methods for micro-coaching dialogs
Elliot Mitchell, Noemie Elhadad, and Lena Mamykina · 2022
Earlier work this paper cites.
Capable but amoral? comparing ai and human expert collaboration in ethical decision making
Suzanne Tolmeijer, Markus Christen, Serhiy Kandul, Markus Kneer, and Abraham Bernstein · 2022
Earlier work this paper cites.
You complete me: Human-ai teams and complementary expertise
Qiaoning Zhang, Matthew L Lee, and Scott Carter · 2022
Earlier work this paper cites.
Exploring the effects of machine learning literacy interventions on laypeople’s reliance on machine learning models
Chun-Wei Chiang and Ming Yin · 2022
Earlier work this paper cites.
Hint: Integration testing for ai-based features with humans in the loop
Quan Ze Chen, Tobias Schnabel, Besmira Nushi, and Saleema Amershi · 2022
Earlier work this paper cites.
Do people engage cognitively with ai? impact of ai assistance on incidental learning
Krzysztof Z Gajos and Lena Mamykina · 2022
Earlier work this paper cites.
Improving human-ai partnerships in child welfare: understanding worker practices, challenges, and desires for algorithmic decision support
Anna Kawakami, Venkatesh Sivaraman, Hao-Fei Cheng, Logan Stapleton, Yanghuidi Cheng, Diana Qing, Adam Perer, Zhiwei Steven Wu, Haiyi Zhu, and Kenneth Holstein · 2022
Earlier work this paper cites.
The deskilling of domain expertise in ai development
Nithya Sambasivan and Rajesh Veeraraghavan · 2022
Earlier work this paper cites.
Two heads are better than one: A dimension space for unifying human and artificial intelligence in shared control
Gabriele Cimolino and TC Nicholas Graham · 2022
Earlier work this paper cites.
Storybuddy: A human-ai collaborative chatbot for parent-child interactive storytelling with flexible parental involvement
Zheng Zhang, Ying Xu, Yanhao Wang, Bingsheng Yao, Daniel Ritchie, Tongshuang Wu, Mo Yu, Dakuo Wang, and Toby Jia-Jun Li · 2022
Earlier work this paper cites.
Interpretable directed diversity: Leveraging model explanations for iterative crowd ideation
Yunlong Wang, Priyadarshini Venkatesh, and Brian Y Lim · 2022
Earlier work this paper cites.
Intuitively assessing ml model reliability through example-based explanations and editing model inputs
Harini Suresh, Kathleen M Lewis, John Guttag, and Arvind Satyanarayan · 2022
Earlier work this paper cites.
Scholastic: Graphical human-ai collaboration for inductive and interpretive text analysis
Matt-Heun Hong, Lauren A Marsh, Jessica L Feuston, Janet Ruppert, Jed R Brubaker, and Danielle Albers Szafir · 2022
Earlier work this paper cites.
Beyond text generation: Supporting writers with continuous automatic text summaries
Hai Dang, Karim Benharrak, Florian Lehmann, and Daniel Buschek · 2022
Earlier work this paper cites.
Human-ai collaboration via conditional delegation: A case study of content moderation
Vivian Lai, Samuel Carton, Rajat Bhatnagar, Q Vera Liao, Yunfeng Zhang, and Chenhao Tan · 2022
Earlier work this paper cites.
Talebrush: Sketching stories with generative pretrained language models
John Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee, Eytan Adar, and Minsuk Chang · 2022
Earlier work this paper cites.
Relative behavioral attributes: Filling the gap between symbolic goal specification and reward learning from human preferences
Lin Guan, Karthik Valmeekam, and Subbarao Kambhampati · 2022
Earlier work this paper cites.
The authenticity gap in human evaluation
Kawin Ethayarajh and Dan Jurafsky · 2022
Earlier work this paper cites.
Simulating bandit learning from user feedback for extractive question answering
Ge Gao, Eunsol Choi, and Yoav Artzi · 2022
Earlier work this paper cites.
In the eye of the beholder: Robust prediction with causal user modeling
Amir Feder, Guy Horowitz, Yoav Wald, Roi Reichart, and Nir Rosenfeld · 2022
Earlier work this paper cites.
Capturing failures of large language models via human cognitive biases
Erik Jones and Jacob Steinhardt · 2022
Earlier work this paper cites.
First contact: Unsupervised human-machine co-adaptation via mutual information maximization
Siddharth Reddy, Sergey Levine, and Anca Dragan · 2022
Cited alongside, same era.
Peer: A collaborative language model
Timo Schick, A Yu Jane, Zhengbao Jiang, Fabio Petroni, Patrick Lewis, Gautier Izacard, Qingfei You, Christoforos Nalmpantis, Edouard Grave, and Sebastian Riedel · 2022
Cited alongside, same era.
Mapping the design space of human-ai interaction in text summarization
Ruijia Cheng, Alison Smith-Renner, Ke Zhang, Joel Tetreault, and Alejandro Jaimes-Larrarte · 2022
Cited alongside, same era.
Towards interpretable deep reinforcement learning with human-friendly prototypes
Eoin M Kenny, Mycal Tucker, and Julie Shah · 2022
Cited alongside, same era.
Are shortest rationales the best explanations for human understanding?
Hua Shen, Tongshuang Wu, Wenbo Guo, and Ting-Hao Huang · 2022
Cited alongside, same era.
Mind the biases: Quantifying cognitive biases in language model prompting
Ruixi Lin and Hwee Tou Ng · 2023
Later among the works it cites.
Learning from active human involvement through proxy value propagation
Zhenghao Mark Peng, Wenjie Mo, Chenda Duan, Quanyi Li, and Bolei Zhou · 2023
Later among the works it cites.
Clear: Continual learning on algorithmic reasoning for human-like intelligence
Bong Gyun Kang, HyunGi Kim, Dahuin Jung, and Sungroh Yoon · 2023
Later among the works it cites.
Camel: Communicative agents for" mind" exploration of large language model society
Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem · 2023
Later among the works it cites.
Learning multi-agent behaviors from distributed and streaming demonstrations
Shicheng Liu and Minghui Zhu · 2023
Later among the works it cites.
Imitation learning from vague feedback
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A Smith, and Daniel S Weld · 2022
Cited alongside, same era.
The why and the how: A survey on natural language interaction in visualization
Henrik Voigt, Özge Alaçam, Monique Meuschke, Kai Lawonn, and Sina Zarrieß · 2022
Cited alongside, same era.
A rationale-centric framework for human-in-the-loop machine learning
Jinghui Lu, Linyi Yang, Brian Namee, and Yue Zhang · 2022
Cited alongside, same era.
Human-machine collaboration approaches to build a dialogue dataset for hate speech countering
Helena Bonaldi, Sara Dellantonio, Serra Sinem Tekiroğlu, and Marco Guerini · 2022
Cited alongside, same era.
Lm-debugger: An interactive tool for inspection and intervention in transformer-based language models
Mor Geva, Avi Caciularu, Guy Dar, Paul Roit, Shoval Sadde, Micah Shlain, Bar Tamir, and Yoav Goldberg · 2022
Cited alongside, same era.
Reframing human-ai collaboration for generating free-text explanations
Sarah Wiegreffe, Jack Hessel, Swabha Swayamdipta, Mark Riedl, and Yejin Choi · 2022
Cited alongside, same era.
Using interactive feedback to improve the accuracy and explainability of question answering systems post-deployment
Zichao Li, Prakhar Sharma, Xing Han Lu, Jackie Chi Kit Cheung, and Siva Reddy · 2022
Cited alongside, same era.
Xin-Qiang Cai, Yu-Jie Zhang, Chao-Kai Chiang, and Masashi Sugiyama · 2023
Later among the works it cites.
Learning zero-shot cooperation with humans, assuming humans are biased
Chao Yu, Jiaxuan Gao, Weilin Liu, Botian Xu, Hao Tang, Jiaqi Yang, Yu Wang, and Yi Wu · 2023
Later among the works it cites.
Learning to simulate natural language feedback for interactive semantic parsing
Hao Yan, Saurabh Srivastava, Yintao Tai, Sida I Wang, Wen-tau Yih, and Ziyu Yao · 2023
Later among the works it cites.
Which spurious correlations impact reasoning in nli models? a visual interactive diagnosis through data-constrained counterfactuals
Robin Chan, Afra Amini, and Mennatallah El-Assady · 2023
Later among the works it cites.
Xmd: An end-to-end framework for interactive explanation-based debugging of nlp models
Dong-Ho Lee, Akshen Kadakia, Brihi Joshi, Aaron Chan, Ziyi Liu, Kiran Narahari, Takashi Shibuya, Ryosuke Mitani, Toshiyuki Sekiya, Jay Pujara, et al · 2023
Later among the works it cites.
Theory of mind for multi-agent collaboration via large language models
Huao Li, Yu Chong, Simon Stepputtis, Joseph P Campbell, Dana Hughes, Charles Lewis, and Katia Sycara · 2023
Later among the works it cites.
Harnessing the power of llms: Evaluating human-ai text co-creation through the lens of news headline generation
Zijian Ding, Alison Smith-Renner, Wenjuan Zhang, Joel Tetreault, and Alejandro Jaimes · 2023
Later among the works it cites.
Urial: Aligning untuned llms with just the’write’amount of in-context learning
Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi · 2023
Later among the works it cites.
Identifying the risks of lm agents with an lm-emulated sandbox
Yangjun Ruan, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J Maddison, and Tatsunori Hashimoto · 2023
Later among the works it cites.
Knowledge of cultural moral norms in large language models
Aida Ramezani and Yang Xu · 2023
Later among the works it cites.
Sociocultural norm similarities and differences via situational alignment and explainable textual entailment
CH-Wang Sky, Arkadiy Saakyan, Oliver Li, Zhou Yu, and Smaranda Muresan · 2023
Later among the works it cites.
This prompt is measuring< mask>: evaluating bias evaluation in language models
Seraphina Goldfarb-Tarrant, Eddie Ungless, Esma Balkir, and Su Lin Blodgett · 2023
Later among the works it cites.
Towards a holistic landscape of situated theory of mind in large language models
Ziqiao Ma, Jacob Sansom, Run Peng, and Joyce Chai · 2023
Later among the works it cites.
Ethical reasoning over moral alignment: A case and framework for in-context ethical policies in llms
Abhinav Rao, Aditi Khandelwal, Kumar Tanmay, Utkarsh Agarwal, and Monojit Choudhury · 2023
Later among the works it cites.
On the humanity of conversational ai: Evaluating the psychological portrayal of llms
Jen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho LAM, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, and Michael Lyu · 2023
Later among the works it cites.
Unmasking and improving data credibility: A study with datasets for training harmless language models
Zhaowei Zhu, Jialu Wang, Hao Cheng, and Yang Liu · 2023
Later among the works it cites.
Alignscore: Evaluating factual consistency with a unified alignment function
Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu · 2023
Later among the works it cites.
From pretraining data to language models to downstream tasks: Tracking the trails of political biases leading to unfair nlp models
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov · 2023
Later among the works it cites.
Are machine rationales (not) useful to humans? measuring and improving human utility of free-text rationales
Brihi Joshi, Ziyi Liu, Sahana Ramnath, Aaron Chan, Zhewei Tong, Shaoliang Nie, Qifan Wang, Yejin Choi, and Xiang Ren · 2023
Later among the works it cites.
Are human explanations always helpful? towards objective evaluation of human natural language explanations
Bingsheng Yao, Prithviraj Sen, Lucian Popa, James Hendler, and Dakuo Wang · 2023
Later among the works it cites.
Finspector: A human-centered visual inspection tool for exploring and comparing biases among foundation models
Bum Chul Kwon and Nandana Mihindukulasooriya · 2023
Later among the works it cites.
A diachronic perspective on user trust in ai under uncertainty
Shehzaad Dhuliawala, Vilém Zouhar, Mennatallah El-Assady, and Mrinmaya Sachan · 2023
Later among the works it cites.
Can large language models capture dissenting human voices?
Noah Lee, Na Min An, and James Thorne · 2023
Later among the works it cites.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Katherine Tian, Eric Mitchell, Allan Zhou, Archit Sharma, Rafael Rafailov, Huaxiu Yao, Chelsea Finn, and Christopher D Manning · 2023
Later among the works it cites.
Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation
Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang · 2023
Later among the works it cites.
Interactive text-to-sql generation via editable step-by-step explanations
Yuan Tian, Zheng Zhang, Zheng Ning, Toby Li, Jonathan K Kummerfeld, and Tianyi Zhang · 2023
Later among the works it cites.
Intelmo: Enhancing models’ adoption of interactive interfaces
Chunxu Yang, Chien-Sheng Wu, Lidiya Murakhovs’ka, Philippe Laban, and Xiang Chen · 2023
Later among the works it cites.
Summhelper: Collaborative human-computer summarization
Aviv Slobodkin, Niv Nachum, Shmuel Amar, Ori Shapira, and Ido Dagan · 2023
Later among the works it cites.
Improving summarization with human edits
Zonghai Yao, Benjamin Schloss, and Sai Selvaraj · 2023
Later among the works it cites.
Increasing diversity while maintaining accuracy: Text data generation with large language models and human interventions
John Chung, Ece Kamar, and Saleema Amershi · 2023
Later among the works it cites.
Pk-icr: Persona-knowledge interactive multi-context retrieval for grounded dialogue
Minsik Oh, Joosung Lee, Jiwei Li, and Guoyin Wang · 2023
Later among the works it cites.
Modelscope-agent: Building your customizable agent system with open-source large language models
Chenliang Li, He Chen, Ming Yan, Weizhou Shen, Haiyang Xu, Zhikai Wu, Zhicheng Zhang, Wenmeng Zhou, Yingda Chen, Chen Cheng, et al · 2023
Later among the works it cites.
The past, present and better future of feedback learning in large language models for subjective human preferences and values
Hannah Kirk, Andrew Bean, Bertie Vidgen, Paul Röttger, and Scott Hale · 2023
Later among the works it cites.
Fairprism: evaluating fairness-related harms in text generation
Eve Fleisig, Aubrie Amstutz, Chad Atalla, Su Lin Blodgett, Hal Daumé III, Alexandra Olteanu, Emily Sheng, Dan Vann, and Hanna Wallach · 2023
Later among the works it cites.
Factually consistent summarization via reinforcement learning with textual entailment feedback
Paul Roit, Johan Ferret, Lior Shani, Roee Aharoni, Geoffrey Cideron, Robert Dadashi, Matthieu Geist, Sertan Girgin, Leonard Hussenot, Orgad Keller, et al · 2023
Later among the works it cites.
Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs
Afra Feyza Akyürek, Ekin Akyürek, Ashwin Kalyan, Peter Clark, Derry Tanti Wijaya, and Niket Tandon · 2023
Later among the works it cites.
Learning to generate equitable text in dialogue from biased training data
Anthony Sicilia and Malihe Alikhani · 2023
Later among the works it cites.
Mitigating label biases for in-context learning
Yu Fei, Yifan Hou, Zeming Chen, and Antoine Bosselut · 2023
Later among the works it cites.
Element-aware summarization with large language models: Expert-aligned evaluation and chain-of-thought method
Yiming Wang, Zhuosheng Zhang, and Rui Wang · 2023
Later among the works it cites.
Improving gender fairness of pre-trained language models without catastrophic forgetting
Zahra Fatemi, Chen Xing, Wenhao Liu, and Caimming Xiong · 2023
Later among the works it cites.
The tail wagging the dog: Dataset construction biases of social bias benchmarks
Nikil Selvam, Sunipa Dev, Daniel Khashabi, Tushar Khot, and Kai-Wei Chang · 2023
Later among the works it cites.
Democratizing reasoning ability: Tailored learning from large language model
Zhaoyang Wang, Shaohan Huang, Yuxuan Liu, Jiahai Wang, Minghui Song, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, et al · 2023
Later among the works it cites.
Learning new skills after deployment: Improving open-domain internet-driven dialogue with human feedback
Jing Xu, Megan Ung, Mojtaba Komeili, Kushal Arora, Y-Lan Boureau, and Jason Weston · 2023
Later among the works it cites.
Continually improving extractive qa via human feedback
Ge Gao, Hung-Ting Chen, Yoav Artzi, and Eunsol Choi · 2023
Later among the works it cites.
Robbie: Robust bias evaluation of large generative language models
David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, and Eric Michael Smith · 2023
Later among the works it cites.
Coannotating: Uncertainty-guided work allocation between human and large language models for data annotation
Minzhi Li, Taiwei Shi, Caleb Ziems, Min-Yen Kan, Nancy Chen, Zhengyuan Liu, and Diyi Yang · 2023
Later among the works it cites.
Actor: Active learning with annotator-specific classification heads to embrace human label variation
Xinpeng Wang and Barbara Plank · 2023
Later among the works it cites.
Interfair: Debiasing with natural language feedback for fair interpretable predictions
Bodhisattwa Majumder, Zexue He, and Julian McAuley · 2023
Later among the works it cites.
Coffee: Counterfactual fairness for personalized text generation in explainable recommendation
Nan Wang, Qifan Wang, Yi-Chia Wang, Maziar Sanjabi, Jingzhou Liu, Hamed Firooz, Hongning Wang, and Shaoliang Nie · 2023
Later among the works it cites.
Be selfish, but wisely: Investigating the impact of agent personality in mixed-motive human-agent interactions
Kushal Chawla, Ian Wu, Yu Rong, Gale Lucas, and Jonathan Gratch · 2023
Later among the works it cites.
Fantom: A benchmark for stress-testing machine theory of mind in interactions
Hyunwoo Kim, Melanie Sclar, Xuhui Zhou, Ronan Bras, Gunhee Kim, Yejin Choi, and Maarten Sap · 2023
Later among the works it cites.
Human-in-the-loop abstractive dialogue summarization
Jiaao Chen, Mohan Dodda, and Diyi Yang · 2023
Later among the works it cites.
Nano: Nested human-in-the-loop reward learning for few-shot language model control
Xiang Fan, Yiwei Lyu, Paul Pu Liang, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2023
Later among the works it cites.
Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback, 2023
Wei Shen, Rui Zheng, Wenyu Zhan, Jun Zhao, Shihan Dou, Tao Gui, Qi Zhang, and Xuan-Jing Huang · 2023
Later among the works it cites.
Ldm2: A large decision model imitating human cognition with dynamic memory enhancement
Xingjin Wang, Linjing Li, and Daniel Zeng · 2023
Later among the works it cites.
Getting more out of mixture of language model reasoning experts
Chenglei Si, Weijia Shi, Chen Zhao, Luke Zettlemoyer, and Jordan Boyd-Graber · 2023
Later among the works it cites.
Improving conversational recommendation systems via bias analysis and language-model-enhanced data augmentation
Xi Wang, Hossein Rahmani, Jiqun Liu, and Emine Yilmaz · 2023
Later among the works it cites.
Harnessing the power of large language models for empathetic response generation: Empirical investigations and improvements
Yushan Qian, Weinan Zhang, and Ting Liu · 2023
Later among the works it cites.
Autoplan: Automatic planning of interactive decision-making tasks with large language models
Siqi Ouyang and Lei Li · 2023
Later among the works it cites.
The iron (ic) melting pot: Reviewing human evaluation in humour, irony and sarcasm generation
Tyler Loakman, Aaron Maladry, and Chenghua Lin · 2023
Later among the works it cites.
Thought cloning: Learning to think while acting by imitating human thinking
Shengran Hu and Jeff Clune · 2023
Later among the works it cites.
Rrhf: Rank responses to align language models with human feedback
Hongyi Yuan, Zheng Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang · 2023
Later among the works it cites.
Learning to influence human behavior with offline reinforcement learning
Joey Hong, Sergey Levine, and Anca Dragan · 2023
Later among the works it cites.
Alignment with human representations supports robust few-shot learning
Ilia Sucholutsky and Tom Griffiths · 2023
Later among the works it cites.
Interactive multi-fidelity learning for cost-effective adaptation of language model with sparse human supervision
Jiaxin Zhang, Zhuohang Li, Kamalika Das, and Sricharan Kumar · 2023
Later among the works it cites.
Collaborative alignment of nlp models
Fereshte Khani and Marco Tulio Ribeiro · 2023
Later among the works it cites.
Diverse conventions for human-ai collaboration
Bidipta Sarkar, Andy Shih, and Dorsa Sadigh · 2023
Later among the works it cites.
Human-aligned calibration for ai-assisted decision making
Nina Corvelo Benz and Manuel Rodriguez · 2023
Later among the works it cites.
Beyond labels: Empowering human annotators with natural language explanations through a novel active-learning architecture
Bingsheng Yao, Ishan Jindal, Lucian Popa, Yannis Katsis, Sayan Ghosh, Lihong He, Yuxuan Lu, Shashank Srivastava, Yunyao Li, James Hendler, et al · 2023
Later among the works it cites.
An efficient end-to-end training approach for zero-shot human-ai coordination
Xue Yan, Jiaxian Guo, Xingzhou Lou, Jun Wang, Haifeng Zhang, and Yali Du · 2023
Later among the works it cites.
Swiftsage: A generative agent with fast and slow thinking for complex interactive tasks
Bill Yuchen Lin, Yicheng Fu, Karina Yang, Faeze Brahman, Shiyu Huang, Chandra Bhagavatula, Prithviraj Ammanabrolu, Yejin Choi, and Xiang Ren · 2023
Later among the works it cites.
Breadcrumbs to the goal: Supervised goal selection from human-in-the-loop feedback
Marcel Torne Villasevil, Balsells I Pamies, Zihan Wang, Samedh Desai, Tao Chen, Pulkit Agrawal, Abhishek Gupta, et al · 2023
Later among the works it cites.
Fine-grained human feedback gives better rewards for language model training
Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Dialguide: Aligning dialogue model behavior with developer guidelines
Prakhar Gupta, Yang Liu, Di Jin, Behnam Hedayatnia, Spandana Gella, Sijia Liu, Patrick L Lange, Julia Hirschberg, and Dilek Hakkani-Tur · 2023
Later among the works it cites.
Effective human-ai teams via learned natural language rules and onboarding
Hussein Mozannar, Jimin Lee, Dennis Wei, Prasanna Sattigeri, Subhro Das, and David Sontag · 2023
Later among the works it cites.
Learning human-compatible representations for case-based decision support
Han Liu, Yizhou Tian, Chacha Chen, Shi Feng, Yuxin Chen, and Chenhao Tan · 2023
Later among the works it cites.
Asymptotic instance-optimal algorithms for interactive decision making
Kefan Dong and Tengyu Ma · 2023
Later among the works it cites.
Human-guided fair classification for natural language processing
Florian E Dorner, Momchil Peychev, Nikola Konstantinov, Naman Goel, Elliott Ash, and Martin Vechev · 2023
Later among the works it cites.
Improving generalization of alignment with human preferences through group invariant learning
Rui Zheng, Wei Shen, Yuan Hua, Wenbin Lai, Shihan Dou, Yuhao Zhou, Zhiheng Xi, Xiao Wang, Haoran Huang, Tao Gui, et al · 2023
Later among the works it cites.
Aligning language models to user opinions
EunJeong Hwang, Bodhisattwa Majumder, and Niket Tandon · 2023
Later among the works it cites.
Towards large language model-based personal agents in the enterprise: Current trends and open problems
Vinod Muthusamy, Yara Rizk, Kiran Kate, Praveen Venkateswaran, Vatche Isahagian, Ashu Gulati, and Parijat Dube · 2023
Later among the works it cites.
Beyond candidates: Adaptive dialogue agent utilizing persona and knowledge
Jungwoo Lim, Myunghoon Kang, Jinsung Kim, Jeongwook Kim, Yuna Hur, and Heui-Seok Lim · 2023
Later among the works it cites.
Large language models as source planner for personalized knowledge-grounded dialogues
WANG Hongru, Minda Hu, Yang Deng, Rui Wang, Fei Mi, Weichao Wang, Yasheng Wang, Wai-Chung Kwan, Irwin King, and Kam-Fai Wong · 2023
Later among the works it cites.
Do llms understand social knowledge? evaluating the sociability of large language models with socket benchmark
Minje Choi, Jiaxin Pei, Sagar Kumar, Chang Shu, and David Jurgens · 2023
Later among the works it cites.
You are what you annotate: Towards better models through annotator representations
Naihao Deng, Xinliang Zhang, Siyang Liu, Winston Wu, Lu Wang, and Rada Mihalcea · 2023
Later among the works it cites.
From instructions to intrinsic human values–a survey of alignment goals for big models
Jing Yao, Xiaoyuan Yi, Xiting Wang, Jindong Wang, and Xing Xie · 2023
Later among the works it cites.
Aligning large language models with human: A survey
Yufei Wang, Wanjun Zhong, Liangyou Li, Fei Mi, Xingshan Zeng, Wenyong Huang, Lifeng Shang, Xin Jiang, and Qun Liu · 2023
Later among the works it cites.
Training language models with language feedback at scale
Jérémy Scheurer, Jon Ander Campos, Tomasz Korbak, Jun Shern Chan, Angelica Chen, Kyunghyun Cho, and Ethan Perez · 2023
Later among the works it cites.
Xiaowei Huang, Wenjie Ruan, Wei Huang, Gaojie Jin, Yi Dong, Changshun Wu, Saddek Bensalem, Ronghui Mu, Yi Qi, Xingyu Zhao, et al · 2023
Later among the works it cites.
Interactive natural language processing
Zekun Wang, Ge Zhang, Kexin Yang, Ning Shi, Wangchunshu Zhou, Shaochun Hao, Guangzheng Xiong, Yizhi Li, Mong Yuan Sim, Xiuying Chen, et al · 2023
Later among the works it cites.
Do users write more insecure code with ai assistants?
Neil Perry, Megha Srivastava, Deepak Kumar, and Dan Boneh · 2023
Later among the works it cites.
Finspector: A human-centered visual inspection tool for exploring and comparing biases among foundation models
Bum Chul Kwon and Nandana Mihindukulasooriya · 2023
Later among the works it cites.
Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark
Alexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li, Steven Basart, Thomas Woodside, Hanlin Zhang, Scott Emmons, and Dan Hendrycks · 2023
Later among the works it cites.
AI alignment — Wikipedia, the free encyclopedia
Wikipedia · 2024
Closest in time.
Levels of agi: Operationalizing progress on the path to agi, 2024
Meredith Ringel Morris, Jascha Sohl-dickstein, Noah Fiedel, Tris Warkentin, Allan Dafoe, Aleksandra Faust, Clement Farabet, and Shane Legg · 2024
Closest in time.
Designing for human-agent alignment: Understanding what humans want from their agents
Nitesh Goyal, Minsuk Chang, and Michael Terry · 2024
Closest in time.
Ai alignment with changing and influenceable reward functions
Micah Carroll, Davis Foote, Anand Siththaranjan, Stuart Russell, and Anca Dragan · 2024
Closest in time.
The value, benefits, and concerns of generative ai-powered assistance in writing
Zhuoyan Li, Chen Liang, Jing Peng, and Ming Yin · 2024
Closest in time.
A design space for intelligent and interactive writing assistants
Mina Lee, Katy Ilonka Gero, John Joon Young Chung, Simon Buckingham Shum, Vipul Raheja, Hua Shen, Subhashini Venugopalan, Thiemo Wambsganss, David Zhou, Emad A Alghamdi, et al · 2024
Closest in time.
Can large language model agents simulate human trust behaviors?
Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye, Kai Shu, Adel Bibi, Ziniu Hu, Philip Torr, Bernard Ghanem, and Guohao Li · 2024
Closest in time.
Foundational challenges in assuring alignment and safety of large language models
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al · 2024
Closest in time.
What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning
Wei Liu, Weihao Zeng, Keqing He, Yong Jiang, and Junxian He · 2024
Closest in time.
A roadmap to pluralistic alignment
Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, et al · 2024
Closest in time.
The alignment problem from a deep learning perspective
Richard Ngo, Lawrence Chan, and Sören Mindermann · 2024
Closest in time.
Constitutionmaker: Interactively critiquing large language models by converting feedback into principles
Savvas Petridis, Benjamin D Wedin, James Wexler, Mahima Pushkarna, Aaron Donsbach, Nitesh Goyal, Carrie J Cai, and Michael Terry · 2024
Closest in time.
Safe rlhf: Safe reinforcement learning from human feedback
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang · 2024
Closest in time.
Social reward: Evaluating and enhancing generative AI through million-user feedback from an online creative community
Arman Isajanyan, Artur Shatveryan, David Kocharian, Zhangyang Wang, and Humphrey Shi · 2024
Closest in time.
Constitutionalexperts: Training a mixture of principle-based prompts
Savvas Petridis, Ben Wedin, Ann Yuan, James Wexler, and Nithum Thain · 2024
Closest in time.
Aligndiff: Aligning diverse human preferences via behavior-customisable diffusion model
Zibin Dong, Yifu Yuan, Jianye HAO, Fei Ni, Yao Mu, YAN ZHENG, Yujing Hu, Tangjie Lv, Changjie Fan, and Zhipeng Hu · 2024
Closest in time.
CRITIC: Large language models can self-correct with tool-interactive critiquing
Zhibin Gou, Zhihong Shao, Yeyun Gong, yelong shen, Yujiu Yang, Nan Duan, and Weizhu Chen · 2024
Closest in time.
Quality-diversity through ai feedback
Herbie Bradley, Andrew Dai, Hannah Teufel, Jenny Zhang, Koen Oostermeijer, Marco Bellagente, Jeff Clune, Kenneth Stanley, Grégory Schott, and Joel Lehman · 2024
Closest in time.
Towards diverse behaviors: A benchmark for imitation learning with human demonstrations
Xiaogang Jia, Denis Blessing, Xinkai Jiang, Moritz Reuss, Atalay Donat, Rudolf Lioutikov, and Gerhard Neumann · 2024
Closest in time.
Uni-rlhf: Universal platform and benchmark suite for reinforcement learning with diverse human feedback
Yifu Yuan, Jianye Hao, Yi Ma, Zibin Dong, Hebin Liang, Jinyi Liu, Zhixin Feng, Kai Zhao, and Yan Zheng · 2024
Closest in time.
How to teach programming in the ai era? using llms as a teachable agent for debugging
Qianou Ma, Hua Shen, Kenneth Koedinger, and Tongshuang Wu · 2024
Closest in time.
Rehearsal: Simulating conflict to teach conflict resolution
Omar Shaikh, Valentino Chai, Michele J Gelfand, Diyi Yang, and Michael S Bernstein · 2024
Closest in time.
Joshua Ashkinaze, Julia Mendelsohn, Li Qiwei, Ceren Budak, and Eric Gilbert · 2024
Closest in time.
Valuecompass: A framework of fundamental values for human-ai alignment
Hua Shen, Tiffany Knearem, Reshmi Ghosh, Yu-Ju Yang, Tanushree Mitra, and Yun Huang · 2024
Closest in time.
Lemur: Harmonizing natural language and code for language agents
Yiheng Xu, Hongjin Su, Chen Xing, Boyu Mi, Qian Liu, Weijia Shi, Binyuan Hui, Fan Zhou, Yitao Liu, Tianbao Xie, et al · 2024
Closest in time.
An ai-resilient text rendering technique for reading and skimming documents
Ziwei Gu, Ian Arawjo, Kenneth Li, Jonathan K Kummerfeld, and Elena L Glassman · 2024
Closest in time.
Human-centered privacy research in the age of large language models
Tianshi Li, Sauvik Das, Hao-Ping Lee, Dakuo Wang, Bingsheng Yao, and Zhiping Zhang · 2024
Closest in time.
The realhumaneval: Evaluating large language models’ abilities to support programmers
Hussein Mozannar, Valerie Chen, Mohammed Alsobay, Subhro Das, Sebastian Zhao, Dennis Wei, Manish Nagireddy, Prasanna Sattigeri, Ameet Talwalkar, and David Sontag · 2024
Closest in time.
Generative artificial intelligence, human creativity, and art
Eric Zhou and Dokyun Lee · 2024
Closest in time.
Enhancing human-AI collaboration through logic-guided reasoning
Chengzhi Cao, Yinghao Fu, Sheng Xu, Ruimao Zhang, and Shuang Li · 2024
Closest in time.
CPPO: Continual learning for reinforcement learning with human feedback
Han Zhang, Yu Lei, Lin Gui, Min Yang, Yulan He, Hui Wang, and Ruifeng Xu · 2024
Closest in time.
FLASK: Fine-grained language model evaluation based on alignment skill sets
Seonghyeon Ye, Doyoung Kim, Sungdong Kim, Hyeonbin Hwang, Seungone Kim, Yongrae Jo, James Thorne, Juho Kim, and Minjoon Seo · 2024
Closest in time.
Preference ranking optimization for human alignment
Feifan Song, Bowen Yu, Minghao Li, Haiyang Yu, Fei Huang, Yongbin Li, and Houfeng Wang · 2024
Closest in time.
Human feedback is not gold standard
Tom Hosking, Phil Blunsom, and Max Bartolo · 2024
Closest in time.
Motif: Intrinsic motivation from artificial intelligence feedback
Martin Klissarov, Pierluca D’Oro, Shagun Sodhani, Roberta Raileanu, Pierre-Luc Bacon, Pascal Vincent, Amy Zhang, and Mikael Henaff · 2024
Closest in time.
Enhancing human experience in human-agent collaboration: A human-centered modeling approach based on positive human gain
Yiming Gao, Feiyu Liu, Liang Wang, Dehua Zheng, Zhenjie Lian, Weixuan Wang, Wenjin Yang, Siqin Li, Xianliang Wang, Wenhui Chen, et al · 2024
Closest in time.
Beyond imitation: Leveraging fine-grained quality signals for alignment
Geyang Guo, Ranchi Zhao, Tianyi Tang, Wayne Xin Zhao, and Ji-Rong Wen · 2024
Closest in time.
Generative judge for evaluating alignment
Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu · 2024
Closest in time.
Tool-augmented reward modeling
Lei Li, Yekun Chai, Shuohuan Wang, Yu Sun, Hao Tian, Ningyu Zhang, and Hua Wu · 2024
Closest in time.
Fine-tuning language models for factuality
Katherine Tian, Eric Mitchell, Huaxiu Yao, Christopher D Manning, and Chelsea Finn · 2024
Closest in time.
Learning from free-text human feedback–collect new datasets or extend existing ones?
Dominic Petrak, Nafise Sadat Moosavi, Ye Tian, Nikolai Rozanov, and Iryna Gurevych · 2024
Closest in time.
The unlocking spell on base LLMs: Rethinking alignment via in-context learning
Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi · 2024
Closest in time.
Query-policy misalignment in preference-based reinforcement learning
Xiao Hu, Jianxiong Li, Xianyuan Zhan, Qing-Shan Jia, and Ya-Qin Zhang · 2024
Closest in time.
Can llms keep a secret? testing privacy implications of language models via contextual integrity theory
Niloofar Mireshghallah, Hyunwoo Kim, Xuhui Zhou, Yulia Tsvetkov, Maarten Sap, Reza Shokri, and Yejin Choi · 2024
Closest in time.
Investigating cultural alignment of large language models
Badr AlKhamissi, Muhammad ElNokrashy, Mai AlKhamissi, and Mona Diab · 2024
Closest in time.
Evallm: Interactive evaluation of large language model prompts on user-defined criteria
Tae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, and Juho Kim · 2024
Closest in time.
Learning from natural language feedback
Angelica Chen, Jérémy Scheurer, Jon Ander Campos, Tomasz Korbak, Jun Shern Chan, Samuel R. Bowman, Kyunghyun Cho, and Ethan Perez · 2024
Closest in time.
Explainability for large language models: A survey
Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du · 2024
Closest in time.
Mechanistic interpretability for ai safety – a review
Leonard Bereska and Efstratios Gavves · 2024
Closest in time.
Managing extreme ai risks amid rapid progress
Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, et al · 2024
Closest in time.
Acm policy on authorship, May 2024
The ACM Director of Publications · 2024
Closest in time.
Mind the value-action gap: Do llms act in alignment with their values?
Hua Shen, Nicholas Clark, and Tanushree Mitra · 2025
Closest in time.
Epistemic alignment: A mediating framework for user-llm knowledge delivery
Nicholas Clark, Hua Shen, Bill Howe, and Tanushree Mitra · 2025
Closest in time.
The siren song of llms: How users perceive and respond to dark patterns in large language models
Yike Shi, Qing Xiao, Hong Shen, Hua Shen, et al · 2025
Closest in time.