Fetching the paper…
Reading the bibliography…
We present metrics for evaluating dialog systems through a psychologically-grounded "human" lens in which conversational agents express a diversity of both states (e.g., emotion) and traits (e.g., personality), just as people do.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
Acute-eval: Improved dialogue evaluation with optimized questions and multi-turn comparisons
Margaret Li, Jason Weston, and Stephen Roller. 2019 · 1909
Earlier work this paper cites.
The concept of traits
HA Carr and FA Kingsbury. 1938 · 1938
Earlier work this paper cites.
A general psychoevolutionary theory of emotion
Robert Plutchik. 1980 · 1980
Earlier work this paper cites.
Measuring individual differences in empathy: Evidence for a multidimensional approach
Mark H Davis. 1983 · 1983
Earlier work this paper cites.
Contexts of accommodation: Developments in applied sociolinguistics
Howard Ed Giles, Justine Ed Coupland, and Nikolas Ed Coupland. 1991 · 1991
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020 · 2001
Earlier work this paper cites.
Linguistic inquiry and word count: Liwc 2001
James W Pennebaker, Martha E Francis, and Roger J Booth. 2001 · 2001
Earlier work this paper cites.
Latent dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan. 2003 · 2003
Earlier work this paper cites.
Usr: An unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and Maxine Eskenazi. 2020b · 2005
Earlier work this paper cites.
An evaluation protocol for generative conversational systems
Seolhwa Lee, Heuiseok Lim, and João Sedoc. 2020 · 2010
Earlier work this paper cites.
Mark my words! linguistic style accommodation in social media
Cristian Danescu-Niculescu-Mizil, Michael Gamon, and Susan Dumais. 2011 · 2011
Earlier work this paper cites.
Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs
Cristian Danescu-Niculescu-Mizil and Lillian Lee. 2011 · 2011
Earlier work this paper cites.
Language style matching predicts relationship initiation and stability
Molly E. Ireland, Richard B. Slatcher, Paul W. Eastwick, Lauren E. Scissors, Eli J. Finkel, and James W. Pennebaker. 2011 · 2011
Earlier work this paper cites.
Echoes of power: Language effects and power differences in social interaction
Cristian Danescu-Niculescu-Mizil, Lillian Lee, Bo Pang, and Jon Kleinberg. 2012 · 2012
Earlier work this paper cites.
Private traits and attributes are predictable from digital records of human behavior
Michal Kosinski, David Stillwell, and Thore Graepel. 2013 · 2013
Earlier work this paper cites.
Convergence of speech rate in conversation predicts cooperation
Joseph H Manson, Gregory A Bryant, Matthew M Gervais, and Michelle A Kline. 2013 · 2013
Earlier work this paper cites.
Exploring demographic language variations to improve multilingual sentiment analysis in social media
Svitlana Volkova, Theresa Wilson, and David Yarowsky. 2013 · 2013
Earlier work this paper cites.
When to use the bonferroni correction
Richard A Armstrong. 2014 · 2014
Earlier work this paper cites.
Demographic factors improve classification performance
Dirk Hovy. 2015 · 2015
Earlier work this paper cites.
Tagging performance correlates with author age
Dirk Hovy and Anders Søgaard. 2015 · 2015
Earlier work this paper cites.
More than reflections: Empathy in motivational interviewing includes language style synchrony between therapist and client
Sarah Peregrine Lord, Elisa Sheng, Zac E Imel, John Baer, and David C Atkins. 2015 · 2015
Earlier work this paper cites.
Using hashtags to capture fine emotion categories from tweets
Saif M Mohammad and Svetlana Kiritchenko. 2015 · 2015
Earlier work this paper cites.
Automatic personality assessment through social media language
Gregory Park, H Andrew Schwartz, Johannes C Eichstaedt, Margaret L Kern, Michal Kosinski, David J Stillwell, Lyle H Ungar, and Martin EP Seligman. 2015 · 2015
Earlier work this paper cites.
Oriol Vinyals and Quoc Le. 2015 · 2015
Earlier work this paper cites.
A persona-based neural conversation model
Jiwei Li, Michel Galley, Chris Brockett, Georgios P Spithourakis, Jianfeng Gao, and Bill Dolan. 2016 · 2016
Earlier work this paper cites.
How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian Vlad Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Cited alongside, same era.
Recognizing pathogenic empathy in social media
Muhammad Abdul-Mageed, Anneke Buffone, Hao Peng, Salvatore Giorgi, Johannes C Eichstaedt, and Lyle H Ungar. 2017 · 2017
Cited alongside, same era.
Reading wikipedia to answer open-domain questions
Danqi Chen, Adam Fisch, Jason Weston, and Antoine Bordes. 2017 · 2017
Cited alongside, same era.
Superagent: A customer service chatbot for e-commerce websites
Lei Cui, Shaohan Huang, Furu Wei, Chuanqi Tan, Chaoqun Duan, and Ming Zhou. 2017 · 2017
Cited alongside, same era.
Frames: a corpus for adding memory to goal-oriented dialogue systems
Layla El Asri, Hannes Schulz, Shikhar Kr Sarma, Jeremie Zumer, Justin Harris, Emery Fine, Rahul Mehrotra, and Kaheer Suleman. 2017 · 2017
Cited alongside, same era.
Designing precise and robust dialogue response evaluators
Tianyu Zhao, Divesh Lala, and Tatsuya Kawahara. 2020 · 2020
Later among the works it cites.
A general language assistant as a laboratory for alignment
Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Ben Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan. 2021 · 2021
Later among the works it cites.
Automatic evaluation and moderation of open-domain dialogue systems
Zhang Chen, João Sedoc, Luis Fernando D’Haro, Rafael Banchs, and Alexander Rudnicky. 2021 · 2021
Later among the works it cites.
Survey on evaluation methods for dialogue systems
Jan Deriu, Alvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak. 2021 · 2021
Later among the works it cites.
Characterizing social spambots by their human traits
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chiori Hori and Takaaki Hori. 2017 · 2017
Cited alongside, same era.
Emotional dialogue generation using image-grounded language models
Bernd Huber, Daniel McDuff, Chris Brockett, Michel Galley, and Bill Dolan. 2018 · 2018
Cited alongside, same era.
Training millions of personalized dialogue agents
Pierre-Emmanuel Mazaré, Samuel Humeau, Martin Raison, and Antoine Bordes. 2018 · 2018
Cited alongside, same era.
Inducing rapport-building behaviors in interaction with an embodied conversational agent
David Novick, Mahdokht Afravi, Adriana Camacho, Laura J Hinojos, and Aaron E Rodriguez. 2018 · 2018
Cited alongside, same era.
Personalizing dialogue agents: I have a dog, do you have pets too?
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018 · 2018
Cited alongside, same era.
Grounded response generation task at dstc7
Michel Galley, Chris Brockett, Xiang Gao, Jianfeng Gao, and Bill Dolan. 2019 · 2019
Cited alongside, same era.
Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them
Hila Gonen and Yoav Goldberg. 2019 · 2019
Cited alongside, same era.
Salvatore Giorgi, Lyle Ungar, and H. Andrew Schwartz. 2021 · 2021
Later among the works it cites.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. 2021 · 2021
Later among the works it cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, et al. 2021 · 2021
Later among the works it cites.
Detoxifying language models risks marginalizing minority voices
Albert Xu, Eshaan Pathak, Eric Wallace, Suchin Gururangan, Maarten Sap, and Dan Klein. 2021 · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
Deep am-fm: Toolkit for automatic dialogue evaluation
Chen Zhang, Luis Fernando D’Haro, Rafael E Banchs, Thomas Friedrichs, and Haizhou Li. 2021 · 2021
Later among the works it cites.
Constitutional ai: Harmlessness from ai feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan. 2022 · 2022
Later among the works it cites.
What is wrong with you?: Leveraging user sentiment for automatic dialog evaluation
Sarik Ghazarian, Behnam Hedayatnia, Alexandros Papangelis, Yang Liu, and Dilek Hakkani-Tur. 2022 · 2022
Later among the works it cites.
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese, Nat McAleese, Maja Trębacz, John Aslanides, Vlad Firoiu, Timo Ewalds, Maribeth Rauh, Laura Weidinger, Martin Chadwick, Phoebe Thacker, Lucy Campbell-Gillingham, Jonathan Uesato, Po-Sen Huang, Ramona Comanescu, Fan Yang, Abigail See, Sumanth Dathathri, Rory Greig, Charlie Chen, Doug Fritz, Jaume Sanchez Elias, Richard Green, Soňa Mokrá, Nicholas Fernando, Boxi Wu, Rachel Foley, Susannah Young, Iason Gabriel, William Isaac, John Mellor, Demis Hassabis, Koray Kavukcuoglu, Lisa Anne Hendricks, and Geoffrey Irving. 2022 · 2022
Later among the works it cites.
Prosocialdialog: A prosocial backbone for conversational agents
Hyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu, Daniel Khashabi, Gunhee Kim, Yejin Choi, and Maarten Sap. 2022 · 2022
Later among the works it cites.
Empathic conversations: A multi-level dataset of contextualized conversations
Damilola Omitaomu, Shabnam Tafreshi, Tingting Liu, Sven Buechel, Chris Callison-Burch, Johannes Eichstaedt, Lyle Ungar, and João Sedoc. 2022 · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Later among the works it cites.
Blenderbot 3: a deployed conversational agent that continually learns to responsibly engage
Kurt Shuster, Jing Xu, Mojtaba Komeili, Da Ju, Eric Michael Smith, Stephen Roller, Megan Ung, Moya Chen, Kushal Arora, Joshua Lane, et al. 2022 · 2022
Later among the works it cites.
Human evaluation of conversations is an open problem: comparing the sensitivity of various methods for evaluating dialogue agents
Eric Smith, Orion Hsu, Rebecca Qian, Stephen Roller, Y-Lan Boureau, and Jason Weston. 2022 · 2022
Later among the works it cites.
The effects of partner extraversion and agreeableness on trust
Olga Stavrova, Anthony M Evans, and Ilja van Beest. 2022 · 2022
Later among the works it cites.
Mirages: On anthropomorphism in dialogue systems
Gavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, and Zeerak Talat. 2023 · 2023
Closest in time.
Using cognitive psychology to understand GPT-3
Marcel Binz and Eric Schulz. 2023 · 2023
Closest in time.
Theory of mind may have spontaneously emerged in large language models
Michal Kosinski. 2023 · 2023
Closest in time.
Whose opinions do language models reflect?
Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023 · 2023
Closest in time.
Artificial intelligence will change the future of psychotherapy: A proposal for responsible, psychologist-led development
Elizabeth Stade, Shannon Wiltsey Stirman, Lyle H Ungar, David Bryce Yaden, H Andrew Schwartz, João Sedoc, Robb Willer, Robert DeRubeis, et al. 2023 · 2023
Closest in time.
Characterizing empathy and compassion using computational linguistic analysis
David B. Yaden, Salvatore Giorgi, Matthew Jordan, Anneke Buffone, Johannes C. Eichstaedt, H. Andrew Schwartz, Lyle H. Ungar, and Paul Bloom. 2023 · 2023
Closest in time.
MojiTalk: Generating emotional responses at scale
Xianda Zhou and William Yang Wang. 2018 · 2023
Closest in time.