Fetching the paper…
Reading the bibliography…
This is a report on the NSF Future Directions Workshop on Automatic Evaluation of Dialog.
Acute-eval: Improved dialogue evaluation with optimized questions and multi-turn comparisons
Margaret Li, Jason Weston, and Stephen Roller · 1909
Earlier work this paper cites.
A model-theoretic coreference scoring scheme
Marc Vilain, John Burger, John Aberdeen, Dennis Connolly, and Lynette Hirschman · 1995
Earlier work this paper cites.
Paradise: A framework for evaluating spoken dialogue agents
Marilyn A Walker, Diane J Litman, Candace A Kamm, and Alicia Abella · 1997
Earlier work this paper cites.
Algorithms for scoring coreference chains
Amit Bagga and Breck Baldwin · 1998
Earlier work this paper cites.
Towards developing general models of usability with PARADISE
Marilyn Walker, Candace Kamm, and Diane Litman · 2000
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
George Doddington · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
On coreference resolution performance metrics
Xiaoqiang Luo · 2005
Earlier work this paper cites.
Quantitative evaluation of user simulation techniques for spoken dialogue systems
Jost Schatzmann, Kallirroi Georgila, and Steve Young · 2005
Earlier work this paper cites.
User simulation for spoken dialogue systems: Learning and evaluation
Kallirroi Georgila, James Henderson, and Oliver Lemon · 2006
Earlier work this paper cites.
Individual and domain adaptation in sentence planning for dialogue
Marilyn Walker, Amanda Stent, François Mairesse, and Rashmi Prasad · 2007
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Metrics and evaluation of spoken dialogue systems
Helen Hastie · 2012
Earlier work this paper cites.
RankME: Reliable human ratings for natural language generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser · 2012
Earlier work this paper cites.
On-line policy optimisation of bayesian spoken dialogue systems via human interaction
M. Gašić, C. Breslin, M. Henderson, D. Kim, M. Szummer, B. Thomson, P. Tsiakoulis, and S. Young · 2013
Earlier work this paper cites.
Cluster-based prediction of user ratings for stylistic surface realisation
Nina Dethlefs, Heriberto Cuayáhuitl, Helen Hastie, Verena Rieser, and Oliver Lemon · 2014
Earlier work this paper cites.
Expert-generated vs. crowd-sourced annotations for evaluating chatting sessions at the turn level
Rafael E Banchs · 2016
Earlier work this paper cites.
Learning end-to-end goal-oriented dialog
Antoine Bordes, Y-Lan Boureau, and Jason Weston · 2016
Earlier work this paper cites.
Learning, adaptive support, student traits, and engagement in scenario-based learning
Mark G. Core, Kallirroi Georgila, Benjamin D. Nye, Daniel Auerbach, Zhi Fei Liu, and Richard DiNinni · 2016
Earlier work this paper cites.
A sequence-to-sequence model for user simulation in spoken dialogue systems
Layla El Asri, Jing He, and Kaheer Suleman · 2016
Earlier work this paper cites.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau · 2016
Earlier work this paper cites.
On-line active reward learning for policy optimisation in spoken dialogue systems
Pei-Hao Su, Milica Gašić, Nikola Mrkšić, Lina M. Rojas-Barahona, Stefan Ultes, David Vandyke, Tsung-Hsien Wen, and Steve Young · 2016
Earlier work this paper cites.
The dialog state tracking challenge series: A review
Jason D Williams, Antoine Raux, and Matthew Henderson · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Towards an automatic Turing test: Learning to evaluate dialogue responses
Ryan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau · 2017
Earlier work this paper cites.
Why we should have seen that coming
K.W Miller, Marty J Wolf, and F.S. Grodzinsky · 2017
Earlier work this paper cites.
Why we need new evaluation metrics for nlg
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser · 2017
Earlier work this paper cites.
Patient and consumer safety risks when using conversational assistants for medical information: An observational study of Siri, Alexa, and Google Assistant
Timothy W Bickmore, Ha Trinh, Stefan Olafsson, Teresa K O’Leary, Reza Asadi, Nathaniel M Rickles, and Ricardo Cruz · 2018
Earlier work this paper cites.
Multiwoz–a large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Inigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić · 2018
Earlier work this paper cites.
Ruber: An unsupervised method for automatic evaluation of open-domain dialog systems
Chongyang Tao, Lili Mou, Dongyan Zhao, and Rui Yan · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman · 2018
Cited alongside, same era.
Sample efficient deep reinforcement learning for dialogue systems with large action spaces
Gellért Weisz, Paweł Budzianowski, Pei-Hao Su, and Milica Gašić · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
The second conversational intelligence challenge (convai2)
Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, et al · 2019
Cited alongside, same era.
Noise and neural natural language generationrubbish in, rubbish out?
Ondrej Dušek, David Howcroft, Karin Sevegnani, and Verena Rieser · 2019
Cited alongside, same era.
Deconstruct to reconstruct a configurable evaluation metric for open-domain dialogue systems
Vitou Phy, Yang Zhao, and Akiko Aizawa · 2020
Later among the works it cites.
Improving dialog evaluation with a multi-reference adversarial dataset and large scale pretraining
Ananya B. Sai, Akash Kumar Mohankumar, Siddhartha Arora, and Mitesh M. Khapra · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh · 2020
Later among the works it cites.
Learning an unreferenced metric for online dialogue evaluation
Koustuv Sinha, Prasanna Parthasarathi, Jasmine Wang, Ryan Lowe, William L. Hamilton, and Joelle Pineau · 2020
Later among the works it cites.
Multiwoz 2.2: A dialogue dataset with additional annotation corrections and state tracking baselines
Xiaoxue Zang, Abhinav Rastogi, Srinivas Sunkara, Raghav Gupta, Jianguo Zhang, and Jindong Chen · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beyond turing: Intelligent agents centered on the user
Maxine Eskenazi, Shikib Mehri, Evgeniia Razumovskaia, and Tiancheng Zhao · 2019
Cited alongside, same era.
Better automatic evaluation of open-domain dialogue systems with contextualized embeddings
Sarik Ghazarian, Johnny Wei, Aram Galstyan, and Nanyun Peng · 2019
Cited alongside, same era.
Investigating evaluation of open-domain dialogue systems with human generated multiple references
Prakhar Gupta, Shikib Mehri, Tiancheng Zhao, Amy Pavel, Maxine Eskenazi, and Jeffrey P Bigham · 2019
Cited alongside, same era.
Exploring social bias in chatbots using stereotype knowledge
Nayeon Lee, Andrea Madotto, and Pascale Fung · 2019
Cited alongside, same era.
Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM 2019) , Minneapolis, Minnesota, June 2019. Association for Computational Linguistics
Rada Mihalcea, Ekaterina Shutova, Lun-Wei Ku, Kilian Evang, and Soujanya Poria, editors · 2019
Cited alongside, same era.
Are training samples correlated? learning to generate dialogue responses with multiple references
Lisong Qiu, Juntao Li, Wei Bi, Dongyan Zhao, and Rui Yan · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Cited alongside, same era.
ConvAbuse: Data, analysis, and benchmarks for nuanced abuse detection in conversational AI
Amanda Cercas Curry, Gavin Abercrombie, and Verena Rieser · 2021
Later among the works it cites.
All that’s ‘human’ is not gold: Evaluating human evaluation of generated text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith · 2021
Later among the works it cites.
Survey on evaluation methods for dialogue systems
Jan Deriu, Alvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak · 2021
Later among the works it cites.
Anticipating safety issues in e2e conversational ai: Framework and tooling
Emily Dinan, Gavin Abercrombie, A Stevie Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser · 2021
Later among the works it cites.
The eval4nlp shared task on explainable quality estimation: Overview and results
Marina Fomicheva, Piyawat Lertvittayakumjorn, Wei Zhao, Steffen Eger, and Yang Gao · 2021
Later among the works it cites.
Experts, errors, and context: A large-scale study of human evaluation for machine translation, 2021
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey · 2021
Later among the works it cites.
What happens if you treat ordinal ratings as interval data? human evaluations in nlp are even more under-powered than you think
David M Howcroft and Verena Rieser · 2021
Later among the works it cites.
A survey of nlp-related crowdsourcing hits: what works and what does not
Jessica Huynh, Jeffrey Bigham, and Maxine Eskenazi · 2021
Later among the works it cites.
Addressing inquiries about history: An efficient and practical framework for evaluating open-domain chatbot consistency
Zekang Li, Jinchao Zhang, Zhengcong Fei, Yang Feng, and Jie Zhou · 2021
Later among the works it cites.
Conversations are not flat: Modeling the dynamic information flow across dialogue utterances
Zekang Li, Jinchao Zhang, Zhengcong Fei, Yang Feng, and Jie Zhou · 2021
Later among the works it cites.
Herald: An annotation efficient method to detect user disengagement in social conversations
Weixin Liang, Kai-Hui Liang, and Zhou Yu · 2021
Later among the works it cites.
Language model augmented relevance score
Ruibo Liu, Jason Wei, and Soroush Vosoughi · 2021
Later among the works it cites.
Better than average: Paired evaluation of NLP systems
Maxime Peyrard, Wei Zhao, Steffen Eger, and Robert West · 2021
Later among the works it cites.
Automatic evaluation of non-task oriented dialog systems by using sentence embeddings projections and their dynamics
Mario Rodríguez-Cantelar, Luis Fernando D’Haro, and Fernando Matía · 2021
Later among the works it cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston · 2021
Later among the works it cites.
Questeval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Patrick Gallinari, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, and Alex Wang · 2021
Later among the works it cites.
On the safety of conversational models: Taxonomy, dataset, and benchmark, 2021
Hao Sun, Guangxuan Xu, Jiawen Deng, Jiale Cheng, Chujie Zheng, Hao Zhou, Nanyun Peng, Xiaoyan Zhu, and Minlie Huang · 2021
Later among the works it cites.
Modeling performance in open-domain dialogue with paradise
Marilyn Walker, Colin Harmon, James Graupera, Davan Harrison, and Steve Whittaker · 2021
Later among the works it cites.
The First Workshop on Evaluations and Assessments of Neural Conversation Systems , Online, November 2021. Association for Computational Linguistics
Wei Wei, Bo Dai, Tuo Zhao, Lihong Li, Diyi Yang, Yun-Nung Chen, Y-Lan Boureau, Asli Celikyilmaz, Alborz Geramifard, Aman Ahuja, and Haoming Jiang, editors · 2021
Later among the works it cites.
Assessing dialogue systems with distribution distances
Jiannan Xiang, Yahui Liu, Deng Cai, Huayang Li, Defu Lian, and Lemao Liu · 2021
Later among the works it cites.
A comprehensive assessment of dialog evaluation metrics
Yi-Ting Yeh, Maxine Eskenazi, and Shikib Mehri · 2021
Later among the works it cites.
DynaEval: Unifying turn and dialogue level evaluation
Chen Zhang, Yiming Chen, Luis Fernando D’Haro, Yan Zhang, Thomas Friedrichs, Grandee Lee, and Haizhou Li · 2021
Later among the works it cites.
D-score: Holistic dialogue evaluation without reference
Chen Zhang, Grandee Lee, Luis Fernando D’Haro, and Haizhou Li · 2021
Later among the works it cites.
Probing the robustness of trained metrics for conversational dialogue systems
Jan Deriu, Don Tuggener, Pius von Daniken, and Mark Cieliebak · 2022
Closest in time.
SafetyKit
Emily Dinan, Gavin Abercrombie, A. Stevie Bergman, Shannon Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser · 2022
Closest in time.
Eric Michael Smith, Orion Hsu, Rebecca Qian, Stephen Roller, Y-Lan Boureau, and Jason Weston · 2022
Closest in time.