Fetching the paper…
Reading the bibliography…
Task-oriented conversational datasets often lack topic variability and linguistic diversity.
A coefficient of agreement for nominal scales
Jacob Cohen. 1960 · 1960
Earlier work this paper cites.
PARADISE: A framework for evaluating spoken dialogue agents
Marilyn A. Walker, Diane J. Litman, Candace A. Kamm, and Alicia Abella. 1997 · 1997
Earlier work this paper cites.
Basic emotions
Paul Ekman. 1999 · 1999
Earlier work this paper cites.
Investigation of governance mechanisms for crowdsourcing initiatives
Radhika Jain. 2010 · 2010
Earlier work this paper cites.
A parameterized and annotated spoken dialog corpus of the CMU let’s go bus information system
Alexander Schmitt, Stefan Ultes, and Wolfgang Minker. 2012 · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
The Ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems
Ryan Lowe, Nissan Pow, Iulian Serban, and Joelle Pineau. 2015 · 2015
Earlier work this paper cites.
Interaction quality: Assessing the quality of ongoing sporiaken dialog interaction by experts—and how it relates to user satisfaction
Alexander Schmitt and Stefan Ultes. 2015 · 2015
Earlier work this paper cites.
Classifying emotions in customer support dialogues in social media
Jonathan Herzig, Guy Feigenblat, Michal Shmueli-Scheuer, David Konopnicki, Anat Rafaeli, Daniel Altman, and David Spivak. 2016 · 2016
Earlier work this paper cites.
Frames: a corpus for adding memory to goal-oriented dialogue systems
Layla El Asri, Hannes Schulz, Shikhar Sharma, Jeremie Zumer, Justin Harris, Emery Fine, Rahul Mehrotra, and Kaheer Suleman. 2017 · 2017
Earlier work this paper cites.
DailyDialog: A manually labelled multi-turn dialogue dataset
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017 · 2017
Cited alongside, same era.
MultiWOZ - a large-scale multi-domain Wizard-of-Oz dataset for task-oriented dialogue modelling
Paweł Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gašić. 2018 · 2018
Cited alongside, same era.
Emotion recognition in conversation: Research challenges, datasets, and recent advances
Soujanya Poria, Navonil Majumder, Rada Mihalcea, and Eduard Hovy. 2019 · 2019
Cited alongside, same era.
What makes a good conversation? how controllable attributes affect human judgments
Abigail See, Stephen Roller, Douwe Kiela, and Jason Weston. 2019 · 2019
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Later among the works it cites.
A comprehensive assessment of dialog evaluation metrics
Yi-Ting Yeh, Maxine Eskenazi, and Shikib Mehri. 2021 · 2021
Later among the works it cites.
Automatic evaluation and moderation of open-domain dialogue systems
Chen Zhang, João Sedoc, L. F. D’Haro, Rafael E. Banchs, and Alexander I. Rudnicky. 2021 · 2021
Later among the works it cites.
Findings of the WMT 2022 shared task on chat translation
Ana C Farinha, M. Amin Farajian, Marianna Buchicchio, Patrick Fernandes, José G. C. de Souza, Helena Moniz, and André F. T. Martins. 2022 · 2022
Later among the works it cites.
A brief survey of textual dialogue corpora
Hugo Gonçalo Oliveira, Patrícia Ferreira, Daniel Martins, Catarina Silva, and Ana Alves. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unsupervised evaluation of interactive dialog with DialoGPT
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
Cited alongside, same era.
Deconstruct to reconstruct a configurable evaluation metric for open-domain dialogue systems
Vitou Phy, Yang Zhao, and Akiko Aizawa. 2020 · 2020
Cited alongside, same era.
Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset
Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan. 2020 · 2020
Cited alongside, same era.
Learning an unreferenced metric for online dialogue evaluation
Koustuv Sinha, Prasanna Parthasarathi, Jasmine Wang, Ryan Lowe, William L. Hamilton, and Joelle Pineau. 2020 · 2020
Cited alongside, same era.
TWEETSUMM - a dialog summarization dataset for customer service
Guy Feigenblat, Chulaka Gunasekara, Benjamin Sznajder, Sachindra Joshi, David Konopnicki, and Ranit Aharonov. 2021 · 2021
Cited alongside, same era.
Human evaluation of conversations is an open problem: comparing the sensitivity of various methods for evaluating dialogue agents
Eric Smith, Orion Hsu, Rebecca Qian, Stephen Roller, Y-Lan Boureau, and Jason Weston. 2022 · 2022
Later among the works it cites.
EnDex: Evaluation of dialogue engagingness at scale
Guangxuan Xu, Ruibo Liu, Fabrice Harel-Canada, Nischal Reddy Chandra, and Nanyun Peng. 2022 · 2022
Later among the works it cites.
Towards multilingual automatic open-domain dialogue evaluation
John Mendonça, Alon Lavie, and Isabel Trancoso. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.