Fetching the paper…
Reading the bibliography…
Evaluation of open-domain dialogue systems is highly challenging and development of better techniques is highlighted time and again as desperately needed.
The second conversational intelligence challenge (convai2)
Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander H. Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, Shrimai Prabhumoye, Alan W. Black, Alexander I. Rudnicky, Jason Williams, Joelle Pineau, Mikhail Burtsev, and Jason Weston. 2019 · 1902
Earlier work this paper cites.
Samuel Humeau, Kurt Shuster, Marie-Anne Lachaux, and Jason Weston. 2019 · 1905
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Acute-eval: Improved dialogue evaluation with optimized questions and multi-turn comparisons
Margaret Li, Jason Weston, and Stephen Roller. 2019 · 1909
Earlier work this paper cites.
A note on averaging correlations
Ralph A Alexander. 1990 · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Bleu: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Automatic evaluation of summaries using n-gram co-occurrence statistics
Chin-Yew Lin and Eduard Hovy. 2003 · 2003
Earlier work this paper cites.
Characteristics of single-item measures in likert scale format
Aliosha Alexandrov. 2010 · 2010
Earlier work this paper cites.
Findings of the 2011 workshop on statistical machine translation
Chris Callison-Burch, Philipp Koehn, Christof Monz, and Omar Zaidan. 2011 · 2011
Earlier work this paper cites.
Meteor 1.3: Automatic metric for reliable optimization and evaluation of machine translation systems
M. Denkowski and A. Lavie. 2011 · 2011
Earlier work this paper cites.
Findings of the 2012 workshop on statistical machine translation
Chris Callison-Burch, Philipp Koehn, Christof Monz, Matt Post, Radu Soricut, and Lucia Specia. 2012 · 2012
Cited alongside, same era.
Findings of the 2013 Workshop on Statistical Machine Translation
Ondřej Bojar, Christian Buck, Chris Callison-Burch, Christian Federmann, Barry Haddow, Philipp Koehn, Christof Monz, Matt Post, Radu Soricut, and Lucia Specia. 2013 · 2013
Cited alongside, same era.
Crowd-sourcing of human judgments of machine translation fluency
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Cited alongside, same era.
Scoring workers in crowdsourcing: How many control questions are enough?
Qiang Liu, Alexander T Ihler, and Mark Steyvers. 2013 · 2013
Cited alongside, same era.
Findings of the 2014 workshop on statistical machine translation
Ondrej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna. 2014 · 2014
Cited alongside, same era.
RankME: Reliable human ratings for natural language generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2018 · 2018
Later among the works it cites.
Conversational AI: the science behind the alexa prize
Ashwin Ram, Rohit Prasad, Chandra Khatri, Anu Venkatesh, Raefer Gabriel, Qing Liu, Jeff Nunn, Behnam Hedayatnia, Ming Cheng, Ashish Nagar, Eric King, Kate Bland, Amanda Wartick, Yi Pan, Han Song, Sk Jayadevan, Gene Hwang, and Art Pettigrue. 2018 · 2018
Later among the works it cites.
Towards best experiment design for evaluating dialogue system output
Sashank Santhanam and Samira Shaikh. 2019 · 2019
Later among the works it cites.
Findings of the 2020 conference on machine translation (wmt20)
Loïc Barrault, Magdalena Biesialska, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Matthias Huck, Eric Joanis, Tom Kocmi, Philipp Koehn, Chi-kiu Lo, Nikola Ljubešić, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Santanu Pal, Matt Post, and Marcos Zampieri. 2020 · 2020
Later among the works it cites.
Towards unified dialogue system evaluation: A comprehensive analysis of current evaluation protocols
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Information extraction and manipulation threats in crowd-powered systems
Walter S. Lasecki, Jaime Teevan, and Ece Kamar. 2014 · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Cited alongside, same era.
Key-value memory networks for directly reading documents
Alexander H. Miller, Adam Fisch, Jesse Dodge, Amir-Hossein Karimi, Antoine Bordes, and Jason Weston. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Cited alongside, same era.
Multimodal topic labelling
Ionut Sorodoc, Jey Han Lau, Nikolaos Aletras, and Timothy Baldwin. 2017 · 2017
Cited alongside, same era.
Wizard of wikipedia: Knowledge-powered conversational agents
Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2018 · 2018
Cited alongside, same era.
Unsupervised evaluation of interactive dialog with DialoGPT
Shikib Mehri and Maxine Eskenazi. 2020a
Cited in the paper.
Sarah E. Finch and Jinho D. Choi. 2020 · 2020
Later among the works it cites.
Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020 · 2020
Later among the works it cites.
Using PRMSE to evaluate automated scoring systems in the presence of label noise
Anastassia Loukina, Nitin Madnani, Aoife Cahill, Lili Yao, Matthew S. Johnson, Brian Riordan, and Daniel F. McCaffrey. 2020 · 2020
Later among the works it cites.
The third multilingual surface realisation shared task (SR’20): Overview and evaluation results
Simon Mille, Anya Belz, Bernd Bohnet, Thiago Castro Ferreira, Yvette Graham, and Leo Wanner. 2020 · 2020
Later among the works it cites.
Towards holistic and automatic evaluation of open-domain dialogue generation
Bo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou, Yixian Liu, and Kewei Tu. 2020 · 2020
Later among the works it cites.
Studying the effects of cognitive biases in evaluation of conversational agents
Sashank Santhanam, Alireza Karduni, and Samira Shaikh. 2020 · 2020
Later among the works it cites.
DIALOGPT : Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020 · 2020
Later among the works it cites.