Fetching the paper…
Reading the bibliography…
While often assumed a gold standard, effective human evaluation of text generation remains an important, open area for research.
Maximum likelihood from incomplete data via the em algorithm
Arthur P Dempster, Nan M Laird, and Donald B Rubin. 1977 · 1977
Earlier work this paper cites.
Conjugate priors for exponential families
Persi Diaconis and Donald Ylvisaker. 1979 · 1979
Earlier work this paper cites.
Generating summaries of multiple news articles
Kathleen McKeown and Dragomir R Radev. 1995 · 1995
Earlier work this paper cites.
A better bound on the variance
Rajendra Bhatia and Chandler Davis. 2000 · 2000
Earlier work this paper cites.
Estimating a dirichlet distribution
Thomas P. Minka. 2000 · 2000
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram cooccurrence statistics
George Doddington. 2002 · 2002
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Correlating automated and human assessments of machine translation quality
Deborah Coughlin. 2003 · 2003
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Evaluating content selection in summarization: The pyramid method
Ani Nenkova and Rebecca Passonneau. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Re-evaluating the role of Bleu in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
An information-theoretic approach to automatic evaluation of summaries
Chin-Yew Lin, Guihong Cao, Jianfeng Gao, and Jian-Yun Nie. 2006 · 2006
Earlier work this paper cites.
Budget-optimal task allocation for reliable crowdsourcing systems
David R. Karger, Sewoong Oh, and Devavrat Shah. 2011 · 2011
Earlier work this paper cites.
Machine Learning: A Probabilistic Perspective
Kevin P. Murphy. 2012 · 2012
Earlier work this paper cites.
Continuous measurement scales in human evaluation of machine translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Earlier work this paper cites.
Learning whom to trust with mace
Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, and Eduard Hovy. 2013 · 2013
Earlier work this paper cites.
Is machine translation getting better over time?
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2014 · 2014
Earlier work this paper cites.
To re (label), or not to re (label)
Christopher H Lin, M Mausam, and Daniel S Weld. 2014 · 2014
Earlier work this paper cites.
The benefits of a model of annotation
Rebecca J Passonneau and Bob Carpenter. 2014 · 2014
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Findings of the 2016 conference on machine translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016 · 2016
Earlier work this paper cites.
Effective crowd annotation for relation extraction
Angli Liu, Stephen Soderland, Jonathan Bragg, Christopher H. Lin, Xiao Ling, and Daniel S. Weld. 2016 · 2016
Cited alongside, same era.
Findings of the 2018 conference on machine translation (WMT18)
Ondřej Bojar, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Philipp Koehn, and Christof Monz. 2018 · 2018
Cited alongside, same era.
Sprout: Crowd-powered task design for crowdsourcing
Jonathan Bragg, Mausam, and Daniel S. Weld. 2018 · 2018
Cited alongside, same era.
The price of debiasing automatic metrics in natural language evalaution
Arun Tejasvi Chaganty, Stephen Mussmann, and Percy Liang. 2018 · 2018
Cited alongside, same era.
Results of the WMT18 metrics shared task: Both characters and embeddings achieve good performance
Qingsong Ma, Ondřej Bojar, and Yvette Graham. 2018 · 2018
Cited alongside, same era.
HYPE: A benchmark for human eye perceptual evaluation of generative models
Sharon Zhou, Mitchell Gordon, Ranjay Krishna, Austin Narcomey, Li F Fei-Fei, and Michael Bernstein. 2019 · 2019
Later among the works it cites.
STORIUM: A Dataset and Evaluation Platform for Machine-in-the-Loop Story Generation
Nader Akoury, Shufan Wang, Josh Whiting, Stephen Hood, Nanyun Peng, and Mohit Iyyer. 2020 · 2020
Later among the works it cites.
Findings of the 2020 conference on machine translation (WMT20)
Loïc Barrault, Magdalena Biesialska, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Matthias Huck, Eric Joanis, Tom Kocmi, Philipp Koehn, Chi-kiu Lo, Nikola Ljubešić, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Santanu Pal, Matt Post, and Marcos Zampieri. 2020 · 2020
Later among the works it cites.
Abductive commonsense reasoning
Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen tau Yih, and Yejin Choi. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shashi Narayan, Shay B Cohen, and Mirella Lapata. 2018 · 2018
Cited alongside, same era.
Comparing bayesian models of annotation
Silviu Paun, Bob Carpenter, Jon Chamberlain, Dirk Hovy, Udo Kruschwitz, and Massimo Poesio. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
Efficient online scalar annotation with bounded support
Keisuke Sakaguchi and Benjamin Van Durme. 2018 · 2018
Cited alongside, same era.
Attaining the unattainable? reassessing claims of human parity in neural machine translation
Antonio Toral, Sheila Castilho, Ke Hu, and Andy Way. 2018 · 2018
Cited alongside, same era.
Findings of the 2019 conference on machine translation (WMT19)
Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019 · 2019
Cited alongside, same era.
Translationese in machine translation evaluation
Yvette Graham, Barry Haddow, and Philipp Koehn. 2019 · 2019
Cited alongside, same era.
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2020
Later among the works it cites.
MOCHA: A dataset for training and evaluating generative reading comprehension metrics
Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2020 · 2020
Later among the works it cites.
On the evaluation of machine translation systems trained with back-translation
Sergey Edunov, Myle Ott, Marc’Aurelio Ranzato, and Michael Auli. 2020 · 2020
Later among the works it cites.
UnifiedQA: Crossing format boundaries with a single QA system
Daniel Khashabi, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, and Hannaneh Hajishirzi. 2020 · 2020
Later among the works it cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Lianhui Qin, Vered Shwartz, Peter West, Chandra Bhagavatula, Jena Hwang, Ronan Le Bras, Antoine Bosselut, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P Parikh. 2020 · 2020
Later among the works it cites.
Findings of the 2021 conference on machine translation (WMT21)
Farhad Akhbardeh, Arkady Arkhangorodsky, Magdalena Biesialska, Ondřej Bojar, Rajen Chatterjee, Vishrav Chaudhary, Marta R. Costa-jussa, Cristina España-Bonet, Angela Fan, Christian Federmann, Markus Freitag, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Leonie Harter, Kenneth Heafield, Christopher Homan, Matthias Huck, Kwabena Amponsah-Kaakyire, Jungo Kasai, Daniel Khashabi, Kevin Knight, Tom Kocmi, Philipp Koehn, Nicholas Lourie, Christof Monz, Makoto Morishita, Masaaki Nagata, Ajay Nagesh, Toshiaki Nakazawa, Matteo Negri, Santanu Pal, Allahsera Auguste Tapo, Marco Turchi, Valentin Vydrin, and Marcos Zampieri. 2021 · 2021
Closest in time.
Sumithra Bhakthavatsalam, Daniel Khashabi, Tushar Khot, Bhavana Dalvi Mishra, Kyle Richardson, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord, and Peter Clark. 2021 · 2021
Closest in time.
All that’s ‘human’is not gold: Evaluating human evaluation of generated text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A Smith. 2021 · 2021
Closest in time.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Closest in time.
Experts, errors, and context: A large-scale study of human evaluation for machine translation
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021 · 2021
Closest in time.
The GEM benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin P. Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Aremu Anuoluwapo, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh D. Dhole, Wanyu Du, Esin Durmus, Ondrej Dusek, Chris Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Rubungo Andre Niyongabo, Salomey Osei, Ankur P. Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021 · 2021
Closest in time.
The perils of using mechanical turk to evaluate open-ended text generation
Marzena Karpinska, Nader Akoury, and Mohit Iyyer. 2021 · 2021
Closest in time.
Bidimensional leaderboards: Generate and evaluate language hand in hand
Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander R. Fabbri, Yejin Choi, and Noah A. Smith. 2021 · 2021
Closest in time.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021 · 2021
Closest in time.
Turingadvice: A generative and dynamic evaluation of language use
Rowan Zellers, Ari Holtzman, Elizabeth Clark, Lianhui Qin, Ali Farhadi, and Yejin Choi. 2021 · 2021
Closest in time.