Fetching the paper…
Reading the bibliography…
Evaluation practices in natural language generation (NLG) have many known flaws, but improved evaluation approaches are rarely widely adopted.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 1919
Earlier work this paper cites.
Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel
J Peter Kincaid, Robert P Fishburne Jr, Richard L Rogers, and Brad S Chissom. 1975 · 1975
Earlier work this paper cites.
Evaluating natural language processing systems: An analysis and review , volume 1083
Karen Sparck Jones and Julia R Galliers. 1995 · 1995
Earlier work this paper cites.
Building applied natural language generation systems
Ehud Reiter and Robert Dale. 1997 · 1997
Earlier work this paper cites.
Towards evaluation in natural language generation
Robert Dale and Chris Mellish. 1998 · 1998
Earlier work this paper cites.
The evolution of evaluation: Lessons from the message understanding conferences
Lynette Hirschman. 1998 · 1998
Earlier work this paper cites.
Generation that exploits corpus-based statistical knowledge
Irene Langkilde and Kevin Knight. 1998 · 1998
Earlier work this paper cites.
The TIPSTER SUMMAC text summarization evaluation
Inderjeet Mani, David House, Gary Klein, Lynette Hirschman, Therese Firmin, and Beth Sundheim. 1999 · 1999
Earlier work this paper cites.
Evaluation metrics for generation
Srinivas Bangalore, Owen Rambow, and Steve Whittaker. 2000 · 2000
Earlier work this paper cites.
Comparison of the quality of assessments using continuous and discrete ordinal rating scales
Elisabeth Svensson. 2000 · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Should corpora texts be gold standards for NLG?
Ehud Reiter and Somayajulu Sripada. 2002 · 2002
Earlier work this paper cites.
The effects of human variation in DUC summarization evaluation
Donna Harman and Paul Over. 2004 · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
ORANGE: a method for evaluating automatic evaluation metrics for machine translation
Chin-Yew Lin and Franz Josef Och. 2004 · 2004
Earlier work this paper cites.
Evaluating content selection in summarization: The pyramid method
Ani Nenkova and Rebecca Passonneau. 2004 · 2004
Earlier work this paper cites.
An introduction to duc-2004
Paul Over and James Yen. 2004 · 2004
Earlier work this paper cites.
Bias decreases in proportion to the number of annotators
Ron Artstein and Massimo Poesio. 2009 · 2005
Earlier work this paper cites.
A catalog of biases in questionnaires
Bernard CK Choi and Anita WP Pak. 2005 · 2005
Earlier work this paper cites.
DUC 2005: Evaluation of question-focused summarization systems
Hoa Trang Dang. 2006 · 2005
Earlier work this paper cites.
A methodology for extrinsic evaluation of text summarization: Does ROUGE correlate?
Bonnie Dorr, Christof Monz, Stacy President, Richard Schwartz, and David Zajic. 2005 · 2005
Earlier work this paper cites.
Evaluating duc 2005 using basic elements
Eduard Hovy, Chin-Yew Lin, and Liang Zhou. 2005 · 2005
Earlier work this paper cites.
Automatic text summarization of newswire: Lessons learned from the document understanding conference
Ani Nenkova. 2005 · 2005
Earlier work this paper cites.
Evaluating evaluation methods for generation in the presence of variation
Amanda Stent, Matthew Marge, and Mohit Singhai. 2005 · 2005
Earlier work this paper cites.
Re-evaluating the role of Bleu in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
Statistical comparisons of classifiers over multiple data sets
Janez Demsar. 2006 · 2006
Earlier work this paper cites.
Manual and automatic evaluation of machine translation between European languages
Philipp Koehn and Christof Monz. 2006 · 2006
Earlier work this paper cites.
An nlg evaluation competition? eight reasons to be cautious
Donia Scott and Johanna Moore. 2007 · 2007
Earlier work this paper cites.
Intrinsic vs. extrinsic evaluation measures for referring expression generation
Anja Belz and Albert Gatt. 2008 · 2008
Earlier work this paper cites.
The TUNA challenge 2008: Overview and evaluation results
Albert Gatt, Anja Belz, and Eric Kow. 2008 · 2008
Earlier work this paper cites.
Get another label? improving data quality and data mining using multiple, noisy labelers
Victor S. Sheng, Foster J. Provost, and Panagiotis G. Ipeirotis. 2008 · 2008
Earlier work this paper cites.
Cheap and fast – but is it good? evaluating non-expert annotations for natural language tasks
Rion Snow, Brendan O’Connor, Daniel Jurafsky, and Andrew Ng. 2008 · 2008
Earlier work this paper cites.
Linguistically naïve != language independent: Why NLP needs linguistic typology
Emily M. Bender. 2009 · 2009
Earlier work this paper cites.
Dataset shift in machine learning
Joaquin Quiñonero-Candela, Masashi Sugiyama, Neil D Lawrence, and Anton Schwaighofer. 2009 · 2009
Earlier work this paper cites.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano. 2020 · 2009
Earlier work this paper cites.
Comparing rating scales and preference judgements in language evaluation
Anja Belz and Eric Kow. 2010 · 2010
Earlier work this paper cites.
Non-expert evaluation of summarization systems is risky
Dan Gillick and Yang Liu. 2010 · 2010
Earlier work this paper cites.
Automatic evaluation of linguistic quality in multi-document summarization
Emily Pitler, Annie Louis, and Ani Nenkova. 2010 · 2010
Earlier work this paper cites.
Multilingual summarization evaluation without human models
Horacio Saggion, Juan-Manuel Torres-Moreno, Iria da Cunha, Eric SanJuan, and Patricia Velázquez-Morales. 2010 · 2010
Earlier work this paper cites.
Discrete vs. continuous rating scales for language evaluation in NLP
Anja Belz and Eric Kow. 2011 · 2011
Earlier work this paper cites.
Amazon Mechanical Turk: Gold Mine or Coal Mine?
Karën Fort, Gilles Adda, and K. Bretonnel Cohen. 2011 · 2011
Earlier work this paper cites.
The handbook of critical intercultural communication
Thomas K Nakayama and Rona Tamiko Halualani. 2011 · 2011
Earlier work this paper cites.
Towards automatic error analysis of machine translation output
Maja Popovic and Hermann Ney. 2011 · 2011
Earlier work this paper cites.
Annotated Gigaword
Courtney Napoles, Matthew Gormley, and Benjamin Van Durme. 2012 · 2012
Earlier work this paper cites.
Assessing the effect of inconsistent assessors on summarization evaluation
Karolina Owczarzak, Peter A. Rankel, Hoa Trang Dang, and John M. Conroy. 2012 · 2012
Earlier work this paper cites.
Data and its (dis)contents: A survey of dataset development and use in machine learning research
Amandalynne Paullada, Inioluwa Deborah Raji, Emily M. Bender, Emily Denton, and Alex Hanna. 2020 · 2012
Earlier work this paper cites.
Better metrics to automatically predict the quality of a text summary
Peter A. Rankel, John M. Conroy, and Judith D. Schlesinger. 2012 · 2012
Earlier work this paper cites.
A decade of automatic content evaluation of news summaries: Reassessing the state of the art
Peter A. Rankel, John M. Conroy, Hoa Trang Dang, and Ani Nenkova. 2013 · 2013
Earlier work this paper cites.
The good, the bad and the ugly: Why crowdsourcing needs ethics
Florian Alexander Schmidt. 2013 · 2013
Earlier work this paper cites.
A snapshot of NLG evaluation practices 2005 - 2014
Dimitra Gkatzia and Saad Mahamood. 2015 · 2014
Earlier work this paper cites.
Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics
Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. 2014 · 2014
Earlier work this paper cites.
Results of the WMT14 metrics shared task
Matouš Macháček and Ondřej Bojar. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Squibs: Evaluating human pairwise preference judgments
Mark Dras. 2015 · 2015
Earlier work this paper cites.
MultiLing 2015: Multilingual summarization of single and multi-documents, on-line fora, and call-center conversations
George Giannakopoulos, Jeff Kubina, John Conroy, Josef Steinberger, Benoit Favre, Mijail Kabadjov, Udo Kruschwitz, and Massimo Poesio. 2015 · 2015
Earlier work this paper cites.
Re-evaluating automatic summarization with BLEU and 192 shades of ROUGE
Yvette Graham. 2015 · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomás Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Fatal or not? finding errors that lead to dialogue breakdowns in chat-oriented dialogue systems
Ryuichiro Higashinaka, Masahiro Mizukami, Kotaro Funakoshi, Masahiro Araki, Hiroshi Tsukahara, and Yuka Kobayashi. 2015 · 2015
Earlier work this paper cites.
Challenges of studying and processing dialects in social media
Anna Jørgensen, Dirk Hovy, and Anders Søgaard. 2015 · 2015
Earlier work this paper cites.
From word embeddings to document distances
Matt J. Kusner, Yu Sun, Nicholas I. Kolkin, and Kilian Q. Weinberger. 2015 · 2015
Earlier work this paper cites.
A neural attention model for abstractive sentence summarization
Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Earlier work this paper cites.
Results of the WMT15 metrics shared task
Miloš Stanojević, Amir Kamran, Philipp Koehn, and Ondřej Bojar. 2015 · 2015
Earlier work this paper cites.
Assessing relative sentence complexity using an incremental CCG parser
Bharat Ram Ambati, Siva Reddy, and Mark Steedman. 2016 · 2016
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of African-American English
Su Lin Blodgett, Lisa Green, and Brendan O’Connor. 2016 · 2016
Earlier work this paper cites.
Results of the WMT16 metrics shared task
Ondřej Bojar, Yvette Graham, Amir Kamran, and Miloš Stanojević. 2016 · 2016
Earlier work this paper cites.
Learning a POS tagger for AAVE-like language
Anna Jørgensen, Dirk Hovy, and Anders Søgaard. 2016 · 2016
Earlier work this paper cites.
Neural text generation from structured data with application to the biography domain
Rémi Lebret, David Grangier, and Michael Auli. 2016 · 2016
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence RNNs and beyond
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Çağlar Guçlçehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
Results of the WMT17 metrics shared task
Ondřej Bojar, Yvette Graham, and Amir Kamran. 2017 · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation
Svetlana Kiritchenko and Saif Mohammad. 2017 · 2017
Earlier work this paper cites.
Summarunner: A recurrent neural network based sequence model for extractive summarization of documents
Ramesh Nallapati, Feifei Zhai, and Bowen Zhou. 2017 · 2017
Earlier work this paper cites.
Why we need new evaluation metrics for NLG
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017a · 2017
Earlier work this paper cites.
Unsettling race and language: Toward a raciolinguistic perspective
Jonathan Rosa and Nelson Flores. 2017 · 2017
Earlier work this paper cites.
Challenges in data-to-document generation
Sam Wiseman, Stuart Shieber, and Alexander Rush. 2017 · 2017
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman. 2018 · 2018
Earlier work this paper cites.
Multiwoz - A large-scale multi-domain wizard-of-oz dataset for task-oriented dialogue modelling
Pawel Budzianowski, Tsung-Hsien Wen, Bo-Hsiang Tseng, Iñigo Casanueva, Stefan Ultes, Osman Ramadan, and Milica Gasic. 2018 · 2018
Earlier work this paper cites.
A semantic qa-based approach for text summarization evaluation
Ping Chen, Fei Wu, Tong Wang, and Wei Ding. 2018 · 2018
Cited alongside, same era.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Cited alongside, same era.
Findings of the E2E NLG challenge
Ondřej Dušek, Jekaterina Novikova, and Verena Rieser. 2018 · 2018
Cited alongside, same era.
Survey of the state of the art in natural language generation: Core tasks, applications and evaluation
Albert Gatt and Emiel Krahmer. 2018 · 2018
Cited alongside, same era.
Bottom-up abstractive summarization
Sebastian Gehrmann, Yuntian Deng, and Alexander Rush. 2018 · 2018
Cited alongside, same era.
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Max Grusky, Mor Naaman, and Yoav Artzi. 2018 · 2018
A gold standard methodology for evaluating accuracy in data-to-text systems
Craig Thomson and Ehud Reiter. 2020 · 2020
Later among the works it cites.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Aspect-controllable opinion summarization
Reinald Kim Amplayo, Stefanos Angelidis, and Mirella Lapata. 2021 · 2021
Later among the works it cites.
Focus attention: Promoting faithfulness and diversity in summarization
Rahul Aralikatte, Shashi Narayan, Joshua Maynez, Sascha Rothe, and Ryan McDonald. 2021 · 2021
Later among the works it cites.
The ReproGen shared task on reproducibility of human evaluations in NLG: Overview and results
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Content selection in deep learning models of summarization
Chris Kedzie, Kathleen McKeown, and Hal Daumé III. 2018 · 2018
Cited alongside, same era.
Results of the WMT18 metrics shared task: Both characters and embeddings achieve good performance
Qingsong Ma, Ondřej Bojar, and Yvette Graham. 2018 · 2018
Cited alongside, same era.
Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Cited alongside, same era.
RankME: Reliable human ratings for natural language generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2018 · 2018
Cited alongside, same era.
Multi-reward reinforced summarization with saliency and entailment
Ramakanth Pasunuru and Mohit Bansal. 2018 · 2018
Cited alongside, same era.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher. 2018 · 2018
Cited alongside, same era.
Anya Belz, Anastasia Shimorina, Shubham Agarwal, and Ehud Reiter. 2021 · 2021
Later among the works it cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Later among the works it cites.
What will it take to fix benchmarking in natural language understanding?
Samuel R. Bowman and George Dahl. 2021 · 2021
Later among the works it cites.
Evaluating the evaluation metrics for style transfer: A case study in multilingual formality transfer
Eleftheria Briakou, Sweta Agrawal, Joel R. Tetreault, and Marine Carpuat. 2021 · 2021
Later among the works it cites.
All that’s ‘human’ is not gold: Evaluating human evaluation of generated text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021 · 2021
Later among the works it cites.
Automatic text evaluation through the lens of wasserstein barycenters
Pierre Colombo, Guillaume Staerman, Chloé Clavel, and Pablo Piantanida. 2021 · 2021
Later among the works it cites.
Compression, transduction, and creation: A unified framework for evaluating natural language generation
Mingkai Deng, Bowen Tan, Zhengzhong Liu, Eric P. Xing, and Zhiting Hu. 2021 · 2021
Later among the works it cites.
Understanding the extent to which content quality metrics measure the information quality of summaries
Daniel Deutsch and Dan Roth. 2021 · 2021
Later among the works it cites.
MSˆ2: Multi-document summarization of medical studies
Jay DeYoung, Iz Beltagy, Madeleine van Zuylen, Bailey Kuehl, and Lucy Wang. 2021 · 2021
Later among the works it cites.
NL-Augmenter: A framework for task-sensitive natural language augmentation
Kaustubh D. Dhole, Varun Gangal, Sebastian Gehrmann, Aadesh Gupta, Zhenhao Li, Saad Mahamood, Abinaya Mahendiran, Simon Mille, Ashish Srivastava, Samson Tan, Tongshuang Wu, Jascha Sohl-Dickstein, Jinho D. Choi, Eduard Hovy, Ondrej Dusek, Sebastian Ruder, Sajant Anand, Nagender Aneja, Rabin Banjade, Lisa Barthe, Hanna Behnke, Ian Berlot-Attwell, Connor Boyle, Caroline Brun, Marco Antonio Sobrevilla Cabezudo, Samuel Cahyawijaya, Emile Chapuis, Wanxiang Che, Mukund Choudhary, Christian Clauss, Pierre Colombo, Filip Cornell, Gautier Dagan, Mayukh Das, Tanay Dixit, Thomas Dopierre, Paul-Alexis Dray, Suchitra Dubey, Tatiana Ekeinhor, Marco Di Giovanni, Rishabh Gupta, Rishabh Gupta, Louanes Hamla, Sang Han, Fabrice Harel-Canada, Antoine Honore, Ishan Jindal, Przemyslaw K. Joniak, Denis Kleyko, Venelin Kovatchev, Kalpesh Krishna, Ashutosh Kumar, Stefan Langer, Seungjae Ryan Lee, Corey James Levinson, Hualou Liang, Kaizhao Liang, Zhexiong Liu, Andrey Lukyanenko, Vukosi Marivate, Gerard de Melo, Simon Meoni, Maxime Meyer, Afnan Mir, Nafise Sadat Moosavi, Niklas Muennighoff, Timothy Sum Hon Mun, Kenton Murray, Marcin Namysl, Maria Obedkova, Priti Oli, Nivranshu Pasricha, Jan Pfister, Richard Plant, Vinay Prabhu, Vasile Pais, Libo Qin, Shahab Raji, Pawan Kumar Rajpoot, Vikas Raunak, Roy Rinberg, Nicolas Roberts, Juan Diego Rodriguez, Claude Roux, Vasconcellos P. H. S., Ananya B. Sai, Robin M. Schmidt, Thomas Scialom, Tshephisho Sefara, Saqib N. Shamsi, Xudong Shen, Haoyue Shi, Yiwen Shi, Anna Shvets, Nick Siegel, Damien Sileo, Jamie Simon, Chandan Singh, Roman Sitelew, Priyank Soni, Taylor Sorensen, William Soto, Aman Srivastava, KV Aditya Srivatsa, Tony Sun, Mukund Varma T, A Tabassum, Fiona Anting Tan, Ryan Teehan, Mo Tiwari, Marie Tolkiehn, Athena Wang, Zijian Wang, Gloria Wang, Zijie J. Wang, Fuxuan Wei, Bryan Wilie, Genta Indra Winata, Xinyi Wu, Witold Wydmański, Tianbao Xie, Usama Yaseen, M. Yee, Jing Zhang, and Yue Zhang. 2021 · 2021
Later among the works it cites.
Anticipating safety issues in E2E conversational AI: framework and tooling
Emily Dinan, Gavin Abercrombie, A. Stevie Bergman, Shannon L. Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser. 2021 · 2021
Later among the works it cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021 · 2021
Later among the works it cites.
SummEval: Re-evaluating Summarization Evaluation
Alexander R. Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Later among the works it cites.
GO FIGURE: A meta evaluation of factuality in summarization
Saadia Gabriel, Asli Celikyilmaz, Rahul Jha, Yejin Choi, and Jianfeng Gao. 2021 · 2021
Later among the works it cites.
Datasheets for datasets
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford. 2021 · 2021
Later among the works it cites.
The GEM benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021 · 2021
Later among the works it cites.
Robustness gym: Unifying the NLP evaluation landscape
Karan Goel, Nazneen Fatema Rajani, Jesse Vig, Zachary Taschdjian, Mohit Bansal, and Christopher Ré. 2021 · 2021
Later among the works it cites.
Annotating and modeling fine-grained factuality in summarization
Tanya Goyal and Greg Durrett. 2021 · 2021
Later among the works it cites.
XL-sum: Large-scale multilingual abstractive summarization for 44 languages
Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021 · 2021
Later among the works it cites.
$qˆ2$: Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021 · 2021
Later among the works it cites.
What happens if you treat ordinal ratings as interval data? human evaluations in NLP are even more under-powered than you think
David M. Howcroft and Verena Rieser. 2021 · 2021
Later among the works it cites.
Towards accountability for machine learning datasets: Practices from software engineering and infrastructure
Ben Hutchinson, Andrew Smart, Alex Hanna, Emily Denton, Christina Greer, Oddur Kjartansson, Parker Barnes, and Margaret Mitchell. 2021 · 2021
Later among the works it cites.
A survey of nlp-related crowdsourcing hits: what works and what does not
Jessica Huynh, Jeffrey Bigham, and Maxine Eskénazi. 2021 · 2021
Later among the works it cites.
The perils of using Mechanical Turk to evaluate open-ended text generation
Marzena Karpinska, Nader Akoury, and Mohit Iyyer. 2021 · 2021
Later among the works it cites.
Bidimensional leaderboards: Generate and evaluate language hand in hand
Jungo Kasai, Keisuke Sakaguchi, Ronan Le Bras, Lavinia Dunagan, Jacob Morrison, Alexander R. Fabbri, Yejin Choi, and Noah A. Smith. 2021 · 2021
Later among the works it cites.
Global explainability of bert-based evaluation metrics by disentangling along linguistic factors
Marvin Kaster, Wei Zhao, and Steffen Eger. 2021 · 2021
Later among the works it cites.
GENIE: A leaderboard for human-in-the-loop evaluation of text generation
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A. Smith, and Daniel S. Weld. 2021 · 2021
Later among the works it cites.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021 · 2021
Later among the works it cites.
Structure-augmented keyphrase generation
Jihyuk Kim, Myeongho Jeong, Seungtaek Choi, and Seung-won Hwang. 2021 · 2021
Later among the works it cites.
Reduced, reused and recycled: The life of a dataset in machine learning research
Bernard Koch, Emily Denton, Alex Hanna, and Jacob G. Foster. 2021 · 2021
Later among the works it cites.
To ship or not to ship: An extensive evaluation of automatic metrics for machine translation
Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021 · 2021
Later among the works it cites.
Generating SOAP notes from doctor-patient conversations using modular summarization techniques
Kundan Krishna, Sopan Khosla, Jeffrey Bigham, and Zachary C. Lipton. 2021 · 2021
Later among the works it cites.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2021 · 2021
Later among the works it cites.
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021 · 2021
Later among the works it cites.
Are we learning yet? a meta review of evaluation failures across machine learning
Thomas Liao, Rohan Taori, Inioluwa Deborah Raji, and Ludwig Schmidt. 2021 · 2021
Later among the works it cites.
ExplainaBoard: An explainable leaderboard for NLP
Pengfei Liu, Jinlan Fu, Yang Xiao, Weizhe Yuan, Shuaichen Chang, Junqi Dai, Yixin Liu, Zihuiwen Ye, and Graham Neubig. 2021 · 2021
Later among the works it cites.
What’s in the box? an analysis of undesirable content in the Common Crawl corpus
Alexandra Luccioni and Joseph Viviano. 2021 · 2021
Later among the works it cites.
Encouraging lexical translation consistency for document-level neural machine translation
Xinglin Lyu, Junhui Li, Zhengxian Gong, and Min Zhang. 2021 · 2021
Later among the works it cites.
Reusable templates and guides for documenting datasets and models for natural language processing and generation: A case study of the HuggingFace and GEM data and model cards
Angelina McMillan-Major, Salomey Osei, Juan Diego Rodriguez, Pawan Sasanka Ammanamanchi, Sebastian Gehrmann, and Yacine Jernite. 2021 · 2021
Later among the works it cites.
Automatic construction of evaluation suites for natural language generation datasets
Simon Mille, Kaustubh Dhole, Saad Mahamood, Laura Perez-Beltrachini, Varun Gangal, Mihir Kale, Emiel van Miltenburg, and Sebastian Gehrmann. 2021 · 2021
Later among the works it cites.
Evaluating the robustness of neural language models to input perturbations
Milad Moradi and Matthias Samwald. 2021 · 2021
Later among the works it cites.
Planning with learned entity prompts for abstractive summarization
Shashi Narayan, Yao Zhao, Joshua Maynez, Gonçalo Simoes, Vitaly Nikolaev, and Ryan McDonald. 2021 · 2021
Later among the works it cites.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021 · 2021
Later among the works it cites.
Models and datasets for cross-lingual summarisation
Laura Perez-Beltrachini and Mirella Lapata. 2021 · 2021
Later among the works it cites.
Adversarially constructed evaluation sets are more challenging, but may not be fair
Jason Phang, Angelica Chen, William Huang, and Samuel R. Bowman. 2021 · 2021
Later among the works it cites.
MAUVE: measuring the gap between neural text and human text using divergence frontiers
Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaïd Harchaoui. 2021 · 2021
Later among the works it cites.
On releasing annotator-level labels and information in datasets
Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz. 2021 · 2021
Later among the works it cites.
Evaluating the morphosyntactic well-formedness of generated texts
Adithya Pratapa, Antonios Anastasopoulos, Shruti Rijhwani, Aditi Chaudhary, David R. Mortensen, Graham Neubig, and Yulia Tsvetkov. 2021 · 2021
Later among the works it cites.
Learning compact metrics for MT
Amy Pu, Hyung Won Chung, Ankur Parikh, Sebastian Gehrmann, and Thibault Sellam. 2021 · 2021
Later among the works it cites.
Measuring attribution in natural language generation models
Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, and David Reitter. 2021 · 2021
Later among the works it cites.
Data-questeval: A referenceless metric for data-to-text semantic evaluation
Clément Rebuffel, Thomas Scialom, Laure Soulier, Benjamin Piwowarski, Sylvain Lamprier, Jacopo Staiano, Geoffrey Scoutheeten, and Patrick Gallinari. 2021 · 2021
Later among the works it cites.
Changing the world by changing the data
Anna Rogers. 2021 · 2021
Later among the works it cites.
‘just what do you think you’re doing, dave?’ a checklist for responsible data use in NLP
Anna Rogers, Timothy Baldwin, and Kobi Leins. 2021 · 2021
Later among the works it cites.
Perturbation checklists for evaluating NLG evaluation metrics
Ananya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, and Mitesh M. Khapra. 2021 · 2021
Later among the works it cites.
”everyone wants to do the model work, not the data work”: Data cascades in high-stakes AI
Nithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong, Praveen K. Paritosh, and Lora Aroyo. 2021 · 2021
Later among the works it cites.
Do datasets have politics? disciplinary values in computer vision dataset development
Morgan Klaus Scheuerman, Alex Hanna, and Emily Denton. 2021 · 2021
Later among the works it cites.
Questeval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021 · 2021
Later among the works it cites.
Anastasia Shimorina and Anya Belz. 2021 · 2021
Later among the works it cites.
Beyond fair pay: Ethical implications of NLP crowdsourcing
Boaz Shmueli, Jan Fell, Soumya Ray, and Lun-Wei Ku. 2021 · 2021
Later among the works it cites.
Reward optimization for neural machine translation with learned metrics
Raphael Shu, Kang Min Yoo, and Jung-Woo Ha. 2021 · 2021
Later among the works it cites.
We need to talk about random splits
Anders Søgaard, Sebastian Ebert, Jasmijn Bastings, and Katja Filippova. 2021 · 2021
Later among the works it cites.
Generation challenges: Results of the accuracy evaluation shared task
Craig Thomson and Ehud Reiter. 2021 · 2021
Later among the works it cites.
Human evaluation of automatically generated text: Current trends and best practice guidelines
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, and Emiel Krahmer. 2021 · 2021
Later among the works it cites.
Underreporting of errors in NLG output, and what to do about it
Emiel van Miltenburg, Miruna Clinciu, Ondřej Dušek, Dimitra Gkatzia, Stephanie Inglis, Leo Leppänen, Saad Mahamood, Emma Manning, Stephanie Schoch, Craig Thomson, and Luou Wen. 2021 · 2021
Later among the works it cites.
Does it capture STEL? a modular, similarity-based linguistic style evaluation framework
Anna Wegmann and Dong Nguyen. 2021 · 2021
Later among the works it cites.
The statistical advantage of automatic NLG metrics at the system level
Johnny Wei and Robin Jia. 2021 · 2021
Later among the works it cites.
Factual consistency evaluation for text summarization via counterfactual estimation
Yuexiang Xie, Fei Sun, Yang Deng, Yaliang Li, and Bolin Ding. 2021 · 2021
Later among the works it cites.
Synthbio: A case study in faster curation of text datasets
Ann Yuan, Daphne Ippolito, Vitaly Nikolaev, Chris Callison-Burch, Andy Coenen, and Sebastian Gehrmann. 2021 · 2021
Later among the works it cites.
Gradient-based adversarial factual consistency evaluation for abstractive summarization
Zhiyuan Zeng, Jiaze Chen, Weiran Xu, and Lei Li. 2021 · 2021
Later among the works it cites.
Finding a balanced degree of automation for summary evaluation
Shiyue Zhang and Mohit Bansal. 2021 · 2021
Later among the works it cites.
Eric Michael Smith, Orion Hsu, Rebecca Qian, Stephen Roller, Y-Lan Boureau, and Jason Weston. 2022 · 2022
Closest in time.
Are factuality checkers reliable? adversarial meta-evaluation of factuality in summarization
Yiran Chen, Pengfei Liu, and Xipeng Qiu. 2021 · 2095
Closest in time.