Fetching the paper…
Reading the bibliography…
The paper surveys evaluation methods of natural language generation (NLG) systems that have been developed in the last few years.
Evaluating the state-of-the-art of end-to-end natural language generation: The E2E NLG challenge
Ondrej Dusek, Jekaterina Novikova, and Verena Rieser · 1901
Earlier work this paper cites.
Cristina Garbacea, Samuel Carton, Shiyan Yan, and Qiaozhu Mei · 1901
Earlier work this paper cites.
Strategies for structuring story generation
Angela Fan, Mike Lewis, and Yann N. Dauphin · 1902
Earlier work this paper cites.
Jointly measuring diversity and quality in text generation models
Ehsan Montahaei, Danial Alihosseini, and Mahdieh Soleymani Baghshah · 1904
Earlier work this paper cites.
Jointly measuring diversity and quality in text generation models
Ehsan Montahaei, Danial Alihosseini, and Mahdieh Soleymani Baghshah · 1904
Earlier work this paper cites.
Effective estimation of deep generative language models
Tom Pelsmaeker and Wilker Aziz · 1904
Earlier work this paper cites.
Generating token-level explanations for natural language inference
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 1904
Earlier work this paper cites.
Survey on evaluation methods for dialogue systems
Jan Deriu, Álvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak · 1905
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon · 1905
Earlier work this paper cites.
Recent advances in neural question generation
Liangming Pan, Wenqiang Lei, Tat-Seng Chua, and Min-Yen Kan · 1905
Earlier work this paper cites.
Qiuyun Zhang, Bin Guo, Hao Wang, Yunji Liang, Shaoyang Hao, and Zhiwen Yu · 1905
Earlier work this paper cites.
ELI5: long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli · 1907
Earlier work this paper cites.
Saadia Gabriel, Antoine Bosselut, Ari Holtzman, Kyle Lo, Asli Celikyilmaz, and Yejin Choi · 1907
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
ERNIE 2.0: A continual pre-training framework for language understanding
Yu Sun, Shuohuan Wang, Yu-Kun Li, Shikun Feng, Hao Tian, Hua Wu, and Haifeng Wang · 1907
Earlier work this paper cites.
Optimizing the factual correctness of a summary: A study of summarizing radiology reports
Yuhao Zhang, Derek Merck, Emily Bao Tsai, Christopher D. Manning, and Curtis P. Langlotz · 1911
Earlier work this paper cites.
Computing Machinary and Intelligence
A. M. Turing · 1950
Earlier work this paper cites.
Reliability of content analysis: The case of nominal scale coding
William A. Scott · 1955
Earlier work this paper cites.
A coefficient of agreement for nominal scales
Jacob Cohen · 1960
Earlier work this paper cites.
On the social psychology of the psychological experiment: With particular reference to demand characteristics and their implications
Martin T. Orne · 1962
Earlier work this paper cites.
Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit
Jacob Cohen · 1968
Earlier work this paper cites.
Estimating the reliability, systematic error and random error of interval data
Klaus Krippendorff · 1970
Earlier work this paper cites.
Measuring nominal scale agreement among many raters
JL Fleiss · 1971
Earlier work this paper cites.
Text generation. using discourse strategies and focus constraints to generate natural language text
Kathleen R. McKeown · 1985
Earlier work this paper cites.
Rhetorical structure theory: Description and construction of text structures
W.C. Mann and S.A. Thompson · 1987
Earlier work this paper cites.
Type/token ratios: what do they really tell us?
Brian Richards · 1987
Earlier work this paper cites.
The sphinx-ii speech recognition system: An overview
Xuedong Huang, Fileno Alleva, Hsiao wuen Hon, Mei yuh Hwang, and Ronald Rosenfeld · 1992
Earlier work this paper cites.
Abstract generation based on rhetorical structure extraction
Kenji Ono, Kazuo Sumita, and Seiji Miike · 1994
Earlier work this paper cites.
Wordnet: A lexical database for english
George A. Miller · 1995
Earlier work this paper cites.
Magnitude estimation of linguistic acceptability
Ellen Gurman Bard, Dan Robertson, and Antonella Sorace · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
From discourse structures to text summaries
Daniel Marcu · 1997
Earlier work this paper cites.
A metric for distributions with applications to image databases
Y. Rubner, C. Tomasi, and L. Guibas · 1998
Earlier work this paper cites.
Dimlex: A lexicon of discourse markers for text generation and understanding
Manfred Stede and Carla Umbach · 1998
Earlier work this paper cites.
Beyond kappa: A review of interrater agreement measures
Mousumi Banerjee, Michelle Hopkins Capozzoli, Laura A. McSweeney, and Debajyoti Sinha · 1999
Earlier work this paper cites.
Using grice’s maxim of quantity to select the content of plan descriptions
R. Michael Young · 1999
Earlier work this paper cites.
Statistics-based summarization – step one: Sentence compression
D. Knight, K.—Marcu · 2000
Earlier work this paper cites.
The limsi arise system
Lori Lamel, Sophie Rosse, Jean-Luc Gauvain, Samir Bennacef, Matine Garnier-Rizet, and Bernard Prouts · 2000
Earlier work this paper cites.
The nist 1999 speaker recognition evaluation - an overview, 2000
A. Martin and M. Przybocki · 2000
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
George Doddington · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Logics of Conversation
N. Asher and A. Lascarides · 2003
Earlier work this paper cites.
Learning to paraphrase: An unsupervised approach using multiple-sequence alignment
Regina Barzilay and Lillian Lee · 2003
Earlier work this paper cites.
Probabilistic text structuring: Experiments with sentence ordering
Mirella Lapata · 2003
Earlier work this paper cites.
Precision and recall of machine translation
I. Melamed, Ryan Green, and Joseph Turian · 2003
Earlier work this paper cites.
Lessons from a failure: Generating tailored smoking cessation letters
Ehud Reiter, Roma Robertson, and Liesl Osman · 2003
Earlier work this paper cites.
A quantitative method for machine translation evaluation
Jesús Tomás, Josep Àngel Mas, and Francisco Casacuberta · 2003
Earlier work this paper cites.
Evaluation of machine translation and its evaluation
L. Shen Turian, J. P. and I. D. Melamed · 2003
Earlier work this paper cites.
Linguistic correlates of style: authorship classification with deep linguistic analysis features
Michael Gamon · 2004
Earlier work this paper cites.
The significance of recall in automatic metrics for mt evaluation
Alon Lavie, Kenji Sagae, and Shyamsundar Jayaraman · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Automatic evaluation of machine translation quality using longest common subsequence and skip-bigram statistics
Chin-Yew Lin and Franz Josef Och · 2004
Earlier work this paper cites.
On the use of information retrieval measures for speech recognition evaluation
Iain Mccowan, Darren Moore, John Dines, Daniel Gatica-Perez, Mike Flynn, Pierre Wellner, and Herve Bourlard · 2004
Earlier work this paper cites.
Evaluating content selection in summarization: The pyramid method
Ani Nenkova and Rebecca Passonneau · 2004
Earlier work this paper cites.
Monolingual machine translation for paraphrase generation
Chris Quirk, Chris Brockett, and William Dolan · 2004
Earlier work this paper cites.
Interpreting bleu/nist scores: How much improvement do we need to have a better system
Ying Zhang, Stephan Vogel, and Alex Waibel · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Viewing referring expression generation as search
Bernd Bohnet and Robert Dale · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
Bill Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Automatic evaluation of text coherence: Models and representations
Mirella Lapata and Regina Barzilay · 2005
Earlier work this paper cites.
Nist 2005 machine translation evaluation official results
Audrey J. Lee and Mark A. Przybocki · 2005
Earlier work this paper cites.
Syntactic features for evaluation of machine translation
Ding Liu and Daniel Gildea · 2005
Earlier work this paper cites.
Choosing words in computer-generated weather forecasts
Ehud Reiter, Somayajulu Sripada, Jim Hunter, Jin Yu, and Ian Davy · 2005
Earlier work this paper cites.
Discourse chunking and its application to sentence compression
M. Sporleder, C.—Lapata · 2005
Earlier work this paper cites.
Comparing automatic and human evaluation of nlg systems
Anja Belz and Ehud Reiter · 2006
Earlier work this paper cites.
Re-evaluating the role of bleu in machine translation research
Chris Callison-Burch and Miles Osborne · 2006
Earlier work this paper cites.
Automatic evaluation of machine translation quality
Cyril Goutte · 2006
Earlier work this paper cites.
Paraphrasing for automatic evaluation
David Kauchak and Regina Barzilay · 2006
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul · 2006
Earlier work this paper cites.
Sentence compression for the lsa-based summarizer
K. Steinberger, J.—Jezek · 2006
Earlier work this paper cites.
Meta- evaluation of machine translation
Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder · 2007
Earlier work this paper cites.
Evaluating algorithms for the generation of referring expressions using a balanced corpus
Albert Gatt, Ielka Van Der Sluis, and Kees Van Deemter · 2007
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgments
Alon Lavie and Abhaya Agarwal · 2007
Earlier work this paper cites.
An architecture for data-to-text systems
Ehud Reiter · 2007
Earlier work this paper cites.
Improved word-level system combination for machine translation
Antti-veikko I. Rosti, Spyros Matsoukas, and Richard Schwartz · 2007
Earlier work this paper cites.
Meteor, mbleu and mter: Evaluation metrics for high-correlation with human rankings of machine translation output
Abhaya Agarwal and Alon Lavie · 2008
Earlier work this paper cites.
Using the f-measure as similarity measure for automatic text summarization
Ramiz Aliguliyev · 2008
Earlier work this paper cites.
Inter-coder agreement for computational linguistics
Ron Artstein and Massimo Poesio · 2008
Earlier work this paper cites.
Learning to sportscast: A test of grounded language acquisition
David L. Chen and Raymond J. Mooney · 2008
Earlier work this paper cites.
Global inference for sentence compression: An integer linear programming approach
J. Clarke and M. Lapata · 2008
Earlier work this paper cites.
Coreference-inspired coherence modeling
M. Elsner and E. Charniak · 2008
Earlier work this paper cites.
Attribute selection for referring expression generation: New algorithms and evaluation methods
Albert Gatt and Anja Belz · 2008
Earlier work this paper cites.
The TUNA challenge 2008: Overview and evaluation results
Albert Gatt, Anja Belz, and Eric Kow · 2008
Earlier work this paper cites.
Correlation between rouge and human evaluation of extractive meeting summaries
Feifan Liu and Yang Liu · 2008
Earlier work this paper cites.
The good-subject effect: investigating participant demand characteristics
Austin Lee Nichols and Jon K. Maner · 2008
Earlier work this paper cites.
The GREC named entity generation challenge 2009: Overview and evaluation results
Anja Belz, Eric Kow, and Jette Viethen · 2009
Earlier work this paper cites.
Evaluation metrics for text summarization
Jan Deriu, Álvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak · 2009
Earlier work this paper cites.
Speech and language processing: An introduction to natural language processing, computational linguistics, and speech recognition
Daniel Jurafsky and James H. Martin · 2009
Earlier work this paper cites.
The software architecture for the first challenge on generating instructions in virtual environments
Alexander Koller, Donna Byron, Justine Cassell, Robert Dale, Johanna Moore, Jon Oberlander, and Kristina Striegnitz · 2009
Earlier work this paper cites.
Learning semantic correspondences with less supervision
Percy Liang, Michael Jordan, and Dan Klein · 2009
Earlier work this paper cites.
Instructional manipulation checks: Detecting satisficing to increase statistical power
Daniel M Oppenheimer, Tom Meyvis, and Nicolas Davidenko · 2009
Earlier work this paper cites.
Textual entailment features for machine translation evaluation
Sebastian Padó, Michel Galley, Dan Jurafsky, and Christoper Manning · 2009
Earlier work this paper cites.
An investigation into the validity of some metrics for automatically evaluating natural language generation systems
Ehud Reiter and Anja Belz · 2009
Earlier work this paper cites.
Evaluation measures for text summarization
Josef Steinberger and Karel Jezek · 2009
Earlier work this paper cites.
The GREC challenges 2010: Overview and evaluation results
Anja Belz and Eric Kow · 2010
Earlier work this paper cites.
Quality management on amazon mechanical turk
Panagiotis G. Ipeirotis, F. Provost, and Jing Wang · 2010
Earlier work this paper cites.
Automatic evaluation of translation quality for distant language pairs
Hideki Isozaki, Tsutomu Hirao, Kevin Duh, Katsuhito Sudoh, and Hajime Tsukada · 2010
Earlier work this paper cites.
Empirical Methods in Natural Language Generation: Data-oriented Methods and Empirical Evaluation
E. Krahmer and M. Theune · 2010
Earlier work this paper cites.
Phrase-based statistical language generation using graphical models and active learning
F. Mairesse, M. Gasic, F. Jurcicek, S. Keizer, B. Thompson, K. Yu, and S. Young · 2010
Earlier work this paper cites.
Mtld, vocd-d, and hd-d: A validation study of sophisticated approaces to lexical diversity assessment
P.M. McCarthy and S. Jarvis · 2010
Earlier work this paper cites.
Spoken dialog challenge 2010: Comparison of live and control test results
Alan Black, Susanne Burger, Alistair Conkie, Helen Hastie, Simon Keizer, Oliver Lemon, Nicolas Merigaud, Gabriel Parent, Gabriel Schubiner, Blaise Thomson, Jason Williams, Kai Yu, Steve Young, and Maxine Eskenazi · 2011
Earlier work this paper cites.
Tesla at wmt 2011: Translation evaluation and tunable metric
Daniel Dahlmeier, Chang Liu, and Hwee Tou Ng · 2011
Cited alongside, same era.
Real user evaluation of spoken dialogue systems using amazon mechanical turk
Filip Jurcicek, Simon Keizer, Milica Gasic, François Mairesse, Blaise Thomson, Kai Yu, and Steve Young · 2011
Cited alongside, same era.
Evaluation without references: Ibm1 scores as evaluation metrics
Maja Popovic, David Vilar, Eleftherios Avramidis, and Aljoscha Burchardt · 2011
Cited alongside, same era.
Data-driven response generation in social media
Alan Ritter, Colin Cherry, and Bill Dolan · 2011
Cited alongside, same era.
Report on the second challenge on generating instructions in virtual environments (give-2.5)
Kristina Striegnitz, Denis Alexandre, Andrew Gargett, Alexander Garoufi, Konstantina Koller, and Mariet Theune · 2011
Cited alongside, same era.
Extraction based automatic text summarization system with hmm tagger
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Later among the works it cites.
Demographics and dynamics of mechanical turk workers
Djellel Eddine Difallah, Elena Filatova, and Panagiotis G. Ipeirotis · 2018
Later among the works it cites.
Findings of the E2E NLG challenge
Ondřej Dušek, Jekaterina Novikova, and Verena Rieser · 2018
Later among the works it cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann N. Dauphin · 2018
Later among the works it cites.
Style transfer in text: Exploration and evaluation
Zhenxin Fu, Xiaoye Tan, Nanyun Peng, Dongyan Zhao, and Rui Yan · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Manne Suneetha and Sheerin Fatima · 2011
Cited alongside, same era.
Pet: a tool for post-editing and assessing machine translation
W. Aziz, Sheila Castilho, and Lucia Specia · 2012
Cited alongside, same era.
The influence of corpus quality on statistical measurements on language resources
Thomas Eckart, Uwe Quasthoff, and Dirk Goldhahn · 2012
Cited alongside, same era.
Unsupervised concept-to-text generation with hypergraphs
I Konstas and M. Lapara · 2012
Cited alongside, same era.
Fully automatic semantic mt evaluation
Chi-Kiu Lo, Anand Karthik Tumuluru, and Dekai Wu · 2012
Cited alongside, same era.
Machine translation of labeled discourse connectives
Thomas Meyer, Andrei Popescu-Belis, Najeh Hajlaoui, and Andrea Gesmundo · 2012
Cited alongside, same era.
RankME: Reliable human ratings for natural language generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser · 2012
Cited alongside, same era.
Survey of the state of the art in natural language generation: Core tasks, applications and evaluation
Albert Gatt and Emiel Krahmer · 2018
Later among the works it cites.
Neural network methods for natural language processing
Yoav Goldberg, Graeme Hirst, Yang Liu, , and Meng Zhang · 2018
Later among the works it cites.
Machine translation evaluation resources and methods: A survey
Lifeng Han · 2018
Later among the works it cites.
A comprehensive survey of deep learning for image captioning
Md. Zakir Hossain, Ferdous Sohel, Mohd Fairuz Shiratuddin, and Hamid Laga · 2018
Later among the works it cites.
A tutorial on deep latent variable models of natural language
Yoon Kim, Sam Wiseman, and Alexander M. Rush · 2018
Later among the works it cites.
Generating reasonable and diversified story ending using sequence to sequence model with adversarial training
Zhongyang Li, Xiao Ding, and Ting Liu · 2018
Later among the works it cites.
An efficient framework for learning sentence representations
Lajanugen Logeswaran and Honglak Lee · 2018
Later among the works it cites.
An entailment-based scoring method for content selection in document summarization
Dang Hoang Long, Minh-Tien Nguyen, Ngo Xuan Bach, Le-Minh Nguyen, and Tu Minh Phuong · 2018
Later among the works it cites.
Neural text generation: Past, present and beyond
Sidi Lu, Yaoming Zhu, Weinan Zhang, Jun Wang, and Yong Yu · 2018
Later among the works it cites.
Ranking sentences for extractive summarization with reinforcement learning
Shashi Narayan, Shay B. Cohen, and Mirella Lapata · 2018
Later among the works it cites.
Towards a better metric for evaluating question generation systems
Preksha Nema and Mitesh M. Khapra · 2018
Later among the works it cites.
Adversarial inference for multi-sentence video description
Jae Sung Park, Marcus Rohrbach, Trevor Darrell, and Anna Rohrbach · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Later among the works it cites.
Data-to-text generation with content selection and planning
Ratish Puduppully, Li Dong, and Mirella Lapata · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Later among the works it cites.
A structured review of the validity of BLEU
Ehud Reiter · 2018
Later among the works it cites.
Anchors: High-precision model-agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2018
Later among the works it cites.
Learning-based composite metrics for improved caption evaluation
Naeha Sharif, Lyndon White, Mohammed Bennamoun, and Syed Afaq Ali Shah · 2018
Later among the works it cites.
Neural abstractive text summarization with sequence-to-sequence models
Tian Shi, Yaser Keneshloo, Naren Ramakrishnan, and Chandan K. Reddy · 2018
Later among the works it cites.
RUSE: Regressor using sentence embeddings for automatic machine translation evaluation
Hiroki Shimanaka, Tomoyuki Kajiwara, and Mamoru Komachi · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Later among the works it cites.
Plan-and-write: Towards better automatic storytelling
Lili Yao, Nanyun Peng, Ralph M. Weischedel, Kevin Knight, Dongyan Zhao, and Rui Yan · 2018
Later among the works it cites.
Texygen: A benchmarking platform for text generation models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu · 2018
Later among the works it cites.
Keyphrase generation: A multi-aspect survey
Erion Cano and Ondrej Bojar · 2019
Later among the works it cites.
WMDO: Fluency-based word mover’s distance for machine translation evaluation
Julian Chow, Lucia Specia, and Pranava Madhyastha · 2019
Later among the works it cites.
Sentence mover’s similarity: Automatic evaluation for multi-sentence texts
Elizabeth Clark, Asli Celikyilmaz, and Noah A. Smith · 2019
Later among the works it cites.
Handling divergent reference texts in table-to-text generation
Bhuwan Dhingra, Manaal Faruqui, Ankur Parikh, Ming-Wei Chang, Dipanjan Das, and William W Cohen · 2019
Later among the works it cites.
Automated rationale generation: A technique for explainable ai and its effects on human perceptions
Upol Ehsan, Pradyumna Tambwekar, Larry Chan, Brent Harrison, and Mark O. Riedl · 2019
Later among the works it cites.
Question answering as an automatic evaluation metric for news article summarization
Matan Eyal, Tal Baumel, and Michael Elhadad · 2019
Later among the works it cites.
Question answering as an automatic evaluation metric for news article summarization
Matan Eyal, Tal Baumel, and Michael Elhadad · 2019
Later among the works it cites.
Ranking generated summaries by correctness: An interesting but challenging application for natural language inference
Tobias Falke, Leonardo F. R. Ribeiro, Prasetya Ajie Utama, Ido Dagan, and Iryna Gurevych · 2019
Later among the works it cites.
Neural approaches to conversational ai
Jianfeng Gao, Michel Galley, and Lihong Li · 2019
Later among the works it cites.
GLTR: Statistical detection and visualization of generated text
Sebastian Gehrmann, Hendrik Strobelt, and Alexander Rush · 2019
Later among the works it cites.
Unifying human and statistical evaluation for natural language generation
Tatsunori Hashimoto, Hugh Zhang, and Percy Liang · 2019
Later among the works it cites.
Tiger: Text-to-image grounding for image caption evaluation
Ming Jiang, Qiuyuan Huang, Lei Zhang, Xin Wang, Pengchuan Zhang, Zhe Gan, Jana Diesner, and Jianfeng Gao · 2019
Later among the works it cites.
Towards neural language evaluators
Hassan Kané, Yusuf Kocyigit, Pelkins Ajanoh, Ali Abdalla, and Mohamed Coulibali · 2019
Later among the works it cites.
Text generation from knowledge graphs with graph transformers
Rik Koncel-Kedziorski, Dhanush Bekal, Yi Luan, Mirella Lapata, and Hannaneh Hajishirzi · 2019
Later among the works it cites.
Neural text summarization: A critical evaluation
Wojciech Kryscinski, Nitish Shirish Keskar, Bryan McCann, Caiming Xiong, and Richard Socher · 2019
Later among the works it cites.
Best practices for the human evaluation of automatically generated text
Chris Van Der Lee, Albert Gatt, Emiel van Miltenburg, Sander Wubben, and Emiel Krahmer · 2019
Later among the works it cites.
Visual to text: Survey of image and video captioning
Sheng Li, Zhiqiang Tao, and Yun Fu · 2019
Later among the works it cites.
Multi-task deep neural networks for natural language understanding
Xiaodong Liu, Pengcheng He, Weizhu Chen, and Jianfeng Gao · 2019
Later among the works it cites.
YiSi - a unified semantic MT quality evaluation and estimation metric for languages with different levels of available resources
Chi-kiu Lo · 2019
Later among the works it cites.
How decoding strategies affect the verifiability of generated text, 2019
Luca Massarelli, Fabio Petroni, Aleksandra Piktus, Myle Ott, Tim Rocktäschel, Vassilis Plachouras, Fabrizio Silvestri, and Sebastian Riedel · 2019
Later among the works it cites.
Putting evaluation in context: Contextual embeddings improve machine translation evaluation
Nitika Mathur, Timothy Baldwin, and Trevor Cohn · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Sentence-bert: Sentence embeddings using siamese bert-networks, 2019
Nils Reimers and Iryna Gurevych · 2019
Later among the works it cites.
Ehud reiter’s blog, 2019
Ehud Reiter · 2019
Later among the works it cites.
Towards debiasing fact verification models
Tal Schuster, Darsh Shah, Yun Jie Serene Yeo, Daniel Roberto Filizzola Ortiz, Enrico Santus, and Regina Barzilay · 2019
Later among the works it cites.
Chateval: A tool for chatbot evaluation
João Sedoc, Daphne Ippolito, Arun Kirubarajan, Jai Thirani, Lyle Ungar, and Chris Callison-Burch · 2019
Later among the works it cites.
On accurate evaluation of GANs for language generation
Stanislau Semeniuta, Aliaksei Severyn, and Sylvain Gelly · 2019
Later among the works it cites.
Evaluating text output in nlp: Bleu at your own risk, 2019
Rachel Tatman · 2019
Later among the works it cites.
Best practices for the human evaluation of automatically generated text
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, Sander Wubben, and Emiel Krahmer · 2019
Later among the works it cites.
Learning from fact-checkers: Analysis and generation of fact-checking language
Nguyen Vo and Kyumin Lee · 2019
Later among the works it cites.
Revisiting challenges in data-to-text generation with fact grounding
Hongmin Wang · 2019
Later among the works it cites.
Neural text generation with unlikelihood training
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston · 2019
Later among the works it cites.
Extending machine translation evaluation metrics with lexical cohesion to document level
Billy T. M. Wong and Chunyu Kit · 2019
Later among the works it cites.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi · 2019
Later among the works it cites.
Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger · 2019
Later among the works it cites.
A survey and taxonomy of adversarial neural networks for text‐to‐image synthesis
Jorge Agnese, Jonathan Herrera, Haicheng Tao, and Xingquan Zhu · 2020
Closest in time.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Closest in time.
Lessons from computational modelling of reference production in Mandarin and English
Guanyi Chen and Kees van Deemter · 2020
Closest in time.
A comprehensive survey of multilingual neural machine translation, 2020
Raj Dabre, Chenhui Chu, and Anoop Kunchukuttan · 2020
Closest in time.
Plug and play language models: A simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu · 2020
Closest in time.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Esin Durmus, He He, and Mona Diab · 2020
Closest in time.
Evaluating semantic accuracy of data-to-text generation with natural language inference
Ondřej Dušek and Zdeněk Kasner · 2020
Closest in time.
Evaluating factuality in generation with dependency-level entailment
Tanya Goyal and Greg Durrett · 2020
Closest in time.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi · 2020
Closest in time.
Twenty years of confusion in human evaluation: Nlg needs evaluation sheets and standardised definitions
David M. Howcroft, Anya Belz, Miruna Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, S. Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser · 2020
Closest in time.
Multi-domain task-oriented dialog challenge ii
Jinchao Li, Qi Zhu, Baolin Peng, Lars Liden, Runze Liang, Ryuichi Takanobu, Shahin Shayandeh, Swadheen Shukla, Zheng Zhang, Minlie Huang, and Jianfeng Gao · 2020
Closest in time.
Learning compact reward for image captioning
Nannan Li and Zhenzhong Chen · 2020
Closest in time.
Kelvin Luu, Rik Koncel-Kedziorski, Kyle Lo, Isabel Cachola, and Noah A. Smith · 2020
Closest in time.
A set of recommendations for assessing human–machine parity in language translation
Samuel Läubli, Sheila Castilho, Graham Neubig, Rico Sennrich, Qinlan Shen, and Antonio Toral · 2020
Closest in time.
Tangled up in bleu: Reevaluating the evaluation of automatic machine translation evaluation metrics
Nitika Mathur, Timothy Baldwin, and Trevor Cohn · 2020
Closest in time.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan T. McDonald · 2020
Closest in time.
The radicalization risks of gpt-3 and advanced neural language models
Kris McGuffie and Alex Newhouse · 2020
Closest in time.
Totto: A controlled table-to-text generation dataset, 2020
Ankur P. Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das · 2020
Closest in time.
Few-shot natural language generation for task-oriented dialog
Baolin Peng, Chenguang Zhu, Chunyuan Li, Xiujun Li, Jinchao Li, Michael Zeng, and Jianfeng Gao · 2020
Closest in time.
Plotmachines: Outline-conditioned generation with dynamic plot state tracking
Hannah Rashkin, Asli Celikyilmaz, Yejin Choi, and Jianfeng Gao · 2020
Closest in time.
Bleurt: Learning robust metrics for text generation, 2020
Thibault Sellam, Dipanjan Das, and Ankur P. Parikh · 2020
Closest in time.
Towards Controllable Biases in Language Generation
Emily Sheng, Kai-Wei Chang, Prem Natarajan, and Nanyun Peng · 2020
Closest in time.
Towards faithful neural table-to-text generation with content-matching constraints
Zhenyi Wang, Xiaoyang Wang, Bang An, Dong Yu, and Changyou Chen · 2020
Closest in time.
Prophetnet: Predicting future n-gram for sequence-to-sequence pre-training
Yu Yan, Weizhen Qi, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, and Ming Zhou · 2020
Closest in time.
Evaluating machines by their real-world language use
Rowan Zellers, Ari Holtzman, Elizabeth Clark, Lianhui Qin, Ali Farhadi, and Yejin Choi · 2020
Closest in time.
Optimizing the factual correctness of a summary: A study of summarizing radiology reports
Yuhao Zhang, Derek Merck, Emily Tsai, Christopher D. Manning, and Curtis Langlotz · 2020
Closest in time.
WebNLG challenge 2020: Language agnostic delexicalisation for multilingual RDF-to-text generation
Giulio Zhou and Gerasimos Lampouras · 2020
Closest in time.
The design and implementation of xiaoice, an empathetic social chatbot
Li Zhou, Jianfeng Gao, Di Li, and Heung-Yeung Shum · 2020
Closest in time.
Learning to compare for better training and evaluation of open domain natural language generation models, 2020
Wangchunshu Zhou and Ke Xu · 2020
Closest in time.
Boosting factual correctness of abstractive summarization, 2020
Chenguang Zhu, William Hinthorn, Ruochen Xu, Qingkai Zeng, Michael Zeng, Xuedong Huang, and Meng Jiang · 2020
Closest in time.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Closest in time.
The gem benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin P. Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Aremu Anuoluwapo, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna Clinciu, Dipanjan Das, Kaustubh D. Dhole, Wanyu Du, Esin Durmus, Ondvrej Duvsek, Chris C. Emezue, Varun Gangal, Cristina Garbacea, T. Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, S. Mille, Emiel van Miltenburg, Moin Nadeem, S. Narayan, V. Nikolaev, Rubungo Andre Niyongabo, S. Osei, Ankur P. Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodríguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, W. Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou · 2021
Closest in time.
Genie: A leaderboard for human-in-the-loop evaluation of text generation
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A. Smith, and Daniel S. Weld · 2021
Closest in time.
Human evaluation of automatically generated text: Current trends and best practice guidelines
C. Lee, Albert Gatt, Emiel van Miltenburg, and E. Krahmer · 2021
Closest in time.
Dart: Open-domain structured data record to text generation
Linyong Nan, Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Mutethia Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, Richard Socher, and Nazneen Fatema Rajani · 2021
Closest in time.
Human vs automatic metrics: on the importance of correlation design
Anastasia Shimorina · 2021
Closest in time.
Trading off diversity and quality in natural language generation
Hugh Zhang, Daniel Duckworth, Daphne Ippolito, and Arvind Neelakantan · 2021
Closest in time.