Fetching the paper…
Reading the bibliography…
The evaluation of natural language generation (NLG) tasks is a significant and longstanding research area.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E. Terry. 1952 · 1952
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Learning to summarize from human feedback
Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F. Christiano. 2020 · 2009
Earlier work this paper cites.
An evaluation protocol for generative conversational systems
Seolhwa Lee, Heuiseok Lim, and João Sedoc. 2020 · 2010
Earlier work this paper cites.
Semantically conditioned lstm-based natural language generation for spoken dialogue systems
Tsung-Hsien Wen, Milica Gasic, Nikola Mrksic, Pei-hao Su, David Vandyke, and Steve J. Young. 2015 · 2015
Earlier work this paper cites.
Crowd-sourcing NLG data: Pictures elicit better data
Jekaterina Novikova, Oliver Lemon, and Verena Rieser. 2016 · 2016
Earlier work this paper cites.
Creating training corpora for NLG micro-planners
Claire Gardent, Anastasia Shimorina, Shashi Narayan, and Laura Perez-Beltrachini. 2017 · 2017
Earlier work this paper cites.
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Max Grusky, Mor Naaman, and Yoav Artzi. 2018 · 2018
Earlier work this paper cites.
Subjective annotation and evaluation of three different chatbots WOCHAT: shared task report
Naomi Kong-Vega, Mingxin Shen, Mo Wang, and Luis Fernando D’Haro. 2018 · 2018
Earlier work this paper cites.
Reliability and learnability of human bandit feedback for sequence-to-sequence reinforcement learning
Julia Kreutzer, Joshua Uyheng, and Stefan Riezler. 2018 · 2018
Earlier work this paper cites.
Rankme: Reliable human ratings for natural language generation
Jekaterina Novikova, Ondrej Dusek, and Verena Rieser. 2018 · 2018
Earlier work this paper cites.
BLEU is not suitable for the evaluation of text simplification
Elior Sulem, Omri Abend, and Ari Rappoport. 2018a · 2018
Earlier work this paper cites.
Semantic structural evaluation for text simplification
Elior Sulem, Omri Abend, and Ari Rappoport. 2018b · 2018
Earlier work this paper cites.
Simple and effective text simplification using semantic and neural methods
Elior Sulem, Omri Abend, and Ari Rappoport. 2018c · 2018
Earlier work this paper cites.
Topical-chat: Towards knowledge-grounded open-domain conversations
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tür. 2019 · 2019
Earlier work this paper cites.
Investigating evaluation of open-domain dialogue systems with human generated multiple references
Prakhar Gupta, Shikib Mehri, Tiancheng Zhao, Amy Pavel, Maxine Eskénazi, and Jeffrey P. Bigham. 2019 · 2019
Earlier work this paper cites.
PARABANK: monolingual bitext generation and sentential paraphrasing via lexically-constrained neural machine translation
J. Edward Hu, Rachel Rudinger, Matt Post, and Benjamin Van Durme. 2019 · 2019
Earlier work this paper cites.
Enabling robust grammatical error correction in new domains: Datasets, metrics, and analyses
Courtney Napoles, Maria Nadejde, and Joel R. Tetreault. 2019 · 2019
Earlier work this paper cites.
Chateval: A tool for chatbot evaluation
João Sedoc, Daphne Ippolito, Arun Kirubarajan, Jai Thirani, Lyle H. Ungar, and Chris Callison-Burch. 2019 · 2019
Earlier work this paper cites.
ASSET: A dataset for tuning and evaluation of sentence simplification models with multiple rewriting transformations
Fernando Alva-Manchego, Louis Martin, Antoine Bordes, Carolina Scarton, Benoît Sagot, and Lucia Specia. 2020 · 2020
Earlier work this paper cites.
The 2020 bilingual, bi-directional WebNLG+ shared task: Overview and evaluation results (WebNLG+ 2020)
Thiago Castro Ferreira, Claire Gardent, Nikolai Ilinykh, Chris van der Lee, Simon Mille, Diego Moussallem, and Anastasia Shimorina. 2020 · 2020
Earlier work this paper cites.
Plug and play language models: A simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2020 · 2020
Earlier work this paper cites.
Evaluating the state-of-the-art of end-to-end natural language generation: The E2E NLG challenge
Ondrej Dusek, Jekaterina Novikova, and Verena Rieser. 2020 · 2020
Earlier work this paper cites.
What have we achieved on text summarization?
Dandan Huang, Leyang Cui, Sen Yang, Guangsheng Bao, Kun Wang, Jun Xie, and Yue Zhang. 2020a · 2020
Earlier work this paper cites.
GRADE: automatic graph-enhanced coherence metric for evaluating open-domain dialogue systems
Lishan Huang, Zheng Ye, Jinghui Qin, Liang Lin, and Xiaodan Liang. 2020b · 2020
Earlier work this paper cites.
Unsupervised evaluation of interactive dialog with dialogpt
Shikib Mehri and Maxine Eskénazi. 2020 · 2020
Earlier work this paper cites.
Human annotated dialogues dataset for natural conversational agents
Erinc Merdivan, Deepika Singh, Sten Hanke, Johannes Kropf, Andreas Holzinger, and Matthieu Geist. 2020 · 2020
Earlier work this paper cites.
Towards holistic and automatic evaluation of open-domain dialogue generation
Bo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou, Yixian Liu, and Kewei Tu. 2020 · 2020
Earlier work this paper cites.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C. Farinha, and Alon Lavie. 2020 · 2020
Earlier work this paper cites.
BLEURT: learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P. Parikh. 2020 · 2020
Cited alongside, same era.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2020
Cited alongside, same era.
SOME: reference-less sub-metrics optimized for manual evaluations of grammatical error correction
Ryoma Yoshimura, Masahiro Kaneko, Tomoyuki Kajiwara, and Mamoru Komachi. 2020 · 2020
Cited alongside, same era.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Cited alongside, same era.
Designing precise and robust dialogue response evaluators
Tianyu Zhao, Divesh Lala, and Tatsuya Kawahara. 2020 · 2020
Cited alongside, same era.
The (un)suitability of automatic evaluation metrics for text simplification
Tigerscore: Towards building explainable metric for all text generation tasks
Dongfu Jiang, Yishan Li, Ge Zhang, Wenhao Huang, Bill Yuchen Lin, and Wenhu Chen. 2023 · 2023
Later among the works it cites.
Pei Ke, Bosi Wen, Zhuoer Feng, Xiao Liu, Xuanyu Lei, Jiale Cheng, Shengyuan Wang, Aohan Zeng, Yuxiao Dong, Hongning Wang, Jie Tang, and Minlie Huang. 2023 · 2023
Later among the works it cites.
Prometheus: Inducing fine-grained evaluation capability in language models
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, and Minjoon Seo. 2023 · 2023
Later among the works it cites.
GEMBA-MQM: detecting translation quality error spans with GPT-4
Tom Kocmi and Christian Federmann. 2023a · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fernando Alva-Manchego, Carolina Scarton, and Lucia Specia. 2021 · 2021
Cited alongside, same era.
Summeval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir R. Radev. 2021 · 2021
Cited alongside, same era.
Experts, errors, and context: A large-scale study of human evaluation for machine translation
Markus Freitag, George F. Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021 · 2021
Cited alongside, same era.
Openmeva: A benchmark for evaluating open-ended story generation metrics
Jian Guan, Zhexin Zhang, Zhuoer Feng, Zitao Liu, Wenbiao Ding, Xiaoxi Mao, Changjie Fan, and Minlie Huang. 2021 · 2021
Cited alongside, same era.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
Artidoro Pagnoni, Vidhisha Balachandran, and Yulia Tsvetkov. 2021 · 2021
Cited alongside, same era.
Improving human text simplification with sentence fusion
Max Schwarzer, Teerapaun Tanprasert, and David Kauchak. 2021 · 2021
Cited alongside, same era.
Rethinking automatic evaluation in sentence simplification
Thomas Scialom, Louis Martin, Jacopo Staiano, Éric Villemonte de la Clergerie, and Benoît Sagot. 2021 · 2021
Cited alongside, same era.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann. 2023b · 2023
Later among the works it cites.
The eval4nlp 2023 shared task on prompting large language models as explainable metrics
Christoph Leiter, Juri Opitz, Daniel Deutsch, Yang Gao, Rotem Dror, and Steffen Eger. 2023 · 2023
Later among the works it cites.
G-eval: NLG evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023b · 2023
Later among the works it cites.
Qingyu Lu, Baopu Qiu, Liang Ding, Liping Xie, and Dacheng Tao. 2023 · 2023
Later among the works it cites.
LENS: A learnable evaluation metric for text simplification
Mounica Maddela, Yao Dou, David Heineman, and Wei Xu. 2023 · 2023
Later among the works it cites.
Summarization is (almost) dead
Xiao Pu, Mingqi Gao, and Xiaojun Wan. 2023 · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023 · 2023
Later among the works it cites.
Opinsummeval: Revisiting automated evaluation for opinion summarization
Yuchen Shen and Xiaojun Wan. 2023 · 2023
Later among the works it cites.
Evaluation metrics in the era of GPT-4: reliably evaluating large language models on sequence to sequence tasks
Andrea Sottana, Bin Liang, Kai Zou, and Zheng Yuan. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurélien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023c · 2023
Later among the works it cites.
Large language models are diverse role-players for summarization evaluation
Ning Wu, Ming Gong, Linjun Shou, Shining Liang, and Daxin Jiang. 2023 · 2023
Later among the works it cites.
The next chapter: A study of large language models in storytelling
Zhuohan Xie, Trevor Cohn, and Jey Han Lau. 2023 · 2023
Later among the works it cites.
INSTRUCTSCORE: towards explainable text generation evaluation with automatic feedback
Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Wang, and Lei Li. 2023 · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 · 2023
Later among the works it cites.
LIMA: less is more for alignment
Chunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy. 2023 · 2023
Later among the works it cites.
Judgelm: Fine-tuned large language models are scalable judges
Lianghui Zhu, Xinggang Wang, and Xinlong Wang. 2023 · 2023
Later among the works it cites.
Llm-based NLG evaluation: Current status and challenges
Mingqi Gao, Xinyu Hu, Jie Ruan, Xiao Pu, and Xiaojun Wan. 2024 · 2024
Closest in time.
Are llm-based evaluators confusing NLG quality criteria?
Xinyu Hu, Mingqi Gao, Sen Hu, Yang Zhang, Yicheng Chen, Teng Xu, and Xiaojun Wan. 2024 · 2024
Closest in time.
Calibrating llm-based evaluator
Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang. 2024a · 2024
Closest in time.
Calibrating llm-based evaluator
Yuxuan Liu, Tianchi Yang, Shaohan Huang, Zihan Zhang, Haizhen Huang, Furu Wei, Weiwei Deng, Feng Sun, and Qi Zhang. 2024b · 2024
Closest in time.
LLM comparative assessment: Zero-shot NLG evaluation through pairwise comparisons using large language models
Adian Liusie, Potsawee Manakul, and Mark J. F. Gales. 2024 · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
Meta. 2024 · 2024
Closest in time.
One prompt to rule them all: Llms for opinion summary evaluation
Tejpalsingh Siledar, Swaroop Nath, Sankara Sri Raghava Ravindra Muddu, Rupasai Rangaraju, Swaprava Nath, Pushpak Bhattacharyya, Suman Banerjee, Amey Patil, Sudhanshu Shekhar Singh, Muthusamy Chelliah, and Nikesh Garera. 2024 · 2024
Closest in time.
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022 · 2038
Closest in time.