Fetching the paper…
Reading the bibliography…
We introduce GEM, a living benchmark for natural language Generation (NLG), its Evaluation, and Metrics.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020 · 1919
Earlier work this paper cites.
Studies in language behavior: A program of research
Wendell Johnson. 1944 · 1944
Earlier work this paper cites.
A mathematical theory of communication
Claude E Shannon and Warren Weaver. 1963 · 1963
Earlier work this paper cites.
LCSTS: A large scale Chinese short text summarization dataset
Baotian Hu, Qingcai Chen, and Fangze Zhu. 2015 · 1972
Earlier work this paper cites.
The plane with parallel coordinates
Alfred Inselberg. 1985 · 1985
Earlier work this paper cites.
Building natural language generation systems
Ehud Reiter and Robert Dale. 2000 · 2000
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Few-shot natural language generation by rewriting templates
Mihir Kale and Abhinav Rastogi. 2020 · 2004
Earlier work this paper cites.
XGLUE: A new benchmark dataset for cross-lingual pre-training, understanding and generation
Yaobo Liang, Nan Duan, Yeyun Gong, Ning Wu, Fenfei Guo, Weizhen Qi, Ming Gong, Linjun Shou, Daxin Jiang, Guihong Cao, Xiaodong Fan, Bruce Zhang, Rahul Agrawal, Edward Cui, Sining Wei, Taroon Bharti, Ying Qiao, Jiun-Hung Chen, Winnie Wu, Shuguang Liu, Fan Yang, Rangan Majumder, and Ming Zhou. 2020 · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: an automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Viktor Schlegel, Goran Nenadic, and Riza Batista-Navarro. 2020 · 2005
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
How complex is that sentence? a proposed revision of the rosenberg and abbeduto d-level scale
Michael A Covington, Congzhou He, Cati Brown, Lorina Naci, and John Brown. 2006 · 2006
Earlier work this paper cites.
Bringing the people back in: Contesting benchmark machine learning datasets
Emily Denton, Alex Hanna, Razvan Amironesei, Andrew Smart, Hilary Nicole, and Morgan Klaus Scheuerman. 2020 · 2007
Earlier work this paper cites.
SummEval: Re-evaluating summarization evaluation
Alexander R. Fabbri, Wojciech Kryscinski, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir R. Radev. 2020 · 2007
Earlier work this paper cites.
DART: open-domain structured data record to text generation
Dragomir R. Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Nazneen Fatema Rajani, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna, Yangxiaokang Liu, Nadia Irwanto, Jessica Pan, Faiaz Rahman, Ahmad Zaidi, Murori Mutuma, Yasin Tarabar, Ankit Gupta, Tao Yu, Yi Chern Tan, Xi Victoria Lin, Caiming Xiong, and Richard Socher. 2020 · 2007
Earlier work this paper cites.
Towards a decomposable metric for explainable evaluation of text generation from amr
Juri Opitz and Anette Frank. 2020 · 2008
Earlier work this paper cites.
KILT: a benchmark for knowledge intensive language tasks
Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick S. H. Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2020 · 2009
Earlier work this paper cites.
Go figure! A meta evaluation of factuality in summarization
Saadia Gabriel, Asli Celikyilmaz, Rahul Jha, Yejin Choi, and Jianfeng Gao. 2020 · 2010
Earlier work this paper cites.
Automatic analysis of syntactic complexity in second language writing
Xiaofei Lu. 2010 · 2010
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2020 · 2010
Earlier work this paper cites.
The first surface realisation shared task: Overview and evaluation results
Anja Belz, Mike White, Dominic Espinosa, Eric Kow, Deirdre Hogan, and Amanda Stent. 2011 · 2011
Earlier work this paper cites.
On achieving and evaluating language-independence in NLP
Emily M. Bender. 2011 · 2011
Earlier work this paper cites.
GLGE: A new general language generation evaluation benchmark
Dayiheng Liu, Yu Yan, Yeyun Gong, Weizhen Qi, Hang Zhang, Jian Jiao, Weizhu Chen, Jie Fu, Linjun Shou, Ming Gong, Pengcheng Wang, Jiusheng Chen, Daxin Jiang, Jiancheng Lv, Ruofei Zhang, Winnie Wu, Ming Zhou, and Nan Duan. 2020a · 2011
Earlier work this paper cites.
Dynasent: A dynamic benchmark for sentiment analysis
Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela. 2020 · 2012
Earlier work this paper cites.
Chinese poetry generation with recurrent neural networks
Xingxing Zhang and Mirella Lapata. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomás Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
The Ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems
Ryan Lowe, Nissan Pow, Iulian Serban, and Joelle Pineau. 2015 · 2015
Earlier work this paper cites.
Results of the WMT15 metrics shared task
Miloš Stanojević, Amir Kamran, Philipp Koehn, and Ondřej Bojar. 2015 · 2015
Earlier work this paper cites.
Results of the WMT16 metrics shared task
Ondřej Bojar, Yvette Graham, Amir Kamran, and Miloš Stanojević. 2016 · 2016
Earlier work this paper cites.
A context-aware natural language generation dataset for dialogue systems
Ondrej Dušek and Filip Jurcıcek. 2016 · 2016
Earlier work this paper cites.
Neural text generation from structured data with application to the biography domain
Rémi Lebret, David Grangier, and Michael Auli. 2016 · 2016
Earlier work this paper cites.
A diversity-promoting objective function for neural conversation models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016 · 2016
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence RNNs and beyond
Ramesh Nallapati, Bowen Zhou, Cicero dos Santos, Çağlar Gu̇lçehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
A dataset and evaluation metrics for abstractive compression of sentences and short paragraphs
Kristina Toutanova, Chris Brockett, Ke M. Tran, and Saleema Amershi. 2016 · 2016
Earlier work this paper cites.
Optimizing statistical machine translation for text simplification
Wei Xu, Courtney Napoles, Ellie Pavlick, Quanze Chen, and Chris Callison-Burch. 2016 · 2016
Earlier work this paper cites.
Results of the WMT17 metrics shared task
Ondřej Bojar, Yvette Graham, and Amir Kamran. 2017 · 2017
Earlier work this paper cites.
Learning to ask: Neural question generation for reading comprehension
Xinya Du, Junru Shao, and Claire Cardie. 2017 · 2017
Earlier work this paper cites.
The WebNLG challenge: Generating text from RDF data
Claire Gardent, Anastasia Shimorina, Shashi Narayan, and Laura Perez-Beltrachini. 2017 · 2017
Cited alongside, same era.
The E2E dataset: New challenges for end-to-end generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
Analysing data-to-text generation benchmarks
Laura Perez-Beltrachini and Claire Gardent. 2017 · 2017
Cited alongside, same era.
Challenges in data-to-document generation
Sam Wiseman, Stuart Shieber, and Alexander Rush. 2017 · 2017
Cited alongside, same era.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman. 2018 · 2018
Cited alongside, same era.
A discourse-aware attention model for abstractive summarization of long documents
Arman Cohan, Franck Dernoncourt, Doo Soon Kim, Trung Bui, Seokhwan Kim, Walter Chang, and Nazli Goharian. 2018 · 2018
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019a · 2019
Later among the works it cites.
STORIUM: A Dataset and Evaluation Platform for Machine-in-the-Loop Story Generation
Nader Akoury, Shufan Wang, Josh Whiting, Stephen Hood, Nanyun Peng, and Mohit Iyyer. 2020 · 2020
Later among the works it cites.
ASSET: A dataset for tuning and evaluation of sentence simplification models with multiple rewriting transformations
Fernando Alva-Manchego, Louis Martin, Antoine Bordes, Carolina Scarton, Benoît Sagot, and Lucia Specia. 2020 · 2020
Later among the works it cites.
Should all cross-lingual embeddings speak English?
Antonios Anastasopoulos and Graham Neubig. 2020 · 2020
Later among the works it cites.
Disentangling the properties of human evaluation methods: A classification system to support comparability, meta-evaluation and reproducibility testing
Anya Belz, Simon Mille, and David M. Howcroft. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Cited alongside, same era.
Survey of the state of the art in natural language generation: Core tasks, applications and evaluation
Albert Gatt and Emiel Krahmer. 2018 · 2018
Cited alongside, same era.
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2018 · 2018
Cited alongside, same era.
Content selection in deep learning models of summarization
Chris Kedzie, Kathleen McKeown, and Hal Daumé III. 2018 · 2018
Cited alongside, same era.
The narrativeQA reading comprehension challenge
Tomáš Kočiskỳ, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette. 2018 · 2018
Cited alongside, same era.
Visual question generation as dual task of visual question answering
Yikang Li, Nan Duan, Bolei Zhou, Xiao Chu, Wanli Ouyang, Xiaogang Wang, and Ming Zhou. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Esin Durmus, He He, and Mona Diab. 2020 · 2020
Later among the works it cites.
Evaluating the state-of-the-art of end-to-end natural language generation: The E2E NLG challenge
Ondrej Dusek, Jekaterina Novikova, and Verena Rieser. 2020 · 2020
Later among the works it cites.
Utility is in the eye of the user: A critique of NLP leaderboards
Kawin Ethayarajh and Dan Jurafsky. 2020 · 2020
Later among the works it cites.
The 2020 bilingual, bi-directional webnlg+ shared task overview and evaluation results (webnlg+ 2020)
Thiago Castro Ferreira, Claire Gardent, Chris van der Lee, Nikolai Ilinykh, Simon Mille, Diego Moussalem, and Anastasia Shimorina. 2020 · 2020
Later among the works it cites.
Participatory research for low-resourced machine translation: A case study in African languages
∀ \forall , Wilhelmina Nekoto, Vukosi Marivate, Tshinondiwa Matsila, Timi Fasubaa, Taiwo Fagbohungbe, Solomon Oluwole Akinola, Shamsuddeen Muhammad, Salomon Kabongo Kabenamualu, Salomey Osei, Freshia Sackey, Rubungo Andre Niyongabo, Ricky Macharm, Perez Ogayo, Orevaoghene Ahia, Musie Meressa Berhe, Mofetoluwa Adeyemi, Masabata Mokgesi-Selinga, Lawrence Okegbemi, Laura Martinus, Kolawole Tajudeen, Kevin Degila, Kelechi Ogueji, Kathleen Siminyu, Julia Kreutzer, Jason Webster, Jamiil Toure Ali, Jade Abbott, Iroro Orife, Ignatius Ezeani, Idris Abdulkadir Dangana, Herman Kamper, Hady Elsahar, Goodness Duru, Ghollah Kioko, Murhabazi Espoir, Elan van Biljon, Daniel Whitenack, Christopher Onyefuluchi, Chris Chinenye Emezue, Bonaventure F. P. Dossou, Blessing Sibanda, Blessing Bassey, Ayodele Olabiyi, Arshath Ramkilowan, Alp Öktem, Adewale Akinfaderin, and Abdallah Bashir. 2020 · 2020
Later among the works it cites.
BLEU might be guilty but references are not innocent
Markus Freitag, David Grangier, and Isaac Caswell. 2020 · 2020
Later among the works it cites.
Interpretable multi-dataset evaluation for named entity recognition
Jinlan Fu, Pengfei Liu, and Graham Neubig. 2020 · 2020
Later among the works it cites.
Findings of the fourth workshop on neural generation and translation
Kenneth Heafield, Hiroaki Hayashi, Yusuke Oda, Ioannis Konstas, Andrew Finch, Graham Neubig, Xian Li, and Alexandra Birch. 2020 · 2020
Later among the works it cites.
Twenty years of confusion in human evaluation: NLG needs evaluation sheets and standardised definitions
David M. Howcroft, Anya Belz, Miruna-Adriana Clinciu, Dimitra Gkatzia, Sadid A. Hasan, Saad Mahamood, Simon Mille, Emiel van Miltenburg, Sashank Santhanam, and Verena Rieser. 2020 · 2020
Later among the works it cites.
XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Later among the works it cites.
Neural CRF model for sentence alignment in text simplification
Chao Jiang, Mounica Maddela, Wuwei Lan, Yang Zhong, and Wei Xu. 2020 · 2020
Later among the works it cites.
The state and fate of linguistic diversity and inclusion in the NLP world
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. 2020 · 2020
Later among the works it cites.
NUBIA: NeUral based interchangeability assessor for text generation
Hassan Kane, Muhammed Yusuf Kocyigit, Ali Abdalla, Pelkins Ajanoh, and Mohamed Coulibali. 2020 · 2020
Later among the works it cites.
WikiLingua: A new benchmark dataset for cross-lingual abstractive summarization
Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen McKeown. 2020 · 2020
Later among the works it cites.
CommonGen: A constrained text generation challenge for generative commonsense reasoning
Bill Yuchen Lin, Wangchunshu Zhou, Ming Shen, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2020 · 2020
Later among the works it cites.
How can we accelerate progress towards human-like linguistic generalization?
Tal Linzen. 2020 · 2020
Later among the works it cites.
A human evaluation of amr-to-english generation systems
Emma Manning, Shira Wein, and Nathan Schneider. 2020 · 2020
Later among the works it cites.
The third multilingual surface realisation shared task (SR’20): Overview and evaluation results
Simon Mille, Anya Belz, Bernd Bohnet, Thiago Castro Ferreira, Yvette Graham, and Leo Wanner. 2020 · 2020
Later among the works it cites.
AmbigQA: Answering ambiguous open-domain questions
Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
ToTTo: A controlled table-to-text generation dataset
Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das. 2020 · 2020
Later among the works it cites.
XCOPA: A multilingual dataset for causal commonsense reasoning
Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Towards scalable multi-domain conversational agents: The schema-guided dialogue dataset
Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan. 2020 · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Later among the works it cites.
MLSUM: The multilingual summarization corpus
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2020 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2020
Later among the works it cites.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le. 2020 · 2020
Later among the works it cites.
MultiWOZ 2.2 : A dialogue dataset with additional annotation corrections and state tracking baselines
Xiaoxue Zang, Abhinav Rastogi, Srinivas Sunkara, Raghav Gupta, Jianguo Zhang, and Jindong Chen. 2020 · 2020
Later among the works it cites.
PEGASUS: pre-training with extracted gap-sentences for abstractive summarization
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter J. Liu. 2020a · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020b · 2020
Later among the works it cites.
GENIE: A leaderboard for human-in-the-loop evaluation of text generation
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A. Smith, and Daniel S. Weld. 2021 · 2021
Closest in time.
Safeval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Gallinari Patrick, Lamprier Sylvain, Piwowarski Benjamin, Staiano Jacopo, and Wang Alex. 2021 · 2021
Closest in time.
Anastasia Shimorina and Anya Belz. 2021 · 2021
Closest in time.
Personalized machine translation: Predicting translational preferences
Shachar Mirkin and Jean-Luc Meunier. 2015 · 2025
Closest in time.
Data-to-text generation with entity modeling
Ratish Puduppully, Li Dong, and Mirella Lapata. 2019 · 2035
Closest in time.