Fetching the paper…
Reading the bibliography…
Recently, there has been a growing interest in designing text generation systems from a discourse coherence perspective, e.g., modeling the interdependence between sentences.
On faithfulness and factuality in abstractive summarization
Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan Thomas Mcdonald. 2020 · 1919
Earlier work this paper cites.
Better summarization evaluation with word embeddings for ROUGE
Jun-Ping Ng and Viktoria Abrecht. 2015 · 1930
Earlier work this paper cites.
The design of experiments
Ronald Aylmer Fisher et al. 1937 · 1937
Earlier work this paper cites.
The representation and use of focus in a system for understanding dialogs
Barbara J Grosz et al. 1977 · 1977
Earlier work this paper cites.
Rhetorical structure theory: Toward a functional theory of text organization
William C Mann and Sandra A Thompson. 1988 · 1988
Earlier work this paper cites.
Centering: A framework for modeling the local coherence of discourse
Barbara J. Grosz, Aravind K. Joshi, and Scott Weinstein. 1995 · 1995
Earlier work this paper cites.
BLEU: A Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with high levels of correlation with human judgments
Alon Lavie and Abhaya Agarwal. 2007 · 2007
Earlier work this paper cites.
Modeling local coherence: An entity-based approach
Regina Barzilay and Mirella Lapata. 2008 · 2008
Earlier work this paper cites.
Revisiting readability: A unified framework for predicting text quality
Emily Pitler and Ani Nenkova. 2008 · 2008
Earlier work this paper cites.
Document-level automatic MT evaluation based on discourse representations
Elisabet Comelles, Jesús Giménez, Lluís Màrquez, Irene Castellón, and Victoria Arranz. 2010 · 2010
Earlier work this paper cites.
Extending the entity grid with entity-specific features
Micha Elsner and Eugene Charniak. 2011 · 2011
Earlier work this paper cites.
Automatically evaluating text coherence using discourse relations
Ziheng Lin, Hwee Tou Ng, and Min-Yen Kan. 2011 · 2011
Earlier work this paper cites.
Extending machine translation evaluation metrics with lexical cohesion to document level
Billy T. M. Wong and Chunyu Kit. 2012 · 2012
Earlier work this paper cites.
Analysing lexical consistency in translation
Liane Guillou. 2013 · 2013
Earlier work this paper cites.
Graph-based local coherence modeling
Camille Guinaudeau and Michael Strube. 2013 · 2013
Earlier work this paper cites.
Using discourse structure improves machine translation evaluation
Francisco Guzmán, Shafiq Joty, Lluís Màrquez, and Preslav Nakov. 2014 · 2014
Earlier work this paper cites.
Dependency-based word embeddings
Omer Levy and Yoav Goldberg. 2014 · 2014
Earlier work this paper cites.
Document-level machine translation evaluation with gist consistency and text cohesion
Zhengxian Gong, Min Zhang, and Guodong Zhou. 2015 · 2015
Earlier work this paper cites.
Pronoun-focused MT and cross-lingual pronoun prediction: Findings of the 2015 DiscoMT shared task on pronoun translation
Christian Hardmeier, Preslav Nakov, Sara Stymne, Jörg Tiedemann, Yannick Versley, and Mauro Cettolo. 2015 · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
From word embeddings to document distances
Matt Kusner, Yu Sun, Nicholas Kolkin, and Kilian Weinberger. 2015 · 2015
Earlier work this paper cites.
Lexical coherence graph modeling using word embeddings
Mohsen Mesgar and Michael Strube. 2016 · 2016
Earlier work this paper cites.
Discourse structure in machine translation evaluation
Shafiq Joty, Francisco Guzmán, Lluís Màrquez, and Preslav Nakov. 2017 · 2017
Earlier work this paper cites.
Learning to score system summaries for better content selection evaluation
Maxime Peyrard, Teresa Botschen, and Iryna Gurevych. 2017 · 2017
Cited alongside, same era.
A neural local coherence model
Dat Tien Nguyen and Shafiq Joty. 2017 · 2017
Cited alongside, same era.
Evaluating discourse phenomena in neural machine translation
Rachel Bawden, Rico Sennrich, Alexandra Birch, and Barry Haddow. 2018 · 2018
Cited alongside, same era.
Machine Translation Evaluation beyond the Sentence Level . Alicante, Spain
Bruno Cartoni, Jindřich Libovický, and Thomas Brovelli (Meyer), editors. 2018 · 2018
Cited alongside, same era.
Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies
Max Grusky, Mor Naaman, and Yoav Artzi. 2018 · 2018
Cited alongside, same era.
A pronoun test suite evaluation of the English–German MT systems at WMT 2018
Liane Guillou, Christian Hardmeier, Ekaterina Lapshinova-Koltunski, and Sharid Loáiciga. 2018 · 2018
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Later among the works it cites.
Using context in neural machine translation training objectives
Danielle Saunders, Felix Stahlberg, and Bill Byrne. 2020 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
Brian Thompson and Matt Post. 2020 · 2020
Later among the works it cites.
Discourse-aware neural extractive text summarization
Jiacheng Xu, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Discourse coherence in the wild: A dataset, evaluation and methods
Alice Lai and Joel Tetreault. 2018 · 2018
Cited alongside, same era.
A neural local coherence model for text quality assessment
Mohsen Mesgar and Michael Strube. 2018 · 2018
Cited alongside, same era.
Document-level neural machine translation with hierarchical attention networks
Lesly Miculicich, Dhananjay Ram, Nikolaos Pappas, and James Henderson. 2018 · 2018
Cited alongside, same era.
Context-aware neural machine translation learns anaphora resolution
Elena Voita, Pavel Serdyukov, Rico Sennrich, and Ivan Titov. 2018 · 2018
Cited alongside, same era.
Evaluation benchmarks and learning criteria for discourse-aware sentence representations
Mingda Chen, Zewei Chu, and Kevin Gimpel. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
On the limitations of cross-lingual encoders as exposed by reference-free machine translation evaluation
Wei Zhao, Goran Glavaš, Maxime Peyrard, Yang Gao, Robert West, and Steffen Eger. 2020 · 2020
Later among the works it cites.
GRUEN for evaluating linguistic quality of generated text
Wanzheng Zhu and Suma Bhat. 2020 · 2020
Later among the works it cites.
Is incoherence surprising? targeted evaluation of coherence prediction from language models
Anne Beyer, Sharid Loáiciga, and David Schlangen. 2021 · 2021
Later among the works it cites.
A training-free and reference-free summarization evaluation metric via centrality-weighted relevance and self-referenced redundancy
Wang Chen, Piji Li, and Irwin King. 2021 · 2021
Later among the works it cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Later among the works it cites.
The Eval4NLP shared task on explainable quality estimation: Overview and results
Marina Fomicheva, Piyawat Lertvittayakumjorn, Wei Zhao, Steffen Eger, and Yang Gao. 2021 · 2021
Later among the works it cites.
Blond: An automatic evaluation metric for document-level machinetranslation
Yuchen Jiang, Shuming Ma, Dongdong Zhang, Jian Yang, Haoyang Huang, and Ming Zhou. 2021 · 2021
Later among the works it cites.
Global explainability of BERT-based evaluation metrics by disentangling along linguistic factors
Marvin Kaster, Wei Zhao, and Steffen Eger. 2021 · 2021
Later among the works it cites.
To ship or not to ship: An extensive evaluation of automatic metrics for machine translation
Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021 · 2021
Later among the works it cites.
Discourse probing of pretrained language models
Fajri Koto, Jey Han Lau, and Timothy Baldwin. 2021 · 2021
Later among the works it cites.
Can transformer models measure coherence in text: Re-thinking the shuffle test
Philippe Laban, Luke Dai, Lucas Bandarkar, and Marti A. Hearst. 2021 · 2021
Later among the works it cites.
Coreference-aware dialogue summarization
Zhengyuan Liu, Ke Shi, and Nancy Chen. 2021 · 2021
Later among the works it cites.
A survey on document-level neural machine translation: Methods and evaluation
Sameen Maruf, Fahimeh Saleh, and Gholamreza Haffari. 2021 · 2021
Later among the works it cites.
A neural graph-based local coherence model
Mohsen Mesgar, Leonardo F. R. Ribeiro, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
Better than average: Paired evaluation of NLP systems
Maxime Peyrard, Wei Zhao, Steffen Eger, and Robert West. 2021 · 2021
Later among the works it cites.
Learning compact metrics for MT
Amy Pu, Hyung Won Chung, Ankur Parikh, Sebastian Gehrmann, and Thibault Sellam. 2021 · 2021
Later among the works it cites.
Perturbation CheckLists for evaluating NLG evaluation metrics
Ananya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, and Mitesh M. Khapra. 2021 · 2021
Later among the works it cites.
BARTScore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
Julius Steen and Katja Markert. 2022 · 2022
Closest in time.