Fetching the paper…
Reading the bibliography…
We introduce Repro, an open-source library which aims at improving the reproducibility and usability of research code.
BLEU: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Meteor Universal: Language Specific Translation Evaluation for Any Target Language
Michael Denkowski and Alon Lavie. 2014 · 2014
Earlier work this paper cites.
AllenNLP: A Deep Semantic Natural Language Processing Platform
Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
A Call for Clarity in Reporting BLEU Scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
Multilingual Constituency Parsing with Self-Attention and Pre-Training
Nikita Kitaev, Steven Cao, and Dan Klein. 2019 · 2019
Earlier work this paper cites.
Text Summarization with Pretrained Encoders
Yang Liu and Mirella Lapata. 2019 · 2019
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Earlier work this paper cites.
Answers Unite! Unsupervised Metrics for Reinforced Summarization Models
Thomas Scialom, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2019 · 2019
Earlier work this paper cites.
MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Earlier work this paper cites.
MOCHA: A Dataset for Training and Evaluating Generative Reading Comprehension Metrics
Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2020 · 2020
Earlier work this paper cites.
SacreROUGE: An Open-Source Library for Using and Developing Summarization Evaluation Metrics
Daniel Deutsch and Dan Roth. 2020 · 2020
Earlier work this paper cites.
RoFT: A Tool for Evaluating Human Detection of Machine-Generated Text
Liam Dugan, Daphne Ippolito, Arun Kirubarajan, and Chris Callison-Burch. 2020 · 2020
Earlier work this paper cites.
FEQA: A Question Answering Evaluation Framework for Faithfulness Assessment in Abstractive Summarization
Esin Durmus, He He, and Mona Diab. 2020 · 2020
Cited alongside, same era.
SUPERT: Towards New Frontiers in Unsupervised Evaluation Metrics for Multi-Document Summarization
Yang Gao, Wei Zhao, and Steffen Eger. 2020 · 2020
Cited alongside, same era.
Evaluating Factuality in Generation with Dependency-level Entailment
Tanya Goyal and Greg Durrett. 2020 · 2020
Cited alongside, same era.
Neural Module Networks for Reasoning over Text
Nitish Gupta, Kevin Lin, Dan Roth, Sameer Singh, and Matt Gardner. 2020 · 2020
Cited alongside, same era.
NUBIA: NeUral Based Interchangeability Assessor for Text Generation
Hassan Kane, Muhammed Yusuf Kocyigit, Ali Abdalla, Pelkins Ajanoh, and Mohamed Coulibali. 2020 · 2020
Cited alongside, same era.
Evaluating the Factual Consistency of Abstractive Text Summarization
Automatic Text Evaluation through the Lens of Wasserstein Barycenters
Pierre Colombo, Guillaume Staerman, Chloé Clavel, and Pablo Piantanida. 2021b · 2021
Later among the works it cites.
Towards Question-Answering as an Automatic Metric for Evaluating the Content Quality of a Summary
Daniel Deutsch, Tania Bedrax-Weiss, and Dan Roth. 2021 · 2021
Later among the works it cites.
GSum: A General Framework for Guided Neural Abstractive Summarization
Zi-Yi Dou, Pengfei Liu, Hiroaki Hayashi, Zhengbao Jiang, and Graham Neubig. 2021 · 2021
Later among the works it cites.
CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Later among the works it cites.
Datasets: A Community Library for Natural Language Processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Br, Simon eis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lys Debut, re, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alex Rush, er, and Thomas Wolf. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wojciech Kryscinski, Bryan McCann, Caiming Xiong, and Richard Socher. 2020 · 2020
Cited alongside, same era.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Cited alongside, same era.
COMET: A Neural Framework for MT Evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
BLEURT: Learning Robust Metrics for Text Generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Cited alongside, same era.
Automatic Machine Translation Evaluation in Many Languages via Zero-Shot Paraphrasing
Brian Thompson and Matt Post. 2020 · 2020
Cited alongside, same era.
Fill in the BLANC: Human-free quality estimation of document summaries
Oleg Vasilyev, Vedant Dharnidharka, and John Bohannon. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-Art Natural Language Processing
Thomas Wolf, Lys Debut, re, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, Alex Rush, and er. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
Asking It All: Generating Contextualized Questions for any Semantic Role
Valentina Pyatkin, Paul Roit, Julian Michael, Yoav Goldberg, Reut Tsarfaty, and Ido Dagan. 2021 · 2021
Later among the works it cites.
QuestEval: Summarization Asks for Fact-based Evaluation
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021 · 2021
Later among the works it cites.
A Pseudo-Metric between Probability Distributions based on Depth-Trimmed Regions
Guillaume Staerman, Pavlo Mozharovskyi, Pierre Colombo, Stéphan Clémençon, and Florence d’Alché Buc. 2021 · 2021
Later among the works it cites.
BARTScore: Evaluating Generated Text as Text Generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021 · 2021
Later among the works it cites.
Finding a Balanced Degree of Automation for Summary Evaluation
Shiyue Zhang and Mohit Bansal. 2021 · 2021
Later among the works it cites.
Large-Scale QA-SRL Parsing
Nicholas FitzGerald, Julian Michael, Luheng He, and Luke Zettlemoyer. 2018 · 2060
Closest in time.
Learning to Capitalize with Character-Level Recurrent Neural Networks: An Empirical Study
Raymond Hendy Susanto, Hai Leong Chieu, and Wei Lu. 2016 · 2095
Closest in time.