Fetching the paper…
Reading the bibliography…
Natural language processing researchers have identified limitations of evaluation methodology for generation tasks, with new questions raised about the validity of automatic metrics and of crowdworker judgments.
Regression shrinkage and selection via the Lasso
Robert Tibshirani. 1994 · 1994
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Re-evaluating the role of Bleu in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Rich Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
(meta-) evaluation of machine translation
Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2007 · 2007
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007
Earlier work this paper cites.
Further meta-evaluation of machine translation
Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2008 · 2008
Earlier work this paper cites.
Natural Language Processing with Python
Steven Bird, Evan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
Non-expert evaluation of summarization systems is risky
Dan Gillick and Yang Liu. 2010 · 2010
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Results of the WMT14 metrics shared task
Matouš Macháček and Ondřej Bojar. 2014 · 2014
Earlier work this paper cites.
The Stanford CoreNLP natural language processing toolkit
Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J. Bethard, and David McClosky. 2014 · 2014
Earlier work this paper cites.
Fitting linear mixed-effects models using lme4
Douglas Bates, Martin Mächler, Ben Bolker, and Steve Walker. 2015 · 2015
Earlier work this paper cites.
Microsoft COCO captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C. Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomás Kociský, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
SPICE: semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence RNNs and beyond
Ramesh Nallapati, Bowen Zhou, Cícero Nogueira dos Santos, Çaglar Gülçehre, and Bing Xiang. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
CharacTer: Translation edit rate on character level
Weiyue Wang, Jan-Thorsten Peter, Hendrik Rosendahl, and Hermann Ney. 2016 · 2016
Earlier work this paper cites.
chrF++: words helping character n-grams
Maja Popović. 2017 · 2017
Earlier work this paper cites.
Get to the point: Summarization with pointer-generator networks
Abigail See, Peter J. Liu, and Christopher D. Manning. 2017 · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Earlier work this paper cites.
Fast abstractive summarization with reinforce-selected sentence rewriting
Yen-Chun Chen and Mohit Bansal. 2018 · 2018
Earlier work this paper cites.
BanditSum: Extractive summarization as a contextual bandit
Yue Dong, Yikang Shen, Eric Crawford, Herke van Hoof, and Jackie Chi Kit Cheung. 2018 · 2018
Earlier work this paper cites.
Bottom-up abstractive summarization
Sebastian Gehrmann, Yuntian Deng, and Alexander Rush. 2018 · 2018
Earlier work this paper cites.
Soft layer-specific multi-task summarization with entailment and question generation
Han Guo, Ramakanth Pasunuru, and Mohit Bansal. 2018 · 2018
Earlier work this paper cites.
Achieving human parity on automatic Chinese to English news translation
Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mengnan Li, Shujie Liu, Tie-Yan Liu, Renqian Luo, Arul Menezes, Tao Qin, Frank Seide, Xu Tan, Fei Tian, Lijun Wu, Shuangzhi Wu, Yingce Xia, Dongdong Zhang, Zhirui Zhang, and Ming Zhou. 2018 · 2018
Earlier work this paper cites.
A unified model for extractive and abstractive summarization using inconsistency loss
Wan-Ting Hsu, Chieh-Kai Lin, Ming-Ying Lee, Kerui Min, Jing Tang, and Min Sun. 2018 · 2018
Cited alongside, same era.
Closed-book training to improve summarization encoder memory
Yichen Jiang and Mohit Bansal. 2018 · 2018
Cited alongside, same era.
Improving abstraction in text summarization
Wojciech Kryściński, Romain Paulus, Caiming Xiong, and Richard Socher. 2018 · 2018
Cited alongside, same era.
Ranking sentences for extractive summarization with reinforcement learning
Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018 · 2018
Cited alongside, same era.
Multi-reward reinforced summarization with saliency and entailment
Utility is in the eye of the user: A critique of NLP leaderboards
Kawin Ethayarajh and Dan Jurafsky. 2020 · 2020
Later among the works it cites.
BLEU might be guilty but references are not innocent
Markus Freitag, David Grangier, and Isaac Caswell. 2020 · 2020
Later among the works it cites.
The University of Edinburgh’s submission to the German-to-English and English-to-German tracks in the WMT 2020 news translation and zero-shot translation robustness tasks
Ulrich Germann. 2020 · 2020
Later among the works it cites.
Twitter a11y: A browser extension to make twitter images accessible
Cole Gleason, Amy Pavel, Emma McCamey, Christina Low, Patrick Carrington, Kris M. Kitani, and Jeffrey P. Bigham. 2020 · 2020
Later among the works it cites.
The AFRL WMT20 news translation systems
Jeremy Gwinnup and Tim Anderson. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ramakanth Pasunuru and Mohit Bansal. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
Attaining the unattainable? reassessing claims of human parity in neural machine translation
Antonio Toral, Sheila Castilho, Ke Hu, and Andy Way. 2018 · 2018
Cited alongside, same era.
Learning to extract coherent summary via deep reinforcement learning
Yuxiang Wu and Baotian Hu. 2018 · 2018
Cited alongside, same era.
Neural latent extractive document summarization
Xingxing Zhang, Mirella Lapata, Furu Wei, and Ming Zhou. 2018 · 2018
Cited alongside, same era.
Neural document summarization by jointly learning to score and select sentences
Qingyu Zhou, Nan Yang, Furu Wei, Shaohan Huang, Ming Zhou, and Tiejun Zhao. 2018 · 2018
Cited alongside, same era.
Better rewards yield better summaries: Learning to summarise without references
Florian Böhm, Yang Gao, Christian M. Meyer, Ori Shapira, Ido Dagan, and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
STRASS: A light and effective method for extractive summarization based on sentence embeddings
Léo Bouscarrat, Antoine Bonnefoy, Thomas Peel, and Cécile Pereira. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Findings of the fourth workshop on neural generation and translation
Kenneth Heafield, Hiroaki Hayashi, Yusuke Oda, Ioannis Konstas, Andrew Finch, Graham Neubig, Xian Li, and Alexandra Birch. 2020 · 2020
Later among the works it cites.
Tohoku-AIP-NTT at WMT 2020 news translation task
Shun Kiyono, Takumi Ito, Ryuto Konno, Makoto Morishita, and Jun Suzuki. 2020 · 2020
Later among the works it cites.
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Zuchao Li, Hai Zhao, Rui Wang, Kehai Chen, Masao Utiyama, and Eiichiro Sumita. 2020 · 2020
Later among the works it cites.
WeChat neural machine translation systems for WMT20
Fandong Meng, Jianhao Yan, Yijin Liu, Yuan Gao, Xianfeng Zeng, Qinsong Zeng, Peng Li, Ming Chen, Jie Zhou, Sifan Liu, and Hao Zhou. 2020 · 2020
Later among the works it cites.
PROMT systems for WMT 2020 shared news translation task
Alexander Molchanov. 2020 · 2020
Later among the works it cites.
eTranslation’s submissions to the WMT 2020 news translation task
Csaba Oravecz, Katina Bontcheva, László Tihanyi, David Kolovratnik, Bhavani Bhaskar, Adrien Lardilleux, Szymon Klocek, and Andreas Eisele. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur P Parikh. 2020 · 2020
Later among the works it cites.
OPPO’s machine translation systems for WMT20
Tingxun Shi, Shiyu Zhao, Xiaopu Li, Xiaoxue Wang, Qian Zhang, Di Ai, Dawei Dang, Xue Zhengshan, and Jie Hao. 2020 · 2020
Later among the works it cites.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
Brian Thompson and Matt Post. 2020 · 2020
Later among the works it cites.
HW-TSC’s participation in the WMT 2020 news translation shared task
Daimeng Wei, Hengchao Shang, Zhanglin Wu, Zhengzhe Yu, Liangyou Li, Jiaxin Guo, Minghan Wang, Hao Yang, Lizhi Lei, Ying Qin, and Shiliang Sun. 2020 · 2020
Later among the works it cites.
The DeepMind Chinese–English document translation system at WMT2020
Lei Yu, Laurent Sartran, Po-Sen Huang, Wojciech Stokowiec, Domenic Donato, Srivatsan Srinivasan, Alek Andreev, Wang Ling, Sona Mokra, Agustin Dal Lago, Yotam Doron, Susannah Young, Phil Blunsom, and Chris Dyer. 2020 · 2020
Later among the works it cites.
Unified vision-language pre-training for image captioning and VQA
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, and Jianfeng Gao. 2020 · 2020
Later among the works it cites.
All that’s ‘human’ is not gold: Evaluating human evaluation of generated text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021 · 2021
Closest in time.
SummEval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev. 2021 · 2021
Closest in time.
Experts, errors, and context: A large-scale study of human evaluation for machine translation
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021 · 2021
Closest in time.
The GEM benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin P. Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Aremu Anuoluwapo, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh D. Dhole, Wanyu Du, Esin Durmus, Ondrej Dusek, Chris Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Rubungo Andre Niyongabo, Salomey Osei, Ankur P. Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou. 2021 · 2021
Closest in time.
CLIPScore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Closest in time.
GENIE: A leaderboard for human-in-the-loop evaluation of text generation
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A. Smith, and Daniel S. Weld. 2021 · 2021
Closest in time.
Scientific credibility of machine translation research: A meta-evaluation of 769 papers
Benjamin Marie, Atsushi Fujita, and Raphael Rubino. 2021 · 2021
Closest in time.
How robust are model rankings : A leaderboard customization approach for equitable evaluation
Swaroop Mishra and Anjana Arunkumar. 2021 · 2021
Closest in time.
Hao Peng, Nikolaos Pappas, Dani Yogatama, Roy Schwartz, Noah A. Smith, and Lingpeng Kong. 2021 · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Closest in time.
VinVL: Making visual representations matter in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao. 2021 · 2021
Closest in time.
Fine-tuning language models from human preferences
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul F. Christiano, and Geoffrey Irving. 2019 · 2021
Closest in time.
Transparent human evaluation for image captioning
Jungo Kasai, Keisuke Sakaguchi, Lavinia Dunagan, Jacob Morrison, Ronan Le Bras, Yejin Choi, and Noah A. Smith. 2022 · 2022
Closest in time.