Fetching the paper…
Reading the bibliography…
The open-ended nature of visual captioning makes it a challenging area for evaluation.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
A survey on bias and fairness in machine learning
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2019 · 1908
Earlier work this paper cites.
Patient knowledge distillation for bert model compression
Siqi Sun, Yu Cheng, Zhe Gan, and Jingjing Liu. 2019 · 1908
Earlier work this paper cites.
Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M Meyer, and Steffen Eger. 2019 · 1909
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Elements of Information Theory
Thomas M Cover. 1999 · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Nubia: Neural based interchangeability assessor for text generation
Hassan Kane, Muhammed Yusuf Kocyigit, Ali Abdalla, Pelkins Ajanoh, and Mohamed Coulibali. 2020 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Data Mining: Practical Machine Learning Tools and Techniques
Ian H. Witten and Eibe Frank. 2005 · 2005
Earlier work this paper cites.
Collecting image annotations using amazon’s mechanical turk
Cyrus Rashtchian, Peter Young, Micah Hodosh, and Julia Hockenmaier. 2010 · 2010
Earlier work this paper cites.
A beam-search decoder for grammatical error correction
Daniel Dahlmeier and Hwee Tou Ng. 2012 · 2012
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier. 2013 · 2013
Earlier work this paper cites.
The CoNLL-2014 shared task on grammatical error correction
Hwee Tou Ng, Siew Mei Wu, Ted Briscoe, Christian Hadiwinoto, Raymond Hendy Susanto, and Christopher Bryant. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Cited alongside, same era.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. 2015 · 2015
Cited alongside, same era.
Towards a standard evaluation method for grammatical error detection and correction
Mariano Felice and Ted Briscoe. 2015 · 2015
Cited alongside, same era.
Human evaluation of grammatical error correction systems
Roman Grundkiewicz, Marcin Junczys-Dowmunt, and Edward Gillian. 2015 · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Reference-less measure of faithfulness for grammatical error correction
Leshem Choshen and Omri Abend. 2018 · 2018
Later among the works it cites.
Learning to evaluate image captioning
Y. Cui, G. Yang, A. Veit, X. Huang, and S. Belongie. 2018 · 2018
Later among the works it cites.
Women also snowboard: Overcoming bias in captioning models
Lisa Anne Hendricks, Kaylee Burns, Kate Saenko, Trevor Darrell, and Anna Rohrbach. 2018 · 2018
Later among the works it cites.
Semstyle: Learning to generate stylised image captions using unaligned text
A. Mathews, L. Xie, and X. He. 2018 · 2018
Later among the works it cites.
Cross entropy of neural language models at infinity—a new bound of the entropy rate
Shuntaro Takahashi and Kumiko Tanaka-Ishii. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ground truth for grammatical error correction metrics
Courtney Napoles, Keisuke Sakaguchi, Matt Post, and Joel Tetreault. 2015 · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh. 2015 · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Cited alongside, same era.
Focused evaluation for image description with binary forced-choice tasks
Micah Hodosh and Julia Hockenmaier. 2016 · 2016
Cited alongside, same era.
There’s no comparison: Reference-less evaluation metrics in grammatical error correction
Courtney Napoles, Keisuke Sakaguchi, and Joel Tetreault. 2016 · 2016
Cited alongside, same era.
Reference-based metrics can be replaced with reference-less metrics in evaluating grammatical error correction systems
Hiroki Asano, Tomoya Mizumoto, and Kentaro Inui. 2017 · 2017
Cited alongside, same era.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
TIGEr: Text-to-image grounding for image caption evaluation
Ming Jiang, Qiuyuan Huang, Lei Zhang, Xin Wang, Pengchuan Zhang, Zhe Gan, Jana Diesner, and Jianfeng Gao. 2019 · 2019
Later among the works it cites.
Lceval: Learned composite metric for caption evaluation
Naeha Sharif, Lyndon White, Mohammed Bennamoun, Wei Liu, and Syed Afaq Ali Shah. 2019 · 2019
Later among the works it cites.
Distilling knowledge learned in bert for text generation
Yen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu, and Jingjing Liu. 2020 · 2020
Later among the works it cites.
Meshed-memory transformer for image captioning
Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, and Rita Cucchiara. 2020 · 2020
Later among the works it cites.
X-linear attention networks for image captioning
Yingwei Pan, Ting Yao, Yehao Li, and Tao Mei. 2020 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
Improving image captioning evaluation by considering inter references variance
Yanzhi Yi, Hangyu Deng, and Jinglu Hu. 2020 · 2020
Later among the works it cites.
Perception score: A learned metric for open-ended text generation evaluation
Jing Gu, Qingyang Wu, and Zhou Yu. 2021 · 2021
Closest in time.
Quality estimation for image captions based on large-scale human evaluations
Tomer Levinboim, Ashish V. Thapliyal, Piyush Sharma, and Radu Soricut. 2021 · 2021
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.