Fetching the paper…
Reading the bibliography…
Traditional automated metrics for evaluating conditional natural language generation use pairwise comparisons between a single generated text and the best-matching gold-standard ground truth text.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 1908
Earlier work this paper cites.
Semantic distance and semantic judgment
R. Durga · 1980
Earlier work this paper cites.
Statistical methods for research workers
R. A. Fisher · 1992
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Meteor, m-bleu and m-ter: Evaluation metrics for high-correlation with human rankings of machine translation output
A. Agarwal and A. Lavie · 2008
Earlier work this paper cites.
A triangle test for equality of distribution functions in high dimensions
Z. Liu and R. Modarres · 2011
Earlier work this paper cites.
Divergence measures for statistical data processing—an annotated bibliography
M. Basseville · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
A note on the evaluation of generative models
L. Theis, A. v. d. Oord, and M. Bethge · 2015
Earlier work this paper cites.
Locally-connected transformations for deep gmms
A. van den Oord and J. Dambre · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
P. Anderson, B. Fernando, M. Johnson, and S. Gould · 2016
Earlier work this paper cites.
C.-W. Liu, R. Lowe, I. V. Serban, M. Noseworthy, L. Charlin, and J. Pineau · 2016
Earlier work this paper cites.
Improved techniques for training gans
T. Salimans, I. Goodfellow, W. Zaremba, V. Cheung, A. Radford, and X. Chen · 2016
Cited alongside, same era.
Diverse beam search: Decoding diverse solutions from neural sequence models
A. K. Vijayakumar, M. Cogswell, R. R. Selvaraju, Q. Sun, S. Lee, D. Crandall, and D. Batra · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Cited alongside, same era.
Enriching word vectors with subword information
P. Bojanowski, E. Grave, A. Joulin, and T. Mikolov · 2017
Cited alongside, same era.
Mmd gan: Towards deeper understanding of moment matching network
C.-L. Li, W.-C. Chang, Y. Cheng, Y. Yang, and B. Póczos · 2017
Cited alongside, same era.
Tvt: Two-view transformer network for video captioning
Gpt-3: Its nature, scope, limits, and consequences
L. Floridi and M. Chiriatti · 2020
Later among the works it cites.
Bleurt: Learning robust metrics for text generation
T. Sellam, D. Das, and A. P. Parikh · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with bert
T. Zhang*, V. Kishore*, F. Wu*, K. Q. Weinberger, and Y. Artzi · 2020
Later among the works it cites.
Unified vision-language pre-training for image captioning and vqa
L. Zhou, H. Palangi, L. Zhang, H. Hu, J. Corso, and J. Gao · 2020
Later among the works it cites.
CLIPScore: a reference-free evaluation metric for image captioning
J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y. Choi · 2021
Later among the works it cites.
O2na: An object-oriented non-autoregressive approach for controllable video captioning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Chen, Y. Li, Z. Zhang, and S. Huang · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Cited alongside, same era.
Ruse: Regressor using sentence embeddings for automatic machine translation evaluation
H. Shimanaka, T. Kajiwara, and M. Komachi · 2018
Cited alongside, same era.
Texygen: A benchmarking platform for text generation models
Y. Zhu, S. Lu, L. Zheng, J. Guo, W. Zhang, J. Wang, and Y. Yu · 2018
Cited alongside, same era.
Sentence mover’s similarity: Automatic evaluation for multi-sentence texts
E. Clark, A. Celikyilmaz, and N. A. Smith · 2019
Cited alongside, same era.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2019
Cited alongside, same era.
F. Liu, X. Ren, X. Wu, B. Yang, S. Ge, and X. Sun · 2021
Later among the works it cites.
Clipcap: Clip prefix for image captioning
R. Mokady, A. Hertz, and A. H. Bermano · 2021
Later among the works it cites.
Improving video captioning with temporal composition of a visual-syntactic embedding
J. Perez-Martin, B. Bustos, and J. Pérez · 2021
Later among the works it cites.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
K. Pillutla, S. Swayamdipta, R. Zellers, J. Thickstun, S. Welleck, Y. Choi, and Z. Harchaoui · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Later among the works it cites.
A comprehensive assessment of dialog evaluation metrics
Y.-T. Yeh, M. Eskenazi, and S. Mehri · 2021
Later among the works it cites.
Retrieval-augmented diffusion models
A. Blattmann, R. Rombach, K. Oktay, and B. Ommer · 2022
Closest in time.
D. M. Chan, A. Myers, S. Vijayanarasimhan, D. A. Ross, B. Seybold, and J. F. Canny · 2022
Closest in time.
Survey of hallucination in natural language generation
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. Bang, A. Madotto, and P. Fung · 2022
Closest in time.
Hierarchical text-conditional image generation with clip latents
A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen · 2022
Closest in time.
Lamda: Language models for dialog applications
R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H.-T. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, et al · 2022
Closest in time.