Fetching the paper…
Reading the bibliography…
A proper evaluation of stories generated for a sequence of images -- the task commonly referred to as visual storytelling -- must consider multiple aspects, such as coherence, grammatical correctness, and visual grounding.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics (extended abstract)
Micah Hodosh, Peter Young, and J. Hockenmaier. 2013 · 2013
Earlier work this paper cites.
Bringing semantics into focus using visual abstraction
C. L. Zitnick and Devi Parikh. 2013 · 2013
Earlier work this paper cites.
Concreteness ratings for 40 thousand generally known english word lemmas
Marc Brysbaert, Amy Beth Warriner, and Victor Kuperman. 2014 · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Visual storytelling
Ting-Hao Kenneth Huang, Francis Ferraro, Nasrin Mostafazadeh, Ishan Misra, Aishwarya Agrawal, Jacob Devlin, Ross Girshick, Xiaodong He, Pushmeet Kohli, Dhruv Batra, C. Lawrence Zitnick, Devi Parikh, Lucy Vanderwende, Michel Galley, and Margaret Mitchell. 2016 · 2016
Cited alongside, same era.
GLAC Net: GLocal Attention Cascading Networks for Multi-image Cued Story Generation
Taehyeong Kim, Min-Oh Heo, Seonil Son, Kyoung-Wha Park, and Byoung-Tak Zhang. 2018 · 2018
Cited alongside, same era.
No metrics are perfect: Adversarial reward learning for visual storytelling
Xin Wang, Wenhu Chen, Yuan-Fang Wang, and William Yang Wang. 2018 · 2018
Cited alongside, same era.
What makes a good story? designing composite rewards for visual storytelling
Junjie Hu, Yu Cheng, Zhe Gan, Jingjing Liu, Jianfeng Gao, and Graham Neubig. 2020 · 2020
Cited alongside, same era.
CLIPScore: A reference-free evaluation metric for image captioning
Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
Cross-lingual and multilingual CLIP
Fredrik Carlsson, Philipp Eisen, Faton Rekathati, and Magnus Sahlgren. 2022 · 2022
Later among the works it cites.
RoViST: Learning robust metrics for visual storytelling
Eileen Wang, Caren Han, and Josiah Poon. 2022 · 2022
Later among the works it cites.
Visual writing prompts: Character-grounded story generation with curated image sequences
Xudong Hong, Asad Sayeed, Khushboo Mehra, Vera Demberg, and Bernt Schiele. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Commonsense knowledge aware concept selection for diverse and informative visual storytelling
Hong Chen, Yifei Huang, Hiroya Takamura, and Hideki Nakayama. 2021 · 2021
Cited alongside, same era.
Aesop: Abstract encoding of stories, objects and pictures
Hareesh Ravi, Kushal Kafle, Scott Cohen, Jonathan Brandt, and Mubbasir Kapadia. 2021 · 2063
Closest in time.