Fetching the paper…
Reading the bibliography…
Image captioning models have been able to generate grammatically correct and human understandable sentences.
Bleu: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2001
Earlier work this paper cites.
ROUGE: A Package For Automatic Evaluation Of Summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results
M. Everingham, L. Van Gool, C. K. I. Williams, J. Winn, and A. Zisserman · 2012
Earlier work this paper cites.
Distributed Representations of Words and Phrases and their Compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Deep Visual-Semantic Alignments for Generating Image Descriptions
Andrej Karpathy and Li Feifei · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, Lubomir D. Bourdev, Ross B. Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár · 2014
Earlier work this paper cites.
CIDEr: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh · 2014
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Jeff Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C. Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
DeepDiary: Automatic Caption Generation for Lifelogging Image Streams
Chenyou Fan and David J. Crandall · 2016
Earlier work this paper cites.
SSD: Single Shot MultiBox Detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott E. Reed, Cheng-Yang Fu, and Alexander C. Berg · 2016
Earlier work this paper cites.
Knowing When to Look: Adaptive Attention via a Visual Sentinel for Image Captioning
Jiasen Lu, Caiming Xiong, Devi Parikh, and Richard Socher · 2016
Earlier work this paper cites.
YOLO9000: Better, Faster, Stronger
Joseph Redmon and Ali Farhadi · 2016
Cited alongside, same era.
Captioning Images with Diverse Objects
Subhashini Venugopalan, Lisa Anne Hendricks, Marcus Rohrbach, Raymond J. Mooney, Trevor Darrell, and Kate Saenko · 2016
Cited alongside, same era.
Image Captioning with Semantic Attention
Quanzeng You, Hailin Jin, Zhaowen Wang, Chen Fang, and Jiebo Luo · 2016
Cited alongside, same era.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2017
Cited alongside, same era.
Word Translation Without Parallel Data
Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou · 2017
Cited alongside, same era.
Image Caption with Global-Local Attention
A Survey on Deep Transfer Learning
Chuanqi Tan, Fuchun Sun, Tao Kong, Wenchang Zhang, Chao Yang, and Chunfang Liu · 2018
Later among the works it cites.
Object Counts! Bringing Explicit Detections Back into Image Captioning
Josiah Wang, Pranava Swaroop Madhyastha, and Lucia Specia · 2018
Later among the works it cites.
Decoupled Novel Object Captioner
Yuehua Wu, Linchao Zhu, Lu Jiang, and Yi Yang · 2018
Later among the works it cites.
nocaps: novel object captioning at scale
Harsh Agrawal, Karan Desai, Yufei Wang, Rishabh Jain, Xinlei Chen, Dhruv Batra, Mark S Johnson, Devi Parikh, Peter Anderson, and Stefan Lee · 2019
Later among the works it cites.
Image Captioning with Unseen Objects
Berkan Demirel, Ramazan Gokberk Cinbis, and Nazli Ikizler-Cinbis · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Linghui Li, Sheng Tang, Lixi Deng, Yongdong Zhang, and Qi Tian · 2017
Cited alongside, same era.
Count-ception: Counting by Fully Convolutional Redundant Counting
Joseph Paul Cohen, Genevieve Boucher, Craig A. Glastonbury, Henry Z. Lo, and Yoshua Bengio · 2017
Cited alongside, same era.
Deep Reinforcement Learning-Based Image Captioning with Embedding Reward
Zhou Ren, Xiaoyu Wang, Ning Zhang, Xutao Lv, and Li-Jia Li · 2017
Cited alongside, same era.
Context-Aware Captions from Context-Agnostic Supervision
Ramakrishna Vedantam, Samy Bengio, Kevin Murphy, Devi Parikh, and Gal Chechik · 2017
Cited alongside, same era.
Unsupervised Image Captioning
Yang Feng, Lin Ma, Wei Liu, and Jiebo Luo · 2018
Cited alongside, same era.
VizWiz Grand Challenge: Answering Visual Questions from Blind People
Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham · 2018
Cited alongside, same era.
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper R. R. Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Tom Duerig, and Vittorio Ferrari · 2018
Cited alongside, same era.
CenterNet: Keypoint Triplets for Object Detection
Kaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi, Qingming Huang, and Qi Tian · 2019
Later among the works it cites.
Leveraging Auxiliary Text for Deep Recognition of Unseen Visual Relationships
Gal Sadeh Kenigsfield and Ran El-Yaniv · 2019
Later among the works it cites.
End-to-End Video Captioning
Silvio Olivastri, Gurkirt Singh, and Fabio Cuzzolin · 2019
Later among the works it cites.
Taking a HINT: Leveraging Explanations to Make Vision and Language Models More Grounded
Ramprasaath R. Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin, Dhruv Batra, and Devi Parikh · 2019
Later among the works it cites.
Active Object Manipulation Facilitates Visual Object Learning: An Egocentric Vision Study
Satoshi Tsutsui, Dian Zhi, Md Alimoor Reza, David J. Crandall, and Chen Yu · 2019
Later among the works it cites.
Navigation Agents for the Visually Impaired: A Sidewalk Simulator and Experiments
Martin Weiss, Simón Chamorro, Roger Girgis, Margaux Luck, Samira Ebrahimi Kahou, Joseph Paul Cohen, Derek Nowrouzezahrai, Doina Precup, Florian Golemo, and Chris Pal · 2019
Later among the works it cites.
Visual-Semantic Graph Attention Network for Human-Object Interaction Detection
Zhijun Liang, Yi-Sheng Guan, and Juan Rojas · 2020
Closest in time.
EGO-TOPO: Environment Affordances from Egocentric Video
Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, and Kristen Grauman · 2020
Closest in time.
Interaction Graphs for Object Importance Estimation in On-road Driving Videos
Zehua Zhang, Ashish Tawari, Sujitha Martin, and David J. Crandall · 2020
Closest in time.