Fetching the paper…
Reading the bibliography…
OCR-based image captioning is an important but under-explored task, aiming to generate descriptions containing visual objects and scene text.
P. Banerjee, T. Gokhale, Y. Yang, C. Baral, Weakly supervised relative spatial reasoning for visual question answering, in: ICCV, 2021, pp. 1908–1918
1918
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, W.-J. Zhu, Bleu: a method for automatic evaluation of machine translation, in: ACL, 2002, pp. 311–318
2002
Earlier work this paper cites.
S. Banerjee, A. Lavie, Meteor: An automatic metric for mt evaluation with improved correlation with human judgments, in: ACL workshop, 2005, pp. 65–72
2005
Earlier work this paper cites.
J. Almazán, A. Gordo, A. Fornés, E. Valveny, Word spotting and recognition with embedded attributes, IEEE transactions on pattern analysis and machine intelligence 36 (12) (2014) 2552–2566
2014
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, J. Sun, Faster r-cnn: Towards real-time object detection with region proposal networks, Advances in neural information processing systems 28 (2015)
2015
Earlier work this paper cites.
O. Vinyals, M. Fortunato, N. Jaitly, Pointer networks, Advances in neural information processing systems 28 (2015)
2015
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, D. Parikh, Cider: Consensus-based image description evaluation, in: CVPR, 2015, pp. 4566–4575
2015
Earlier work this paper cites.
D. Kinga, J. B. Adam, et al., A method for stochastic optimization, in: ICLR, Vol. 5, 2015, pp. 1–15
2015
Earlier work this paper cites.
P. Anderson, B. Fernando, M. Johnson, S. Gould, Spice: Semantic propositional image caption evaluation, in: ECCV, Springer, 2016, pp. 382–398
2016
Earlier work this paper cites.
Q. Huang, Y. Liang, J. Wei, Y. Cai, H. Liang, H.-f. Leung, Q. Li, Image difference captioning with instance-level fine-grained feature representation, IEEE Transactions on Multimedia 24 (2021) 2004–2017
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, Attention is all you need, Advances in neural information processing systems 30 (2017)
2017
Earlier work this paper cites.
P. Bojanowski, E. Grave, A. Joulin, T. Mikolov, Enriching word vectors with subword information, Transactions of the association for computational linguistics 5 (2017) 135–146
2017
Earlier work this paper cites.
M. Liao, B. Shi, X. Bai, X. Wang, W. Liu, Textboxes: A fast text detector with a single deep neural network, in: AAAI, Vol. 31, 2017
2017
Earlier work this paper cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, L. Zhang, Bottom-up and top-down attention for image captioning and visual question answering, in: CVPR, 2018, pp. 6077–6086
2018
Earlier work this paper cites.
F. Borisyuk, A. Gordo, V. Sivakumar, Rosetta: Large scale system for text detection and recognition in images, in: ACM SIGKDD, 2018, pp. 71–79
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
L. Huang, W. Wang, J. Chen, X.-Y. Wei, Attention on attention for image captioning, in: ICCV, 2019, pp. 4634–4643
2019
Cited alongside, same era.
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, M. Rohrbach, Towards vqa models that can read, in: CVPR, 2019, pp. 8317–8326
2019
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional transformers for language understanding, in: NAACL, 2019, pp. 4171–4186
2019
Cited alongside, same era.
O. Sidorov, R. Hu, M. Rohrbach, A. Singh, Textcaps: a dataset for image captioning with reading comprehension, in: ECCV, Springer, 2020, pp. 742–758
2020
Cited alongside, same era.
J. Wang, J. Tang, J. Luo, Multimodal attention with image text spatial relationship for ocr-based image captioning, in: ACM Multimedia, 2020, pp. 4337–4345
2020
Z.-X. Jin, M. Z. Shou, F. Zhou, S. Tsutsui, J. Qin, X.-C. Yin, From token to word: Ocr token evolution via contrastive learning and semantic matching for text-vqa, in: ACM Multimedia, 2022, pp. 4564–4572
2022
Later among the works it cites.
W. Zhang, H. Shi, J. Guo, S. Zhang, Q. Cai, J. Li, S. Luo, Y. Zhuang, Magic: Multimodal relational graph adversarial inference for diverse and unpaired text-based image captioning, in: AAAI, Vol. 36, 2022, pp. 3335–3343
2022
Later among the works it cites.
J. Wang, Z. Yang, X. Hu, L. Li, K. Lin, Z. Gan, Z. Liu, C. Liu, L. Wang, Git: A generative image-to-text transformer for vision and language, Trans. Mach. Learn. Res. 2022 (2022) 1–49
2022
Later among the works it cites.
W. Tang, Z. Hu, Z. Song, R. Hong, Ocr-oriented master object for text image captioning, in: ICMR, 2022, pp. 39–43
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
R. Hu, A. Singh, T. Darrell, M. Rohrbach, Iterative answer prediction with pointer-augmented multimodal transformers for textvqa, in: CVPR, 2020, pp. 9992–10002
2020
Cited alongside, same era.
Y. Qiu, Y. Satoh, R. Suzuki, K. Iwata, H. Kataoka, 3d-aware scene change captioning from multiview images, IEEE Robotics and Automation Letters 5 (3) (2020) 4743–4750
2020
Cited alongside, same era.
Z. Wang, R. Bao, Q. Wu, S. Liu, Confidence-aware non-repetitive multimodal transformers for textcaps, in: AAAI, Vol. 35, 2021, pp. 2835–2843
2021
Cited alongside, same era.
J. Wang, J. Tang, M. Yang, X. Bai, J. Luo, Improving ocr-based image captioning by incorporating geometrical relationship, in: CVPR, 2021, pp. 1306–1315
2021
Cited alongside, same era.
Z. Yang, Y. Lu, J. Wang, X. Yin, D. Florencio, L. Wang, C. Zhang, L. Zhang, J. Luo, Tap: Text-aware pre-training for text-vqa and text-caption, in: CVPR, 2021, pp. 8751–8761
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., Learning transferable visual models from natural language supervision, in: ICML, PMLR, 2021, pp. 8748–8763
2021
Cited alongside, same era.
G. Zeng, Y. Zhang, Y. Zhou, X. Yang, Beyond ocr+ vqa: involving ocr into the flow for robust and accurate textvqa, in: ACM Multimedia, 2021, pp. 376–385
2021
Cited alongside, same era.
Y. Liu, W. Wei, D. Peng, X.-L. Mao, Z. He, P. Zhou, Depth-aware and semantic guided relational attention network for visual question answering, IEEE Transactions on Multimedia (2022) 1–14
2022
Later among the works it cites.
B. Wan, W. Jiang, Y.-M. Fang, M. Zhu, Q. Li, Y. Liu, Revisiting image captioning via maximum discrepancy competition, Pattern Recognition 122 (2022) 108358
2022
Later among the works it cites.
R. Zhang, Z. Guo, W. Zhang, K. Li, X. Miao, B. Cui, Y. Qiao, P. Gao, H. Li, Pointclip: Point cloud understanding by clip, in: CVPR, 2022, pp. 8552–8562
2022
Later among the works it cites.
Y. Ma, G. Xu, X. Sun, M. Yan, J. Zhang, R. Ji, X-clip: End-to-end multi-grained contrastive learning for video-text retrieval, in: ACM Multimedia, 2022, pp. 638–647
2022
Later among the works it cites.
2022
Later among the works it cites.
Z. Yang, P. Wang, T. Chu, J. Yang, Human-centric image captioning, Pattern Recognition 126 (2022) 108545
2022
Later among the works it cites.
Q. Wang, H. Deng, X. Wu, Z. Yang, Y. Liu, Y. Wang, G. Hao, Lcm-captioner: A lightweight text-based image captioning method with collaborative mechanism between vision and text, Neural Networks 162 (2023) 318–329
2023
Closest in time.
D. Xu, W. Zhao, Y. Cai, Q. Huang, Zero-textcap: Zero-shot framework for text-based image captioning, in: ACM Multimedia, 2023, pp. 4949–4957
2023
Closest in time.
R. Li, D. Xue, S. Su, X. He, Q. Mao, Y. Zhu, J. Sun, Y. Zhang, Learning depth via leveraging semantics: Self-supervised monocular depth estimation with both implicit and explicit semantic guidance, Pattern Recognition 137 (2023) 109297
2023
Closest in time.
J. Zhou, C. Yang, Y. Zhu, Y. Zhang, Cross-region feature fusion with geometrical relationship for ocr-based image captioning, Neurocomputing 601 (2024) 128197
2024
Closest in time.
Y. Li, J. Ji, X. Sun, Y. Zhou, Y. Luo, R. Ji, M3ixup: A multi-modal data augmentation approach for image captioning, Pattern Recognition 158 (2025) 110941
2025
Closest in time.