Fetching the paper…
Reading the bibliography…
Remote sensing image change captioning (RSICC) aims to automatically generate sentences that describe content differences in remote sensing bitemporal images.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics , 2002, pp. 311–318
2002
Earlier work this paper cites.
L. C. Rouge, “A package for automatic evaluation of summaries,” in Proceedings of Workshop on Text Summarization of ACL, Spain , vol. 5, 2004
2004
Earlier work this paper cites.
S. Banerjee, A. Lavie et al. , “An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the ACL-2005 Workshop on Intrinsic and Extrinsic Evaluation Measures for MT and/or Summarization , 2005, pp. 65–72
2005
Earlier work this paper cites.
A. Graves and A. Graves, “Long short-term memory,” Supervised sequence labelling with recurrent neural networks , pp. 37–45, 2012
2012
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term recurrent convolutional networks for visual recognition and description,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 2625–2634
2015
Earlier work this paper cites.
R. Vedantam, C. Lawrence Zitnick, and D. Parikh, “Cider: Consensus-based image description evaluation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 4566–4575
2015
Earlier work this paper cites.
J. Lu, C. Xiong, D. Parikh, and R. Socher, “Knowing when to look: Adaptive attention via a visual sentinel for image captioning,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 375–383
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 6077–6086
2018
Earlier work this paper cites.
D. H. Park, L. A. Hendricks, Z. Akata, A. Rohrbach, B. Schiele, T. Darrell, and M. Rohrbach, “Multimodal explanations: Justifying decisions and pointing to the evidence,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 8779–8788
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. H. Park, T. Darrell, and A. Rohrbach, “Robust change captioning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 4624–4633
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Z. Huang, X. Wang, L. Huang, C. Huang, Y. Wei, and W. Liu, “Ccnet: Criss-cross attention for semantic segmentation,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 603–612
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2020
Cited alongside, same era.
2020
Cited alongside, same era.
C. Liu, R. Zhao, H. Chen, Z. Zou, and Z. Shi, “Remote sensing image change captioning with dual-branch transformers: A new method and a large scale dataset,” IEEE Transactions on Geoscience and Remote Sensing , vol. 60, pp. 1–20, 2022
2022
Later among the works it cites.
Z. Pan, J. Cai, and B. Zhuang, “Fast vision transformers with hilo attention,” Advances in Neural Information Processing Systems , vol. 35, pp. 14 541–14 554, 2022
2022
Later among the works it cites.
X. Dong, J. Bao, D. Chen, W. Zhang, N. Yu, L. Yuan, D. Chen, and B. Guo, “Cswin transformer: A general vision transformer backbone with cross-shaped windows,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 12 124–12 134
2022
Later among the works it cites.
J. Fang, L. Xie, X. Wang, X. Zhang, W. Liu, and Q. Tian, “Msg-transformer: Exchanging local spatial information by manipulating messenger tokens,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 12 063–12 072
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Huang, J. Chen, W. Ouyang, W. Wan, and Y. Xue, “Image captioning with end-to-end attribute detection and subsequent attributes prediction,” IEEE Transactions on Image processing , vol. 29, pp. 4013–4026, 2020
2020
Cited alongside, same era.
S. Chouaf, G. Hoxha, Y. Smara, and F. Melgani, “Captioning changes in bi-temporal remote sensing images,” in 2021 IEEE International Geoscience and Remote Sensing Symposium IGARSS . IEEE, 2021, pp. 2891–2894
2021
Cited alongside, same era.
B. Graham, A. El-Nouby, H. Touvron, P. Stock, A. Joulin, H. Jégou, and M. Douze, “Levit: a vision transformer in convnet’s clothing for faster inference,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 12 259–12 269
2021
Cited alongside, same era.
2021
Cited alongside, same era.
X. Chu, Z. Tian, Y. Wang, B. Zhang, H. Ren, X. Wei, H. Xia, and C. Shen, “Twins: Revisiting the design of spatial attention in vision transformers,” Advances in neural information processing systems , vol. 34, pp. 9355–9366, 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
C.-F. R. Chen, Q. Fan, and R. Panda, “Crossvit: Cross-attention multi-scale vision transformer for image classification,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 357–366
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 10 012–10 022
2021
Cited alongside, same era.
2022
Later among the works it cites.
M. Stefanini, M. Cornia, L. Baraldi, S. Cascianelli, G. Fiameni, and R. Cucchiara, “From show to tell: A survey on deep learning-based image captioning,” IEEE transactions on pattern analysis and machine intelligence , vol. 45, no. 1, pp. 539–559, 2022
2022
Later among the works it cites.
J. Ji, Y. Ma, X. Sun, Y. Zhou, Y. Wu, and R. Ji, “Knowing what to learn: a metric-oriented focal mechanism for image captioning,” IEEE Transactions on Image Processing , vol. 31, pp. 4321–4335, 2022
2022
Later among the works it cites.
K. E. Ak, Y. Sun, and J. H. Lim, “Learning by imagination: A joint framework for text-based image manipulation and change captioning,” IEEE Transactions on Multimedia , vol. 25, pp. 3006–3016, 2022
2022
Later among the works it cites.
S. Chang and P. Ghamisi, “Changes to captions: An attentive network for remote sensing change captioning,” IEEE Transactions on Image Processing , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
D. Han, X. Pan, Y. Han, S. Song, and G. Huang, “Flatten transformer: Vision transformer using focused linear attention,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 5961–5971
2023
Later among the works it cites.
C. Liu, J. Yang, Z. Qi, Z. Zou, and Z. Shi, “Progressive scale-aware network for remote sensing image change captioning,” in IGARSS 2023-2023 IEEE International Geoscience and Remote Sensing Symposium . IEEE, 2023, pp. 6668–6671
2023
Later among the works it cites.
C. Liu, R. Zhao, J. Chen, Z. Qi, Z. Zou, and Z. Shi, “A decoupling paradigm with prompt learning for remote sensing image change captioning,” IEEE Transactions on Geoscience and Remote Sensing , 2023
2023
Later among the works it cites.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in International conference on machine learning . PMLR, 2015, pp. 2048–2057
2057
Closest in time.