Fetching the paper…
Reading the bibliography…
We introduce a method called the Expansion mechanism that processes the input unconstrained by the number of elements in the sequence.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
R. Socher and L. Fei-Fei, “Connecting modalities: Semi-supervised segmentation and annotation of images using unaligned text corpora,” in 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition . IEEE, 2010, pp. 966–973
2010
Earlier work this paper cites.
B. Z. Yao, X. Yang, L. Lin, M. W. Lee, and S.-C. Zhu, “I2t: Image parsing to text description,” Proceedings of the IEEE , vol. 98, no. 8, pp. 1485–1508, 2010
2010
Earlier work this paper cites.
M. Mitchell et al. , “Midge: Generating image descriptions from computer vision detections,” in Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics , 2012, pp. 747–756
2012
Earlier work this paper cites.
G. Kulkarni et al. , “Babytalk: Understanding and generating simple image descriptions,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 35, no. 12, pp. 2891–2903, 2013
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
T.-Y. Lin et al. , “Microsoft coco: Common objects in context,” in European conference on computer vision . Springer, 2014, pp. 740–755
2014
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3156–3164
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
A. Karpathy and L. Fei-Fei, “Deep visual-semantic alignments for generating image descriptions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 3128–3137
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani et al. , “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
S. J. Rennie, E. Marcheret, Y. Mroueh, J. Ross, and V. Goel, “Self-critical sequence training for image captioning,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 7008–7024
2017
Earlier work this paper cites.
P. Anderson et al. , “Bottom-up and top-down attention for image captioning and visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2018, pp. 6077–6086
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Yao, Y. Pan, Y. Li, and T. Mei, “Exploring visual relationship for image captioning,” in Proceedings of the European conference on computer vision (ECCV) , 2018, pp. 684–699
2018
Earlier work this paper cites.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
L. Huang, W. Wang, J. Chen, and X.-Y. Wei, “Attention on attention for image captioning,” in Proceedings of the IEEE International Conference on Computer Vision , 2019, pp. 4634–4643
2019
Cited alongside, same era.
Y. Tay, D. Bahri, D. Metzler, D.-C. Juan, Z. Zhao, and C. Zheng, “Synthesizer: Rethinking self-attention for transformer models,” in International conference on machine learning . PMLR, 2021, pp. 10 183–10 192
2021
Later among the works it cites.
2021
Later among the works it cites.
I. O. Tolstikhin et al. , “Mlp-mixer: An all-mlp architecture for vision,” Advances in neural information processing systems , vol. 34, pp. 24 261–24 272, 2021
2021
Later among the works it cites.
J. Ji et al. , “Improving image captioning by leveraging intra-and inter-layer global representation in transformer network,” in Proceedings of the AAAI conference on artificial intelligence , vol. 35, no. 2, 2021, pp. 1655–1663
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
H. Agrawal et al. , “Nocaps: Novel object captioning at scale,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 8948–8957
2019
Cited alongside, same era.
X. Yang, K. Tang, H. Zhang, and J. Cai, “Auto-encoding scene graphs for image captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 10 685–10 694
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y. Pan, T. Yao, Y. Li, and T. Mei, “X-linear attention networks for image captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 10 971–10 980
2020
Cited alongside, same era.
2020
Cited alongside, same era.
H. Ramsauer et al. , “Hopfield networks is all you need,” arXiv preprint arXiv:2008.02217 , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
X. Zhang et al. , “Rstnet: Captioning with adaptive attention on visual and non-visual words,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 15 465–15 474
2021
Later among the works it cites.
S. Changpinyo, P. Sharma, N. Ding, and R. Soricut, “Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 3558–3568
2021
Later among the works it cites.
2021
Later among the works it cites.
2022
Closest in time.
V.-Q. Nguyen, M. Suganuma, and T. Okatani, “Grit: Faster and better image captioning transformer using dual visual features,” in Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part XXXVI . Springer, 2022, pp. 167–184
2022
Closest in time.
P. Wang et al. , “Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,” in International Conference on Machine Learning . PMLR, 2022, pp. 23 318–23 340
2022
Closest in time.
2022
Closest in time.
P. Zeng, H. Zhang, J. Song, and L. Gao, “S2 transformer for image captioning,” in Proceedings of the International Joint Conferences on Artificial Intelligence , vol. 5, 2022
2022
Closest in time.
2022
Closest in time.
X. Hu et al. , “Scaling up vision-language pre-training for image captioning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 17 980–17 989
2022
Closest in time.
2022
Closest in time.
J. C. Hu, R. Cavicchioli, and A. Capotondi, “Exploring the sequence length bottleneck in the transformer for image captioning,” 2022
2022
Closest in time.
J. Hu, R. Cavicchioli, and A. Capotondi, “A request for clarity over the end of sequence token in the self-critical sequence training”,” Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) , vol. 14233 LNCS, p. 39 – 50, 2023
2023
Closest in time.
K. Xu et al. , “Show, attend and tell: Neural image caption generation with visual attention,” in International conference on machine learning , 2015, pp. 2048–2057
2057
Closest in time.