Fetching the paper…
Reading the bibliography…
Image captioning models aim at connecting Vision and Language by providing natural language descriptions of input images.
BLEU: a method for automatic evaluation of machine translation. In
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context. In
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description. In
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions. In
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization. In
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
CIDEr: Consensus-based Image Description Evaluation. In
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator. In
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Earlier work this paper cites.
SPICE: Semantic Propositional Image Caption Evaluation. In
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units. In
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Particular object retrieval with integral max-pooling of CNN activations. In
Giorgos Tolias, Ronan Sicre, and Hervé Jégou. 2016 · 2016
Earlier work this paper cites.
Self-critical sequence training for image captioning. In
Steven J Rennie, Etienne Marcheret, Youssef Mroueh, Jarret Ross, and Vaibhava Goel. 2017 · 2017
Earlier work this paper cites.
Attention is all you need. In
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering. In
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Earlier work this paper cites.
Convolutional image captioning. In
Jyoti Aneja, Aditya Deshpande, and Alexander G Schwing. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding. In
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Mixed Precision Training. In
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2018 · 2018
Earlier work this paper cites.
Exploring Visual Relationship for Image Captioning. In
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei. 2018 · 2018
Earlier work this paper cites.
Image Captioning: Transforming Objects into Words. In
Simao Herdade, Armin Kappeler, Kofi Boakye, and Joao Soares. 2019 · 2019
Cited alongside, same era.
Attention on Attention for Image Captioning. In
Lun Huang, Wenmin Wang, Jie Chen, and Xiao-Yong Wei. 2019 · 2019
Cited alongside, same era.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Augmenting Self-Attention with Persistent Memory
Sainbayar Sukhbaatar, Edouard Grave, Guillaume Lample, Herve Jegou, and Armand Joulin. 2019 · 2019
Cited alongside, same era.
Explore and Explain: Self-supervised Navigation and Recounting. In
Learning to Select: A Fully Attentive Approach for Novel Object Captioning. In
Marco Cagrandi, Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, and Rita Cucchiara. 2021 · 2021
Later among the works it cites.
Extracting training data from large language models. In
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Later among the works it cites.
Explaining transformer-based image captioning models: An empirical analysis
Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2021 · 2021
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Later among the works it cites.
Vision Transformer Hashing for Image Retrieval. In
Shiv Ram Dubey, Satish Kumar Singh, and Wei-Ta Chu. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Roberto Bigazzi, Federico Landi, Marcella Cornia, Silvia Cascianelli, Lorenzo Baraldi, and Rita Cucchiara. 2020 · 2020
Cited alongside, same era.
Language models are few-shot learners. In
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Normalized and Geometry-Aware Self-Attention Network for Image Captioning. In
Longteng Guo, Jing Liu, Xinxin Zhu, Peng Yao, Shichen Lu, and Hanqing Lu. 2020 · 2020
Cited alongside, same era.
REALM: Retrieval-Augmented Language Model Pre-Training. In
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020 · 2020
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2020
Cited alongside, same era.
Generalization through Memorization: Nearest Neighbor Language Models. In
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020 · 2020
Cited alongside, same era.
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020 · 2020
Cited alongside, same era.
Training Vision Transformers for Image Retrieval
Alaaeldin El-Nouby, Natalia Neverova, Ivan Laptev, and Hervé Jégou. 2021 · 2021
Later among the works it cites.
Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering. In
Gautier Izacard and Edouard Grave. 2021 · 2021
Later among the works it cites.
Working Memory Connections for LSTM
Federico Landi, Lorenzo Baraldi, Marcella Cornia, and Rita Cucchiara. 2021 · 2021
Later among the works it cites.
CPTR: Full Transformer Network for Image Captioning
Wei Liu, Sihan Chen, Longteng Guo, Xinxin Zhu, and Jing Liu. 2021 · 2021
Later among the works it cites.
Dual-Level Collaborative Transformer for Image Captioning. In
Yunpeng Luo, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao, Yongjian Wu, Feiyue Huang, Chia-Wen Lin, and Rongrong Ji. 2021 · 2021
Later among the works it cites.
Learning Transferable Visual Models From Natural Language Supervision. In
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
How Much Can CLIP Benefit Vision-and-Language Tasks?
Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer. 2021 · 2021
Later among the works it cites.
Training data-efficient image transformers & distillation through attention. In
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou. 2021 · 2021
Later among the works it cites.
Adaptive Semiparametric Language Models
Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong. 2021 · 2021
Later among the works it cites.
RSTNet: Captioning with Adaptive Attention on Visual and Non-Visual Words. In
Xuying Zhang, Xiaoshuai Sun, Yunpeng Luo, Jiayi Ji, Yiyi Zhou, Yongjian Wu, Feiyue Huang, and Rongrong Ji. 2021 · 2021
Later among the works it cites.
Flamingo: a Visual Language Model for Few-Shot Learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al · 2022
Closest in time.
Universal Captioner: Inducing Content-Style Separation in Vision-and-Language Model Training
Marcella Cornia, Lorenzo Baraldi, Giuseppe Fiameni, and Rita Cucchiara. 2022 · 2022
Closest in time.
Scaling Up Vision-Language Pre-Training for Image Captioning. In
Xiaowei Hu, Zhe Gan, Jianfeng Wang, Zhengyuan Yang, Zicheng Liu, Yumao Lu, and Lijuan Wang. 2022 · 2022
Closest in time.
Memorizing Transformers. In
Yuhuai Wu, Markus N Rabe, DeLesley Hutchins, and Christian Szegedy. 2022 · 2022
Closest in time.