Fetching the paper…
Reading the bibliography…
People say, "A picture is worth a thousand words".
Object detection in 20 years: A survey. https://arxiv.org/abs/1905.05055
Zou, Z · 1905
Earlier work this paper cites.
Convolutional auto-encoding of sentence topics for image paragraph generation
Wang, J · 1908
Earlier work this paper cites.
Wordnet: a lexical database for english
Miller, G. A · 1995
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K · 2002
Earlier work this paper cites.
Accurate unlexicalized parsing
Klein, D · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Maynez, J · 2005
Earlier work this paper cites.
Vision: A computational investigation into the human representation and processing of visual information
Marr, D · 2010
Earlier work this paper cites.
Detecting hallucinated content in conditional neural sequence generation
Zhou, C · 2011
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Denkowski, M · 2014
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
Johnson, J · 2015
Earlier work this paper cites.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
Schuster, S · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R · 2015
Earlier work this paper cites.
Spice: Semantic propositional image caption evaluation
Anderson, P · 2016
Earlier work this paper cites.
Visual storytelling
Huang, T.-H · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Rajpurkar, P · 2016
Earlier work this paper cites.
Towards diverse and natural image descriptions via a conditional gan
Dai, B · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y · 2017
Earlier work this paper cites.
A hierarchical approach for generating descriptive image paragraphs
Krause, J · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R · 2017
Cited alongside, same era.
Dailydialog: A manually labelled multi-turn dialogue dataset
Li, Y · 2017
Cited alongside, same era.
Recurrent topic-transition gan for visual paragraph generation
Liang, X · 2017
Cited alongside, same era.
Diverse and coherent paragraph generation from images
Chatterjee, M · 2018
Cited alongside, same era.
Beyond narrative description: Generating poetry from images by multi-adversarial training
Liu, B · 2018
Scaling up visual and vision-language representation learning with noisy text supervision
Jia, C · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A · 2021
Later among the works it cites.
S2td: A tree-structured decoder for image paragraph captioning
Shi, Y · 2021
Later among the works it cites.
Flava: A foundational language and vision alignment model
Singh, A · 2021
Later among the works it cites.
Simvlm: Simple visual language model pretraining with weak supervision
Wang, Z · 2021
Later among the works it cites.
An empirical study of gpt-3 for few-shot knowledge-based vqa
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Show and tell more: Topic-oriented multi-sentence image captioning
Mao, Y · 2018
Cited alongside, same era.
Training for diversity in image paragraph captioning
Melas-Kyriazi, L · 2018
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A · 2019
Cited alongside, same era.
Curiosity-driven reinforcement learning for diverse visual paragraph generation
Luo, Y · 2019
Cited alongside, same era.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino, K · 2019
Cited alongside, same era.
Inverse cooking: Recipe generation from food images
Salvador, A · 2019
Cited alongside, same era.
Yang, Z · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Yuan, L · 2021
Later among the works it cites.
Zhu, X · 2021
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B · 2022
Closest in time.
Visualgpt: Data-efficient adaptation of pretrained language models for image captioning
Chen, J · 2022
Closest in time.
What is optical character recognition? - azure cognitive services. https://docs.microsoft.com/en-us/azure/cognitive-services/computer-vision/overview-ocr
Farley, P · 2022
Closest in time.
Searching for computer vision north stars
Fei-Fei, L · 2022
Closest in time.
Li, J · 2022
Closest in time.
Language models can see: Plugging visual controls in text generation
Su, Y · 2022
Closest in time.
Winoground: Probing vision and language models for visio-linguistic compositionality
Thrush, T · 2022
Closest in time.
Wang, P · 2022
Closest in time.
Coca: Contrastive captioners are image-text foundation models
Yu, J · 2022
Closest in time.
Socratic models: Composing zero-shot multimodal reasoning with language
Zeng, A · 2022
Closest in time.