Fetching the paper…
Reading the bibliography…
Shouldn't language and vision features be treated equally in vision-language (VL) tasks? Many VL approaches treat the language component as an afterthought, using simple language models that are either built upon fixed word embeddings trained on text-only data or are learned from scratch.
Wordnet: a lexical database for english
George A. Miller · 1995
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin · 2003
Earlier work this paper cites.
The IAPR TC-12 benchmark – a new evaluation resource for visual information systems, 2006
Michael Grubinger, Paul Clough, Henning Müller, and Thomas Deselaers · 2006
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tom Mikolov, Wen-tau Yih, and Geoffrey Zweig · 2013
Earlier work this paper cites.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John Platt, et al · 2014
Earlier work this paper cites.
ReferItGame: Referring to objects in photographs of natural scenes
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg · 2014
Earlier work this paper cites.
Learning image embeddings using convolutional neural networks for improved multi-modal semantics
Douwe Kiela and Léon Bottou · 2014
Earlier work this paper cites.
What are you talking about? text-to-image coreference
Chen Kong, Dahua Lin, Mohit Bansal, Raquel Urtasun, and Sanja Fidler · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Improving lexical embeddings with semantic knowledge
Mo Yu and Mark Dredze · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Retrofitting word vectors to semantic lexicons
Manaal Faruqui, Jesse Dodge, Sujay K. Jauhar, Chris Dyer, Eduard Hovy, and Noah A. Smith · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Fisher vectors derived from hybrid gaussian-laplacian mixture models for image annotation
Benjamin Klein, Guy Lev, Gil Sadeh, and Lior Wolf · 2015
Earlier work this paper cites.
Combining language and vision with a multimodal skip-gram model
Angeliki Lazaridou, Nghia The Pham, and Marco Baroni · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Sequence to sequence – video to text
Subhashini Venugopalan, Marcus Rohrbach, Jeff Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Aaron Courville, Ruslan Salakhutdinov, Richard Zemel, and Yoshua Bengio · 2015
Cited alongside, same era.
Visual Madlibs: Fill in the blank Image Generation and Question Answering
Licheng Yu, Eunbyung Park, Alexander C. Berg, and Tamara L. Berg · 2015
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Cited alongside, same era.
Natural language object retrieval
Ronghang Hu, Huazhe Xu, Marcus Rohrbach, Jiashi Feng, Kate Saenko, and Trevor Darrell · 2016
Cited alongside, same era.
Visual word2vec (vis-w2v): Learning visually grounded word embeddings using abstract scenes
Satwik Kottur, Ramakrishna Vedantam, Jose ´ M. F. Moura, and Devi Parikh · 2016
Cited alongside, same era.
Phrase localization and visual relationship detection with comprehensive image-language cues
Bryan A. Plummer, Arun Mallya, Christopher M. Cervantes, Julia Hockenmaier, and Svetlana Lazebnik · 2017
Later among the works it cites.
Flickr30K Entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Learning two-branch neural networks for image-text matching tasks
Liwei Wang, Yin Li, Jing Huang, and Svetlana Lazebnik · 2017
Later among the works it cites.
R-C3D: Region convolutional 3d network for temporal activity detection
Huijuan Xu, Abir Das, and Kate Saenko · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
RNN fisher vectors for action recognition and image annotation
Guy Lev, Gil Sadeh, Benjamin Klein, and Lior Wolf · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Cited alongside, same era.
Grounding of textual phrases in images by reconstruction
Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, and Bernt Schiele · 2016
Cited alongside, same era.
Training region-based object detectors with online hard example mining
Abhinav Shrivastava, Abhinav Gupta, and Ross Girshick · 2016
Cited alongside, same era.
Solving Visual Madlibs with Multiple Cues
Tatiana Tommasi, Arun Mallya, Bryan A. Plummer, Svetlana Lazebnik, Alex C. Berg, and Tamara L. Berg · 2016
Cited alongside, same era.
Order embeddings of images and language
Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun · 2016
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Later among the works it cites.
Temporally grounding natural sentence in video
Jingyuan Chen, Xinpeng Chen, Lin Ma, Zequn Jie, and Tat-Seng Chua · 2018
Later among the works it cites.
Regularizing rnns for caption generation by reconstructing the past with the present
Xinpeng Chen, Lin Ma, Wenhao Jiang, Jian Yao, and Wei Liu · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Later among the works it cites.
VSE++: improving visual-semantic embeddings with hard negatives
Fartash Faghri, David J. Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2018
Later among the works it cites.
Image captioning with word level attention
Fang Fang, Hanli Wang, and Pengjie Tang · 2018
Later among the works it cites.
Discriminative learning of open-vocabulary object retrieval and localization by negative phrase augmentation
Ryota Hinami and Shin’ichi Satoh · 2018
Later among the works it cites.
Learning semantic concepts and order for image and sentence matching
Yan Huang, Qi Wu, and Liang Wang · 2018
Later among the works it cites.
Bilinear attention networks
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang · 2018
Later among the works it cites.
Stacked cross attention for image-text matching
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He · 2018
Later among the works it cites.
Temporal modular networks for retrieving complex compositional activities in videos
Bingbin Liu, Serena Yeung, Edward Chou, De-An Huang, Li Fei-Fei, and Juan Carlos Niebles · 2018
Later among the works it cites.
Packnet: Adding multiple tasks to a single network by iterative pruning
Arun Mallya and Svetlana Lazebnik · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Later among the works it cites.
Conditional image-text embedding networks
Bryan A. Plummer, Paige Kordas, M. Hadi Kiapour, Shuai Zheng, Robinson Piramuthu, and Svetlana Lazebnik · 2018
Later among the works it cites.
Revisiting image-language embeddings for open-ended phrase detection
Bryan A. Plummer, Kevin J. Shih, Yichen Li, Ke Xu, Svetlana Lazebnik, Stan Sclaroff, and Kate Saenko · 2018
Later among the works it cites.
Deep cross-modal projection learning for image-text matching
Ying Zhang and Huchuan Lu · 2018
Later among the works it cites.
Multilevel language and vision integration for text-to-clip retrieval
Huijuan Xu, Kun He, Bryan A. Plummer, Leonid Sigal, Stan Sclaroff, and Kate Saenko · 2019
Closest in time.