Fetching the paper…
Reading the bibliography…
We propose Unified Visual-Semantic Embeddings (UniVSE) for learning a joint space of visual and textual concepts.
Hierarchical Multimodal LSTM for Dense Visual-Semantic Embedding
Zhenxing Niu, Mo Zhou, Le Wang, Xinbo Gao, and Gang Hua. 2017 · 1907
Earlier work this paper cites.
Universal Grammar
Richard Montague. 1970 · 1970
Earlier work this paper cites.
Luke S. Zettlemoyer and Michael Collins. 2005 · 2005
Earlier work this paper cites.
A probabilistic computational model of cross-situational word learning
Afsaneh Fazly, Afra Alishahi, and Suzanne Stevenson. 2010 · 2010
Earlier work this paper cites.
Abstract Meaning Representation for Sembanking
Laura Banarescu, Claire Bonial, Shu Cai, Madalina Georgescu, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Philipp Koehn, Martha Palmer, and Nathan Schneider. 2013 · 2013
Earlier work this paper cites.
Devise: A Deep Visual-Semantic Embedding Model
Andrea Frome, Greg S Corrado, Jon Shlens, Samy Bengio, Jeff Dean, Tomas Mikolov, et al. 2013 · 2013
Earlier work this paper cites.
Learning Dependency-Based Compositional Semantics
Percy Liang, Michael I. Jordan, and Dan Klein. 2013 · 2013
Earlier work this paper cites.
Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
Andrej Karpathy, Armand Joulin, and Li F Fei-Fei. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
GloVe: Global Vectors for Word Representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Long-Term Recurrent Convolutional Networks for Visual Recognition and Description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell. 2015 · 2015
Earlier work this paper cites.
Image Retrieval using Scene Graphs
Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Fei-Fei Li. 2015 · 2015
Earlier work this paper cites.
Deep Visual-Semantic Alignments for Generating Image Descriptions
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Multimodal Convolutional Neural Networks for Matching Image and Sentence
Lin Ma, Zhengdong Lu, Lifeng Shang, and Hang Li. 2015 · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. 2015 · 2015
Cited alongside, same era.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
Sebastian Schuster, Ranjay Krishna, Angel Chang, Li Fei-Fei, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Show and Tell: A Neural Image Caption Generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Cited alongside, same era.
Linking Image and Text with 2-way Nets
Aviv Eisenschtat and Lior Wolf. 2017 · 2017
Later among the works it cites.
Instance-Aware Image and Sentence Matching with Selective Multimodal LSTM
Yan Huang, Wei Wang, and Liang Wang. 2017 · 2017
Later among the works it cites.
Learning a Recurrent Residual Fusion Network for Multimodal Matching
Yu Liu, Yanming Guo, Erwin M Bakker, and Michael S Lew. 2017 · 2017
Later among the works it cites.
Multiple Instance Visual-Semantic Embedding
Zhou Ren, Hailin Jin, Zhe Lin, Chen Fang, and Alan Yuille. 2017 · 2017
Later among the works it cites.
FOIL it! Find One Mismatch between Image and Language Caption
Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurélie Herbelot, Moin Nabi, Enver Sangineto, and Raffaella Bernardi. 2017 · 2017
Later among the works it cites.
Weakly-Supervised Visual Grounding of Phrases with Linguistic Structures
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiajun Wu, Yinan Yu, Chang Huang, and Kai Yu. 2015 · 2015
Cited alongside, same era.
Gordon Christie, Ankit Laddha, Aishwarya Agrawal, Stanislaw Antol, Yash Goyal, Kevin Kochersberger, and Dhruv Batra. 2016 · 2016
Cited alongside, same era.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei. 2016 · 2016
Cited alongside, same era.
Joint Image-Text Representation by Gaussian Visual-Semantic Embedding
Zhou Ren, Hailin Jin, Zhe Lin, Chen Fang, and Alan Yuille. 2016 · 2016
Cited alongside, same era.
Training Region-Based Object Detectors with Online Hard Example Mining
Abhinav Shrivastava, Abhinav Gupta, and Ross Girshick. 2016 · 2016
Cited alongside, same era.
Learning Deep Structure-Preserving Image-Text Embeddings
Liwei Wang, Yin Li, and Svetlana Lazebnik. 2016 · 2016
Cited alongside, same era.
Fanyi Xiao, Leonid Sigal, and Yong Jae Lee. 2017 · 2017
Later among the works it cites.
Finding Beans in Burgers: Deep Semantic-Visual Embedding with Localization
Martin Engilberge, Louis Chevallier, Patrick Pérez, and Matthieu Cord. 2018 · 2018
Later among the works it cites.
VSE++: Improving Visual-Semantic Embeddings with Hard Negatives
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2018 · 2018
Later among the works it cites.
Image Generation from Scene Graphs
Justin Johnson, Agrim Gupta, and Li Fei-Fei. 2018 · 2018
Later among the works it cites.
Long Short-Term Memory as a Dynamically Computed Element-wise Weighted Sum
Omer Levy, Kenton Lee, Nicholas FitzGerald, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Learning Visually-Grounded Semantics from Contrastive Adversarial Samples
Haoyue Shi, Jiayuan Mao, Tete Xiao, Yuning Jiang, and Jian Sun. 2018 · 2018
Later among the works it cites.
End-to-End Convolutional Semantic Embeddings
Quanzeng You, Zhengyou Zhang, and Jiebo Luo. 2018 · 2018
Later among the works it cites.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.