Fetching the paper…
Reading the bibliography…
To advance models of multimodal context, we introduce a simple yet powerful neural architecture for data that combines vision and natural language.
Gqa: a new dataset for compositional question answering over real-world images
Drew A Hudson and Christopher D Manning. 2019 · 1902
Earlier work this paper cites.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019 · 1904
Earlier work this paper cites.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Daxin Jiang, and Ming Zhou. 2019a · 1908
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019b · 1908
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 1908
Earlier work this paper cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2019 · 1908
Earlier work this paper cites.
Distributional structure
Zellig S Harris. 1954 · 1954
Earlier work this paper cites.
A synopsis of linguistic theory, 1930-1955
John R Firth. 1957 · 1955
Earlier work this paper cites.
Indexing by latent semantic analysis
Scott Deerwester, Susan T Dumais, George W Furnas, Thomas K Landauer, and Richard Harshman. 1990 · 1990
Earlier work this paper cites.
Wsabie: Scaling up to large vocabulary image annotation
Jason Weston, Samy Bengio, and Nicolas Usunier. 2011 · 2011
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Cited alongside, same era.
Multimodal distributional semantics
Elia Bruni, Nam-Khanh Tran, and Marco Baroni. 2014 · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Cited alongside, same era.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Identity mappings in deep residual networks
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
End-to-end retrieval in continuous space
Daniel Gillick, Alessandro Presta, and Gaurav Singh Tomar. 2018 · 2018
Later among the works it cites.
Compositional attention networks for machine reasoning
Drew A Hudson and Christopher D Manning. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Yin and Yang: Balancing and answering binary visual questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2016 · 2016
Cited alongside, same era.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. 2017 · 2017
Cited alongside, same era.
Learning visually grounded sentence representations
Douwe Kiela, Alexis Conneau, Allan Jabri, and Maximilian Nickel. 2017 · 2017
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Later among the works it cites.
Starspace: Embed all the things!
Ledell Yu Wu, Adam Fisch, Sumit Chopra, Keith Adams, Antoine Bordes, and Jason Weston. 2018 · 2018
Later among the works it cites.
The neuro-symbolic concept learner: Interpreting scenes, words, and sentences from natural supervision
Jiayuan Mao, Chuang Gan, Pushmeet Kohli, Joshua B Tenenbaum, and Jiajun Wu. 2019 · 2019
Closest in time.
Like a baby: Visually situated neural language acquisition
Alexander G. Ororbia, Ankur Mali, Matthew A. Kelly, and David Reitter. 2019 · 2019
Closest in time.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Closest in time.