Fetching the paper…
Reading the bibliography…
We present a novel unsupervised feature representation learning method, Visual Commonsense Region-based Convolutional Neural Network (VC R-CNN), to serve as an improved visual region encoder for high-level tasks such as captioning and VQA.
The theory of affordances
James J Gibson · 1977
Earlier work this paper cites.
Common sense concepts about motion
Ibrahim Abou Halloun and David Hestenes · 1985
Earlier work this paper cites.
The structures of the common-sense world
Barry Smith · 1995
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Who killed the directed model?
Justin Domke, Alap Karapurkar, and Yiannis Aloimonos · 2008
Earlier work this paper cites.
Visualizing data using t-sne
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Beyond categories: The visual memex model for reasoning about object relationships
Tomasz Malisiewicz and Alyosha Efros · 2009
Earlier work this paper cites.
Common sense
Sophia A Rosenfeld · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Deep generative stochastic networks trainable by backprop
Yoshua Bengio, Eric Laufer, Guillaume Alain, and Jason Yosinski · 2014
Earlier work this paper cites.
Visual causal feature learning
Krzysztof Chalupka, Pietro Perona, and Frederick Eberhardt · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Interpretation and identification of causal mediation
Judea Pearl · 2014
Earlier work this paper cites.
Reasoning about object affordances in a knowledge base representation
Yuke Zhu, Alireza Fathi, and Li Fei-Fei · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Unsupervised visual representation learning by context prediction
Carl Doersch, Abhinav Gupta, and Alexei A Efros · 2015
Earlier work this paper cites.
The visual object tracking vot2015 challenge results
Matej Kristan, Jiri Matas, Ales Leonardis, Michael Felsberg, Luka Cehovin, Gustavo Fernandez, Tomas Vojir, Gustav Hager, Georg Nebehay, and Roman Pflugfelder · 2015
Earlier work this paper cites.
Don’t just listen, use your imagination: Leveraging visual common sense for non-visual tasks
Xiao Lin and Devi Parikh · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Viske: Visual knowledge extraction and question answering by visual verification of relation phrases
Fereshteh Sadeghi, Santosh K Kumar Divvala, and Ali Farhadi · 2015
Earlier work this paper cites.
Generative image modeling using spatial lstms
Lucas Theis and Matthias Bethge · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Cited alongside, same era.
Learning common sense through visual abstraction
Ramakrishna Vedantam, Xiao Lin, Tanmay Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Cited alongside, same era.
Spice: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Rubi: Reducing unimodal biases for visual question answering
Remi Cadene, Corentin Dancette, Matthieu Cord, Devi Parikh, et al · 2019
Later among the works it cites.
Uniter: Learning universal image-text representations
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2019
Later among the works it cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Causal reasoning from meta-reinforcement learning
Ishita Dasgupta, Jane Wang, Silvia Chiappa, Jovana Mitrovic, Pedro Ortega, David Raposo, Edward Hughes, Peter Battaglia, Matthew Botvinick, and Zeb Kurth-Nelson · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Ssd: Single shot multibox detector
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg · 2016
Cited alongside, same era.
Causal inference in statistics: A primer
Judea Pearl, Madelyn Glymour, and Nicholas P Jewell · 2016
Cited alongside, same era.
Ask me anything: Free-form visual question answering based on knowledge from external sources
Qi Wu, Peng Wang, Chunhua Shen, Anthony Dick, and Anton van den Hengel · 2016
Cited alongside, same era.
Stating the obvious: Extracting visual common sense knowledge
Mark Yatskar, Vicente Ordonez, and Ali Farhadi · 2016
Cited alongside, same era.
The” something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Cited alongside, same era.
Later among the works it cites.
On multi-cause approaches to causal inference with unobserved counfounding: Two cautionary failure cases and a promising alternative
Alexander D’Amour · 2019
Later among the works it cites.
Dynamic fusion with intra-and inter-modality attention flow for visual question answering
Peng Gao, Zhengkai Jiang, Haoxuan You, Pan Lu, Steven CH Hoi, Xiaogang Wang, and Hongsheng Li · 2019
Later among the works it cites.
Attention on attention for image captioning
Lun Huang, Wenmin Wang, Jie Chen, and Xiao-Yong Wei · 2019
Later among the works it cites.
Revisiting self-supervised visual representation learning
Alexander Kolesnikov, Xiaohua Zhai, and Lucas Beyer · 2019
Later among the works it cites.
Rethinking data augmentation: Self-supervision and self-distillation
Hankook Lee, Sung Ju Hwang, and Jinwoo Shin · 2019
Later among the works it cites.
Siamrpn++: Evolution of siamese visual tracking with very deep networks
Bo Li, Wei Wu, Qiang Wang, Fangyi Zhang, Junliang Xing, and Junjie Yan · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Later among the works it cites.
Explicit bias discovery in visual question answering models
Varun Manjunatha, Nirat Saini, and Larry S Davis · 2019
Later among the works it cites.
Causal induction from visual observations for goal directed tasks
Suraj Nair, Yuke Zhu, Silvio Savarese, and Li Fei-Fei · 2019
Later among the works it cites.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Later among the works it cites.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Later among the works it cites.
Auto-encoding scene graphs for image captioning
Xu Yang, Kaihua Tang, Hanwang Zhang, and Jianfei Cai · 2019
Later among the works it cites.
Learning to collocate neural modules for image captioning
Xu Yang, Hanwang Zhang, and Jianfei Cai · 2019
Later among the works it cites.
Deep modular co-attention networks for visual question answering
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian · 2019
Later among the works it cites.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Later among the works it cites.
S4l: Self-supervised semi-supervised learning
Xiaohua Zhai, Avital Oliver, Alexander Kolesnikov, and Lucas Beyer · 2019
Later among the works it cites.
The open images dataset v4
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al · 2020
Closest in time.
Two causal principles for improving visual dialog
Jiaxin Qi, Yulei Niu, Jianqiang Huang, and Hanwang Zhang · 2020
Closest in time.
Unbiased scene graph generation from biased training
Kaihua Tang, Yulei Niu, Jianqiang Huang, Jiaxin Shi, and Hanwang Zhang · 2020
Closest in time.
Deconfounded image captioning: A causal retrospect
Xu Yang, Hanwang Zhang, and Jianfei Cai · 2020
Closest in time.