Fetching the paper…
Reading the bibliography…
Existing visual reasoning datasets such as Visual Question Answering (VQA), often suffer from biases conditioned on the question, image or answer distributions.
A natural logic inference system
Y. Fyodorov, Y. Winter, and N. Francez · 2000
Earlier work this paper cites.
Entailment, intensionality and text understanding
C. Condoravdi, D. Crouch, V. De Paiva, R. Stolle, and D. G. Bobrow · 2003
Earlier work this paper cites.
Recognising textual entailment with logical inference
J. Bos and K. Markert · 2005
Earlier work this paper cites.
The pascal recognising textual entailment challenge
I. Dagan, O. Glickman, and B. Magnini · 2006
Earlier work this paper cites.
An extended model of natural logic
B. MacCartney and C. D. Manning · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Amazon mechanical turk
A. M. Turk · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Earlier work this paper cites.
Fast and scalable polynomial kernels via explicit feature maps
N. Pham and R. Pagh · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Improving image-sentence embeddings using large weakly annotated photo collections
Y. Gong, L. Wang, M. Hodosh, J. Hockenmaier, and S. Lazebnik · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. Manning · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Earlier work this paper cites.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. Lawrence Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
S. R. Bowman, G. Angeli, C. Potts, and C. D. Manning · 2015
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
J. Johnson, R. Krishna, M. Stark, L.-J. Li, D. Shamma, M. Bernstein, and L. Fei-Fei · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Bilinear cnn models for fine-grained visual recognition
T.-Y. Lin, A. RoyChowdhury, and S. Maji · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Earlier work this paper cites.
Reasoning about entailment with neural attention
T. Rocktäschel, E. Grefenstette, K. M. Hermann, T. Kočiskỳ, and P. Blunsom · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
J. Weston, A. Bordes, S. Chopra, A. M. Rush, B. van Merriënboer, A. Joulin, and T. Mikolov · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio · 2015
Earlier work this paper cites.
Enhanced lstm for natural language inference
Q. Chen, X. Zhu, Z. Ling, S. Wei, H. Jiang, and D. Inkpen · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
Compact bilinear pooling
Y. Gao, O. Beijbom, N. Zhang, and T. Darrell · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Modeling relationships in referential expressions with compositional modular networks
R. Hu, M. Rohrbach, J. Andreas, T. Darrell, and K. Saenko · 2016
Cited alongside, same era.
Explaining nonlinear classification decisions with deep taylor decomposition
G. Montavon, S. Lapuschkin, A. Binder, W. Samek, and K.-R. Müller · 2017
Later among the works it cites.
Shortcut-stacked sentence encoders for multi-domain inference
Y. Nie and M. Bansal · 2017
Later among the works it cites.
Film: Visual reasoning with a general conditioning layer
E. Perez, F. Strub, H. De Vries, V. Dumoulin, and A. Courville · 2017
Later among the works it cites.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
M. Raghu, J. Gilmer, J. Yosinski, and J. Sohl-Dickstein · 2017
Later among the works it cites.
A simple neural network module for relational reasoning
A. Santoro, D. Raposo, D. G. Barrett, M. Malinowski, R. Pascanu, P. Battaglia, and T. Lillicrap · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Johnson, A. Karpathy, and L. Fei-Fei · 2016
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, M. Bernstein, and L. Fei-Fei · 2016
Cited alongside, same era.
Attentive explanations: Justifying decisions and pointing to the evidence
D. H. Park, L. A. Hendricks, Z. Akata, B. Schiele, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
Areas of attention for image captioning
M. Pedersoli, T. Lucas, C. Schmid, and J. Verbeek · 2016
Cited alongside, same era.
Why should i trust you?: Explaining the predictions of any classifier
M. T. Ribeiro, S. Singh, and C. Guestrin · 2016
Cited alongside, same era.
Movieqa: Understanding stories in movies through question-answering
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Cited alongside, same era.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra, et al · 2017
Later among the works it cites.
Tips and tricks for visual question answering: Learnings from the 2017 challenge
D. Teney, P. Anderson, X. He, and A. van den Hengel · 2017
Later among the works it cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Later among the works it cites.
Show and tell: Lessons learned from the 2015 mscoco image captioning challenge
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2017
Later among the works it cites.
Visual question answering: A survey of methods and datasets
Q. Wu, D. Teney, P. Wang, C. Shen, A. Dick, and A. van den Hengel · 2017
Later among the works it cites.
Scene graph generation by iterative message passing
D. Xu, Y. Zhu, C. B. Choy, and L. Fei-Fei · 2017
Later among the works it cites.
Visual translation embedding network for visual relation detection
H. Zhang, Z. Kyaw, S.-F. Chang, and T.-S. Chua · 2017
Later among the works it cites.
Relationship proposal networks
J. Zhang, M. Elhoseiny, S. Cohen, W. Chang, and A. Elgammal · 2017
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Later among the works it cites.
Image captioning pytorch implementation
Y. Choi · 2018
Later among the works it cites.
VQA challenge leaderboard 2018
EvalAI · 2018
Later among the works it cites.
Annotation artifacts in natural language inference data
S. Gururangan, S. Swayamdipta, O. Levy, R. Schwartz, S. R. Bowman, and N. A. Smith · 2018
Later among the works it cites.
Compositional attention networks for machine reasoning
D. A. Hudson and C. D. Manning · 2018
Later among the works it cites.
To trust or not to trust a classifier
H. Jiang, B. Kim, and M. Gupta · 2018
Later among the works it cites.
Pythia v0. 1: the winning entry to the vqa challenge 2018
Y. Jiang, V. Natarajan, X. Chen, M. Rohrbach, D. Batra, and D. Parikh · 2018
Later among the works it cites.
J.-H. Kim, J. Jun, and B.-T. Zhang · 2018
Later among the works it cites.
Progressive reasoning by module composition
S. W. Kim, M. Tapaswi, and S. Fidler · 2018
Later among the works it cites.
Mask rcnn pytorch implementation
I. Matterport · 2018
Later among the works it cites.
Reinforced self-attention network: a hybrid of hard and soft attention for sequence modeling
T. Shen, T. Zhou, G. Long, J. Jiang, S. Wang, and C. Zhang · 2018
Later among the works it cites.
Attention on attention: Architectures for visual question answering (vqa)
J. Singh, V. Ying, and A. Nutkiewicz · 2018
Later among the works it cites.
H. T. Vu, C. Greco, A. Erofeeva, S. Jafaritazehjan, G. Linders, M. Tanti, A. Testoni, R. Bernardi, and A. Gatt · 2018
Later among the works it cites.
Top-down neural attention by excitation backprop
J. Zhang, S. A. Bargal, Z. Lin, J. Brandt, X. Shen, and S. Sclaroff · 2018
Later among the works it cites.