Fetching the paper…
Reading the bibliography…
The world is fundamentally compositional, so it is natural to think of visual recognition as the recognition of basic visually primitives that are composed according to well-defined rules.
Sense and reference
G. Frege · 1948
Earlier work this paper cites.
Handbook of Boolean algebras
J. Monk and R. Bonnet · 1989
Earlier work this paper cites.
The Theaetetus of Plato
M. Burnyeat et al · 1990
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
J. Platt et al · 1999
Earlier work this paper cites.
Image parsing: Unifying segmentation, detection, and recognition
Z. Tu, X. Chen, A. L. Yuille, and S.-C. Zhu · 2005
Earlier work this paper cites.
Similarity by composition
O. Boiman and M. Irani · 2007
Earlier work this paper cites.
A stochastic grammar of images
S.-C. Zhu, D. Mumford, et al · 2007
Earlier work this paper cites.
Max margin and/or graph learning for parsing the human body
L. Zhu, Y. Chen, Y. Lu, C. Lin, and A. Yuille · 2008
Earlier work this paper cites.
Neural networks and learning machines , volume 3
S. S. Haykin, S. S. Haykin, S. S. Haykin, and S. S. Haykin · 2009
Earlier work this paper cites.
Learning to detect unseen object classes by between-class attribute transfer
C. H. Lampert, H. Nickisch, and S. Harmeling · 2009
Earlier work this paper cites.
Zero-shot learning with semantic output codes
M. Palatucci, D. Pomerleau, G. E. Hinton, and T. M. Mitchell · 2009
Earlier work this paper cites.
Object detection with discriminatively trained part-based models
P. F. Felzenszwalb, R. B. Girshick, D. McAllester, and D. Ramanan · 2010
Cited alongside, same era.
Latent hierarchical structural learning for object detection
L. Zhu, Y. Chen, A. Yuille, and W. Freeman · 2010
Cited alongside, same era.
Object detection with grammar models
R. B. Girshick, P. F. Felzenszwalb, and D. A. Mcallester · 2011
Cited alongside, same era.
Parsing natural scenes and natural language with recursive neural networks
R. Socher, C. C. Lin, C. Manning, and A. Y. Ng · 2011
Cited alongside, same era.
The Caltech-UCSD Birds-200-2011 Dataset
C. Wah, S. Branson, P. Welinder, P. Perona, and S. Belongie · 2011
Cited alongside, same era.
A numerical study of the bottom-up and top-down inference processes in and-or graphs
T. Wu and S.-C. Zhu · 2011
Cited alongside, same era.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Later among the works it cites.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. A. Efros · 2015
Later among the works it cites.
Predicting deep zero-shot convolutional neural networks using textual descriptions
J. Lei Ba, K. Swersky, S. Fidler, et al · 2015
Later among the works it cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Later among the works it cites.
Neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2016
Later among the works it cites.
Neural programmer: Inducing latent programs with gradient descent
A. Neelakantan, Q. V. Le, and I. Sutskever · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“clustering by composition”–unsupervised discovery of image categories
A. Faktor and M. Irani · 2012
Cited alongside, same era.
Co-segmentation by composition
A. Faktor and M. Irani · 2013
Cited alongside, same era.
Devise: A deep visual-semantic embedding model
A. Frome, G. S. Corrado, J. Shlens, S. Bengio, J. Dean, T. Mikolov, et al · 2013
Cited alongside, same era.
Learning and-or templates for object recognition and detection
Z. Si and S.-C. Zhu · 2013
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Self-supervised video representation learning with odd-one-out networks
B. Fernando, H. Bilen, E. Gavves, and S. Gould · 2017
Later among the works it cites.
Learning to reason: End-to-end module networks for visual question answering
R. Hu, J. Andreas, M. Rohrbach, T. Darrell, and K. Saenko · 2017
Later among the works it cites.
From Red Wine to Red Tomato: Composition with Context
I. Misra, A. Gupta, and M. Hebert · 2017
Later among the works it cites.
Deeppermnet: Visual permutation learning
R. Santa Cruz, B. Fernando, A. Cherian, and S. Gould · 2017
Later among the works it cites.
Towards a unified compositional model for visual pattern modeling
W. Tang, P. Yu, J. Zhou, and Y. Wu · 2017
Later among the works it cites.
Zero-shot learning-a comprehensive evaluation of the good, the bad and the ugly
Y. Xian, C. H. Lampert, B. Schiele, and Z. Akata · 2017
Later among the works it cites.