Understand
A fundamental problem in object recognition is the development of image representations that are invariant to common transformations such as translation, rotation, and small deformations.
- There are multiple hypotheses regarding the source of translation invariance in CNNs.
- One idea is that translation invariance is due to the increasing receptive field size of neurons in successive convolution layers.
- Another possibility is that invariance is due to the pooling operation.
Built on
LeCun, Yann, et al. ”Gradient-based learning applied to document recognition.” Proceedings of the IEEE 86.11 (1998): 2278-2324
1998
Earlier work this paper cites.
LeCun, Yann, Corinna Cortes, and Christopher JC Burges. ”The MNIST database of handwritten digits.” (1998)
1998
Earlier work this paper cites.
Sundaramoorthi, Ganesh, et al. ”On the set of images modulo viewpoint and contrast changes.” Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on. IEEE, 2009
2009
Earlier work this paper cites.
Krizhevsky, Alex, Ilya Sutskever, and Geoffrey E. Hinton. ”Imagenet classification with deep convolutional neural networks.” Advances in neural information processing systems. 2012
2012
Earlier work this paper cites.
LeCun, Yann. ”Learning invariant feature hierarchies.” Computer vision?ECCV 2012. Workshops and demonstrations. Springer Berlin Heidelberg, 2012
2012
Earlier work this paper cites.
Similar
Mallat, Stephane. ”Group invariant scattering.” Communications on Pure and Applied Mathematics 65.10 (2012): 1331-1398
2012
Cited alongside, same era.
Bruna, Joan, and Stephane Mallat. ”Invariant scattering convolution networks.” Pattern Analysis and Machine Intelligence, IEEE Transactions on 35.8 (2013): 1872-1886
2013
Cited alongside, same era.
2013
Cited alongside, same era.
Gens, Robert, and Pedro M. Domingos. ”Deep symmetry networks.” Advances in neural information processing systems. 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
Then
Jaderberg, Max, Karen Simonyan, and Andrew Zisserman. ”Spatial transformer networks.” Advances in Neural Information Processing Systems. 2015
2015
Later among the works it cites.
Lenc, Karel, and Andrea Vedaldi. ”Understanding image representations by measuring their equivariance and equivalence.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2015
2015
Later among the works it cites.
2015
Later among the works it cites.
Soatta, Stegano, and Alessandro Chiuso. ”Visual Representations: Defining Properties and Deep Approximations.” ICLR (2016)
2016
Later among the works it cites.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…