Fetching the paper…
Reading the bibliography…
Spatial transformer networks (STNs) were designed to enable convolutional neural networks (CNNs) to learn invariance to image transformations.
B. D. Lukas and T. Kanade, “An iterative image registration technique with an application to stereo vision,” in Image Understanding Workshop , 1981
1981
Earlier work this paper cites.
J. R. Bergen, P. Anandan, K. J. Hanna, and R. Hingorani, “Hierarchical model-based motion estimation,” in European Conference on Computer Vision (ECCV) . Springer, 1992, pp. 237–252
1992
Earlier work this paper cites.
T. Lindeberg and J. Gårding, “Shape-adapted smoothing in estimation of 3-D depth cues from affine distortions of local 2-D structure,” Image and Vision Computing , vol. 15, pp. 415–434, 1997
1997
Earlier work this paper cites.
A. Baumberg, “Reliable feature matching across widely separated views,” in Proc. Computer Vision and Pattern Recognition (CVPR) , Hilton Head, SC, 2000, pp. I:1774–1781
2000
Earlier work this paper cites.
K. Mikolajczyk and C. Schmid, “Scale & affine invariant interest point detectors,” International Journal of Computer Vision , vol. 60, no. 1, pp. 63–86, 2004
2004
Earlier work this paper cites.
K. Mikolajczyk, T. Tuytelaars, C. Schmid, A. Zisserman, J. Matas, F. Schaffalitzky, T. Kadir, and L. van Gool, “A comparison of affine region detectors,” International Journal of Computer Vision , vol. 65, no. 1–2, pp. 43–72, 2005
2005
Earlier work this paper cites.
Y. LeCun and C. Cortes, “MNIST handwritten digit database,” 2010. [Online]. Available: http://yann.lecun.com/exdb/mnist/
2010
Earlier work this paper cites.
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, and A. Y. Ng, “Reading digits in natural images with unsupervised feature learning,” in NIPS Workshop on Deep Learning and Unsupervised Feature Learning 2011 , 2011
2011
Earlier work this paper cites.
L. Sifre and S. Mallat, “Rotation, scaling and deformation invariant scattering for texture discrimination,” in Proc. Conference on Computer Vision and Pattern Recognition (CVPR) , 2013, pp. 1233–1240
2013
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Spatial pyramid pooling in deep convolutional networks for visual recognition,” in European Conference on Computer Vision (ECCV) . Springer, 2014, pp. 346–361
2014
Earlier work this paper cites.
R. K. Cowen, S. Sponaugle, K. Robinson, J. Luo, O. S. University, and H. M. S. Center, “PlanktonSet 1.0: Plankton imagery data collected from F.G. Walton Smith in Straits of Florida from 2014-06-03 to 2014-06-06 and used in the 2015 National Data Science Bowl (NCEI accession 0127422).” [Online]. Available: https://accession.nodc.noaa.gov/0127422
2015
Cited alongside, same era.
K. Lenc and A. Vedaldi, “Understanding image representations by measuring their equivariance and equivalence,” in Proc. Conference on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 991–999
2015
Cited alongside, same era.
2015
Cited alongside, same era.
C. B. Choy, J. Gwak, S. Savarese, and M. Chandraker, “Universal correspondence network,” in Advances in Neural Information Processing Systems (NIPS) , 2016, pp. 2414–2422
S. Kim, S. Lin, S. R. JEON, D. Min, and K. Sohn, “Recurrent transformer networks for semantic correspondence,” in Advances in Neural Information Processing Systems (NIPS) , 2018, pp. 6126–6136
2018
Later among the works it cites.
Z. Zheng, L. Zheng, and Y. Yang, “Pedestrian alignment network for large-scale person re-identification,” IEEE Transactions on Circuits and Systems for Video Technology , 2018
2018
Later among the works it cites.
R. Kondor and S. Trivedi, “On the generalization of equivariance and convolution in neural networks to the action of compact groups,” in International Conference on Machine Learning (ICML) , 2018, pp. 2752–2760
2018
Later among the works it cites.
Á. Arcos-García, J. A. Alvarez-Garcia, and L. M. Soria-Morillo, “Deep neural network for traffic sign recognition systems: An analysis of spatial transformers and stochastic optimisation methods,” Neural Networks , vol. 99, pp. 158–165, 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
T. Cohen and M. Welling, “Group equivariant convolutional networks,” in International Conference on Machine Learning (ICML) , 2016, pp. 2990–2999
2016
Cited alongside, same era.
F. Yu and V. Koltun, “Multi-scale context aggregation by dilated convolutions,” in Int. Conf. on Learning Representations (ICLR) , 2016
2016
Cited alongside, same era.
B.-I. Cîrstea and L. Likforman-Sulem, “Tied spatial transformer networks for digit recognition,” in 2016 15th International Conference on Frontiers in Handwriting Recognition (ICFHR) . IEEE, 2016, pp. 524–529
2016
Cited alongside, same era.
2017
Cited alongside, same era.
J. Dai, H. Qi, Y. Xiong, Y. Li, G. Zhang, H. Hu, and Y. Wei, “Deformable convolutional networks,” in Proc. International Conference on Computer Vision (ICCV) , 2017, pp. 764–773
2017
Cited alongside, same era.
C.-H. Lin and S. Lucey, “Inverse compositional spatial transformer networks,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2017, pp. 2568–2576
2017
Cited alongside, same era.
T. S. Cohen, M. Geiger, and M. Weiler, “A general theory of equivariant CNNs on homogeneous spaces,” in Advances in Neural Information Processing Systems (NIPS) , 2019, pp. 9142–9153
2019
Later among the works it cites.
T. V. Souza and C. Zanchettin, “Improving deep image clustering with spatial transformer layers,” in International Conference on Artificial Neural Networks (ICANN) . Springer, 2019, pp. 641–654
2019
Later among the works it cites.
T. Lindeberg, “Provably scale-covariant continuous hierarchical networks based on scale-normalized differential expressions coupled in cascade,” Journal of Mathematical Imaging and Vision , vol. 62, no. 1, pp. 120–148, 2020
2020
Closest in time.
Y. Jansson, M. Maydanskiy, L. Finnveden, and T. Lindeberg, “Inability of spatial transformations of CNN feature maps to support invariant recognition,” arXiv preprint:2004.14716 , 2020
2020
Closest in time.
M. Jaderberg, K. Simonyan, A. Zisserman, and K. Kavukcuoglu, “Spatial transformer networks,” in Advances in Neural Information Processing Systems (NIPS) , 2015, pp. 2017–2025
2025
Closest in time.