Fetching the paper…
Reading the bibliography…
We present an attention-based modular neural framework for computer vision.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
R. Caruana, “Multitask learning,” Machine learning , vol. 28, no. 1, pp. 41–75, 1997
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
C. Schüldt, I. Laptev, and B. Caputo, “Recognizing human actions: a local svm approach,” in Pattern Recognition, 2004. ICPR 2004. Proceedings of the 17th International Conference on , vol. 3. IEEE, 2004, pp. 32–36
2004
Earlier work this paper cites.
N. Dalal and B. Triggs, “Histograms of oriented gradients for human detection,” in International Conference on Computer Vision & Pattern Recognition , C. Schmid, S. Soatto, and C. Tomasi, Eds., vol. 2, INRIA Rhône-Alpes, ZIRST-655, av. de l’Europe, Montbonnot-38334, June 2005, pp. 886–893. [Online]. Available: http://lear.inrialpes.fr/pubs/2005/DT05
2005
Earlier work this paper cites.
A. Ess, B. Leibe, K. Schindler, , and L. van Gool, “A mobile vision system for robust multi-person tracking,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR’08) . IEEE Press, June 2008
2008
Earlier work this paper cites.
I. Sutskever, G. E. Hinton, and G. W. Taylor, “The recurrent temporal restricted boltzmann machine,” in Advances in Neural Information Processing Systems , 2009, pp. 1601–1608
2009
Earlier work this paper cites.
H. Larochelle and G. E. Hinton, “Learning to combine foveal glimpses with a third-order boltzmann machine,” in Advances in neural information processing systems , 2010, pp. 1243–1251
2010
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted boltzmann machines,” in Proceedings of the 27th International Conference on Machine Learning (ICML-10) , 2010, pp. 807–814
2010
Earlier work this paper cites.
J. Bergstra, O. Breuleux, F. Bastien, P. Lamblin, R. Pascanu, G. Desjardins, J. Turian, D. Warde-Farley, and Y. Bengio, “Theano: a cpu and gpu math expression compiler,” in Proceedings of the Python for scientific computing conference (SciPy) , vol. 4. Austin, TX, 2010, p. 3
2010
Earlier work this paper cites.
M. Everingham, L. Van Gool, C. K. Williams, J. Winn, and A. Zisserman, “The pascal visual object classes (voc) challenge,” International journal of computer vision , vol. 88, no. 2, pp. 303–338, 2010
2010
Earlier work this paper cites.
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
2012
Cited alongside, same era.
2012
Cited alongside, same era.
Z. Jiang, Z. Lin, and L. S. Davis, “Recognizing human actions by learning and matching shape-motion prototype trees,” Pattern Analysis and Machine Intelligence, IEEE Transactions on , vol. 34, no. 3, pp. 533–547, 2012
2012
Cited alongside, same era.
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” in Conference on Computer Vision and Pattern Recognition (CVPR) , 2012
2012
Cited alongside, same era.
2014
Later among the works it cites.
2015
Closest in time.
2015
Closest in time.
2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. Wang and D.-Y. Yeung, “Learning a deep compact image representation for visual tracking,” in Advances in neural information processing systems , 2013, pp. 809–817
2013
Cited alongside, same era.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in Acoustics, Speech and Signal Processing (ICASSP), 2013 IEEE International Conference on . IEEE, 2013, pp. 6645–6649
2013
Cited alongside, same era.
2014
Cited alongside, same era.
V. Mnih, N. Heess, A. Graves et al. , “Recurrent models of visual attention,” in Advances in Neural Information Processing Systems , 2014, pp. 2204–2212
2014
Cited alongside, same era.
M. Ranzato, “On learning where to look,” arXiv preprint arXiv:1405.5488 , 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
A. W. Smeulders, D. M. Chu, R. Cucchiara, S. Calderara, A. Dehghan, and M. Shah, “Visual tracking: An experimental survey,” Pattern Analysis and Machine Intelligence, IEEE Transactions on , vol. 36, no. 7, pp. 1442–1468, 2014
2014
Cited alongside, same era.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Advances in neural information processing systems , 2014, pp. 3104–3112
2014
Cited alongside, same era.
S. K. Sønderby, C. K. Sønderby, H. Nielsen, and O. Winther, “Convolutional lstm networks for subcellular localization of proteins,” in Algorithms for computational biology . Springer, 2015, pp. 68–80
2015
Closest in time.
2015
Closest in time.
2015
Closest in time.
2015
Closest in time.
S. Ebrahimi Kahou, V. Michalski, K. Konda, R. Memisevic, and C. Pal, “Recurrent neural networks for emotion recognition in video,” in Proceedings of the 2015 ACM on International Conference on Multimodal Interaction , ser. ICMI ’15. New York, NY, USA: ACM, 2015, pp. 467–474. [Online]. Available: http://doi.acm.org/10.1145/2818346.2830596
2015
Closest in time.
Y. Wu, J. Lim, and M.-H. Yang, “Object tracking benchmark,” Pattern Analysis and Machine Intelligence, IEEE Transactions on , vol. 37, no. 9, pp. 1834–1848, 2015
2015
Closest in time.
2015
Closest in time.
M. Jaderberg, K. Simonyan, A. Zisserman et al. , “Spatial transformer networks,” in Advances in Neural Information Processing Systems , 2015, pp. 2008–2016
2016
Closest in time.