Fetching the paper…
Reading the bibliography…
Object Permanence allows people to reason about the location of non-visible objects, by understanding that they continue to exist even when not perceived directly.
Piaget, J.: The construction of reality in the child (1954)
1954
Earlier work this paper cites.
Baillargeon, R., DeVos, J.: Object permanence in young infants: Further evidence. Child development 62
1991
Earlier work this paper cites.
Hochreiter, S., Schmidhuber, J.: Long short-term memory. Neural computation 9
1997
Earlier work this paper cites.
Aguiar, A., Baillargeon, R.: 2.5-month-old infants’ reasoning about when objects should and should not be occluded. Cognitive psychology 39
1999
Earlier work this paper cites.
Huang, Y., Essa, I.: Tracking multiple objects through occlusions. In: 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05). vol. 2, pp. 1051–1058. IEEE (2005)
2005
Earlier work this paper cites.
Smitsman, A.W., Dejonckheere, P.J., De Wit, T.C.: The significance of event information for 6-to 16-month-old infants’ perception of containment. Developmental psychology 45
2009
Earlier work this paper cites.
Grabner, H., Matas, J., Van Gool, L., Cattin, P.: Tracking the invisible: Learning where the object might be. In: 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. pp. 1285–1292. IEEE (2010)
2010
Earlier work this paper cites.
Papadourakis, V., Argyros, A.: Multiple objects tracking in the presence of long-term occlusions. Computer Vision and Image Understanding 114
2010
Earlier work this paper cites.
Sadeghi, M.A., Farhadi, A.: Recognition using visual phrases. In: CVPR 2011. pp. 1745–1752. IEEE (2011)
2011
Earlier work this paper cites.
Simonyan, K., Zisserman, A.: Two-stream convolutional networks for action recognition in videos. In: Advances in neural information processing systems. pp. 568–576 (2014)
2014
Earlier work this paper cites.
Ren, S., He, K., Girshick, R., Sun, J.: Faster r-cnn: Towards real-time object detection with region proposal networks. In: Advances in neural information processing systems. pp. 91–99 (2015)
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Tran, D., Bourdev, L., Fergus, R., Torresani, L., Paluri, M.: Learning spatiotemporal features with 3d convolutional networks. In: Proceedings of the IEEE international conference on computer vision. pp. 4489–4497 (2015)
2015
Cited alongside, same era.
Wu, Y., Lim, J., Yang, M.H.: Object tracking benchmark. IEEE Transactions on Pattern Analysis and Machine Intelligence 37
2015
Cited alongside, same era.
Yue-Hei Ng, J., Hausknecht, M., Vijayanarasimhan, S., Vinyals, O., Monga, R., Toderici, G.: Beyond short snippets: Deep networks for video classification. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 4694–4702 (2015)
2015
Cited alongside, same era.
2016
Cited alongside, same era.
Kristan, M., Leonardis, A., et al., J.M.: The sixth visual object tracking vot2018 challenge results. In: ECCV Workshops (2018)
2018
Later among the works it cites.
Li, B., Yan, J., Wu, W., Zhu, Z., Hu, X.: High performance visual tracking with siamese region proposal network. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition pp. 8971–8980 (2018)
2018
Later among the works it cites.
Liang, W., Zhu, Y., Zhu, S.C.: Tracking occluded objects and recovering incomplete trajectories by reasoning about containment relations and human actions. In: Thirty-Second AAAI Conference on Artificial Intelligence (2018)
2018
Later among the works it cites.
Zhou, B., Andonian, A., Oliva, A., Torralba, A.: Temporal relational reasoning in videos. European Conference on Computer Vision (2018)
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lu, C., Krishna, R., Bernstein, M., Fei-Fei, L.: Visual relationship detection with language priors. In: European conference on computer vision. pp. 852–869. Springer (2016)
2016
Cited alongside, same era.
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 6299–6308 (2017)
2017
Cited alongside, same era.
Gao, L., Guo, Z., Zhang, H., Xu, X., Shen, H.T.: Video captioning with attention-based lstm and semantic consistency. IEEE Transactions on Multimedia 19
2017
Cited alongside, same era.
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Lawrence Zitnick, C., Girshick, R.: Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 2901–2910 (2017)
2017
Cited alongside, same era.
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.J., Shamma, D.A., et al.: Visual genome: Connecting language and vision using crowdsourced dense image annotations. International Journal of Computer Vision 123
2017
Cited alongside, same era.
Song, S., Lan, C., Xing, J., Zeng, W., Liu, J.: An end-to-end spatio-temporal attention model for human action recognition from skeleton data. In: Thirty-first AAAI conference on artificial intelligence (2017)
2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need (2017)
2017
Cited alongside, same era.
2018
Later among the works it cites.
Fan, H., Lin, L., Yang, F., Chu, P., Deng, G., Yu, S., Bai, H., Xu, Y., Liao, C., Ling, H.: Lasot: A high-quality benchmark for large-scale single object tracking. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 5374–5383 (2019)
2019
Later among the works it cites.
Fan, H., Ling, H.: Siamese cascaded region proposal networks for real-time visual tracking. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (Jun 2019)
2019
Later among the works it cites.
2019
Later among the works it cites.
Mojtaba Marvasti-Zadeh, S., Cheng, L., Ghanei-Yakhdan, H., Kasaei, S.: Deep learning for visual tracking: A comprehensive survey. arXiv pp. arXiv–1912 (2019)
2019
Later among the works it cites.
Ullman, S., Dorfman, N., Harari, D.: A model for discovering ‘containment’relations. Cognition 183
2019
Later among the works it cites.
Yi, K., Gan, C., Li, Y., Kohli, P., Wu, J., Torralba, A., Tenenbaum, J.B.: Clevrer: Collision events for video representation and reasoning (2019)
2019
Later among the works it cites.