Hochreiter, S., Schmidhuber, J.: Long short-term memory. Neural computation 9
1997
Earlier work this paper cites.
Chopra, S., Hadsell, R., LeCun, Y.: Learning a similarity metric discriminatively, with application to face verification. In: Proc. CVPR (2005)
2005
Earlier work this paper cites.
Zach, C., Pock, T., Bischof, H.: A duality based approach for realtime TV-L1 optical flow. In: Pattern Recognition (2007)
2007
Earlier work this paper cites.
Gutmann, M.U., Hyvärinen, A.: Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In: AISTATS (2010)
2010
Earlier work this paper cites.
Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., Serre, T.: HMDB: A large video database for human motion recognition. In: Proc. ICCV. pp. 2556–2563 (2011)
2011
Earlier work this paper cites.
Soomro, K., Zamir, A.R., Shah, M.: UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012)
Original
2012
Earlier work this paper cites.
Chung, J., Gulcehre, C., Cho, K., Bengio, Y.: Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555 (2014)
Original
2014
Earlier work this paper cites.
Graves, A., Wayne, G., Danihelka, I.: Neural turing machines. arXiv preprint arXiv:1410.5401 (2014)
Original
2014
Earlier work this paper cites.
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)
Original
2014
Earlier work this paper cites.
Simonyan, K., Zisserman, A.: Two-stream convolutional networks for action recognition in videos. In: NeurIPS (2014)
2014
Earlier work this paper cites.
Agrawal, P., Carreira, J., Malik, J.: Learning to see by moving. In: Proc. ICCV. pp. 37–45. IEEE (2015)
2015
Earlier work this paper cites.
Bahdanau, D., Cho, K., Bengio, Y.: Neural machine translation by jointly learning to align and translate. In: Proc. ICLR (2015)
2015
Earlier work this paper cites.
Isola, P., Zoran, D., Krishnan, D., Adelson, E.H.: Learning visual groups from co-occurrences in space and time. In: Proc. ICLR (2015)
2015
Earlier work this paper cites.
Jayaraman, D., Grauman, K.: Learning image representations tied to ego-motion. In: Proc. ICCV (2015)
2015
Earlier work this paper cites.
Sukhbaatar, S., Szlam, A., Weston, J., Fergus, R.: End-to-end memory networks. In: NeurIPS (2015)
2015
Earlier work this paper cites.
Wang, X., Gupta, A.: Unsupervised learning of visual representations using videos. In: Proc. ICCV (2015)
2015
Earlier work this paper cites.
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., Bengio, Y.: Show, attend and tell: Neural image caption generation with visual attention. In: Proc. ICML (2015)
2015
Earlier work this paper cites.
Arandjelović, R., Gronat, P., Torii, A., Pajdla, T., Sivic, J.: NetVLAD: CNN architecture for weakly supervised place recognition. In: Proc. CVPR (2016)
2016
Earlier work this paper cites.
Brabandere, B.D., Jia, X., Tuytelaars, T., Gool, L.V.: Dynamic filter networks. In: NeurIPS (2016)
2016
Earlier work this paper cites.
Feichtenhofer, C., Pinz, A., Zisserman, A.: Convolutional two-stream network fusion for video action recognition. In: Proc. CVPR (2016)
2016
Earlier work this paper cites.
Jayaraman, D., Grauman, K.: Slow and steady feature analysis: higher order temporal coherence in video. In: Proc. CVPR (2016)
2016
Earlier work this paper cites.
Kumar, A., Irsoy, O., Ondruska, P., Iyyer, M., Bradbury, J., Gulrajani, I., Zhong, V., Paulus, R., Socher, R.: Ask me anything: Dynamic memory networks for natural language processing. In: Proc. ICML (2016)
2016
Earlier work this paper cites.
Misra, I., Zitnick, C.L., Hebert, M.: Shuffle and learn: Unsupervised learning using temporal order verification. In: Proc. ECCV (2016)
2016
Earlier work this paper cites.
Noroozi, M., Favaro, P.: Unsupervised learning of visual representations by solving jigsaw puzzles. In: Proc. ECCV (2016)
2016
Earlier work this paper cites.