Fetching the paper…
Reading the bibliography…
To understand the world, we humans constantly need to relate the present to the past, and put events in context.
Learning process in an asymmetric threshold network
Y. LeCun · 1986
Earlier work this paper cites.
Object bank: A high-level image representation for scene classification & semantic feature sparsification
L.-J. Li, H. Su, L. Fei-Fei, and E. P. Xing · 2010
Earlier work this paper cites.
Detection bank: an object detection based video representation for multimedia event recognition
T. Althoff, H. O. Song, and T. Darrell · 2012
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
J. Donahue, L. Anne Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell · 2015
Earlier work this paper cites.
Fast R-CNN
R. Girshick · 2015
Earlier work this paper cites.
Finding action tubes
G. Gkioxari and J. Malik · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Earlier work this paper cites.
End-to-end memory networks
S. Sukhbaatar, A. Szlam, J. Weston, and R. Fergus · 2015
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri · 2015
Earlier work this paper cites.
Learning to track for spatio-temporal action localization
P. Weinzaepfel, Z. Harchaoui, and C. Schmid · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
J. Yue-Hei Ng, M. Hausknecht, S. Vijayanarasimhan, O. Vinyals, R. Monga, and G. Toderici · 2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Spatiotemporal residual networks for video action recognition
C. Feichtenhofer, A. Pinz, and R. Wildes · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Multi-region two-stream r-cnn for action detection
X. Peng and C. Schmid · 2016
Cited alongside, same era.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
J. Carreira and A. Zisserman · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Cited alongside, same era.
Mask R-CNN
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang · 2018
Closest in time.
Object level visual reasoning in videos
F. Baradel, N. Neverova, C. Wolf, J. Mille, and G. Mori · 2018
Closest in time.
Scaling egocentric vision: The EPIC-kitchens dataset
D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al · 2018
Closest in time.
R. Girdhar, J. Carreira, C. Doersch, and A. Zisserman · 2018
Closest in time.
Detectron, 2018
R. Girshick, I. Radosavovic, G. Gkioxari, P. Dollár, and K. He · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. He, G. Gkioxari, P. Dollár, and R. Girshick · 2017
Cited alongside, same era.
Tube convolutional neural network (T-CNN) for action detection in videos
R. Hou, C. Chen, and M. Shah · 2017
Cited alongside, same era.
Action tubelet detector for spatio-temporal action localization
V. Kalogeiton, P. Weinzaepfel, V. Ferrari, and C. Schmid · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al · 2017
Cited alongside, same era.
Temporal modeling approaches for large-scale youtube-8m video understanding
F. Li, C. Gan, X. Liu, Y. Bian, X. Long, Y. Li, Z. Li, J. Zhou, and S. Wen · 2017
Cited alongside, same era.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie · 2017
Cited alongside, same era.
Learnable pooling with context gating for video classification
A. Miech, I. Laptev, and J. Sivic · 2017
Cited alongside, same era.
Closest in time.
AVA: A video dataset of spatio-temporally localized atomic visual actions
C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankar, et al · 2018
Closest in time.
Human centric spatio-temporal action localization
J. Jiang, Y. Cao, L. Song, S. Z. Y. Li, Z. Xu, Q. Wu, C. Gan, C. Zhang, and G. Yu · 2018
Closest in time.
Recurrent tubelet proposal and recognition networks for action detection
D. Li, Z. Qiu, Q. Dai, T. Yao, and T. Mei · 2018
Closest in time.
Videolstm convolves, attends and flows for action recognition
Z. Li, K. Gavrilyuk, E. Gavves, M. Jain, and C. G. Snoek · 2018
Closest in time.
Attend and interact: Higher-order object interactions for video understanding
C.-Y. Ma, A. Kadav, I. Melvin, Z. Kira, G. AlRegib, and H. P. Graf · 2018
Closest in time.
Actor-centric relation network
C. Sun, A. Shrivastava, C. Vondrick, K. Murphy, R. Sukthankar, and C. Schmid · 2018
Closest in time.
Non-local netvlad encoding for video classification
Y. Tang, X. Zhang, J. Wang, S. Chen, L. Ma, and Y.-G. Jiang · 2018
Closest in time.
Long-term temporal convolutions for action recognition
G. Varol, I. Laptev, and C. Schmid · 2018
Closest in time.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2018
Closest in time.
Videos as space-time region graphs
X. Wang and A. Gupta · 2018
Closest in time.
Compressed video action recognition
C.-Y. Wu, M. Zaheer, H. Hu, R. Manmatha, A. J. Smola, and P. Krähenbühl · 2018
Closest in time.
Rethinking spatiotemporal feature learning for video understanding
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy · 2018
Closest in time.
Temporal relational reasoning in videos
B. Zhou, A. Andonian, A. Oliva, and A. Torralba · 2018
Closest in time.