Fetching the paper…
Reading the bibliography…
Despite the steady progress in video analysis led by the adoption of convolutional neural networks (CNNs), the relative improvement has been less drastic as that in 2D static image classification.
A duality based approach for realtime tv-l1 optical flow
Zach, C., Pock, T., Bischof, H.: · 2007
Earlier work this paper cites.
Visualizing data using t-SNE
Maaten, L.v.d., Hinton, G.: · 2008
Earlier work this paper cites.
HMDB: A large video database for human motion recognition
Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., Serre, T.: · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., Hinton, G.E.: · 2012
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A., Shah, M.: · 2012
Earlier work this paper cites.
Towards understanding action recognition
Jhuang, H., Gall, J., Zuffi, S., Schmid, C., Black, M.: · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., Fei-Fei, L.: · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Simonyan, K., Zisserman, A.: · 2014
Earlier work this paper cites.
Seeing the arrow of time
Pickup, L.C., Pan, Z., Wei, D., Shih, Y., Zhang, C., Zisserman, A., Scholkopf, B., Freeman, W.T.: · 2014
Earlier work this paper cites.
C3D: Generic features for video analysis
Tran, D., Bourdev, L.D., Fergus, R., Torresani, L., Paluri, M.: · 2014
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A.: · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K., Zisserman, A.: · 2015
Earlier work this paper cites.
ActivityNet: A large-scale video benchmark for human activity understanding
Caba Heilbron, F., Escorcia, V., Ghanem, B., Niebles, J.C.: · 2015
Earlier work this paper cites.
Human action recognition using factorized spatio-temporal convolutional networks
Sun, L., Jia, K., Yeung, D.Y., Shi, B.E.: · 2015
Earlier work this paper cites.
Towards good practices for very deep two-stream convnets
Wang, L., Xiong, Y., Wang, Z., Qiao, Y.: · 2015
Earlier work this paper cites.
Beyond short snippets: Deep networks for video classification
Ng, J.Y., Hausknecht, M.J., Vijayanarasimhan, S., Vinyals, O., Monga, R., Toderici, G.: · 2015
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Donahue, J., Hendricks, L.A., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., Darrell, T.: · 2015
Earlier work this paper cites.
Faster R-CNN: Towards real-time object detection with region proposal networks
Ren, S., He, K., Girshick, R., Sun, J.: · 2015
Earlier work this paper cites.
Finding action tubes
Gkioxari, G., Malik, J.: · 2015
Earlier work this paper cites.
Learning to track for spatio-temporal action localization
Weinzaepfel, P., Harchaoui, Z., Schmid, C.: · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., Sun, J.: · 2016
Cited alongside, same era.
Hollywood in homes: Crowdsourcing data collection for activity understanding
Sigurdsson, G.A., Varol, G., Wang, X., Farhadi, A., Laptev, I., Gupta, A.: · 2016
Cited alongside, same era.
Dynamic image networks for action recognition
Bilen, H., Fernando, B., Gavves, E., Vedaldi, A., Gould, S.: · 2016
Cited alongside, same era.
Convolutional two-stream network fusion for video action recognition
Feichtenhofer, C., Pinz, A., Zisserman, A.: · 2016
Cited alongside, same era.
Spatiotemporal residual networks for video action recognition
Feichtenhofer, C., Pinz, A., Wildes, R.P.: · 2016
Cited alongside, same era.
Learnable pooling with context gating for video classification
Miech, A., Laptev, I., Sivic, J.: · 2017
Closest in time.
Action recognition with dynamic image networks
Bilen, H., Fernando, B., Gavves, E., Vedaldi, A.: · 2017
Closest in time.
Temporal residual networks for dynamic scene recognition
Feichtenhofer, C., Pinz, A., Wildes, R.P.: · 2017
Closest in time.
Spatiotemporal multiplier networks for video action recognition
Feichtenhofer, C., Pinz, A., Wildes, R.: · 2017
Closest in time.
Chained multi-stream networks exploiting pose, motion, and appearance for action classification and detection
Zolfaghari, M., Oliveira, G.L., Sedaghat, N., Brox, T.: · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Actions ~ transformations
Wang, X., Farhadi, A., Gupta, A.: · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., Van Gool, L.: · 2016
Cited alongside, same era.
Multi-region two-stream r-cnn for action detection
Peng, X., Schmid, C.: · 2016
Cited alongside, same era.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al.: · 2017
Cited alongside, same era.
The “something something” video database for learning and evaluating visual common sense
Goyal, R., Kahou, S.E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al.: · 2017
Cited alongside, same era.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J., Zisserman, A.: · 2017
Cited alongside, same era.
Zhou, B., Andonian, A., Torralba, A.: · 2017
Closest in time.
Bian, Y., Gan, C., Liu, X., Li, F., Long, X., Li, Y., Qi, H., Zhou, J., Wen, S., Lin, Y.: · 2017
Closest in time.
Learning spatio-temporal representation with pseudo-3d residual networks
Qiu, Z., Yao, T., Mei, T.: · 2017
Closest in time.
Convnet architecture search for spatiotemporal feature learning
Tran, D., Ray, J., Shou, Z., Chang, S., Paluri, M.: · 2017
Closest in time.
AMTnet: Action-micro-tube regression by end-to-end trainable deep architecture
Saha, S., G.Sing, Cuzzolin, F.: · 2017
Closest in time.
Speed/accuracy trade-offs for modern convolutional object detectors
Huang, J., Rathod, V., Sun, C., Zhu, M., Korattikara, A., Fathi, A., Fischer, I., Wojna, Z., Song, Y., Guadarrama, S., et al.: · 2017
Closest in time.
Action Tubelet Detector for Spatio-Temporal Action Localization
Kalogeiton, V., Weinzaepfel, P., Ferrari, V., Schmid, C.: · 2017
Closest in time.
A closer look at spatiotemporal convolutions for action recognition
Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., Paluri, M.: · 2018
Closest in time.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., de Vries, H., Dumoulin, V., Courville, A.: · 2018
Closest in time.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
Elfwing, S., Uchibe, E., Doya, K.: · 2018
Closest in time.
Squeeze-and-excitation networks
Hu, J., Shen, L., Sun, G.: · 2018
Closest in time.
Non-local neural networks
Wang, X., Girshick, R., Gupta, A., He, K.: · 2018
Closest in time.
Appearance-and-relation networks for video classification
Wang, L., Li, W., Li, W., Gool, L.V.: · 2018
Closest in time.
AVA: A video dataset of spatio-temporally localized atomic visual actions
Gu, C., Sun, C., Ross, D.A., Vondrick, C., Pantofaru, C., Li, Y., Vijayanarasimhan, S., Toderici, G., Ricco, S., Sukthankar, R., Schmid, C., Malik, J.: · 2018
Closest in time.