Fetching the paper…
Reading the bibliography…
Temporal relational reasoning, the ability to link meaningful transformations of objects or entities over time, is a fundamental property of intelligent species.
Ucf101: A dataset of 101 human actions classes from videos in the wild
Soomro, K., Zamir, A.R., Shah, M.: · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., Hinton, G.E.: · 2012
Earlier work this paper cites.
Temporal localization of actions with actoms
Gaidon, A., Harchaoui, Z., Schmid, C.: · 2013
Earlier work this paper cites.
Action recognition with improved trajectories
Wang, H., Schmid, C.: · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., Fei-Fei, L.: · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Simonyan, K., Zisserman, A.: · 2014
Earlier work this paper cites.
Learning deep features for scene recognition using places database
Zhou, B., Lapedriza, A., Xiao, J., Torralba, A., Oliva, A.: · 2014
Earlier work this paper cites.
Parsing videos of actions with segmental grammars
Pirsiavash, H., Ramanan, D.: · 2014
Earlier work this paper cites.
Activity representation with motion hierarchies
Gaidon, A., Harchaoui, Z., Schmid, C.: · 2014
Earlier work this paper cites.
Cnn features off-the-shelf: an astounding baseline for recognition
Sharif Razavian, A., Azizpour, H., Sullivan, J., Carlsson, S.: · 2014
Earlier work this paper cites.
Thumos challenge: Action recognition with a large number of classes
Gorban, A., Idrees, H., Jiang, Y., Zamir, A.R., Laptev, I., Shah, M., Sukthankar, R.: · 2015
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Donahue, J., Anne Hendricks, L., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., Darrell, T.: · 2015
Cited alongside, same era.
Learning spatiotemporal features with 3d convolutional networks
Tran, D., Bourdev, L., Fergus, R., Torresani, L., Paluri, M.: · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S., Szegedy, C.: · 2015
Cited alongside, same era.
Building machines that learn and think like people
Lake, B.M., Ullman, T.D., Tenenbaum, J.B., Gershman, S.J.: · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., Sun, J.: · 2016
Later among the works it cites.
What actions are needed for understanding human actions in videos?
Sigurdsson, G.A., Russakovsky, O., Gupta, A.: · 2017
Closest in time.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J., Zisserman, A.: · 2017
Closest in time.
A simple neural network module for relational reasoning
Santoro, A., Raposo, D., Barrett, D.G., Malinowski, M., Pascanu, R., Battaglia, P., Lillicrap, T.: · 2017
Closest in time.
The” something something” video database for learning and evaluating visual common sense
Goyal, R., Kahou, S., Michalski, V., Materzyńska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al.: · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hollywood in homes: Crowdsourcing data collection for activity understanding
Sigurdsson, G.A., Varol, G., Wang, X., Farhadi, A., Laptev, I., Gupta, A.: · 2016
Cited alongside, same era.
Temporal segment networks: Towards good practices for deep action recognition
Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., Van Gool, L.: · 2016
Cited alongside, same era.
Mofap: A multi-level representation for action recognition
Wang, L., Qiao, Y., Tang, X.: · 2016
Cited alongside, same era.
Learning to poke by poking: Experiential learning of intuitive physics
Agrawal, P., Nair, A.V., Abbeel, P., Malik, J., Levine, S.: · 2016
Cited alongside, same era.
The curious robot: Learning visual representations via physical interactions
Pinto, L., Gandhi, D., Han, Y., Park, Y.L., Gupta, A.: · 2016
Cited alongside, same era.
Closest in time.
Twentybn jester dataset: a hand gesture dataset
: · 2017
Closest in time.
The kinetics human action video dataset
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., et al.: · 2017
Closest in time.
Time-contrastive networks: Self-supervised learning from multi-view observation
Sermanet, P., Lynch, C., Hsu, J., Levine, S.: · 2017
Closest in time.
Asynchronous temporal fields for action recognition
Sigurdsson, G.A., Divvala, S., Farhadi, A., Gupta, A.: · 2017
Closest in time.
Fine-grained video classification and captioning
Mahdisoltani, F., Berger, G., Gharbieh, W., Fleet, D., Memisevic, R.: · 2018
Closest in time.