Fetching the paper…
Reading the bibliography…
Understanding temporal dynamics of video is an essential aspect of learning better video representations.
Two-frame motion estimation based on polynomial expansion
Farnebäck, G · 2003
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Simonyan, K. and Zisserman, A · 2014
Earlier work this paper cites.
Learning spatiotemporal features with 3d convolutional networks
Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M · 2015
Earlier work this paper cites.
Shuffle and learn: unsupervised learning using temporal order verification
Misra, I., Zitnick, C. L., and Hebert, M · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J. and Zisserman, A · 2017
Earlier work this paper cites.
The ”something something” video database for learning and evaluating visual common sense
Goyal, R., Kahou, S. E., Michalski, V., Materzyńska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., Hoppe, F., Thurau, C., Bax, I., and Memisevic, R · 2017
Earlier work this paper cites.
The kinetics human action video dataset, 2017
Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., Natsev, P., Suleyman, M., and Zisserman, A · 2017
Earlier work this paper cites.
Unsupervised representation learning by sorting sequences
Lee, H.-Y., Huang, J.-B., Singh, M., and Yang, M.-H · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Noise or signal: The role of image backgrounds in object recognition
Xiao, K., Engstrom, L., Ilyas, A., and Madry, A · 2017
Earlier work this paper cites.
What makes a video a video: Analyzing temporal information in video understanding models and datasets
Huang, D.-A., Ramanathan, V., Mahajan, D., Torresani, L., Paluri, M., Fei-Fei, L., and Niebles, J. C · 2018
Cited alongside, same era.
Training confidence-calibrated classifiers for detecting out-of-distribution samples
Lee, K., Lee, H., Lee, K., and Shin, J · 2018
Cited alongside, same era.
Fixing weight decay regularization in adam
Loshchilov, I. and Hutter, F · 2018
Cited alongside, same era.
A closer look at spatiotemporal convolutions for action recognition
Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., and Paluri, M · 2018
Cited alongside, same era.
Slowfast networks for video recognition
Feichtenhofer, C., Fan, H., Malik, J., and He, K · 2019
Cited alongside, same era.
Deep anomaly detection with outlier exposure
Hendrycks, D., Mazeika, M., and Dietterich, T · 2019
Is space-time attention all you need for video understanding?
Bertasius, G., Wang, H., and Torresani, L · 2021
Later among the works it cites.
Space-time mixing attention for video transformer
Bulat, A., Perez-Rua, J.-M., Sudhakaran, S., Martinez, B., and Tzimiropoulos, G · 2021
Later among the works it cites.
Mask2former for video instance segmentation
Cheng, B., Choudhuri, A., Misra, I., Kirillov, A., Girdhar, R., and Schwing, A. G · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
An image classifier can suffice for video understanding
Fan, Q., Panda, R., et al · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Repair: Removing representation bias by dataset resampling
Li, Y. and Vasconcelos, N · 2019
Cited alongside, same era.
Tsm: Temporal shift module for efficient video understanding
Lin, J., Gan, C., and Han, S · 2019
Cited alongside, same era.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V · 2020
Cited alongside, same era.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2020
Cited alongside, same era.
Vivit: A video vision transformer
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lučić, M., and Schmid, C · 2021
Cited alongside, same era.
Contrast and order representations for video self-supervised learning
Hu, K., Shao, J., Liu, Y., Raj, B., Savvides, M., and Shen, Z · 2021
Later among the works it cites.
Masker: Masked keyword regularization for reliable text classification
Moon, S. J., Mo, S., Lee, K., Lee, J., and Shin, J · 2021
Later among the works it cites.
Neimark, D., Bar, O., Zohar, M., and Asselmann, D · 2021
Later among the works it cites.
Keeping your eye on the ball: Trajectory attention in video transformers
Patrick, M., Campbell, D., Asano, Y. M., Metze, I. M. F., Feichtenhofer, C., Vedaldi, A., Henriques, J., et al · 2021
Later among the works it cites.
Only time can tell: Discovering temporal data for temporal modeling
Sevilla-Lara, L., Zha, S., Yan, Z., Goswami, V., Feiszli, M., and Torresani, L · 2021
Later among the works it cites.
Uniformer: Unified transformer for efficient spatiotemporal representation learning
Li, K., Wang, Y., Gao, P., Song, G., Liu, Y., Li, H., and Qiao, Y · 2022
Closest in time.