Fetching the paper…
Reading the bibliography…
In this paper, we introduce a novel visual representation learning which relies on a handful of adaptively learned tokens, and which is applicable to both image and video understanding tasks.
C. Feichtenhofer, A. Pinz, and A. Zisserman, “Convolutional two-stream network fusion for video action recognition,” in
1941
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in
2009
Earlier work this paper cites.
S. Ji, W. Xu, M. Yang, and K. Yu, “3d convolutional neural networks for human action recognition,”
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” in
2014
Earlier work this paper cites.
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta, “Hollywood in homes: Crowdsourcing data collection for activity understanding,” in
2016
Earlier work this paper cites.
C. Feichtenhofer, A. Pinz, and R. Wildes, “Spatiotemporal residual networks for video action recognition,” in
2016
Earlier work this paper cites.
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
K. Hara, H. Kataoka, and Y. Satoh, “Learning spatio-temporal features with 3d residual networks for action recognition,” in
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
C. Sun, A. Shrivastava, S. Singh, and A. Gupta, “Revisiting unreasonable effectiveness of data in deep learning era,” in
2017
Earlier work this paper cites.
W. Rawat and Z. Wang, “Deep convolutional neural networks for image classification: A comprehensive review,”
2017
Earlier work this paper cites.
C. Feichtenhofer, A. Pinz, and R. P. Wildes, “Spatiotemporal multiplier networks for video action recognition,” in
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” in
2018
Earlier work this paper cites.
J. Hu, L. Shen, S. Albanie, G. Sun, and E. Wu, “Squeeze-and-excitation networks,” in
2018
Earlier work this paper cites.
X. Wang, R. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in
2018
Earlier work this paper cites.
X. Wang and A. Gupta, “Videos as space-time region graphs,” in
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
B. Zhou, A. Andonian, A. Oliva, and A. Torralba, “Temporal relational reasoning in videos,” in
2018
Earlier work this paper cites.
M. Zolfaghari, K. Singh, and T. Brox, “Eco: efficient convolutional network for online video understanding,” in
2018
Cited alongside, same era.
S. Xie, C. Sun, J. Huang, Z. Tu, and K. Murphy, “Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification,” in
2018
Cited alongside, same era.
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” in
2019
Cited alongside, same era.
A. Piergiovanni, A. Angelova, A. Toshev, and M. S. Ryoo, “Evolving space-time neural architectures for videos,”
2019
Cited alongside, same era.
N. Hussein, E. Gavves, and A. W. Smeulders, “Timeception for complex action recognition,” in
2019
Cited alongside, same era.
D. Weissenborn, O. Täckström, and J. Uszkoreit, “Scaling autoregressive video models,” in
2020
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in
2021
Closest in time.
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić, and C. Schmid, “ViViT: A video vision transformer,” in
2021
Closest in time.
G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” in
2021
Closest in time.
M. S. Ryoo, A. Piergiovanni, A. Arnab, M. Dehghani, and A. Angelova, “Tokenlearner: Adaptive space-time tokenization for videos,” in
2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Korbar, D. Tran, and L. Torresani, “Scsampler: Sampling salient clips from video for efficient action recognition,” in
2019
Cited alongside, same era.
2019
Cited alongside, same era.
P. Ramachandran, N. Parmar, A. Vaswani, I. Bello, A. Levskaya, and J. Shlens, “Stand-alone self-attention in vision models,” in
2019
Cited alongside, same era.
R. Girdhar, J. Carreira, C. Doersch, and A. Zisserman, “Video action transformer network,” in
2019
Cited alongside, same era.
A. Piergiovanni, A. Angelova, and M. S. Ryoo, “Evolving losses for unsupervised video representation learning,” in
2020
Cited alongside, same era.
J.-B. Alayrac, A. Recasens, R. Schneider, R. Arandjelović, J. Ramapuram, J. D. Fauw, L. Smaira, S. Dieleman, and A. Zisserman, “Self-supervised multimodal versatile networks,” in
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Closest in time.
2021
Closest in time.
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer, “Scaling vision transformers,”
2021
Closest in time.
2021
Closest in time.
Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,”
2021
Closest in time.
2021
Closest in time.
D. Kondratyuk, L. Yuan, Y. Li, L. Zhang, M. Tan, M. Brown, and B. Gong, “Movinets: Mobile video networks for efficient video recognition,” in
2021
Closest in time.
C.-Y. Wu and P. Krahenbuhl, “Towards long-form video understanding,” in
2021
Closest in time.
A. Piergiovanni, A. Angelova, and M. S. Ryoo, “Tiny video networks,” in
2021
Closest in time.
L. Xu, Y. Guan, S. Jin, W. Liu, C. Qian, P. Luo, W. Ouyang, and X. Wang, “Vipnas: Efficient video pose estimation via neural architecture search,” in
2021
Closest in time.
H. Touvron, M. Cord, A. Sablayrolles, G. Synnaev, and H. Jegou, “Going deeper with image transformers,” in
2021
Closest in time.
B. Graham, A. El-Nouby, H. Touvron, P. Stock, A. Joulin, H. Jegou, and M. Douze, “Levit: A vision transformer in convnet’s clothing for faster inference,” in
2021
Closest in time.
L. Yuan, Y. Chen, T. Wang, W. Yu, Y. Shi, F. E. Tay, J. Feng, and S. Yan, “Tokensto-token vit: Training vision transformers from scratch on imagenet,” in
2021
Closest in time.
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in
2021
Closest in time.
A. Jaegle, F. Gimeno, A. Brock, A. Zisserman, O. Vinyals, and J. Carreira, “Perceiver: General perception with iterative attention,” in
2021
Closest in time.
J. Lee, Y. Lee, J. Kim, A. R. Kosiorek, S. Choi, and Y. W. Teh, “Set transformer: A framework for attention-based permutation-invariant neural networks,” in
2021
Closest in time.