Fetching the paper…
Reading the bibliography…
Adapting image models to the video domain has emerged as an efficient paradigm for solving video recognition tasks.
Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., Serre, T.: Hmdb: a large video database for human motion recognition. In: Int. Conf. Comput. Vis. pp. 2556–2563. IEEE (2011)
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
Ba, L.J., Kiros, J.R., Hinton, G.E.: Layer normalization. arXiv preprint arXiv:1607.06450 (2016)
2016
Earlier work this paper cites.
Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., Gool, L.V.: Temporal segment networks: Towards good practices for deep action recognition. In: Eur. Conf. Comput. Vis. vol. 9912, pp. 20–36 (2016)
2016
Earlier work this paper cites.
Carreira, J., Zisserman, A.: Quo vadis, action recognition? A new model and the kinetics dataset. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 4724–4733 (2017)
2017
Earlier work this paper cites.
Goyal, R., Kahou, S.E., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fründ, I., Yianilos, P., Mueller-Freitag, M., Hoppe, F., Thurau, C., Bax, I., Memisevic, R.: The "something something" video database for learning and evaluating visual common sense. In: Int. Conf. Comput. Vis. pp. 5843–5851. IEEE Computer Society (2017)
2017
Earlier work this paper cites.
Gu, C., Sun, C., Ross, D.A., Vondrick, C., Pantofaru, C., Li, Y., Vijayanarasimhan, S., Toderici, G., Ricco, S., Sukthankar, R., et al.: Ava: A video dataset of spatio-temporally localized atomic visual actions. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 6047–6056 (2018)
2018
Earlier work this paper cites.
Li, Y., Li, Y., Vasconcelos, N.: Resound: Towards action recognition without representation bias. In: Eur. Conf. Comput. Vis. pp. 513–528 (2018)
2018
Earlier work this paper cites.
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al.: Improving language understanding by generative pre-training. OpenAI blog (2018)
2018
Earlier work this paper cites.
Xie, S., Sun, C., Huang, J., Tu, Z., Murphy, K.: Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. In: Eur. Conf. Comput. Vis. pp. 305–321 (2018)
2018
Earlier work this paper cites.
Zhou, B., Andonian, A., Oliva, A., Torralba, A.: Temporal relational reasoning in videos. In: Eur. Conf. Comput. Vis. vol. 11205, pp. 831–846 (2018)
2018
Earlier work this paper cites.
Devlin, J., Chang, M., Lee, K., Toutanova, K.: Bert: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of NAACL-HLT. pp. 4171–4186 (2019)
2019
Earlier work this paper cites.
Feichtenhofer, C., Fan, H., Malik, J., He, K.: Slowfast networks for video recognition. In: Int. Conf. Comput. Vis. pp. 6201–6210 (2019)
2019
Earlier work this paper cites.
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., Attariyan, M., Gelly, S.: Parameter-efficient transfer learning for NLP. In: Int. Conf. Mach. Learn. vol. 97, pp. 2790–2799 (2019)
2019
Earlier work this paper cites.
Michel, P., Levy, O., Neubig, G.: Are sixteen heads really better than one? In: Adv. Neural Inform. Process. Syst. pp. 14014–14024 (2019)
2019
Earlier work this paper cites.
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al.: Language models are unsupervised multitask learners. OpenAI blog 1
2019
Earlier work this paper cites.
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J.D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.: Language models are few-shot learners. In: Adv. Neural Inform. Process. Syst. vol. 33, pp. 1877–1901 (2020)
2020
Earlier work this paper cites.
Cubuk, E.D., Zoph, B., Shlens, J., Le, Q.: Randaugment: Practical automated data augmentation with a reduced search space. In: Adv. Neural Inform. Process. Syst. (2020)
2020
Earlier work this paper cites.
Feichtenhofer, C.: X3D: expanding architectures for efficient video recognition. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 200–210 (2020)
2020
Earlier work this paper cites.
He, K., Fan, H., Wu, Y., Xie, S., Girshick, R.B.: Momentum contrast for unsupervised visual representation learning. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 9726–9735 (2020)
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Li, Y., Ji, B., Shi, X., Zhang, J., Kang, B., Wang, L.: TEA: temporal excitation and aggregation for action recognition. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 906–915 (2020)
2020
Earlier work this paper cites.
Pfeiffer, J., Rücklé, A., Poth, C., Kamath, A., Vulić, I., Ruder, S., Cho, K., Gurevych, I.: Adapterhub: A framework for adapting transformers. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. pp. 46–54 (2020)
2020
Earlier work this paper cites.
Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., Batra, D.: Grad-cam: Visual explanations from deep networks via gradient-based localization. Int. J. Comput. Vis. 128
2020
Earlier work this paper cites.
Zhong, Z., Zheng, L., Kang, G., Li, S., Yang, Y.: Random erasing data augmentation. In: AAAI Conf. Artif. Intell. pp. 13001–13008 (2020)
2020
Earlier work this paper cites.
Arnab, A., Dehghani, M., Heigold, G., Sun, C., Lucic, M., Schmid, C.: Vivit: A video vision transformer. In: Int. Conf. Comput. Vis. pp. 6816–6826 (2021)
2021
Earlier work this paper cites.
Bertasius, G., Wang, H., Torresani, L.: Is space-time attention all you need for video understanding? In: Int. Conf. Mach. Learn. vol. 139, pp. 813–824 (2021)
2021
Earlier work this paper cites.
Bulat, A., Pérez-Rúa, J., Sudhakaran, S., Martínez, B., Tzimiropoulos, G.: Space-time mixing attention for video transformer. In: Adv. Neural Inform. Process. Syst. pp. 19594–19607 (2021)
2021
Earlier work this paper cites.
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., Joulin, A.: Emerging properties in self-supervised vision transformers. In: Int. Conf. Comput. Vis. pp. 9630–9640 (2021)
2021
Cited alongside, same era.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., Houlsby, N.: An image is worth 16x16 words: Transformers for image recognition at scale. In: Int. Conf. Learn. Represent. (2021)
2021
Cited alongside, same era.
Fan, H., Xiong, B., Mangalam, K., Li, Y., Yan, Z., Malik, J., Feichtenhofer, C.: Multiscale vision transformers. In: Int. Conf. Comput. Vis. pp. 6804–6815 (2021)
2021
Cited alongside, same era.
Fan, Q., Chen, C.F., Panda, R.: Can an image classifier suffice for action recognition? In: Int. Conf. Learn. Represent. (2021)
2021
Cited alongside, same era.
Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., Dong, L., Wei, F., Guo, B.: Swin transformer V2: scaling up capacity and resolution. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 11999–12009 (2022)
2022
Later among the works it cites.
Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., Hu, H.: Video swin transformer. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 3192–3201 (2022)
2022
Later among the works it cites.
Ni, B., Peng, H., Chen, M., Zhang, S., Meng, G., Fu, J., Xiang, S., Ling, H.: Expanding language-image pretrained models for general video recognition. In: Eur. Conf. Comput. Vis. vol. 13664, pp. 1–18 (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lester, B., Al-Rfou, R., Constant, N.: The power of scale for parameter-efficient prompt tuning. In: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. pp. 3045–3059 (2021)
2021
Cited alongside, same era.
Li, X.L., Liang, P.: Prefix-tuning: Optimizing continuous prompts for generation. In: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). pp. 4582–4597 (2021)
2021
Cited alongside, same era.
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., Guo, B.: Swin transformer: Hierarchical vision transformer using shifted windows. In: Int. Conf. Comput. Vis. pp. 9992–10002 (2021)
2021
Cited alongside, same era.
Liu, Z., Wang, L., Wu, W., Qian, C., Lu, T.: TAM: temporal adaptive module for video recognition. In: Int. Conf. Comput. Vis. pp. 13688–13698 (2021)
2021
Cited alongside, same era.
Pfeiffer, J., Kamath, A., Rücklé, A., Cho, K., Gurevych, I.: Adapterfusion: Non-destructive task composition for transfer learning. In: Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. pp. 487–503 (2021)
2021
Cited alongside, same era.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., Sutskever, I.: Learning transferable visual models from natural language supervision. In: Int. Conf. Mach. Learn. vol. 139, pp. 8748–8763 (2021)
2021
Cited alongside, same era.
Wang, L., Tong, Z., Ji, B., Wu, G.: TDN: temporal difference networks for efficient action recognition. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 1895–1904 (2021)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Pan, J., Lin, Z., Zhu, X., Shao, J., Li, H.: St-adapter: Parameter-efficient image-to-video transfer learning. In: Adv. Neural Inform. Process. Syst. (2022)
2022
Later among the works it cites.
Tan, J., Zhao, X., Shi, X., Kang, B., Wang, L.: Pointtad: Multi-label temporal action detection with learnable query points. NIPS 35
2022
Later among the works it cites.
Tong, Z., Song, Y., Wang, J., Wang, L.: Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. In: Adv. Neural Inform. Process. Syst. (2022)
2022
Later among the works it cites.
2022
Later among the works it cites.
Wang, R., Chen, D., Wu, Z., Chen, Y., Dai, X., Liu, M., Jiang, Y., Zhou, L., Yuan, L.: BEVT: BERT pretraining of video transformers. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 14713–14723 (2022)
2022
Later among the works it cites.
Xiang, W., Li, C., Wang, B., Wei, X., Hua, X., Zhang, L.: Spatiotemporal self-attention modeling with temporal patch shift for action recognition. In: Eur. Conf. Comput. Vis. vol. 13663, pp. 627–644 (2022)
2022
Later among the works it cites.
Yan, S., Xiong, X., Arnab, A., Lu, Z., Zhang, M., Sun, C., Schmid, C.: Multiview transformers for video recognition. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 3323–3333 (2022)
2022
Later among the works it cites.
Zaken, E.B., Goldberg, Y., Ravfogel, S.: Bitfit: Simple parameter-efficient fine-tuning for transformer-based masked language-models. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). pp. 1–9 (2022)
2022
Later among the works it cites.
Zhai, X., Kolesnikov, A., Houlsby, N., Beyer, L.: Scaling vision transformers. In: IEEE Conf. Comput. Vis. Pattern Recog. pp. 1204–1213 (2022)
2022
Later among the works it cites.
Zhang, Y., Zhou, K., Liu, Z.: Neural prompt search. arXiv preprint arXiv:2206.04673 (2022)
2022
Later among the works it cites.
Li, K., Wang, Y., He, Y., Li, Y., Wang, Y., Wang, L., Qiao, Y.: Uniformerv2: Unlocking the potential of image vits for video understanding. In: Int. Conf. Comput. Vis. pp. 1632–1643 (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Tu, S., Dai, Q., Wu, Z., Cheng, Z., Hu, H., Jiang, Y.: Implicit temporal modeling with learnable alignment for video recognition. In: Int. Conf. Comput. Vis. (2023)
2023
Closest in time.
Wang, L., Huang, B., Zhao, Z., Tong, Z., He, Y., Wang, Y., Wang, Y., Qiao, Y.: Videomae V2: scaling video masked autoencoders with dual masking. In: IEEE Conf. Comput. Vis. Pattern Recog. (2023)
2023
Closest in time.
Wu, W., Sun, Z., Ouyang, W.: Revisiting classifier: Transferring vision-language models for video recognition. In: AAAI Conf. Artif. Intell. pp. 2847–2855 (2023)
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
Yang, T., Zhu, Y., Xie, Y., Zhang, A., Chen, C., Li, M.: Aim: Adapting image models for efficient video action recognition. In: Int. Conf. Learn. Represent. (2023)
2023
Closest in time.
Zhang, G., Zhu, Y., Wang, H., Chen, Y., Wu, G., Wang, L.: Extracting motion and appearance via inter-frame attention for efficient video frame interpolation. In: IEEE Conf. Comput. Vis. Pattern Recog. (2023)
2023
Closest in time.
2024
Closest in time.
Wang, Y., He, Y., Li, Y., Li, K., Yu, J., Ma, X., Li, X., Chen, G., Chen, X., Wang, Y., et al.: Internvid: A large-scale video-text dataset for multimodal understanding and generation. In: ICLR (2024)
2024
Closest in time.
2024
Closest in time.
Zhu, Y., Zhang, G., Tan, J., Wu, G., Wang, L.: Dual detrs for multi-label temporal action detection. In: CVPR. pp. 18559–18569 (2024)
2024
Closest in time.