Fetching the paper…
Reading the bibliography…
Human action understanding is crucial for the advancement of multimodal systems.
Semi-supervised learning literature survey
Zhu, X. J. 2005 · 2005
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A. 2020 · 2010
Earlier work this paper cites.
HMDB: A large video database for human motion recognition
Kuehne, H.; Jhuang, H.; Garrote, E.; Poggio, T.; and Serre, T. 2011 · 2011
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Soomro, K. 2012 · 2012
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. 2015 · 2015
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
Wang, L.; Xiong, Y.; Wang, Z.; Qiao, Y.; Lin, D.; Tang, X.; and Van Gool, L. 2016 · 2016
Earlier work this paper cites.
Quo vadis, action recognition? a new model and the kinetics dataset
Carreira, J.; and Zisserman, A. 2017 · 2017
Earlier work this paper cites.
Improved Regularization of Convolutional Neural Networks with Cutout
DeVries, T. 2017 · 2017
Earlier work this paper cites.
The” something something” video database for learning and evaluating visual common sense
Goyal, R.; Ebrahimi Kahou, S.; Michalski, V.; Materzynska, J.; Westphal, S.; Kim, H.; Haenel, V.; Fruend, I.; Yianilos, P.; Mueller-Freitag, M.; et al. 2017 · 2017
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Tarvainen, A.; and Valpola, H. 2017 · 2017
Earlier work this paper cites.
Find and focus: Retrieve and localize video events with natural language queries
Shao, D.; Xiong, Y.; Zhao, Y.; Huang, Q.; Qiao, Y.; and Lin, D. 2018 · 2018
Earlier work this paper cites.
Temporal segment networks for action recognition in videos
Wang, L.; Xiong, Y.; Wang, Z.; Qiao, Y.; Lin, D.; Tang, X.; and Van Gool, L. 2018 · 2018
Earlier work this paper cites.
Temporal relational reasoning in videos
Zhou, B.; Andonian, A.; Oliva, A.; and Torralba, A. 2018 · 2018
Earlier work this paper cites.
Learning motion in feature space: Locally-consistent deformable convolution networks for fine-grained action detection
Mac, K.-N. C.; Joshi, D.; Yeh, R. A.; Xiong, J.; Feris, R. S.; and Do, M. N. 2019 · 2019
Earlier work this paper cites.
Cutmix: Regularization strategy to train strong classifiers with localizable features
Yun, S.; Han, D.; Oh, S. J.; Chun, S.; Choe, J.; and Yoo, Y. 2019 · 2019
Earlier work this paper cites.
Randaugment: Practical automated data augmentation with a reduced search space
Cubuk, E. D.; Zoph, B.; Shlens, J.; and Le, Q. V. 2020 · 2020
Earlier work this paper cites.
Memory-augmented dense predictive coding for video representation learning
Han, T.; Xie, W.; and Zisserman, A. 2020 · 2020
Earlier work this paper cites.
Remixmatch: Semi-supervised learning with distribution matching and augmentation anchoring
Kurakin, A.; Raffel, C.; Berthelot, D.; Cubuk, E. D.; Zhang, H.; Sohn, K.; and Carlini, N. 2020 · 2020
Earlier work this paper cites.
Fixmatch: Simplifying semi-supervised learning with consistency and confidence
Sohn, K.; Berthelot, D.; Carlini, N.; Zhang, Z.; Zhang, H.; Raffel, C. A.; Cubuk, E. D.; Kurakin, A.; and Li, C.-L. 2020 · 2020
Earlier work this paper cites.
Unsupervised data augmentation for consistency training
Xie, Q.; Dai, Z.; Hovy, E.; Luong, T.; and Le, Q. 2020 · 2020
Earlier work this paper cites.
Temporal pyramid network for action recognition
Yang, C.; Xu, Y.; Shi, J.; Dai, B.; and Zhou, B. 2020 · 2020
Cited alongside, same era.
Is space-time attention all you need for video understanding?
Bertasius, G.; Wang, H.; and Torresani, L. 2021 · 2021
Cited alongside, same era.
Video pose distillation for few-shot, fine-grained sports action recognition
Hong, J.; Fisher, M.; Gharbi, M.; and Fatahalian, K. 2021 · 2021
Cited alongside, same era.
Videossl: Semi-supervised learning for video classification
Jing, L.; Parag, T.; Wu, Z.; Tian, Y.; and Wang, H. 2021 · 2021
Cited alongside, same era.
Joint learning on the hierarchy representation for fine-grained human action recognition
Leong, M. C.; Tan, H. L.; Zhang, H.; Li, L.; Lin, F.; and Lim, J. H. 2021 · 2021
Cited alongside, same era.
M3net: multi-view encoding, matching, and fusion for few-shot fine-grained action recognition
Tang, H.; Liu, J.; Yan, S.; Yan, R.; Li, Z.; and Tang, J. 2023 · 2023
Later among the works it cites.
Semi-Supervised Action Recognition From Temporal Augmentation Using Curriculum Learning
Tong, A.; Tang, C.; and Wang, W. 2023 · 2023
Later among the works it cites.
Svformer: Semi-supervised video transformer for action recognition
Xing, Z.; Dai, Q.; Hu, H.; Chen, J.; Wu, Z.; and Jiang, Y.-G. 2023 · 2023
Later among the works it cites.
Videoglue: Video general understanding evaluation of foundation models
Yuan, L.; Gundavarapu, N. B.; Zhao, L.; Zhou, H.; Cui, Y.; Jiang, L.; Yang, X.; Jia, M.; Weyand, T.; Friedman, L.; et al. 2023 · 2023
Later among the works it cites.
Learning representational invariances for data-efficient action recognition
Zou, Y.; Choi, J.; Wang, Q.; and Huang, J.-B. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rizve, M. N.; Duarte, K.; Rawat, Y. S.; and Shah, M. 2021 · 2021
Cited alongside, same era.
Semi-supervised action recognition with temporal contrastive learning
Singh, A.; Chakraborty, O.; Varshney, A.; Panda, R.; Feris, R.; Saenko, K.; and Das, A. 2021 · 2021
Cited alongside, same era.
Few-shot fine-grained action recognition via bidirectional attention and contrastive meta-learning
Wang, J.; Wang, Y.; Liu, S.; and Li, A. 2021 · 2021
Cited alongside, same era.
Multiview pseudo-labeling for semi-supervised learning from video
Xiong, B.; Fan, H.; Grauman, K.; and Feichtenhofer, C. 2021 · 2021
Cited alongside, same era.
TCLR: Temporal contrastive learning for video representation
Dave, I.; Gupta, R.; Rizve, M. N.; and Shah, M. 2022 · 2022
Cited alongside, same era.
Learn2augment: learning to composite videos for data augmentation in action recognition
Gowda, S. N.; Rohrbach, M.; Keller, F.; and Sevilla-Lara, L. 2022 · 2022
Cited alongside, same era.
Combined CNN transformer encoder for enhanced fine-grained human action recognition
Leong, M. C.; Zhang, H.; Tan, H. L.; Li, L.; and Lim, J. H. 2022 · 2022
Cited alongside, same era.
Sphinx-x: Scaling data and parameters for a family of multi-modal large language models
Gao, P.; Zhang, R.; Liu, C.; Qiu, L.; Huang, S.; Lin, W.; Zhao, S.; Geng, S.; Lin, Z.; Jin, P.; et al. 2024 · 2024
Later among the works it cites.
Crest: Cross-modal resonance through evidential deep learning for enhanced zero-shot learning
Huang, H.; Qiao, X.; Chen, Z.; Chen, H.; Li, B.; Sun, Z.; Chen, M.; and Li, X. 2024 · 2024
Later among the works it cites.
Mvbench: A comprehensive multi-modal video understanding benchmark
Li, K.; Wang, Y.; He, Y.; Li, Y.; Wang, Y.; Liu, Y.; Wang, Z.; Xu, J.; Chen, G.; Luo, P.; et al. 2024 · 2024
Later among the works it cites.
Beyond uncertainty: Evidential deep learning for robust video temporal grounding
Ma, K.; Huang, H.; Chen, J.; Chen, H.; Ji, P.; Zang, X.; Fang, H.; Ban, C.; Sun, H.; Chen, M.; et al. 2024 · 2024
Later among the works it cites.
GPT-4 System Card
OpenAI. 2024 · 2024
Later among the works it cites.
CMT: Cross Modulation Transformer with Hybrid Loss for Pansharpening
Shu, W.-J.; Dou, H.-X.; Wen, R.; Wu, X.; and Deng, L.-J. 2024 · 2024
Later among the works it cites.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Tong, S.; Liu, Z.; Zhai, Y.; Ma, Y.; LeCun, Y.; and Xie, S. 2024 · 2024
Later among the works it cites.
Chatgpt for robotics: Design principles and model abilities
Vemprala, S. H.; Bonatti, R.; Bucker, A.; and Kapoor, A. 2024 · 2024
Later among the works it cites.
FSC: Few-point Shape Completion
Wu, X.; Wu, X.; Luan, T.; Bai, Y.; Lai, Z.; and Yuan, J. 2024 · 2024
Later among the works it cites.
Urbanclip: Learning text-enhanced urban region profiling with contrastive language-image pretraining from the web
Yan, Y.; Wen, H.; Zhong, S.; Chen, W.; Chen, H.; Wen, Q.; Zimmermann, R.; and Liang, Y. 2024b · 2024
Later among the works it cites.
Zhang, P.; Dong, X.; Zang, Y.; Cao, Y.; Qian, R.; Chen, L.; Guo, Q.; Duan, H.; Wang, B.; Ouyang, L.; Zhang, S.; Zhang, W.; Li, Y.; Gao, Y.; Sun, P.; Zhang, X.; Li, W.; Li, J.; Wang, W.; Yan, H.; He, C.; Zhang, X.; Chen, K.; Dai, J.; Qiao, Y.; Lin, D.; and Wang, J. 2024 · 2024
Later among the works it cites.
VideoPrism: A Foundational Visual Encoder for Video Understanding
Zhao, L.; Gundavarapu, N. B.; Yuan, L.; Zhou, H.; Yan, S.; Sun, J. J.; Friedman, L.; Qian, R.; Weyand, T.; Zhao, Y.; et al. 2024 · 2024
Later among the works it cites.
VideoGen-of-Thought: A Collaborative Framework for Multi-Shot Video Generation
Zheng, M.; Xu, Y.; Huang, H.; Ma, X.; Liu, Y.; Shu, W.; Pang, Y.; Tang, F.; Chen, Q.; Yang, H.; et al. 2024 · 2024
Later among the works it cites.
Finepseudo: improving pseudo-labelling through temporal-alignablity for semi-supervised fine-grained action recognition
Dave, I. R.; Rizve, M. N.; and Shah, M. 2025 · 2025
Closest in time.