Fetching the paper…
Reading the bibliography…
Our world is full of varied actions and moves across specialized domains that we, as humans, strive to identify and understand.
VATEX: A large-scale, high-quality multilingual dataset for video-and-language research
X. Wang, J. Wu, J. Chen, L. Li, Y. Wang, and W. Y. Wang · 1904
Earlier work this paper cites.
Recognizing human actions: a local svm approach
C. Schuldt, I. Laptev, and B. Caputo · 2004
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2009
Earlier work this paper cites.
Hmdb: a large video database for human motion recognition
H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre · 2011
Earlier work this paper cites.
American time use survey, 2013
U.S. Department of Labor · 2013
Earlier work this paper cites.
The pascal visual object classes challenge: A retrospective
M. Everingham, S. M. A. Eslami, L. V. Gool, C. K. I. Williams, J. M. Winn, and A. Zisserman · 2014
Earlier work this paper cites.
The language of actions: Recovering the syntax and semantics of goal-directed human activities
H. Kuehne, A. B. Arslan, and T. Serre · 2014
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
B. G. Fabian Caba Heilbron, Victor Escorcia and J. C. Niebles · 2015
Earlier work this paper cites.
Recognizing fine-grained and composite activities using hand-centric features and script data
M. Rohrbach, A. Rohrbach, M. Regneri, S. Amin, M. Andriluka, M. Pinkal, and B. Schiele · 2015
Earlier work this paper cites.
Action recognition with trajectory-pooled deep-convolutional descriptors
L. Wang, Y. Qiao, and X. Tang · 2015
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering, 2016
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Earlier work this paper cites.
Temporal segment networks: Towards good practices for deep action recognition
L. Wang, Y. Xiong, Z. Wang, Y. Qiao, D. Lin, X. Tang, and L. Van Gool · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
J. Xu, T. Mei, T. Yao, and Y. Rui · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. Berg, and T. Berg · 2016
Earlier work this paper cites.
The" something something" video database for learning and evaluating visual common sense
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al · 2017
Earlier work this paper cites.
The kinetics human action video dataset, 2017
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman · 2017
Earlier work this paper cites.
Dense-captioning events in videos
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
D. Xu, Z. Zhao, J. Xiao, F. Wu, H. Zhang, X. He, and Y. Zhuang · 2017
Earlier work this paper cites.
Towards structured analysis of broadcast badminton videos
A. Ghosh, S. Singh, and C. V. Jawahar · 2018
Earlier work this paper cites.
Ava: A video dataset of spatio-temporally localized atomic visual actions, 2018
C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankar, C. Schmid, and J. Malik · 2018
Earlier work this paper cites.
Localizing moments in video with temporal language
L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell · 2018
Earlier work this paper cites.
Resound: Towards action recognition without representation bias
Y. Li, Y. Li, and N. Vasconcelos · 2018
Earlier work this paper cites.
Sport action recognition with siamese spatio-temporal cnns: Application to table tennis
P.-E. Martin, J. Benois-Pineau, R. Péteri, and J. Morlier · 2018
Earlier work this paper cites.
Towards automatic learning of procedures from web instructional videos
L. Zhou, C. Xu, and J. J. Corso · 2018
Earlier work this paper cites.
Slowfast networks for video recognition
C. Feichtenhofer, H. Fan, J. Malik, and K. He · 2019
Cited alongside, same era.
Tvqa: Localized, compositional video question answering, 2019
J. Lei, L. Yu, M. Bansal, and T. L. Berg · 2019
Cited alongside, same era.
Tsm: Temporal shift module for efficient video understanding
J. Lin, C. Gan, and S. Han · 2019
Cited alongside, same era.
Moments in time dataset: one million videos for event understanding, 2019
M. Monfort, A. Andonian, B. Zhou, K. Ramakrishnan, S. A. Bargal, T. Yan, L. Brown, Q. Fan, D. Gutfruend, C. Vondrick, and A. Oliva · 2019
Cited alongside, same era.
Long-Term Feature Banks for Detailed Video Understanding
C.-Y. Wu, C. Feichtenhofer, H. Fan, K. He, P. Krähenbühl, and R. Girshick · 2019
Cited alongside, same era.
Activitynet-qa: A dataset for understanding complex web videos via question answering, 2019
Z. Yu, D. Xu, J. Yu, T. Yu, Z. Zhao, Y. Zhuang, and D. Tao · 2019
Elasticsearch
Elastic · 2023
Later among the works it cites.
Large language models are zero-shot reasoners, 2023
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa · 2023
Later among the works it cites.
Video-chatgpt: Towards detailed video understanding via large vision and language models, 2023
M. Maaz, H. Rasheed, S. Khan, and F. S. Khan · 2023
Later among the works it cites.
Egoschema: A diagnostic benchmark for very long-form video language understanding, 2023
K. Mangalam, R. Akshulakov, and J. Malik · 2023
Later among the works it cites.
Video-bench: A comprehensive benchmark and toolkit for evaluating video-based large language models, 2023
M. Ning, B. Zhu, Y. Xie, B. Lin, J. Cui, L. Yuan, D. Chen, and L. Yuan · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Condensed movies: Story based retrieval with contextual embeddings, 2020
M. Bain, A. Nagrani, A. Brown, and A. Zisserman · 2020
Cited alongside, same era.
X3d: Expanding architectures for efficient video recognition
C. Feichtenhofer · 2020
Cited alongside, same era.
Fine-grained action recognition on a novel basketball dataset
X. Gu, X. Xue, and F. Wang · 2020
Cited alongside, same era.
Finegym: A hierarchical video dataset for fine-grained action understanding, 2020
D. Shao, Y. Zhao, B. Dai, and D. Lin · 2020
Cited alongside, same era.
A short note on the kinetics-700-2020 human action dataset
L. Smaira, J. Carreira, E. Noland, E. Clancy, A. Wu, and A. Zisserman · 2020
Cited alongside, same era.
Training verifiers to solve math word problems, 2021
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Cited alongside, same era.
V. Pătrăucean, L. Smaira, A. Gupta, A. R. Continente, L. Markeeva, D. Banarse, S. Koppula, J. Heyward, M. Malinowski, Y. Yang, C. Doersch, T. Matejovicova, Y. Sulsky, A. Miech, A. Frechette, H. Klimczak, R. Koster, J. Zhang, S. Winkler, Y. Aytar, S. Osindero, D. Damen, A. Zisserman, and J. Carreira · 2023
Later among the works it cites.
Eva-clip: Improved training techniques for clip at scale
Q. Sun, Y. Fang, L. Wu, X. Wang, and Y. Cao · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi, et al · 2023
Later among the works it cites.
Video-llama: An instruction-tuned audio-visual language model for video understanding, 2023
H. Zhang, X. Li, and L. Bing · 2023
Later among the works it cites.
https://cloud.google.com/vision/docs/ocr
Optical character recognition (ocr) | cloud vision api | google cloud · 2024
Closest in time.
Sharegpt4video: Improving video understanding and generation with better captions, 2024
L. Chen, X. Wei, J. Li, X. Dong, P. Zhang, Y. Zang, Z. Chen, H. Duan, B. Lin, Z. Tang, L. Yuan, Y. Qiao, D. Lin, F. Zhao, and J. Wang · 2024
Closest in time.
Ego-exo4d: Understanding skilled human activity from first- and third-person perspectives, 2024
K. G. et al · 2024
Closest in time.
Google generative ai
Google · 2024
Closest in time.
Video-lavit: Unified video-language pre-training with decoupled visual-motional tokenization, 2024
Y. Jin, Z. Sun, K. Xu, K. Xu, L. Chen, H. Jiang, Q. Huang, C. Song, Y. Liu, D. Zhang, Y. Song, K. Gai, and Y. Mu · 2024
Closest in time.
Mvbench: A comprehensive multi-modal video understanding benchmark, 2024
K. Li, Y. Wang, Y. He, Y. Li, Y. Wang, Y. Liu, Z. Wang, J. Xu, G. Chen, P. Luo, L. Wang, and Y. Qiao · 2024
Closest in time.
Morevqa: Exploring modular reasoning models for video question answering, 2024
J. Min, S. Buch, A. Nagrani, M. Cho, and C. Schmid · 2024
Closest in time.
Calculating costs for vision api, 2024
OpenAI · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
G. Team · 2024
Closest in time.
Qwen2-vl, 2024
Q. team · 2024
Closest in time.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024
P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, Y. Fan, K. Dang, M. Du, X. Ren, R. Men, D. Liu, C. Zhou, J. Zhou, and J. Lin · 2024
Closest in time.
Finesports: A multi-person hierarchical sports video dataset for fine-grained action understanding
J. Xu, G. Zhao, S. Yin, W. Zhou, and Y. Peng · 2024
Closest in time.
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen · 2024
Closest in time.
Llava-next: A strong zero-shot video understanding model, April 2024b
Y. Zhang, B. Li, h. Liu, Y. j. Lee, L. Gui, D. Fu, J. Feng, Z. Liu, and C. Li · 2024
Closest in time.