Fetching the paper…
Reading the bibliography…
Existing research of video understanding still struggles to achieve in-depth comprehension and reasoning in complex videos, primarily due to the under-exploration of two key bottlenecks: fine-grained spatial-temporal perceptive understanding and cognitive-level video scene comprehension.
Activitynet: A large-scale video benchmark for human activity understanding
Heilbron, F. C., Escorcia, V., Ghanem, B., and Niebles, J. C · 2015
Earlier work this paper cites.
MSR-VTT: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., and Rui, Y · 2016
Earlier work this paper cites.
Real-time video super-resolution with spatio-temporal networks and motion compensation
Caballero, J., Ledig, C., Aitken, A., Acosta, A., Totz, J., Wang, Z., and Shi, W · 2017
Earlier work this paper cites.
Image generation from scene graphs
Johnson, J., Gupta, A., and Fei-Fei, L · 2018
Earlier work this paper cites.
Tvqa: Localized, compositional video question answering
Lei, J., Yu, L., Bansal, M., and Berg, T · 2018
Earlier work this paper cites.
Eco: Efficient convolutional network for online video understanding
Zolfaghari, M., Singh, K., and Brox, T · 2018
Earlier work this paper cites.
Tsm: Temporal shift module for efficient video understanding
Lin, J., Gan, C., and Han, S · 2019
Earlier work this paper cites.
Social-iq: A question answering benchmark for artificial social intelligence
Zadeh, A., Chan, M., Liang, P. P., Tong, E., and Morency, L.-P · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Earlier work this paper cites.
A generalization of transformer networks to graphs
Dwivedi, V. P. and Bresson, X · 2020
Earlier work this paper cites.
Action genome: Actions as compositions of spatio-temporal scene graphs
Ji, J., Krishna, R., Fei-Fei, L., and Niebles, J. C · 2020
Earlier work this paper cites.
What is more likely to happen next? video-and-language future event prediction
Lei, J., Yu, L., Berg, T., and Bansal, M · 2020
Earlier work this paper cites.
Gps-net: Graph property sensing network for scene graph generation
Lin, X., Ding, C., Zeng, J., and Tao, D · 2020
Cited alongside, same era.
Frozen in time: A joint video and image encoder for end-to-end retrieval
Bain, M., Nagrani, A., Varol, G., and Zisserman, A · 2021
Cited alongside, same era.
Is space-time attention all you need for video understanding?
Bertasius, G., Wang, H., and Torresani, L · 2021
Cited alongside, same era.
Spatial-temporal transformer for dynamic scene graph generation
Cong, Y., Liao, W., Ackermann, H., Rosenhahn, B., and Yang, M. Y · 2021
Cited alongside, same era.
Video transformer network
Neimark, D., Bar, O., Zohar, M., and Asselmann, D · 2021
Cited alongside, same era.
Star: A benchmark for situated reasoning in real-world videos
Wu, B., Yu, S., Chen, Z., Tenenbaum, J. B., and Gan, C · 2021
Cited alongside, same era.
Automatic chain of thought prompting in large language models
Zhang, Z., Zhang, A., Li, M., and Smola, A · 2022
Later among the works it cites.
Vicuna: An open-source chatbot impressing gpt-4 with 90
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Later among the works it cites.
Large language models are temporal and causal reasoners for video question answering
Ko, D., Lee, J., Kang, W.-Y., Roh, B., and Kim, H · 2023
Later among the works it cites.
Video-llava: Learning united visual representation by alignment before projection
Lin, B., Ye, Y., Zhu, B., Cui, J., Ning, M., Jin, P., and Yuan, L · 2023
Later among the works it cites.
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Next-qa: Next phase of question-answering to explaining temporal actions
Xiao, J., Shang, X., Yao, A., and Chua, T · 2021
Cited alongside, same era.
Tokens-to-token vit: Training vision transformers from scratch on imagenet
Yuan, L., Chen, Y., Wang, T., Yu, W., Shi, Y., Jiang, Z.-H., Tay, F. E., Feng, J., and Yan, S · 2021
Cited alongside, same era.
Matching structure for dual learning
Fei, H., Wu, S., Ren, Y., and Zhang, M · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., and Kalyan, A · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Later among the works it cites.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Maaz, M., Rasheed, H. A., Khan, S. H., and Khan, F. S · 2023
Later among the works it cites.
Wang, X., Liang, J., Wang, C.-K., Deng, K., Lou, Y., Lin, M., and Yang, S · 2023
Later among the works it cites.
Hitea: Hierarchical temporal-aware video-language pre-training
Ye, Q., Xu, G., Yan, M., Xu, H., Qian, Q., Zhang, J., and Huang, F · 2023
Later among the works it cites.
Self-chained image-language model for video localization and question answering
Yu, S., Cho, J., Yadav, P., and Bansal, M · 2023
Later among the works it cites.
Enhancing video-language representations with structural spatio-temporal alignment
Fei, H., Wu, S., Zhang, M., Zhang, M., Chua, T.-S., and Yan, S · 2024
Closest in time.
Imagine that! abstract-to-intricate text-to-image synthesis with scene graph hallucination diffusion
Wu, S., Fei, H., Zhang, H., and Chua, T.-S · 2024
Closest in time.
Reverse multi-choice dialogue commonsense inference with graph-of-thought
Zheng, L., Fei, H., Li, F., Li, B., Liao, L., Ji, D., and Teng, C · 2024
Closest in time.