Fetching the paper…
Reading the bibliography…
Rapid advancements have been made in extending Large Language Models (LLMs) to Large Multi-modal Models (LMMs).
B. K. Horn and B. G. Schunck, “Determining optical flow,” Artificial intelligence , vol. 17, no. 1-3, pp. 185–203, 1981
1981
Earlier work this paper cites.
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta, “Hollywood in homes: Crowdsourcing data collection for activity understanding,” in Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14 . Springer, 2016, pp. 510–526
2016
Earlier work this paper cites.
Y. Jang, Y. Song, Y. Yu, Y. Kim, and G. Kim, “Tgif-qa: Toward spatio-temporal reasoning in visual question answering,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 2758–2766
2017
Earlier work this paper cites.
R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag et al. , “The” something something” video database for learning and evaluating visual common sense,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 5842–5850
2017
Earlier work this paper cites.
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
L. Zhou, C. Xu, and J. Corso, “Towards automatic learning of procedures from web instructional videos,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 32, no. 1, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
X. Wang, J. Wu, J. Chen, L. Li, Y.-F. Wang, and W. Y. Wang, “Vatex: A large-scale, high-quality multilingual dataset for video-and-language research,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 4581–4591
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
M. Bain, A. Nagrani, G. Varol, and A. Zisserman, “Frozen in time: A joint video and image encoder for end-to-end retrieval,” in IEEE International Conference on Computer Vision , 2021
2021
Earlier work this paper cites.
J. Xiao, X. Shang, A. Yao, and T.-S. Chua, “Next-qa: Next phase of question-answering to explaining temporal actions,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 9777–9786
2021
Earlier work this paper cites.
A. Yang, A. Miech, J. Sivic, I. Laptev, and C. Schmid, “Just ask: Learning to answer questions from millions of narrated videos,” in ICCV , 2021
2021
Earlier work this paper cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” Advances in neural information processing systems , vol. 35, pp. 23 716–23 736, 2022
2022
Earlier work this paper cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al. , “Laion-5b: An open large-scale dataset for training next generation image-text models,” Advances in Neural Information Processing Systems , vol. 35, pp. 25 278–25 294, 2022
2022
Earlier work this paper cites.
J. Gu, X. Meng, G. Lu, L. Hou, N. Minzhe, X. Liang, L. Yao, R. Huang, W. Zhang, X. Jiang et al. , “Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark,” Advances in Neural Information Processing Systems , vol. 35, pp. 26 418–26 431, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu et al. , “Ego4d: Around the world in 3,000 hours of egocentric video,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 18 995–19 012
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
OpenAI, “Introducing chatgpt,” https://openai.com/chatgpt/ , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
OpenAI, “Gpt-4v(ision) system card,” https://cdn.openai.com/papers/GPTV_System_Card.pdf , 2023
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in International conference on machine learning . PMLR, 2023, pp. 19 730–19 742
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
H. Zhang, X. Li, and L. Bing, “Video-llama: An instruction-tuned audio-visual language model for video understanding,” EMNLP Demo Track , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
H. Liu and P. Abbeel, “Blockwise parallel transformer for large context models,” Advances in neural information processing systems , 2023
2023
Cited alongside, same era.
2024
Closest in time.
E. Song, W. Chai, G. Wang, Y. Zhang, H. Zhou, F. Wu, H. Chi, X. Guo, T. Ye, Y. Zhang et al. , “Moviechat: From dense token to sparse memory for long video understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 18 221–18 232
2024
Closest in time.
S. Ren, L. Yao, S. Li, X. Sun, and L. Hou, “Timechat: A time-sensitive multimodal large language model for long video understanding,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 14 313–14 323
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
A. Chen, Z. Wang, C. Dong, K. Tian, R. Zhao, X. Liang, Z. Kang, and X. Li, “Chinaopen: A dataset for open-world multimodal learning,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 6432–6440
2023
Cited alongside, same era.
G. Jocher, A. Chaurasia, and J. Qiu, “Ultralytics yolo,” https://github.com/ultralytics/ultralytics , 2023
2023
Cited alongside, same era.
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi, “Instructblip: Towards general-purpose vision-language models with instruction tuning,” arXiv preprint arkiv:2305.06500 , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Y. Fang, W. Wang, B. Xie, Q. Sun, L. Wu, X. Wang, T. Huang, X. Wang, and Y. Cao, “Eva: Exploring the limits of masked visual representation learning at scale,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 19 358–19 369
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
W. Wu, Y. Zhao, Z. Li, J. Li, H. Zhou, M. Z. Shou, and X. Bai, “A large cross-modal video retrieval dataset with reading comprehension,” Pattern Recognition , p. 110818, 2024
2024
Closest in time.
E. Cui, Y. He, Z. Ma, Z. Chen, H. Tian, W. Wang, K. Li, Y. Wang, W. Wang, X. Zhu, L. Lu, T. Lu, Y. Wang, L. Wang, Y. Qiao, and J. Dai, “Sharegpt-4o: Comprehensive multimodal annotations with gpt-4o,” https://sharegpt4o.github.io/ , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
B. Li, Y. Ge, Y. Ge, G. Wang, R. Wang, R. Zhang, and Y. Shan, “Seed-bench: Benchmarking multimodal large language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 13 299–13 308
2024
Closest in time.
K. Mangalam, R. Akshulakov, and J. Malik, “Egoschema: A diagnostic benchmark for very long-form video language understanding,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
J. Lin, H. Yin, W. Ping, P. Molchanov, M. Shoeybi, and S. Han, “Vila: On pre-training for visual language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 26 689–26 699
2024
Closest in time.
2024
Closest in time.
OpenAI, “Hello gpt-4o,” https://openai.com/index/hello-gpt-4o/ , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
X. Chu, J. Su, B. Zhang, and C. Shen, “Visionllama: A unified llama backbone for vision tasks,” in European Conference on Computer Vision , 2024
2024
Closest in time.
Anthropic, “Introducing the next generation of claude,” https://www.anthropic.com/news/claude-3-family , 2024
2024
Closest in time.