Fetching the paper…
Reading the bibliography…
Video-Language Models (VLMs), powered by the advancements in Large Language Models (LLMs), are charting new frontiers in video understanding.
Hierarchical models of object recognition in cortex
Riesenhuber, M. and Poggio, T · 1999
Earlier work this paper cites.
Visual objects in context
Bar, M · 2004
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation
Chen, D. and Dolan, W. B · 2011
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding
Caba Heilbron, F., Escorcia, V., Ghanem, B., and Carlos Niebles, J · 2015
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Xu, J., Mei, T., Yao, T., and Rui, Y · 2016
Earlier work this paper cites.
Monet: Unsupervised scene decomposition and representation
Burgess, C. P., Matthey, L., Watters, N., Kabra, R., Higgins, I., Botvinick, M., and Lerchner, A · 2019
Earlier work this paper cites.
Slowfast networks for video recognition
Feichtenhofer, C., Fan, H., Malik, J., and He, K · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Object-centric learning with slot attention
Locatello, F., Weissenborn, D., Unterthiner, T., Mahendran, A., Heigold, G., Uszkoreit, J., Dosovitskiy, A., and Kipf, T · 2020
Earlier work this paper cites.
Emerging properties in self-supervised vision transformers
Caron, M., Touvron, H., Misra, I., Jégou, H., Mairal, J., Bojanowski, P., and Joulin, A · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Cited alongside, same era.
Bridging the gap to real-world object-centric learning
Seitzer, M., Horn, M., Zadaianchuk, A., Zietlow, D., Xiao, T., Simon-Gabriel, C.-J., He, T., Zhang, Z., Schölkopf, B., Brox, T., et al · 2022
Cited alongside, same era.
Jin, P., Takanobu, R., Zhang, C., Cao, X., and Yuan, L · 2023
Later among the works it cites.
Video-llava: Learning united visual representation by alignment before projection
Lin, B., Zhu, B., Ye, Y., Ning, M., Jin, P., and Yuan, L · 2023
Later among the works it cites.
Video-ChatGPT: Towards detailed video understanding via large vision and language models
Maaz, M., Rasheed, H., Khan, S., and Khan, F. S · 2023
Later among the works it cites.
Chatgpt: Large language model for human style conversation
OpenAI · 2023
Later among the works it cites.
Moviechat: From dense token to sparse memory for long video understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Illiterate DALL-E learns to compose
Singh, G., Deng, F., and Ahn, S · 2022
Cited alongside, same era.
Minigpt-v2: large language model as a unified interface for vision-language multi-task learning
Chen, J., Zhu, D., Shen, X., Li, X., Liu, Z., Zhang, P., Krishnamoorthi, R., Chandra, V., Xiong, Y., and Elhoseiny, M · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., Stoica, I., and Xing, E. P · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Dai, W., Li, J., Li, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Cited alongside, same era.
Li, J., Li, D., Savarese, S., and Hoi, S
Cited in the paper.
Videochat: Chat-centric video understanding
Li, K., He, Y., Wang, Y., Li, Y., Wang, W., Luo, P., Wang, Y., Wang, L., and Qiao, Y
Cited in the paper.
Mvbench: A comprehensive multi-modal video understanding benchmark
Li, K., Wang, Y., He, Y., Li, Y., Wang, Y., Liu, Y., Wang, Z., Xu, J., Chen, G., Luo, P., Wang, L., and Qiao, Y
Cited in the paper.
Song, E., Chai, W., Wang, G., Zhang, Y., Zhou, H., Wu, F., Guo, X., Ye, T., Lu, Y., Hwang, J.-N., et al · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Later among the works it cites.
Internvid: A large-scale video-text dataset for multimodal understanding and generation
Wang, Y., He, Y., Li, Y., Li, K., Yu, J., Ma, X., Li, X., Chen, G., Chen, X., Wang, Y., et al · 2023
Later among the works it cites.
Chatbridge: Bridging modalities with large language model as a language catalyst
Zhao, Z., Guo, L., Yue, T., Chen, S., Shao, S., Zhu, X., Yuan, Z., and Liu, J · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.