Fetching the paper…
Reading the bibliography…
Video Language Models (VLMs) are crucial for generalizing across diverse tasks and using language cues to enhance learning.
A. Vaswani, “Attention is all you need,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. Patrick, D. Campbell, Y. Asano, I. Misra, F. Metze, C. Feichtenhofer, A. Vedaldi, and J. F. Henriques, “Keeping your eye on the ball: Trajectory attention in video transformers,” Advances in neural information processing systems , vol. 34, pp. 12 493–12 506, 2021
2021
Earlier work this paper cites.
R. Girdhar and K. Grauman, “Anticipative video transformer,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 13 505–13 515
2021
Earlier work this paper cites.
M. Minderer, A. Gritsenko, A. Stone, M. Neumann, D. Weissenborn, A. Dosovitskiy, A. Mahendran, A. Arnab, M. Dehghani, Z. Shen et al. , “Simple open-vocabulary object detection,” in European Conference on Computer Vision . Springer, 2022, pp. 728–755
2022
Earlier work this paper cites.
H.-H. Wu, P. Seetharaman, K. Kumar, and J. P. Bello, “Wav2clip: Learning robust audio representations from clip,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2022, pp. 4563–4567
2022
Earlier work this paper cites.
Y. Jin, X. Wang, R. Yang, Y. Sun, W. Wang, H. Liao, and X. Xie, “Towards fine-grained reasoning for fake news detection,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 5, 2022, pp. 5746–5754
2022
Earlier work this paper cites.
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu et al. , “Ego4d: Around the world in 3,000 hours of egocentric video,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 18 995–19 012
2022
Earlier work this paper cites.
D. Damen, H. Doughty, G. M. Farinella, A. Furnari, E. Kazakos, J. Ma, D. Moltisanti, J. Munro, T. Perrett, W. Price et al. , “Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100,” International Journal of Computer Vision , pp. 1–23, 2022
2022
Earlier work this paper cites.
Y. Ma, G. Xu, X. Sun, M. Yan, J. Zhang, and R. Ji, “X-clip: End-to-end multi-grained contrastive learning for video-text retrieval,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 638–647
2022
Earlier work this paper cites.
R. Girdhar, M. Singh, N. Ravi, L. Van Der Maaten, A. Joulin, and I. Misra, “Omnivore: A single model for many visual modalities,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 16 102–16 112
2022
Cited alongside, same era.
2022
Cited alongside, same era.
C.-Y. Wu, Y. Li, K. Mangalam, H. Fan, B. Xiong, J. Malik, and C. Feichtenhofer, “Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 13 587–13 597
2022
Cited alongside, same era.
M. M. Islam and G. Bertasius, “Long movie clip classification with state-space video models,” in European Conference on Computer Vision . Springer, 2022, pp. 87–104
2022
Z. Zhong, D. Schneider, M. Voit, R. Stiefelhagen, and J. Beyerer, “Anticipative feature fusion transformer for multi-modal action anticipation,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2023, pp. 6068–6077
2023
Later among the works it cites.
Z. Dong, X. Liu, B. Chen, P. Polak, and P. Zhang, “Musechat: A conversational music recommendation system for videos,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 12 775–12 785
2024
Closest in time.
X. Liu, Z. Dong, and P. Zhang, “Tackling data bias in music-avqa: Crafting a balanced dataset for unbiased question-answering,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 4478–4487
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
E. Nguyen, K. Goel, A. Gu, G. Downs, P. Shah, T. Dao, S. Baccus, and C. Ré, “S4nd: Modeling images and videos as multidimensional signals with state spaces,” Advances in neural information processing systems , vol. 35, pp. 2846–2861, 2022
2022
Cited alongside, same era.
R. Herzig, E. Ben-Avraham, K. Mangalam, A. Bar, G. Chechik, A. Rohrbach, T. Darrell, and A. Globerson, “Object-region video transformers,” in Proceedings of the ieee/cvf conference on computer vision and pattern recognition , 2022, pp. 3148–3159
2022
Cited alongside, same era.
S. Pramanick, Y. Song, S. Nag, K. Q. Lin, H. Shah, M. Z. Shou, R. Chellappa, and P. Zhang, “Egovlpv2: Egocentric video-language pre-training with fusion in the backbone,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 5285–5297
2023
Cited alongside, same era.
Y. Zhao, I. Misra, P. Krähenbühl, and R. Girdhar, “Learning video representations from large language models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 6586–6597
2023
Cited alongside, same era.
Z. Tang, J. Cho, J. Lei, and M. Bansal, “Perceiver-vl: Efficient vision-and-language modeling with iterative latent attention,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2023, pp. 4410–4420
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
C. Ryali, Y.-T. Hu, D. Bolya, C. Wei, H. Fan, P.-Y. Huang, V. Aggarwal, A. Chowdhury, O. Poursaeed, J. Hoffman et al. , “Hiera: A hierarchical vision transformer without the bells-and-whistles,” in International Conference on Machine Learning . PMLR, 2023, pp. 29 441–29 454
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
OpenAI, “Gpt-4o,” August 2024, accessed: 2024-08-29. [Online]. Available: https://www.openai.com/gpt-4o
2024
Closest in time.