Fetching the paper…
Reading the bibliography…
The integration of long-context capabilities with visual understanding unlocks unprecedented potential for Vision Language Models (VLMs).
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 1904
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations
Tillet, P., Kung, H.-T., and Cox, D · 2019
Earlier work this paper cites.
Activitynet-qa: A dataset for understanding complex web videos via question answering
Yu, Z., Xu, D., Yu, J., Yu, T., Zhao, Z., Zhuang, Y., and Tao, D · 2019
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B · 2021
Earlier work this paper cites.
Next-qa: Next phase of question-answering to explaining temporal actions
Xiao, J., Shang, X., Yao, A., and Chua, T.-S · 2021
Earlier work this paper cites.
Dynamic sparse attention for scalable transformer acceleration
Liu, L., Qu, Z., Chen, Z., Tu, F., Ding, Y., and Xie, Y · 2022
Earlier work this paper cites.
Token merging: Your vit but faster
Bolya, D., Fu, C.-Y., Dai, X., Zhang, P., Feichtenhofer, C., and Hoffman, J · 2023
Earlier work this paper cites.
Neighborhood attention transformer
Hassani, A., Walton, S., Li, J., Li, S., and Shi, H · 2023
Earlier work this paper cites.
Gaia-1: A generative world model for autonomous driving
Hu, A., Russell, L., Yeo, H., Murez, Z., Fedoseev, G., Kendall, A., Shotton, J., and Corrado, G · 2023
Earlier work this paper cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Earlier work this paper cites.
Egoschema: A diagnostic benchmark for very long-form video language understanding
Mangalam, K., Akshulakov, R., and Malik, J · 2023
Earlier work this paper cites.
Perception test: A diagnostic benchmark for multimodal video models
Patraucean, V., Smaira, L., Gupta, A., Recasens, A., Markeeva, L., Banarse, D., Koppula, S., Heyward, J., Malinowski, M., Yang, Y., Doersch, C., Matejovicova, T., Sulsky, Y., Miech, A., Fréchette, A., Klimczak, H., Koster, R., Zhang, J., Winkler, S., Aytar, Y., Osindero, S., Damen, D., Zisserman, A., and Carreira, J · 2023
Earlier work this paper cites.
Dao, tri and haziza, daniel and massa, francisco and sizov, grigory, 2023
Qwen, T · 2023
Earlier work this paper cites.
Pit: Optimization of dynamic sparse deep learning models via permutation invariant transformation
Zheng, N., Jiang, H., Zhang, Q., Han, Z., Ma, L., Yang, Y., Yang, F., Zhang, C., Qiu, L., Yang, M., et al · 2023
Earlier work this paper cites.
Star attention: Efficient llm inference over long sequences
Acharya, S., Jia, F., and Ginsburg, B · 2024
Earlier work this paper cites.
Zero-shot robotic manipulation with pre-trained image-editing diffusion models
Black, K., Nakamoto, M., Atreya, P., Walke, H. R., Finn, C., Kumar, A., and Levine, S · 2024
Earlier work this paper cites.
Gr-2: A generative video-language-action model with web-scale knowledge for robot manipulation
Cheang, C.-L., Chen, G., Jing, Y., Kong, T., Li, H., Li, Y., Liu, Y., Wu, H., Xu, J., Yang, Y., et al · 2024
Earlier work this paper cites.
An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Chen, L., Zhao, H., Liu, T., Bai, S., Lin, J., Zhou, C., and Chang, B · 2024
Cited alongside, same era.
Chu, Y., Xu, J., Yang, Q., Wei, H., Wei, X., Guo, Z., Leng, Y., Lv, Y., He, J., Lin, J., et al · 2024
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Dao, T · 2024
Cited alongside, same era.
Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Fu, C., Dai, Y., Luo, Y., Li, L., Ren, S., Zhang, R., Wang, Z., Zhou, C., Shen, Y., Zhang, M., et al · 2024
Cited alongside, same era.
Vista: A generalizable driving world model with high fidelity and versatile controllability
Gao, S., Yang, J., Chen, L., Chitta, K., Qiu, Y., Geiger, A., Zhang, J., and Li, H · 2024
Efficient vision-language models by summarizing visual tokens into compact registers
Wen, Y., Cao, Q., Fu, Q., Mehta, S., and Najibi, M · 2024
Later among the works it cites.
Longvlm: Efficient long video understanding via large language models
Weng, Y., Han, M., He, H., Chang, X., and Zhuang, B · 2024
Later among the works it cites.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2024
Later among the works it cites.
Visionzip: Longer is better but not necessary in vision language models
Yang, S., Chen, Y., Tian, Z., Wang, C., Li, J., Yu, B., and Jia, J · 2024
Later among the works it cites.
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., et al · 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
He, Y., Chen, F., Liu, J., Shao, W., Zhou, H., Zhang, K., and Zhuang, B · 2024
Cited alongside, same era.
Dialoggen: Multi-modal interactive dialogue system for multi-turn text-to-image generation
Huang, M., Long, Y., Deng, X., Chu, R., Xiong, J., Liang, X., Cheng, H., Lu, Q., and Liu, W · 2024
Cited alongside, same era.
MInference 1.0: Accelerating pre-filling for long-context LLMs via dynamic sparse attention
Jiang, H., Li, Y., Zhang, C., Wu, Q., Luo, X., Ahn, S., Han, Z., Abdi, A. H., Li, D., Lin, C.-Y., Yang, Y., and Qiu, L · 2024
Cited alongside, same era.
Video detail caption, 2024
Lab, L · 2024
Cited alongside, same era.
SnapKV: LLM knows what you are looking for before generation
Li, Y., Huang, Y., Yang, B., Venkitesh, B., Locatelli, A., Ye, H., Cai, T., Lewis, P., and Chen, D · 2024
Cited alongside, same era.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Maaz, M., Rasheed, H. A., Khan, S., and Khan, F · 2024
Cited alongside, same era.
Perception test: A diagnostic benchmark for multimodal video models
Patraucean, V., Smaira, L., Gupta, A., Recasens, A., Markeeva, L., Banarse, D., Koppula, S., Malinowski, M., Yang, Y., Doersch, C., et al · 2024
Cited alongside, same era.
Closest in time.
LongVILA: Scaling long-context visual language models for long videos
Chen, Y., Xue, F., Li, D., Hu, Q., Zhu, L., Li, X., Fang, Y., Tang, H., Yang, S., Liu, Z., He, Y., Yin, H., Molchanov, P., Kautz, J., Fan, L., Zhu, Y., Lu, Y., and Han, S · 2025
Closest in time.
Efficient-vdit: Efficient video diffusion transformers with attention tile
Ding, H., Li, D., Su, R., Zhang, P., Deng, Z., Stoica, I., and Zhang, H · 2025
Closest in time.
Flexprefill: A context-aware sparse attention mechanism for efficient long-sequence inference
Lai, X., Lu, J., Luo, Y., Ma, Y., and Zhou, X · 2025
Closest in time.
Videochat-flash: Hierarchical compression for long-context video modeling
Li, X., Wang, Y., Yu, J., Zeng, X., Zhu, Y., Huang, H., Gao, J., Li, K., He, Y., Wang, C., et al · 2025
Closest in time.
SCBench: A KV cache-centric analysis of long-context methods
LI, Y., Jiang, H., Wu, Q., Luo, X., Ahn, S., Zhang, C., Abdi, A. H., Li, D., Gao, J., Yang, Y., and Qiu, L · 2025
Closest in time.
Baichuan-omni-1.5 technical report
Li, Y., Liu, J., Zhang, T., Chen, S., Li, T., Li, Z., Liu, L., Ming, L., Dong, G., Pan, D., et al · 2025
Closest in time.
Moba: Mixture of block attention for long-context llms
Lu, E., Jiang, Z., Liu, J., Du, Y., Jiang, T., Hong, C., Liu, S., He, W., Yuan, E., Wang, Y., et al · 2025
Closest in time.
VL-cache: Sparsity and modality-aware KV cache compression for vision-language model inference acceleration
Tu, D., Vashchilenko, D., Lu, Y., and Xu, P · 2025
Closest in time.
Retrieval head mechanistically explains long-context factuality
Wu, W., Wang, Y., Xiao, G., Peng, H., and Fu, Y · 2025
Closest in time.
Sparse videogen: Accelerating video diffusion transformers with spatial-temporal sparsity
Xi, H., Yang, S., Zhao, Y., Xu, C., Li, M., Li, X., Lin, Y., Cai, H., Zhang, J., Li, D., et al · 2025
Closest in time.
Native sparse attention: Hardware-aligned and natively trainable sparse attention
Yuan, J., Gao, H., Dai, D., Luo, J., Zhao, L., Zhang, Z., Xie, Z., Wei, Y., Wang, L., Xiao, Z., et al · 2025
Closest in time.
Fast video generation with sliding tile attention
Zhang, P., Chen, Y., Su, R., Ding, H., Stoica, I., Liu, Z., and Zhang, H · 2025
Closest in time.