Fetching the paper…
Reading the bibliography…
We introduce Long-VITA, a simple yet effective large multi-modal model for long-context visual-language understanding tasks.
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi, “a Diagram Is Worth a Dozen Images,”
2016
Earlier work this paper cites.
Q. Huang, Y. Xiong, A. Rao, J. Wang, and D. Lin, “MovieNet: A Holistic Dataset for Movie Understanding,”
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan, “Flamingo: A Visual Language Model for Few-Shot Learning,” 2022
2022
Earlier work this paper cites.
OpenAI, “GPT-4,” 2023. [Online]. Available:
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual Instruction Tuning,” 2023. [Online]. Available:
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
Teknium, “OpenHermes 2.5: An Open Dataset of Synthetic Data for Generalist LLM Assistants,” 2023. [Online]. Available:
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
M. Conover, M. Hayes, A. Mathur, J. Xie, J. Wan, S. Shah, A. Ghodsi, P. Wendell, M. Zaharia, and R. Xin, “Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM,” 2023. [Online]. Available:
2023
Earlier work this paper cites.
A. Unified, “Atlas math sets,” 2023. [Online]. Available:
2023
Earlier work this paper cites.
L. Tiedong, “Goat,” 2023. [Online]. Available:
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
T. Computer, “Long Data Collections,” 2023. [Online]. Available:
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Later among the works it cites.
“ShareGemini: Scaling up Video Caption Data for Multimodal Large Language Models,” 2024. [Online]. Available:
2024
Later among the works it cites.
L. Yu, W. Jiang, H. Shi, J. Yu, Z. Liu, Y. Zhang, J. T. Kwok, Z. Li, A. Weller, and W. Liu, “Metamath: Bootstrap Your Own Mathematical Questions for Large Language Models,”
2024
Later among the works it cites.
X. Yue, X. Qu, G. Zhang, Y. Fu, W. Huang, H. Sun, Y. Su, and W. Chen, “Mammoth: Building Math Generalist Models Through Hybrid Instruction Tuning,”
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
Google, “Introducing Gemini 2.0: Our New AI Model for the Agentic Era,” 2024. [Online]. Available:
2024
Cited alongside, same era.
Anthropic, “Claude 3.5 Sonnet,” 2024. [Online]. Available:
2024
Cited alongside, same era.
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Google, “Gemini 1.5.” [Online]. Available:
2024
Cited alongside, same era.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Qwen Team, “Qwen2.5: A Party of Foundation Models,” 2024. [Online]. Available:
2024
Later among the works it cites.
2024
Later among the works it cites.
P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K. W. Chang, M. Galley, and J. Gao, “Mathvista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts,”
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Meta, “Llama 3.2: Revolutionizing Edge AI and Vision with Open, Customizable Models,” 2024. [Online]. Available:
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Alipay, “bailingMM-Mini,” 2024. [Online]. Available:
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
H. Liu, M. Zaharia, and P. Abbeel, “Ring Attention with Blockwise Transformers for Near-Infinite Context,”
2024
Later among the works it cites.
H. W. Chen and Zhenzhong, “Visual Context Window Extension: A New Perspective for Long Video Understanding,” 2024
2024
Later among the works it cites.