Fetching the paper…
Reading the bibliography…
We present DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL, through two key major upgrades.
Icfhr 2014 competition on handwritten digit string recognition in challenging datasets (hdsrc 2014)
M. Diem, S. Fiel, F. Kleber, R. Sablatnig, J. M. Saavedra, D. Contreras, J. M. Barrios, and L. S. Oliveira · 2014
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
S. Kazemzadeh, V. Ordonez, M. Matten, and T. Berg · 2014
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
B. A. Plummer, L. Wang, C. M. Cervantes, J. C. Caicedo, J. Hockenmaier, and S. Lazebnik · 2015
Earlier work this paper cites.
Large-scale classification of fine-art paintings: Learning the right metric on the right feature
B. Saleh and A. Elgammal · 2015
Earlier work this paper cites.
A diagram is worth a dozen images
A. Kembhavi, M. Salvato, E. Kolve, M. Seo, H. Hajishirzi, and A. Farhadi · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li · 2016
Earlier work this paper cites.
Modeling context in referring expressions
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg · 2016
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Wikihow: A large scale text summarization dataset
M. Koupaee and W. Y. Wang · 2018
Earlier work this paper cites.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
A. Amini, S. Gabriel, P. Lin, R. Koncel-Kedziorski, Y. Choi, and H. Hajishirzi · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Kvqa: Knowledge-aware visual question answering
S. Shah, A. Mishra, N. Yadati, and P. P. Talukdar · 2019
Earlier work this paper cites.
Objects365: A large-scale, high-quality dataset for object detection
S. Shao, Z. Li, T. Zhang, C. Peng, G. Yu, X. Zhang, J. Li, and J. Sun · 2019
Earlier work this paper cites.
Towards vqa models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
A. Kuznetsova, H. Rom, N. Alldrin, J. Uijlings, I. Krasin, J. Pont-Tuset, S. Kamali, S. Popov, M. Malloci, A. Kolesnikov, T. Duerig, and V. Ferrari · 2020
Earlier work this paper cites.
Image-based table recognition: Data, model, and evaluation
X. Zhong, E. ShafieiBavani, and A. Jimeno-Yepes · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning
P. Lu, L. Qiu, J. Chen, T. Xia, Y. Zhao, W. Zhang, Z. Yu, X. Liang, and S.-C. Zhu · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
M. Mathew, D. Karatzas, and C. Jawahar · 2021
Earlier work this paper cites.
Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning
K. Srinivasan, K. Raman, J. Chen, M. Bendersky, and M. Najork · 2021
Earlier work this paper cites.
Global table extractor (gte): A framework for joint table identification and cell structure recognition using visual context
X. Zheng, D. Burdick, L. Popa, P. Zhong, and N. X. R. Wang · 2021
Earlier work this paper cites.
Are deep neural networks smarter than second graders?
A. Cherian, K.-C. Peng, S. Lohit, K. Smith, and J. B. Tenenbaum · 2022
Earlier work this paper cites.
Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark
J. Gu, X. Meng, G. Lu, L. Hou, N. Minzhe, X. Liang, L. Yao, R. Huang, W. Zhang, X. Jiang, C. Xu, and H. Xu · 2022
Earlier work this paper cites.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque · 2022
Cited alongside, same era.
Infographicvqa
M. Mathew, V. Bagal, R. Tito, D. Karatzas, E. Valveny, and C. Jawahar · 2022
Cited alongside, same era.
Syntax-aware network for handwritten mathematical expression recognition
Y. Yuan, X. Liu, W. Dikubab, H. Liu, Z. Ji, Z. Wu, and X. Bai · 2022
Cited alongside, same era.
Latex-ocr — a tool to convert images of latex equations into latex code
L. Blecher · 2023
Cited alongside, same era.
Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023
M. Conover, M. Hayes, A. Mathur, J. Xie, J. Wan, S. Shah, A. Ghodsi, P. Wendell, M. Zaharia, and R. Xin · 2023
Cited alongside, same era.
Vip-llava: Making large multimodal models understand arbitrary visual prompts
M. Cai, H. Liu, S. K. Mustikovela, G. P. Meyer, Y. Chai, D. Park, and Y. J. Lee · 2024
Closest in time.
Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models
M. Deitke, C. Clark, S. Lee, R. Tripathi, Y. Yang, J. S. Park, M. Salehi, N. Muennighoff, K. Lo, L. Soldaini, et al · 2024
Closest in time.
Mind2web: Towards a generalist agent for the web
X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su · 2024
Closest in time.
Dalle3 1 Million+ High Quality Captions, May 2024
B. Egan, A. Redden, XWAVE, and SilentAntagonist · 2024
Closest in time.
Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024
C. Fu, P. Chen, Y. Shen, Y. Qin, M. Zhang, X. Lin, J. Yang, X. Zheng, K. Li, X. Sun, Y. Wu, and R. Ji · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. He, Z. Jin, C. Xu, J. Qiu, B. Wang, W. Li, H. Yan, J. Wang, and D. Lin · 2023
Cited alongside, same era.
HAI-LLM: Efficient and lightweight training tool for large models, 2023
High-flyer · 2023
Cited alongside, same era.
Gpt-4v(ision) system card
A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al · 2023
Cited alongside, same era.
Pick-a-pic: An open dataset of user preferences for text-to-image generation
Y. Kirstain, A. Polyak, U. Singer, S. Matiana, J. Penna, and O. Levy · 2023
Cited alongside, same era.
Laion-aesthetics, 2023
LAION · 2023
Cited alongside, same era.
OBELICS: an open web-scale filtered dataset of interleaved image-text documents
H. Laurençon, L. Saulnier, L. Tronchon, S. Bekman, A. Singh, A. Lozhkov, T. Wang, S. Karamcheti, A. M. Rush, D. Kiela, M. Cord, and V. Sanh · 2023
Cited alongside, same era.
Uninext: Exploring a unified architecture for vision recognition
F. Lin, J. Yuan, S. Wu, F. Wang, and Z. Wang · 2023
Cited alongside, same era.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024b
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.
Deepseek-vl: towards real-world vision-language understanding
H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang, et al · 2024
Closest in time.
Y. Ma, X. Liu, X. Chen, W. Liu, C. Wu, Z. Wu, Z. Pan, Z. Xie, H. Zhang, L. Zhao, et al · 2024
Closest in time.
Orca-math: Unlocking the potential of slms in grade school math, 2024
A. Mitra, H. Khanpour, C. Rosset, and A. Awadallah · 2024
Closest in time.
Math-llava: Bootstrapping mathematical reasoning for multimodal large language models
W. Shi, Z. Hu, Y. Bin, J. Liu, Y. Yang, S.-K. Ng, L. Bing, and R. K.-W. Lee · 2024
Closest in time.
From pixels to prose: A large dataset of dense image captions
V. Singla, K. Yue, S. Paul, R. Shirkavand, M. Jayawardhana, A. Ganjdanesh, H. Huang, A. Bhatele, G. Somepalli, and T. Goldstein · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
G. Team, P. Georgiev, V. I. Lei, R. Burnell, L. Bai, A. Gulati, G. Tanzer, D. Vincent, Z. Pan, S. Wang, et al · 2024
Closest in time.
Openmathinstruct-2: Accelerating ai for math with massive open-source instruction data
S. Toshniwal, W. Du, I. Moshkov, B. Kisacanin, A. Ayrapetyan, and I. Gitman · 2024
Closest in time.
Magicoder: Empowering code generation with OSS-instruct
Y. Wei, Z. Wang, J. Liu, Y. Ding, and L. Zhang · 2024
Closest in time.
Janus: Decoupling visual encoding for unified multimodal understanding and generation
C. Wu, X. Chen, Z. Wu, Y. Ma, X. Liu, Z. Pan, W. Liu, Z. Xie, X. Yu, C. Ruan, et al · 2024
Closest in time.
Grok-1.5 vision preview
xAI · 2024
Closest in time.
Florence-2: Advancing a unified representation for a variety of vision tasks
B. Xiao, H. Wu, W. Xu, X. Dai, H. Hu, Y. Lu, M. Zeng, C. Liu, and L. Yuan · 2024
Closest in time.
Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing
Z. Xu, F. Jiang, L. Niu, Y. Deng, R. Poovendran, Y. Choi, and B. Y. Lin · 2024
Closest in time.
Minicpm-v: A gpt-4v level mllm on your phone
Y. Yao, T. Yu, A. Zhang, C. Wang, J. Cui, H. Zhu, T. Cai, H. Li, W. Zhao, Z. He, et al · 2024
Closest in time.
Texthawk2: A large vision-language model excels in bilingual ocr and grounding with 16x fewer tokens
Y.-Q. Yu, M. Liao, J. Zhang, and J. Wu · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al · 2024
Closest in time.
Gpt-4v(ision) is a generalist web agent, if grounded
B. Zheng, B. Gou, J. Kil, H. Sun, and Y. Su · 2024
Closest in time.
Groma: Localized visual tokenization for grounding multimodal large language models
C. Ma, Y. Jiang, J. Wu, Z. Yuan, and X. Qi · 2025
Closest in time.
The all-seeing project v2: Towards general relation comprehension of the open world
W. Wang, Y. Ren, H. Luo, T. Li, C. Yan, Z. Chen, W. Wang, Q. Li, L. Lu, X. Zhu, et al · 2025
Closest in time.