Fetching the paper…
Reading the bibliography…
We present DeepSeek-VL, an open-source Vision-Language (VL) Model designed for real-world vision and language understanding applications.
Generation and comprehension of unambiguous object descriptions
J. Mao, J. Huang, A. Toshev, O. Camburu, A. L. Yuille, and K. Murphy · 2016
Earlier work this paper cites.
Coco-text: Dataset and benchmark for text detection and recognition in natural images
A. Veit, T. Matera, L. Neumann, J. Matas, and S. Belongie · 2016
Earlier work this paper cites.
Icdar2017 robust reading challenge on multi-lingual scene text detection and script identification-rrc-mlt
N. Nayef, F. Yin, I. Bizid, H. Choi, Y. Feng, D. Karatzas, Z. Luo, U. Pal, C. Rigaud, J. Chazalon, et al · 2017
Earlier work this paper cites.
Icdar2017 competition on reading chinese text in the wild (rctw-17)
B. Shi, C. Yao, M. Liao, M. Yang, P. Xu, L. Cui, S. Belongie, S. Lu, and X. Bai · 2017
Earlier work this paper cites.
Uber-text: A large-scale dataset for optical character recognition from street-level imagery
Y. Zhang, L. Gueguen, I. Zharkov, P. Zhang, K. Seifert, and B. Kadlec · 2017
Earlier work this paper cites.
Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art
C. K. Chng, Y. Liu, Y. Sun, C. C. Ng, C. Luo, Z. Ni, C. Fang, S. Zhang, J. Han, E. Ding, et al · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Earlier work this paper cites.
Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt
Y. Sun, Z. Ni, C.-K. Chng, Y. Liu, C. Luo, C. C. Ng, J. Han, E. Ding, J. Liu, D. Karatzas, et al · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Earlier work this paper cites.
Root mean square layer normalization
B. Zhang and R. Sennrich · 2019
Earlier work this paper cites.
Icdar 2019 robust reading challenge on reading chinese text on signboard
R. Zhang, Y. Zhou, Q. Jiang, Q. Song, N. Li, K. Zhou, L. Wang, D. Wang, M. Liao, M. Yang, et al · 2019
Earlier work this paper cites.
The Pile: An 800GB dataset of diverse text for language modeling
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Widget captioning: Generating natural language description for mobile user interface elements
Y. Li, G. Li, L. He, J. Zheng, H. Li, and Z. Guan · 2020
Earlier work this paper cites.
Glu variants improve transformer
N. Shazeer · 2020
Earlier work this paper cites.
Program synthesis with large language models
J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Open images v5 text annotation and yet another mask text spotter
I. Krylov, S. Nosov, and V. Sovrasov · 2021
Earlier work this paper cites.
Iconqa: A new benchmark for abstract diagram understanding and visual language reasoning
P. Lu, L. Qiu, J. Chen, T. Xia, Y. Zhao, W. Zhang, Z. Yu, X. Liang, and S.-C. Zhu · 2021
Earlier work this paper cites.
Efficient large-scale language model training on gpu clusters using megatron-lm
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro, et al · 2021
Earlier work this paper cites.
Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text
A. Singh, G. Pang, M. Toh, J. Huang, W. Galuba, and T. Hassner · 2021
Earlier work this paper cites.
Screen2words: Automatic mobile ui summarization with multimodal learning
B. Wang, G. Li, X. Zhou, Z. Chen, T. Grossman, and Y. Li · 2021
Earlier work this paper cites.
Visual goal-step inference using wikihow
Y. Yang, A. Panagopoulou, Q. Lyu, L. Zhang, M. Yatskar, and C. Callison-Burch · 2021
Earlier work this paper cites.
Screenqa: Large-scale question-answer pairs over mobile app screenshots
Y.-C. Hsiao, F. Zubach, M. Wang, et al · 2022
Earlier work this paper cites.
Chart-to-text: A large-scale benchmark for chart summarization
S. Kantharaj, R. T. Leong, X. Lin, A. Masry, M. Thakkar, E. Hoque, and S. Joty · 2022
Cited alongside, same era.
Exploring plain vision transformer backbones for object detection
Y. Li, H. Mao, R. Girshick, and K. He · 2022
Cited alongside, same era.
Taisu: A 166m large-scale high-quality dataset for chinese vision-language pre-training
Y. Liu, G. Zhu, B. Zhu, Q. Song, G. Ge, H. Chen, G. Qiao, R. Peng, L. Wu, and J. Wang · 2022
Cited alongside, same era.
Towards end-to-end unified scene text detection and layout analysis
S. Long, S. Qin, D. Panteleev, A. Bissacco, Y. Fujii, and M. Raptis · 2022
Cited alongside, same era.
Chatgpt: Optimizing language models for dialogue
OpenAI · 2022
Cited alongside, same era.
Introducing Claude, 2023
Anthropic · 2023
Cited alongside, same era.
Generative pretraining in multimodality
Q. Sun, Q. Yu, Y. Cui, F. Zhang, X. Zhang, Y. Wang, H. Gao, J. Liu, T. Huang, and X. Wang · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
G. Team, R. Anil, S. Borgeaud, Y. Wu, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, et al · 2023
Later among the works it cites.
Vary: Scaling up the vision vocabulary for large vision-language models
H. Wei, L. Kong, J. Chen, L. Zhao, Z. Ge, J. Yang, J. Sun, C. Han, and X. Zhang · 2023
Later among the works it cites.
J. Ye, A. Hu, H. Xu, Q. Ye, M. Yan, G. Xu, C. Li, J. Tian, Q. Qian, J. Zhang, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Cited alongside, same era.
Introducing our multimodal models, 2023
R. Bavishi, E. Elsen, C. Hawthorne, M. Nye, A. Odena, A. Somani, and S. Taşırlar · 2023
Cited alongside, same era.
Nougat: Neural optical understanding for academic documents
L. Blecher, G. Cucurull, T. Scialom, and R. Stojnic · 2023
Cited alongside, same era.
A suite of generative tasks for multi-level multimodal webpage understanding
A. Burns, K. Srinivasan, J. Ainslie, G. Brown, B. A. Plummer, K. Saenko, J. Ni, and M. Guo · 2023
Cited alongside, same era.
Sharegpt4v: Improving large multi-modal models with better captions
L. Chen, J. Li, X. Dong, P. Zhang, C. He, J. Wang, F. Zhao, and D. Lin · 2023
Cited alongside, same era.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi · 2023
Cited alongside, same era.
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al · 2023
Later among the works it cites.
AGIEval: A human-centric benchmark for evaluating foundation models
W. Zhong, R. Cui, Y. Guo, Y. Liang, S. Lu, Y. Wang, A. Saied, W. Chen, and N. Duan · 2023
Later among the works it cites.
Yi-34B vision language model
01-ai · 2024
Closest in time.
Screenshot to code
Abi · 2024
Closest in time.
Anna’s archive
Anna’s Archive · 2024
Closest in time.
Latex-ocr
L. Blecher · 2024
Closest in time.
Textocr-gpt4v
J. Carter · 2024
Closest in time.
Deepseek llm: Scaling open-source language models with longtermism
DeepSeek-AI · 2024
Closest in time.
X. Dong, P. Zhang, Y. Zang, Y. Cao, B. Wang, L. Ouyang, X. Wei, S. Zhang, H. Duan, M. Cao, et al · 2024
Closest in time.
Detailed caption dataset
echo840 · 2024
Closest in time.
Websight dataset
HuggingFaceM4 · 2024
Closest in time.
wkhtmltopdf
A. Kulkarni and J. Truelsen · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a
H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee · 2024
Closest in time.
Are emergent abilities of large language models a mirage?
R. Schaeffer, B. Miranda, and S. Koyejo · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
S. Tong, Z. Liu, Y. Zhai, Y. Ma, Y. LeCun, and S. Xie · 2024
Closest in time.
Cmmmu: A chinese massive multi-discipline multimodal understanding benchmark
G. Zhang, X. Du, B. Chen, Y. Liang, T. Luo, T. Zheng, K. Zhu, Y. Cheng, C. Xu, S. Guo, et al · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al · 2024
Closest in time.
Multimodal c4: An open, billion-scale corpus of images interleaved with text
W. Zhu, J. Hessel, A. Awadalla, S. Y. Gadre, J. Dodge, A. Fang, Y. Yu, L. Schmidt, W. Y. Wang, and Y. Choi · 2024
Closest in time.