Fetching the paper…
Reading the bibliography…
We propose SPHINX-X, an extensive Multimodality Large Language Model (MLLM) series developed upon SPHINX.
Im2text: Describing images using 1 million captioned photographs
Ordonez, V., Kulkarni, G., and Berg, T · 2011
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes
Kazemzadeh, S., Ordonez, V., andre Matten, M., and Berg, T. L · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S. J., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M. S., Berg, A. C., and Fei-Fei, L · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Agrawal, A., Lu, J., Antol, S., Mitchell, M., Zitnick, C. L., Parikh, D., and Batra, D · 2015
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions
Mao, J., Huang, J., Toshev, A., Camburu, O.-M., Yuille, A. L., and Murphy, K. P · 2015
Earlier work this paper cites.
Compositional semantic parsing on semi-structured tables
Pasupat, P. and Liang, P · 2015
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C. L., and Girshick, R. B · 2016
Earlier work this paper cites.
A diagram is worth a dozen images
Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., and Farhadi, A · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalantidis, Y., Li, L.-J., Shamma, D. A., et al · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P · 2018
Earlier work this paper cites.
Dvqa: Understanding data visualizations via question answering
Kafle, K., Cohen, S. D., Price, B. L., and Kanan, C · 2018
Earlier work this paper cites.
The open images dataset v4
Kuznetsova, A., Rom, H., Alldrin, N. G., Uijlings, J. R. R., Krasin, I., Pont-Tuset, J., Kamali, S., Popov, S., Malloci, M., Kolesnikov, A., Duerig, T., and Ferrari, V · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Sharma, P., Ding, N., Goodman, S., and Soricut, R · 2018
Earlier work this paper cites.
Tabfact: A large-scale dataset for table-based fact verification
Chen, W., Wang, H., Chen, J., Zhang, Y., Wang, H., LI, S., Zhou, X., and Wang, W. Y · 2019
Earlier work this paper cites.
Lvis: A dataset for large vocabulary instance segmentation
Gupta, A., Dollár, P., and Girshick, R. B · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino, K., Rastegari, M., Farhadi, A., and Mottaghi, R · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
Mishra, A., Shekhar, S., Singh, A. K., and Chakraborty, A · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Objects365: A large-scale, high-quality dataset for object detection
Shao, S., Li, Z., Zhang, T., Peng, C., Yu, G., Zhang, X., Li, J., and Sun, J · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2020
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Deepform, 2020
Svetlichnaya, S · 2020
Earlier work this paper cites.
Donut: Document understanding transformer without ocr
Kim, G., Hong, T., Yim, M., Park, J., Yim, J., Hwang, W., Yun, S., Han, D., and Park, S · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A · 2021
Cited alongside, same era.
Kleister: Key information extraction datasets involving long documents with complex layouts
Stanislawek, T., Grali’nski, F., Wr’oblewska, A., Lipi’nski, D., Kaliska, A., Rosalska, P., Topolski, B., and Biecek, P · 2021
Cited alongside, same era.
Guo, Z., Zhang, R., Zhu, X., Tang, Y., Ma, X., Han, J., Chen, K., Gao, P., Li, X., Li, H., et al · 2023
Later among the works it cites.
mplug-paperowl: Scientific diagram analysis with the multimodal large language model
Hu, A., Shi, Y., Xu, H., Ye, J., Ye, Q., Yan, M., Li, C., Qian, Q., Zhang, J., and Huang, F · 2023
Later among the works it cites.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Later among the works it cites.
Jin, P., Takanobu, R., Zhang, C., Cao, X., and Yuan, L · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tanaka, R., Nishida, K., and Yoshida, S · 2021
Cited alongside, same era.
Synthtiger: Synthetic text image generator towards better text recognition models
Yim, M., Kim, Y., Cho, H.-C., and Park, S · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Cited alongside, same era.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Bao, H., Wang, W., Dong, L., Liu, Q., Mohammed, O. K., Aggarwal, K., Som, S., Piao, S., and Wei, F · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Cited alongside, same era.
A convnet for the 2020s
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S · 2022
Cited alongside, same era.
Openorca: An open dataset of gpt augmented flan reasoning traces
Lian, W., Goodson, B., Pentland, E., Cook, A., Vong, C., and ”Teknium” · 2023
Later among the works it cites.
Lin, Z., Liu, C., Zhang, R., Gao, P., Qiu, L., Xiao, H., Qiu, H., Lin, C., Shao, W., Chen, K., et al · 2023
Later among the works it cites.
Lu, P., Bansal, H., Xia, T., Liu, J., yue Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J · 2023
Later among the works it cites.
Wizardcoder: Empowering code large language models with evol-instruct
Luo, Z., Xu, C., Zhao, P., Sun, Q., Geng, X., Hu, W., Tao, C., Ma, J., Lin, Q., and Jiang, D · 2023
Later among the works it cites.
Kosmos-2.5: A multimodal literate model
Lv, T., Huang, Y., Chen, J., Cui, L., Ma, S., Chang, Y., Huang, S., Wang, W., Dong, L., Luo, W., et al · 2023
Later among the works it cites.
Video-chatgpt: Towards detailed video understanding via large vision and language models
Maaz, M., Rasheed, H., Khan, S., and Khan, F. S · 2023
Later among the works it cites.
Phi-2, 2023
Microsoft · 2023
Later among the works it cites.
Video-bench: A comprehensive benchmark and toolkit for evaluating video-based large language models
Ning, M., Zhu, B., Xie, Y., Lin, B., Cui, J., Yuan, L., Chen, D., and Yuan, L · 2023
Later among the works it cites.
GPT-4V(ision) system card, 2023
OpenAI · 2023
Later among the works it cites.
Dinov2: Learning robust visual features without supervision
Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al · 2023
Later among the works it cites.
Tiny lvlm-ehub: Early multimodal experiments with bard
Shao, W., Hu, Y., Gao, P., Lei, M., Zhang, K., Meng, F., Xu, P., Huang, S., Li, H., Qiao, Y., et al · 2023
Later among the works it cites.
Pandagpt: One model to instruction-follow them all
Su, Y., Lan, T., Li, H., Xu, J., Wang, Y., and Cai, D · 2023
Later among the works it cites.
Internlm: A multilingual language model with progressively enhanced capabilities, 2023
Team, I · 2023
Later among the works it cites.
Dilu: A knowledge-driven approach to autonomous driving with large language models
Wen, L., Fu, D., Li, X., Cai, X., Ma, T., Cai, P., Dou, M., Shi, B., He, L., and Qiao, Y · 2023
Later among the works it cites.
Prompt, generate, then cache: Cascade of foundation models makes strong few-shot learners
Zhang, R., Hu, X., Li, B., Huang, S., Deng, H., Li, H., Qiao, Y., and Gao, P · 2023
Later among the works it cites.
Pointclip v2: Prompting clip and gpt for powerful 3d open-world learning
Zhu, X., Zhang, R., He, B., Guo, Z., Zeng, Z., Qin, Z., Zhang, S., and Gao, P · 2023
Later among the works it cites.
Seeclick: Harnessing gui grounding for advanced visual gui agents
Cheng, K., Sun, Q., Chu, Y., Xu, F., Li, Y., Zhang, J., and Wu, Z · 2024
Closest in time.
Cmmmu: A chinese massive multi-discipline multimodal understanding benchmark
Ge, Z., Xinrun, D., Bei, C., Yiming, L., Tongxu, L., Tianyu, Z., Kang, Z., Yuyang, C., Chunpu, X., Shuyue, G., Haoran, Z., Xingwei, Q., Junjie, W., Ruibin, Y., Yizhi, L., Zekun, W., Yudong, L., Yu-Hsuan, T., Fengji, Z., Chenghua, L., Wenhao, H., Wenhu, C., and Jie, F · 2024
Closest in time.
Sciverse
Guo, Z., Zhang, R., Chen, H., Gao, J., Gao, P., Li, H., and Heng, P.-A · 2024
Closest in time.
Webvoyager: Building an end-to-end web agent with large multimodal models
He, H., Yao, W., Ma, K., Yu, W., Dai, Y., Zhang, H., Lan, Z., and Yu, D · 2024
Closest in time.
Aesbench: An expert benchmark for multimodal large language models on image aesthetics perception
Huang, Y., Yuan, Q., Sheng, X., Yang, Z., Wu, H., Chen, P., Yang, Y., Li, L., and Lin, W · 2024
Closest in time.
Llava-next: Stronger llms supercharge multimodal capabilities in the wild, May 2024a
Li, B., Zhang, K., Zhang, H., Guo, D., Zhang, R., Li, F., Zhang, Y., Liu, Z., and Li, C · 2024
Closest in time.
Llava-next: Tackling multi-image, video, and 3d in large multimodal models, June 2024b
Li, F., Zhang, R., Zhang, H., Zhang, Y., Li, B., Li, W., Ma, Z., and Li, C · 2024
Closest in time.
Meng, F., Shao, W., Lu, Q., Gao, P., Zhang, K., Qiao, Y., and Luo, P · 2024
Closest in time.
Eyes wide shut? exploring the visual shortcomings of multimodal llms
Tong, S., Liu, Z., Zhai, Y., Ma, Y., LeCun, Y., and Xie, S · 2024
Closest in time.