Fetching the paper…
Reading the bibliography…
Recent advances demonstrate that scaling Large Vision-Language Models (LVLMs) effectively improves downstream task performances.
Liii. on lines and planes of closest fit to systems of points in space
Pearson, K · 1901
Earlier work this paper cites.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
Eigen, D., Ranzato, M., and Sutskever, I · 2013
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Adaptive input representations for neural language modeling
Baevski, A. and Auli, M · 2018
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P · 2018
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
Learning deep transformer models for machine translation
Wang, Q., Li, B., Xiao, T., Zhu, J., Li, C., Wong, D. F., and Chao, L. S · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
Earlier work this paper cites.
Glm: General language model pretraining with autoregressive blank infilling
Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J · 2021
Earlier work this paper cites.
Beyond distillation: Task-level mixture-of-experts for efficient inference
Kudugunta, S., Huang, Y., Bapna, A., Krikun, M., Lepikhin, D., Luong, M.-T., and Firat, O · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Scaling vision with sparse mixture of experts
Riquelme, C., Puigcerver, J., Mustafa, B., Neumann, M., Jenatton, R., Susano Pinto, A., Keysers, D., and Houlsby, N · 2021
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al · 2022
Earlier work this paper cites.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Bao, H., Wang, W., Dong, L., Liu, Q., Mohammed, O. K., Aggarwal, K., Som, S., Piao, S., and Wei, F · 2022
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Earlier work this paper cites.
Sparse upcycling: Training mixture-of-experts from dense checkpoints
Komatsuzaki, A., Puigcerver, J., Lee-Thorp, J., Ruiz, C. R., Mustafa, B., Ainslie, J., Tay, Y., Dehghani, M., and Houlsby, N · 2022
Earlier work this paper cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J., Li, D., Xiong, C., and Hoi, S · 2022
Earlier work this paper cites.
Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning
Liang, V. W., Zhang, Y., Kwon, Y., Yeung, S., and Zou, J. Y · 2022
Earlier work this paper cites.
Swin transformer v2: Scaling up capacity and resolution
Liu, Z., Hu, H., Lin, Y., Yao, Z., Xie, Z., Wei, Y., Ning, J., Cao, Y., Zhang, Z., Dong, L., et al · 2022
Earlier work this paper cites.
Learn to explain: Multimodal reasoning via thought chains for science question answering
Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.-W., Zhu, S.-C., Tafjord, O., Clark, P., and Kalyan, A · 2022
Earlier work this paper cites.
Multimodal contrastive learning with limoe: the language-image mixture of experts
Mustafa, B., Riquelme, C., Puigcerver, J., Jenatton, R., and Houlsby, N · 2022
Cited alongside, same era.
Rome: Role-aware mixture-of-expert transformer for text-to-video retrieval
Satar, B., Zhu, H., Zhang, H., and Lim, J. H · 2022
Cited alongside, same era.
Bloom: A 176b-parameter open-access multilingual language model
Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilić, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., et al · 2022
Cited alongside, same era.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks
Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., et al · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Multiway-adapater: Adapting large-scale multi-modal models for scalable image-text retrieval
Long, Z., Killick, G., McCreadie, R., and Camarasa, G. A · 2023
Later among the works it cites.
Ma, G., Wu, X., Wang, P., and Hu, S · 2023
Later among the works it cites.
Phi-2: The surprising power of small language models
Microsoft · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E., and Launay, J · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Glm-130b: An open bilingual pre-trained model
Zeng, A., Liu, X., Du, Z., Wang, Z., Lai, H., Ding, M., Yang, Z., Xu, Y., Zheng, W., Xia, X., et al · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2022
Cited alongside, same era.
Uni-perceiver-moe: Learning sparse generalist models with conditional moes
Zhu, J., Zhu, X., Wang, W., Wang, X., Li, H., Wang, X., and Dai, J · 2022
Cited alongside, same era.
St-moe: Designing stable and transferable sparse expert models
Zoph, B., Bello, I., Kumar, S., Du, N., Huang, Y., Dean, J., Shazeer, N., and Fedus, W · 2022
Cited alongside, same era.
Building the next generation of open-source and bilingual llms
01-ai · 2023
Cited alongside, same era.
Honeybee: Locality-enhanced projector for multimodal llm
Cha, J., Kang, W., Mun, J., and Roh, B · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Cited alongside, same era.
Kosmos-2: Grounding multimodal large language models to the world
Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., and Wei, F · 2023
Later among the works it cites.
Detgpt: Detect what you need via reasoning
Pi, R., Gao, J., Diao, S., Pan, R., Dong, H., Zhang, J., Yao, L., Han, J., Xu, H., and Zhang, L. K. T · 2023
Later among the works it cites.
Glamm: Pixel grounding large multimodal model
Rasheed, H., Maaz, M., Shaji, S., Shaker, A., Khan, S., Cholakkal, H., Anwer, R. M., Xing, E., Yang, M.-H., and Khan, F. S · 2023
Later among the works it cites.
Scaling vision-language models with sparse mixture of experts
Shen, S., Yao, Z., Li, C., Darrell, T., Keutzer, K., and He, Y · 2023
Later among the works it cites.
Moss: Training conversational language models from synthetic data
Sun, T., Zhang, X., He, Z., Li, P., Cheng, Q., Yan, H., Liu, X., Shao, Y., Tang, Q., Zhao, X., et al · 2023
Later among the works it cites.
Sus-chat: Instruction tuning done right
SUSTech-IDEA · 2023
Later among the works it cites.
Alpaca: A strong, replicable instruction-following model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Internlm: A multilingual language model with progressively enhanced capabilities, 2023
Team, I · 2023
Later among the works it cites.
Baichuan 2: Open large-scale language models
Yang, A., Xiao, B., Wang, B., Zhang, B., Bian, C., Yin, C., Lv, C., Pan, D., Wang, D., Yan, D., et al · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Ye, Q., Xu, H., Xu, G., Ye, J., Yan, M., Zhou, Y., Wang, J., Hu, A., Shi, P., Shi, Y., et al · 2023
Later among the works it cites.
A survey on multimodal large language models
Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., and Chen, E · 2023
Later among the works it cites.
Mm-vet: Evaluating large multimodal models for integrated capabilities
Yu, W., Yang, Z., Li, L., Wang, J., Lin, K., Liu, Z., Wang, X., and Wang, L · 2023
Later among the works it cites.
Tinygpt-v: Efficient multimodal large language model via small backbones
Yuan, Z., Li, Z., and Sun, L · 2023
Later among the works it cites.
Xuanyuan 2.0: A large chinese financial chat model with hundreds of billions parameters
Zhang, X. and Yang, Q · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D., Chen, J., Shen, X., Li, X., and Elhoseiny, M · 2023
Later among the works it cites.
Deepseek llm: Scaling open-source language models with longtermism
DeepSeek-AI · 2024
Closest in time.
Mixtral of experts, 2024
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., de las Casas, D., Hanna, E. B., Bressand, F., Lengyel, G., Bour, G., Lample, G., Lavaud, L. R., Saulnier, L., Lachaux, M.-A., Stock, P., Subramanian, S., Yang, S., Antoniak, S., Scao, T. L., Gervet, T., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2024
Closest in time.
Llava-phi: Efficient multi-modal assistant with small language model, 2024
Zhu, Y., Zhu, M., Liu, N., Ou, Z., Mou, X., and Tang, J · 2024
Closest in time.