Fetching the paper…
Reading the bibliography…
Information comes in diverse modalities.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Earlier work this paper cites.
LAION-400M: open dataset of clip-filtered 400 million image-text pairs
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki · 2021
Earlier work this paper cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2022
Earlier work this paper cites.
BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation
J. Li, D. Li, C. Xiong, and S. C. H. Hoi · 2022
Earlier work this paper cites.
St-moe: Designing stable and transferable sparse expert models
B. Zoph, I. Bello, S. Kumar, N. Du, Y. Huang, J. Dean, N. Shazeer, and W. Fedus · 2022
Earlier work this paper cites.
Qwen-VL: A versatile vision-language model for understanding, localization, text reading, and beyond
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Earlier work this paper cites.
Patch n’ pack: Navit, a vision transformer for any aspect ratio and resolution
M. Dehghani, B. Mustafa, J. Djolonga, J. Heek, M. Minderer, M. Caron, A. Steiner, J. Puigcerver, R. Geirhos, I. M. Alabdulmohsin, A. Oliver, P. Padlewski, A. A. Gritsenko, M. Lucic, and N. Houlsby · 2023
Cited alongside, same era.
Llmtest_needleinahaystack, 2023
G. Kamradt · 2023
Cited alongside, same era.
Pix2struct: Screenshot parsing as pretraining for visual language understanding
K. Lee, M. Joshi, I. R. Turc, H. Hu, F. Liu, J. M. Eisenschlos, U. Khandelwal, P. Shaw, M. Chang, and K. Toutanova · 2023
Cited alongside, same era.
BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models
J. Li, D. Li, S. Savarese, and S. C. H. Hoi · 2023
Cited alongside, same era.
Scaling vision-language models with sparse mixture of experts
S. Shen, Z. Yao, C. Li, T. Darrell, K. Keutzer, and Y. He · 2023
Cited alongside, same era.
Mixtral of Experts: A sparse mixture of experts language model
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de Las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed · 2024
Closest in time.
What matters when building vision-language models?
H. Laurençon, L. Tronchon, M. Cord, and V. Sanh · 2024
Closest in time.
Llava-onevision: Easy visual task transfer
B. Li, Y. Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, Y. Li, Z. Liu, and C. Li · 2024
Closest in time.
Improved baselines with visual instruction tuning
H. Liu, C. Li, Y. Li, and Y. J. Lee · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
W. Shi, S. Min, M. Lomeli, C. Zhou, M. Li, G. Szilvasy, R. James, X. V. Lin, N. A. Smith, L. Zettlemoyer, S. Yih, and M. Lewis · 2023
Cited alongside, same era.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Cited alongside, same era.
DeepSeekMoE: Towards ultimate expert specialization in mixture-of-experts language models
D. Dai, C. Deng, C. Zhao, R. Xu, H. Gao, D. Chen, J. Li, W. Zeng, X. Yu, Y. Wu, et al · 2024
Cited alongside, same era.
The llama 3 herd of models
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Rozière, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. M. Kloumann, I. Misra, I. Evtimov, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, and et al · 2024
Cited alongside, same era.
MoE-LLaVA: Mixture of experts for large vision-language models
B. Lin, Z. Tang, Y. Ye, J. Cui, B. Zhu, P. Jin, J. Zhang, M. Ning, and L. Yuan
Cited in the paper.
MoMa: Efficient early-fusion pre-training with mixture of modality-aware experts
X. V. Lin, A. Shrivastava, L. Luo, S. Iyer, M. Lewis, G. Ghosh, L. Zettlemoyer, and A. Aghajanyan
Cited in the paper.
J. Ludziejewski, J. Krajewski, K. Adamczewski, M. Pióro, M. Krutul, S. Antoniak, K. Ciebiera, K. Król, T. Odrzygózdz, P. Sankowski, M. Cygan, and S. Jaszczur · 2024
Closest in time.
Pixtral 12b - the first-ever multimodal mistral model, 2024
Mixtral · 2024
Closest in time.
Mia-bench: Towards better instruction following evaluation of multimodal llms
Y. Qian, H. Ye, J. Fauconnier, P. Grasch, Y. Yang, and Z. Gan · 2024
Closest in time.
Longvideobench: A benchmark for long-context interleaved video-language understanding
H. Wu, D. Li, B. Chen, and J. Li · 2024
Closest in time.