Fetching the paper…
Reading the bibliography…
We introduce Pixtral-12B, a 12--billion-parameter multimodal language model.
Glu variants improve transformer
Shazeer, N. (2020) · 2002
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A. (2020) · 2010
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K. (2016) · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. (2021) · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al. (2021) · 2021
Earlier work this paper cites.
Lepard: Learning partial point cloud matching in rigid and deformable scenes
Li, Y. and Harada, T. (2022) · 2022
Earlier work this paper cites.
Fuyu-8b: A multimodal architecture for ai agents
Bavishi, R., Elsen, E., Hawthorne, C., Nye, M., Odena, A., Somani, A., and Taşırlar, S. (2023) · 2023
Earlier work this paper cites.
An inverse scaling law for clip training
Li, X., Wang, Z., and Xie, C. (2023) · 2023
Cited alongside, same era.
Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J. (2023) · 2023
Cited alongside, same era.
Gpt-4 technical report
OpenAI, R. et al. (2023) · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023) · 2023
Cited alongside, same era.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. arxiv
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., et al. (2023) · 2023
Cited alongside, same era.
Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution
Dehghani, M., Mustafa, B., Djolonga, J., Heek, J., Minderer, M., Caron, M., Steiner, A., Puigcerver, J., Geirhos, R., Alabdulmohsin, I. M., et al. (2024) · 2024
Closest in time.
Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models
Deitke, M., Clark, C., Lee, S., Tripathi, R., Yang, Y., Park, J. S., Salehi, M., Muennighoff, N., Lo, K., Soldaini, L., et al. (2024) · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al. (2024) · 2024
Closest in time.
Llava-onevision: Easy visual task transfer
Li, B., Zhang, Y., Guo, D., Zhang, R., Li, F., Zhang, H., Zhang, K., Li, Y., Liu, Z., and Li, C. (2024) · 2024
Closest in time.
Mistral NeMo 12B
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. (2023) · 2023
Cited alongside, same era.
Agieval: A human-centric benchmark for evaluating foundation models
Zhong, W., Cui, R., Guo, Y., Liang, Y., Lu, S., Wang, Y., Saied, A., Chen, W., and Duan, N. (2023) · 2023
Cited alongside, same era.
The Claude 3 Model Family: Opus, Sonnet, Haiku
Anthropic (2024) · 2024
Cited alongside, same era.
Improved baselines with visual instruction tuning
Liu, H., Li, C., Li, Y., and Lee, Y. J. (2024a)
Cited in the paper.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J. (2024b)
Cited in the paper.
MistralAI (2024) · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reid, M., Savinov, N., Teplyashin, D., Lepikhin, D., Lillicrap, T., Alayrac, J.-b., Soricut, R., Lazaridou, A., Firat, O., Schrittwieser, J., et al. (2024) · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y. (2024) · 2024
Closest in time.
Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution
Wang, P., Bai, S., Tan, S., Wang, S., Fan, Z., Bai, J., Chen, K., Liu, X., Wang, J., Ge, W., Fan, Y., Dang, K., Du, M., Ren, X., Men, R., Liu, D., Zhou, C., Zhou, J., and Lin, J. (2024) · 2024
Closest in time.