Fetching the paper…
Reading the bibliography…
Vision Language Models (VLMs) are impressive at visual question answering and image captioning.
Ocrbench: on the hidden mystery of ocr in large multimodal models
Liu, Y., Li, Z., Huang, M., Yang, B., Yu, W., Li, C., Yin, X.-C., Liu, C.-L., Jin, L., and Bai, X · 1919
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Zero-shot learning through cross-modal transfer
Socher, R., Ganjoo, M., Manning, C. D., and Ng, A · 2013
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Chen, X., Fang, H., Lin, T.-Y., Vedantam, R., Gupta, S., Dollár, P., and Zitnick, C. L · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
Agrawal, A., Batra, D., and Parikh, D · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Earlier work this paper cites.
Measuring abstract reasoning in neural networks
Barrett, D., Hill, F., Santoro, A., Morcos, A., and Lillicrap, T · 2018
Earlier work this paper cites.
Lectures on convex optimization
Nesterov, Y · 2018
Earlier work this paper cites.
Learning to make analogies by contrasting abstract relational structure
Hill, F., Santoro, A., Barrett, D., Morcos, A., and Lillicrap, T · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
On the ability and limitations of transformers to recognize formal languages
Bhattamishra, S., Ahuja, K., and Goyal, N · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Vokenization: Improving language understanding with contextualized, visual-grounded supervision
Tan, H. and Bansal, M · 2020
Earlier work this paper cites.
What makes training multi-modal classification networks hard?
Wang, W., Tran, D., and Feiszli, M · 2020
Earlier work this paper cites.
Cross-modal generalization: Learning in low resource modalities via meta-alignment
Liang, P. P., Wu, P., Ziyin, L., Morency, L.-P., and Salakhutdinov, R · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Earlier work this paper cites.
Thinking like transformers
Weiss, G., Goldberg, Y., and Yahav, E · 2021
Earlier work this paper cites.
Exploring length generalization in large language models
Anil, C., Wu, Y., Andreassen, A., Lewkowycz, A., Misra, V., Ramasesh, V., Slone, A., Gur-Ari, G., Dyer, E., and Neyshabur, B · 2022
Earlier work this paper cites.
Modality competition: What makes joint training of multi-modal network fail in deep learning? (Provably)
Huang, Y., Lin, J., Zhou, C., Yang, H., and Huang, L · 2022
Earlier work this paper cites.
A convnet for the 2020s
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S · 2022
Earlier work this paper cites.
Prioritized training on points that are learnable, worth learning, and not yet learnt
Mindermann, S., Brauner, J. M., Razzak, M. T., Sharma, M., Kirsch, A., Xu, W., Höltgen, B., Gomez, A. N., Morisot, A., Farquhar, S., et al · 2022
Earlier work this paper cites.
Balanced multimodal learning via on-the-fly gradient modulation
Peng, X., Wei, Y., Deng, A., Wang, D., and Hu, D · 2022
Earlier work this paper cites.
Winoground: Probing vision and language models for visio-linguistic compositionality
Thrush, T., Jiang, R., Bartolo, M., Singh, A., Williams, A., Kiela, D., and Ross, C · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al · 2023
Cited alongside, same era.
Are deep neural networks smarter than second graders?
Cherian, A., Peng, K.-C., Lohit, S., Smith, K. A., and Tenenbaum, J. B · 2023
Cited alongside, same era.
Pmr: Prototypical modal rebalance for multimodal learning
Fan, Y., Xu, W., Wang, H., Wang, J., and Guo, S · 2023
Cited alongside, same era.
Mme: A comprehensive evaluation benchmark for multimodal large language models
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Qiu, Z., Lin, W., Yang, J., Zheng, X., Li, K., Sun, X., and Ji, R · 2023
Cited alongside, same era.
T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation
Huang, K., Sun, K., Xie, E., Li, Z., and Liu, X · 2023
How to train long-context language models (effectively)
Gao, T., Wettig, A., Yen, H., and Chen, D · 2024
Later among the works it cites.
Ghosal, D., Han, V. T. Y., Ken, C. Y., and Poria, S · 2024
Later among the works it cites.
Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality
Hsieh, C.-Y., Zhang, J., Ma, Z., Kembhavi, A., and Krishna, R · 2024
Later among the works it cites.
The impact of positional encoding on length generalization in transformers
Kazemnejad, A., Padhi, I., Natesan Ramamurthy, K., Das, P., and Reddy, S · 2024
Later among the works it cites.
II-MMR: Identifying and improving multi-modal multi-hop reasoning in visual question answering
Kil, J., Tavazoee, F., Kang, D., and Kim, J.-K · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Length generalization in arithmetic transformers
Jelassi, S., d’Ascoli, S., Domingo-Enrich, C., Wu, Y., Li, Y., and Charton, F · 2023
Cited alongside, same era.
Representations and computations in transformers that support generalization on structured tasks
Li, Y. and McClelland, J · 2023
Cited alongside, same era.
Evaluating object hallucination in large vision-language models
Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, X. Z., and Wen, J.-R · 2023
Cited alongside, same era.
Improved baselines with visual instruction tuning
Liu, H., Li, C., Li, Y., and Lee, Y. J · 2023
Cited alongside, same era.
MetaVL: Transferring in-context learning ability from language models to vision-language models
Monajatipoor, M., Li, L. H., Rouhsedaghat, M., Yang, L., and Chang, K.-W · 2023
Cited alongside, same era.
Trak: Attributing model behavior at scale
Park, S. M., Georgiev, K., Ilyas, A., Leclerc, G., and Madry, A · 2023
Cited alongside, same era.
Achieving cross modal generalization with multimodal unified representation
Xia, Y., Huang, H., Zhu, J., and Zhao, Z · 2023
Cited alongside, same era.
How language models extrapolate outside the training data: A case study in textualized gridworld
Kim, D., Lee, J., Park, J., and Seo, M · 2024
Later among the works it cites.
Teaching arithmetic to small transformers
Lee, N., Sreenivasan, K., Lee, J. D., Lee, K., and Papailiopoulos, D · 2024
Later among the works it cites.
Suppress and rebalance: Towards generalized multi-modal face anti-spoofing
Lin, X., Wang, S., Cai, R., Liu, Y., Fu, Y., Tang, W., Yu, Z., and Kot, A · 2024
Later among the works it cites.
Benchmarking chatgpt on algorithmic reasoning
McLeish, S., Schwarzschild, A., and Goldstein, T · 2024
Later among the works it cites.
Ada2i: Enhancing modality balance for multimodal conversational emotion recognition
Nguyen, C.-V. T., Le, T.-S., Mai, A.-T., and Le, D.-T · 2024
Later among the works it cites.
Vision language models are blind
Rahmanzadehgervi, P., Bolton, L., Taesiri, M. R., and Nguyen, A. T · 2024
Later among the works it cites.
Understanding transformer reasoning capabilities via graph algorithms
Sanford, C., Fatemi, B., Hall, E., Tsitsulin, A., Kazemi, M., Halcrow, J., Perozzi, B., and Mirrokni, V · 2024
Later among the works it cites.
Easy-to-hard generalization: Scalable alignment beyond human supervision
Sun, Z., Yu, L., Shen, Y., Liu, W., Yang, Y., Welleck, S., and Gan, C · 2024
Later among the works it cites.
Are large-language models graph algorithmic reasoners?
Taylor, A. K., Cuturrufo, A., Yathish, V., Ma, M. D., and Wang, W · 2024
Later among the works it cites.
Cambrian-1: A fully open, vision-centric exploration of multimodal LLMs
Tong, S., II, E. L. B., Wu, P., Woo, S., IYER, A. J., Akula, S. C., Yang, S., Yang, J., Middepogu, M., Wang, Z., Pan, X., Fergus, R., LeCun, Y., and Xie, S · 2024
Later among the works it cites.
Enhancing multimodal cooperation via sample-level modality valuation
Wei, Y., Feng, R., Wang, Z., and Hu, D · 2024
Later among the works it cites.
LESS: Selecting influential data for targeted instruction tuning
Xia, M., Malladi, S., Gururangan, S., Arora, S., and Chen, D · 2024
Later among the works it cites.
Doremi: Optimizing data mixtures speeds up language model pretraining
Xie, S. M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P. S., Le, Q. V., Ma, T., and Yu, A. W · 2024
Later among the works it cites.
SKILL-MIX: a flexible and expandable family of evaluations for AI models
Yu, D., Kaur, S., Gupta, A., Brown-Cohen, J., Goyal, A., and Arora, S · 2024
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., et al · 2024
Later among the works it cites.
Can models learn skill composition from examples?
Zhao, H., Kaur, S., Yu, D., Goyal, A., and Arora, S · 2024
Later among the works it cites.
Qwen2.5-vl technical report, 2025
Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., and Lin, J · 2025
Closest in time.
Looped transformers for length generalization
Fan, Y., Du, Y., Ramchandran, K., and Lee, K · 2025
Closest in time.
Mmbench: Is your multi-modal model an all-around player?
Liu, Y., Duan, H., Zhang, Y., Li, B., Zhang, S., Zhao, W., Yuan, Y., Wang, J., He, C., Liu, Z., et al · 2025
Closest in time.
Eagle: Exploring the design space for multimodal LLMs with mixture of encoders
Shi, M., Liu, F., Wang, S., Liao, S., Radhakrishnan, S., Zhao, Y., Huang, D.-A., Yin, H., Sapra, K., Yacoob, Y., Shi, H., Catanzaro, B., Tao, A., Kautz, J., Yu, Z., and Liu, G · 2025
Closest in time.
Dualformer: Controllable fast and slow thinking by learning with randomized reasoning traces
Su, D., Sukhbaatar, S., Rabbat, M., Tian, Y., and Zheng, Q · 2025
Closest in time.