Fetching the paper…
Reading the bibliography…
Many real-world tasks require an agent to reason jointly over text and visual objects, (e.g., navigating in public spaces), which we refer to as context-sensitive text-rich visual reasoning.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D · 2015
Earlier work this paper cites.
A diagram is worth a dozen images
Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., and Farhadi, A · 2016
Earlier work this paper cites.
Rico: A mobile app dataset for building data-driven design applications
Deka, B., Huang, Z., Franzen, C., Hibschman, J., Afergan, D., Li, Y., Nichols, J., and Kumar, R · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people
Gurari, D., Li, Q., Stangl, A. J., Guo, A., Lin, C., Grauman, K., Luo, J., and Bigham, J. P · 2018
Earlier work this paper cites.
Scene text visual question answering
Biten, A. F., Tito, R., Mafla, A., Gomez, L., Rusinol, M., Valveny, E., Jawahar, C., and Karatzas, D · 2019
Earlier work this paper cites.
Openwebtext corpus
Gokaslan, A. and Cohen, V · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Li, L. H., Yatskar, M., Yin, D., Hsieh, C.-J., and Chang, K.-W · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
Mishra, A., Shekhar, S., Singh, A. K., and Chakraborty, A · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., Natarajan, V., Shah, M., Jiang, Y., Chen, X., Batra, D., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Tan, H. and Bansal, M · 2019
Earlier work this paper cites.
From recognition to cognition: Visual commonsense reasoning
Zellers, R., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Kiela, D., Firooz, H., Mohan, A., Goswami, V., Singh, A., Ringshia, P., and Testuggine, D · 2020
Earlier work this paper cites.
Bleurt: Learning robust metrics for text generation
Sellam, T., Das, D., and Parikh, A. P · 2020
Earlier work this paper cites.
Memeify: A large-scale meme generation system
Vyalla, S. R. and Udandarao, V · 2020
Cited alongside, same era.
On the general value of evidence, and bilingual scene-text visual question answering
Wang, X., Liu, Y., Shen, C., Ng, C. C., Luo, C., Jin, L., Chan, C. S., Hengel, A. v. d., and Wang, L · 2020
Cited alongside, same era.
Layoutlm: Pre-training of text and layout for document image understanding
Xu, Y., Li, M., Cui, L., Huang, S., Wei, F., and Zhou, M · 2020
Cited alongside, same era.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Changpinyo, S., Sharma, P., Ding, N., and Soricut, R · 2021
Cited alongside, same era.
Docvqa: A dataset for vqa on document images
Mathew, M., Karatzas, D., and Jawahar, C · 2021
Cited alongside, same era.
Broaden the vision: Geo-diverse visual commonsense reasoning
Llama-adapter v2: Parameter-efficient visual instruction model, 2023
Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., Li, H., and Qiao, Y · 2023
Later among the works it cites.
Hallusionbench: An advanced diagnostic suite for entangled language hallucination & visual illusion in large vision-language models
Guan, T., Liu, F., Li, X. W. R. X. Z., Wang, X. L. X., Yacoob, L. C. F. H. Y., and Zhou, D. M. T · 2023
Later among the works it cites.
Visual programming: Compositional visual reasoning without training
Gupta, T. and Kembhavi, A · 2023
Later among the works it cites.
Bliva: A simple multimodal llm for better handling of text-rich visual questions
Hu, W., Xu, Y., Li, Y., Li, W., Chen, Z., and Tu, Z · 2023
Later among the works it cites.
Introducing idefics: An open reproduction of state-of-the-art visual language model
HuggingFace · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yin, D., Li, L. H., Hu, Z., Peng, N., and Chang, K.-W · 2021
Cited alongside, same era.
Clip retrieval: Easily compute clip embeddings and build a clip retrieval system with them
Beaumont, R · 2022
Cited alongside, same era.
Unigeo: Unifying geometry logical reasoning via reformulating mathematical expression
Chen, J., Li, T., Qin, J., Lu, P., Lin, L., Chen, C., and Liang, X · 2022
Cited alongside, same era.
Ocr-free document understanding transformer
Kim, G., Hong, T., Yim, M., Nam, J., Park, J., Yim, J., Hwang, W., Yun, S., Han, D., and Park, S · 2022
Cited alongside, same era.
Matcha: Enhancing visual language pretraining with math reasoning and chart derendering
Liu, F., Piccinno, F., Krichene, S., Pang, C., Lee, K., Joshi, M., Altun, Y., Collier, N., and Eisenschlos, J. M · 2022
Cited alongside, same era.
Infographicvqa
Mathew, M., Bagal, V., Tito, R., Karatzas, D., Valveny, E., and Jawahar, C · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., et al · 2022
Cited alongside, same era.
Memecap: A dataset for captioning and interpreting memes
Hwang, E. and Shwartz, V · 2023
Later among the works it cites.
Pix2struct: Screenshot parsing as pretraining for visual language understanding
Lee, K., Joshi, M., Turc, I. R., Hu, H., Liu, F., Eisenschlos, J. M., Khandelwal, U., Shaw, P., Chang, M.-W., and Toutanova, K · 2023
Later among the works it cites.
Unichart: A universal vision-language pretrained model for chart comprehension and reasoning
Masry, A., Kavehzadeh, P., Do, X. L., Hoque, E., and Joty, S · 2023
Later among the works it cites.
Paddleocr: Multilingual ocr toolkit based on paddlepaddle
paddlepadle · 2023
Later among the works it cites.
Vipergpt: Visual inference via python execution for reasoning
Surís, D., Menon, S., and Vondrick, C · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G., Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Document understanding dataset and evaluation (dude)
Van Landeghem, J., Tito, R., Borchmann, Ł., Pietruszka, M., Joziak, P., Powalski, R., Jurkiewicz, D., Coustaty, M., Anckaert, B., Valveny, E., et al · 2023
Later among the works it cites.
Visual chatgpt: Talking, drawing and editing with visual foundation models
Wu, C., Yin, S., Qi, W., Wang, X., Tang, Z., and Duan, N · 2023
Later among the works it cites.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., Sun, Y., et al · 2023
Later among the works it cites.
Llavar: Enhanced visual instruction tuning for text-rich image understanding
Zhang, Y., Zhang, R., Gu, J., Zhou, Y., Lipka, N., Yang, D., and Sun, T · 2023
Later among the works it cites.
Generalized decoding for pixel, image, and language
Zou, X., Dou, Z.-Y., Yang, J., Gan, Z., Li, L., Li, C., Dai, X., Behl, H., Wang, J., Yuan, L., et al · 2023
Later among the works it cites.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Liu, H., Li, C., Li, Y., Li, B., Zhang, Y., Shen, S., and Lee, Y. J · 2024
Closest in time.