Fetching the paper…
Reading the bibliography…
In the era of Large Language Models (LLMs), tremendous strides have been made in the field of multimodal understanding.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Scene text recognition using higher order language priors
Mishra, A.; Alahari, K.; and Jawahar, C. 2012 · 2012
Earlier work this paper cites.
Detecting texts of arbitrary orientations in natural images
Yao, C.; Bai, X.; Liu, W.; Ma, Y.; and Tu, Z. 2012 · 2012
Earlier work this paper cites.
Total-text: A comprehensive dataset for scene text detection and recognition
Ch’ng, C. K.; and Chan, C. S. 2017 · 2017
Earlier work this paper cites.
Towards end-to-end text spotting with convolutional recurrent neural networks
Li, H.; Wang, P.; and Shen, C. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I.; and Hutter, F. 2017 · 2017
Earlier work this paper cites.
Curved scene text detection via transverse and longitudinal sequence connection
Liu, Y.; Jin, L.; Zhang, S.; Luo, C.; and Zhang, S. 2019 · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A.; Natarajan, V.; Shah, M.; Jiang, Y.; Chen, X.; Batra, D.; Parikh, D.; and Rohrbach, M. 2019 · 2019
Earlier work this paper cites.
Super-convergence: Very fast training of neural networks using large learning rates
Smith, L. N.; and Topin, N. 2019 · 2019
Cited alongside, same era.
Pix2seq: A Language Modeling Framework for Object Detection
Chen, T.; Saxena, S.; Li, L.; Fleet, D. J.; and Hinton, G. 2021 · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021 · 2021
Cited alongside, same era.
Bloom: A 176b-parameter open-access multilingual language model
Scao, T. L.; Fan, A.; Akiki, C.; Pavlick, E.; Ilić, S.; Hesslow, D.; Castagné, R.; Luccioni, A. S.; Yvon, F.; Gallé, M.; et al. 2022 · 2022
Cited alongside, same era.
Laion-5b: An open large-scale dataset for training next generation image-text models
Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; et al. 2022 · 2022
Cited alongside, same era.
Shikra: Unleashing Multimodal LLM’s Referential Dialogue Magic
Chen, K.; Zhang, Z.; Zeng, W.; Zhang, R.; Zhu, F.; and Zhao, R. 2023 · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; et al. 2023 · 2023
Closest in time.
Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023 · 2023
Closest in time.
Stanford Alpaca: An Instruction-following LLaMA model
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Closest in time.
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Zhang, Y.; Zhang, R.; Gu, J.; Zhou, Y.; Lipka, N.; Yang, D.; and Sun, T. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Self-instruct: Aligning language model with self generated instructions
Wang, Y.; Kordi, Y.; Mishra, S.; Liu, A.; Smith, N. A.; Khashabi, D.; and Hajishirzi, H. 2022 · 2022
Cited alongside, same era.
Toward understanding wordart: Corner-guided transformer for scene text recognition
Xie, X.; Fu, L.; Zhang, Z.; Wang, Z.; and Bai, X. 2022 · 2022
Cited alongside, same era.
Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023a
Cited in the paper.
On the hidden mystery of ocr in large multimodal models
Liu, Y.; Li, Z.; Li, H.; Yu, W.; Huang, M.; Peng, D.; Liu, M.; Chen, M.; Li, C.; Jin, L.; et al. 2023b
Cited in the paper.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023a
Cited in the paper.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023b
Cited in the paper.
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Ye, J.; Hu, A.; Xu, H.; Ye, Q.; Yan, M.; Dan, Y.; Zhao, C.; Xu, G.; Li, C.; Tian, J.; et al. 2023a
Cited in the paper.
Closest in time.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E. P.; Zhang, H.; Gonzalez, J. E.; and Stoica, I. 2023 · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Zhu, D.; Chen, J.; Shen, X.; Li, X.; and Elhoseiny, M. 2023 · 2023
Closest in time.