Fetching the paper…
Reading the bibliography…
Large models have recently played a dominant role in natural language processing and multimodal vision-language learning.
The iam-database: an english sentence database for offline handwriting recognition
U-V Marti and Horst Bunke · 2002
Earlier work this paper cites.
Top-down and bottom-up cues for scene text recognition
Anand Mishra, Karteek Alahari, and CV Jawahar · 2012
Earlier work this paper cites.
ICDAR 2013 robust reading competition
Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, M. Iwamura, Lluís Gómez i Bigorda, Sergi Robles Mestre, Joan Mas Romeu, David Fernández Mota, Jon Almazán, and Lluís-Pere de las Heras · 2013
Earlier work this paper cites.
Recognizing text with perspective distortion in natural scenes
Trung Quy Phan, Palaiahnakote Shivakumara, Shangxuan Tian, and Chew Lim Tan · 2013
Earlier work this paper cites.
End-to-end scene text recognition using tree-structured models
Cunzhao Shi, Chunheng Wang, Baihua Xiao, Song Gao, and Jinlong Hu · 2014
Earlier work this paper cites.
A robust arbitrary text detection system for natural scene images
Anhar Risnumawan, Palaiahnakote Shivakumara, Chee Seng Chan, and Chew Lim Tan · 2014
Earlier work this paper cites.
Icfhr 2014 competition on handwritten digit string recognition in challenging datasets (hdsrc 2014)
Markus Diem, Stefan Fiel, Florian Kleber, Robert Sablatnig, Jose M. Saavedra, David Contreras, Juan Manuel Barrios, and Luiz S. Oliveira · 2014
Earlier work this paper cites.
Synthetic data and artificial neural networks for natural scene text recognition
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
ICDAR 2015 competition on robust reading
Dimosthenis Karatzas, Lluís Gómez i Bigorda, Anguelos Nicolaou, Suman K. Ghosh, Andrew D. Bagdanov, M. Iwamura, Jiri Matas, Lukás Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, Faisal Shafait, Seiichi Uchida, and Ernest Valveny · 2015
Earlier work this paper cites.
COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images
Andreas Veit, Tomas Matera, Lukás Neumann, Jiri Matas, and Serge J. Belongie · 2016
Earlier work this paper cites.
Synthetic data for text localisation in natural images
Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman · 2016
Earlier work this paper cites.
Total-Text: A comprehensive dataset for scene text detection and recognition
Chee-Kheng Chng and Chee Seng Chan · 2017
Earlier work this paper cites.
Scene text visual question answering
Ali Furkan Biten, Rubèn Pérez Tito, Andrés Mafla, Lluís Gómez, Marçal Rusiñol, Ernest Valveny, C. V. Jawahar, and Dimosthenis Karatzas · 2019
Earlier work this paper cites.
Curved scene text detection via transverse and longitudinal sequence connection
Yuliang Liu, Lianwen Jin, Shuaitao Zhang, Canjie Luo, and Sheng Zhang · 2019
Earlier work this paper cites.
Icdar 2019 robust reading challenge on reading chinese text on signboard
Rui Zhang, Yongsheng Zhou, Qianyi Jiang, Qi Song, Nan Li, Kai Zhou, Lei Wang, Dong Wang, Minghui Liao, Mingkun Yang, et al · 2019
Earlier work this paper cites.
Icdar 2019 competition on scene text visual question answering
Ali Furkan Biten, Rubèn Tito, Andrés Mafla, Lluis Gomez, Marçal Rusiñol, Minesh Mathew, C.V. Jawahar, Ernest Valveny, and Dimosthenis Karatzas · 2019
Earlier work this paper cites.
Towards VQA models that can read, 2019
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach · 2019
Earlier work this paper cites.
OCR-VQA: visual question answering by reading text in images
Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty · 2019
Earlier work this paper cites.
ICDAR2019 competition on scanned receipt OCR and information extraction
Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and C. V. Jawahar · 2019
Earlier work this paper cites.
FUNSD: A dataset for form understanding in noisy scanned documents, 2019
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran · 2019
Earlier work this paper cites.
Cord: A consolidated receipt dataset for post-ocr parsing
Seunghyun Park, Seung Shin, Bado Lee, Junyeop Lee, Jaeheung Surh, Minjoon Seo, and Hwalsuk Lee · 2019
Earlier work this paper cites.
On the general value of evidence, and bilingual scene-text visual question answering
Xinyu Wang, Yuliang Liu, Chunhua Shen, Chun Chet Ng, Canjie Luo, Lianwen Jin, Chee Seng Chan, Anton van den Hengel, and Liangwei Wang · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Earlier work this paper cites.
Florence: A new foundation model for computer vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al · 2021
Earlier work this paper cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig · 2021
Earlier work this paper cites.
From two to one: A new scene text recognizer with visual language modeling network
Yuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang, Shenggao Zhu, and Yongdong Zhang · 2021
Earlier work this paper cites.
DocVQA: A dataset for vqa on document images, 2021
Minesh Mathew, Dimosthenis Karatzas, and C. V. Jawahar · 2021
Earlier work this paper cites.
Winner team mia at textvqa challenge 2021: Vision-and-language representation learning with pre-trained sequence-to-sequence model, 2021
Yixuan Qiao, Hao Chen, Jun Wang, Yihao Chen, Xianbin Ye, Ziliang Li, Xianbiao Qi, Peng Gao, and Guotong Xie · 2021
Earlier work this paper cites.
Structext: Structured text understanding with multi-modal transformers
Yulin Li, Yuxi Qian, Yuchen Yu, Xiameng Qin, Chengquan Zhang, Yang Liu, Kun Yao, Junyu Han, Jingtuo Liu, and Errui Ding · 2021
Earlier work this paper cites.
Vision-language pre-training: Basics, recent advances, and future trends
Zhe Gan, Linjie Li, Chunyuan Li, Lijuan Wang, Zicheng Liu, Jianfeng Gao, et al · 2022
Earlier work this paper cites.
ELEVATER: A benchmark and toolkit for evaluating language-augmented visual models
Chunyuan Li, Haotian Liu, Liunian Harold Li, Pengchuan Zhang, Jyoti Aneja, Jianwei Yang, Ping Jin, Houdong Hu, Zicheng Liu, Yong Jae Lee, and Jianfeng Gao · 2022
Earlier work this paper cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al · 2022
Earlier work this paper cites.
Git: A generative image-to-text transformer for vision and language
Jianfeng Wang, Zhengyuan Yang, Xiaowei Hu, Linjie Li, Kevin Lin, Zhe Gan, Zicheng Liu, Ce Liu, and Lijuan Wang · 2022
Earlier work this paper cites.
Toward understanding WordArt: Corner-guided transformer for scene text recognition
Xudong Xie, Ling Fu, Zhifei Zhang, Zhaowen Wang, and Xiang Bai · 2022
Earlier work this paper cites.
Infographicvqa
Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar · 2022
Earlier work this paper cites.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Ahmed Masry, Do Xuan Long, Jia Qing Tan, Shafiq Joty, and Enamul Hoque · 2022
Earlier work this paper cites.
Syntax-aware network for handwritten mathematical expression recognition
Ye Yuan, Xiao Liu, Wondimu Dikubab, Hui Liu, Zhilong Ji, Zhongqin Wu, and Xiang Bai · 2022
Earlier work this paper cites.
Glm: General language model pretraining with autoregressive blank infilling
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang · 2022
Earlier work this paper cites.
OPT: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Cited alongside, same era.
Scene text recognition with permuted autoregressive sequence models
Darwin Bautista and Rowel Atienza · 2022
Cited alongside, same era.
Attentionhtr: Handwritten text recognition based on attention encoder-decoder networks, 2022
Dmitrijs Kass and Ekta Vats · 2022
Cited alongside, same era.
An efficient prototype-based model for handwritten text recognition with multi-loss fusion
Ming-Ming Yu, Heng Zhang, Fei Yin, and Cheng-Lin Liu · 2022
Cited alongside, same era.
Ernie-layout: Layout knowledge enhanced pre-training for visually-rich document understanding
Qiming Peng, Yinxu Pan, Wenjin Wang, Bin Luo, Zhenyu Zhang, Zhengjie Huang, Teng Hu, Weichong Yin, Yongfeng Chen, Yin Zhang, Shi Feng, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang · 2022
Jiabang He, Lei Wang, Yingpeng Hu, Ning Liu, Hui juan Liu, Xingdong Xu, and Hengtao Shen · 2023
Closest in time.
Bo Li, Gexiang Fang, Yang Yang, Quansen Wang, Wei Ye, Wen Zhao, and Shikun Zhang · 2023
Closest in time.
Towards expert-level medical question answering with large language models, 2023
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, Mike Schaekermann, Amy Wang, Mohamed Amin, Sami Lachgar, Philip Mansfield, Sushant Prakash, Bradley Green, Ewa Dominowska, Blaise Aguera y Arcas, Nenad Tomasev, Yun Liu, Renee Wong, Christopher Semturs, S. Sara Mahdavi, Joelle Barral, Dale Webster, Greg S. Corrado, Yossi Matias, Shekoofeh Azizi, Alan Karthikesalingam, and Vivek Natarajan · 2023
Closest in time.
Llava-med: Training a large language-and-vision assistant for biomedicine in one day, 2023
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cross-modal attention networks with modality disentanglement for scene-text vqa
Zeqin Fang, Lin Li, Zhongwei Xie, and Jingling Yuan · 2022
Cited alongside, same era.
Deplot: One-shot visual language reasoning by plot-to-table translation
Fangyu Liu, Julian Martin Eisenschlos, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Wenhu Chen, Nigel Collier, and Yasemin Altun · 2022
Cited alongside, same era.
When counting meets hmer: Counting-aware network for handwritten mathematical expression recognition
Bohan Li, Ye Yuan, Dingkang Liang, Xiao Liu, Zhilong Ji, Jinfeng Bai, Wenyu Liu, and Xiang Bai · 2022
Cited alongside, same era.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Vicuna · 2023
Cited alongside, same era.
Closest in time.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al · 2023
Closest in time.
Dublin–document understanding by language-image network
Kriti Aggarwal, Aditi Khandelwal, Kumar Tanmay, Owais Mohammed Khan, Qiang Liu, Monojit Choudhury, Subhojit Som, Vishrav Chaudhary, and Saurabh Tiwary · 2023
Closest in time.
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2024
Closest in time.
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee · 2024
Closest in time.
Textmonkey: An ocr-free large multimodal model for understanding document, 2024
Yuliang Liu, Biao Yang, Qiang Liu, Zhang Li, Zhiyin Ma, Shuo Zhang, and Xiang Bai · 2024
Closest in time.
Vlmevalkit: An open-source toolkit for evaluating large multi-modality models, 2024
Haodong Duan, Junming Yang, Yuxuan Qiao, Xinyu Fang, Lin Chen, Yuan Liu, Xiaoyi Dong, Yuhang Zang, Pan Zhang, Jiaqi Wang, Dahua Lin, and Kai Chen · 2024
Closest in time.
Lmms-eval: Accelerating the development of large multimoal models, March 2024
Bo Li*, Peiyuan Zhang*, Kaichen Zhang*, Fanyi Pu*, Yuhao Dong Xinrun Du, Haotian Liu, Yuanhan Zhang, Ge Zhang, Chunyuan Li, and Ziwei Liu · 2024
Closest in time.
Minicpm-v2.6
OpenBMB · 2024
Closest in time.
Cambrian-1: A fully open, vision-centric exploration of multimodal llms, 2024
Shengbang Tong, Ellis Brown, Penghao Wu, Sanghyun Woo, Manoj Middepogu, Sai Charitha Akula, Jihan Yang, Shusheng Yang, Adithya Iyer, Xichen Pan, Austin Wang, Rob Fergus, Yann LeCun, and Saining Xie · 2024
Closest in time.
How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites, 2024
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhangwei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, Ji Ma, Jiaqi Wang, Xiaoyi Dong, Hang Yan, Hewei Guo, Conghui He, Botian Shi, Zhenjiang Jin, Chao Xu, Bin Wang, Xingjian Wei, Wei Li, Wenjian Zhang, Bo Zhang, Pinlong Cai, Licheng Wen, Xiangchao Yan, Min Dou, Lewei Lu, Xizhou Zhu, Tong Lu, Dahua Lin, Yu Qiao, Jifeng Dai, and Wenhai Wang · 2024
Closest in time.
Paligemma: A versatile 3b vlm for transfer, 2024
Lucas Beyer, Andreas Steiner, André Susano Pinto, Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, Thomas Unterthiner, Daniel Keysers, Skanda Koppula, Fangyu Liu, Adam Grycner, Alexey Gritsenko, Neil Houlsby, Manoj Kumar, Keran Rong, Julian Eisenschlos, Rishabh Kabra, Matthias Bauer, Matko Bošnjak, Xi Chen, Matthias Minderer, Paul Voigtlaender, Ioana Bica, Ivana Balazevic, Joan Puigcerver, Pinelopi Papalampidi, Olivier Henaff, Xi Xiong, Radu Soricut, Jeremiah Harmsen, and Xiaohua Zhai · 2024
Closest in time.
Congrong
CloudWalk · 2024
Closest in time.
Minicpm-v-2
OpenBMB · 2024
Closest in time.
Mini-monkey: Alleviate the sawtooth effect by multi-scale adaptive cropping, 2024
Mingxin Huang, Yuliang Liu, Dingkang Liang, Lianwen Jin, and Xiang Bai · 2024
Closest in time.
Claude3.5-sonnet
Anthropic · 2024
Closest in time.
Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee · 2024
Closest in time.
Gpt-4o-mini-20240718
OpenAI · 2024
Closest in time.
Internlm-xcomposer2: Mastering free-form text-image composition and comprehension in vision-language large model, 2024
Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Xilin Wei, Songyang Zhang, Haodong Duan, Maosong Cao, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Xinyue Zhang, Wei Li, Jingwen Li, Kai Chen, Conghui He, Xingcheng Zhang, Yu Qiao, Dahua Lin, and Jiaqi Wang · 2024
Closest in time.
Rekaflash
Reka AI · 2024
Closest in time.
Gemini models
Google · 2024
Closest in time.
Xverse-v
XVERSE · 2024
Closest in time.
Ovis: Structural embedding alignment for multimodal large language model, 2024
Shiyin Lu, Yang Li, Qing-Guo Chen, Zhao Xu, Weihua Luo, Kaifu Zhang, and Han-Jia Ye · 2024
Closest in time.
Minicpm-llama3-v-2.5
OpenBMB · 2024
Closest in time.
Generative multimodal models are in-context learners, 2024
Quan Sun, Yufeng Cui, Xiaosong Zhang, Fan Zhang, Qiying Yu, Zhengxiong Luo, Yueze Wang, Yongming Rao, Jingjing Liu, Tiejun Huang, and Xinlong Wang · 2024
Closest in time.
Deepseek-vl: Towards real-world vision-language understanding, 2024
Haoyu Lu, Wen Liu, Bo Zhang, Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, Yaofeng Sun, Chengqi Deng, Hanwei Xu, Zhenda Xie, and Chong Ruan · 2024
Closest in time.
Omnilmm-12b
OpenBMB · 2024
Closest in time.
Transcore-m
PCI Research · 2024
Closest in time.
Internlm-xcomposer-2.5: A versatile large vision language model supporting long-contextual input and output, 2024
Pan Zhang, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Rui Qian, Lin Chen, Qipeng Guo, Haodong Duan, Bin Wang, Linke Ouyang, Songyang Zhang, Wenwei Zhang, Yining Li, Yang Gao, Peng Sun, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Hang Yan, Conghui He, Xingcheng Zhang, Kai Chen, Jifeng Dai, Yu Qiao, Dahua Lin, and Jiaqi Wang · 2024
Closest in time.
Internlm-xcomposer2-4khd: A pioneering large vision-language model handling resolutions from 336 pixels to 4k hd, 2024
Xiaoyi Dong, Pan Zhang, Yuhang Zang, Yuhang Cao, Bin Wang, Linke Ouyang, Songyang Zhang, Haodong Duan, Wenwei Zhang, Yining Li, Hang Yan, Yang Gao, Zhe Chen, Xinyue Zhang, Wei Li, Jingwen Li, Wenhai Wang, Kai Chen, Conghui He, Xingcheng Zhang, Jifeng Dai, Yu Qiao, Dahua Lin, and Jiaqi Wang · 2024
Closest in time.
Minicpm-v
OpenBMB · 2024
Closest in time.
Yi-vl-34b
01-AI · 2024
Closest in time.
Phi-3-vision
Microsoft · 2024
Closest in time.
What matters when building vision-language models?, 2024
Hugo Laurençon, Léo Tronchon, Matthieu Cord, and Victor Sanh · 2024
Closest in time.
Chameleon: Mixed-modal early-fusion foundation models, 2024
Chameleon Team · 2024
Closest in time.