Fetching the paper…
Reading the bibliography…
Currently, vision encoder models like Vision Transformers (ViTs) typically excel at image recognition tasks but cannot simultaneously support text recognition like human visual recognition.
End-to-end scene text recognition
K. Wang, B. Babenko, and S. Belongie · 2011
Earlier work this paper cites.
End-to-end text recognition with convolutional neural networks
T. Wang, D. J. Wu, A. Coates, and A. Y. Ng · 2012
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Earlier work this paper cites.
Full-page text recognition: Learning where to start and when to stop
B. Moysset, C. Kermorvant, and C. Wolf · 2017
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
P. Sharma, N. Ding, S. Goodman, and R. Soricut · 2018
Earlier work this paper cites.
MMDetection: Open mmlab detection toolbox and benchmark
K. Chen, J. Wang, J. Pang, Y. Cao, Y. Xiong, X. Li, S. Sun, W. Feng, Z. Liu, J. Xu, Z. Zhang, D. Cheng, C. Zhu, T. Cheng, Q. Zhao, B. Li, X. Lu, R. Zhu, Y. Wu, J. Dai, J. Wang, J. Shi, W. Ouyang, C. C. Loy, and D. Lin · 2019
Earlier work this paper cites.
GQA: a new dataset for compositional question answering over real-world images
D. A. Hudson and C. D. Manning · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
G. Jaume, H. K. Ekenel, and J.-P. Thiran · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
K. Marino, M. Rastegari, A. Farhadi, and R. Mottaghi · 2019
Earlier work this paper cites.
End-to-end object detection with transformers
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko · 2020
Earlier work this paper cites.
MMSegmentation: Openmmlab semantic segmentation toolbox and benchmark
M. Contributors · 2020
Earlier work this paper cites.
High-performance large-scale image recognition without normalization
A. Brock, S. De, S. L. Smith, and K. Simonyan · 2021
Earlier work this paper cites.
Text recognition in the wild: A survey
X. Chen, L. Jin, Y. Zhu, C. Luo, and T. Wang · 2021
Earlier work this paper cites.
Dynamic detr: End-to-end object detection with dynamic attention
X. Dai, Y. Chen, J. Yang, P. Zhang, L. Yuan, and L. Zhang · 2021
Earlier work this paper cites.
Rethinking text line recognition models
D. H. Diaz, S. Qin, R. Ingle, Y. Fujii, and A. Bissacco · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Earlier work this paper cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo · 2021
Earlier work this paper cites.
Docvqa: A dataset for vqa on document images
M. Mathew, D. Karatzas, and C. Jawahar · 2021
Earlier work this paper cites.
An end-to-end transformer model for 3d object detection
I. Misra, R. Girdhar, and A. Joulin · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki · 2021
Earlier work this paper cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
W. Wang, E. Xie, X. Li, D.-P. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao · 2021
Earlier work this paper cites.
Cvt: Introducing convolutions to vision transformers
H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, L. Yuan, and L. Zhang · 2021
Earlier work this paper cites.
Segformer: Simple and efficient design for semantic segmentation with transformers
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
J. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. L. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan · 2022
Cited alongside, same era.
Scene text recognition with permuted autoregressive sequence models
D. Bautista and R. Atienza · 2022
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2022
Cited alongside, same era.
Ocr-free document understanding transformer
G. Kim, T. Hong, M. Yim, J. Nam, J. Park, J. Yim, W. Hwang, S. Yun, D. Han, and S. Park · 2022
Cited alongside, same era.
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
J. Li, D. Li, S. Savarese, and S. Hoi · 2023
Later among the works it cites.
Dinov2: Learning robust visual features without supervision, 2023
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y. Huang, S.-W. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski · 2023
Later among the works it cites.
B. Peng, C. Li, P. He, M. Galley, and J. Gao · 2023
Later among the works it cites.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Y. Shen, K. Song, X. Tan, D. Li, W. Lu, and Y. Zhuang · 2023
Later among the works it cites.
Vita-clip: Video and text adaptive clip via multimodal prompting
S. T. Wasim, M. Naseer, S. Khan, F. S. Khan, and M. Shah · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation
J. Li, D. Li, C. Xiong, and S. C. H. Hoi · 2022
Cited alongside, same era.
Mvitv2: Improved multiscale vision transformers for classification and detection
Y. Li, C.-Y. Wu, H. Fan, K. Mangalam, B. Xiong, J. Malik, and C. Feichtenhofer · 2022
Cited alongside, same era.
Swin transformer v2: Scaling up capacity and resolution
Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang, L. Dong, et al · 2022
Cited alongside, same era.
A convnet for the 2020s, 2022
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie · 2022
Cited alongside, same era.
Chartqa: A benchmark for question answering about charts with visual and logical reasoning
A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque · 2022
Cited alongside, same era.
Infographicvqa
M. Mathew, V. Bagal, R. Tito, D. Karatzas, E. Valveny, and C. Jawahar · 2022
Cited alongside, same era.
Maxvit: Multi-axis vision transformer, 2022
Z. Tu, H. Talebi, H. Zhang, F. Yang, P. Milanfar, A. Bovik, and Y. Li · 2022
Cited alongside, same era.
Later among the works it cites.
Vary: Scaling up the vision vocabulary for large vision-language models
H. Wei, L. Kong, J. Chen, L. Zhao, Z. Ge, J. Yang, J. Sun, C. Han, and X. Zhang · 2023
Later among the works it cites.
Visual chatgpt: Talking, drawing and editing with visual foundation models
C. Wu, S. Yin, W. Qi, X. Wang, Z. Tang, and N. Duan · 2023
Later among the works it cites.
Gpt4tools: Teaching large language model to use tools via self-instruction
R. Yang, L. Song, Y. Li, S. Zhao, Y. Ge, X. Li, and Y. Shan · 2023
Later among the works it cites.
mplug-docowl: Modularized multimodal large language model for document understanding
J. Ye, A. Hu, H. Xu, Q. Ye, M. Yan, Y. Dan, C. Zhao, G. Xu, C. Li, J. Tian, et al · 2023
Later among the works it cites.
J. Ye, A. Hu, H. Xu, Q. Ye, M. Yan, G. Xu, C. Li, J. Tian, Q. Qian, J. Zhang, et al · 2023
Later among the works it cites.
mplug-owl: Modularization empowers large language models with multimodality
Q. Ye, H. Xu, G. Xu, J. Ye, M. Yan, Y. Zhou, J. Wang, A. Hu, P. Shi, Y. Shi, et al · 2023
Later among the works it cites.
Clip2: Contrastive language-image-point pretraining from real-world point cloud data
Y. Zeng, C. Jiang, J. Mao, J. Han, C. Ye, Q. Huang, D.-Y. Yeung, Z. Yang, X. Liang, and H. Xu · 2023
Later among the works it cites.
Sigmoid loss for language image pre-training
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer · 2023
Later among the works it cites.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2023
Later among the works it cites.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny · 2023
Later among the works it cites.
Biformer: Vision transformer with bi-level routing attention
L. Zhu, X. Wang, Z. Ke, W. Zhang, and R. W. Lau · 2023
Later among the works it cites.
Coda: Collaborative novel box discovery and cross-modal alignment for open-vocabulary 3d object detection
Y. Cao, Z. Yihan, H. Xu, and D. Xu · 2024
Closest in time.
Visual instruction tuning
H. Liu, C. Li, Q. Wu, and Y. J. Lee · 2024
Closest in time.
Mm1: Methods, analysis & insights from multimodal llm pre-training
B. McKinzie, Z. Gan, J.-P. Fauconnier, S. Dodge, B. Zhang, P. Dufter, D. Shah, X. Du, F. Peng, F. Weers, et al · 2024
Closest in time.
Am-radio: Agglomerative vision foundation model reduce all domains into one
M. Ranzinger, G. Heinrich, J. Kautz, and P. Molchanov · 2024
Closest in time.
Small language model meets with reinforced vision vocabulary
H. Wei, L. Kong, J. Chen, L. Zhao, Z. Ge, E. Yu, J. Sun, C. Han, and X. Zhang · 2024
Closest in time.
mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration
Q. Ye, H. Xu, J. Ye, M. Yan, A. Hu, H. Liu, Q. Qian, J. Zhang, and F. Huang · 2024
Closest in time.
Vision transformer with quadrangle attention
Q. Zhang, J. Zhang, Y. Xu, and D. Tao · 2024
Closest in time.