Fetching the paper…
Reading the bibliography…
Pre-trained vision-language models~(VLMs) are the de-facto foundation models for various downstream tasks.
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in ICML , 2006, pp. 369–376
2006
Earlier work this paper cites.
A. Krizhevsky, G. Hinton et al. , “Learning multiple layers of features from tiny images,” Technical Report , 2009
2009
Earlier work this paper cites.
K. Wang, B. Babenko, and S. J. Belongie, “End-to-end scene text recognition,” in ICCV , 2011, pp. 1457–1464
2011
Earlier work this paper cites.
A. Mishra, K. Alahari, and C. V. Jawahar, “Scene text recognition using higher order language priors,” in BMVC , 2012
2012
Earlier work this paper cites.
D. Karatzas, F. Shafait, S. Uchida, M. Iwamura, L. G. i Bigorda, S. R. Mestre, J. Mas, D. F. Mota, J. A. Almazan, and L. P. De Las Heras, “Icdar 2013 robust reading competition,” in ICDAR , 2013, pp. 1484–1493
2013
Earlier work this paper cites.
T. Q. Phan, P. Shivakumara, S. Tian, and C. L. Tan, “Recognizing text with perspective distortion in natural scenes,” in ICCV , 2013, pp. 569–576
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” JMLR , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
A. Risnumawan, P. Shivakumara, C. S. Chan, and C. L. Tan, “A robust arbitrary text detection system for natural scene images,” Expert Syst. Appl. , vol. 41, no. 18, pp. 8027–8048, 2014
2014
Earlier work this paper cites.
D. Karatzas, L. Gomez-Bigorda, A. Nicolaou, S. K. Ghosh, A. D. Bagdanov, M. Iwamura, J. Matas, L. Neumann, V. R. Chandrasekhar, S. Lu, F. Shafait, S. Uchida, and E. Valveny, “ICDAR 2015 competition on robust reading,” in ICDAR , 2015, pp. 1156–1160
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and L. Fei-Fei, “Imagenet large scale visual recognition challenge,” IJCV , pp. 211–252, 2015
2015
Earlier work this paper cites.
P. He, W. Huang, Y. Qiao, C. C. Loy, and X. Tang, “Reading scene text in deep convolutional sequences,” in AAAI , vol. 30, no. 1, 2016
2016
Earlier work this paper cites.
B. Shi, X. Bai, and C. Yao, “An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,” T-PAMI , vol. 39, no. 11, pp. 2298–2304, 2016
2016
Earlier work this paper cites.
C. Lee and S. Osindero, “Recursive recurrent nets with attention modeling for OCR in the wild,” in CVPR , 2016, pp. 2231–2239
2016
Earlier work this paper cites.
A. Gupta, A. Vedaldi, and A. Zisserman, “Synthetic data for text localisation in natural images,” in CVPR , 2016, pp. 2315–2324
2016
Earlier work this paper cites.
J. Lei Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” ArXiv e-prints , pp. arXiv–1607, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Y. Zhang, L. Gueguen, I. Zharkov, P. Zhang, K. Seifert, and B. Kadlec, “Uber-text: A large-scale dataset for optical character recognition from street-level imagery,” in SUNw: Scene Understanding Workshop-CVPR , vol. 2017, 2017, p. 5
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017, p. 5998–6008
2017
Earlier work this paper cites.
Z. Cheng, F. Bai, Y. Xu, G. Zheng, S. Pu, and S. Zhou, “Focusing attention: Towards accurate text recognition in natural images,” in ICCV , 2017, pp. 5076–5084
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
B. Shi, C. Yao, M. Liao, M. Yang, P. Xu, L. Cui, S. Belongie, S. Lu, and X. Bai, “Icdar2017 competition on reading chinese text in the wild (rctw-17),” in ICDAR , vol. 1, 2017, pp. 1429–1434
2017
Earlier work this paper cites.
I. Krasin, T. Duerig, N. Alldrin, V. Ferrari, S. Abu-El-Haija, A. Kuznetsova, H. Rom, J. Uijlings, S. Popov, A. Veit, S. Belongie, V. Gomes, A. Gupta, C. Sun, G. Chechik, D. Cai, Z. Feng, D. Narayanan, and K. Murphy, “Openimages: A public dataset for large-scale multi-label and multi-class image classification.” Dataset available from https://github.com/openimages , 2017
2017
Earlier work this paper cites.
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning filters for efficient convnets,” in ICLR . OpenReview.net, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
F. Borisyuk, A. Gordo, and V. Sivakumar, “Rosetta: Large scale system for text detection and recognition in images,” in ACM SIGKDD , 2018, pp. 71–79
2018
Earlier work this paper cites.
P. Micikevicius, S. Narang, J. Alben, G. F. Diamos, E. Elsen, D. García, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu, “Mixed precision training,” in ICLR , 2018
2018
Earlier work this paper cites.
P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan et al. , “Ray: A distributed framework for emerging { \{ AI } \} applications,” in OSDI , 2018, pp. 561–577
2018
Earlier work this paper cites.
C. K. Chng, Y. Liu, Y. Sun, C. C. Ng, C. Luo, Z. Ni, C. Fang, S. Zhang, J. Han, E. Ding et al. , “Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art,” in ICDAR , 2019, pp. 1571–1576
2019
Earlier work this paper cites.
M. Liao, J. Zhang, Z. Wan, F. Xie, J. Liang, P. Lyu, C. Yao, and X. Bai, “Scene text recognition from two-dimensional perspective,” in AAAI , vol. 33, no. 01, 2019, pp. 8714–8721
2019
Earlier work this paper cites.
B. Shi, M. Yang, X. Wang, P. Lyu, C. Yao, and X. Bai, “ASTER: an attentional scene text recognizer with flexible rectification,” T-PAMI , vol. 41, no. 9, pp. 2035–2048, 2019
2019
Earlier work this paper cites.
F. Sheng, Z. Chen, and B. Xu, “Nrtr: A no-recurrence sequence-to-sequence model for scene text recognition,” in ICDAR , 2019, pp. 781–786
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT , 2019
2019
Earlier work this paper cites.
Y. Sun, Z. Ni, C.-K. Chng, Y. Liu, C. Luo, C. C. Ng, J. Han, E. Ding, J. Liu, D. Karatzas et al. , “Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt,” in ICDAR , 2019, pp. 1557–1562
2019
Earlier work this paper cites.
N. Nayef, Y. Patel, M. Busta, P. N. Chowdhury, D. Karatzas, W. Khlif, J. Matas, U. Pal, J.-C. Burie, C.-l. Liu et al. , “Icdar2019 robust reading challenge on multi-lingual scene text detection and recognition—rrc-mlt-2019,” in ICDAR , 2019, pp. 1582–1587
2019
Earlier work this paper cites.
R. Zhang, Y. Zhou, Q. Jiang, Q. Song, N. Li, K. Zhou, L. Wang, D. Wang, M. Liao, M. Yang et al. , “Icdar 2019 robust reading challenge on reading chinese text on signboard,” in ICDAR , 2019, pp. 1577–1581
2019
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra, “Grad-cam: visual explanations from deep networks via gradient-based localization,” IJCV , vol. 128, pp. 336–359, 2020
2020
Cited alongside, same era.
Z. Wan, M. He, H. Chen, X. Bai, and C. Yao, “Textscanner: Reading characters in order for robust scene text recognition,” in AAAI , vol. 34, no. 07, 2020, pp. 12 120–12 127
Z. Raisi and J. Zelek, “Occluded text detection and recognition in the wild,” in CRV , 2022, pp. 140–150
2022
Later among the works it cites.
D. Bautista and R. Atienza, “Scene text recognition with permuted autoregressive sequence models,” in ECCV , 2022, pp. 178–196
2022
Later among the works it cites.
S. Zhao, L. Zhu, X. Wang, and Y. Yang, “Centerclip: Token clustering for efficient text-video retrieval,” in SIGIR , 2022, pp. 970–981
2022
Later among the works it cites.
X. Wang, L. Zhu, Z. Zheng, M. Xu, and Y. Yang, “Align and tell: Boosting text-video retrieval with local alignment and fine-grained supervision,” T-MM , vol. 25, pp. 6079–6089, 2022
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
P. Dai, H. Zhang, and X. Cao, “Sloan: Scale-adaptive orientation attention network for scene text recognition,” T-IP , vol. 30, pp. 1687–1701, 2020
2020
Cited alongside, same era.
D. Yu, X. Li, C. Zhang, T. Liu, J. Han, J. Liu, and E. Ding, “Towards accurate scene text recognition with semantic reasoning networks,” in CVPR , 2020, pp. 12 113–12 122
2020
Cited alongside, same era.
E. D. Cubuk, B. Zoph, J. Shlens, and Q. Le, “Randaugment: Practical automated data augmentation with a reduced search space,” in NeurIPS , 2020, pp. 702–703
2020
Cited alongside, same era.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” NeurIPS , vol. 33, pp. 1877–1901, 2020
2020
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021, pp. 8748–8763
2021
Cited alongside, same era.
C. Jia, Y. Yang, Y. Xia, Y. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in ICML , 2021, pp. 4904–4916
2021
Cited alongside, same era.
J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y. Choi, “Clipscore: A reference-free evaluation metric for image captioning,” in EMNLP , 2021, pp. 7514–7528
2021
Cited alongside, same era.
2022
Later among the works it cites.
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu, “Coca: Contrastive captioners are image-text foundation models.” TMLR , vol. 2022, 2022
2022
Later among the works it cites.
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang, “OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,” in ICML , 2022, pp. 23 318–23 340
2022
Later among the works it cites.
L. Yao, R. Huang, L. Hou, G. Lu, M. Niu, H. Xu, X. Liang, Z. Li, X. Jiang, and C. Xu, “FILIP: fine-grained interactive language-image pre-training,” in ICLR . OpenReview.net, 2022
2022
Later among the works it cites.
M. Byeon, B. Park, H. Kim, S. Lee, W. Baek, and S. Kim, “Coyo-700m: Image-text pair dataset,” https://github.com/kakaobrain/coyo-dataset , 2022
2022
Later among the works it cites.
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman et al. , “Laion-5b: An open large-scale dataset for training next generation image-text models,” NeurIPS , vol. 35, pp. 25 278–25 294, 2022
2022
Later among the works it cites.
S. Shen, L. H. Li, H. Tan, M. Bansal, A. Rohrbach, K. Chang, Z. Yao, and K. Keutzer, “How much can CLIP benefit vision-and-language tasks?” in ICLR . OpenReview.net, 2022
2022
Later among the works it cites.
L. Zhao, Z. Wu, X. Wu, G. Wilsbacher, and S. Wang, “Background-insensitive scene text recognition with text semantic segmentation,” in ECCV , 2022, pp. 163–182
2022
Later among the works it cites.
C. Da, P. Wang, and C. Yao, “Levenshtein OCR,” in ECCV , 2022, pp. 322–338
2022
Later among the works it cites.
B. Na, Y. Kim, and S. Park, “Multi-modal text recognition networks: Interactive enhancements between visual and semantic features,” in ECCV , 2022, pp. 446–463
2022
Later among the works it cites.
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds et al. , “Flamingo: a visual language model for few-shot learning,” NeurIPS , vol. 35, pp. 23 716–23 736, 2022
2022
Later among the works it cites.
Y. Wang, H. Xie, S. Fang, M. Xing, J. Wang, S. Zhu, and Y. Zhang, “Petr: Rethinking the capability of transformer-based language model in scene text recognition,” T-IP , vol. 31, pp. 5585–5598, 2022
2022
Later among the works it cites.
H. Bao, L. Dong, S. Piao, and F. Wei, “Beit: BERT pre-training of image transformers,” in ICLR , 2022
2022
Later among the works it cites.
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Learning to prompt for vision-language models,” IJCV , vol. 130, no. 9, pp. 2337–2348, 2022
2022
Later among the works it cites.
Y.-L. Sung, J. Cho, and M. Bansal, “Lst: Ladder side-tuning for parameter and memory efficient transfer learning,” NeurIPS , vol. 35, pp. 12 991–13 005, 2022
2022
Later among the works it cites.
M. Li, T. Lv, J. Chen, L. Cui, Y. Lu, D. Florencio, C. Zhang, Z. Li, and F. Wei, “Trocr: Transformer-based optical character recognition with pre-trained models,” in AAAI , vol. 37, no. 11, 2023, pp. 13 094–13 102
2023
Closest in time.
J. Zhang, T. Lin, Y. Xu, K. Chen, and R. Zhang, “Relational contrastive learning for scene text recognition,” in ACM MM , 2023, pp. 5764–5775
2023
Closest in time.
M. V. Ty and R. Atienza, “Scene text recognition models explainability using local features,” in ICIP , 2023, pp. 645–649
2023
Closest in time.
Q. Jiang, J. Wang, D. Peng, C. Liu, and L. Jin, “Revisiting scene text recognition: A data perspective,” in ICCV , 2023, pp. 20 543–20 554
2023
Closest in time.
A. Aberdam, D. Bensaïd, A. Golts, R. Ganz, O. Nuriel, R. Tichauer, S. Mazor, and R. Litman, “Clipter: Looking at the bigger picture in scene text recognition,” in ICCV , 2023, pp. 21 706–21 717
2023
Closest in time.
Z. Wang, H. Xie, Y. Wang, J. Xu, B. Zhang, and Y. Zhang, “Symmetrical linguistic feature distillation with clip for scene text recognition,” in ACM MM , 2023, pp. 509–518
2023
Closest in time.
T. Guan, C. Gu, J. Tu, X. Yang, Q. Feng, Y. Zhao, and W. Shen, “Self-supervised implicit glyph attention for text recognition,” in CVPR , 2023, pp. 15 285–15 294
2023
Closest in time.
C. Cheng, P. Wang, C. Da, Q. Zheng, and C. Yao, “Lister: Neighbor decoding for length-insensitive scene text recognition,” in ICCV , 2023, pp. 19 541–19 551
2023
Closest in time.
T. Guan, W. Shen, X. Yang, Q. Feng, Z. Jiang, and X. Yang, “Self-supervised character-to-character distillation for text recognition,” in ICCV , 2023, pp. 19 473–19 484
2023
Closest in time.
2023
Closest in time.
X. Yang, Z. Qiao, J. Wei, D. Yang, and Y. Zhou, “Masked and permuted implicit context learning for scene text recognition,” IEEE Signal Processing Letters , vol. 31, pp. 964–968, 2023
2023
Closest in time.
X. Yang, L. Zhu, X. Wang, and Y. Yang, “Dgl: Dynamic global-local prompt tuning for text-video retrieval,” in AAAI , vol. 38, no. 7, 2024, pp. 6540–6548
2024
Closest in time.
S. Zhao, X. Wang, L. Zhu, and Y. Yang, “Test-time adaptation with CLIP reward for zero-shot generalization in vision-language models,” in ICLR , 2024
2024
Closest in time.
M. Fujitake, “Dtrocr: Decoder-only transformer for optical character recognition,” in WACV , 2024, pp. 8025–8035
2024
Closest in time.
M. Rang, Z. Bi, C. Liu, Y. Wang, and K. Han, “An empirical study of scaling law for ocr,” in CVPR , 2024
2024
Closest in time.
S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang et al. , “Datacomp: In search of the next generation of multimodal datasets,” NeurIPS , vol. 36, 2024
2024
Closest in time.
P. Gao, S. Geng, R. Zhang, T. Ma, R. Fang, Y. Zhang, H. Li, and Y. Qiao, “Clip-adapter: Better vision-language models with feature adapters,” IJCV , vol. 132, no. 2, pp. 581–595, 2024
2024
Closest in time.