Fetching the paper…
Reading the bibliography…
The Segment Anything Model (SAM), a profound vision foundation model pretrained on a large-scale dataset, breaks the boundaries of general segmentation and sparks various downstream applications.
X. Zu, H. Yu, B. Li, and X. Xue, “Weakly-supervised text instance segmentation,” in ACM MM , 2023, pp. 1915–1923
1923
Earlier work this paper cites.
C. Yao, X. Bai, W. Liu, Y. Ma, and Z. Tu, “Detecting texts of arbitrary orientations in natural images,” in CVPR , 2012, pp. 1083–1090
2012
Earlier work this paper cites.
D. Karatzas, L. Gomez-Bigorda, A. Nicolaou, S. Ghosh, A. Bagdanov, M. Iwamura, J. Matas, L. Neumann, V. R. Chandrasekhar, S. Lu et al. , “Icdar 2015 competition on robust reading,” in ICDAR , 2015, pp. 1156–1160
2015
Earlier work this paper cites.
F. Milletari, N. Navab, and S.-A. Ahmadi, “V-net: Fully convolutional neural networks for volumetric medical image segmentation,” in 3DV , 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
C. K. Ch’ng and C. S. Chan, “Total-text: A comprehensive dataset for scene text detection and recognition,” in ICDAR , vol. 1, 2017, pp. 935–942
2017
Earlier work this paper cites.
M. Liao, B. Shi, X. Bai, X. Wang, and W. Liu, “Textboxes: A fast text detector with a single deep neural network,” in AAAI , vol. 31, no. 1, 2017
2017
Earlier work this paper cites.
S. Schreiber, S. Agne, I. Wolf, A. Dengel, and S. Ahmed, “Deepdesrt: Deep learning for detection and structure recognition of tables in document images,” in ICDAR , vol. 1, 2017, pp. 1162–1167
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask r-cnn,” in ICCV , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in ICCV , 2017, pp. 2980–2988
2017
Earlier work this paper cites.
M. Liao, B. Shi, and X. Bai, “Textboxes++: A single-shot oriented scene text detector,” IEEE TIP , vol. 27, no. 8, pp. 3676–3690, 2018
2018
Earlier work this paper cites.
L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-decoder with atrous separable convolution for semantic image segmentation,” in ECCV , 2018, pp. 801–818
2018
Earlier work this paper cites.
Y. Liu, L. Jin, S. Zhang, C. Luo, and S. Zhang, “Curved scene text detection via transverse and longitudinal sequence connection,” PR , vol. 90, pp. 337–345, 2019
2019
Earlier work this paper cites.
C. K. Chng, Y. Liu, Y. Sun, C. C. Ng, C. Luo, Z. Ni, C. Fang, S. Zhang, J. Han, E. Ding et al. , “Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art,” in ICDAR , 2019, pp. 1571–1576
2019
Earlier work this paper cites.
Y. Sun, Z. Ni, C.-K. Chng, Y. Liu, C. Luo, C. C. Ng, J. Han, E. Ding, J. Liu, D. Karatzas et al. , “Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt,” in ICDAR , 2019, pp. 1557–1562
2019
Earlier work this paper cites.
Y. Baek, B. Lee, D. Han, S. Yun, and H. Lee, “Character region awareness for text detection,” in CVPR , 2019, pp. 9365–9374
2019
Earlier work this paper cites.
W. Wang, E. Xie, X. Li, W. Hou, T. Lu, G. Yu, and S. Shao, “Shape robust text detection with progressive scale expansion network,” in CVPR , 2019, pp. 9336–9345
2019
Earlier work this paper cites.
W. Wang, E. Xie, X. Song, Y. Zang, W. Wang, T. Lu, G. Yu, and C. Shen, “Efficient and accurate arbitrary-shaped text detection with pixel aggregation network,” in ICCV , 2019, pp. 8440–8449
2019
Earlier work this paper cites.
J. Lee, H. Hayashi, W. Ohyama, and S. Uchida, “Page segmentation using a convolutional neural network with trainable co-occurrence features,” in ICDAR , 2019, pp. 1023–1028
2019
Earlier work this paper cites.
X. Zhong, J. Tang, and A. J. Yepes, “Publaynet: largest dataset ever for document layout analysis,” in ICDAR , 2019, pp. 1015–1022
2019
Earlier work this paper cites.
S. Bonechi, P. Andreini, M. Bianchini, and F. Scarselli, “Coco_ts dataset: pixel–level annotations based on weak supervision for scene text segmentation,” in ICANN , 2019, pp. 238–250
2019
Earlier work this paper cites.
A. Kirillov, K. He, R. Girshick, C. Rother, and P. Dollár, “Panoptic segmentation,” in CVPR , 2019, pp. 9404–9413
2019
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR , 2019
2019
Earlier work this paper cites.
J. Liu, Z. Chen, B. Du, and D. Tao, “Asts: A unified framework for arbitrary shape text spotting,” IEEE TIP , vol. 29, pp. 5924–5936, 2020
2020
Earlier work this paper cites.
M. Liao, Z. Wan, C. Yao, K. Chen, and X. Bai, “Real-time scene text detection with differentiable binarization,” in AAAI , vol. 34, no. 07, 2020, pp. 11 474–11 481
2020
Earlier work this paper cites.
J. Ye, Z. Chen, J. Liu, and B. Du, “Textfusenet: Scene text detection with richer fused features.” in IJCAI , vol. 20, 2020, pp. 516–522
2020
Earlier work this paper cites.
Y. Liu, H. Chen, C. Shen, T. He, L. Jin, and L. Wang, “Abcnet: Real-time scene text spotting with adaptive bezier-curve network,” in CVPR , 2020, pp. 9809–9818
2020
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV , 2020, pp. 213–229
2020
Earlier work this paper cites.
X. Wang, R. Zhang, T. Kong, L. Li, and C. Shen, “Solov2: Dynamic and fast instance segmentation,” in NeurIPS , vol. 33, 2020, pp. 17 721–17 732
2020
Earlier work this paper cites.
J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y. Zhao, D. Liu, Y. Mu, M. Tan, X. Wang et al. , “Deep high-resolution representation learning for visual recognition,” IEEE TPAMI , vol. 43, no. 10, pp. 3349–3364, 2020
2020
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in NeurIPS , vol. 33, 2020, pp. 6840–6851
2020
Cited alongside, same era.
X. Xu, Z. Zhang, Z. Wang, B. Price, Z. Wang, and H. Shi, “Rethinking text segmentation: A novel dataset and a text-specific refinement approach,” in CVPR , 2021, pp. 12 045–12 055
2021
Cited alongside, same era.
S. Appalaraju, B. Jasani, B. U. Kota, Y. Xie, and R. Manmatha, “Docformer: End-to-end transformer for document understanding,” in ICCV , 2021, pp. 993–1003
2021
Cited alongside, same era.
A. Singh, G. Pang, M. Toh, J. Huang, W. Galuba, and T. Hassner, “Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text,” in CVPR , 2021, pp. 8802–8812
M. Ye, J. Zhang, S. Zhao, J. Liu, B. Du, and D. Tao, “Dptext-detr: Towards better scene text detection with dynamic points in transformer,” in AAAI , vol. 37, no. 3, 2023, pp. 3241–3249
2023
Later among the works it cites.
M. Ye, J. Zhang, S. Zhao, J. Liu, T. Liu, B. Du, and D. Tao, “Deepsolo: Let transformer decoder with explicit points solo for text spotting,” in CVPR , 2023, pp. 19 348–19 357
2023
Later among the works it cites.
J. Liu, H. Ding, Z. Cai, Y. Zhang, R. K. Satzoda, V. Mahadevan, and R. Manmatha, “Polyformer: Referring image segmentation as sequential polygon generation,” in CVPR , 2023, pp. 18 653–18 663
2023
Later among the works it cites.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo et al. , “Segment anything,” in ICCV , 2023, pp. 4015–4026
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
S.-X. Zhang, X. Zhu, C. Yang, H. Wang, and X.-C. Yin, “Adaptive boundary proposal network for arbitrary shape text detection,” in CVPR , 2021, pp. 1305–1314
2021
Cited alongside, same era.
P. Dai, S. Zhang, H. Zhang, and X. Cao, “Progressive contour regression for arbitrary-shape scene text detection,” in CVPR , 2021, pp. 7393–7402
2021
Cited alongside, same era.
Y. Zhu, J. Chen, L. Liang, Z. Kuang, L. Jin, and W. Zhang, “Fourier contour embedding for arbitrary-shaped text detection,” in CVPR , 2021, pp. 3123–3131
2021
Cited alongside, same era.
S. Biswas, P. Riba, J. Lladós, and U. Pal, “Beyond document object detection: instance-level segmentation of complex layouts,” IJDAR , vol. 24, no. 3, pp. 269–281, 2021
2021
Cited alongside, same era.
C. Li, B. Bi, M. Yan, W. Wang, S. Huang, F. Huang, and L. Si, “Structurallm: Structural pre-training for form understanding,” in ACL , 2021, pp. 6309–6318
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021, pp. 8748–8763
2021
Cited alongside, same era.
M. Ryoo, A. Piergiovanni, A. Arnab, M. Dehghani, and A. Angelova, “Tokenlearner: Adaptive space-time tokenization for videos,” in NeurIPS , vol. 34, 2021, pp. 12 786–12 797
2021
Cited alongside, same era.
L. Ke, M. Ye, M. Danelljan, Y. Liu, Y.-W. Tai, C.-K. Tang, and F. Yu, “Segment anything in high quality,” in NeurIPS , vol. 36, 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
W. Yue, J. Zhang, K. Hu, Y. Xia, J. Luo, and Z. Wang, “Surgicalsam: Efficient class promptable surgical instrument segmentation,” in AAAI , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
D. Wang, J. Zhang, B. Du, M. Xu, L. Liu, D. Tao, and L. Zhang, “SAMRS: Scaling-up remote sensing segmentation dataset with segment anything model,” in NeurIPS Datasets and Benchmarks Track , 2023. [Online]. Available: https://openreview.net/forum?id=jHrgq55ftl
2023
Later among the works it cites.
X. Wang, C. Wu, H. Yu, B. Li, and X. Xue, “Textformer: Component-aware text segmentation with transformer,” in ICME , 2023, pp. 1877–1882
2023
Later among the works it cites.
D. Coquenet, C. Chatelain, and T. Paquet, “Dan: a segmentation-free document attention network for handwritten document recognition,” IEEE TPAMI , vol. 45, no. 7, pp. 8227–8243, 2023
2023
Later among the works it cites.
K. Lee, M. Joshi, I. R. Turc, H. Hu, F. Liu, J. M. Eisenschlos, U. Khandelwal, P. Shaw, M.-W. Chang, and K. Toutanova, “Pix2struct: Screenshot parsing as pretraining for visual language understanding,” in ICML , 2023, pp. 18 893–18 912
2023
Later among the works it cites.
W. Yu, Y. Liu, W. Hua, D. Jiang, B. Ren, and X. Bai, “Turning a clip model into a scene text detector,” in CVPR , 2023, pp. 6978–6988
2023
Later among the works it cites.
Z. Wang, H. Xie, Y. Wang, J. Xu, B. Zhang, and Y. Zhang, “Symmetrical linguistic feature distillation with clip for scene text recognition,” in ACM MM , 2023, pp. 509–518
2023
Later among the works it cites.
J. Cen, Z. Zhou, J. Fang, C. Yang, W. Shen, L. Xie, X. Zhang, and Q. Tian, “Segment anything in 3d with nerfs,” in NeurIPS , vol. 36, 2023, pp. 25 971–25 990
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
X. Chen, X. Wang, J. Zhou, Y. Qiao, and C. Dong, “Activating more pixels in image super-resolution transformer,” in CVPR , 2023, pp. 22 367–22 377
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
D. Peng, Z. Yang, J. Zhang, C. Liu, Y. Shi, K. Ding, F. Guo, and L. Jin, “Upocr: Towards unified pixel-level ocr interface,” in ICML , 2024
2024
Closest in time.
S. Long, S. Qin, Y. Fujii, A. Bissacco, and M. Raptis, “Hierarchical text spotter for joint text spotting and layout analysis,” in WACV , 2024, pp. 903–913
2024
Closest in time.
F. Li, H. Zhang, P. Sun, X. Zou, S. Liu, J. Yang, C. Li, L. Zhang, and J. Gao, “Semantic-sam: Segment and recognize anything at any granularity,” in ECCV , 2024
2024
Closest in time.
W. Yu, Y. Liu, X. Zhu, H. Cao, X. Sun, and X. Bai, “Turning a clip model into a scene text spotter,” IEEE TPAMI , 2024
2024
Closest in time.
J. Li, J. Jain, and H. Shi, “Matting anything,” in CVPR , 2024, pp. 1775–1785
2024
Closest in time.
R. Zhang, Z. Jiang, Z. Guo, S. Yan, J. Pan, H. Dong, P. Gao, and H. Li, “Personalize segment anything model with one shot,” in ICLR , 2024
2024
Closest in time.
Y. Liu, M. Zhu, H. Li, H. Chen, X. Wang, and C. Shen, “Matcher: Segment anything with one shot using all-purpose feature matching,” in ICLR , 2024
2024
Closest in time.