Fetching the paper…
Reading the bibliography…
A large amount of document data exists in unstructured form such as raw images without any text information.
LayoutLMv2: Multi-modal pre-training for visually-rich document understanding
Xu, Y.; Xu, Y.; Lv, T.; Cui, L.; Wei, F.; Wang, G.; Lu, Y.; Florencio, D.; Zhang, C.; Che, W.; et al. 2020b · 2012
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I.; Vinyals, O.; and Le, Q. V. 2014 · 2014
Earlier work this paper cites.
Robust Scene Text Recognition with Automatic Rectification
Shi, B.; Wang, X.; Lyu, P.; Cong, Y.; and Xiang, B. 2016 · 2016
Earlier work this paper cites.
Feature pyramid networks for object detection
Lin, T.-Y.; Dollár, P.; Girshick, R.; He, K.; Hariharan, B.; and Belongie, S. 2017 · 2017
Earlier work this paper cites.
Shufflenet v2: Practical guidelines for efficient cnn architecture design
Ma, N.; Zhang, X.; Zheng, H.-T.; and Sun, J. 2018 · 2018
Earlier work this paper cites.
Searching for mobilenetv3
Howard, A.; Sandler, M.; Chu, G.; Chen, L.-C.; Chen, B.; Tan, M.; Wang, W.; Zhu, Y.; Pang, R.; Vasudevan, V.; et al. 2019 · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
Jaume, G.; Ekenel, H. K.; and Thiran, J.-P. 2019 · 2019
Earlier work this paper cites.
PaddlePaddle: An open-source deep learning platform from industrial practice
Ma, Y.; Yu, D.; Wu, T.; and Wang, H. 2019 · 2019
Earlier work this paper cites.
Publaynet: largest dataset ever for document layout analysis
Zhong, X.; Tang, J.; and Yepes, A. J. 2019 · 2019
Cited alongside, same era.
Ghostnet: More features from cheap operations
Han, K.; Wang, Y.; Tian, Q.; Guo, J.; Xu, C.; and Xu, C. 2020 · 2020
Cited alongside, same era.
Table structure recognition using top-down and bottom-up cues
Raja, S.; Mondal, A.; and Jawahar, C. 2020 · 2020
Cited alongside, same era.
CDLA: A Chinese document layout analysis (CDLA) dataset
Hang, L. 2021 · 2021
Cited alongside, same era.
PP-YOLOv2: A practical object detector
Huang, X.; Wang, X.; Lv, W.; Bai, X.; Long, X.; Deng, K.; Dang, Q.; Han, S.; Liu, Q.; Hu, X.; et al. 2021 · 2021
Cited alongside, same era.
Lgpma: Complicated table structure recognition with local and global pyramid mask alignment
Ye, J.; Qi, X.; He, Y.; Chen, Y.; Gu, D.; Gao, P.; and Xiao, R. 2021 · 2021
Later among the works it cites.
VSR: a unified framework for document layout analysis combining vision, semantics and relations
Zhang, P.; Li, C.; Qiao, L.; Cheng, Z.; Pu, S.; Niu, Y.; and Wu, F. 2021 · 2021
Later among the works it cites.
Global table extractor (gte): A framework for joint table identification and cell structure recognition using visual context
Zheng, X.; Burdick, D.; Popa, L.; Zhong, X.; and Wang, N. X. R. 2021 · 2021
Later among the works it cites.
PULC_text_image_orientation
Cui, C. 2022 · 2022
Closest in time.
Xylayoutlm: Towards layout-aware multimodal networks for visually-rich document understanding
Gu, Z.; Meng, C.; Wang, K.; Lan, J.; Wang, W.; Gu, M.; and Zhang, L. 2022 · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qiao, L.; Li, Z.; Cheng, Z.; Zhang, P.; Pu, S.; Niu, Y.; Ren, W.; Tan, W.; and Wu, F. 2021 · 2021
Cited alongside, same era.
LayoutParser: A Unified Toolkit for Deep Learning Based Document Image Analysis
Shen, Z.; Zhang, R.; Dell, M.; Lee, B. C. G.; Carlson, J.; and Li, W. 2021 · 2021
Cited alongside, same era.
Layoutxlm: Multimodal pre-training for multilingual visually-rich document understanding
Xu, Y.; Lv, T.; Cui, L.; Wang, G.; Lu, Y.; Florencio, D.; Zhang, C.; and Wei, F. 2021 · 2021
Cited alongside, same era.
PP-LCNet: A Lightweight CPU Convolutional Neural Network
Cui, C.; Gao, T.; Wei, S.; Du, Y.; Guo, R.; Dong, S.; Lu, B.; Zhou, Y.; Lv, X.; Liu, Q.; et al. 2021a
Cited in the paper.
Cui, C.; Guo, R.; Du, Y.; He, D.; Li, F.; Wu, Z.; Liu, Q.; Wen, S.; Huang, J.; Hu, X.; et al. 2021b
Cited in the paper.
PP-OCRv2: bag of tricks for ultra lightweight OCR system
Du, Y.; Li, C.; Guo, R.; Cui, C.; Liu, W.; Zhou, J.; Lu, B.; Yang, Y.; Liu, Q.; Hu, X.; et al. 2021a
Cited in the paper.
TableRec-RARE
Du, Y.; Li, C.; Zhou, J.; et al. 2021b
Cited in the paper.
XFUND: A Benchmark Dataset for Multilingual Visually Rich Form Understanding
Xu, Y.; Lv, T.; Cui, L.; Wang, G.; Lu, Y.; Florencio, D.; Zhang, C.; and Wei, F. 2022 · 2022
Closest in time.
Focal and global knowledge distillation for detectors
Yang, Z.; Li, Z.; Jiang, X.; Gong, Y.; Yuan, Z.; Zhao, D.; and Yuan, C. 2022 · 2022
Closest in time.