Fetching the paper…
Reading the bibliography…
Multi-modal document pre-trained models have proven to be very effective in a variety of visually-rich document understanding (VrDU) tasks.
D. Lewis, G. Agam, and S. Argamon, “Building a test collection for complex document information processing,” in ACM SIGIR , 2006, pp. 665–666
2006
Earlier work this paper cites.
Y. Yang, Y. Zhuang, “Harmonizing hierarchical manifolds for multimedia document semantics understanding and cross-media retrieval,” IEEE Trans. Multim , vol. 10, pp. 437–446, 2008
2008
Earlier work this paper cites.
J. Bian, Y. Yang, “Multimedia Summarization for Social Events in Microblog Stream,” IEEE Trans. on Multimedia , vol. 17, pp. 216–228, 2014
2014
Earlier work this paper cites.
S. Ren, K. He, and R. Girshick, “Faster R-CNN: Towards real-time object detection with region proposal networks,” Advances in neural information processing systems , vol. 28, pp. 91–99, 2015
2015
Earlier work this paper cites.
A. W. Harley, A. Ufkes, and K. G. Derpanis, “Evaluation of deep convolutional nets for document image classification and retrieval,” in ICDAR , 2015, pp. 991–995
2015
Earlier work this paper cites.
Q. Fang, C. Xu, “Word-of-mouth understanding: Entity-centric multimodal aspect-opinion mining in social media,” IEEE Trans. Multim , vol. 17, pp. 2281–2296, 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in ICLR , 2015
2015
Earlier work this paper cites.
E. Baralis, L. Cagliero, “Learning from summaries: Supporting e-learning activities by means of document summarization,” IEEE Trans. Vis. Comput. Graph , vol. 4, pp. 416–428, 2016
2016
Earlier work this paper cites.
Y. Wu, M. Schuster, Z. Chen, and Q. V. Le, “Google’s neural machine translation system: Bridging the gap between human and machine translation,” arXiv , pp. arXiv–1609, 2016
2016
Earlier work this paper cites.
X. Yang, E. Yumer, and P. Asente, “Learning to extract semantic structure from documents using multimodal fully convolutional neural networks,” in CVPR , 2017, pp. 5315–5324
2017
Earlier work this paper cites.
M. Z. Afzal, A. Kölsch, and Ahmed, “Cutting the error by half: Investigation of very deep cnn and advanced training strategies for document image classification,” in ICDAR , vol. 1, 2017, pp. 883–888
2017
Earlier work this paper cites.
C. Szegedy, S. Ioffe, and V. Vanhoucke, “Inception-v4, inception-resnet and the impact of residual connections on learning,” in AAAI , 2017
2017
Earlier work this paper cites.
C. Luo, W. Li, and Q. Chen, “A part-of-speech enhanced neural conversation model,” in ECIR , 2017, pp. 173–185
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. Girshick, “Mask R-CNN,” in ICCV , 2017, pp. 2961–2969
2017
Earlier work this paper cites.
A. R. Katti, C. Reisswig, C. Guder, and S. Brarda, “Chargrid: Towards understanding 2d documents,” in EMNLP , 2018, pp. 4459–4469
2018
Earlier work this paper cites.
A. Das, S. Roy, U. Bhattacharya, and S. K. Parui, “Document image classification with intra-domain transfer learning and stacked generalization of deep convolutional neural networks,” in ICPR . IEEE, 2018, pp. 3180–3185
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, and K. Lee, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT , 2019, pp. 4171–4186
2019
Earlier work this paper cites.
C. Soto and S. Yoo, “Visual detection with context for document layout analysis,” in EMNLP-IJCNLP , 2019, pp. 3464–3470
2019
Earlier work this paper cites.
T. I. Denk and C. Reisswig, “BERTgrid: Contextualized embedding for 2d document representation and understanding,” in NIPS , 2019
2019
Earlier work this paper cites.
X. Liu, F. Gao, and Q. Zhang, “Graph convolution for multimodal information extraction from visually rich documents,” in NAACL-HLT , 2019, pp. 32–39
2019
Earlier work this paper cites.
Y. Qian, E. Santus, and Z. Jin, “GraphIE: A graph-based framework for information extraction,” in NAACL-HLT , 2019, pp. 751–761
2019
Cited alongside, same era.
R. Jain and C. Wigington, “Multimodal document image classification,” in ICDAR , 2019, pp. 71–77
2019
Cited alongside, same era.
Y. Liu, M. Ott, N. Goyal, and J. Du, “RoBERTa: A robustly optimized bert pretraining approach,” in arXiv , 2019
2019
Cited alongside, same era.
S. Badam, Z. Liu, and N. Elmqvist, “Elastic documents: Coupling text and tables through contextual visualizations for enhanced document reading,” IEEE Trans. Vis. Comput. Graph , vol. 25, pp. 661–671, 2019
2019
Cited alongside, same era.
G. Jaume, H. K. Ekenel, and J.-P. Thiran, “FUNSD: A dataset for form understanding in noisy scanned documents,” in ICDARW , vol. 2, 2019, pp. 1–6
2019
Cited alongside, same era.
W. Hwang, J. Yim, and S. Park, “Spatial dependency parsing for semi-structured document information extraction,” arXiv , 2020
2020
Later among the works it cites.
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai, “VL-BERT: Pre-training of generic visual-linguistic representations,” in ICLR , 2020
2020
Later among the works it cites.
Y.-C. Chen, L. Li, L. Yu, A. El Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, “UNITER: Learning universal image-text representations,” in ECCV , 2020
2020
Later among the works it cites.
X. Li, X. Yin, C. Li, P. Zhang, X. Hu, L. Zhang, L. Wang, H. Hu, L. Dong, F. Wei et al. , “Oscar: Object-semantics aligned pre-training for vision-language tasks,” in ECCV , 2020, pp. 121–137
2020
Later among the works it cites.
C. Li, B. Bi, and M. Yan, “StructuralLM: Structural pre-training for form understanding,” in ACL , 2021, pp. 6309-6318
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Park, S. Shin, B. Lee, and J. Lee, “CORD: A consolidated receipt dataset for post-ocr parsing,” in NIPS , 2019
2019
Cited alongside, same era.
R. Sarkhel and A. Nandi, “Deterministic routing between layout abstractions for multi-scale classification of visually rich documents,” in IJCAI , 2019, pp. 3360–3366
2019
Cited alongside, same era.
T. Dauphinee, N. Patel, and M. Rashidi, “Modular multimodal architecture for document classification,” In arXiv , 2019
2019
Cited alongside, same era.
M. L. Villegas, W. J. Giraldo y C. Collazos, Modeling Interactive Systems at the Business Level: Inter-Action Diagram, IEEE Latin America Transactions, vol. 17, no 03, pp. 462-472, 2019
2019
Cited alongside, same era.
X. Zhong, J. Tang, and A. J. Yepes, “PubLayNet: largest dataset ever for document layout analysis,” in ICDAR , 2019, pp. 1015–1022
2019
Cited alongside, same era.
J. Lu, D. Batra, D. Parikh, and S. Lee, “ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in NeurIPS , 2019, pp. 13–23
2019
Cited alongside, same era.
Y. Xu, M. Li, L. Cui, and S. Huang, “LayoutLM: Pre-training of text and layout for document image understanding,” in KDD , 2020, pp. 1192–1200
2020
Cited alongside, same era.
2021
Later among the works it cites.
Y. Xu, Y. Xu, T. Lv, and L. Cui, “LayoutLM v2: Multi-modal pre-training for visually-rich document understanding,” in ACL , 2021, pp. 2579-2591
2021
Later among the works it cites.
P. Li, J. Gu, and J. Kuen, “SelfDoc: Self-supervised document representation learning,” in CVPR , 2021, pp. 5652–5660
2021
Later among the works it cites.
S. Appalaraju, B. Jasani, and B. U. Kota, “DocFormer: End-to-end transformer for document understanding,” in ICCV , 2021, pp. 4171–4186
2021
Later among the works it cites.
T. A. N. Dang, D. T. Hoang, and Q. B. Tran, “End-to-end hierarchical relation extraction for generic form understanding,” in ICPR , 2021, pp. 5238–5245
2021
Later among the works it cites.
G. Tang, L. Xie, L. Jin, and Wang, “MatchVIE: Exploiting match relevancy between entities for visual information extraction,” in IJCAI , 2021, pp. 1039–1045
2021
Later among the works it cites.
S. Frank, E. Bugliarello, and D. ElliottC, “Vision-and-language or vision-for-language? on cross-modal influence in multimodal transformers,” in EMNLP , 2021, pp. 9847–9857
2021
Later among the works it cites.
Y. Lai, Y. Liu, and Y. Feng, “Lattice-BERT: Leveraging multi-granularity representations in chinese pre-trained language models,” in NAACL , 2021, pp. 1716–1731
2021
Later among the works it cites.
M. Mathew, D. Karatzas, and C. Jawahar, “DocVQA: A dataset for vqa on document images,” in WACV , 2021, pp. 2200–2209
2021
Later among the works it cites.
W. Li, C. Gao, G. Niu, X. Xiao, H. Liu, J. Liu, H. Wu, and H. Wang, “UNIMO: Towards unified-modal understanding and generation via cross-modal contrastive learning,” in ACL , 2021, pp. 2592–2607
2021
Later among the works it cites.
C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in ICML , 2021, pp. 4904–4916
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021, pp. 8748-8763
2021
Later among the works it cites.
W. Yu, N. Lu, X. Qi, P. Gong, and R. Xiao, “PICK: Processing key information extraction from documents using improved graph learning-convolutional networks,” in ICPR , 2021, pp. 4363–4370
2021
Later among the works it cites.
W. Lin, Q. Gao, L. Sun, Z. Zhong, K. Hu, Q. Ren, and Q. Huo, “ViBERTgrid: A jointly trained multi-modal 2d document representation for key information extraction from documents,” in ICDAR , 2021, pp. 548–563
2021
Later among the works it cites.
Powalski, R., Borchmann, Ł., Jurkiewicz, D., Dwojak, T., Pietruszka, M., Pałka, G., “Going full-tilt boogie on document understanding with text-image-layout transformer,” in ICDAR , 2021, pp. 732–747
2021
Later among the works it cites.