Fetching the paper…
Reading the bibliography…
Information extraction from semi-structured documents is crucial for frictionless business-to-business (B2B) communication.
1903
Earlier work this paper cites.
Berge, J.: The EDIFACT standards. Blackwell Publishers, Inc. (1994)
1994
Earlier work this paper cites.
Yi, J., Sundaresan, N.: A classifier for semi-structured documents. In: Proceedings of the sixth ACM SIGKDD international conference on Knowledge discovery and data mining. pp. 340–344 (2000)
2000
Earlier work this paper cites.
Cesarini, F., Francesconi, E., Gori, M., Soda, G.: Analysis and understanding of multi-class invoices. Document Analysis and Recognition 6
2003
Earlier work this paper cites.
Ford, G., Thoma, G.R.: Ground truth data for document image analysis. In: Symposium on Document Image Understanding and Technology. pp. 199–205. Citeseer (2003)
2003
Earlier work this paper cites.
Meadows, B., Seaburg, L.: Universal business language 1.0. Organization for the Advancement of Structured Information Standards (OASIS) (2004)
2004
Earlier work this paper cites.
Bosak, J., McGrath, T., Holman, G.K.: Universal business language v2. 0. Organization for the Advancement of Structured Information Standards (OASIS), Standard (2006)
2006
Earlier work this paper cites.
Lewis, D., Agam, G., Argamon, S., Frieder, O., Grossman, D., Heard, J.: Building a test collection for complex document information processing. In: Proceedings of the 29th annual international ACM SIGIR conference on Research and development in information retrieval. pp. 665–666 (2006)
2006
Earlier work this paper cites.
Hamza, H., Belaïd, Y., Belaïd, A.: Case-based reasoning for invoice analysis and recognition. In: International conference on case-based reasoning. pp. 404–418. Springer (2007)
2007
Earlier work this paper cites.
Nadeau, D., Sekine, S.: A survey of named entity recognition and classification. Lingvisticæ Investigationes pp. 3–26 (2007). https://doi.org/https://doi.org/10.1075/li.30.1.03nad
2007
Earlier work this paper cites.
Smith, R.: An overview of the tesseract ocr engine. In: Ninth international conference on document analysis and recognition (ICDAR 2007). vol. 2, pp. 629–633. IEEE (2007)
2007
Earlier work this paper cites.
Antonacopoulos, A., Bridson, D., Papadopoulos, C., Pletschacher, S.: A realistic dataset for performance evaluation of document layout analysis. In: Proceedings of ICDAR. pp. 296–300. IEEE (2009)
2009
Earlier work this paper cites.
Shahab, A., Shafait, F., Kieninger, T., Dengel, A.: An open approach towards the benchmarking of table structure recognition systems. In: Doermann, D.S., Govindaraju, V., Lopresti, D.P., Natarajan, P. (eds.) The Ninth IAPR International Workshop on Document Analysis Systems, DAS. pp. 113–120 (2010). https://doi.org/10.1145/1815330.1815345
2010
Earlier work this paper cites.
Medvet, E., Bartoli, A., Davanzo, G.: A probabilistic approach to printed document understanding. Int. J. Document Anal. Recognit. 14
2011
Earlier work this paper cites.
Fang, J., Tao, X., Tang, Z., Qiu, R., Liu, Y.: Dataset, ground-truth and performance metrics for table detection evaluation. In: Blumenstein, M., Pal, U., Uchida, S. (eds.) Proceedings of IAPR International Workshop on Document Analysis Systems, DAS. pp. 445–449. IEEE (2012). https://doi.org/10.1109/DAS.2012.29
2012
Earlier work this paper cites.
Jiang, J.: Information extraction from text. In: Mining text data, pp. 11–41. Springer (2012)
2012
Earlier work this paper cites.
Göbel, M.C., Hassan, T., Oro, E., Orsi, G.: ICDAR 2013 table competition. In: Proceedings of ICDAR. pp. 1449–1453. IEEE Computer Society (2013). https://doi.org/10.1109/ICDAR.2013.292
2013
Earlier work this paper cites.
Rusinol, M., Benkhelfallah, T., Poulain dAndecy, V.: Field extraction from administrative documents by incremental structural templates. In: 2013 12th International Conference on Document Analysis and Recognition. pp. 1100–1104. IEEE (2013)
2013
Earlier work this paper cites.
Schuster, D., Muthmann, K., Esser, D., Schill, A., Berger, M., Weidling, C., Aliyev, K., Hofmeier, A.: Intellix–end-user trained information extraction for document archiving. In: 2013 12th International Conference on Document Analysis and Recognition. pp. 101–105. IEEE (2013)
2013
Earlier work this paper cites.
Directive 2014/55/eu of the european parliament and of the council on electronic invoicing in public procurement (Apr 2014), https://eur-lex.europa.eu/eli/dir/2014/55/oj
2014
Earlier work this paper cites.
Harley, A.W., Ufkes, A., Derpanis, K.G.: Evaluation of deep convolutional nets for document image classification and retrieval. In: International Conference on Document Analysis and Recognition (ICDAR) (2015)
2015
Earlier work this paper cites.
Stockerl, M., Ringlstetter, C., Schubert, M., Ntoutsi, E., Kriegel, H.P.: Online template matching over a stream of digitized documents. In: Proceedings of the 27th International Conference on Scientific and Statistical Database Management. pp. 1–12 (2015)
2015
Earlier work this paper cites.
Hamad, K.A., Mehmet, K.: A detailed analysis of optical character recognition technology. International Journal of Applied Mathematics Electronics and Computers 1
2016
Earlier work this paper cites.
Kumar, A., Irsoy, O., Ondruska, P., Iyyer, M., Bradbury, J., Gulrajani, I., Zhong, V., Paulus, R., Socher, R.: Ask me anything: Dynamic memory networks for natural language processing. In: Balcan, M., Weinberger, K.Q. (eds.) Proceedings of ICML. vol. 48, pp. 1378–1387. JMLR.org (2016)
2016
Earlier work this paper cites.
Lample, G., Ballesteros, M., Subramanian, S., Kawakami, K., Dyer, C.: Neural architectures for named entity recognition. In: Knight, K., Nenkova, A., Rambow, O. (eds.) Proceedings of NAACL HLT. pp. 260–270 (2016). https://doi.org/10.18653/v1/n16-1030
2016
Earlier work this paper cites.
Liu, W., Zhang, Y., Wan, B.: Unstructured document recognition on business invoice. Mach. Learn., Stanford iTunes Univ., Stanford, CA, USA, Tech. Rep (2016)
2016
Earlier work this paper cites.
Gao, L., Yi, X., Jiang, Z., Hao, L., Tang, Z.: ICDAR2017 competition on page object detection. In: Proceedings of ICDAR. pp. 1417–1422 (2017). https://doi.org/10.1109/ICDAR.2017.231
2017
Earlier work this paper cites.
He, S., Schomaker, L.: Beyond ocr: Multi-faceted understanding of handwritten document characteristics. Pattern Recognition 63
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
Palm, R.B., Winther, O., Laws, F.: Cloudscan - A configuration-free invoice analysis system using recurrent neural networks. In: Proceedings of ICDAR. pp. 406–413. IEEE (2017). https://doi.org/10.1109/ICDAR.2017.74
2017
Earlier work this paper cites.
Schreiber, S., Agne, S., Wolf, I., Dengel, A., Ahmed, S.: DeepDeSRT: Deep Learning for Detection and Structure Recognition of Tables in Document Images. In: Proceedings of ICDAR. pp. 1162–1167 (2017). https://doi.org/10.1109/ICDAR.2017.192
2017
Earlier work this paper cites.
Cho, M., Amplayo, R.K., Hwang, S., Park, J.: Adversarial TableQA: Attention Supervision for Question Answering on Tables. In: Zhu, J., Takeuchi, I. (eds.) Proceedings of ACML. Proceedings of Machine Learning Research, vol. 95, pp. 391–406 (2018)
2018
Earlier work this paper cites.
Cristani, M., Bertolaso, A., Scannapieco, S., Tomazzoli, C.: Future paradigms of automated processing of business documents. International Journal of Information Management 40
2018
Earlier work this paper cites.
d’Andecy, V.P., Hartmann, E., Rusinol, M.: Field extraction by hybrid incremental and a-priori structural templates. In: 2018 13th IAPR International Workshop on Document Analysis Systems (DAS). pp. 251–256. IEEE (2018)
2018
Earlier work this paper cites.
Holt, X., Chisholm, A.: Extracting structured data from invoices. In: Proceedings of the Australasian Language Technology Association Workshop 2018. pp. 53–59 (2018)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
McCann, B., Keskar, N.S., Xiong, C., Socher, R.: The natural language decathlon: Multitask learning as question answering. CoRR (2018)
2018
Cited alongside, same era.
Siegel, N., Lourie, N., Power, R., Ammar, W.: Extracting scientific figures with distantly supervised neural networks. In: Chen, J., Gonçalves, M.A., Allen, J.M., Fox, E.A., Kan, M., Petras, V. (eds.) Proceedings of the 18th ACM/IEEE on Joint Conference on Digital Libraries, JCDL. pp. 223–232 (2018). https://doi.org/10.1145/3197026.3197040
2018
Cited alongside, same era.
Xu, Y., Li, M., Cui, L., Huang, S., Wei, F., Zhou, M.: Layoutlm: Pre-training of text and layout for document image understanding. In: Gupta, R., Liu, Y., Tang, J., Prakash, B.A. (eds.) Proceedings on KDD. pp. 1192–1200 (2020). https://doi.org/10.1145/3394486.3403172
2020
Later among the works it cites.
Yu, D., Li, X., Zhang, C., Liu, T., Han, J., Liu, J., Ding, E.: Towards accurate scene text recognition with semantic reasoning networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 12113–12122 (2020)
2020
Later among the works it cites.
Zhong, X., ShafieiBavani, E., Jimeno-Yepes, A.: Image-based table recognition: Data, model, and evaluation. In: Vedaldi, A., Bischof, H., Brox, T., Frahm, J. (eds.) Proceedings of ECCV. pp. 564–580. Springer (2020). https://doi.org/10.1007/978-3-030-58589-1_34
2020
Later among the works it cites.
Baviskar, D., Ahirrao, S., Kotecha, K.: Multi-layout invoice document dataset (MIDD): A dataset for named entity recognition. Data (2021). https://doi.org/10.3390/data6070078
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Baek, Y., Lee, B., Han, D., Yun, S., Lee, H.: Character region awareness for text detection. In: Proceedings of the IEEE/CVF CVPR. pp. 9365–9374 (2019)
2019
Cited alongside, same era.
Clausner, C., Antonacopoulos, A., Pletschacher, S.: ICDAR 2019 competition on recognition of documents with complex layouts-rdcl2019. In: Proceedings of ICDAR. pp. 1521–1526. IEEE (2019)
2019
Cited alongside, same era.
Deng, Y., Rosenberg, D.S., Mann, G.: Challenges in end-to-end neural scientific table recognition. In: Proceedings of ICDAR. pp. 894–901. IEEE (2019). https://doi.org/10.1109/ICDAR.2019.00148
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Dhakal, P., Munikar, M., Dahal, B.: One-shot template matching for automatic document data capture. In: Proceeedings of Artificial Intelligence for Transforming Business and Society (AITB). vol. 1, pp. 1–6. IEEE (2019)
2019
Cited alongside, same era.
Guillaume Jaume, Hazim Kemal Ekenel, J.P.T.: Funsd: A dataset for form understanding in noisy scanned documents. In: Accepted to ICDAR-OST (2019)
2019
Cited alongside, same era.
Holeček, M., Hoskovec, A., Baudiš, P., Klinger, P.: Table understanding in structured documents. In: 2019 International Conference on Document Analysis and Recognition Workshops (ICDARW). vol. 5, pp. 158–164. IEEE (2019)
2019
Cited alongside, same era.
Huang, Z., Chen, K., He, J., Bai, X., Karatzas, D., Lu, S., Jawahar, C.V.: ICDAR2019 competition on scanned receipt OCR and information extraction. In: Proceedings of ICDAR. pp. 1516–1520. IEEE (2019). https://doi.org/10.1109/ICDAR.2019.00244
2019
Cited alongside, same era.
2021
Later among the works it cites.
Bensch, O., Popa, M., Spille, C.: Key information extraction from documents: Evaluation and generator. In: Abbès, S.B., Hantach, R., Calvez, P., Buscaldi, D., Dessì, D., Dragoni, M., Recupero, D.R., Sack, H. (eds.) Proceedings of DeepOntoNLP and X-SENTIMENT. CEUR Workshop Proceedings, vol. 2918, pp. 47–53. CEUR-WS.org (2021)
2021
Later among the works it cites.
Borchmann, Ł., Pietruszka, M., Stanislawek, T., Jurkiewicz, D., Turski, M., Szyndler, K., Graliński, F.: DUE: End-to-end document understanding benchmark. In: Proceeedings of NeurIPS (2021)
2021
Later among the works it cites.
Chen, L., Chen, X., Zhao, Z., Zhang, D., Ji, J., Luo, A., Xiong, Y., Yu, K.: Websrc: A dataset for web-based structural reading comprehension. CoRR (2021)
2021
Later among the works it cites.
Chen, W., Chang, M., Schlinger, E., Wang, W.Y., Cohen, W.W.: Open question answering over tables and text. In: Proceedings of ICLR (2021)
2021
Later among the works it cites.
Garncarek, L., Powalski, R., Stanislawek, T., Topolski, B., Halama, P., Turski, M., Gralinski, F.: LAMBERT: layout-aware language modeling for information extraction. In: Lladós, J., Lopresti, D., Uchida, S. (eds.) Proceedings of ICDAR. vol. 12821, pp. 532–547. Springer (2021). https://doi.org/10.1007/978-3-030-86549-8_34
2021
Later among the works it cites.
Holeček, M.: Learning from similarity and information extraction from structured documents. International Journal on Document Analysis and Recognition (IJDAR) pp. 1–17 (2021)
2021
Later among the works it cites.
Krieger, F., Drews, P., Funk, B., Wobbe, T.: Information extraction from invoices: A graph neural network approach for datasets with high layout variety. In: International Conference on Wirtschaftsinformatik. pp. 5–20. Springer (2021)
2021
Later among the works it cites.
Mathew, M., Karatzas, D., Jawahar, C.V.: Docvqa: A dataset for VQA on document images. In: Proceedings of WACV. pp. 2199–2208. IEEE (2021). https://doi.org/10.1109/WACV48630.2021.00225
2021
Later among the works it cites.
Stanisławek, T., Graliński, F., Wróblewska, A., Lipiński, D., Kaliska, A., Rosalska, P., Topolski, B., Biecek, P.: Kleister: key information extraction datasets involving long documents with complex layouts. In: International Conference on Document Analysis and Recognition. pp. 564–579. Springer (2021)
2021
Later among the works it cites.
2021
Later among the works it cites.
Wang, J., Liu, C., Jin, L., Tang, G., Zhang, J., Zhang, S., Wang, Q., Wu, Y., Cai, M.: Towards robust visual information extraction in real world: New dataset and novel solution. In: Proceedings of the AAAI Conference on Artificial Intelligence (2021)
2021
Later among the works it cites.
Xu, Y., Lv, T., Cui, L., Wang, G., Lu, Y., Florêncio, D., Zhang, C., Wei, F.: LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding. CoRR (2021)
2021
Later among the works it cites.
Yu, W., Lu, N., Qi, X., Gong, P., Xiao, R.: PICK: processing key information extraction from documents using improved graph learning-convolutional networks. In: Proceedings of ICPR. pp. 4363–4370. IEEE (2020). https://doi.org/10.1109/ICPR48806.2021.9412927
2021
Later among the works it cites.
Zheng, X., Burdick, D., Popa, L., Zhong, X., Wang, N.X.R.: Global table extractor (GTE): A framework for joint table identification and cell structure recognition using visual context. In: Proceedings of WACV. pp. 697–706. IEEE (2021). https://doi.org/10.1109/WACV48630.2021.00074
2021
Later among the works it cites.
Zhu, F., Lei, W., Huang, Y., Wang, C., Zhang, S., Lv, J., Feng, F., Chua, T.: TAT-QA: A question answering benchmark on a hybrid of tabular and textual content in finance. In: Zong, C., Xia, F., Li, W., Navigli, R. (eds.) Proceedings International Joint Conference on Natural Language Processing. pp. 3277–3287 (2021). https://doi.org/10.18653/v1/2021.acl-long.254
2021
Later among the works it cites.
Mathew, M., Bagal, V., Tito, R., Karatzas, D., Valveny, E., Jawahar, C.: Infographicvqa. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. pp. 1697–1706 (2022)
2022
Closest in time.
2022
Closest in time.
Web: Annual reports. https://www.annualreports.com/ , accessed: 2022-04-28
2022
Closest in time.
Web: Charity Commission for England and Wales. https://apps.charitycommission.gov.uk/showcharity/registerofcharities/RegisterHomePage.aspx , accessed: 2022-04-22
2022
Closest in time.
Web: EDGAR. https://www.sec.gov/edgar.shtml , accessed: 2022-04-22
2022
Closest in time.
Web: Industry Documents Library. https://www.industrydocuments.ucsf.edu/ , accessed: 2022-04-22
2022
Closest in time.
Web: NIST Special Database 2. https://www.nist.gov/srd/nist-special-database-2 , accessed: 2022-04-25
2022
Closest in time.
Web: Open Government Data (OGD) Platform India. https://visualize.data.gov.in/ , accessed: 2022-04-22
2022
Closest in time.
Web: Public Inspection Files. https://publicfiles.fcc.gov/ , accessed: 2022-04-22
2022
Closest in time.
Web: Scitsr. https://github.com/Academic-Hammer/SciTSR , accessed: 2022-04-26
2022
Closest in time.
Web: S&P 500 Companies with Financial Information. https://www.spglobal.com/spdji/en/indices/equity/sp-500/#data , accessed: 2022-04-25
2022
Closest in time.
Web: Statistics of Common Crawl Monthly Archives — MIME Types. https://commoncrawl.github.io/cc-crawl-statistics/plots/mimetypes , accessed: 2022-04-22
2022
Closest in time.
Web: Tablebank. https://github.com/doc-analysis/TableBank , accessed: 2022-04-26
2022
Closest in time.
Web: World Bank Open Data. https://data.worldbank.org/ , accessed: 2022-04-22
2022
Closest in time.