Fetching the paper…
Reading the bibliography…
We introduce a simple new approach to the problem of understanding documents where non-trivial layout influences the local semantics.
Ishitani, Y.: Model-based information extraction method tolerant of ocr errors for document images. Int. J. Comput. Process. Orient. Lang. 15
2002
Earlier work this paper cites.
2002
Earlier work this paper cites.
Cesarini, F., Francesconi, E., Gori, M., Soda, G.: Analysis and understanding of multi-class invoices. IJDAR 6
2003
Earlier work this paper cites.
Lewis, D., Agam, G., Argamon, S., Frieder, O., Grossman, D., Heard, J.: Building a test collection for complex document information processing. In: Proceedings of the 29th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (2006)
2006
Earlier work this paper cites.
Hamza, H., Belaïd, Y., Belaïd, A., Chaudhuri, B.: An end-to-end administrative document analysis system. In: 2008 The Eighth IAPR International Workshop on Document Analysis Systems. pp. 175–182 (2008)
2008
Earlier work this paper cites.
Bart, E., Sarkar, P.: Information extraction by finding repeated structure. In: DAS ’10 (2010)
2010
Earlier work this paper cites.
Medvet, E., Bartoli, A., Davanzo, G.: A probabilistic approach to printed document understanding. IJDAR 14
2011
Earlier work this paper cites.
Peanho, C., Stagni, H., Silva, F.: Semantic information extraction from images of complex documents. Applied Intelligence 37
2012
Earlier work this paper cites.
Rusinol, M., Benkhelfallah, T., Poulain d’Andecy, V.: Field extraction from administrative documents by incremental structural templates. In: ICDAR (2013)
2013
Earlier work this paper cites.
Gehring, J., Auli, M., Grangier, D., Yarats, D., Dauphin, Y.N.: Convolutional sequence to sequence learning. In: ICML (2017)
2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. In: Advances in Neural Information Processing Systems 30 (2017)
2017
Earlier work this paper cites.
Katti, A.R., Reisswig, C., Guder, C., Brarda, S., Bickel, S., Höhne, J., Faddoul, J.B.: Chargrid: Towards understanding 2D documents. In: EMNLP (2018)
2018
Earlier work this paper cites.
Shaw, P., Uszkoreit, J., Vaswani, A.: Self-attention with relative position representations. In: NAACL-HLT (2018)
2018
Cited alongside, same era.
Amazon: Amazon Textract. https://aws.amazon.com/textract/ (accessed November 25, 2019) (2019)
2019
Cited alongside, same era.
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q.V., Salakhutdinov, R.: Transformer-XL: Attentive language models beyond a fixed-length context. In: ACL (2019)
2019
Cited alongside, same era.
Denk, T.I., Reisswig, C.: BERTgrid: Contextualized Embedding for 2D Document Representation and Understanding. In: Workshop on Document Intelligence at NeurIPS 2019 (2019)
2019
Cited alongside, same era.
Devlin, J., Chang, M.W., Lee, K., Toutanova, K.: BERT: Pre-training of deep bidirectional transformers for language understanding. In: NAACL-HLT (2019)
2019
Cited alongside, same era.
Park, S., Shin, S., Lee, B., Lee, J., Surh, J., Seo, M., Lee, H.: CORD: A Consolidated Receipt Dataset for Post-OCR Parsing. In: Document Intelligence Workshop at Neural Information Processing Systems (2019)
2019
Later among the works it cites.
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., Bowman, S.R.: GLUE: A multi-task benchmark and analysis platform for natural language understanding. In: Proceedings of ICLR (2019), https://gluebenchmark.com/ (accessed November 26, 2019)
2019
Later among the works it cites.
Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., Smith, N.A.: Don’t stop pretraining: Adapt language models to domains and tasks. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. pp. 8342–8360. Association for Computational Linguistics (2020). https://doi.org/10.18653/v1/2020.acl-main.740
2020
Closest in time.
ICDAR: Leaderboard of the Information Extraction Task, Robust Reading Competition. https://rrc.cvc.uab.es/?ch=13&com=evaluation&task=3 (accessed April 7, 2020) (2020)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Google: Cloud Document Understanding AI. https://cloud.google.com/document-understanding/docs/ (accessed November 25, 2019) (2019)
2019
Cited alongside, same era.
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, M., Chen, D., Lee, H., Ngiam, J., Le, Q.V., Wu, Y., Chen, Z.: Gpipe: Efficient training of giant neural networks using pipeline parallelism. In: NeurIPS (2019)
2019
Cited alongside, same era.
ICDAR: Competition on Scanned Receipts OCR and Information Extraction. https://rrc.cvc.uab.es/?ch=13 (accessed February 21, 2021) (2019)
2019
Cited alongside, same era.
Liu, X., Gao, F., Zhang, Q., Zhao, H.: Graph convolution for multimodal information extraction from visually rich documents. In: NAACL-HLT (2019)
2019
Cited alongside, same era.
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., Stoyanov, V.: RoBERTa: A Robustly Optimized BERT Pretraining Approach. ArXiv 1907.11692
2019
Cited alongside, same era.
Microsoft: Cognitive Services. https://azure.microsoft.com/en-us/services/cognitive-services/ (accessed November 25, 2019) (2019)
2019
Cited alongside, same era.
2020
Closest in time.
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., Liu, P.J.: Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research 21
2020
Closest in time.
Rahman, W., Hasan, M., Lee, S., Zadeh, A., Mao, C., Morency, L.P., Hoque, E.: Integrating multimodal information in large pretrained transformers. In: ACL (2020)
2020
Closest in time.
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T.L., Gugger, S., Drame, M., Lhoest, Q., Rush, A.M.: Transformers: State-of-the-art natural language processing. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. pp. 38–45. Association for Computational Linguistics, Online (Oct 2020), https://www.aclweb.org/anthology/2020.emnlp-demos.6
2020
Closest in time.
Xu, Y., Xu, Y., Lv, T., Cui, L., Wei, F., Wang, G., Lu, Y., Florencio, D., Zhang, C., Che, W., Zhang, M., Zhou, L.: LayoutLMv2: Multi-modal pre-training for visually-rich document understanding. arXiv 2012.14740
2020
Closest in time.
Xu, Y., Li, M., Cui, L., Huang, S., Wei, F., Zhou, M.: LayoutLM: Pre-training of text and layout for document image understanding. In: Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. p. 1192–1200 (2020)
2020
Closest in time.
Stanisławek, T., Graliński, F., Wróblewska, A., Lipiński, D., Kaliska, A., Rosalska, P., Topolski, B., Biecek, P.: Kleister: A novel task for information extraction involving long documents with complex layout. ArXiv (accepted to ICDAR 2021) 2105.05796
2021
Closest in time.
Yu, W., Lu, N., Qi, X., Gong, P., Xiao, R.: PICK: Processing key information extraction from documents using improved graph learning-convolutional networks. In: 2020 25th International Conference on Pattern Recognition (ICPR). pp. 4363–4370 (2021). https://doi.org/10.1109/ICPR48806.2021.9412927
2021
Closest in time.