Understand
For understanding generic documents, information like font sizes, column layout, and generally the positioning of words may carry semantic information that is crucial for solving a downstream document intelligence task.
- Our novel BERTgrid, which is based on Chargrid by Katti et al.
- (2018), represents a document as a grid of contextualized word piece embedding vectors, thereby making its spatial structure and semantics accessible to the processing neural network.
- The contextualized embedding vectors are retrieved from a BERT language model.
Built on
CUTIE: Learning to Understand Documents with Convolutional Universal Text Information Extractor
1903
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space
2013
Earlier work this paper cites.
Similar
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
2015
Cited alongside, same era.
Chargrid: Towards Understanding 2D Documents
2018
Cited alongside, same era.
Then
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
2019
Closest in time.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…