Fetching the paper…
Reading the bibliography…
Document intelligence as a relatively new research topic supports many business applications.
A. W. Harley, A. Ufkes, and K. G. Derpanis, “Evaluation of deep convolutional nets for document image classification and retrieval,” in ICDAR , 2015
2015
Earlier work this paper cites.
R. Girshick, “Fast r-cnn,” in ICCV , 2015
2015
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in ICLR , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
P.-Y. Huang, F. Liu, S.-R. Shiang, J. Oh, and C. Dyer, “Attention-based multimodal neural machine translation,” in WMT , 2016
2016
Earlier work this paper cites.
D. Hendrycks and K. Gimpel, “Gaussian error linear units (gelus),” arXiv , vol. abs/1606.08415, 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” NeurIPS , 2017
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollar, and R. Girshick, “Mask r-cnn,” in ICCV , 2017
2017
Earlier work this paper cites.
H. Nam, J.-W. Ha, and J. Kim, “Dual attention networks for multimodal reasoning and matching,” in CVPR , 2017
2017
Earlier work this paper cites.
Z. Jin, J. Cao, H. Guo, Y. Zhang, and J. Luo, “Multimodal fusion with recurrent neural networks for rumor detection on microblogs,” in ACM-MM , 2017
2017
Earlier work this paper cites.
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, “Feature pyramid networks for object detection,” in CVPR , 2017
2017
Earlier work this paper cites.
A. Williams, N. Nangia, and S. R. Bowman, “A broad-coverage challenge corpus for sentence understanding through inference,” in NAACL , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
X. Yang, E. Yumer, P. Asente, M. Kraley, D. Kifer, and C. Lee Giles, “Learning to extract semantic structure from documents using multimodal fully convolutional neural networks,” in CVPR , 2017
2017
Earlier work this paper cites.
P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio, “Graph Attention Networks,” ICLR , 2018
2018
Earlier work this paper cites.
K. L. Jacob Devlin, Ming-Wei Chang and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in NAACL , 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
G. Jaume, H. K. Ekenel, and J.-P. Thiran, “Funsd: A dataset for form understanding in noisy scanned documents,” in ICDAR Workshop , 2019
2019
Earlier work this paper cites.
S. Park, S. Shin, B. Lee, J. Lee, J. Surh, M. Seo, and H. Lee, “Cord: a consolidated receipt dataset for post-ocr parsing,” in NeurIPS Workshop , 2019
2019
Earlier work this paper cites.
2019
Cited alongside, same era.
Q. Guo, X. Qiu, P. Liu, Y. Shao, X. Xue, and Z. Zhang, “Star-transformer,” in NAACL , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
X. Zhong, J. Tang, and A. J. Yepes, “Publaynet: largest dataset ever for document layout analysis,” in ICDAR , 2019
2019
Cited alongside, same era.
Y. Xu, Y. Xu, T. Lv, L. Cui, F. Wei, G. Wang, Y. Lu, D. Florencio, C. Zhang, W. Che, M. Zhang, and L. Zhou, “Layoutlmv2: Multi-modal pre-training for visually-rich document understanding,” in IJCNLP , 2021
2021
Later among the works it cites.
P. Li, J. Gu, J. Kuen, V. I. Morariu, H. Zhao, R. Jain, V. Manjunatha, and H. Liu, “Selfdoc: Self-supervised document representation learning,” in CVPR , 2021
2021
Later among the works it cites.
J. Gu, J. Kuen, V. Morariu, H. Zhao, R. Jain, N. Barmpalios, A. Nenkova, and T. Sun, “Unidoc: Unified pretraining framework for document understanding,” NeurIPS , 2021
2021
Later among the works it cites.
S. Appalaraju, B. Jasani, B. U. Kota, Y. Xie, and R. Manmatha, “Docformer: End-to-end transformer for document understanding,” in ICCV , 2021
2021
Later among the works it cites.
C. Li, B. Bi, M. Yan, W. Wang, S. Huang, F. Huang, and L. Si, “Structurallm: Structural pre-training for form understanding,” in IJCNLP , 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Z. Huang, K. Chen, J. He, X. Bai, D. Karatzas, S. Lu, and C. Jawahar, “Icdar2019 competition on scanned receipt ocr and information extraction,” in ICDAR , 2019
2019
Cited alongside, same era.
X. Zhou, D. Wang, and P. Krähenbühl, “Objects as points,” arXiv , vol. abs/1904.07850, 2019
2019
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
Y. Xu, M. Li, L. Cui, S. Huang, F. Wei, and M. Zhou, “Layoutlm: Pre-training of text and layout for document image understanding,” in KDD , 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” NeurIPS , 2020
2020
Cited alongside, same era.
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” ICCV , 2021
2021
Later among the works it cites.
X. Shu, L. Zhang, G.-J. Qi, W. Liu, and J. Tang, “Spatiotemporal co-attention recurrent neural networks for human-skeleton motion prediction,” IEEE TPAMI , 2021
2021
Later among the works it cites.
H. Akbari, L. Yuan, R. Qian, W.-H. Chuang, S.-F. Chang, Y. Cui, and B. Gong, “Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,” NIPS , 2021
2021
Later among the works it cites.
B. Chen, A. Rouditchenko, K. Duarte, H. Kuehne, S. Thomas, A. Boggust, R. Panda, B. Kingsbury, R. Feris, D. Harwath et al. , “Multimodal clustering networks for self-supervised learning from unlabeled videos,” in ICCV , 2021
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Later among the works it cites.
Y. Li, Y. Qian, Y. Yu, X. Qin, C. Zhang, Y. Liu, K. Yao, J. Han, J. Liu, and E. Ding, “Structext: Structured text understanding with multi-modal transformers,” in MM , 2021
2021
Later among the works it cites.
W. Lin, Q. Gao, L. Sun, Z. Zhong, K. Hu, Q. Ren, and Q. Huo, “Vibertgrid: a jointly trained multi-modal 2d document representation for key information extraction from documents,” in ICDAR , 2021
2021
Later among the works it cites.
T. Zhu, L. Li, J. Yang, S. Zhao, H. Liu, and J. Qian, “Multimodal sentiment analysis with image-text interaction network,” IEEE TMM , 2022
2022
Closest in time.
Z. Zhang, J. Zhang, J. Du, and F. Wang, “Split, embed and merge: An accurate table structure recognizer,” PR , 2022
2022
Closest in time.
Y. Liu, J. Wu, L. Qu, T. Gan, J. Yin, and L. Nie, “Self-supervised correlation learning for cross-modal retrieval,” IEEE TMM , 2022
2022
Closest in time.
T. Hong, D. Kim, M. Ji, W. Hwang, D. Nam, and S. Park, “Bros: A pre-trained language model focusing on text and layout for better key information extraction from documents,” in AAAI , 2022
2022
Closest in time.
X. Shu, J. Yang, R. Yan, and Y. Song, “Expansion-squeeze-excitation fusion network for elderly activity recognition,” IEEE TCSVT , 2022
2022
Closest in time.
J. Wang, L. Jin, and K. Ding, “Lilt: A simple yet effective language-independent layout transformer for structured document understanding,” in ACL , 2022
2022
Closest in time.