Fetching the paper…
Reading the bibliography…
In this work, we propose a new framework, called Document Image Transformer (DocTr), to address the issue of geometry and illumination distortion of the document images.
Binary codes capable of correcting deletions, insertions, and reversals
Vladimir I Levenshtein. 1966 · 1966
Earlier work this paper cites.
Document restoration using 3D shape: a general deskewing algorithm for arbitrarily warped documents. In Proceedings of the IEEE International Conference on Computer Vision , Vol. 2. 367–374
Michael S Brown and W Brent Seales. 2001 · 2001
Earlier work this paper cites.
Active contours network to straighten distorted text lines. In Proceedings of the International Conference on Image Processing , Vol. 3. 748–751
Olivier Lavialle, X Molines, Franck Angella, and Pierre Baylou. 2001 · 2001
Earlier work this paper cites.
Document image de-warping for text/graphics recognition. In Joint IAPR International Workshops on Statistical Techniques in Pattern Recognition (SPR) and Structural and Syntactic Pattern Recognition (SSPR) . Springer, 348–357
Changhua Wu and Gady Agam. 2002 · 2002
Earlier work this paper cites.
Rectifying the bound document image captured by the camera: a model based approach. In Proceedings of the International Conference on Document Analysis and Recognition , Vol. 1. 71–75
Huaigu Cao, Xiaoqing Ding, and Changsong Liu. 2003 · 2003
Earlier work this paper cites.
Multiscale structural similarity for image quality assessment. In Proceedings of the Asilomar Conference on Signals, Systems Computers , Vol. 2. 1398–1402
Zhou Wang, Eero P. Simoncelli, and Alan C. Bovik. 2003 · 2003
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. 2004 · 2004
Earlier work this paper cites.
A tutorial on the cross-entropy method
Pieter-Tjerk De Boer, Dirk P Kroese, Shie Mannor, and Reuven Y Rubinstein. 2005 · 2005
Earlier work this paper cites.
Document image de-warping based on detection of distorted text lines. In Proceedings of the International Conference on Image Analysis and Processing . Springer, 1068–1075
Lothar Mischke and Wolfram Luther. 2005 · 2005
Earlier work this paper cites.
Restoring warped document images through 3D shape modeling
Chew Lim Tan, Li Zhang, Zheng Zhang, and Tao Xia. 2006 · 2006
Earlier work this paper cites.
An Overview of the Tesseract OCR Engine. In Proceedings of the International Conference on Document Analysis and Recognition , Vol. 2. 629–633
Ray Smith. 2007 · 2007
Earlier work this paper cites.
An Improved Physically-Based Method for Geometric Restoration of Distorted Document Images
Li Zhang, Yu Zhang, and Chew Tan. 2008 · 2008
Earlier work this paper cites.
Composition of a Dewarped and Enhanced Document Image From Two View Images
Hyung Il Koo, Jinho Kim, and Nam Ik Cho. 2009 · 2009
Earlier work this paper cites.
SIFT Flow: Dense Correspondence across Scenes and Its Applications
Ce Liu, Jenny Yuen, and Antonio Torralba. 2011 · 2011
Earlier work this paper cites.
Pre-Trained Image Processing Transformer
Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. 2020b · 2012
Cited alongside, same era.
A Book Dewarping System by Boundary-Based 3D Surface Reconstruction. In Proceedings of the International Conference on Document Analysis and Recognition . 403–407
Yuan He, Pan Pan, Shufu Xie, Jun Sun, and Satoshi Naoi. 2013 · 2013
Cited alongside, same era.
Describing Textures in the Wild. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 3606–3613
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. 2014 · 2014
Cited alongside, same era.
Active Flattening of Curved Document Images via Two Structured Beams. In Proceedings of the IEEE International Conference on Computer Vision . 3890–3897
Gaofeng Meng, Ying Wang, Shenquan Qu, Shiming Xiang, and Chunhong Pan. 2014 · 2014
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Document rectification and illumination correction using a patch-based CNN
Xiaoyu Li, Bo Zhang, Jing Liao, and Pedro V Sander. 2019 · 2019
Later among the works it cites.
Decoupled Weight Decay Regularization. In Proceedings of the International Conference on Learning Representations
I. Loshchilov and F. Hutter. 2019 · 2019
Later among the works it cites.
Super-convergence: Very fast training of neural networks using large learning rates. In Artificial Intelligence and Machine Learning for Multi-Domain Operations Applications , Vol. 11006. International Society for Optics and Photonics, 1100612
Leslie N Smith and Nicholay Topin. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer-assisted Intervention . Springer, 234–241
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Cited alongside, same era.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017 · 2017
Cited alongside, same era.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
DocUNet: Document Image Unwarping via a Stacked U-Net. In Proceedings of the IEEE International Conference on Computer Vision . 4700–4709
Ke Ma, Zhixin Shu, Xue Bai, Jue Wang, and Dimitris Samaras. 2018 · 2018
Cited alongside, same era.
Image Transformer. In Proceedings of the 35th International Conference on Machine Learning , Vol. 80. 4055–4064
Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran. 2018 · 2018
Cited alongside, same era.
Improving Language Understanding by Generative Pre-Training. Technical report, OpenAI
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Cited alongside, same era.
XLNet: Generalized Autoregressive Pretraining for Language Understanding. In Proceedings of the Neural Information Processing Systems
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
End-to-End Object Detection with Transformers. In Proceedings of the European Conference on Computer Vision . Springer International Publishing, 213–229
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020 · 2020
Later among the works it cites.
Transformer-Based Label Set Generation for Multi-Modal Multi-Label Emotion Detection. In Proceedings of the 28th ACM International Conference on Multimedia (Seattle, WA, USA) (MM ’20) . Association for Computing Machinery, New York, NY, USA, 512–520
Xincheng Ju, Dong Zhang, Junhui Li, and Guodong Zhou. 2020 · 2020
Later among the works it cites.
Semi-Supervised Multi-Modal Emotion Recognition with Cross-Modal Distribution Matching. In Proceedings of the 28th ACM International Conference on Multimedia (Seattle, WA, USA) (MM ’20) . Association for Computing Machinery, New York, NY, USA, 2852–2861
Jingjun Liang, Ruichen Li, and Qin Jin. 2020 · 2020
Later among the works it cites.
Can You Read Me Now? Content Aware Rectification Using Angle Supervision. In Proceedings of the European Conference on Computer Vision . Springer, 208–223
Amir Markovitz, Inbal Lavi, Or Perel, Shai Mazor, and Roee Litman. 2020 · 2020
Later among the works it cites.
U2-Net: Going deeper with nested U-structure for salient object detection
Xuebin Qin, Zichen Zhang, Chenyang Huang, Masood Dehghan, Osmar R. Zaiane, and Martin Jagersand. 2020 · 2020
Later among the works it cites.
Learning Semantic Concepts and Temporal Alignment for Narrated Video Procedural Captioning. In Proceedings of the 28th ACM International Conference on Multimedia (Seattle, WA, USA) (MM ’20) . Association for Computing Machinery, New York, NY, USA, 4355–4363
Botian Shi, Lei Ji, Zhendong Niu, Nan Duan, Ming Zhou, and Xilin Chen. 2020 · 2020
Later among the works it cites.
Dewarping Document Image by Displacement Flow Estimation with Fully Convolutional Network. In International Workshop on Document Analysis Systems . Springer, 131–144
Guowang Xie, Fei Yin, Xuyao Zhang, and Chenglin Liu. 2020 · 2020
Later among the works it cites.
DeVLBert: Learning Deconfounded Visio-Linguistic Representations. In Proceedings of the 28th ACM International Conference on Multimedia (Seattle, WA, USA) (MM ’20) . Association for Computing Machinery, New York, NY, USA, 4373–4382
Shengyu Zhang, Tan Jiang, Tan Wang, Kun Kuang, Zhou Zhao, Jianke Zhu, Jin Yu, Hongxia Yang, and Fei Wu. 2020 · 2020
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Closest in time.