Fetching the paper…
Reading the bibliography…
Form understanding depends on both textual contents and organizational structure.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
M-bert: Injecting multimodal information in the bert structure
Wasifur Rahman, Md Kamrul Hasan, Amir Zadeh, Louis-Philippe Morency, and Mohammed Ehsan Hoque. 2019 · 1908
Earlier work this paper cites.
A fast algorithm for bottom-up document layout analysis
Anikó Simon, J-C Pret, and A Peter Johnson. 1997 · 1997
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015 · 2015
Earlier work this paper cites.
A table detection method for pdf documents based on convolutional neural networks
Leipeng Hao, Liangcai Gao, Xiaohan Yi, and Zhi Tang. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Star-net: A spatial attention residue network for scene text recognition
Wei Liu, Chaofeng Chen, Kwan-Yee K Wong, Zhizhong Su, and Junyu Han. 2016 · 2016
Cited alongside, same era.
An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Baoguang Shi, Xiang Bai, and Cong Yao. 2016 · 2016
Cited alongside, same era.
Detecting text in natural image with connectionist text proposal network
Zhi Tian, Weilin Huang, Tong He, Pan He, and Yu Qiao. 2016 · 2016
Cited alongside, same era.
Multi-scale multi-task fcn for semantic page segmentation and table detection
Dafang He, Scott Cohen, Brian Price, Daniel Kifer, and C Lee Giles. 2017 · 2017
Cited alongside, same era.
Textboxes: A fast text detector with a single deep neural network
Minghui Liao, Baoguang Shi, Xiang Bai, Xinggang Wang, and Wenyu Liu. 2017 · 2017
Cited alongside, same era.
Gated recurrent convolution neural network for ocr
Jianfeng Wang and Xiaolin Hu. 2017 · 2017
Later among the works it cites.
Pixellink: Detecting scene text via instance segmentation
Dan Deng, Haifeng Liu, Xuelong Li, and Deng Cai. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Multimodal sentiment analysis using hierarchical fusion with context modeling
Navonil Majumder, Devamanyu Hazarika, Alexander Gelbukh, Erik Cambria, and Soujanya Poria. 2018 · 2018
Later among the works it cites.
Funsd: A dataset for form understanding in noisy scanned documents
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Maximillian Nickel and Douwe Kiela. 2017 · 2017
Cited alongside, same era.
Context-dependent sentiment analysis in user-generated videos
Soujanya Poria, Erik Cambria, Devamanyu Hazarika, Navonil Majumder, Amir Zadeh, and Louis-Philippe Morency. 2017 · 2017
Cited alongside, same era.
Document page decomposition by the bounding-box project
Jaekyu Ha, Robert M Haralick, and Ihsin T Phillips. 1995a
Cited in the paper.
Recursive xy cut using bounding boxes of connected components
Jaekyu Ha, Robert M Haralick, and Ihsin T Phillips. 1995b
Cited in the paper.
Words can shift: Dynamically adjusting word representations using nonverbal behaviors
Yansen Wang, Ying Shen, Zhun Liu, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2019 · 2019
Later among the works it cites.
Layoutlm: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020 · 2020
Closest in time.