Fetching the paper…
Reading the bibliography…
While many NLP pipelines assume raw, clean texts, many texts we encounter in the wild, including a vast majority of legal documents, are not so clean, with many of them being visually structured documents (VSDs) such as PDFs.
Random Forests
Leo Breiman. 2001 · 2001
Earlier work this paper cites.
Automatic Paragraph Identification: A Study across Languages and Domains
Caroline Sporleder and Mirella Lapata. 2004 · 2004
Earlier work this paper cites.
Sentence Boundary Detection and the Problem with the U.S
Dan Gillick. 2009 · 2009
Earlier work this paper cites.
Scikit-learn: Machine Learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011 · 2011
Earlier work this paper cites.
About the Panama Papers
Frederik Obermaier, Bastian Obermayer, Vanessa Wormer, and Wolfgang Jaschensky. 2016 · 2016
Earlier work this paper cites.
A Benchmark and Evaluation for Text Extraction from PDF
Hannah Bast and Claudius Korzen. 2017 · 2017
Earlier work this paper cites.
Deep Biaffine Attention for Neural Dependency Parsing
Timothy Dozat and Christopher D. Manning. 2017 · 2017
Cited alongside, same era.
Estimating Legal Document Structure by Considering Style Information and Table of Contents
Yoichi Hatsutori, Katsumasa Yoshikawa, and Haruki Imai. 2017 · 2017
Cited alongside, same era.
PDFdigest: an Adaptable Layout-Aware PDF-to-XML Textual Content Extractor for Scientific Articles
Daniel Ferrés, Horacio Saggion, Francesco Ronzano, and Àlex Bravo. 2018 · 2018
Cited alongside, same era.
DeepPDF: A Deep Learning Approach to Extracting Text from PDFs
Christopher Stahl, Steven Young, Drahomira Herrmannova, Robert Patton, and Jack Wells. 2018 · 2018
Cited alongside, same era.
EDGAR® Public Dissemination Service Technical Specification
The U.S. Securities and Exchange Commission. 2018 · 2018
Cited alongside, same era.
FinDSE@FinTOC-2019 Shared Task
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Visual Detection with Context for Document Layout Analysis
Carlos Soto and Shinjae Yoo. 2019 · 2019
Later among the works it cites.
AMR Parsing as Sequence-to-Graph Transduction
Sheng Zhang, Xutai Ma, Kevin Duh, and Benjamin Van Durme. 2019 · 2019
Later among the works it cites.
Layoutlm: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou. 2020 · 2020
Later among the works it cites.
FinSBD-2021: The 3rd Shared Task on Structure Boundary Detection in Unstructured Text in the Financial Domain
Willy Au, Abderrahim Ait-Azzi, and Juyeon Kang. 2021 · 2021
Closest in time.
LayoutLMv2: Multi-modal pre-training for visually-rich document understanding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Carla Abreu, Henrique Cardoso, and Eugénio Oliveira. 2019 · 2019
Cited alongside, same era.
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, and Lidong Zhou. 2021 · 2021
Closest in time.