Fetching the paper…
Reading the bibliography…
The web contains countless semi-structured websites, which can be a rich source of information for populating knowledge bases.
Binary codes capable of correcting deletions, insertions, and reversals
V. I. Levenshtein · 1966
Earlier work this paper cites.
Wrapper induction for information extraction
N. Kushmerick, D. S. Weld, and R. B. Doorenbos · 1997
Earlier work this paper cites.
A hierarchical approach to wrapper induction
I. Muslea, S. Minton, and C. Knoblock · 1999
Earlier work this paper cites.
Roadrunner: Towards automatic data extraction from large web sites
V. Crescenzi, G. Mecca, P. Merialdo, et al · 2001
Earlier work this paper cites.
Xpath: Looking forward
D. Olteanu, H. Meuss, T. Furche, and F. Bry · 2002
Earlier work this paper cites.
Extracting structured data from web pages
A. Arasu and H. Garcia-Molina · 2003
Earlier work this paper cites.
Web data extraction based on partial tree alignment
Y. Zhai and B. Liu · 2005
Earlier work this paper cites.
Extracting web data using instance-based learning
Y. Zhai and B. Liu · 2007
Earlier work this paper cites.
Freebase: a collaboratively created graph database for structuring human knowledge
K. D. Bollacker, C. Evans, P. Paritosh, T. Sturge, and J. Taylor · 2008
Earlier work this paper cites.
Webtables: exploring the power of tables on the web
M. J. Cafarella, A. Y. Halevy, D. Z. Wang, E. Wu, and Y. Zhang · 2008
Earlier work this paper cites.
Distant supervision for relation extraction without labeled data
M. Mintz, S. Bills, R. Snow, and D. Jurafsky · 2009
Earlier work this paper cites.
Toward an architecture for never-ending language learning
A. Carlson, J. Betteridge, B. Kisiel, B. Settles, E. R. H. Jr., and T. M. Mitchell · 2010
Earlier work this paper cites.
Exploiting content redundancy for web information extraction
P. Gulhane, R. Rastogi, S. H. Sengamedu, and A. Tengli · 2010
Cited alongside, same era.
Modeling relations and their mentions without labeled text
S. Riedel, L. Yao, and A. McCallum · 2010
Cited alongside, same era.
Learning to adapt web information extraction knowledge and discovering new attributes via a bayesian approach
T.-L. Wong and W. Lam · 2010
Cited alongside, same era.
Automatic wrappers for large scale web extraction
N. N. Dalvi, R. Kumar, and M. A. Soliman · 2011
Cited alongside, same era.
Web-scale information extraction with Vertex
P. Gulhane, A. Madaan, R. R. Mehta, J. Ramamirtham, R. Rastogi, S. Satpal, S. H. Sengamedu, A. Tengli, and C. Tiwari · 2011
Cited alongside, same era.
From one tree to a forest: a unified solution for structured web data extraction
Q. Hao, R. Cai, Y. Pang, and L. Zhang · 2011
Diadem: Thousands of websites to a single database
T. Furche, G. Gottlob, G. Grasso, X. Guo, G. Orsi, C. Schallhart, and C. Wang · 2014
Later among the works it cites.
Semi-supervised web wrapper repair via recursive tree matching
J. P. Cohen, W. Ding, and A. Bagherjeiran · 2015
Later among the works it cites.
Knowledge-based trust: estimating the trustworthiness of web sources
X. L. Dong, E. Gabrilovich, K. Murphy, V. Dang, W. Horn, C. Lugaresi, S. Sun, and W. Zhang · 2015
Later among the works it cites.
Big data integration (synthesis lectures on data management)
X. L. Dong and D. Srivastava · 2015
Later among the works it cites.
Early steps towards web scale information extraction with lodie
A. L. Gentile, Z. Zhang, and F. Ciravegna · 2015
Later among the works it cites.
Wadar: Joint wrapper and data repair
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Knowledge-based weak supervision for information extraction of overlapping relations
R. Hoffmann, C. Zhang, X. Ling, L. S. Zettlemoyer, and D. S. Weld · 2011
Cited alongside, same era.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Cited alongside, same era.
Information extraction
R. Grishman · 2012
Cited alongside, same era.
Extraction and integration of partially overlapping web sources
M. Bronzi, V. Crescenzi, P. Merialdo, and P. Papotti · 2013
Cited alongside, same era.
Knowledge vault: a web-scale approach to probabilistic knowledge fusion
X. Dong, E. Gabrilovich, G. Heitz, W. Horn, N. Lao, K. Murphy, T. Strohmann, S. Sun, and W. Zhang · 2014
Cited alongside, same era.
From data fusion to knowledge fusion
X. Dong, E. Gabrilovich, G. Heitz, W. Horn, K. Murphy, S. Sun, and W. Zhang · 2014
Cited alongside, same era.
S. Ortona, G. Orsi, M. Buoncristiano, and T. Furche · 2015
Later among the works it cites.
Joint repairs for web wrappers
S. Ortona, G. Orsi, T. Furche, and M. Buoncristiano · 2016
Later among the works it cites.
Data programming: Creating large training sets, quickly
A. J. Ratner, C. M. De Sa, S. Wu, D. Selsam, and C. Ré · 2016
Later among the works it cites.
http://www.commoncrawl.com, 2017
Commoncrawl corpus · 2017
Later among the works it cites.
Knowledge verification for long tail verticals
F. Li, X. L. Dong, A. Largen, and Y. Li · 2017
Later among the works it cites.
Heterogeneous supervision for relation extraction: A representation learning approach
L. Liu, X. Ren, Q. Zhu, S. Zhi, H. Gui, H. Ji, and J. Han · 2017
Later among the works it cites.
The biggrams: the semi-supervised information extraction system from html: an improvement in the wrapper induction
M. Mironczuk · 2017
Later among the works it cites.