Fetching the paper…
Reading the bibliography…
Web search is an essential way for humans to obtain information, but it's still a great challenge for machines to understand the contents of web pages.
Introduction to neural network based approaches for question answering over knowledge graphs
Nilesh Chakraborty, Denis Lukovnikov, Gaurav Maheshwari, Priyansh Trivedi, Jens Lehmann, and Asja Fischer. 2019 · 1907
Earlier work this paper cites.
Semistructured data
Peter Buneman. 1997 · 1997
Earlier work this paper cites.
Wrapper induction for information extraction
Nicholas Kushmerick, Daniel S Weld, and Robert Doorenbos. 1997 · 1997
Earlier work this paper cites.
A hierarchical approach to wrapper induction
Ion Muslea, Steve Minton, and Craig Knoblock. 1999 · 1999
Earlier work this paper cites.
Wrapper induction: Efficiency and expressiveness
Nicholas Kushmerick. 2000 · 2000
Earlier work this paper cites.
Web wrapper induction: a brief survey
Sergio Flesca, Giuseppe Manco, Elio Masciari, Eugenio Rende, and Andrea Tagarelli. 2004 · 2004
Earlier work this paper cites.
A survey of table recognition
Richard Zanibbi, Dorothea Blostein, and James R Cordy. 2004 · 2004
Earlier work this paper cites.
2d conditional random fields for web information extraction
Jun Zhu, Zaiqing Nie, Ji-Rong Wen, Bo Zhang, and Wei-Ying Ma. 2005 · 2005
Earlier work this paper cites.
A survey of web information extraction systems
Chia-Hui Chang, Mohammed Kayed, Moheb R Girgis, and Khaled F Shaalan. 2006 · 2006
Earlier work this paper cites.
Building a test collection for complex document information processing
David D. Lewis, Gady Agam, Shlomo Argamon, Ophir Frieder, David A. Grossman, and Jefferson Heard. 2006 · 2006
Earlier work this paper cites.
Simultaneous record detection and attribute labeling in web data extraction
Jun Zhu, Zaiqing Nie, Ji-Rong Wen, Bo Zhang, and Wei-Ying Ma. 2006 · 2006
Earlier work this paper cites.
From one tree to a forest: a unified solution for structured web data extraction
Qiang Hao, Rui Cai, Yanwei Pang, and Lei Zhang. 2011 · 2011
Earlier work this paper cites.
Extraction and integration of partially overlapping web sources
Mirko Bronzi, Valter Crescenzi, Paolo Merialdo, and Paolo Papotti. 2013 · 2013
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Evaluation of deep convolutional nets for document image classification and retrieval
Adam W. Harley, Alex Ufkes, and Konstantinos G. Derpanis. 2015 · 2015
Earlier work this paper cites.
Dataset and neural recurrent sequence labeling model for open-domain factoid question answering
Peng Li, Wei Li, Zhengyan He, Xuguang Wang, Ying Cao, Jie Zhou, and Wei Xu. 2016 · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2016 · 2016
Cited alongside, same era.
Quasar: Datasets for question answering by search and reading
Bhuwan Dhingra, Kathryn Mazaitis, and William W Cohen. 2017 · 2017
Cited alongside, same era.
Searchqa: A new q&a dataset augmented with context from a search engine
Matthew Dunn, Levent Sagun, Mike Higgins, V Ugur Guney, Volkan Cirik, and Kyunghyun Cho. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
ICDAR2019 competition on scanned receipt OCR and information extraction
Zheng Huang, Kai Chen, Jianhua He, Xiang Bai, Dimosthenis Karatzas, Shijian Lu, and C. V. Jawahar. 2019 · 2019
Later among the works it cites.
FUNSD: A dataset for form understanding in noisy scanned documents
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran. 2019 · 2019
Later among the works it cites.
Openceres: When open information extraction meets the semi-structured web
Colin Lockard, Prashant Shiralkar, and Xin Luna Dong. 2019 · 2019
Later among the works it cites.
Ocr-vqa: Visual question answering by reading text in images
Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The amazing mysteries of the gutter: Drawing inferences between panels in comic book narratives
Mohit Iyyer, Varun Manjunatha, Anupam Guha, Yogarshi Vyas, Jordan Boyd-Graber, Hal Daume, and Larry S Davis. 2017 · 2017
Cited alongside, same era.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017 · 2017
Cited alongside, same era.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi. 2017 · 2017
Cited alongside, same era.
Race: Large-scale reading comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Neural reading comprehension and beyond
Danqi Chen. 2018 · 2018
Cited alongside, same era.
Quac: Question answering in context
Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Siva Reddy, Danqi Chen, and Christopher D Manning. 2019 · 2019
Later among the works it cites.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach. 2019 · 2019
Later among the works it cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Later among the works it cites.
Publaynet: Largest dataset ever for document layout analysis
Xu Zhong, Jianbin Tang, and Antonio Jimeno-Yepes. 2019 · 2019
Later among the works it cites.
ELECTRA: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Later among the works it cites.
Docbank: A benchmark dataset for document layout analysis
Minghao Li, Yiheng Xu, Lei Cui, Shaohan Huang, Furu Wei, Zhoujun Li, and Ming Zhou. 2020 · 2020
Later among the works it cites.
A survey on machine reading comprehension—tasks, evaluation metrics and benchmark datasets
Changchang Zeng, Shaobo Li, Qin Li, Jie Hu, and Jianjun Hu. 2020 · 2020
Later among the works it cites.
A graph representation of semi-structured data for web question answering
Xingyao Zhang, Linjun Shou, Jian Pei, Ming Gong, Lijie Wen, and Daxin Jiang. 2020 · 2020
Later among the works it cites.
Web Question Answering with Neurosymbolic Program Synthesis , page 328–343. Association for Computing Machinery, New York, NY, USA
Qiaochu Chen, Aaron Lamoreaux, Xinyu Wang, Greg Durrett, Osbert Bastani, and Isil Dillig. 2021 · 2021
Closest in time.
Docvqa: A dataset for vqa on document images
Minesh Mathew, Dimosthenis Karatzas, and C.V. Jawahar. 2021 · 2021
Closest in time.