Fetching the paper…
Reading the bibliography…
We call on the Document AI (DocAI) community to reevaluate current methodologies and embrace the challenge of creating more practically-oriented benchmarks.
Measurement of diversity
E. H. SIMPSON · 1949
Earlier work this paper cites.
Binary codes capable of correcting deletions, insertions, and reversals
Vladimir I Levenshtein et al · 1966
Earlier work this paper cites.
The well-calibrated bayesian
A Philip Dawid · 1982
Earlier work this paper cites.
The comparison and evaluation of forecasters
Morris H DeGroot and Stephen E Fienberg · 1983
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates
Bianca Zadrozny and Charles Elkan · 2002
Earlier work this paper cites.
Predicting good probabilities with supervised learning
Alexandru Niculescu-Mizil and Rich Caruana · 2005
Earlier work this paper cites.
Vqa: Visual question answering, 2015
Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C. Lawrence Zitnick, Dhruv Batra, and Devi Parikh · 2015
Earlier work this paper cites.
Evaluation of deep convolutional nets for document image classification and retrieval
Adam W Harley, Alex Ufkes, and Konstantinos G Derpanis · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using Bayesian binning
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht · 2015
Earlier work this paper cites.
Compositional semantic parsing on semi-structured tables
Panupong Pasupat and Percy Liang · 2015
Earlier work this paper cites.
WikiQA: A challenge dataset for open-domain question answering
Yi Yang, Wen-tau Yih, and Christopher Meek · 2015
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Newsqa: A machine comprehension dataset
Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman · 2016
Earlier work this paper cites.
Selective classification for deep neural networks
Yonatan Geifman and Ran El-Yaniv · 2017
Earlier work this paper cites.
A question answering approach for emotion cause extraction
Lin Gui, Jiannan Hu, Yulan He, Ruifeng Xu, Qin Lu, and Jiachen Du · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Lc-quad: A corpus for complex question answering over knowledge graphs
Priyansh Trivedi, Gaurav Maheshwari, Mohnish Dubey, and Jens Lehmann · 2017
Earlier work this paper cites.
STAIR captions: Constructing a large-scale Japanese image caption dataset
Yuya Yoshikawa, Yutaro Shigeto, and Akikazu Takeuchi · 2017
Earlier work this paper cites.
Vizwiz grand challenge: Answering visual questions from blind people, 2018
Danna Gurari, Qing Li, Abigale J. Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P. Bigham · 2018
Earlier work this paper cites.
Verification of the Expected Answer Type for Biomedical Question Answering
Sanjay Kamath, Brigitte Grau, and Yue Ma · 2018
Earlier work this paper cites.
KT-speech-crawler: Automatic dataset construction for speech recognition from YouTube videos
Egor Lakomkin, Sven Magg, Cornelius Weber, and Stefan Wermter · 2018
Earlier work this paper cites.
TVQA: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara Berg · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
Places: A 10 million image database for scene recognition
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2018
Earlier work this paper cites.
JuICe: A large scale distantly supervised dataset for open domain context-based code generation
Rajas Agashe, Srinivasan Iyer, and Luke Zettlemoyer · 2019
Earlier work this paper cites.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019
Aida Amini, Saadia Gabriel, Peter Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi · 2019
Earlier work this paper cites.
Icdar 2019 competition on scene text visual question answering
Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marçal Rusinol, Minesh Mathew, CV Jawahar, Ernest Valveny, and Dimosthenis Karatzas · 2019
Earlier work this paper cites.
Scene text visual question answering
Ali Furkan Biten, Ruben Tito, Andres Mafla, Lluis Gomez, Marçal Rusinol, Ernest Valveny, CV Jawahar, and Dimosthenis Karatzas · 2019
Earlier work this paper cites.
Quoref: A reading comprehension dataset with questions requiring coreferential reasoning
Pradeep Dasigi, Nelson F Liu, Ana Marasović, Noah A Smith, and Matt Gardner · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
SemEval-2019 task 10: Math question answering
Mark Hopkins, Ronan Le Bras, Cristian Petrescu-Prahova, Gabriel Stanovsky, Hannaneh Hajishirzi, and Rik Koncel-Kedziorski · 2019
Earlier work this paper cites.
Funsd: A dataset for form understanding in noisy scanned documents
Guillaume Jaume, Hazim Kemal Ekenel, and Jean-Philippe Thiran · 2019
Earlier work this paper cites.
PubMedQA: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu · 2019
Earlier work this paper cites.
Verified uncertainty calibration
Ananya Kumar, Percy Liang, and Tengyu Ma · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Earlier work this paper cites.
XQA: A cross-lingual open-domain question answering dataset
Jiahua Liu, Yankai Lin, Zhiyuan Liu, and Maosong Sun · 2019
Earlier work this paper cites.
Ocr-vqa: Visual question answering by reading text in images
Anand Mishra, Shashank Shekhar, Ajeet Kumar Singh, and Anirban Chakraborty · 2019
Cited alongside, same era.
Measuring calibration in deep learning
Jeremy Nixon, Michael W Dusenberry, Linchuan Zhang, Ghassen Jerfel, and Dustin Tran · 2019
Cited alongside, same era.
RUN through the streets: A new dataset and baseline models for realistic urban navigation
Tzuf Paz-Argaman and Reut Tsarfaty · 2019
Cited alongside, same era.
Evaluating model calibration in classification
Juozas Vaicenavicius, David Widmann, Carl Andersson, Fredrik Lindsten, Jacob Roll, and Thomas Schön · 2019
Cited alongside, same era.
Publaynet: largest dataset ever for document layout analysis
Xu Zhong, Jianbin Tang, and Antonio Jimeno Yepes · 2019
Cited alongside, same era.
Question answering over temporal knowledge graphs
Apoorv Saxena, Soumen Chakrabarti, and Partha Talukdar · 2021
Later among the works it cites.
Kleister: Key information extraction datasets involving long documents with complex layouts
Tomasz Stanislawek, Filip Gralinski, Anna Wróblewska, Dawid Lipinski, Agnieszka Kaliska, Paulina Rosalska, Bartosz Topolski, and Przemyslaw Biecek · 2021
Later among the works it cites.
Visualmrc: Machine reading comprehension on document images
Ryota Tanaka, Kyosuke Nishida, and Sen Yoshida · 2021
Later among the works it cites.
Document collection visual question answering
Rubèn Tito, Dimosthenis Karatzas, and Ernest Valveny · 2021
Later among the works it cites.
Icdar 2021 competition on document visual question answering
Rubèn Tito, Minesh Mathew, CV Jawahar, Ernest Valveny, and Dimosthenis Karatzas · 2021
Later among the works it cites.
ParsFEVER: a dataset for Farsi fact extraction and verification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Cited alongside, same era.
SubjQA: A Dataset for Subjectivity and Review Comprehension
Johannes Bjerva, Nikita Bhutani, Behzad Golshan, Wang-Chiew Tan, and Isabelle Augenstein · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
LifeQA: A real-life dataset for video question answering
Santiago Castro, Mahmoud Azab, Jonathan Stroud, Cristina Noujaim, Ruoyao Wang, Jia Deng, and Rada Mihalcea · 2020
Cited alongside, same era.
TutorialVQA: Question answering dataset for tutorial videos
Anthony Colas, Seokhwan Kim, Franck Dernoncourt, Siddhesh Gupte, Zhe Wang, and Doo Soon Kim · 2020
Cited alongside, same era.
Selective question answering under domain shift
Amita Kamath, Robin Jia, and Percy Liang · 2020
Cited alongside, same era.
Docbank: A benchmark dataset for document layout analysis, 2020
Minghao Li, Yiheng Xu, Lei Cui, Shaohan Huang, Furu Wei, Zhoujun Li, and Ming Zhou · 2020
Cited alongside, same era.
Majid Zarharan, Mahsa Ghaderan, Amin Pourdabiri, Zahra Sayedi, Behrouz Minaei-Bidgoli, Sauleh Eetemadi, and Mohammad Taher Pilehvar · 2021
Later among the works it cites.
NOAHQA: Numerical reasoning with interpretable graph question answering dataset
Qiyuan Zhang, Lei Wang, Sicheng Yu, Shuohang Wang, Yang Wang, Jing Jiang, and Ee-Peng Lim · 2021
Later among the works it cites.
Knowing more about questions can help: Improving calibration in question answering
Shujian Zhang, Chengyue Gong, and Eunsol Choi · 2021
Later among the works it cites.
Global table extractor (gte): A framework for joint table identification and cell structure recognition using visual context
Xinyi Zheng, Douglas Burdick, Lucian Popa, Xu Zhong, and Nancy Xin Ru Wang · 2021
Later among the works it cites.
In-the-wild video question answering
Santiago Castro, Naihao Deng, Pingxuan Huang, Mihai Burzo, and Rada Mihalcea · 2022
Later among the works it cites.
Mapqa: A dataset for question answering on choropleth maps, 2022
Shuaichen Chang, David Palzer, Jialin Li, Eric Fosler-Lussier, and Ningchuan Xiao · 2022
Later among the works it cites.
PerKGQA: Question answering over personalized knowledge graphs
Ritam Dutt, Kasturi Bhattacharjee, Rashmi Gangadharaiah, Dan Roth, and Carolyn Rose · 2022
Later among the works it cites.
Overview of the MedVidQA 2022 shared task on medical video question-answering
Deepak Gupta and Dina Demner-Fushman · 2022
Later among the works it cites.
CHEF: A pilot Chinese dataset for evidence-based fact-checking
Xuming Hu, Zhijiang Guo, GuanYu Wu, Aiwei Liu, Lijie Wen, and Philip Yu · 2022
Later among the works it cites.
Layoutlmv3: Pre-training for document ai with unified text and image masking
Yupan Huang, Tengchao Lv, Lei Cui, Yutong Lu, and Furu Wei · 2022
Later among the works it cites.
Multispanqa: A dataset for multi-span question answering
Haonan Li, Martin Tomko, Maria Vasardani, and Timothy Baldwin · 2022
Later among the works it cites.
Dit: Self-supervised pre-training for document image transformer
Junlong Li, Yiheng Xu, Tengchao Lv, Lei Cui, Cha Zhang, and Furu Wei · 2022
Later among the works it cites.
Qalayout: Question answering layout based on multimodal attention for visual question answering on corporate document
Ibrahim Souleiman Mahamoud, Mickaël Coustaty, Aurélie Joseph, Vincent Poulain d’Andecy, and Jean-Marc Ogier · 2022
Later among the works it cites.
Infographicvqa
Minesh Mathew, Viraj Bagal, Rubèn Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar · 2022
Later among the works it cites.
NumGLUE: A suite of fundamental yet challenging mathematical reasoning tasks
Swaroop Mishra, Arindam Mitra, Neeraj Varshney, Bhavdeep Sachdeva, Peter Clark, Chitta Baral, and Ashwin Kalyan · 2022
Later among the works it cites.
Towards improving calibration in object detection under domain shift
Muhammad Akhtar Munir, Muhammad Haris Khan, M Saquib Sarfraz, and Mohsen Ali · 2022
Later among the works it cites.
Overview of BioASQ 2022: The tenth BioASQ challenge on large-scale biomedical semantic indexing and question answering
Anastasios Nentidis, Georgios Katsimpras, Eirini Vandorou, Anastasia Krithara, Antonio Miranda-Escalada, Luis Gasco, Martin Krallinger, and Georgios Paliouras · 2022
Later among the works it cites.
Stable: Table generation framework for encoder-decoder models, 2022
Michał Pietruszka, Michał Turski, Łukasz Borchmann, Tomasz Dwojak, Gabriela Pałka, Karolina Szyndler, Dawid Jurkiewicz, and Łukasz Garncarek · 2022
Later among the works it cites.
DuReader vis \textrm{DuReader}_{\textrm{vis}} : A Chinese dataset for open-domain document visual question answering
Le Qi, Shangwen Lv, Hongyu Li, Jing Liu, Yu Zhang, Qiaoqiao She, Hua Wu, Haifeng Wang, and Ting Liu · 2022
Later among the works it cites.
T5score: Discriminative fine-tuning of generative evaluation metrics
Yiwei Qin, Weizhe Yuan, Graham Neubig, and Pengfei Liu · 2022
Later among the works it cites.
Mitigating bias in calibration error estimation
Rebecca Roelofs, Nicholas Cain, Jonathon Shlens, and Michael C Mozer · 2022
Later among the works it cites.
Pubtables-1m: Towards comprehensive table extraction from unstructured documents
Brandon Smock, Rohith Pesala, and Robin Abraham · 2022
Later among the works it cites.
Hierarchical multimodal transformers for multi-page docvqa
Rubèn Tito, Dimosthenis Karatzas, and Ernest Valveny · 2022
Later among the works it cites.
Towards fine-grained causal reasoning and qa, 2022
Linyi Yang, Zhen Wang, Yuxiang Wu, Jie Yang, and Yue Zhang · 2022
Later among the works it cites.
On multi-domain long-tailed recognition, generalization and beyond
Yuzhe Yang, Hao Wang, and Dina Katabi · 2022
Later among the works it cites.
End-to-end spoken conversational question answering: Task, dataset and model
Chenyu You, Nuo Chen, Fenglin Liu, Shen Ge, Xian Wu, and Yuexian Zou · 2022
Later among the works it cites.
NAIL: A challenging benchmark for na\”ive logical reasoning, 2022
Xinbo Zhang, Changzhi Sun, Yue Zhang, Lei Li, and Hao Zhou · 2022
Later among the works it cites.
Towards complex document understanding by discrete reasoning
Fengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang, Haozhou Zhang, and Tat-Seng Chua · 2022
Later among the works it cites.
https://spacy.io/models/en
SpaCy · 2023
Closest in time.
A call to reflect on evaluation practices for failure detection in image classification
Paul F Jaeger, Carsten Tim Lüth, Lukas Klein, and Till J. Bungert · 2023
Closest in time.
Player of jeopardy: Chatgpt evaluation, 2023
Andreas Kirsch · 2023
Closest in time.
Slidevqa: A dataset for document visual question answering on multiple images, 2023
Ryota Tanaka, Kyosuke Nishida, Kosuke Nishida, Taku Hasegawa, Itsumi Saito, and Kuniko Saito · 2023
Closest in time.