Fetching the paper…
Reading the bibliography…
Scholastic trivia competitions test knowledge and intelligence through mastery of question answering.
A survey of cross-validation procedures for model selection
Sylvain Arlot and Alain Celisse · 1935
Earlier work this paper cites.
Distributional structure
Zellig S Harris · 1954
Earlier work this paper cites.
Decision analysis: introductory lectures on choices under uncertainty
Howard Raiffa · 1968
Earlier work this paper cites.
The use of questions in teaching
Meredith D Gall · 1970
Earlier work this paper cites.
A statistical interpretation of term specificity and its application in retrieval
Karen Spärck Jones · 1972
Earlier work this paper cites.
A vector space model for automatic indexing
G Salton, A Wong, and C S Yang · 1975
Earlier work this paper cites.
Finding structure in time
Jeffrey L. Elman · 1990
Earlier work this paper cites.
Murax: A robust linguistic approach for question answering using an on-line encyclopedia
Julian Kupiec · 1993
Earlier work this paper cites.
Some simple effective approximations to the 2-poisson model for probabilistic weighted retrieval
Stephen E. Robertson and Steve Walker · 1994
Earlier work this paper cites.
Confucianism and democracy
Francis Fukuyama · 1995
Earlier work this paper cites.
Deep blue system overview
Feng hsiung Hsu, Murray Campbell, and A. Joseph Hoane · 1995
Earlier work this paper cites.
Computers & thought
Alan M. Turing · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Opponent modeling in poker
Darse Billings, Denis Papp, Jonathan Schaeffer, and Duane Szafron · 1998
Earlier work this paper cites.
Pay enough or don’t pay at all *
Uri Gneezy and Aldo Rustichini · 2000
Earlier work this paper cites.
Do batch and user evaluations give the same results?
William Hersh, Andrew Turpin, Susan Price, Benjamin Chan, Dale Kramer, Lynetta Sacherek, and Daniel Olson · 2000
Earlier work this paper cites.
Building a question answering test collection
Ellen M. Voorhees and Dawn M. Tice · 2000
Earlier work this paper cites.
Exploiting redundancy in question answering
Charles L A Clarke, Gordon V Cormack, and Thomas R Lynam · 2001
Earlier work this paper cites.
The foundations of cost-sensitive learning
Charles Elkan · 2001
Earlier work this paper cites.
Discovery of inference rules for question-answering
Dekang Lin and Patrick Pantel · 2001
Earlier work this paper cites.
Overview of the trec 2001 question answering track
Ellen M. Voorhees · 2001
Earlier work this paper cites.
Fenomen “Chto? Gde? Kogda?”
A. Korin · 2002
Earlier work this paper cites.
Learning question classifiers
Xin Li and Dan Roth · 2002
Earlier work this paper cites.
Pruning improves heuristic search for cost-sensitive learning
Valentina Bayer Zubek and Thomas G. Dietterich · 2002
Earlier work this paper cites.
Reinforcement learning as classification: Leveraging modern classifiers
Michail G Lagoudakis and Ronald Parr · 2003
Earlier work this paper cites.
1 billion pages = 1 million dollars? mining the web to play “who wants to be a millionaire?”
S K Lam, D Pennock, Dan Cosley, and S Lawrence · 2003
Earlier work this paper cites.
Budgeted learning of Naive-Bayes classifiers
Daniel J Lizotte, Omid Madani, and Russell Greiner · 2003
Earlier work this paper cites.
Coreference-based summarization and question answering: a case for high precision anaphor resolution
Roland Stuckardt · 2003
Earlier work this paper cites.
Evaluating the evaluation: A case study using the TREC 2002 question answering track
Ellen M Voorhees · 2003
Earlier work this paper cites.
Test-cost sensitive naive bayes classification
Xiaoyong Chai, Lin Deng, Qiang Yang, and Charles X. Ling · 2004
Earlier work this paper cites.
Labeling images with a computer game
Luis von Ahn and Laura Dabbish · 2004
Earlier work this paper cites.
Overview of the TREC 2002 question answering track
Ellen M Voorhees · 2004
Earlier work this paper cites.
An expected utility approach to active feature-value acquisition
P Melville, M Saar-Tsechansky, F Provost, and R Mooney · 2005
Earlier work this paper cites.
Prisoner of Trebekistan: a decade in Jeopardy!
Bob Harris · 2006
Earlier work this paper cites.
Games with a purpose
Luis von Ahn · 2006
Earlier work this paper cites.
Frustratingly easy domain adaptation
Hal Daume III · 2007
Earlier work this paper cites.
Opponent modeling in scrabble
Mark Richards and Eyal Amir · 2007
Earlier work this paper cites.
Opponent modeling in Real-Time strategy games
Frederik Schadd, Sander Bakkes, and Pieter Spronck · 2007
Earlier work this paper cites.
The MAIN model : A heuristic approach to understanding technology effects on credibility
S. Shyam Sundar · 2007
Earlier work this paper cites.
Learning for control from multiple demonstrations
Adam Coates, Pieter Abbeel, and Andrew Y Ng · 2008
Earlier work this paper cites.
A unified architecture for natural language processing: deep neural networks with multitask learning
Ronan Collobert and Jason Weston · 2008
Earlier work this paper cites.
Question and answer Test-Train overlap in Open-Domain question answering datasets
Patrick Lewis, Pontus Stenetorp, and Sebastian Riedel · 2008
Earlier work this paper cites.
Investigating statistical machine learning as a tool for software development
Kayur Patel, James Fogarty, James A. Landay, and Beverly L. Harrison · 2008
Earlier work this paper cites.
Cheap and fast – but is it good? evaluating non-expert annotations for natural language tasks
Rion Snow, Brendan O’Connor, Daniel Jurafsky, and Andrew Ng · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J Deng, W Dong, R Socher, L Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Financial incentives and the “performance of crowds”
Winter Mason and Duncan J Watts · 2009
Earlier work this paper cites.
Dataset shift in machine learning
Joaquin Quionero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D Lawrence · 2009
Earlier work this paper cites.
Predicting protein structures with a multiplayer online game
Seth Cooper, Firas Khatib, Adrien Treuille, Janos Barbero, Jeehyung Lee, Michael Beenen, Andrew Leaver-Fay, David Baker, Zoran Popović, and Foldit Players · 2010
Earlier work this paper cites.
Building watson: An overview of the deepqa project
David A. Ferrucci, Eric W. Brown, Jennifer Chu-Carroll, James Fan, David Gondek, Aditya Kalyanpur, Adam Lally, J. William Murdock, Eric Nyberg, John M. Prager, Nico Schlaefer, and Christopher A. Welty · 2010
Earlier work this paper cites.
Supervised noun phrase coreference research: The first fifteen years
Vincent Ng · 2010
Earlier work this paper cites.
Efficient reductions for imitation learning
Stephane Ross and Drew Bagnell · 2010
Earlier work this paper cites.
Data Mining , page 1–17
Anand Rajaraman and Jeffrey David Ullman · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to No-Regret online learning
Stephane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Calibration of confidence measures in speech recognition
D Yu, J Li, and L Deng · 2011
Earlier work this paper cites.
Besting the quiz master: Crowdsourcing incremental classification games
Jordan Boyd-Graber, Brianna Satinoff, He He, and Hal Daumé III · 2012
Cited alongside, same era.
Calibrating predictive model estimates to support personalized medicine
Xiaoqian Jiang, Melanie Osl, Jihoon Kim, and Lucila Ohno-Machado · 2012
Cited alongside, same era.
Your starter for ten: 50 years of university challenge, 2012
David Taylor, Colin McNulty, and Jo Meek · 2012
Cited alongside, same era.
Semantic parsing on freebase from question-answer pairs
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy S. Liang · 2013
Cited alongside, same era.
Findings of the 2013 Workshop on Statistical Machine Translation
Ondřej Bojar, Christian Buck, Chris Callison-Burch, Christian Federmann, Barry Haddow, Philipp Koehn, Christof Monz, Matt Post, Radu Soricut, and Lucia Specia · 2013
Cited alongside, same era.
Relational inference for wikification
Estimating uncertainty online against an adversary
Volodymyr Kuleshov and Stefano Ermon · 2017
Later among the works it cites.
Newsqa: A machine comprehension dataset
Adam Trischler, Tong Wang, Xingdi Yuan, Justin Harris, Alessandro Sordoni, Philip Bachman, and Kaheer Suleman · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
The minds of many: Opponent modeling in a stochastic game
F B Von Der Osten, M Kirley, and T Miller · 2017
Later among the works it cites.
Learning distributed representations of texts and entities from knowledge base
Ikuya Yamada, Hiroyuki Shindo, Hideaki Takeda, and Yoshiyasu Takefuji · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiao Cheng and Dan Roth · 2013
Cited alongside, same era.
Analysis of watson’s strategies for playing jeopardy!
Gerald Tesauro, David Gondek, Jonathan Lenchner, James Fan, and John M. Prager · 2013
Cited alongside, same era.
Supervised sequential classification under budget constraints
Kirill Trapeznikov and Venkatesh Saligrama · 2013
Cited alongside, same era.
Turing Test as a Defining Feature of AI-Completeness , pages 3–17
Roman V. Yampolskiy · 2013
Cited alongside, same era.
A reliable effective terascale linear learning system
Alekh Agarwal, Olivier Chapelle, Miroslav Dudík, and John Langford · 2014
Cited alongside, same era.
On the properties of neural machine translation: Encoder-decoder approaches
Kyunghyun Cho, Bart van Merrienboer, Dzmitry Bahdanau, and Yoshua Bengio · 2014
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
R Girshick, J Donahue, T Darrell, and J Malik · 2014
Cited alongside, same era.
Christina Boyce-Jacinoand Simon DeDeo · 2018
Later among the works it cites.
Multinomial adversarial networks for multi-domain text classification
Xilun Chen and Claire Cardie · 2018
Later among the works it cites.
Dataset and baselines for sequential open-domain question answering
Ahmed Elgohary, Chen Zhao, and Jordan Boyd-Graber · 2018
Later among the works it cites.
Pathologies of neural models make interpretations difficult
Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan Boyd-Graber · 2018
Later among the works it cites.
To trust or not to trust a classifier
Heinrich Jiang, Been Kim, Melody Guan, and Maya Gupta · 2018
Later among the works it cites.
How much reading does reading comprehension require? a critical investigation of popular benchmarks
Divyansh Kaushik and Zachary Chase Lipton · 2018
Later among the works it cites.
Writing good quizbowl questions: A quick primer
Paul Lujan and Seth Teitler · 2018
Later among the works it cites.
Subash maddipoti’s tips on question writing
Subash Maddipoti · 2018
Later among the works it cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Later among the works it cites.
Simplequestions nearly solved: A new upperbound and baseline approach
Michael Petrochuk and Luke S. Zettlemoyer · 2018
Later among the works it cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang · 2018
Later among the works it cites.
Semantically equivalent adversarial rules for debugging nlp models
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2018
Later among the works it cites.
What makes reading comprehension questions easier?
Saku Sugawara, Kentaro Inui, Satoshi Sekine, and Akiko Aizawa · 2018
Later among the works it cites.
How to write questions
Jerry Vinokurov · 2018
Later among the works it cites.
Constructing datasets for multi-hop reading comprehension across documents
Johannes Welbl, Pontus Stenetorp, and Sebastian Riedel · 2018
Later among the works it cites.
Zero-shot learning - a comprehensive evaluation of the good, the bad and the ugly
Yongqin Xian, Christoph H. Lampert, Bernt Schiele, and Zeynep Akata · 2018
Later among the works it cites.
Studio ousia’s quiz bowl question answering system
Ikuya Yamada, Ryuji Tamaki, Hiroyuki Shindo, and Yoshiyasu Takefuji · 2018
Later among the works it cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Later among the works it cites.
SWAG: A Large-Scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi · 2018
Later among the works it cites.
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James R. Glass · 2019
Closest in time.
Superhuman AI for multiplayer poker
Noam Brown and Tuomas Sandholm · 2019
Closest in time.
Addressing failure prediction by learning model confidence
Charles Corbière, Nicolas Thome, Avner Bar-Hen, Matthieu Cord, and Patrick Pérez · 2019
Closest in time.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Closest in time.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner · 2019
Closest in time.
What can ai do for me: Evaluating machine learning interpretations in cooperative play
Shi Feng and Jordan Boyd-Graber · 2019
Closest in time.
Misleading failures of partial-input baselines
Shi Feng, Eric Wallace, and Jordan Boyd-Graber · 2019
Closest in time.
Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets
Mor Geva, Yoav Goldberg, and Jonathan Berant · 2019
Closest in time.
Natural adversarial examples
Dan Hendrycks, Kevin Zhao, Steven Basart, Jacob Steinhardt, and Dawn Song · 2019
Closest in time.
Qasc: A dataset for question answering via sentence composition
Tushar Khot, Peter Clark, Michal Guerquin, Paul Edward Jansen, and Ashish Sabharwal · 2019
Closest in time.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Matthew Kelcey, Jacob Devlin, Kenton Lee, Kristina N. Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov · 2019
Closest in time.
How uncle Jamie broke Jeopardy
Kenny Malone · 2019
Closest in time.
Compositional questions do not necessitate multi-hop reasoning
Sewon Min, Eric Wallace, Sameer Singh, Matt Gardner, Hannaneh Hajishirzi, and Luke S. Zettlemoyer · 2019
Closest in time.
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru · 2019
Closest in time.
Human vs. muppet: A conservative estimate of human performance on the GLUE benchmark
Nikita Nangia and Samuel R Bowman · 2019
Closest in time.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D Sculley, Sebastian Nowozin, Joshua V Dillon, Balaji Lakshminarayanan, and Jasper Snoek · 2019
Closest in time.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Closest in time.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Closest in time.
The FEVER2.0 shared task
James Thorne, Andreas Vlachos, Oana Cocarascu, Christos Christodoulopoulos, and Arpit Mittal · 2019
Closest in time.
2014-15 packet submission guidelines
Jerry Vinokurov, Gautam Kandlikar, Gaurav Kandlikar, Matthew Jackson, Ryan Westbrook, and Rob Carson · 2019
Closest in time.
Trick me if you can: Human-in-the-loop generation of adversarial question answering examples
Eric Wallace, Pedro Rodriguez, Shi Feng, and Jordan Boyd-Graber · 2019
Closest in time.
Beat the AI: Investigating adversarial human annotation for reading comprehension
Max Bartolo, Alastair Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp · 2020
Closest in time.
What question answering can learn from trivia nerds
Jordan Boyd-Graber and Benjamin Börschinger · 2020
Closest in time.
Open-Domain question answering
Danqi Chen and Wen-Tau Yih · 2020
Closest in time.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Victoria Basmov, Jonathan Berant, Ben Bogin, Sihao Chen, Pradeep Dasigi, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A Smith, Sanjay Subramanian, Reut Tsarfaty, Eric Wallace, Ally Zhang, and Ben Zhou · 2020
Closest in time.
Selective question answering under domain shift
Amita Kamath, Robin Jia, and Percy Liang · 2020
Closest in time.
Naqt | press guide
LLC National Academic Quiz Tournaments · 2020
Closest in time.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 2020
Closest in time.
Break it down: A question understanding benchmark
Tomer Wolfson, Mor Geva, Ankit Gupta, Matt Gardner, Yoav Goldberg, Daniel Deutch, and Jonathan Berant · 2020
Closest in time.
Robustness gym: Unifying the NLP evaluation landscape
Karan Goel, Nazneen Rajani, Jesse Vig, Samson Tan, Jason Wu, Stephan Zheng, Caiming Xiong, Mohit Bansal, and Christopher Ré · 2021
Closest in time.