Fetching the paper…
Reading the bibliography…
Despite the successes of pretrained language models, there are still few high-quality, general-purpose QA systems that are freely available.
The winograd schema challenge
H. Levesque, E. Davis, and L. Morgenstern · 2011
Earlier work this paper cites.
Combining retrieval, statistics, and inference to answer elementary science questions
P. Clark, O. Etzioni, T. Khot, A. Sabharwal, O. Tafjord, P. D. Turney, and D. Khashabi · 2016
Earlier work this paper cites.
How to write science questions that are easy for people and hard for computers
E. Davis · 2016
Earlier work this paper cites.
Tracking the world state with recurrent entity networks
M. Henaff, J. Weston, A. D. Szlam, A. Bordes, and Y. LeCun · 2016
Earlier work this paper cites.
Towards AI-Complete question answering: A set of prerequisite toy tasks
J. Weston, A. Bordes, S. Chopra, and T. Mikolov · 2016
Earlier work this paper cites.
RACE: Large-scale reading comprehension dataset from examinations
G. Lai, Q. Xie, H. Liu, Y. Yang, and E. Hovy · 2017
Earlier work this paper cites.
Think you have solved question answering? Try ARC, the AI2 Reasoning Challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Tracking state changes in procedural text: a challenge dataset and models for process paragraph comprehension
B. Dalvi, L. Huang, N. Tandon, W. tau Yih, and P. Clark · 2018
Earlier work this paper cites.
P. A. Jansen, E. Wainwright, S. Marmorstein, and C. T. Morrison · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Cited alongside, same era.
Reasoning about actions and state changes by injecting commonsense knowledge
N. Tandon, B. Dalvi, J. Grus, W. tau Yih, A. Bosselut, and P. Clark · 2018
Cited alongside, same era.
Natural questions: a benchmark for question answering research
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, M. Kelcey, J. Devlin, K. Lee, K. N. Toutanova, L. Jones, M.-W. Chang, A. Dai, J. Uszkoreit, Q. Le, and S. Petrov · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Unqovering stereotyping biases via underspecified questions
T. Li, D. Khashabi, T. Khot, A. Sabharwal, and V. Srikumar · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Later among the works it cites.
How much knowledge can you pack into the parameters of a language model?
A. Roberts, C. Raffel, and N. M. Shazeer · 2020
Later among the works it cites.
S. Bhakthavatsalam, D. Khashabi, T. Khot, B. D. Mishra, K. Richardson, A. Sabharwal, C. Schoenick, O. Tafjord, and P. Clark · 2021
Closest in time.
How much coffee was consumed during emnlp 2019? fermi problems: A new reasoning challenge for ai
A. Kalyan, A. Kumar, A. Chandrasekaran, A. Sabharwal, and P. Clark · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Leakage-adjusted simulatability: Can models generate non-trivial explanations of their behavior in natural language?
P. Hase, S. Zhang, H. Xie, and M. Bansal · 2020
Cited alongside, same era.
The curious case of neural text degeneration
A. Holtzman, J. Buys, M. Forbes, and Y. Choi · 2020
Cited alongside, same era.
Unifiedqa: Crossing format boundaries with a single qa system
D. Khashabi, S. Min, T. Khot, A. Sabharwal, O. Tafjord, P. Clark, and H. Hajishirzi
Cited in the paper.
Unifiedqa: Crossing format boundaries with a single QA system
D. Khashabi, S. Min, T. Khot, A. Sabharwal, O. Tafjord, P. Clark, and H. Hajishirzi
Cited in the paper.
Experiments testing gpt-3’s ability at commonsense reasoning: Results
G. Marcus and E. Davis
Cited in the paper.
Gpt-3, bloviator: Openai’s language generator has no idea what it’s talking about
G. Marcus and E. Davis
Cited in the paper.
Which linguist invented the lightbulb? presupposition verification for question-answering
N. Kim, E. Pavlick, B. K. Ayan, and D. Ramachandran · 2021
Closest in time.
Challenges in automated debiasing for toxic language detection
X. Zhou, M. Sap, S. Swayamdipta, N. A. Smith, and Y. Choi · 2021
Closest in time.