Space/time trade-offs in hash coding with allowable errors
Burton H. Bloom. 1970 · 1970
Earlier work this paper cites.
A dataset for statutory reasoning in tax law entailment and question answering
Original
Nils Holzenberger, Andrew Blair-Stanek, and Benjamin Van Durme. 2020 · 2005
Earlier work this paper cites.
Comparing automatic and human evaluation of nlg systems
Anja Belz and Ehud Reiter. 2006 · 2006
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Original
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2020 · 2009
Earlier work this paper cites.
An investigation into the validity of some metrics for automatically evaluating natural language generation systems
Ehud Reiter and Anja Belz. 2009 · 2009
Earlier work this paper cites.
chrF: character n-gram f-score for automatic mt evaluation
Maja Popović. 2015 · 2015
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Original
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
chrF++: words helping character n-grams
Maja Popović. 2017 · 2017
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
Bridging the gap between consumers’ medication questions and trusted answers
Asma Ben Abacha, Yassine Mrabet, Mark Sharp, Travis R Goodwin, Sonya E Shooshan, and Dina Demner-Fushman. 2019 · 2019
Earlier work this paper cites.
ELI5: Long Form Question Answering
Original
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019 · 2019
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. 2019 · 2019
Earlier work this paper cites.
Natural Questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019 · 2019
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Original
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018 · 2019
Earlier work this paper cites.
Extracting training data from large language models
Original
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom B. Brown, Dawn Xiaodong Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel. 2020 · 2020
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Original
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Original
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020 · 2020
Earlier work this paper cites.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021 · 2021
Earlier work this paper cites.
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Original
Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, and Matt Gardner. 2021 · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Original
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Earlier work this paper cites.
Investigating memorization of conspiracy theories in text generation
Sharon Levy, Michael Saxon, and William Yang Wang. 2021 · 2021
Earlier work this paper cites.