Fetching the paper…
Reading the bibliography…
Despite its importance to experimental design, statistical power (the probability that, given a real effect, an experiment will reject the null hypothesis) has largely been ignored by the NLP community.
ROBERTA: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni. 2019 · 1907
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019 · 1910
Earlier work this paper cites.
R. Thomas McCoy, Junghyun Min, and Tal Linzen. 2019 · 1911
Earlier work this paper cites.
Approximate statistical tests for comparing supervised classification learning algorithms
Thomas G. Dietterich. 1998 · 1923
Earlier work this paper cites.
Fixed-sample-size analysis of sequential observations
Frank J. Anscombe. 1954 · 1954
Earlier work this paper cites.
The statistical power of abnormal-social psychological research: A review
Jacob Cohen. 1962 · 1962
Earlier work this paper cites.
Case-control studies: Design, conduct, analysis
James J. Schlesselman. 1982 · 1982
Earlier work this paper cites.
Asymptotic and exact power for the McNemar test and its analogue with R controls per case
Stephen W. Duffy. 1984 · 1984
Earlier work this paper cites.
Sample size and power for pair-matched case-control studies
John E. Connett, Judith A. Smith, and Richard B. McHugh. 1987 · 1987
Earlier work this paper cites.
The 2 x 2 matched-pairs trial: Exact unconditional design and analysis
Samy Suissa and Jonathan J. Shuster. 1991 · 1991
Earlier work this paper cites.
On the sample size for studies based upon McNemar’s test
Peter A Lachenbruch. 1992 · 1992
Earlier work this paper cites.
Operating characteristics of a rank correlation test for publication bias
Colin B. Begg and Madhuchhanda Mazumdar. 1994 · 1994
Earlier work this paper cites.
Publication bias: The “file-drawer” problem in scientific inference
Jeffrey D. Scargle. 1999 · 1999
Earlier work this paper cites.
The abuse of power: The pervasive fallacy of power calculations for data analysis
John M. Hoenig and Dennis M. Heisey. 2001 · 2001
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
The PASCAL recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini. 2005 · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
On some pitfalls in automatic evaluation and significance testing for MT
Stefan Riezler and John T. Maxwell. 2005 · 2005
Earlier work this paper cites.
The second PASCAL recognising textual entailment challenge
Roy Bar-Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor. 2006 · 2006
Earlier work this paper cites.
Statistical comparisons of classifiers over multiple data sets
Janez Demšar. 2006 · 2006
Earlier work this paper cites.
Data Analysis Using Regression and Multilevel/Hierarchical Models
Andrew Gelman and Jennifer Hill. 2006 · 2006
Earlier work this paper cites.
Comparing spoken dialog corpora collected with recruited subjects versus real users
Hua Ai, Antoine Raux, Dan Bohus, Maxine Eskenazi, and Diane Litman. 2007 · 2007
Earlier work this paper cites.
The third PASCAL recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and Bill Dolan. 2007 · 2007
Cited alongside, same era.
Brief report: Post hoc power, observed power, a priori power, retrospective power, prospective power, achieved power: Sorting out appropriate uses of statistical power analyses
Daniel J. O’Keefe. 2007 · 2007
Cited alongside, same era.
A practical solution to the pervasive problems of p p values
Eric-Jan Wagenmakers. 2007 · 2007
Cited alongside, same era.
The fifth PASCAL recognizing textual entailment challenge
Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo. 2009 · 2009
Cited alongside, same era.
An empirical investigation of statistical significance in NLP
Taylor Berg-Kirkpatrick, David Burkett, and Dan Klein. 2012 · 2012
Cited alongside, same era.
The Winograd schema challenge
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli. 2018 · 2018
Later among the works it cites.
Sentence encoders on STILTs: Supplementary training on intermediate labeled-data tasks
Jason Phang, Thibault Févry, and Samuel R Bowman. 2018 · 2018
Later among the works it cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Later among the works it cites.
Know what you don’t know: Unanswerable questions for SQuAD
Pranav Rajpurkar, Robin Jia, and Percy Liang. 2018 · 2018
Later among the works it cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
A broad-coverage challenge corpus for sentence understanding through inference
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hector J. Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Cited alongside, same era.
Random effects structure for confirmatory hypothesis testing: Keep it maximal
Dale J. Barr, Roger Levy, Christoph Scheepers, and Harry J. Tily. 2013 · 2013
Cited alongside, same era.
Power failure: Why small sample size undermines the reliability of neuroscience
Katherine S. Button, John P. A. Ioannidis, Claire Mokrysz, Brian A. Nosek, Jonathan Flint, Emma S. J. Robinson, and Marcus R. Munafò. 2013 · 2013
Cited alongside, same era.
The McNemar test for binary matched-pairs data: mid- p p and asymptotic are better than exact conditional
Morten W. Fagerland, Stian Lydersen, and Petter Laake. 2013 · 2013
Cited alongside, same era.
The garden of forking paths: Why multiple comparisons can be a problem, even when there is no “fishing expedition” or “p-hacking” and the research hypothesis was posited ahead of time
Andrew Gelman and Eric Loken. 2013 · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Y. Ng, and Christopher Potts. 2013 · 2013
Cited alongside, same era.
A systematic comparison of smoothing techniques for sentence-level BLEU
Boxing Chen and Colin Cherry. 2014 · 2014
Cited alongside, same era.
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Later among the works it cites.
Statistical power in two-level models: A tutorial based on Monte Carlo simulation
Matthias G. Arend and Thomas Schäfer. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Show your work: Improved reporting of experimental results
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Judge the judges: A large-scale evaluation study of neural language models for online review generation
Cristina Garbacea, Samuel Carton, Shiyan Yan, and Qiaozhu Mei. 2019 · 2019
Later among the works it cites.
Don’t calculate post-hoc power using observed estimate of effect size
Andrew Gelman. 2019 · 2019
Later among the works it cites.
Unifying human and statistical evaluation for natural language generation
Tatsunori B Hashimoto, Hugh Zhang, and Percy Liang. 2019 · 2019
Later among the works it cites.
What have we (not) learnt from millions of scientific papers with P P values?
John P. A. Ioannidis. 2019 · 2019
Later among the works it cites.
Best practices for the human evaluation of automatically generated text
Chris van der Lee, Albert Gatt, Emiel van Miltenburg, Sander Wubben, and Emiel Krahmer. 2019 · 2019
Later among the works it cites.
Abandon statistical significance
Blakeley B. McShane, David Gal, Andrew Gelman, Christian Robert, and Jennifer L. Tackett. 2019 · 2019
Later among the works it cites.
Facebook FAIR’s WMT19 news translation task submission
Nathan Ng, Kyra Yee, Alexei Baevski, Myle Ott, Michael Auli, and Sergey Edunov. 2019 · 2019
Later among the works it cites.
FAIRSEQ: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R. Salakhutdinov, and Quoc V. Le. 2019 · 2019
Later among the works it cites.
Not all claims are created equal: Choosing the right statistical approach to assess hypotheses
Erfan Sadeqi Azer, Daniel Khashabi, Ashish Sabharwal, and Dan Roth. 2020 · 2020
Closest in time.
Plug and play language models: A simple approach to controlled text generation
Sumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung, Eric Frank, Piero Molino, Jason Yosinski, and Rosanne Liu. 2020 · 2020
Closest in time.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Closest in time.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Closest in time.
Statistical power and optimal design in experiments in which samples of participants respond to samples of stimuli
Jacob Westfall, David A. Kenny, and Charles M. Judd. 2014 · 2045
Closest in time.