Fetching the paper…
Reading the bibliography…
The field of Deep Learning (DL) has undergone explosive growth during the last decade, with a substantial impact on Natural Language Processing (NLP) as well.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Karl Popper: Logik der Forschung
Karl Popper. 1934 · 1934
Earlier work this paper cites.
Teoria statistica delle classi e calcolo delle probabilita
Carlo Bonferroni. 1936 · 1936
Earlier work this paper cites.
The structure of scientific revolutions , volume 111
Thomas S Kuhn. 1970 · 1970
Earlier work this paper cites.
Testing a point null hypothesis: The irreconcilability of p values and evidence
James O Berger and Thomas Sellke. 1987 · 1987
Earlier work this paper cites.
Individual comparisons by ranking methods
Frank Wilcoxon. 1992 · 1992
Earlier work this paper cites.
Artificial intelligence: an empirical science
Herbert A Simon. 1995 · 1995
Earlier work this paper cites.
Bootstrap approach to inference and power analysis based on three test statistics for covariance structure models
Ke-Hai Yuan and Kentaro Hayashi. 2003 · 2003
Earlier work this paper cites.
Why most published research findings are false
John P. A. Ioannidis. 2005 · 2005
Earlier work this paper cites.
On some pitfalls in automatic evaluation and significance testing for mt
Stefan Riezler and John T Maxwell III. 2005 · 2005
Earlier work this paper cites.
Extracting COVID-19 events from twitter
Shi Zong, Ashutosh Baheti, Wei Xu, and Alan Ritter. 2020 · 2006
Earlier work this paper cites.
Carbontracker: Tracking and predicting the carbon footprint of training deep learning models
Lasse F Wolff Anthony, Benjamin Kanding, and Raghavendra Selvan. 2020 · 2007
Earlier work this paper cites.
On the appropriateness of statistical tests in machine learning
Janez Demšar. 2008 · 2008
Earlier work this paper cites.
Last words: Empiricism is not a matter of faith
Ted Pedersen. 2008 · 2008
Earlier work this paper cites.
Clarin: Common language resources and technology infrastructure
Tamás Váradi, Peter Wittenburg, Steven Krauwer, Martin Wynne, and Kimmo Koskenniemi. 2008 · 2008
Earlier work this paper cites.
The cult of statistical significance: How the standard error costs us jobs, justice, and lives
Steve Ziliak and Deirdre Nansen McCloskey. 2008 · 2008
Earlier work this paper cites.
Replicability is not reproducibility: Nor is it good science
Chris Drummond. 2009 · 2009
Earlier work this paper cites.
Best practices for managing data annotation projects
Tina Tseng, Amanda Stent, and Domenic Maida. 2020 · 2009
Earlier work this paper cites.
Bayesian data analysis
John K Kruschke. 2010 · 2010
Earlier work this paper cites.
Handbook of markov chain monte carlo
Steve Brooks, Andrew Gelman, Galin Jones, and Xiao-Li Meng. 2011 · 2011
Earlier work this paper cites.
Underspecification presents challenges for credibility in modern machine learning
Alexander D’Amour, Katherine A. Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yi-An Ma, Cory Y. McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Zachary Nado, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Xuezhi Wang, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai, and D. Sculley. 2020 · 2011
Earlier work this paper cites.
Evaluating learning algorithms: a classification perspective
Nathalie Japkowicz and Mohak Shah. 2011 · 2011
Earlier work this paper cites.
Reproducible research in computational science
Roger D Peng. 2011 · 2011
Earlier work this paper cites.
Data mining and statistics for decision making
Stéphane Tufféry. 2011 · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio. 2012 · 2012
Earlier work this paper cites.
Measuring the prevalence of questionable research practices with incentives for truth telling
Leslie K John, George Loewenstein, and Drazen Prelec. 2012 · 2012
Earlier work this paper cites.
WILDS: A benchmark of in-the-wild distribution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Sara Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang. 2020 · 2012
Earlier work this paper cites.
Natural Language Annotation for Machine Learning: A guide to corpus-building for applications
James Pustejovsky and Amber Stubbs. 2012 · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P. Adams. 2012 · 2012
Earlier work this paper cites.
Offspring from reproduction problems: What replication failure teaches us
Antske Fokkens, Marieke van Erp, Marten Postma, Ted Pedersen, Piek Vossen, and Nuno Freire. 2013 · 2013
Earlier work this paper cites.
Bayesian data analysis
Andrew Gelman, John B Carlin, Hal S Stern, David B Dunson, Aki Vehtari, and Donald B Rubin. 2013 · 2013
Earlier work this paper cites.
Bayesian estimation supersedes the t test
John K Kruschke. 2013 · 2013
Earlier work this paper cites.
Research commentary—too big to fail: large samples and the p-value problem
Mingfeng Lin, Henry C Lucas Jr, and Galit Shmueli. 2013 · 2013
Earlier work this paper cites.
A bayesian wilcoxon signed-rank test based on the dirichlet process
Alessio Benavoli, Giorgio Corani, Francesca Mangili, Marco Zaffalon, and Fabrizio Ruggeri. 2014 · 2014
Earlier work this paper cites.
A bayesian approach for comparing cross-validated algorithms on multiple data sets
Giorgio Corani and Alessio Benavoli. 2015 · 2015
Earlier work this paper cites.
Last words: Computational linguistics and deep learning
Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Simple baseline for visual question answering
Bolei Zhou, Yuandong Tian, Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus. 2015 · 2015
Earlier work this paper cites.
Statistical tests, p values, confidence intervals, and power: a guide to misinterpretations
Sander Greenland, Stephen J Senn, Kenneth J Rothman, John B Carlin, Charles Poole, Steven N Goodman, and Douglas G Altman. 2016 · 2016
Earlier work this paper cites.
The social impact of natural language processing
Dirk Hovy and Shannon L. Spruit. 2016 · 2016
Earlier work this paper cites.
The Hitchhiker’s guide to Python: best practices for development
Kenneth Reitz and Tanya Schlusser. 2016 · 2016
Earlier work this paper cites.
Time for a change: a tutorial for comparing multiple classifiers through bayesian analysis
Alessio Benavoli, Giorgio Corani, Janez Demsar, and Marco Zaffalon. 2017 · 2017
Earlier work this paper cites.
Replicability analysis for natural language processing: Testing significance with multiple datasets
Rotem Dror, Gili Baumer, Marina Bogomolov, and Roi Reichart. 2017 · 2017
Earlier work this paper cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar. 2017 · 2017
Earlier work this paper cites.
Results blind science publishing
Joseph J Locascio. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Proceedings of the 12th international workshop on semantic evaluation
Marianna Apidianaki, Saif Mohammad, Jonathan May, Ekaterina Shutova, Steven Bethard, and Marine Carpuat. 2018 · 2018
Earlier work this paper cites.
Data statements for natural language processing: Toward mitigating system bias and enabling better science
Emily M. Bender and Batya Friedman. 2018 · 2018
Earlier work this paper cites.
Hard numbers: Language exclusion in computational linguistics and natural language processing
Martin Benjamin. 2018 · 2018
Earlier work this paper cites.
Experimental evidence for tipping points in social convention
Damon Centola, Joshua Becker, Devon Brackbill, and Andrea Baronchelli. 2018 · 2018
Earlier work this paper cites.
Three dimensions of reproducibility in natural language processing
K Bretonnel Cohen, Jingbo Xia, Pierre Zweigenbaum, Tiffany J Callahan, Orin Hargraves, Foster Goss, Nancy Ide, Aurélie Névéol, Cyril Grouin, and Lawrence E Hunter. 2018 · 2018
Earlier work this paper cites.
The hitchhiker’s guide to testing statistical significance in natural language processing
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018 · 2018
Earlier work this paper cites.
State of the art: Reproducibility in artificial intelligence
Odd Erik Gundersen and Sigbjørn Kjensmo. 2018 · 2018
Earlier work this paper cites.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger. 2018 · 2018
Earlier work this paper cites.
Bayesian data analysis for newcomers
John K Kruschke and Torrin M Liddell. 2018 · 2018
Earlier work this paper cites.
A performance evaluation of federated learning algorithms
Adrian Nilsson, Simon Smith, Gregor Ulm, Emil Gustavsson, and Mats Jirstrand. 2018 · 2018
Cited alongside, same era.
The preregistration revolution
Brian A Nosek, Charles R Ebersole, Alexander C DeHaven, and David T Mellor. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
Model evaluation, model selection, and algorithm selection in machine learning
Sebastian Raschka. 2018 · 2018
Cited alongside, same era.
Reproducibility in computational linguistics: are we willing to share?
Martijn Wieling, Josine Rawee, and Gertjan van Noord. 2018 · 2018
Cited alongside, same era.
Design challenges and misconceptions in neural sequence labeling
Jie Yang, Shuailong Liang, and Yue Zhang. 2018 · 2018
Really doing great at estimating cate? a critical look at ml benchmarking practices in treatment effect estimation
Alicia Curth, David Svensson, James Weatherall, and Mihaela van der Schaar. 2021 · 2021
Later among the works it cites.
The benchmark lottery
Mostafa Dehghani, Yi Tay, Alexey A Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals. 2021 · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Later among the works it cites.
How to write a bias statement: Recommendations for submissions to the workshop on gender bias in NLP
Christian Hardmeier, Marta R. Costa-jussà, Kellie Webster, Will Radford, and Su Lin Blodgett. 2021 · 2021
Later among the works it cites.
The hardware lottery
Sara Hooker. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Evaluating the underlying gender bias in contextualized word embeddings
Christine Basta, Marta R Costa-jussà, and Noe Casas. 2019 · 2019
Cited alongside, same era.
Beyond falsifiability: Normal science in a multiverse
Sean M Carroll. 2019 · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Show your work: Improved reporting of experimental results
Jesse Dodge, Suchin Gururangan, Dallas Card, Roy Schwartz, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
Deep dominance - how to properly compare deep neural models
Rotem Dror, Segev Shlomov, and Roi Reichart. 2019 · 2019
Cited alongside, same era.
Diversity in machine learning
Zhiqiang Gong, Ping Zhong, and Weidong Hu. 2019 · 2019
Cited alongside, same era.
Five sources of bias in natural language processing
Dirk Hovy and Shrimai Prabhumoye. 2021 · 2021
Later among the works it cites.
Is there a replication crisis in finance?
Theis Ingerslev Jensen, Bryan T Kelly, and Lasse Heje Pedersen. 2021 · 2021
Later among the works it cites.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021 · 2021
Later among the works it cites.
Bayesian analysis reporting guidelines
John K Kruschke. 2021 · 2021
Later among the works it cites.
Publishing fast and slow: A path toward generalizability in psychology and ai
Andrew Kyle Lampinen, Stephanie CY Chan, Adam Santoro, and Felix Hill. 2021 · 2021
Later among the works it cites.
huggingface/datasets: 1.12.1
Quentin Lhoest, Albert Villanova del Moral, Patrick von Platen, Thomas Wolf, Yacine Jernite, Abhishek Thakur, Lewis Tunstall, Suraj Patil, Mariama Drame, Julien Chaumond, Julien Plu, Joe Davison, Simon Brandeis, Teven Le Scao, Victor Sanh, Kevin Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Steven Liu, Nathan Raw, Sylvain Lesage, Théo Matussière, Lysandre Debut, Stas Bekman, and Clément Delangue. 2021 · 2021
Later among the works it cites.
Deep learning for road traffic forecasting: Does it make a difference?
Eric L Manibardo, Ibai Laña, and Javier Del Ser. 2021 · 2021
Later among the works it cites.
Scientific credibility of machine translation research: A meta-evaluation of 769 papers
Benjamin Marie, Atsushi Fujita, and Raphael Rubino. 2021 · 2021
Later among the works it cites.
Saif M. Mohammad. 2021 · 2021
Later among the works it cites.
Understanding the failure modes of out-of-distribution generalization
Vaishnavh Nagarajan, Anders Andreassen, and Behnam Neyshabur. 2021 · 2021
Later among the works it cites.
Do transformer modifications transfer across implementations and applications?
Sharan Narang, Hyung Won Chung, Yi Tay, Liam Fedus, Thibault Févry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li, Nan Ding, Jake Marcus, Adam Roberts, and Colin Raffel. 2021 · 2021
Later among the works it cites.
On releasing annotator-level labels and information in datasets
Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz. 2021 · 2021
Later among the works it cites.
Domain divergences: A survey and empirical analysis
Abhinav Ramesh Kashyap, Devamanyu Hazarika, Min-Yen Kan, and Roger Zimmermann. 2021 · 2021
Later among the works it cites.
Ml and nlp publications in 2021
Marek Rei. 2022 · 2021
Later among the works it cites.
Validity, reliability, and significance
Stefan Riezler and Michael Hagmann. 2021 · 2021
Later among the works it cites.
Annotators with attitudes: How annotator beliefs and identities bias toxic language detection
Maarten Sap, Swabha Swayamdipta, Laura Vianna, Xuhui Zhou, Yejin Choi, and Noah A Smith. 2021 · 2021
Later among the works it cites.
Targeting the benchmark: On methodology in current natural language processing research
David Schlangen. 2021 · 2021
Later among the works it cites.
CodeCarbon: Estimate and Track Carbon Emissions from Machine Learning Computing
Victor Schmidt, Kamal Goyal, Aditya Joshi, Boris Feld, Liam Conell, Nikolas Laskaris, Doug Blank, Jonathan Wilson, Sorelle Friedler, and Sasha Luccioni. 2021 · 2021
Later among the works it cites.
The multiberts: Bert reproductions for robustness analysis
Thibault Sellam, Steve Yadlowsky, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Turc, Jacob Eisenstein, Dipanjan Das, et al. 2021 · 2021
Later among the works it cites.
An error analysis framework for shallow surface realization
Anastasia Shimorina, Yannick Parmentier, and Claire Gardent. 2021 · 2021
Later among the works it cites.
We need to talk about random splits
Anders Søgaard, Sebastian Ebert, Jasmijn Bastings, and Katja Filippova. 2021 · 2021
Later among the works it cites.
Disagreement in human evaluation: blame the task not the annotators
Lucia Specia. 2021 · 2021
Later among the works it cites.
You are the best reviewer of your own papers: An owner-assisted scoring mechanism
Weijie Su. 2021 · 2021
Later among the works it cites.
Learning from disagreement: A survey
Alexandra N Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2021 · 2021
Later among the works it cites.
We need to talk about train-dev-test splits
Rob van der Goot. 2021 · 2021
Later among the works it cites.
Massive choice, ample tasks (MaChAmp): A toolkit for multi-task learning in NLP
Rob van der Goot, Ahmet Üstün, Alan Ramponi, Ibrahim Sharaf, and Barbara Plank. 2021 · 2021
Later among the works it cites.
Preregistering NLP research
Emiel van Miltenburg, Chris van der Lee, and Emiel Krahmer. 2021 · 2021
Later among the works it cites.
MassiveSumm: a very large-scale, very multilingual, news summarisation dataset
Daniel Varab and Natalie Schluter. 2021 · 2021
Later among the works it cites.
Disembodied machine learning: On the illusion of objectivity in nlp
Zeerak Waseem, Smarika Lulz, Joachim Bingel, and Isabelle Augenstein. 2021 · 2021
Later among the works it cites.
Examining the inductive bias of neural language models with artificial languages
Jennifer C. White and Ryan Cotterell. 2021 · 2021
Later among the works it cites.
Ood-bench: Benchmarking and understanding out-of-distribution generalization datasets and algorithms
Nanyang Ye, Kaican Li, Lanqing Hong, Haoyue Bai, Yiting Chen, Fengwei Zhou, and Zhenguo Li. 2021 · 2021
Later among the works it cites.
Cartography active learning
Mike Zhang and Barbara Plank. 2021 · 2021
Later among the works it cites.
Reproducibility criteria
Association for Computational Linguistics. 2022 · 2022
Closest in time.
Acm code of ethics and professional conduct
Association for Computing Machinery. 2022 · 2022
Closest in time.
CrossRE: A Cross-Domain Dataset for Relation Extraction
Elisa Bassignana and Barbara Plank. 2022 · 2022
Closest in time.
The european language resources association (elra)
ELRA. 1995 · 2022
Closest in time.
Repairing the cracked foundation: A survey of obstacles in evaluation practices for generated text
Sebastian Gehrmann, Elizabeth Clark, and Thibault Sellam. 2022 · 2022
Closest in time.
Towards climate awareness in nlp research
Daniel Hershcovich, Nicolas Webersinke, Mathias Kraus, Julia Anna Bingler, and Markus Leippold. 2022 · 2022
Closest in time.
Deep learning reproducibility and explainable ai (xai)
A-M Leventi-Peetz and T Östreich. 2022 · 2022
Closest in time.
Replicability vs. reproducibility — or is it the other way around?
Mark Liberman. 2015 · 2022
Closest in time.
Don’t blame the annotator: Bias already starts in the annotation instructions
Mihir Parmar, Swaroop Mishra, Mor Geva, and Chitta Baral. 2022 · 2022
Closest in time.
Statistical methods for annotation analysis
Silviu Paun, Ron Artstein, and Massimo Poesio. 2022 · 2022
Closest in time.
Does the market of citations reward reproducible work?
Edward Raff. 2022 · 2022
Closest in time.
How to review for acl rolling review?
Anna Rogers and Isabelle Augenstein. 2021 · 2022
Closest in time.
Compute trends across three eras of machine learning
Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn, and Pablo Villalobos. 2022 · 2022
Closest in time.
Submission guidelines and editorial policies
TMLR. 2022 · 2022
Closest in time.
Twitter post (@eturner303): Reading the openai gpt-3 paper
Elliot Turner. 2020 · 2022
Closest in time.
deep-significance-easy and meaningful statistical significance testing in the age of neural networks
Dennis Ulmer, Christian Hardmeier, and Jes Frellsen. 2022 · 2022
Closest in time.