Fetching the paper…
Reading the bibliography…
In this paper, we introduce Ranger - a toolkit to facilitate the easy use of effect-size-based meta-analysis for multi-task evaluation in NLP and IR.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
The Cranfield Tests on Index Language Devices . San Francisco, CA, USA
Cyril Cleverdon. 1967 · 1967
Earlier work this paper cites.
Distribution theory for glass’s estimator of effect size and related estimators
Larry V Hedges. 1981 · 1981
Earlier work this paper cites.
Introduction to meta-analysis
Michael Borenstein, Larry V Hedges, Julian PT Higgins, and Hannah R Rothstein. 2009 · 2009
Earlier work this paper cites.
Meta-analysis in clinical trials revisited
Rebecca DerSimonian and Nan Laird. 2015 · 2015
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Meta-analysis for retrieval experiments involving multiple test collections
Ian Soboroff. 2018 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Cwl_eval: An evaluation tool for information retrieval
Leif Azzopardi, Paul Thomas, and Alistair Moffat. 2019 · 2019
Cited alongside, same era.
What will it take to fix benchmarking in natural language understanding?
Samuel R. Bowman and George Dahl. 2021 · 2021
Cited alongside, same era.
Benchmarking: Past, present and future
Kenneth Church, Mark Liberman, and Valia Kordoni. 2021 · 2021
Cited alongside, same era.
Overview of the trec 2021 deep learning track
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Campos, and Jimmy Lin. 2022 · 2021
Cited alongside, same era.
Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware Sampling
Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021 · 2021
Later among the works it cites.
On the quality of the TREC-COVID IR test collections
Ellen M. Voorhees and Kirk Roberts. 2021 · 2021
Later among the works it cites.
The dangers of underclaiming: Reasons for caution when reporting how NLP systems fail
Samuel Bowman. 2022 · 2022
Later among the works it cites.
What are the best systems? new perspectives on nlp benchmarking
Pierre Colombo, Nathan Noiry, Ekhine Irurozki, and Stephan Clémençon. 2022 · 2022
Later among the works it cites.
Introducing neural bag of whole-words with colberter: Contextualized late interactions using enhanced reduction
Sebastian Hofstätter, Omar Khattab, Sophia Althammer, Mete Sertkan, and Allan Hanbury. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin, and Allan Hanbury. 2021 · 2021
Cited alongside, same era.
Streamlining evaluation with ir-measures
Sean MacAvaney, Craig Macdonald, and Iadh Ounis. 2022 · 2022
Later among the works it cites.