Fetching the paper…
Reading the bibliography…
The development of state-of-the-art systems in different applied areas of machine learning (ML) is driven by benchmarks, which have shaped the paradigm of evaluating generalisation capabilities from multiple perspectives.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
The theory of committees and elections
Duncan Black et al. 1958 · 1958
Earlier work this paper cites.
Preference, utility and subjective probability. inhandbook of mathematical psychology, ed. rd luce, rr bush and eh galanter, 3, 249–410
P Suppes. 1965 · 1965
Earlier work this paper cites.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020 · 1967
Earlier work this paper cites.
Voting procedures: A summary analysis
Hannu Nurmi. 1983 · 1983
Earlier work this paper cites.
Voting schemes for which it can be difficult to tell who won the election
John Bartholdi, Craig A Tovey, and Michael A Trick. 1989 · 1989
Earlier work this paper cites.
Theory of Choice , volume 38
Mark Aizerman and Fuad Aleskerov. 1995 · 1995
Earlier work this paper cites.
An introduction to vote-counting schemes
Jonathan Levin and Barry Nalebuff. 1995 · 1995
Earlier work this paper cites.
MultiBoosting: A Technique for Combining Boosting and Wagging
Geoffrey I Webb. 2000 · 2000
Earlier work this paper cites.
Rank Aggregation Methods for the Web
Cynthia Dwork, Ravi Kumar, Moni Naor, and Dandapani Sivakumar. 2001 · 2001
Earlier work this paper cites.
Three brief proofs of arrow’s impossibility theorem
John Geanakoplos. 2005 · 2005
Earlier work this paper cites.
Statistical Comparisons of Classifiers over Multiple Data Sets
Janez Demšar. 2006 · 2006
Earlier work this paper cites.
On the Robustness of Preference Aggregation in Noisy Environments
Ariel D Procaccia, Jeffrey S Rosenschein, and Gal A Kaminka. 2007 · 2007
Earlier work this paper cites.
The threshold aggregation
Fuad Aleskerov, Vyacheslav V Chistyakov, and Valery Kalyagin. 2010 · 2010
Earlier work this paper cites.
Group Recommender Systems: Combining Individual Models
Judith Masthoff. 2011 · 2011
Earlier work this paper cites.
Social choice and individual values
Kenneth J Arrow. 2012 · 2012
Earlier work this paper cites.
Electoral Systems: Paradoxes, Assumptions, and Procedures
Dan S Felsenthal and Moshé Machover. 2012 · 2012
Earlier work this paper cites.
Choosing aggregation rules for composite indicators
Giuseppe Munda. 2012 · 2012
Earlier work this paper cites.
On the Discriminative Power of Tournament Solutions
Felix Brandt and Hans Georg Seedig. 2016 · 2014
Earlier work this paper cites.
The myth of the condorcet winner
Paul H Edelman. 2015 · 2015
Earlier work this paper cites.
Should We Really Use Post-Hoc Tests Based on Mean-Ranks?
Alessio Benavoli, Giorgio Corani, and Francesca Mangili. 2016 · 2016
Earlier work this paper cites.
Learning to score system summaries for better content selection evaluation
Maxime Peyrard, Teresa Botschen, and Iryna Gurevych. 2017 · 2017
Earlier work this paper cites.
Ranking with Fairness Constraints
L Elisa Celis, Damian Straszak, and Nisheeth K Vishnoi. 2018 · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Cited alongside, same era.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
What will it take to fix benchmarking in natural language understanding?
Samuel R. Bowman and George Dahl. 2021 · 2021
Later among the works it cites.
How Linguistically Fair are Multilingual Pre-trained Language Models
Monojit Choudhury and Amit Deshpande. 2021 · 2021
Later among the works it cites.
Mostafa Dehghani, Yi Tay, Alexey A Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals. 2021 · 2021
Later among the works it cites.
Memorization vs. generalization : Quantifying data leakage in NLP performance evaluation
Aparna Elangovan, Jiayuan He, and Karin Verspoor. 2021 · 2021
Later among the works it cites.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Systems, procedures and voting rules in context: A primer for voting rule selection , volume 9
Adiel Teixeira De Almeida, Danielle Costa Morais, and Hannu Nurmi. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
ALBERT: A Lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Cited alongside, same era.
How the Transformers Broke NLP Leaderboards
Anna Rogers. 2019 · 2019
Cited alongside, same era.
ERNIE: Enhanced language representation with informative entities
Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu. 2019 · 2019
Cited alongside, same era.
The probability of violating arrow’s conditions
Keith L Dougherty and Jac C Heckelman. 2020 · 2020
Cited alongside, same era.
A CLIP-Enhanced Method for Video-Language Understanding
Guohao Li, Feng He, and Zhifan Feng. 2021 · 2021
Later among the works it cites.
Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation Benchmarking
Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain, Ledell Wu, Robin Jia, Christopher Potts, Adina Williams, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
How Robust are Model Rankings: A Leaderboard Customization Approach for Equitable Evaluation
Swaroop Mishra and Anjana Arunkumar. 2021 · 2021
Later among the works it cites.
StereoSet: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021 · 2021
Later among the works it cites.
AI and the Everything in the Whole Wide World Benchmark
Inioluwa Deborah Raji, Emily Denton, Emily M Bender, Alex Hanna, and Amandalynne Paullada. 2021 · 2021
Later among the works it cites.
Evaluation examples are not equally informative: How should that change NLP leaderboards?
Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, and Jordan Boyd-Graber. 2021 · 2021
Later among the works it cites.
Challenges and Opportunities in NLP Benchmarking
Sebastian Ruder. 2021 · 2021
Later among the works it cites.
How not to lie with a benchmark: rearranging NLP leaderboards
Tatiana Shavrina and Valentin Malykh. 2021 · 2021
Later among the works it cites.
Minchul Shin, Jonghwan Mun, Kyoung-Woon On, Woo-Young Kang, Gunsoo Han, and Eun-Sol Kim. 2021 · 2021
Later among the works it cites.
Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models
Boxin Wang, Chejian Xu, Shuohang Wang, Shuohang Wang, Zhe Gan, Yu Cheng, Jianfeng Gao, Ahmed Awadallah, and Bo Li. 2021 · 2021
Later among the works it cites.
Disembodied Machine Learning: On the Illusion of Objectivity in NLP
Zeerak Waseem, Smarika Lulz, Joachim Bingel, and Isabelle Augenstein. 2021 · 2021
Later among the works it cites.
HULK: An energy efficiency benchmark platform for responsible natural language processing
Xiyou Zhou, Zhiyu Chen, Xiaoyong Jin, and William Yang Wang. 2021 · 2021
Later among the works it cites.
Your fairness may vary: Pretrained language model fairness in toxic text classification
Ioana Baldini, Dennis Wei, Karthikeyan Natesan Ramamurthy, Moninder Singh, and Mikhail Yurochkin. 2022 · 2022
Closest in time.
InfoLM: A New Metric to Evaluate Summarization & Data2Text Generation
Pierre Jean A Colombo, Chloé Clavel, and Pablo Piantanida. 2022 · 2022
Closest in time.
Over-optimism in Benchmark Studies and the Multiplicity of Design and Analysis Options when Interpreting Their Results
Christina Nießl, Moritz Herrmann, Chiara Wiedemann, Giuseppe Casalicchio, and Anne-Laure Boulesteix. 2022 · 2022
Closest in time.
Mapping Global Dynamics of Benchmark Creation and Saturation in Artificial Intelligence
Simon Ott, Adriano Barbosa-Silva, Kathrin Blagec, Jan Brauner, and Matthias Samwald. 2022 · 2022
Closest in time.
ILDAE: Instance-level difficulty analysis of evaluation data
Neeraj Varshney, Swaroop Mishra, and Chitta Baral. 2022 · 2022
Closest in time.
HERO: Hierarchical encoder for Video+Language omni-representation pre-training
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu. 2020 · 2065
Closest in time.