Fetching the paper…
Reading the bibliography…
Dynamic benchmarks interweave model fitting and data collection in an attempt to mitigate the limitations of static benchmarks.
Improved boosting algorithms using confidence-rated predictions
Robert E Schapire and Yoram Singer · 1999
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
A theory of learning from different domains
Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman Vaughan · 2010
Earlier work this paper cites.
The ladder: A reliable leaderboard for machine learning competitions
Avrim Blum and Moritz Hardt · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Preserving statistical validity in adaptive data analysis
Cynthia Dwork, Vitaly Feldman, Moritz Hardt, Toniann Pitassi, Omer Reingold, and Aaron Leon Roth · 2015
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández · 2016
Earlier work this paper cites.
SWAG: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi · 2018
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston · 2019
Earlier work this paper cites.
Inherent disagreements in human textual inferences
Ellie Pavlick and Tom Kwiatkowski · 2019
Earlier work this paper cites.
Do imagenet classifiers generalize to imagenet?
Benjamin Recht, Rebecca Roelofs, Ludwig Schmidt, and Vaishaal Shankar · 2019
Cited alongside, same era.
Adversarial filters of dataset biases
Ronan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew Peters, Ashish Sabharwal, and Yejin Choi · 2020
Cited alongside, same era.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 2020
Cited alongside, same era.
Measuring robustness to natural distribution shifts in image classification
Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, and Ludwig Schmidt · 2020
Cited alongside, same era.
Models in the loop: Aiding crowdworkers with generative annotation assistants
Max Bartolo, Tristan Thrush, Sebastian Riedel, Pontus Stenetorp, Robin Jia, and Douwe Kiela · 2021
Cited alongside, same era.
Dynaboard: An evaluation-as-a-service platform for holistic next-generation benchmarking
Zhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain, Ledell Wu, Robin Jia, Christopher Potts, Adina Williams, and Douwe Kiela · 2021
Later among the works it cites.
Personalized benchmarking with the ludwig benchmarking toolkit
Avanika Narayan, Piero Molino, Karan Goel, Willie Neiswanger, and Christopher Re · 2021
Later among the works it cites.
DynaSent: A dynamic benchmark for sentiment analysis
Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela · 2021
Later among the works it cites.
On releasing annotator-level labels and information in datasets
Vinodkumar Prabhakaran, Aida Mostafazadeh Davani, and Mark Diaz · 2021
Later among the works it cites.
TuringAdvice: A generative and dynamic evaluation of language use
Rowan Zellers, Ari Holtzman, Elizabeth Clark, Lianhui Qin, Ali Farhadi, and Yejin Choi · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What will it take to fix benchmarking in natural language understanding?
Samuel Bowman and George Dahl · 2021
Cited alongside, same era.
The GEM benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, Wanyu Du, Esin Durmus, Ondřej Dušek, Chris Chinenye Emezue, Varun Gangal, Cristina Garbacea, Tatsunori Hashimoto, Yufang Hou, Yacine Jernite, Harsh Jhamtani, Yangfeng Ji, Shailza Jolly, Mihir Kale, Dhruv Kumar, Faisal Ladhak, Aman Madaan, Mounica Maddela, Khyati Mahajan, Saad Mahamood, Bodhisattwa Prasad Majumder, Pedro Henrique Martins, Angelina McMillan-Major, Simon Mille, Emiel van Miltenburg, Moin Nadeem, Shashi Narayan, Vitaly Nikolaev, Andre Niyongabo Rubungo, Salomey Osei, Ankur Parikh, Laura Perez-Beltrachini, Niranjan Ramesh Rao, Vikas Raunak, Juan Diego Rodriguez, Sashank Santhanam, João Sedoc, Thibault Sellam, Samira Shaikh, Anastasia Shimorina, Marco Antonio Sobrevilla Cabezudo, Hendrik Strobelt, Nishant Subramani, Wei Xu, Diyi Yang, Akhila Yerukola, and Jiawei Zhou · 2021
Cited alongside, same era.
On the efficacy of adversarial data collection for question answering: Results from a large-scale randomized study
Divyansh Kaushik, Douwe Kiela, Zachary C. Lipton, and Wen-tau Yih · 2021
Cited alongside, same era.
Genie: A leaderboard for human-in-the-loop evaluation of text generation
Daniel Khashabi, Gabriel Stanovsky, Jonathan Bragg, Nicholas Lourie, Jungo Kasai, Yejin Choi, Noah A Smith, and Daniel S Weld · 2021
Cited alongside, same era.
Dynabench: Rethinking benchmarking in NLP
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adina Williams · 2021
Cited alongside, same era.
Dealing with disagreements: Looking beyond the majority vote in subjective annotations
Aida Mostafazadeh Davani, Mark Díaz, and Vinodkumar Prabhakaran · 2022
Closest in time.
Patterns, predictions, and actions: Foundations of machine learning
Moritz Hardt and Benjamin Recht · 2022
Closest in time.
Adversarially constructed evaluation sets are more challenging, but may not be fair
Jason Phang, Angelica Chen, William Huang, and Samuel R. Bowman · 2022
Closest in time.
Sequential nature of recommender systems disrupts the evaluation process
Ali Shirali · 2022
Closest in time.
Analyzing dynamic adversarial training data in the limit
Eric Wallace, Adina Williams, Robin Jia, and Douwe Kiela · 2022
Closest in time.