Fetching the paper…
Reading the bibliography…
Modern language models often exhibit powerful but brittle behavior, leading to the development of larger and more diverse benchmarks to reliably assess their behavior.
Information-based objective functions for active data selection
David J. C. MacKay. 1992 · 1992
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett. 2004 · 2004
Earlier work this paper cites.
Finding groups in data: an introduction to cluster analysis
Leonard Kaufman and Peter J Rousseeuw. 2009 · 2009
Earlier work this paper cites.
Facility location: concepts, models, algorithms and case studies
Masoud Hekmatfar Reza Zanjirani Farahani. 2009 · 2009
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
Neil Houlsby, Ferenc Huszár, Zoubin Ghahramani, and Máté Lengyel. 2011 · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment tree-bank
Richard Socher, Alex Perelygin, Jean Wu, Christopher D Manning Jason Chuang, Andrew Ng, , and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Understanding Machine Learning: From Theory to Algorithms
Shai Ben-David and Shai Shalev-Shwartz. 2014 · 2014
Earlier work this paper cites.
Submodularity in data subset selection and active learning
Jeff Bilmes Kai Wei, Rishabh Iyer. 2015 · 2015
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, , and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017 · 2017
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
Ozan Sener and Silvio Savarese. 2017 · 2017
Earlier work this paper cites.
A theory of learning from different domains
Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman Vaughan. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Umap: Uniform manifold approximation and projection for dimension reduction
Leland McInnes, John Healy, and James Melville. 2018 · 2018
Earlier work this paper cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018 · 2018
Earlier work this paper cites.
Active learning at the imagenet scale
Zeyad Ali Sami Emam, Hong-Min Chu, Ping-Yeh Chiang, Wojciech Czaja, Richard Leapman, Micah Goldblum, and Tom Goldstein. 2019 · 2019
Earlier work this paper cites.
Batchbald: Efficient and diverse batch acquisition for deep bayesian active learning
Andreas Kirsch, Joost van Amersfoort, and Yarin Gal. 2019 · 2019
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019 · 2019
Cited alongside, same era.
Coresets via bilevel optimization for continual learning and streaming
Zalán Borsos, Mojmír Mutný, and Andreas Krause. 2020 · 2020
Cited alongside, same era.
What will it take to fix benchmarking in natural language understanding?
Samuel R. Bowman and George E. Dahl. 2020 · 2020
Cited alongside, same era.
With little power comes great responsibility
Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky. 2020 · 2020
Cited alongside, same era.
Computing the testing error without a testing set
Ciprian Corneanu, Meysam Madadi, Sergio Escalera, and Aleix Martinez. 2020 · 2020
Cited alongside, same era.
Introduction to core-sets: an updated survey
Dan Feldman. 2020 · 2020
True few-shot learning with language models
Ethan Perez, Douwe Kiela, and Kyunghyun Cho. 2021 · 2021
Later among the works it cites.
A survey of deep active learning
Pengzhen Ren, Yun Xiao, Xiaojun Chang, Po-Yao Huang, Zhihui Li, Brij B. Gupta, Xiaojiang Chen, and Xin Wang. 2021 · 2021
Later among the works it cites.
Evaluation examples are not equally informative: How should that change nlp leaderboards?
Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, and Jordan Boyd-Graber. 2021 · 2021
Later among the works it cites.
Beyond marginal uncertainty: How accurately can bayesian regression models estimate posterior predictive correlations?
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2021 · 2021
Later among the works it cites.
Logme: Practical assessment of pre-trained models for transfer learning
Kaichao You, Yong Liu, Jianmin Wang, and Mingsheng Long. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cartography active learning
Barbara Plank Mike Zhang. 2020 · 2020
Cited alongside, same era.
Coresets for data-efficient training of machine learning models
Baharan Mirzasoleiman, Jeff Bilmes, and Jure Leskovec. 2020 · 2020
Cited alongside, same era.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Imitation attacks and defenses for black-box machine translation systems
Eric Wallace, Mitchell Stern, and Dawn Song. 2020 · 2020
Cited alongside, same era.
Exploring the limits of large scale pre-training
Samira Abnar, Mostafa Dehghani, Behnam Neyshabur, and Hanie Sedghi. 2021 · 2021
Cited alongside, same era.
Are labels always necessary for classifier accuracy evaluation?
Weijian Deng and Liang Zheng. 2021 · 2021
Cited alongside, same era.
Are larger pretrained language models uniformly better? comparing performance at the instance level
Ruiqi Zhong, Dhruba Ghosh, Dan Klein, and Jacob Steinhardt. 2021 · 2021
Later among the works it cites.
Promptsource: An integrated development environment and repository for natural language prompts
Stephen H. Bach, Victor Sanh, Zheng-Xin Yong, Albert Webson, Colin Raffel, Nihal V. Nayak, Abheesht Sharma, Taewoon Kim, M Saiful Bari, Thibault Fevry, Zaid Alyafeai, Manan Dey, Andrea Santilli, Zhiqing Sun, Srulik Ben-David, Canwen Xu, Gunjan Chhablani, Han Wang, Jason Alan Fries, Maged S. Al-shaibani, Shanya Sharma, Urmish Thakker, Khalid Almubarak, Xiangru Tang, Dragomir Radev, Mike Tian-Jian Jiang, and Alexander M. Rush. 2022 · 2022
Later among the works it cites.
Agreement-on-the-line: Predicting the performance of neural networks under distribution shift
Christina Baek, Yiding Jiang, Aditi Raghunathan, and Zico Kolter. 2022 · 2022
Later among the works it cites.
Evidence > intuition: Transferability estimation for encoder selection
Elisa Bassignana, Max Müller-Eberstein, Mike Zhang, and Barbara Plank. 2022 · 2022
Later among the works it cites.
Understanding dataset difficulty with v-usable information
Kawin Ethayarajh, Yejin Choi, and Swabha Swayamdipta. 2022 · 2022
Later among the works it cites.
Deepcore: A comprehensive library for coreset selection in deep learning
Chengcheng Guo, Bo Zhao, and Yanbing Bai. 2022 · 2022
Later among the works it cites.
Language models in the loop: Incorporating prompting into weak supervision
Ryan Smith, Jason Fries, Braden Hancock, and Stephen Bach. 2022 · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava and Abhinav Rastogi. 2022 · 2022
Later among the works it cites.
Simfluence: Modeling the influence of individual training examples by simulating training runs
Kelvin Guu, Albert Webson, Ellie Pavlick, Lucas Dixon, Ian Tenney, and Tolga Bolukbasi. 2023 · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Closest in time.
Self-instruct: Aligning language models with self-generated instructions
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023 · 2023
Closest in time.
Recogs: How incidental details of a logical form overshadow an evaluation of semantic interpretation
Zhengxuan Wu, Christopher Manning, and Christopher Potts. 2023 · 2023
Closest in time.
Data selection for language models via importance resampling
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy Liang. 2023 · 2023
Closest in time.