Fetching the paper…
Reading the bibliography…
Estimating the difficulty of a dataset typically involves comparing state-of-the-art models to humans; the bigger the performance gap, the harder the dataset is said to be.
Towards efficient data valuation based on the shapley value
Jia, R., Dao, D., Wang, B., Hubis, F. A., Hynes, N., Gürel, N. M., Li, B., Zhang, C., Song, D., and Spanos, C. J · 1902
Earlier work this paper cites.
Unsupervised label noise modeling and loss correction
Arazo, E., Ortego, D., Albert, P., O’Connor, N., and McGuinness, K · 1904
Earlier work this paper cites.
Data shapley: Equitable valuation of data for machine learning
Ghorbani, A. and Zou, J · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach, 2019
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Verified uncertainty calibration
Kumar, A., Liang, P., and Ma, T · 1909
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2019
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 1910
Earlier work this paper cites.
A mathematical theory of communication
Shannon, C. E · 1948
Earlier work this paper cites.
Heterogeneous uncertainty sampling for supervised learning
Lewis, D. D. and Catlett, J · 1994
Earlier work this paper cites.
A sequential algorithm for training text classifiers
Lewis, D. D. and Gale, W. A · 1994
Earlier work this paper cites.
Text classification from labeled and unlabeled documents using EM
Nigam, K., McCallum, A. K., Thrun, S., and Mitchell, T · 2000
Earlier work this paper cites.
Identifying mislabeled data using the area under the margin ranking
Pleiss, G., Zhang, T., Elenberg, E. R., and Weinberger, K. Q · 2001
Earlier work this paper cites.
Dodge, J., Ilharco, G., Schwartz, R., Farhadi, A., Hajishirzi, H., and Smith, N · 2002
Earlier work this paper cites.
Adversarial filters of dataset biases
Le Bras, R., Swayamdipta, S., Bhagavatula, C., Zellers, R., Peters, M., Sabharwal, A., and Choi, Y · 2002
Earlier work this paper cites.
On issues of instance selection
Liu, H. and Motoda, H · 2002
Earlier work this paper cites.
A theory of usable information under computational constraints
Xu, Y., Zhao, S., Song, J., Stewart, R., and Ermon, S · 2002
Earlier work this paper cites.
On the stability of fine-tuning bert: Misconceptions, explanations, and strong baselines
Mosbach, M., Andriushchenko, M., and Klakow, D · 2006
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J · 2009
Earlier work this paper cites.
From baby steps to leapfrog: How “less is more” in unsupervised dependency parsing
Spitkovsky, V. I., Alshawi, H., and Jurafsky, D · 2010
Earlier work this paper cites.
Unbiased look at dataset bias
Torralba, A. and Efros, A. A · 2011
Earlier work this paper cites.
Item response theory
Embretson, S. E. and Reise, S. P · 2013
Earlier work this paper cites.
A survey on instance selection for active learning
Fu, Y., Zhu, X., and Li, B · 2013
Earlier work this paper cites.
Learning whom to trust with MACE
Hovy, D., Berg-Kirkpatrick, T., Vaswani, A., and Hovy, E · 2013
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Bowman, S., Angeli, G., Potts, C., and Manning, C. D · 2015
Cited alongside, same era.
Active bias: Training more accurate neural networks by emphasizing high variance samples
Chang, H.-S., Learned-Miller, E., and McCallum, A · 2017
Cited alongside, same era.
Automated hate speech detection and the problem of offensive language
Davidson, T., Warmsley, D., Macy, M., and Weber, I · 2017
Cited alongside, same era.
Annotation artifacts in natural language inference data
Gururangan, S., Swayamdipta, S., Levy, O., Schwartz, R., Bowman, S., and Smith, N. A · 2017
Cited alongside, same era.
Do imagenet classifiers generalize to imagenet?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2019
Later among the works it cites.
The risk of racial bias in hate speech detection
Sap, M., Card, D., Gabriel, S., Choi, Y., and Smith, N. A · 2019
Later among the works it cites.
Learning with bad training data via iterative trimmed loss minimization
Shen, Y. and Sanghavi, S · 2019
Later among the works it cites.
Language (technology) is power: A critical survey of “bias” in NLP
Blodgett, S. L., Barocas, S., Daumé III, H., and Wallach, H · 2020
Later among the works it cites.
Utility is in the eye of the user: A critique of nlp leaderboards
Ethayarajh, K. and Jurafsky, D · 2020
Later among the works it cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
Honnibal, M. and Montani, I · 2017
Cited alongside, same era.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Cited alongside, same era.
Measuring and mitigating unintended bias in text classification
Dixon, L., Li, J., Sorensen, J., Thain, N., and Vasserman, L · 2018
Cited alongside, same era.
Co-teaching: Robust training of deep neural networks with extremely noisy labels
Han, B., Yao, Q., Yu, X., Niu, G., Xu, M., Hu, W., Tsang, I. W., and Sugiyama, M · 2018
Cited alongside, same era.
Understanding deep learning performance through an examination of test set difficulty: A psychometric case study
Lalor, J. P., Wu, H., Munkhdalai, T., and Yu, H · 2018
Cited alongside, same era.
Explanation in artificial intelligence: Insights from the social sciences
Miller, T · 2018
Cited alongside, same era.
Systematic error analysis of the stanford question answering dataset
Rondeau, M.-A. and Hazen, T. J · 2018
Cited alongside, same era.
Later among the works it cites.
What do models learn from question answering datasets?
Sen, P. and Saffari, A · 2020
Later among the works it cites.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
Swayamdipta, S., Schwartz, R., Lourie, N., Wang, Y., Hajishirzi, H., Smith, N. A., and Choi, Y · 2020
Later among the works it cites.
DIME: An Information-Theoretic difficulty measure for AI datasets
Zhang, P., Wang, H., Naik, N., Xiong, C., and Socher, R · 2020
Later among the works it cites.
Competency problems: On finding and removing artifacts in language data, 2021
Gardner, M., Merrill, W., Dodge, J., Peters, M. E., Ross, A., Singh, S., and Smith, N · 2021
Closest in time.
Conditional probing: measuring usable information beyond a baseline
Hewitt, J., Ethayarajh, K., Liang, P., and Manning, C. D · 2021
Closest in time.
Dynaboard: An evaluation-as-a-service platform for holistic next-generation benchmarking, 2021
Ma, Z., Ethayarajh, K., Thrush, T., Jain, S., Wu, L., Jia, R., Potts, C., Williams, A., and Kiela, D · 2021
Closest in time.
What context features can transformer language models use?
O’Connor, J. and Andreas, J · 2021
Closest in time.
Rissanen data analysis: Examining dataset characteristics via description length
Perez, E., Kiela, D., and Cho, K · 2021
Closest in time.
Combining feature and instance attribution to detect artifacts, 2021
Pezeshkpour, P., Jain, S., Singh, S., and Wallace, B. C · 2021
Closest in time.
A bayesian framework for information-theoretic probing
Pimentel, T. and Cotterell, R · 2021
Closest in time.
Evaluation examples are not equally informative: How should that change NLP leaderboards?
Rodriguez, P., Barrow, J., Hoyle, A. M., Lalor, J. P., Jia, R., and Boyd-Graber, J · 2021
Closest in time.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al · 2022
Closest in time.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Closest in time.
Hypothesis only baselines in natural language inference
Poliak, A., Naradowsky, J., Haldar, A., Rudinger, R., and Van Durme, B · 2023
Closest in time.