Fetching the paper…
Reading the bibliography…
In machine learning, we traditionally evaluate the performance of a single model, averaged over a collection of test inputs.
On the power of curriculum learning in training deep networks
Hacohen, G. and Weinshall, D · 1904
Earlier work this paper cites.
SGD on neural networks learns functions of increasing complexity
Nakkiran, P., Kaplun, G., Kalimeris, D., Yang, T., Edelman, B. L., Zhang, F., and Barak, B · 1905
Earlier work this paper cites.
A constructive prediction of the generalization error across scales
Rosenfeld, J. S., Rosenfeld, A., Belinkov, Y., and Shavit, N · 1909
Earlier work this paper cites.
Deep double descent: Where bigger models and more data hurt
Nakkiran, P., Kaplun, G., Bansal, Y., Yang, T., Barak, B., and Sutskever, I · 1912
Earlier work this paper cites.
Sample Selection Bias as a Specification Error (With an Application to the Estimation of Labor Supply Functions), 1977
Heckman, J. J · 1977
Earlier work this paper cites.
Forest before trees: The precedence of global features in visual perception
Navon, D · 1977
Earlier work this paper cites.
A theory of the learnable
Valiant, L. G · 1984
Earlier work this paper cites.
Improving Predictive Inference Under Covariate Shift by Weighting the Log-Likelihood Function
Shimodaira, H · 2000
Earlier work this paper cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
Identifying mislabeled data using the area under the margin ranking
Pleiss, G., Zhang, T., Elenberg, E. R., and Weinberger, K. Q · 2001
Earlier work this paper cites.
Exploring the memorization-generalization continuum in deep learning
Jiang, Z., Zhang, C., Talwar, K., and Mozer, M. C · 2002
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Estimating example difficulty using variance of gradients
Agarwal, C. and Hooker, S · 2008
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2009
Earlier work this paper cites.
Introduction to nonparametric estimation., 2009
Tsybakov, A. B · 2009
Earlier work this paper cites.
A theory of learning from different domains
Ben-David, S., Blitzer, J., Crammer, K., Kulesza, A., Pereira, F., and Vaughan, J. W · 2010
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., Hallacy, C., Mann, B., Radford, A., Ramesh, A., Ryder, N., Ziegler, D. M., Schulman, J., Amodei, D., and McCandlish, S · 2010
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Lakshminarayanan, B., Pritzel, A., and Blundell, C · 2016
Cited alongside, same era.
Training region-based object detectors with online hard example mining
Shrivastava, A., Gupta, A., and Girshick, R · 2016
Cited alongside, same era.
A downsampled variant of imagenet as an alternative to the cifar datasets
Chrabaszcz, P., Loshchilov, I., and Hutter, F · 2017
Cited alongside, same era.
Selective classification for deep neural networks
Geifman, Y. and El-Yaniv, R · 2017
Cited alongside, same era.
An analysis of machine learning intelligence
Lalor, J. P., Wu, H., Munkhdalai, T., and Yu, H · 2017
Pretrained transformers improve out-of-distribution robustness
Hendrycks, D., Liu, X., Wallace, E., Dziedzic, A., Krishnan, R., and Song, D · 2020
Later among the works it cites.
Racial disparities in automated speech recognition
Koenecke, A., Nam, A., Lake, E., Nudell, J., Quartey, M., Mengesha, Z., Toups, C., Rickford, J. R., Jurafsky, D., and Goel, S · 2020
Later among the works it cites.
The deep bootstrap framework: Good online learners are good offline generalizers
Nakkiran, P., Neyshabur, B., and Sedghi, H · 2020
Later among the works it cites.
Why are bootstrapped deep ensembles not better?
Nixon, J., Lakshminarayanan, B., and Tran, D · 2020
Later among the works it cites.
A neural scaling law from the dimension of the data manifold
Sharma, U. and Kaplan, J · 2020
Later among the works it cites.
Hybrid models for open set recognition
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Enhancing the reliability of out-of-distribution image detection in neural networks
Liang, S., Li, Y., and Srikant, R · 2017
Cited alongside, same era.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Buolamwini, J. and Gebru, T · 2018
Cited alongside, same era.
CINIC-10 is not imagenet or CIFAR-10
Darlow, L. N., Crowley, E. J., Antoniou, A., and Storkey, A. J · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
To trust or not to trust a classifier
Jiang, H., Kim, B., Guan, M. Y., and Gupta, M. R · 2018
Cited alongside, same era.
Detecting and Correcting for Label Shift with Black Box Predictors
Lipton, Z. C., Wang, Y.-X., and Smola, A · 2018
Cited alongside, same era.
An empirical study of example forgetting during deep neural network learning
Toneva, M., Sordoni, A., des Combes, R. T., Trischler, A., Bengio, Y., and Gordon, G. J · 2018
Cited alongside, same era.
Zhang, H., Li, A., Guo, J., and Guo, Y · 2020
Later among the works it cites.
Explaining neural scaling laws
Bahri, Y., Dyer, E., Kaplan, J., Lee, J., and Sharma, U · 2021
Later among the works it cites.
Deep learning through the lens of example difficulty
Baldock, R. J. N., Maennel, H., and Neyshabur, B · 2021
Later among the works it cites.
Are labels always necessary for classifier accuracy evaluation?
Deng, W. and Zheng, L · 2021
Later among the works it cites.
The three stages of learning dynamics in high-dimensional kernel methods, 2021
Ghosh, N., Mei, S., and Yu, B · 2021
Later among the works it cites.
No one representation to rule them all: Overlapping features of training methods
Gontijo-Lopes, R., Dauphin, Y., and Cubuk, E. D · 2021
Later among the works it cites.
Predicting with confidence on unseen distributions
Guillory, D., Shankar, V., Ebrahimi, S., Darrell, T., and Schmidt, L · 2021
Later among the works it cites.
Accuracy on the line: on the strong correlation between out-of-distribution and in-distribution generalization
Miller, J. P., Taori, R., Raghunathan, A., Sagawa, S., Koh, P. W., Shankar, V., Liang, P., Carmon, Y., and Schmidt, L · 2021
Later among the works it cites.
Pervasive label errors in test sets destabilize machine learning benchmarks, 2021
Northcutt, C. G., Athalye, A., and Mueller, J · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Later among the works it cites.
Are larger pretrained language models uniformly better? comparing performance at the instance level
Zhong, R., Ghosh, D., Klein, D., and Steinhardt, J · 2021
Later among the works it cites.
Leveraging unlabeled data to predict out-of-distribution performance
Garg, S., Balakrishnan, S., Lipton, Z., Neyshabur, B., and Sedghi, H · 2022
Closest in time.
Distributional generalization: Structure beyond test error, 2022
Nakkiran, P. and Bansal, Y · 2022
Closest in time.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2022
Closest in time.