Fetching the paper…
Reading the bibliography…
Data pruning algorithms are commonly used to reduce the memory and computational cost of the optimization process.
Selection via proxy: Efficient data selection for deep learning
Coleman, C., Yeh, C., Mussmann, S., Mirzasoleiman, B., Bailis, P., Liang, P., Leskovec, J., and Zaharia, M. (2019) · 1906
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S. (1951) · 1951
Earlier work this paper cites.
On the convergence of sample probability distributions
Varadarajan, V. S. (1958) · 1960
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Hornik, K. (1991) · 1991
Earlier work this paper cites.
Fundamentals of general topology: Problems and exercises
Arkhangel’skiǐ, A. V. (2001) · 2001
Earlier work this paper cites.
Herding dynamical weights to learn
Welling, M. (2009) · 2009
Earlier work this paper cites.
Nonlinear functional analysis
Deimling, K. (2010) · 2010
Earlier work this paper cites.
Scalable training of mixture models via coresets
Feldman, D., Faulkner, M., and Krause, A. (2011) · 2011
Earlier work this paper cites.
Super-samples from kernel herding
Chen, Y., Welling, M., and Smola, A. (2012) · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Coresets for nonparametric estimation-the case of dp-means
Bachem, O., Lucic, M., and Krause, A. (2015) · 2015
Earlier work this paper cites.
Coresets for scalable bayesian logistic regression
Huggins, J., Campbell, T., and Broderick, T. (2016) · 2016
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M. M. A., Yang, Y., and Zhou, Y. (2017) · 2017
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
Sener, O. and Savarese, S. (2017) · 2017
Earlier work this paper cites.
Adversarial active learning for deep networks: a margin based approach
Ducoffe, M. and Precioso, F. (2018) · 2018
Cited alongside, same era.
On coresets for logistic regression
Munteanu, A., Schwiegelshohn, C., Sohler, C., and Woodruff, D. P. (2018) · 2018
Cited alongside, same era.
An empirical study of example forgetting during deep neural network learning
Toneva, M., Sordoni, A., Combes, R. T. d., Trischler, A., Bengio, Y., and Gordon, G. J. (2018) · 2018
Cited alongside, same era.
Wang, T., Zhu, J.-Y., Torralba, A., and Efros, A. A. (2018) · 2018
Cited alongside, same era.
Gradient based sample selection for online continual learning
Aljundi, R., Lin, M., Goujaud, B., and Bengio, Y. (2019) · 2019
Cited alongside, same era.
Scaling laws for transfer
Hernandez, D., Kaplan, J., Henighan, T., and McCandlish, S. (2021) · 2021
Later among the works it cites.
Kaushal, V., Kothawade, S., Ramakrishnan, G., Bilmes, J., and Iyer, R. (2021) · 2021
Later among the works it cites.
Similar: Submodular information measures based active learning in realistic scenarios
Kothawade, S., Beck, N., Killamsetty, K., and Iyer, R. (2021) · 2021
Later among the works it cites.
Active learning by acquiring contrastive examples
Margatina, K., Vernikos, G., Barrault, L., and Aletras, N. (2021) · 2021
Later among the works it cites.
Deep learning on a data diet: Finding important examples early in training
Paul, M., Ganguli, S., and Dziugaite, G. K. (2021) · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Automated scalable bayesian inference via hilbert coresets
Campbell, T. and Broderick, T. (2019) · 2019
Cited alongside, same era.
Contextual diversity for active learning
Agarwal, S., Arora, H., Anand, S., and Arora, C. (2020) · 2020
Cited alongside, same era.
Selection via proxy: Efficient data selection for deep learning
Coleman, C., Yeh, C., Mussmann, S., Mirzasoleiman, B., Bailis, P., Liang, P., Leskovec, J., and Zaharia, M. (2020) · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020) · 2020
Cited alongside, same era.
Universal approximation with deep narrow networks
Kidger, P. and Lyons, T. (2020) · 2020
Cited alongside, same era.
Coresets for data-efficient training of machine learning models
Mirzasoleiman, B., Bilmes, J., and Leskovec, J. (2020) · 2020
Cited alongside, same era.
A constructive prediction of the generalization error across scales
Rosenfeld, J. S., Rosenfeld, A., Belinkov, Y., and Shavit, N. (2020) · 2020
Cited alongside, same era.
Later among the works it cites.
Dataset condensation with differentiable siamese augmentation
Zhao, B. and Bilen, H. (2021) · 2021
Later among the works it cites.
Dataset condensation with gradient matching
Zhao, B., Mopuri, K. R., and Bilen, H. (2021) · 2021
Later among the works it cites.
The curse of (non)convexity: The case of an optimization-inspired data pruning algorithm
Ayed, F. and Hayou, S. (2022) · 2022
Later among the works it cites.
Deepcore: A comprehensive library for coreset selection in deep learning
Guo, C., Zhao, B., and Bai, Y. (2022) · 2022
Later among the works it cites.
An empirical analysis of compute-optimal large language model training
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J. W., and Sifre, L. (2022) · 2022
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
Sorscher, B., Geirhos, R., Shekhar, S., Ganguli, S., and Morcos, A. S. (2022) · 2022
Later among the works it cites.
Scaling vision transformers
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L. (2022) · 2022
Later among the works it cites.
Universality of deep convolutional neural networks
Zhou, D.-X. (2020) · 2022
Later among the works it cites.