Fetching the paper…
Reading the bibliography…
Training on web-scale data can take months.
Backpropagation applied to handwritten zip code recognition
LeCun, Y., Boser, B., Denker, J. S., Henderson, D., Howard, R. E., Hubbard, W., and Jackel, L. D · 1989
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Word frequency distributions , volume 18
Baayen, R. H · 2001
Earlier work this paper cites.
Large scale online learning
Bottou, L. and LeCun, Y · 2004
Earlier work this paper cites.
Confidence-based active learning
Li, M. and Sethi, I. K · 2006
Earlier work this paper cites.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Active learning literature survey
Settles, B · 2009
Earlier work this paper cites.
Bayesian active learning for classification and preference learning
Houlsby, N., Huszár, F., Ghahramani, Z., and Lengyel, M · 2011
Earlier work this paper cites.
Efficient backprop
LeCun, Y. A., Bottou, L., Orr, G. B., and Müller, K.-R · 2012
Earlier work this paper cites.
Variance reduction in sgd by distributed importance sampling
Alain, G., Lamb, A., Sankar, C., Courville, A., and Bengio, Y · 2015
Earlier work this paper cites.
Weight uncertainty in neural network
Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Online batch selection for faster training of neural networks
Loshchilov, I. and Hutter, F · 2015
Earlier work this paper cites.
Schaul, T., Quan, J., Antonoglou, I., and Silver, D · 2015
Earlier work this paper cites.
Learning from massive noisy labeled data for image classification
Xiao, T., Xia, T., Yang, Y., Huang, C., and Wang, X · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2016
Earlier work this paper cites.
In-datacenter performance analysis of a tensor processing unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al · 2017
Earlier work this paper cites.
Biased importance sampling for deep neural network training
Katharopoulos, A. and Fleuret, F · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Cited alongside, same era.
Deep learning is robust to massive label noise
Rolnick, D., Veit, A., Belongie, S., and Shavit, N · 2017
Cited alongside, same era.
Large scale distributed neural network training through online distillation
Anil, R., Pereyra, G., Passos, A., Ormandi, R., Dahl, G. E., and Hinton, G. E · 2018
Cited alongside, same era.
Cinic-10 is not imagenet or cifar-10
Darlow, L. N., Crowley, E. J., Antoniou, A., and Storkey, A. J · 2018
Cited alongside, same era.
Training deep models faster with robust, approximate importance sampling
Johnson, T. B. and Guestrin, C · 2018
Coresets via bilevel optimization for continual learning and streaming
Borsos, Z., Mutnỳ, M., and Krause, A · 2020
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
Selection via proxy: Efficient data selection for deep learning
Coleman, C., Yeh, C., Mussmann, S., Mirzasoleiman, B., Bailis, P., Liang, P., Leskovec, J., and Zaharia, M · 2020
Later among the works it cites.
Ordered sgd: A new stochastic optimization framework for empirical risk minimization
Kawaguchi, K. and Lu, H · 2020
Later among the works it cites.
Glister: Generalization based data subset selection for efficient and robust learning
Killamsetty, K., Sivasubramanian, D., Ramakrishnan, G., and Iyer, R · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Not all samples are created equal: Deep learning with importance sampling
Katharopoulos, A. and Fleuret, F · 2018
Cited alongside, same era.
An empirical model of large-batch training
McCandlish, S., Kaplan, J., Amodei, D., and Team, O. D · 2018
Cited alongside, same era.
Learning to reweight examples for robust deep learning
Ren, M., Zeng, W., Yang, B., and Urtasun, R · 2018
Cited alongside, same era.
An empirical study of example forgetting during deep neural network learning
Toneva, M., Sordoni, A., Combes, R. T. d., Trischler, A., Bengio, Y., and Gordon, G. J · 2018
Cited alongside, same era.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Cited alongside, same era.
Understanding and utilizing deep neural networks trained with noisy labels
Chen, P., Liao, B. B., Chen, G., and Zhang, S · 2019
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Cited alongside, same era.
Later among the works it cites.
Coresets for data-efficient training of machine learning models
Mirzasoleiman, B., Bilmes, J., and Leskovec, J · 2020
Later among the works it cites.
Identifying mislabeled data using the area under the margin ranking
Pleiss, G., Zhang, T., Elenberg, E., and Weinberger, K. Q · 2020
Later among the works it cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Later among the works it cites.
Bayesian deep learning and a probabilistic perspective of generalization
Wilson, A. G. and Izmailov, P · 2020
Later among the works it cites.
Image classification with deep learning in the presence of noisy labels: A survey
Algan, G. and Ulusoy, I · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
On statistical bias in active learning: How and when to fix it
Farquhar, S., Gal, Y., and Rainforth, T · 2021
Later among the works it cites.
Active learning under pool set distribution shift and noisy data
Kirsch, A., Rainforth, T., and Gal, Y · 2021
Later among the works it cites.
A survey on bias and fairness in machine learning
Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., and Galstyan, A · 2021
Later among the works it cites.
Mukhoti, J., Kirsch, A., van Amersfoort, J., Torr, P. H., and Gal, Y · 2021
Later among the works it cites.
Deep learning on a data diet: Finding important examples early in training
Paul, M., Ganguli, S., and Dziugaite, G. K · 2021
Later among the works it cites.
huyvnphan/pytorch_cifar10
Phan, H · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision, 2021
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Later among the works it cites.
Divide and contrast: Self-supervised learning from uncurated data
Tian, Y., Henaff, O. J., and Oord, A. v. d · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2022
Closest in time.