Fetching the paper…
Reading the bibliography…
Previous work has proposed many new loss functions and regularizers that improve test accuracy on image classification tasks.
Updating quasi-newton matrices with limited storage
J. Nocedal · 1980
Earlier work this paper cites.
For valid generalization, the size of the weights is more important than the size of the network
P. L. Bartlett · 1997
Earlier work this paper cites.
On kernel-target alignment
N. Cristianini, J. Shawe-Taylor, A. Elisseeff, and J. S. Kandola · 2002
Earlier work this paper cites.
Automated flower classification over a large number of classes
M.-E. Nilsback and A. Zisserman · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
An investigation of how label smoothing affects generalization
B. Chen, L. Ziyin, Z. Wang, and P. P. Liang · 2010
Earlier work this paper cites.
Sun database: Large-scale scene recognition from abbey to zoo
J. Xiao, J. Hays, K. A. Ehinger, A. Oliva, and A. Torralba · 2010
Earlier work this paper cites.
Algorithms for learning kernels based on centered alignment
C. Cortes, M. Mohri, and A. Rostamizadeh · 2012
Earlier work this paper cites.
Cats and dogs
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. Jawahar · 2012
Earlier work this paper cites.
Collecting a large-scale dataset of fine-grained cars
J. Krause, J. Deng, M. Stark, and L. Fei-Fei · 2013
Earlier work this paper cites.
Dropout training as adaptive regularization
S. Wager, S. Wang, and P. S. Liang · 2013
Earlier work this paper cites.
Birdsnap: Large-scale fine-grained visual categorization of birds
T. Berg, J. Liu, S. W. Lee, M. L. Alexander, D. W. Jacobs, and P. N. Belhumeur · 2014
Earlier work this paper cites.
Food-101 — mining discriminative components with random forests
L. Bossard, M. Guillaumin, and L. Van Gool · 2014
Earlier work this paper cites.
Decaf: A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell · 2014
Earlier work this paper cites.
Cnn features off-the-shelf: an astounding baseline for recognition
A. Sharif Razavian, H. Azizpour, J. Sullivan, and S. Carlsson · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Altitude training: Strong bounds for single-layer dropout
S. Wager, W. Fithian, S. Wang, and P. S. Liang · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
J. Yosinski, J. Clune, Y. Bengio, and H. Lipson · 2014
Earlier work this paper cites.
Norm-based capacity control in neural networks
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, et al · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
G. Alain and Y. Bengio · 2016
Earlier work this paper cites.
Training and investigating residual nets
S. Gross and M. Wilber · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
On loss functions for deep neural networks in classification
K. Janocha and W. M. Czarnecki · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Earlier work this paper cites.
Spectrally-normalized margin bounds for neural networks
P. L. Bartlett, D. J. Foster, and M. Telgarsky · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: training ImageNet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Earlier work this paper cites.
On calibration of modern neural networks
C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger · 2017
Cited alongside, same era.
Sphereface: Deep hypersphere embedding for face recognition
W. Liu, Y. Wen, Z. Yu, M. Li, B. Raj, and L. Song · 2017
Cited alongside, same era.
SGDR: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2017
Cited alongside, same era.
Regularizing neural networks by penalizing confident output distributions
G. Pereyra, G. Tucker, J. Chorowski, Ł. Kaiser, and G. Hinton · 2017
Cited alongside, same era.
L2-constrained softmax loss for discriminative face verification
R. Ranjan, C. D. Castillo, and R. Chellappa · 2017
Cited alongside, same era.
Do ImageNet classifiers generalize to ImageNet?
B. Recht, R. Roelofs, L. Schmidt, and V. Shankar · 2019
Later among the works it cites.
On the information bottleneck theory of deep learning
A. M. Saxe, Y. Bansal, J. Dapello, M. Advani, A. Kolchinsky, B. D. Tracey, and D. D. Cox · 2019
Later among the works it cites.
Learning robust global representations by penalizing local predictive power
H. Wang, S. Ge, E. P. Xing, and Z. C. Lipton · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
C. Wei, J. D. Lee, Q. Liu, and T. Ma · 2019
Later among the works it cites.
L. Beyer, O. J. Hénaff, A. Kolesnikov, X. Zhai, and A. v. d. Oord · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Shwartz-Ziv and N. Tishby · 2017
Cited alongside, same era.
Prototypical networks for few-shot learning
J. Snell, K. Swersky, and R. Zemel · 2017
Cited alongside, same era.
Normface: L2 hypersphere embedding for face verification
F. Wang, X. Xiang, J. Cheng, and A. L. Yuille · 2017
Cited alongside, same era.
Estimating information flow in deep neural networks
Z. Goldfeld, E. v. d. Berg, K. Greenewald, I. Melnyk, N. Nguyen, B. Kingsbury, and Y. Polyanskiy · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
S. Gunasekar, J. D. Lee, D. Soudry, and N. Srebro · 2018
Cited alongside, same era.
A pac-bayesian approach to spectrally-normalized margin bounds for neural networks
B. Neyshabur, S. Bhojanapalli, and N. Srebro · 2018
Cited alongside, same era.
Tadam: Task dependent adaptive metric for improved few-shot learning
B. N. Oreshkin, P. Rodriguez, and A. Lacoste · 2018
Cited alongside, same era.
Separability and geometry of object manifolds in deep neural networks
U. Cohen, S. Chung, D. D. Lee, and H. Sompolinsky · 2020
Closest in time.
Crosstransformers: spatially-aware few-shot transfer
C. Doersch, A. Gupta, and A. Zisserman · 2020
Closest in time.
Unraveling meta-learning: Understanding feature representations for few-shot tasks
M. Goldblum, S. Reich, L. Fowl, R. Ni, V. Cherepanova, and T. Goldstein · 2020
Closest in time.
The many faces of robustness: A critical analysis of out-of-distribution generalization
D. Hendrycks, S. Basart, N. Mu, S. Kadavath, F. Wang, E. Dorundo, R. Desai, T. Zhu, S. Parajuli, M. Guo, et al · 2020
Closest in time.
Supervised contrastive learning
P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Isola, A. Maschinot, C. Liu, and D. Krishnan · 2020
Closest in time.
Negative margin matters: Understanding margin in few-shot classification
B. Liu, Y. Cao, Y. Lin, Q. Li, Z. Zhang, M. Long, and H. Hu · 2020
Closest in time.
Does label smoothing mitigate label noise?
M. Lukasik, S. Bhojanapalli, A. Menon, and S. Kumar · 2020
Closest in time.
Individual differences among deep neural network models
J. Mehrer, C. J. Spoerer, N. Kriegeskorte, and T. C. Kietzmann · 2020
Closest in time.
Generalized entropy regularization or: There’s nothing special about label smoothing
C. Meister, E. Salesky, and R. Cotterell · 2020
Closest in time.
On convergence and generalization of dropout training
P. Mianjy and R. Arora · 2020
Closest in time.
What is being transferred in transfer learning?
B. Neyshabur, H. Sedghi, and C. Zhang · 2020
Closest in time.
Prevalence of neural collapse during the terminal phase of deep learning training
V. Papyan, X. Han, and D. L. Donoho · 2020
Closest in time.
Revisiting training strategies and generalization performance in deep metric learning
K. Roth, T. Milbich, S. Sinha, P. Gupta, B. Ommer, and J. P. Cohen · 2020
Closest in time.
Measuring robustness to natural distribution shifts in image classification
R. Taori, A. Dave, V. Shankar, N. Carlini, B. Recht, and L. Schmidt · 2020
Closest in time.
Rethinking few-shot image classification: a good embedding is all you need?
Y. Tian, Y. Wang, D. Krishnan, J. B. Tenenbaum, and P. Isola · 2020
Closest in time.
The implicit and explicit regularization effects of dropout
C. Wei, S. Kakade, and T. Ma · 2020
Closest in time.
Towards understanding label smoothing
Y. Xu, Y. Xu, Q. Qian, H. Li, and R. Jin · 2020
Closest in time.
Exploring the limits of large scale pre-training
S. Abnar, M. Dehghani, B. Neyshabur, and H. Sedghi · 2021
Closest in time.
Deconstructing the regularization of batchnorm
Y. Dauphin and E. D. Cubuk · 2021
Closest in time.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
L. Hui and M. Belkin · 2021
Closest in time.
Do wide and deep networks learn the same things? uncovering how neural network representations vary with width and depth
T. Nguyen, M. Raghu, and S. Kornblith · 2021
Closest in time.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Closest in time.
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2021
Closest in time.