Fetching the paper…
Reading the bibliography…
Bregman proximal point algorithm (BPPA) has witnessed emerging machine learning applications, yet its theoretical understanding has been largely unexplored.
Distillation strategies for proximal policy optimization
Green, S · 1901
Earlier work this paper cites.
Li, Y · 1907
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovski, A. S · 1983
Earlier work this paper cites.
Nonlinear proximal point algorithms using bregman functions, with applications to convex programming
Eckstein, J · 1993
Earlier work this paper cites.
Proximal minimization methods with generalized bregman functions
Kiwiel, K. C · 1997
Earlier work this paper cites.
Error bounds for proximal point subproblems and associated inexact proximal point algorithms
Solodov, M. V · 2000
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L · 2002
Earlier work this paper cites.
Empirical margin distributions and bounding the generalization error of combined classifiers
Koltchinskii, V · 2002
Earlier work this paper cites.
The large learning rate phase of deep learning: the catapult mechanism
Lewkowycz, A · 2003
Earlier work this paper cites.
Boosting as a regularized path to a maximum margin classifier
Rosset, S · 2004
Earlier work this paper cites.
On the duality of strong convexity and strong smoothness: Learning applications and matrix regularization
Kakade, S · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Convergence rate of inexact proximal point methods with relative error criteria for convex optimization
Monteiro, R. D · 2010
Earlier work this paper cites.
Convergence of a proximal point method in the presence of computational errors in hilbert spaces
Zaslavski, A. J · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J · 2011
Earlier work this paper cites.
Robustness and generalization
Xu, H · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I · 2013
Cited alongside, same era.
Margins, shrinkage, and boosting
Telgarsky, M · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
Hinton, G · 2015
Cited alongside, same era.
Trust region policy optimization
Schulman, J · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K · 2016
Cited alongside, same era.
Relatively smooth convex optimization by first-order methods, and applications
Lu, H · 2018
Later among the works it cites.
Shufflenet v2: Practical guidelines for efficient cnn architecture design
Ma, N · 2018
Later among the works it cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M · 2018
Later among the works it cites.
Smith, L. N · 2018
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Smith, S. L · 2018
Later among the works it cites.
The implicit bias of gradient descent on separable data
Soudry, D · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
First-order methods in optimization
Beck, A · 2017
Cited alongside, same era.
Sgd learns over-parameterized networks that provably generalize on linearly separable data
Brutzkus, A · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P · 2017
Cited alongside, same era.
Learning without forgetting
Li, Z · 2017
Cited alongside, same era.
Tarvainen, A · 2017
Cited alongside, same era.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z · 2018
Cited alongside, same era.
Later among the works it cites.
Stochastic gradient/mirror descent: Minimax optimality and implicit regularization
Azizan, N · 2019
Later among the works it cites.
Bag of tricks for image classification with convolutional neural networks
He, T · 2019
Later among the works it cites.
Convergence of gradient descent on separable data
Nacson, M. S · 2019
Later among the works it cites.
Revisit knowledge distillation: a teacher-free framework
Yuan, L · 2019
Later among the works it cites.
Efficient meta learning via minibatch proximal update
Zhou, P · 2019
Later among the works it cites.
SMART: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization
Jiang, H · 2020
Later among the works it cites.
Implicit bias of gradient descent based adversarial training on separable data
Li, Y · 2020
Later among the works it cites.
Implicit regularization in matrix sensing via mirror descent
Wu, F · 2021
Closest in time.
Bregman proximal point algorithm revisited: a new inexact version and its variant
Yang, L · 2021
Closest in time.