Fetching the paper…
Reading the bibliography…
We initiate a formal study of reproducibility in optimization.
Problem complexity and method efficiency in optimization
A. S. Nemirovski and D. B. Yudin · 1983
Earlier work this paper cites.
On characterizations of the input-to-state stability property
E. D. Sontag and Y. Wang · 1995
Earlier work this paper cites.
On smooth activation functions
H. Mhaskar · 1997
Earlier work this paper cites.
Ensemble methods in machine learning
T. G. Dietterich · 2000
Earlier work this paper cites.
Algorithmic stability and generalization performance
O. Bousquet and A. Elisseeff · 2001
Earlier work this paper cites.
Almost-everywhere algorithmic stability and generalization error
S. Kutin and P. Niyogi · 2002
Earlier work this paper cites.
Minimum variance in biased estimation: Bounds and asymptotically optimal estimators
Y. C. Eldar · 2004
Earlier work this paper cites.
Testing statistical hypotheses
E. Lehmann and J. Romano · 2005
Earlier work this paper cites.
Calibrating noise to sensitivity in private data analysis
C. Dwork, F. McSherry, K. Nissim, and A. D. Smith · 2006
Earlier work this paper cites.
Mechanism design via differential privacy
F. McSherry and K. Talwar · 2007
Earlier work this paper cites.
Smooth optimization with approximate gradient
A. d’Aspremont · 2008
Earlier work this paper cites.
Rectified linear units improve restricted Boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Convex analysis and monotone operator theory in Hilbert spaces , volume 408
H. H. Bauschke, P. L. Combettes, et al · 2011
Earlier work this paper cites.
Convex optimization: algorithms and complexity
S. Bubeck · 2014
Earlier work this paper cites.
First-order methods of smooth convex optimization with inexact oracle
O. Devolder, F. Glineur, and Y. Nesterov · 2014
Earlier work this paper cites.
Critical learning periods in deep neural networks
A. Achille, M. Rovere, and S. Soatto · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
B. Lakshminarayanan, A. Pritzel, and C. Blundell · 2017
Cited alongside, same era.
Large scale distributed neural network training through online distillation
R. Anil, G. Pereyra, A. Passos, R. Ormandi, G. E. Dahl, and G. E. Hinton · 2018
Cited alongside, same era.
On acceleration with noise-corrupted gradients
M. Cohen, J. Diakonikolas, and L. Orecchia · 2018
Cited alongside, same era.
Lectures on convex optimization , volume 137
Y. Nesterov · 2018
Cited alongside, same era.
Analyzing the role of model uncertainty for electronic health records
M. W. Dusenberry, D. Tran, E. Choi, J. Kemp, J. Nixon, G. Jerfel, K. Heller, and A. M. Dai · 2020
Later among the works it cites.
Deep ensembles: A loss landscape perspective
S. Fort, H. Hu, and B. Lakshminarayanan · 2020
Later among the works it cites.
Linear mode connectivity and the lottery ticket hypothesis
J. Frankle, G. K. Dziugaite, D. Roy, and M. Carbin · 2020
Later among the works it cites.
Anti-distillation: Improving reproducibility of deep networks
G. I. Shamir and L. Coviello · 2020
Later among the works it cites.
Smooth activations and reproducibility in deep networks
G. I. Shamir, D. Lin, and L. Coviello · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. J. Shallue, J. Lee, J. Antognini, J. Sohl-Dickstein, R. Frostig, and G. E. Dahl · 2018
Cited alongside, same era.
Systems and methods for improved generalization, reproducibility, and stabilization of neural networks via error control code constraints, 2018
G. I. Shamir · 2018
Cited alongside, same era.
Potential-function proofs for gradient methods
N. Bansal and A. Gupta · 2019
Cited alongside, same era.
Gradient Descent for Non-convex Problems in Modern Machine Learning
S. Du · 2019
Cited alongside, same era.
The complexity of making the gradient small in stochastic convex optimization
D. J. Foster, A. Sekhari, O. Shamir, N. Srebro, K. Sridharan, and B. Woodworth · 2019
Cited alongside, same era.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Z. Allen-Zhu and Y. Li · 2020
Cited alongside, same era.
Robust accelerated gradient methods for smooth strongly convex functions
N. S. Aybat, A. Fallah, M. Gurbuzbalaban, and A. Ozdaglar · 2020
Cited alongside, same era.
Algorithmic instabilities of accelerated gradient descent
A. Attia and T. Koren · 2021
Later among the works it cites.
On the reproducibility of neural network predictions
S. Bhojanapalli, K. Wilber, A. Veit, A. S. Rawat, S. Kim, A. Menon, and S. Kumar · 2021
Later among the works it cites.
Improving reproducibility in machine learning research: a report from the NeurIPS 2019 reproducibility program
J. Pineau, P. Vincent-Lamarre, K. Sinha, V. Larivière, A. Beygelzimer, F. d’Alché Buc, E. Fox, and H. Larochelle · 2021
Later among the works it cites.
Synthesizing irreproducibility in deep networks
R. R. Snapp and G. I. Shamir · 2021
Later among the works it cites.
On nondeterminism and instability in neural network optimization, 2021
C. Summers and M. J. Dinneen · 2021
Later among the works it cites.
Dropout prediction variation estimation using neuron activation strength
H. Yu, Z. Chen, D. Lin, G. Shamir, and J. Han · 2021
Later among the works it cites.
Randomness in neural network training: Characterizing the impact of tooling
D. Zhuang, X. Zhang, S. L. Song, and S. Hooker · 2021
Later among the works it cites.
R. Impagliazzo, R. Lei, T. Pitassi, and J. Sorrell · 2022
Closest in time.
On the sample complexity of stability constrained imitation learning
S. Tu, A. Robey, T. Zhang, and N. Matni · 2022
Closest in time.