Fetching the paper…
Reading the bibliography…
The Lipschitz constant is an important quantity that arises in analysing the convergence of gradient-based optimization methods.
Neural stochastic differential equations: Deep latent gaussian models in the diffusion limit
Tzen, B. and Raginsky, M · 1905
Earlier work this paper cites.
Neural SDE: stabilizing neural ODE networks with stochastic noise
Liu, X., Xiao, T., Si, S., Cao, Q., Kumar, S., and Hsieh, C · 1906
Earlier work this paper cites.
Latent odes for irregularly-sampled time series
Rubanova, Y., Chen, R. T. Q., and Duvenaud, D · 1907
Earlier work this paper cites.
Deep neural networks, generic universal interpolation, and controlled ODEs, 2019
Cuchiero, C., Larsson, M., and Teichmann, J · 1908
Earlier work this paper cites.
Stochastic integration and differential equations
Protter, P · 1992
Earlier work this paper cites.
MNIST Handwritten Digit Database
LeCun, Y. and Cortes, C · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming, 2013
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Analysis 2
Königsberger, K · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Nesterov, Y · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Earlier work this paper cites.
Stochastic calculus and applications
Cohen, S. N. and Elliott, R. J · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y · 2015
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Ge, R., Lee, J. D., and Ma, T · 2016
Cited alongside, same era.
Natasha: Faster non-convex stochastic optimization via strongly non-convex parameter, 2017
Allen-Zhu, Z · 2017
Cited alongside, same era.
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M. J · 2017
Cited alongside, same era.
Non-convex finite-sum optimization via scsg methods, 2017
Lei, L., Ju, C., Chen, J., and Jordan, M. I · 2017
Cited alongside, same era.
On the state of the art of evaluation in neural language models
Adagrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization, 2018
Ward, R., Wu, X., and Bottou, L · 2018
Later among the works it cites.
Stochastic nested variance reduction for nonconvex optimization, 2018
Zhou, D., Xu, P., and Gu, Q · 2018
Later among the works it cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q · 2018
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z · 2019
Later among the works it cites.
Low-rank plus sparse decomposition of covariance matrices using neural network parametrization, 2019
Baes, M., Herrera, C., Neufeld, A., and Ruyssen, P · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Melis, G., Dyer, C., and Blunsom, P · 2017
Cited alongside, same era.
Stronger generalization bounds for deep nets via a compression approach
Arora, S., Ge, R., Neyshabur, B., and Zhang, Y · 2018
Cited alongside, same era.
Neural ordinary differential equations
Chen, T. Q., Rubanova, Y., Bettencourt, J., and Duvenaud, D. K · 2018
Cited alongside, same era.
Spider: Near-optimal non-convex optimization via stochastic path integrated differential estimator, 2018
Fang, C., Li, C. J., Lin, Z., and Zhang, T · 2018
Cited alongside, same era.
Stability-certified reinforcement learning: A control-theoretic perspective
Jin, M. and Lavaei, J · 2018
Cited alongside, same era.
On tighter generalization bound for deep neural networks: Cnns, resnets, and beyond
Li, X., Lu, J., Wang, Z., Haupt, J., and Zhao, T · 2018
Cited alongside, same era.
Certified defenses against adversarial examples
Raghunathan, A., Steinhardt, J., and Liang, P · 2018
Cited alongside, same era.
GRU-ODE-Bayes: Continuous modeling of sporadically-observed time series
Brouwer, E. D., Simm, J., Arany, A., and Moreau, Y · 2019
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y. and Gu, Q · 2019
Later among the works it cites.
Lipschitz certificates for neural network structures driven by averaged activation operators
Combettes, P. L. and Pesquet, J.-C · 2019
Later among the works it cites.
Efficient and accurate estimation of lipschitz constants for deep neural networks
Fazlyab, M., Robey, A., Hassani, H., Morari, M., and Pappas, G. J · 2019
Later among the works it cites.
Neural jump stochastic differential equations
Jia, J. and Benson, A. R · 2019
Later among the works it cites.
On the convergence of stochastic gradient descent with adaptive stepsizes
Li, X. and Orabona, F · 2019
Later among the works it cites.
Infinitely deep neural networks as diffusion processes
Peluchetti, S. and Favaro, S · 2019
Later among the works it cites.
Lipschitz constant estimation for neural networks via sparse polynomial optimization
Latorre, F., Rolland, P., and Cevher, V · 2020
Closest in time.