Fetching the paper…
Reading the bibliography…
We propose the particle dual averaging (PDA) method, which generalizes the dual averaging method in convex optimization to the optimization over probability distributions with quantitative runtime guarantee.
Analysis of a two-layer neural network via displacement convexity
Javanmard, A., Mondelli, M., and Montanari, A. (2019) · 1901
Earlier work this paper cites.
Linearized two-layers neural networks in high dimension
Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A. (2019b) · 1904
Earlier work this paper cites.
Mean-field langevin dynamics and energy landscape of neural networks
Hu, K., Ren, Z., Siska, D., and Szpruch, L. (2019) · 1905
Earlier work this paper cites.
Nitanda, A., Chinot, G., and Suzuki, T. (2019) · 1905
Earlier work this paper cites.
A mean-field limit for certain deep neural networks
Araújo, D., Oliveira, R. I., and Yukimura, D. (2019) · 1906
Earlier work this paper cites.
Sparse optimization on measures with over-parameterized gradient descent
Chizat, L. (2019) · 1907
Earlier work this paper cites.
Ji, Z. and Telgarsky, M. (2019) · 1909
Earlier work this paper cites.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Bai, Y. and Lee, J. D. (2019) · 1910
Earlier work this paper cites.
Ensemble kalman sampling: Mean-field limit and convergence analysis
Ding, Z. and Li, Q. (2019) · 1910
Earlier work this paper cites.
Mean-field neural odes via relaxed optimal control
Jabir, J.-F., Šiška, D., and Szpruch, Ł. (2019) · 1912
Earlier work this paper cites.
Diffusions hypercontractives in sem. probab. xix lnm 1123
Bakry, D. and Émery, M. (1985) · 1985
Earlier work this paper cites.
Logarithmic sobolev inequalities and stochastic ising models
Holley, R. and Stroock, D. (1987) · 1987
Earlier work this paper cites.
Sample estimate of the entropy of a random vector
Kozachenko, L. and Leonenko, N. N. (1987) · 1987
Earlier work this paper cites.
18. an extended variational
Jordan, R. and Kinderlehrer, D. (1996) · 1996
Earlier work this paper cites.
Exponential convergence of langevin distributions and their discrete approximations
Roberts, G. O. and Tweedie, R. L. (1996) · 1996
Earlier work this paper cites.
The variational formulation of the fokker–planck equation
Jordan, R., Kinderlehrer, D., and Otto, F. (1998) · 1998
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Laurent, B. and Massart, P. (2000) · 2000
Earlier work this paper cites.
Generalization of an inequality by talagrand and links with the logarithmic sobolev inequality
Otto, F. and Villani, C. (2000) · 2000
Earlier work this paper cites.
Backward feature correction: How deep learning performs deep learning
Allen-Zhu, Z. and Li, Y. (2020) · 2001
Earlier work this paper cites.
A rigorous framework for the mean field limit of multilayer neural networks
Nguyen, P.-M. and Pham, H. T. (2020) · 2001
Earlier work this paper cites.
A generalized neural tangent kernel analysis for two-layer neural networks
Chen, Z., Cao, Y., Gu, Q., and Zhang, T. (2020) · 2002
Earlier work this paper cites.
Learning parities with neural networks
Daniely, A. and Malach, E. (2020) · 2002
Earlier work this paper cites.
Ergodicity for sdes and approximations: locally lipschitz vector fields and degenerate noise
Mattingly, J. C., Stuart, A. M., and Higham, D. J. (2002) · 2002
Earlier work this paper cites.
Lu, Y., Ma, C., Lu, Y., Lu, J., and Ying, L. (2020) · 2003
Earlier work this paper cites.
Convex neural networks
Bengio, Y., Le Roux, N., Vincent, P., Delalleau, O., and Marcotte, P. (2005) · 2005
Earlier work this paper cites.
On the convergence of langevin monte carlo: The interplay between tail growth and smoothness
Erdogdu, M. A. and Hosseinzadeh, R. (2020) · 2005
Earlier work this paper cites.
Smooth minimization of non-smooth functions
Nesterov, Y. (2005) · 2005
Earlier work this paper cites.
When do neural networks outperform kernel methods?
Ghorbani, B., Mei, S., Misiakiewicz, T., and Montanari, A. (2020) · 2006
Earlier work this paper cites.
Convergence of unadjusted hamiltonian monte carlo for mean-field models
Bou-Rabee, N. and Schuh, K. (2020) · 2009
Cited alongside, same era.
Primal-dual subgradient methods for convex problems
Nesterov, Y. (2009) · 2009
Cited alongside, same era.
Dual averaging method for regularized stochastic learning and online optimization
Xiao, L. (2009) · 2009
Cited alongside, same era.
Couplings and quantitative contraction rates for langevin dynamics
Eberle, A., Guillin, A., Zimmer, R., et al. (2019) · 2010
Cited alongside, same era.
Advantage of deep neural networks for estimating functions with singularity on curves
Imaizumi, M. and Fukumizu, K. (2020) · 2011
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z., Li, Y., and Song, Z. (2019) · 2019
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y. and Gu, Q. (2019) · 2019
Later among the works it cites.
Probability functional descent: A unifying perspective on gans, variational inference, and reinforcement learning
Chu, C., Blanchet, J., and Glynn, P. (2019) · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Du, S. S., Zhai, X., Poczos, B., and Singh, A. (2019) · 2019
Later among the works it cites.
Analysis of langevin monte carlo via convex optimization
Durmus, A., Majewski, S., and Miasojedow, B. (2019) · 2019
Later among the works it cites.
Finding mixed nash equilibria of generative adversarial networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. (2011) · 2011
Cited alongside, same era.
Feature learning in infinite-width neural networks
Yang, G. and Hu, E. J. (2020) · 2011
Cited alongside, same era.
Foundations of Machine Learning
Mohri, M., Rostamizadeh, A., and Talwalkar, A. (2012) · 2012
Cited alongside, same era.
Theoretical guarantees for approximate sampling from smooth and log-concave densities
Dalalyan, A. S. (2014) · 2014
Cited alongside, same era.
Poincaré and logarithmic sobolev inequalities by decomposition of the energy landscape
Menz, G. and Schlichting, A. (2014) · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S. (2014) · 2014
Cited alongside, same era.
Provable bayesian inference via particle mirror descent
Dai, B., He, N., Dai, H., and Song, L. (2016) · 2016
Cited alongside, same era.
Hsieh, Y.-P., Liu, C., and Cevher, V. (2019) · 2019
Later among the works it cites.
Stochastic runge-kutta accelerates langevin monte carlo and beyond
Li, X., Wu, Y., Mackey, L., and Erdogdu, M. A. (2019) · 2019
Later among the works it cites.
Scaling limit of the stein variational gradient descent: The mean field regime
Lu, J., Lu, Y., and Nolen, J. (2019) · 2019
Later among the works it cites.
Global convergence of neuron birth-death dynamics
Rotskoff, G. M., Jelassi, S., Bruna, J., and Vanden-Eijnden, E. (2019) · 2019
Later among the works it cites.
Adaptivity of deep relu network for learning in besov and mixed smooth besov spaces: optimal rate and curse of dimensionality
Suzuki, T. (2019) · 2019
Later among the works it cites.
Rapid convergence of the unadjusted langevin algorithm: Isoperimetry suffices
Vempala, S. and Wibisono, A. (2019) · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Wei, C., Lee, J. D., Liu, Q., and Ma, T. (2019) · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
Yehudai, G. and Shamir, O. (2019) · 2019
Later among the works it cites.
Interacting langevin diffusions: Gradient structure and ensemble kalman sampler
Garbuno-Inigo, A., Hoffmann, F., Li, W., and Stuart, A. M. (2020) · 2020
Closest in time.
Learning over-parametrized two-layer neural networks beyond ntk
Li, Y., Ma, T., and Zhang, H. R. (2020) · 2020
Closest in time.
Nonparametric regression using deep neural networks with relu activation function
Schmidt-Hieber, J. (2020) · 2020
Closest in time.
Mean field analysis of neural networks: A central limit theorem
Sirignano, J. and Spiliopoulos, K. (2020) · 2020
Closest in time.
Generalization bound of globally optimal non-convex neural network training: Transportation map estimation by infinite dimensional langevin dynamics
Suzuki, T. (2020) · 2020
Closest in time.
Mirror descent algorithms for minimizing interacting free energy
Ying, L. (2020) · 2020
Closest in time.
Gradient descent optimizes over-parameterized deep relu networks
Zou, D., Cao, Y., Zhou, D., and Gu, Q. (2020) · 2020
Closest in time.
Mixing time guarantees for unadjusted hamiltonian monte carlo
Bou-Rabee, N. and Eberle, A. (2021) · 2021
Closest in time.
Convergence rates of gradient methods for convex optimization in the space of measures
Chizat, L. (2021) · 2021
Closest in time.
Frank-wolfe methods in probability space
Kent, C., Blanchet, J., and Glynn, P. (2021) · 2021
Closest in time.
Optimal rates for averaged stochastic gradient descent under neural tangent kernel regime
Nitanda, A. and Suzuki, T. (2021) · 2021
Closest in time.
Global convergence of three-layer neural networks in the mean field regime
Pham, H. T. and Nguyen, P.-M. (2021) · 2021
Closest in time.
Deep learning is adaptive to intrinsic dimensionality of model smoothness in anisotropic besov space
Suzuki, T. and Nitanda, A. (2021) · 2021
Closest in time.
Sampling as optimization in the space of measures: The langevin dynamics as a composite optimization problem
Wibisono, A. (2018) · 2093
Closest in time.