Fetching the paper…
Reading the bibliography…
Excessive computational cost for learning large data and streaming data can be alleviated by using stochastic algorithms, such as stochastic gradient descent and its variants.
Algorithms of robust stochastic optimization based on mirror descent method
Juditsky, A., Nazin, A., Nemirovsky, A., and Tsybakov, A. (2019) · 1907
Earlier work this paper cites.
Algorithms of robust stochastic optimization based on mirror descent method
Juditsky, A., Nazin, A., Nemirovsky, A., and Tsybakov, A. (2019) · 1907
Earlier work this paper cites.
Statistical and computational trade-offs in estimation of sparse principal components
Wang, T., Berthet, Q., and Samworth, R. J. (2016) · 1930
Earlier work this paper cites.
Statistical and computational trade-offs in estimation of sparse principal components
Wang, T., Berthet, Q., and Samworth, R. J. (2016) · 1930
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S. (1951) · 1951
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S. (1951) · 1951
Earlier work this paper cites.
Ordinary Differential Equations
Hale, J. K. (1969) · 1969
Earlier work this paper cites.
Ordinary Differential Equations
Hale, J. K. (1969) · 1969
Earlier work this paper cites.
Convex analysis
Rockafellar, R. T. (1970) · 1970
Earlier work this paper cites.
Convex analysis
Rockafellar, R. T. (1970) · 1970
Earlier work this paper cites.
Some useful functions for functional limit theorems
Whitt, W. (1980) · 1980
Earlier work this paper cites.
Some useful functions for functional limit theorems
Whitt, W. (1980) · 1980
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovski, A. S. and Yudin, D. B. (1983) · 1983
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Nemirovski, A. S. and Yudin, D. B. (1983) · 1983
Earlier work this paper cites.
On stochastic approximation of the eigenvectors and eigenvalues of the expectation of a random matrix
Oja, E. and Karhunen, J. (1985) · 1985
Earlier work this paper cites.
On stochastic approximation of the eigenvectors and eigenvalues of the expectation of a random matrix
Oja, E. and Karhunen, J. (1985) · 1985
Earlier work this paper cites.
Markov Processes: Characterization and Convergence
Ethier, S. N. and Kurtz, T. G. (1986) · 1986
Earlier work this paper cites.
Markov Processes: Characterization and Convergence
Ethier, S. N. and Kurtz, T. G. (1986) · 1986
Earlier work this paper cites.
Efficient estimations from a slowly convergent Robbins–Monro process
Ruppert, D. (1988) · 1988
Earlier work this paper cites.
Efficient estimations from a slowly convergent Robbins–Monro process
Ruppert, D. (1988) · 1988
Earlier work this paper cites.
Adaptive Algorithms and Stochastic Approximations
Benveniste, A., Métivier, M., and Priouret, P. (1990) · 1990
Earlier work this paper cites.
Adaptive Algorithms and Stochastic Approximations
Benveniste, A., Métivier, M., and Priouret, P. (1990) · 1990
Earlier work this paper cites.
Numerical Solution of Stochastic Differential Equations
Kloeden, P. E. and Platen, E. (1992) · 1992
Earlier work this paper cites.
Asymptotics for m m -estimators defined by convex minimization
Niemiro, W. (1992) · 1992
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B. (1992) · 1992
Earlier work this paper cites.
Perturbation bounds for matrix square roots and pythagorean sums
Schmitt, B. A. (1992) · 1992
Earlier work this paper cites.
Numerical Solution of Stochastic Differential Equations
Kloeden, P. E. and Platen, E. (1992) · 1992
Earlier work this paper cites.
Asymptotics for m m -estimators defined by convex minimization
Niemiro, W. (1992) · 1992
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Polyak, B. T. and Juditsky, A. B. (1992) · 1992
Earlier work this paper cites.
Perturbation bounds for matrix square roots and pythagorean sums
Schmitt, B. A. (1992) · 1992
Earlier work this paper cites.
Weak convergence and local stability properties of fixed step size recursive algorithms
Bucklew, J. A., Kurtz, T. G., and Sethares, W. A. (1993) · 1993
Earlier work this paper cites.
Variable selection via gibbs sampling
George, E. I. and McCulloch, R. E. (1993) · 1993
Earlier work this paper cites.
Real and Functional Analysis
Lang, S. (1993) · 1993
Earlier work this paper cites.
Weak convergence and local stability properties of fixed step size recursive algorithms
Bucklew, J. A., Kurtz, T. G., and Sethares, W. A. (1993) · 1993
Earlier work this paper cites.
Variable selection via gibbs sampling
George, E. I. and McCulloch, R. E. (1993) · 1993
Earlier work this paper cites.
Real and Functional Analysis
Lang, S. (1993) · 1993
Earlier work this paper cites.
Probability and Measure
Billingsley, P. (1995) · 1995
Earlier work this paper cites.
Lyapunov functions for convergence of principal component algorithms
Plumbley, M. D. (1995) · 1995
Earlier work this paper cites.
Probability and Measure
Billingsley, P. (1995) · 1995
Earlier work this paper cites.
Lyapunov functions for convergence of principal component algorithms
Plumbley, M. D. (1995) · 1995
Earlier work this paper cites.
Asymptotic methods in the theory of Gaussian processes and fields
Piterbarg, V. I. (1996) · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Tibshirani, R. (1996) · 1996
Earlier work this paper cites.
Weak convergence and empirical processes
van der Vaart, A. W. and Wellner, J. A. (1996) · 1996
Earlier work this paper cites.
Asymptotic methods in the theory of Gaussian processes and fields
Piterbarg, V. I. (1996) · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Tibshirani, R. (1996) · 1996
Earlier work this paper cites.
Weak convergence and empirical processes
van der Vaart, A. W. and Wellner, J. A. (1996) · 1996
Earlier work this paper cites.
Foundations of Modern Probability
Kallenberg, O. (1997) · 1997
Earlier work this paper cites.
Foundations of Modern Probability
Kallenberg, O. (1997) · 1997
Earlier work this paper cites.
Brownian Motion and Stochastic Calculus
Karatzas, I. and Shreve, S. (1998) · 1998
Earlier work this paper cites.
Limiting distributions for L 1 L_{1} regression estimators under general conditions
Knight, K. (1998) · 1998
Earlier work this paper cites.
Brownian Motion and Stochastic Calculus
Karatzas, I. and Shreve, S. (1998) · 1998
Earlier work this paper cites.
Limiting distributions for L 1 L_{1} regression estimators under general conditions
Knight, K. (1998) · 1998
Earlier work this paper cites.
Convergence of Probability Measures
Billingsley, P. (1999) · 1999
Earlier work this paper cites.
Convergence of Probability Measures
Billingsley, P. (1999) · 1999
Earlier work this paper cites.
Asymptotics for Lasso-type estimators
Knight, K. and Fu, W. (2000) · 2000
Earlier work this paper cites.
Asymptotics for Lasso-type estimators
Knight, K. and Fu, W. (2000) · 2000
Earlier work this paper cites.
Relative loss bounds for on-line density estimation with the exponential family of distributions
Azoury, K. S. and Warmuth, M. K. (2001) · 2001
Earlier work this paper cites.
Variable selection via nonconcave penalized likelihood and its oracle properties
Fan, J. and Li, R. (2001) · 2001
Earlier work this paper cites.
Competitive on-line statistics
Vovk, V. (2001) · 2001
Earlier work this paper cites.
Relative loss bounds for on-line density estimation with the exponential family of distributions
Azoury, K. S. and Warmuth, M. K. (2001) · 2001
Earlier work this paper cites.
Variable selection via nonconcave penalized likelihood and its oracle properties
Fan, J. and Li, R. (2001) · 2001
Earlier work this paper cites.
Competitive on-line statistics
Vovk, V. (2001) · 2001
Earlier work this paper cites.
Nonlinear Systems
Khalil, H. K. (2002) · 2002
Earlier work this paper cites.
Convex analysis in general vector spaces
Zǎlinescu, C. (2002) · 2002
Earlier work this paper cites.
Nonlinear Systems
Khalil, H. K. (2002) · 2002
Earlier work this paper cites.
Convex analysis in general vector spaces
Zǎlinescu, C. (2002) · 2002
Earlier work this paper cites.
A modified principal component technique based on the lasso
Jolliffe, I. T., Trendafilov, N. T., and Uddin, M. (2003) · 2003
Earlier work this paper cites.
Stochastic Approximation and Recursive Algorithms and Applications
Kushner, H. J. and Yin, G. (2003) · 2003
Earlier work this paper cites.
A modified principal component technique based on the lasso
Jolliffe, I. T., Trendafilov, N. T., and Uddin, M. (2003) · 2003
Cited alongside, same era.
Stochastic Approximation and Recursive Algorithms and Applications
Kushner, H. J. and Yin, G. (2003) · 2003
Cited alongside, same era.
Elementary Differential Equations and Boundary Value Problems
Boyce, W. E. and DiPrima, R. C. (2005) · 2005
Cited alongside, same era.
Probability: Theory and Examples
Durrett, R. (2005) · 2005
Cited alongside, same era.
Regularization and variable selection via the elastic net
Zou, H. and Hastie, T. (2005) · 2005
Cited alongside, same era.
Elementary Differential Equations and Boundary Value Problems
Boyce, W. E. and DiPrima, R. C. (2005) · 2005
Cited alongside, same era.
Ordinary Differential Equations and Dynamical Systems
Teschl, G. (2012) · 2012
Later among the works it cites.
Non-strongly-convex smooth stochastic approximation with convergence rate O ( 1 / N ) O(1/N)
Bach, F. and Moulines, E. (2013) · 2013
Later among the works it cites.
Minimax bounds for sparse PCA with noisy high-dimensional data
Birnbaum, A., Johnstone, I. M., Nadler, B., and Paul, D. (2013) · 2013
Later among the works it cites.
Sparse PCA: Optimal rates and adaptive estimation
Cai, T. T., Ma, Z., and Wu, Y. (2013) · 2013
Later among the works it cites.
Sparse principal component analysis and iterative thresholding
Ma, Z. (2013) · 2013
Later among the works it cites.
Consistency of sparse pca in high dimension, low sample size contexts
Shen, D., Shen, H., and Marron, J. (2013) · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Probability: Theory and Examples
Durrett, R. (2005) · 2005
Cited alongside, same era.
Regularization and variable selection via the elastic net
Zou, H. and Hastie, T. (2005) · 2005
Cited alongside, same era.
Model selection and estimation in regression with grouped variables
Yuan, M. and Lin, Y. (2006) · 2006
Cited alongside, same era.
Sparse principal component analysis
Zou, H., Hastie, T., and Tibshirani, R. (2006) · 2006
Cited alongside, same era.
Model selection and estimation in regression with grouped variables
Yuan, M. and Lin, Y. (2006) · 2006
Cited alongside, same era.
Sparse principal component analysis
Zou, H., Hastie, T., and Tibshirani, R. (2006) · 2006
Cited alongside, same era.
Fantope projection and selection: A near-optimal convex relaxation of sparse PCA
Vu, V. Q., Cho, J., Lei, J., and Rohe, K. (2013) · 2013
Later among the works it cites.
Minimax sparse principal subspace estimation in high dimensions
Vu, V. Q. and Lei, J. (2013) · 2013
Later among the works it cites.
Non-strongly-convex smooth stochastic approximation with convergence rate O ( 1 / N ) O(1/N)
Bach, F. and Moulines, E. (2013) · 2013
Later among the works it cites.
Minimax bounds for sparse PCA with noisy high-dimensional data
Birnbaum, A., Johnstone, I. M., Nadler, B., and Paul, D. (2013) · 2013
Later among the works it cites.
Sparse PCA: Optimal rates and adaptive estimation
Cai, T. T., Ma, Z., and Wu, Y. (2013) · 2013
Later among the works it cites.
Sparse principal component analysis and iterative thresholding
Ma, Z. (2013) · 2013
Later among the works it cites.
Consistency of sparse pca in high dimension, low sample size contexts
Shen, D., Shen, H., and Marron, J. (2013) · 2013
Later among the works it cites.
Fantope projection and selection: A near-optimal convex relaxation of sparse PCA
Vu, V. Q., Cho, J., Lei, J., and Rohe, K. (2013) · 2013
Later among the works it cites.
Minimax sparse principal subspace estimation in high dimensions
Vu, V. Q. and Lei, J. (2013) · 2013
Later among the works it cites.
Proximal reinforcement learning: A new theory of sequential decision making in primal-dual spaces
Mahadevan, S., Liu, B., Thomas, P., Dabney, W., Giguere, S., Jacek, N., Gemp, I., and Liu, J. (2014) · 2014
Later among the works it cites.
An easy path to convex analysis and applications
Mordukhovich, B. S. and Nam, N. M. (2014) · 2014
Later among the works it cites.
Proximal algorithms
Parikh, N. and Boyd, S. (2014) · 2014
Later among the works it cites.
Proximal reinforcement learning: A new theory of sequential decision making in primal-dual spaces
Mahadevan, S., Liu, B., Thomas, P., Dabney, W., Giguere, S., Jacek, N., Gemp, I., and Liu, J. (2014) · 2014
Later among the works it cites.
An easy path to convex analysis and applications
Mordukhovich, B. S. and Nam, N. M. (2014) · 2014
Later among the works it cites.
Proximal algorithms
Parikh, N. and Boyd, S. (2014) · 2014
Later among the works it cites.
A generalized online mirror descent with applications to classification and regression
Orabona, F., Crammer, K., and Cesa-Bianchi, N. (2015) · 2015
Later among the works it cites.
Streaming sparse principal component analysis
Yang, W. and Xu, H. (2015) · 2015
Later among the works it cites.
A generalized online mirror descent with applications to classification and regression
Orabona, F., Crammer, K., and Cesa-Bianchi, N. (2015) · 2015
Later among the works it cites.
Streaming sparse principal component analysis
Yang, W. and Xu, H. (2015) · 2015
Later among the works it cites.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A. (2016) · 2016
Later among the works it cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Han, S., Mao, H., and Dally, W. (2016) · 2016
Later among the works it cites.
Online learning for sparse pca in high dimensions: Exact dynamics and phase transitions
Wang, C. and Lu, Y. M. (2016) · 2016
Later among the works it cites.
Deep Learning
Goodfellow, I., Bengio, Y., and Courville, A. (2016) · 2016
Later among the works it cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Han, S., Mao, H., and Dally, W. (2016) · 2016
Later among the works it cites.
Online learning for sparse pca in high dimensions: Exact dynamics and phase transitions
Wang, C. and Lu, Y. M. (2016) · 2016
Later among the works it cites.
Stochastic composite least-squares regression with convergence rate O ( 1 / n ) O(1/n)
Flammarion, N. and Bach, F. (2017) · 2017
Later among the works it cites.
Diffusion approximations for online principal component estimation and global convergence
Li, C. J., Wang, M., Liu, H., and Zhang, T. (2017) · 2017
Later among the works it cites.
Stochastic mirror descent in variationally coherent optimization problems
Zhou, Z., Mertikopoulos, P., Bambos, N., Boyd, S., and Glynn, P. (2017) · 2017
Later among the works it cites.
Stochastic composite least-squares regression with convergence rate O ( 1 / n ) O(1/n)
Flammarion, N. and Bach, F. (2017) · 2017
Later among the works it cites.
Diffusion approximations for online principal component estimation and global convergence
Li, C. J., Wang, M., Liu, H., and Zhang, T. (2017) · 2017
Later among the works it cites.
Stochastic mirror descent in variationally coherent optimization problems
Zhou, Z., Mertikopoulos, P., Bambos, N., Boyd, S., and Glynn, P. (2017) · 2017
Later among the works it cites.
Online principal component analysis in high dimension: Which algorithm to choose?
Cardot, H. and Degras, D. (2018) · 2018
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Chaudhari, P. and Soatto, S. (2018) · 2018
Later among the works it cites.
Convergence diagnostics for stochastic gradient descent with constant learning rate
Chee, J. and Toulis, P. (2018) · 2018
Later among the works it cites.
Bridging the Gap between Constant Step Size Stochastic Gradient Descent and Markov Chains
Dieuleveut, A., Durmus, A., and Bach, F. (2018) · 2018
Later among the works it cites.
Sparse principal component analysis via variable projection
Erichson, N. B., Zheng, P., Manohar, K., Brunton, S. L., Kutz, J. N., and Aravkin, A. Y. (2018) · 2018
Later among the works it cites.
Sparse principal component analysis via random projections
Gataric, M., Wang, T., and Samworth, R. J. (2018) · 2018
Later among the works it cites.
De-biased sparse PCA: Inference and testing for eigenstructure of large covariance matrices
Janková, J. and van de Geer, S. (2018) · 2018
Later among the works it cites.
Modified regularized dual averaging method for training sparse convolutional neural networks
Jia, X., Zhao, L., Zhang, L., He, J., and Xu, J. (2018) · 2018
Later among the works it cites.
Convergence of online mirror descent
Lei, Y. and Zhou, D.-X. (2018) · 2018
Later among the works it cites.
Learning sparse neural networks through L 0 L_{0} regularization
Louizos, C., Welling, M., and Kingma, D. P. (2018) · 2018
Later among the works it cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A. (2018) · 2018
Later among the works it cites.
Uncertainty quantification for online learning and stochastic approximation via hierarchical incremental gradient descent
Su, W. J. and Zhu, Y. (2018) · 2018
Later among the works it cites.
On convergence of some gradient-based temporal-differences algorithms for off-policy learning
Yu, H. (2018) · 2018
Later among the works it cites.
On the convergence rate of stochastic mirror descent for nonsmooth nonconvex optimization
Zhang, S. and He, N. (2018) · 2018
Later among the works it cites.
Online principal component analysis in high dimension: Which algorithm to choose?
Cardot, H. and Degras, D. (2018) · 2018
Later among the works it cites.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Chaudhari, P. and Soatto, S. (2018) · 2018
Later among the works it cites.
Convergence diagnostics for stochastic gradient descent with constant learning rate
Chee, J. and Toulis, P. (2018) · 2018
Later among the works it cites.
Bridging the Gap between Constant Step Size Stochastic Gradient Descent and Markov Chains
Dieuleveut, A., Durmus, A., and Bach, F. (2018) · 2018
Later among the works it cites.
Sparse principal component analysis via variable projection
Erichson, N. B., Zheng, P., Manohar, K., Brunton, S. L., Kutz, J. N., and Aravkin, A. Y. (2018) · 2018
Later among the works it cites.
Sparse principal component analysis via random projections
Gataric, M., Wang, T., and Samworth, R. J. (2018) · 2018
Later among the works it cites.
De-biased sparse PCA: Inference and testing for eigenstructure of large covariance matrices
Janková, J. and van de Geer, S. (2018) · 2018
Later among the works it cites.
Modified regularized dual averaging method for training sparse convolutional neural networks
Jia, X., Zhao, L., Zhang, L., He, J., and Xu, J. (2018) · 2018
Later among the works it cites.
Convergence of online mirror descent
Lei, Y. and Zhou, D.-X. (2018) · 2018
Later among the works it cites.
Learning sparse neural networks through L 0 L_{0} regularization
Louizos, C., Welling, M., and Kingma, D. P. (2018) · 2018
Later among the works it cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A. (2018) · 2018
Later among the works it cites.
Uncertainty quantification for online learning and stochastic approximation via hierarchical incremental gradient descent
Su, W. J. and Zhu, Y. (2018) · 2018
Later among the works it cites.
On convergence of some gradient-based temporal-differences algorithms for off-policy learning
Yu, H. (2018) · 2018
Later among the works it cites.
On the convergence rate of stochastic mirror descent for nonsmooth nonconvex optimization
Zhang, S. and He, N. (2018) · 2018
Later among the works it cites.
Statistical inference for model parameters in stochastic gradient descent
Chen, X., Lee, J. D., Tong, X. T., and Zhang, Y. (2019) · 2019
Closest in time.
Statistical inference for model parameters in stochastic gradient descent
Chen, X., Lee, J. D., Tong, X. T., and Zhang, Y. (2019) · 2019
Closest in time.