Fetching the paper…
Reading the bibliography…
Many convex optimization problems with important applications in machine learning are formulated as empirical risk minimization (ERM).
The regression analysis of binary sequences
David R Cox · 1958
Earlier work this paper cites.
On estimating regression
Elizbar A Nadaraya · 1964
Earlier work this paper cites.
Smooth regression analysis
Geoffrey S Watson · 1964
Earlier work this paper cites.
An algorithm for the machine calculation of complex Fourier series
James W Cooley and John W Tukey · 1965
Earlier work this paper cites.
A new polynomial-time algorithm for linear programming
Narendra Karmarkar · 1984
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik · 1995
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
Introductory lectures on convex programming volume i: Basic course
Yurii Nesterov · 1998
Earlier work this paper cites.
The space complexity of approximating the frequency moments
Noga Alon, Yossi Matias, and Mario Szegedy · 1999
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Finding frequent items in data streams
Moses Charikar, Kevin Chen, and Martin Farach-Colton · 2002
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Introduction to algorithms
Thomas H Cormen, Charles E Leiserson, Ronald L Rivest, and Clifford Stein · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Alexander J Smola, and Lihong Li · 2010
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
Feng Niu, Benjamin Recht, Christopher Ré, and Stephen J Wright · 2011
Earlier work this paper cites.
Nearly optimal sparse fourier transform
Haitham Hassanieh, Piotr Indyk, Dina Katabi, and Eric Price · 2012
Earlier work this paper cites.
Simple and practical algorithm for sparse Fourier transform
Haitham Hassanieh, Piotr Indyk, Dina Katabi, and Eric Price · 2012
Earlier work this paper cites.
Low rank approximation and regression in input sparsity time
Kenneth L. Clarkson and David P. Woodruff · 2013
Earlier work this paper cites.
Applied logistic regression
David W Hosmer Jr, Stanley Lemeshow, and Rodney X Sturdivant · 2013
Earlier work this paper cites.
Faster ridge regression via the subsampled randomized hadamard transform
Yichao Lu, Paramveer Dhillon, Dean P Foster, and Lyle Ungar · 2013
Earlier work this paper cites.
A parallel sgd method with strong convergence
Dhruv Mahajan, S Sathiya Keerthi, S Sundararajan, and Léon Bottou · 2013
Earlier work this paper cites.
Low-distortion subspace embeddings in input-sparsity time and applications to robust linear regression
Xiangrui Meng and Michael W Mahoney · 2013
Earlier work this paper cites.
Osnap: Faster numerical linear algebra algorithms via sparser subspace embeddings
Jelani Nelson and Huy L Nguyên · 2013
Earlier work this paper cites.
Sparse recovery and Fourier sampling
Eric C. Price · 2013
Earlier work this paper cites.
Trading computation for communication: Distributed stochastic dual coordinate ascent
Tianbao Yang · 2013
Earlier work this paper cites.
Optimal cur matrix decompositions
Christos Boutsidis and David P Woodruff · 2014
Earlier work this paper cites.
Sample-optimal fourier sampling in any constant dimension
Piotr Indyk and Michael Kapralov · 2014
Earlier work this paper cites.
(Nearly) Sample-optimal sparse Fourier transform
Piotr Indyk, Michael Kapralov, and Eric Price · 2014
Earlier work this paper cites.
Communication-efficient distributed dual coordinate ascent
Martin Jaggi, Virginia Smith, Martin Takáč, Jonathan Terhorst, Sanjay Krishnan, Thomas Hofmann, and Michael I Jordan · 2014
Earlier work this paper cites.
Distributed stochastic optimization and learning
Ohad Shamir and Nathan Srebro · 2014
Earlier work this paper cites.
Communication-efficient distributed optimization using an approximate newton-type method
Ohad Shamir, Nati Srebro, and Tong Zhang · 2014
Earlier work this paper cites.
Adding vs. averaging in distributed primal-dual optimization
Chenxin Ma, Virginia Smith, Martin Jaggi, Michael Jordan, Peter Richtárik, and Martin Takác · 2015
Earlier work this paper cites.
A robust sparse Fourier transform in the continuous setting
Eric Price and Zhao Song · 2015
Earlier work this paper cites.
On variance reduction in stochastic gradient descent and its asynchronous variants
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex Smola · 2015
Earlier work this paper cites.
Disco: Distributed optimization for self-concordant empirical loss
Yuchen Zhang and Xiao Lin · 2015
Earlier work this paper cites.
Optimal principal component analysis in distributed and streaming models
Christos Boutsidis, David P Woodruff, and Peilin Zhong · 2016
Earlier work this paper cites.
Fourier-sparse interpolation without a frequency gap
Xue Chen, Daniel M Kane, Eric Price, and Zhao Song · 2016
Earlier work this paper cites.
Optimization in high dimensions via accelerated, parallel, and proximal coordinate descent
Olivier Fercoq and Peter Richtárik · 2016
Earlier work this paper cites.
Sparse Fourier transform in any constant dimension with nearly-optimal sample complexity in sublinear time
Michael Kapralov · 2016
Cited alongside, same era.
Federated learning: Strategies for improving communication efficiency
Jakub Konečnỳ, H Brendan McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Cited alongside, same era.
Aide: Fast and communication efficient distributed optimization
Sashank J Reddi, Jakub Konečnỳ, Peter Richtárik, Barnabás Póczós, and Alex Smola · 2016
Cited alongside, same era.
Distributed coordinate descent method for learning with big data
Peter Richtárik and Martin Takáč · 2016
Cited alongside, same era.
Distributed low rank approximation of implicit functions of a matrix
David P Woodruff and Peilin Zhong · 2016
Cited alongside, same era.
Fl-ntk: A neural tangent kernel-based framework for federated learning analysis
Baihe Huang, Xiaoxiao Li, Zhao Song, and Xin Yang · 2021
Later among the works it cites.
Faster dynamic matrix inverse for faster lps
Shunhua Jiang, Zhao Song, Omri Weinstein, and Hengjie Zhang · 2021
Later among the works it cites.
Fedbn: Federated learning on non-iid features via local batch normalization
Xiaoxiao Li, Meirui Jiang, Xiaofei Zhang, Michael Kamp, and Qi Dou · 2021
Later among the works it cites.
Oblivious sketching-based central path method for linear programming
Zhao Song and Zheng Yu · 2021
Later among the works it cites.
Does preprocessing help training over-parameterized neural networks?
Zhao Song, Shuo Yang, and Ruizhe Zhang · 2021
Later among the works it cites.
Training multi-layer over-parametrized neural network in subquadratic time
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sample efficient estimation and recovery in sparse fft via isolation on average
Michael Kapralov · 2017
Cited alongside, same era.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J Dally · 2017
Cited alongside, same era.
Distributed stochastic variance reduced gradient methods by sampling extra data with replacement
Jason D Lee, Qihang Lin, Tengyu Ma, and Tianbao Yang · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Cited alongside, same era.
Fast regression with an ℓ ∞ {\ell}_{\infty} guarantee
Eric Price, Zhao Song, and David P. Woodruff · 2017
Cited alongside, same era.
Low rank approximation with entrywise ℓ 1 \ell_{1} -norm error
Zhao Song, David P Woodruff, and Peilin Zhong · 2017
Cited alongside, same era.
A general distributed dual coordinate optimization framework for regularized loss minimization
Shun Zheng, Jialei Wang, Fen Xia, Wei Xu, and Tong Zhang · 2017
Cited alongside, same era.
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2021
Later among the works it cites.
Josh Alman, Jiehao Liang, Zhao Song, Ruizhe Zhang, and Danyang Zhuo · 2022
Later among the works it cites.
Symmetric sparse boolean matrix factorization and applications
Sitan Chen, Zhao Song, Runzhou Tao, and Ruizhe Zhang · 2022
Later among the works it cites.
Yichuan Deng, Wenyu Jin, Zhao Song, Xiaorui Sun, and Omri Weinstein · 2022
Later among the works it cites.
Discrepancy minimization in input-sparsity time
Yichuan Deng, Zhao Song, and Omri Weinstein · 2022
Later among the works it cites.
Fast distance oracles for any symmetric norm
Yichuan Deng, Zhao Song, Omri Weinstein, and Ruizhe Zhang · 2022
Later among the works it cites.
A sublinear adversarial training algorithm
Yeqi Gao, Lianke Qin, Zhao Song, and Yitan Wang · 2022
Later among the works it cites.
A faster small treewidth sdp solver
Yuzhou Gu and Zhao Song · 2022
Later among the works it cites.
An o ( k log n ) o(k\log n) time Fourier set query algorithm
Yeqi Gao, Zhao Song, and Baocheng Sun · 2022
Later among the works it cites.
A faster quantum algorithm for semidefinite programming via robust ipm framework
Baihe Huang, Shunhua Jiang, Zhao Song, Runzhou Tao, and Ruizhe Zhang · 2022
Later among the works it cites.
Solving sdp faster: A robust ipm framework and efficient implementation
Baihe Huang, Shunhua Jiang, Zhao Song, Runzhou Tao, and Ruizhe Zhang · 2022
Later among the works it cites.
Sublinear time algorithm for online weighted bipartite matching
Hang Hu, Zhao Song, Runzhou Tao, Zhaozhuo Xu, and Danyang Zhuo · 2022
Later among the works it cites.
Training overparametrized neural networks in sublinear time
Hang Hu, Zhao Song, Omri Weinstein, and Danyang Zhuo · 2022
Later among the works it cites.
Adore: Differentially oblivious relational database operators
Lianke Qin, Rajesh Jayaram, Elaine Shi, Zhao Song, Danyang Zhuo, and Shumo Chu · 2022
Later among the works it cites.
Adaptive and dynamic multi-resolution hashing for pairwise summations
Lianke Qin, Aravind Reddy, Zhao Song, Zhaozhuo Xu, and Danyang Zhuo · 2022
Later among the works it cites.
Quartic samples suffice for fourier interpolation
Zhao Song, Baocheng Sun, Omri Weinstein, and Ruizhe Zhang · 2022
Later among the works it cites.
Sparse fourier transform over lattices: A unified approach to signal reconstruction
Zhao Song, Baocheng Sun, Omri Weinstein, and Ruizhe Zhang · 2022
Later among the works it cites.
Speeding up optimizations via data structures: Faster search, sample and maintenance
Lichen Zhang · 2022
Later among the works it cites.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Closest in time.
Attention scheme inspired softmax regression
Yichuan Deng, Zhihang Li, and Zhao Song · 2023
Closest in time.
An improved sample complexity for rank-1 matrix sensing
Yichuan Deng, Zhihang Li, and Zhao Song · 2023
Closest in time.
Yichuan Deng, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
An iterative algorithm for rescaled hyperbolic functions regression
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Closest in time.
Low rank matrix completion via robust alternating minimization in nearly linear time
Yuzhou Gu, Zhao Song, Junze Yin, and Lichen Zhang · 2023
Closest in time.
Convex minimization with integer minima in O ~ ( n 4 ) \widetilde{O}(n^{4}) time
Haotian Jiang, Yin Tat Lee, Zhao Song, and Lichen Zhang · 2023
Closest in time.
The closeness of in-context learning and weight shifting for softmax regression
Shuai Li, Zhao Song, Yu Xia, Tong Yu, and Tianyi Zhou · 2023
Closest in time.
Federated adversarial learning: A framework with convergence analysis
Xiaoxiao Li, Zhao Song, and Jiaming Yang · 2023
Closest in time.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Fast submodular function maximization
Lianke Qin, Zhao Song, and Yitan Wang · 2023
Closest in time.
A general algorithm for solving rank-one matrix sensing
Lianke Qin, Zhao Song, and Ruizhe Zhang · 2023
Closest in time.
An online and unified algorithm for projection matrix vector multiplication with application to empirical risk minimization
Lianke Qin, Zhao Song, Lichen Zhang, and Danyang Zhuo · 2023
Closest in time.
Fast and efficient matching algorithm with deadline instances
Zhao Song, Weixin Wang, and Chenbo Yin · 2023
Closest in time.
Sketching for first order method: Efficient algorithm for low-bandwidth channel and vulnerability
Zhao Song, Yitan Wang, Zheng Yu, and Lichen Zhang · 2023
Closest in time.
A nearly-optimal bound for fast regression with ℓ ∞ \ell_{\infty} guarantee
Zhao Song, Mingquan Ye, Junze Yin, and Lichen Zhang · 2023
Closest in time.