Fetching the paper…
Reading the bibliography…
Many convex problems in machine learning and computer science share the same form: \begin{align*} \min_{x} \sum_{i} f_i( A_i x + b_i), \end{align*} where $f_i$ are convex functions on $\mathbb{R}^{n_i}$ with constant $n_i$, $A_i \in \mathbb{R}^{n_i \times d}$, $b_i \in \mathbb{R}^{n_i}$ and $\sum_i n_i = n$.
The regression analysis of binary sequences
David R Cox · 1958
Earlier work this paper cites.
On estimating regression
Elizbar A Nadaraya · 1964
Earlier work this paper cites.
Smooth regression analysis
Geoffrey S Watson · 1964
Earlier work this paper cites.
An algorithm for the machine calculation of complex Fourier series
James W Cooley and John W Tukey · 1965
Earlier work this paper cites.
Gaussian elimination is not optimal
Volker Strassen · 1969
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate O ( 1 / k 2 ) {O}(1/k^{2})
Yurii E Nesterov · 1983
Earlier work this paper cites.
Matrix multiplication via arithmetic progressions
Don Coppersmith and Shmuel Winograd · 1987
Earlier work this paper cites.
Speeding-up linear programming using fast matrix multiplication
Pravin M Vaidya · 1989
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Principles of risk minimization for learning theory
Vladimir Vapnik · 1992
Earlier work this paper cites.
Support-vector networks
Corinna Cortes and Vladimir Vapnik · 1995
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
A decision-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1997
Earlier work this paper cites.
Introductory lectures on convex programming volume i: Basic course
Yurii Nesterov · 1998
Earlier work this paper cites.
Galton, edgeworth, frisch, and prospects for quantile regression in econometrics
Roger Koenker · 2000
Earlier work this paper cites.
Adaptive estimation of a quadratic functional by model selection
Beatrice Laurent and Pascal Massart · 2000
Earlier work this paper cites.
Quantile regression
Roger Koenker and Kevin F Hallock · 2001
Earlier work this paper cites.
A generalized mean field algorithm for variational inference in exponential families
Eric P Xing, Michael I Jordan, and Stuart Russell · 2002
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Yurii Nesterov · 2004
Earlier work this paper cites.
Local rademacher complexities
Peter L Bartlett, Olivier Bousquet, and Shahar Mendelson · 2005
Earlier work this paper cites.
Subgradient and sampling algorithms for ℓ 1 \ell_{1} regression
Kenneth L Clarkson · 2005
Earlier work this paper cites.
Quantile Regression
Roger Koenker · 2005
Earlier work this paper cites.
Regularization and variable selection via the elastic net
Hui Zou and Trevor Hastie · 2005
Earlier work this paper cites.
The tradeoffs of large scale learning
Léon Bottou and Olivier Bousquet · 2008
Earlier work this paper cites.
Sampling algorithms and coresets for ℓ p \ell_{p} regression
Anirban Dasgupta, Petros Drineas, Boulos Harb, Ravi Kumar, and Michael W Mahoney · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R Bach · 2011
Earlier work this paper cites.
Agnostic learning of monomials by halfspaces is hard
Vitaly Feldman, Venkatesan Guruswami, Prasad Raghavendra, and Yi Wu · 2012
Earlier work this paper cites.
Nearly optimal sparse fourier transform
Haitham Hassanieh, Piotr Indyk, Dina Katabi, and Eric Price · 2012
Earlier work this paper cites.
Simple and practical algorithm for sparse Fourier transform
Haitham Hassanieh, Piotr Indyk, Dina Katabi, and Eric Price · 2012
Cited alongside, same era.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas Le Roux, Mark W Schmidt, and Francis R Bach · 2012
Cited alongside, same era.
Multiplying matrices faster than coppersmith-winograd
Virginia Vassilevska Williams · 2012
Cited alongside, same era.
Applied logistic regression
David W Hosmer Jr, Stanley Lemeshow, and Rodney X Sturdivant · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Faster ridge regression via the subsampled randomized hadamard transform
Yichao Lu, Paramveer Dhillon, Dean P Foster, and Lyle Ungar · 2013
Leverage score sampling for faster accelerated regression and erm
Naman Agarwal, Sham Kakade, Rahul Kidambi, Yin Tat Lee, Praneeth Netrapalli, and Aaron Sidford · 2017
Later among the works it cites.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Later among the works it cites.
Natasha: Faster Non-Convex Stochastic Optimization via Strongly Non-Convex Parameter
Zeyuan Allen-Zhu · 2017
Later among the works it cites.
Fast rates for empirical risk minimization of strict saddle problems
Alon Gonen and Shai Shalev-Shwartz · 2017
Later among the works it cites.
Sample efficient estimation and recovery in sparse fft via isolation on average
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sparse recovery and Fourier sampling
Eric C. Price · 2013
Cited alongside, same era.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shai Shalev-Shwartz and Tong Zhang · 2013
Cited alongside, same era.
The nature of statistical learning theory
Vladimir Vapnik · 2013
Cited alongside, same era.
Constant step size least-mean-square: Bias-variance trade-offs and optimal sampling distributions
Alexandre Défossez and Francis Bach · 2014
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Cited alongside, same era.
Sample-optimal fourier sampling in any constant dimension
Piotr Indyk and Michael Kapralov · 2014
Cited alongside, same era.
Michael Kapralov · 2017
Later among the works it cites.
Uniform sampling and inverse maintenance
Yin Tat Lee · 2017
Later among the works it cites.
Non-convex finite-sum optimization via scsg methods
Lihua Lei, Cheng Ju, Jianbo Chen, and Michael I Jordan · 2017
Later among the works it cites.
Catalyst acceleration for first-order convex optimization: from theory to practice
Hongzhou Lin, Julien Mairal, and Zaid Harchaoui · 2017
Later among the works it cites.
Doubly accelerated stochastic variance reduced dual averaging method for regularized empirical risk minimization
Tomoya Murata and Taiji Suzuki · 2017
Later among the works it cites.
Efficiency of the accelerated coordinate descent method on structured optimization problems
Yurii Nesterov and Sebastian U Stich · 2017
Later among the works it cites.
Fast regression with an ℓ ∞ {\ell}_{\infty} guarantee
Eric Price, Zhao Song, and David P. Woodruff · 2017
Later among the works it cites.
Newton sketch: A near linear-time optimization algorithm with linear-quadratic convergence
Mert Pilanci and Martin J Wainwright · 2017
Later among the works it cites.
Variance reduced stochastic gradient descent with sufficient decrease
Fanhua Shang, Yuanyuan Liu, James Cheng, KW Ng, and Yuichi Yoshida · 2017
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2017
Later among the works it cites.
A general distributed dual coordinate optimization framework for regularized loss minimization
Shun Zheng, Jialei Wang, Fen Xia, Wei Xu, and Tong Zhang · 2017
Later among the works it cites.
Stochastic primal-dual coordinate method for regularized empirical risk minimization
Yuchen Zhang and Lin Xiao · 2017
Later among the works it cites.
Lijun Zhang, Tianbao Yang, and Rong Jin · 2017
Later among the works it cites.
Katyusha X: Practical Momentum Method for Stochastic Sum-of-Nonconvex Optimization
Zeyuan Allen-Zhu · 2018
Later among the works it cites.
Natasha 2: Faster Non-Convex Optimization Than SGD
Zeyuan Allen-Zhu · 2018
Later among the works it cites.
Limits on the universal method for matrix multiplication
Josh Alman · 2018
Later among the works it cites.
Further limitations of the known approaches for matrix multiplication
Josh Alman and Virginia Vassilevska Williams · 2018
Later among the works it cites.
Limits on all known (and some unknown) approaches to matrix multiplication
Josh Alman and Virginia Vassilevska Williams · 2018
Later among the works it cites.
An homotopy method for ℓ p \ell_{p} regression provably beyond self-concordance and in input-sparsity time
Sébastien Bubeck, Michael B Cohen, Yin Tat Lee, and Yuanzhi Li · 2018
Later among the works it cites.
Data sampling strategies in stochastic algorithms for empirical risk minimization
Dominik Csiba · 2018
Later among the works it cites.
On the local minima of the empirical risk
Chi Jin, Lydia T Liu, Rong Ge, and Michael I Jordan · 2018
Later among the works it cites.
Improved rectangular matrix multiplication using powers of the coppersmith-winograd tensor
Francois Le Gall and Florent Urrutia · 2018
Later among the works it cites.
Iterative refinement for ℓ p \ell_{p} -norm regression
Deeksha Adil, Rasmus Kyng, Richard Peng, and Sushant Sachdeva · 2019
Closest in time.
Solving linear programs in the current matrix multiplication time
Michael B Cohen, Yin Tat Lee, and Zhao Song · 2019
Closest in time.
(Nearly) sample-optimal sparse Fourier transform in any dimension; RIPless and Filterless
Vasileios Nakos, Zhao Song, and Zhengyu Wang · 2019
Closest in time.