Fetching the paper…
Reading the bibliography…
Logistic regression is one of the most popular methods in binary classification, wherein estimation of model parameters is carried out by solving the maximum likelihood (ML) optimization problem, and the ML estimator is defined to be the optimal solution of this problem.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition
T. M. Cover · 1965
Earlier work this paper cites.
On the existence of maximum likelihood estimators for the binomial response models
M. J. Silvapulle · 1981
Earlier work this paper cites.
On the existence of maximum likelihood estimates in logistic regression models
A. Albert and J. A. Anderson · 1984
Earlier work this paper cites.
Introduction to Optimization
B. Polyak · 1987
Earlier work this paper cites.
Some perturbation theory for linear programming
J. Renegar · 1994
Earlier work this paper cites.
Special invited paper. additive logistic regression: A statistical view of boosting
J. Friedman, T. Hastie, and R. Tibshirani · 2000
Earlier work this paper cites.
Introductory lectures on convex optimization: a basic course
Y. E. Nesterov · 2003
Earlier work this paper cites.
Smooth minimization of non-smooth functions
Y. E. Nesterov · 2005
Earlier work this paper cites.
Kernel logistic regression and the import vector machine
J. Zhu and T. Hastie · 2005
Earlier work this paper cites.
Elements of Statistical Learning: Data Mining, Inference, and Prediction
T. Hastie, R. Tibshirani, and J. Friedman · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Cited alongside, same era.
Lectures on stochastic programming: modeling and theory
A. Shapiro, D. Dentcheva, and A. Ruszczyński · 2009
Cited alongside, same era.
Self-concordant analysis for logistic regression
F. Bach · 2010
Cited alongside, same era.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Cited alongside, same era.
An optimal method for stochastic composite optimization
G. Lan · 2012
Cited alongside, same era.
Efficiency of coordinate descent methods on huge-scale optimization problems
Y. Nesterov · 2012
Cited alongside, same era.
Iteration complexity of randomized block-coordinate descent methods for minimizing a composite function
P. Richtárik and M. Takáč · 2014
Later among the works it cites.
New analysis and results for the Frank-Wolfe method
R. Freund and P. Grigas · 2016
Later among the works it cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Later among the works it cites.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, and N. Srebro · 2017
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic gradient method with an exponential convergence rate for finite training sets
N. L. Roux, M. Schmidt, and F. Bach · 2012
Cited alongside, same era.
Non-strongly-convex smooth stochastic approximation with convergence rate O(1/n)
F. Bach and E. Moulines · 2013
Cited alongside, same era.
On the convergence of block coordinate descent type methods
A. Beck and L. Tetruashvili · 2013
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
Adaptivity of averaged stochastic gradient descent to local strong convexity for logistic regression
F. Bach · 2014
Cited alongside, same era.
E. J. Candès and P. Sur · 2018
Closest in time.
Characterizing implicit bias in terms of optimization geometry
S. Gunasekar, J. Lee, D. Soudry, and N. Srebro · 2018
Closest in time.
Risk and parameter convergence of logistic regression
Z. Ji and M. Telgarsky · 2018
Closest in time.
Convergence of gradient descent on separable data
M. S. Nacson, J. Lee, S. Gunasekar, P. H. P. Savarese, N. Srebro, and D. Soudry · 2018
Closest in time.
Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate
M. S. Nacson, N. Srebro, and D. Soudry · 2018
Closest in time.