Fetching the paper…
Reading the bibliography…
We study optimization algorithms for the finite sum problems frequently arising in machine learning applications.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Methods of conjugate gradients for solving linear systems
Magnus R. Hestenes and Eduard Stiefel · 1952
Earlier work this paper cites.
Fonctions convexes duales et points proximaux dans un espace hilbertien
Jean-Jacques Moreau · 1962
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Convex Analysis
Ralph T Rockafellar · 1970
Earlier work this paper cites.
On solutions of stochastic programming problems by descent procedures with stochastic and deterministic directions
Kurt Marti and Erich Fuchs · 1979
Earlier work this paper cites.
Distributed asynchronous computation of fixed points
Dimitri P Bertsekas · 1983
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
Problems in decentralized decision making and computation
John N Tsitsiklis · 1984
Earlier work this paper cites.
Rates of convergence of semi-stochastic approximation procedures for solving stochastic optimization problems
Kurt Marti and Erich Fuchs · 1986
Earlier work this paper cites.
Two-point step size gradient methods
Jonathan Barzilai and Jonathan M Borwein · 1988
Earlier work this paper cites.
Parallel and distributed computation: numerical methods
Dimitri P Bertsekas and John N Tsitsiklis · 1989
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Dong C Liu and Jorge Nocedal · 1989
Earlier work this paper cites.
A limited memory algorithm for bound constrained optimization
Richard H Byrd, Peihuang Lu, Jorge Nocedal, and Ciyou Zhu · 1995
Earlier work this paper cites.
Statistical Learning Theory
Vladimir N Vapnik · 1998
Earlier work this paper cites.
An overview of statistical learning theory
Vladimir N Vapnik · 1999
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course
Yurii Nesterov · 2004
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
Tong Zhang · 2004
Earlier work this paper cites.
Quantized incremental algorithms for distributed optimization
M.G. Rabbat and R.D. Nowak · 2005
Earlier work this paper cites.
Convergence theory for nonconvex stochastic programming with an application to mixed logit
Fabian Bastin, Cinzia Cirillo, and Philippe L Toint · 2006
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen J. Wright · 2006
Earlier work this paper cites.
Regularization tools version 4.0 for matlab 7.3
Per Christian Hansen · 2007
Earlier work this paper cites.
Gradient methods for minimizing composite objective function
Yurii Nesterov · 2007
Earlier work this paper cites.
The tradeoffs of large scale learning
Olivier Bousquet and Léon Bottou · 2008
Earlier work this paper cites.
Lazy sparse stochastic gradient descent for regularized multinomial logistic regression
Bob Carpenter · 2008
Earlier work this paper cites.
MapReduce: Simplified data processing on large clusters
Jeffrey Dean and Sanjay Ghemawat · 2008
Earlier work this paper cites.
A dual coordinate descent method for large-scale linear svm
Cho-Jui Hsieh, Kai-Wei Chang, Chih-Jen Lin, S Sathiya Keerthi, and Sellamanickam Sundararajan · 2008
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Amir Beck and Marc Teboulle · 2009
Earlier work this paper cites.
SGD-QN: Careful quasi-Newton stochastic gradient descent
Antoine Bordes, Léon Bottou, and Patrick Gallinari · 2009
Earlier work this paper cites.
Curiously fast convergence of some stochastic gradient descent algorithms
Léon Bottou · 2009
Earlier work this paper cites.
Variable-number sample-path optimization
Geng Deng and Michael C Ferris · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Sparse online learning via truncated gradient
John Langford, Lihong Li, and Tong Zhang · 2009
Earlier work this paper cites.
Efficient large-scale distributed training of conditional maximum entropy models
Ryan Mcdonald, Mehryar Mohri, Nathan Silberman, Dan Walker, and Gideon S Mann · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
A Randomized Kaczmarz Algorithm with Exponential Convergence
Thomas Strohmer and Roman Vershynin · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Stephen Boyd, Neal Parikh, Eric Chu, Borja Peleato, and Jonathan Eckstein · 2010
Earlier work this paper cites.
Consensus-based distributed support vector machines
Pedro A Forero, Alfonso Cano, and Georgios B Giannakis · 2010
Earlier work this paper cites.
Randomized methods for linear constraints: convergence rates and conditioning
Dennis Leventhal and Adrian S Lewis · 2010
Earlier work this paper cites.
Randomized kaczmarz solver for noisy linear systems
Deanna Needell · 2010
Earlier work this paper cites.
Exascale computing technology challenges
John Shalf, Sudip Dosanjh, and John Morrison · 2010
Earlier work this paper cites.
Spark: cluster computing with working sets
Matei Zaharia, Mosharaf Chowdhury, Michael J Franklin, Scott Shenker, and Ion Stoica · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Scaling up machine learning: Parallel and distributed approaches
Ron Bekkerman, Mikhail Bilenko, and John Langford · 2011
Earlier work this paper cites.
Parallel coordinate descent for L1-regularized loss minimization
Joseph Bradley, Aapo Kyrola, Daniel Bickson, and Carlos Guestrin · 2011
Earlier work this paper cites.
Differentially private empirical risk minimization
Kamalika Chaudhuri, Claire Monteleoni, and Anand D Sarwate · 2011
Earlier work this paper cites.
Proximal splitting methods in signal processing
Patrick Louis Combettes and Jean-Christophe Pesquet · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R Bach · 2011
Earlier work this paper cites.
On optimization methods for deep learning
Jiquan Ngiam, Adam Coates, Ahbik Lahiri, Bobby Prochnow, Quoc V Le, and Andrew Y Ng · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Feng Niu, Benjamin Recht, Christopher Ré, and Stephen J Wright · 2011
Earlier work this paper cites.
Solving large scale linear svm with distributed block minimization
Dmitry Pechyony, Libin Shen, and Rosie Jones · 2011
Earlier work this paper cites.
Pegasos: Primal estimated sub-gradient solver for SVM
Shai Shalev-Shwartz, Yoram Singer, Nathan Srebro, and Andrew Cotter · 2011
Earlier work this paper cites.
Distributed learning, communication complexity and privacy
Maria-Florina Balcan, Avrim Blum, Shai Fine, and Yishay Mansour · 2012
Earlier work this paper cites.
Stochastic gradient descent tricks
Léon Bottou · 2012
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Ofer Dekel, Ran Gilad-Bachrach, Ohad Shamir, and Lin Xiao · 2012
Earlier work this paper cites.
Dual averaging for distributed optimization: convergence analysis and network scaling
John C Duchi, Alekh Agarwal, and Martin J Wainwright · 2012
Earlier work this paper cites.
Hybrid deterministic-stochastic methods for data fitting
Michael P Friedlander and Mark Schmidt · 2012
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Efficiency of coordinate descent methods on huge-scale optimization problems
Yurii Nesterov · 2012
Earlier work this paper cites.
A stochastic gradient method with an exponential convergence rate for finite training sets
Nicolas Le Roux, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Large linear classification when data cannot fit in memory
Hsiang-Fu Yu, Cho-Jui Hsieh, Kai-Wei Chang, and Chih-Jen Lin · 2012
Earlier work this paper cites.
Recent advances of large-scale linear classification
Guo-Xun Yuan, Chia-Hua Ho, and Chih-Jen Lin · 2012
Earlier work this paper cites.
Efficient distributed linear classification algorithms via the alternating direction method of multipliers
C Zhang, H Lee, and K G Shin · 2012
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Yuchen Zhang, Martin J. Wainwright, and John C Duchi · 2012
Earlier work this paper cites.
Where (and when) do you use your smartphone: Bedroom? Church?
CNN · 2013
Earlier work this paper cites.
Predicting parameters in deep learning
Misha Denil, Babak Shakibi, Laurent Dinh, Marc’Aurelio Ranzato, and Nando de Freitas · 2013
Cited alongside, same era.
Estimation, optimization, and parallelism when data is sparse
John C Duchi, Michael I Jordan, and Brendan H McMahan · 2013
Cited alongside, same era.
Smooth minimization of nonsmooth functions with parallel coordinate descent methods
Olivier Fercoq and Peter Richtárik · 2013
Cited alongside, same era.
Large-scale learning with less ram via randomization
Daniel Golovin, D Sculley, Brendan H McMahan, and Michael Young · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Partitioning data on features or samples in communication-efficient distributed optimization?
Chenxin Ma and Martin Takáč · 2015
Later among the works it cites.
Chenxin Ma, Rachael Tappenden, and Martin Takáč · 2015
Later among the works it cites.
Perturbed iterate analysis for asynchronous stochastic optimization
Horia Mania, Xinghao Pan, Dimitris Papailiopoulos, Benjamin Recht, Kannan Ramchandran, and Michael I Jordan · 2015
Later among the works it cites.
Distributed block coordinate descent for minimizing partially separable functions
Jakub Mareček, Peter Richtárik, and Martin Takáč · 2015
Later among the works it cites.
Convergence analysis for Kaczmarz-type methods in a Hilbert space framework
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jakub Konečný and Peter Richtárik · 2013
Cited alongside, same era.
Efficient accelerated coordinate descent methods and faster algorithms for solving linear systems
Yin Tat Lee and Aaron Sidford · 2013
Cited alongside, same era.
An efficient distributed learning algorithm based on effective local functional approximations
Dhruv Mahajan, Nikunj Agrawal, S Sathiya Keerthi, Sundararajan Sellamanickam, and Léon Bottou · 2013
Cited alongside, same era.
Ad click prediction: a view from the trenches
H Brendan McMahan, Gary Holt, David Sculley, Michael Young, Dietmar Ebner, Julian Grady, Lan Nie, Todd Phillips, Eugene Davydov, Daniel Golovin, et al · 2013
Cited alongside, same era.
Distributed coordinate descent methods for composite minimization
Ion Necoara and Dragos Clipici · 2013
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
Mark Schmidt, Nicolas Le Roux, and Francis Bach · 2013
Cited alongside, same era.
Stochastic dual coordinate ascent methods for regularized loss
Shai Shalev-Shwartz and Tong Zhang · 2013
Cited alongside, same era.
Peter Oswald and Weiqi Zhou · 2015
Later among the works it cites.
Quartz: Randomized dual coordinate ascent with arbitrary sampling
Zheng Qu, Peter Richtárik, and Tong Zhang · 2015
Later among the works it cites.
On variance reduction in stochastic gradient descent and its asynchronous variants
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabás Póczós, and Alex Smola · 2015
Later among the works it cites.
Non-uniform stochastic average gradient method for training conditional random fields
Mark Schmidt, Reza Babanezhad, Mohamed Ahmed, Aaron Defazio, Ann Clifton, and Anoop Sarkar · 2015
Later among the works it cites.
Shai Shalev-Shwartz · 2015
Later among the works it cites.
Martin Takáč, Peter Richtárik, and Nathan Srebro · 2015
Later among the works it cites.
Separable approximations and decomposition methods for the augmented Lagrangian
Rachael Tappenden, Peter Richtárik, and Burak Büke · 2015
Later among the works it cites.
On the complexity of parallel coordinate descent
Rachael Tappenden, Martin Takáč, and Peter Richtárik · 2015
Later among the works it cites.
Coordinate descent algorithms
Stephen J Wright · 2015
Later among the works it cites.
DiSCO: Distributed optimization for self-concordant empirical loss
Yuchen Zhang and Xiao Lin · 2015
Later among the works it cites.
Stochastic optimization with importance sampling for regularized loss minimization
Peilin Zhao and Tong Zhang · 2015
Later among the works it cites.
Distributed newton methods for regularized logistic regression
Yong Zhuang, Wei-Sheng Chin, Yu-Chin Juan, and Chih-Jen Lin · 2015
Later among the works it cites.
http://www.speedtest.net/reports/united-states/ , August 2016
Speedtest market report · 2016
Later among the works it cites.
http://www.tensorflow.org/tutorials/deep_cnn , 2016
Tensorflow convolutional neural networks tutorial · 2016
Later among the works it cites.
Deep learning with differential privacy
Martín Abadi, Andy Chu, Ian Goodfellow, Brendan H McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang · 2016
Later among the works it cites.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2016
Later among the works it cites.
Exploiting the structure: Stochastic gradient methods using raw clusters
Zeyuan Allen-Zhu, Yang Yuan, and Karthik Sridharan · 2016
Later among the works it cites.
Practical secure aggregation for federated learning on user-held data
Keith Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone, Brendan H McMahan, Sarvar Patel, Daniel Ramage, Aaron Segal, and Karn Seth · 2016
Later among the works it cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Later among the works it cites.
A stochastic quasi-Newton method for large-scale optimization
Richard H Byrd, Samantha L Hansen, Jorge Nocedal, and Yoram Singer · 2016
Later among the works it cites.
Revisiting distributed synchronous SGD
Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Later among the works it cites.
Revisiting distributed synchronous sgd
Jianmin Chen, Rajat Monga, Samy Bengio, and Rafal Jozefowicz · 2016
Later among the works it cites.
Coordinate descent face-off: primal or dual?
Dominik Csiba and Peter Richtárik · 2016
Later among the works it cites.
Importance sampling for minibatches
Dominik Csiba and Peter Richtárik · 2016
Later among the works it cites.
A simple practical accelerated method for finite sums
Aaron Defazio · 2016
Later among the works it cites.
On the global and linear convergence of the generalized alternating direction method of multipliers
Wei Deng and Wotao Yin · 2016
Later among the works it cites.
Primal-Dual Rates and Certificates
Celestine Dünner, Simone Forte, Martin Takáč, and Martin Jaggi · 2016
Later among the works it cites.
Optimization in high dimensions via accelerated, parallel, and proximal coordinate descent
Olivier Fercoq and Peter Richtárik · 2016
Later among the works it cites.
On randomized distributed coordinate descent with quantized updates
Mostafa El Gamal and Lifeng Lai · 2016
Later among the works it cites.
Stochastic block bfgs: squeezing more curvature out of data
Robert M Gower, Donald Goldfarb, and Peter Richtárik · 2016
Later among the works it cites.
Randomized quasi-Newton updates are linearly convergent matrix inversion algorithms
Robert M Gower and Peter Richtárik · 2016
Later among the works it cites.
DUAL-LOCO: Distributing Statistical Estimation Using Random Projections
Christina Heinze, Brian McWilliams, and Nicolai Meinshausen · 2016
Later among the works it cites.
Mini-batch semi-stochastic gradient descent in the proximal setting
Jakub Konečný, Jie Liu, Peter Richtárik, and Martin Takáč · 2016
Later among the works it cites.
Federated optimization: Distributed machine learning for on-device intelligence
Jakub Konečný, Brendan H McMahan, Daniel Ramage, and Peter Richtárik · 2016
Later among the works it cites.
Randomized distributed mean estimation: Accuracy vs communication
Jakub Konečný and Peter Richtárik · 2016
Later among the works it cites.
Federated learning: Strategies for improving communication efficiency
Jakub Konečný, Brendan H McMahan, Felix X Yu, Peter Richtárik, Ananda Theertha Suresh, and Dave Bacon · 2016
Later among the works it cites.
ASAGA: Asynchronous parallel saga
Rémi Leblond, Fabian Pedregosa, and Simon Lacoste-Julien · 2016
Later among the works it cites.
Distributed inexact damped newton method: Data partitioning and load-balancing
Chenxin Ma and Martin Takáč · 2016
Later among the works it cites.
Federated learning of deep networks using model averaging
Brendan H McMahan, Eider Moore, Daniel Ramage, and Blaise Aguera y Arcas · 2016
Later among the works it cites.
A linearly-convergent stochastic L-BFGS algorithm
Philipp Moritz, Robert Nishihara, and Michael Jordan · 2016
Later among the works it cites.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
Deanna Needell, Nathan Srebro, and Rachel Ward · 2016
Later among the works it cites.
ARock: an algorithmic framework for asynchronous parallel coordinate updates
Zhimin Peng, Yangyang Xu, Ming Yan, and Wotao Yin · 2016
Later among the works it cites.
Coordinate descent with arbitrary sampling I: Algorithms and complexity
Zheng Qu and Peter Richtárik · 2016
Later among the works it cites.
Coordinate descent with arbitrary sampling II: Expected separable overapproximation
Zheng Qu and Peter Richtárik · 2016
Later among the works it cites.
SDNA: Stochastic dual newton ascent for empirical risk minimization
Zheng Qu, Peter Richtárik, Martin Takáč, and Olivier Fercoq · 2016
Later among the works it cites.
Stochastic variance reduction for nonconvex optimization
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabás Póczós, and Alex Smola · 2016
Later among the works it cites.
AIDE: Fast and communication efficient distributed optimization
Sashank J Reddi, Jakub Konečný, Peter Richtárik, Barnabás Póczós, and Alex Smola · 2016
Later among the works it cites.
Distributed coordinate descent method for learning with big data
Peter Richtárik and Martin Takáč · 2016
Later among the works it cites.
On optimal probabilities in stochastic coordinate descent methods
Peter Richtárik and Martin Takáč · 2016
Later among the works it cites.
Parallel coordinate descent methods for big data optimization
Peter Richtárik and Martin Takáč · 2016
Later among the works it cites.
Stochastic reformulation of linear systems and fast stochastic iterative methods
Peter Richtárik and Martin Takáč · 2016
Later among the works it cites.
SDCA without duality, regularization, and individual convexity
Shai Shalev-Shwartz · 2016
Later among the works it cites.
Distributed mean estimation with limited communication
Ananda Theertha Suresh, Felix X Yu, Brendan H McMahan, and Sanjiv Kumar · 2016
Later among the works it cites.
Inexact coordinate descent: complexity and preconditioning
Rachael Tappenden, Peter Richtárik, and Jacek Gondzio · 2016
Later among the works it cites.
Variable-length quantity, 2016
Wikipedia · 2016
Later among the works it cites.
Tight complexity bounds for optimizing composite objectives
Blake Woodworth and Nathan Srebro · 2016
Later among the works it cites.
Orthogonal random features
Felix X Yu, Ananda Theertha Suresh, Krzysztof Choromanski, Daniel Holtmann-Rice, and Sanjiv Kumar · 2016
Later among the works it cites.
Practical secure aggregation for privacy preserving machine learning
Keith Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone, H Brendan McMahan, Sarvar Patel, Daniel Ramage, Aaron Segal, and Karn Seth · 2017
Closest in time.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Closest in time.
Privacy preserving randomized gossip algorithms
Filip Hanzely, Jakub Konečný, Nicolas Loizou, Peter Richtárik, and Dmitry Grishchenko · 2017
Closest in time.
Federated learning: Collaborative machine learning without centralized training data
H. Brendan McMahan and Daniel Ramage · 2017
Closest in time.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
Lam Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Closest in time.
Exploiting strong convexity from data with primal-dual first-order algorithms
Jialei Wang and Lin Xiao · 2017
Closest in time.