Fetching the paper…
Reading the bibliography…
It has been experimentally observed that distributed implementations of mini-batch stochastic gradient descent (SGD) algorithms exhibit speedup saturation and decaying generalization ability beyond a particular batch-size.
Stability and generalization
O. Bousquet and A. Elisseeff · 2002
Earlier work this paper cites.
Training linear svms in linear time
T. Joachims · 2006
Earlier work this paper cites.
Efficient large-scale distributed training of conditional maximum entropy models
R. Mcdonald, M. Mohri, N. Silberman, D. Walker, and G. S. Mann · 2009
Earlier work this paper cites.
Learnability, stability and uniform convergence
S. Shalev-Shwartz, O. Shamir, N. Srebro, and K. Sridharan · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola · 2010
Earlier work this paper cites.
Better mini-batch algorithms via accelerated gradient methods
A. Cotter, O. Shamir, N. Srebro, and K. Sridharan · 2011
Earlier work this paper cites.
Large-scale matrix factorization with distributed stochastic gradient descent
R. Gemulla, E. Nijkamp, P. J. Haas, and Y. Sismanis · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
F. Niu, B. Recht, C. Re, and S. Wright · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
M. Welling and Y. W. Teh · 2011
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le, et al · 2012
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao · 2012
Earlier work this paper cites.
Hybrid deterministic-stochastic methods for data fitting
M. P. Friedlander and M. Schmidt · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Communication-efficient algorithms for statistical optimization
Y. Zhang, M. J. Wainwright, and J. C. Duchi · 2012
Earlier work this paper cites.
Estimation, optimization, and parallelism when data is sparse
J. Duchi, M. I. Jordan, and B. McMahan · 2013
Earlier work this paper cites.
Accelerated mini-batch stochastic dual coordinate ascent
S. Shalev-Shwartz and T. Zhang · 2013
Earlier work this paper cites.
Mini-batch primal and dual methods for svms
M. Takác, A. S. Bijral, P. Richtárik, and N. Srebro · 2013
Earlier work this paper cites.
Regularization of neural networks using dropconnect
L. Wan, M. Zeiler, S. Zhang, Y. L. Cun, and R. Fergus · 2013
Cited alongside, same era.
H. Yun, H.-F. Yu, C.-J. Hsieh, S. Vishwanathan, and I. Dhillon · 2013
Cited alongside, same era.
Project adam: Building an efficient and scalable deep learning training system
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman · 2014
Cited alongside, same era.
Constant step size least-mean-square: Bias-variance trade-offs and optimal sampling distributions
A. Dfossez and F. Bach · 2014
Cited alongside, same era.
Communication-efficient distributed dual coordinate ascent
M. Jaggi, V. Smith, M. Takác, J. Terhorst, S. Krishnan, T. Hofmann, and M. I. Jordan · 2014
Cited alongside, same era.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2016
Later among the works it cites.
Revisiting distributed synchronous sgd
J. Chen, R. Monga, S. Bengio, and R. Jozefowicz · 2016
Later among the works it cites.
Big batch sgd: Automated inference using adaptive batch sizes
S. De, A. Yadav, D. Jacobs, and T. Goldstein · 2016
Later among the works it cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
S. Ghadimi and G. Lan · 2016
Later among the works it cites.
Train faster, generalize better: Stability of stochastic gradient descent
M. Hardt, B. Recht, and Y. Singer · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient mini-batch training for stochastic optimization
M. Li, T. Zhang, Y. Chen, and A. J. Smola · 2014
Cited alongside, same era.
An asynchronous parallel stochastic coordinate descent algorithm
J. Liu, S. Wright, C. Re, V. Bittorf, and S. Sridhar · 2014
Cited alongside, same era.
Stochastic proximal gradient descent with acceleration techniques
A. Nitanda · 2014
Cited alongside, same era.
Communication-efficient distributed optimization using an approximate newton-type method
O. Shamir, N. Srebro, and T. Zhang · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Escaping from saddle points-online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Cited alongside, same era.
J. D. Lee, Q. Lin, T. Ma, and T. Yang · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Parallelizing stochastic approximation through mini-batching and tail-averaging
P. Jain, S. M. Kakade, R. Kidambi, P. Netrapalli, and A. Sidford · 2016
Later among the works it cites.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Lojasiewicz condition
H. Karimi, J. Nutini, and M. Schmidt · 2016
Later among the works it cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2016
Later among the works it cites.
Mini-batch semi-stochastic gradient descent in the proximal setting
J. Konečnỳ, J. Liu, P. Richtárik, and M. Takáč · 2016
Later among the works it cites.
Batched stochastic gradient descent with weighted sampling
D. Needell and R. Ward · 2016
Later among the works it cites.
Paleo: A performance model for deep neural networks, 2016
H. Qi, E. Sparks, and A. Talwalkar · 2016
Later among the works it cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Closest in time.
Algorithmic stability and hypothesis complexity
T. Liu, G. Lugosi, G. Neu, and D. Tao · 2017
Closest in time.
Memory and communication efficient distributed stochastic optimization with minibatch prox
J. Wang, W. Wang, and N. Srebro · 2017
Closest in time.
Stochastic learning on imbalanced data: Determinantal point processes for mini-batch diversification
C. Zhang, H. Kjellstrom, and S. Mandt · 2017
Closest in time.