Fetching the paper…
Reading the bibliography…
In this paper, we propose a Distributed Accumulated Newton Conjugate gradiEnt (DANCE) method in which sample size is gradually increasing to quickly obtain a solution whose empirical loss is under satisfactory statistical accuracy.
Quasi-newton methods for deep learning: Forget the past, just sample
Berahas, A. S., Jahani, M., and Takáč, M. (2019) · 1901
Earlier work this paper cites.
Scaling up quasi-newton algorithms: Communication efficient distributed sr1
Jahani, M., Nazari, M., Rusakov, S., Berahas, A. S., and Takáč, M. (2019) · 1905
Earlier work this paper cites.
Concentration Inequalities and Empirical Processes Theory Applied to the Analysis of Learning Algorithms
Bousquet, o. (2002) · 2002
Earlier work this paper cites.
Convex optimization
Boyd, S. and Vandenberghe, L. (2004) · 2004
Earlier work this paper cites.
Convexity, classification, and risk bounds
Bartlett, P. L., Jordan, M. I., and McAuliffe, J. D. (2006) · 2006
Earlier work this paper cites.
Sequential quadratic programming
Nocedal, J. and Wright, S. J. (2006) · 2006
Earlier work this paper cites.
Numerical recipes: the art of scientific computing, 3rd Edition
Press, W. H., Teukolsky, S. A., Vetterling, W. T., and Flannery, B. P. (2007) · 2007
Earlier work this paper cites.
A stochastic quasi-Newton method for online convex optimization
Schraudolph, N. N., Yu, J., and Günter, S. (2007) · 2007
Earlier work this paper cites.
The tradeoffs of large scale learning
Bousquet, O. and Bottou, L. (2008) · 2008
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Beck, A. and Teboulle, M. (2009) · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L. (2010) · 2010
Earlier work this paper cites.
LIBSVM: A library for support vector machines
Chang, C.-C. and Lin, C.-J. (2011) · 2011
Earlier work this paper cites.
Parallel distributed computing using python
Dalcin, L. D., Paz, R. R., Kler, P. A., and Cosimo, A. (2011) · 2011
Cited alongside, same era.
A stochastic gradient method with an exponential convergence rate for finite training sets
Roux, N. L., Schmidt, M., and Bach, F. R. (2012) · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T. (2013) · 2013
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course
Nesterov, Y. (2013) · 2013
Cited alongside, same era.
Stochastic dual coordinate ascent methods for regularized loss minimization
Shalev-Shwartz, S. and Zhang, T. (2013) · 2013
Cited alongside, same era.
The nature of statistical learning theory
Vapnik, V. (2013) · 2013
Cited alongside, same era.
A stochastic quasi-Newton method for large-scale optimization
Byrd, R. H., Hansen, S. L., Nocedal, J., and Singer, Y. (2016) · 2016
Later among the works it cites.
Distributed inexact damped Newton method: Data partitioning and load-balancing
Ma, C. and Takáč, M. (2016) · 2016
Later among the works it cites.
Adaptive Newton method for empirical risk minimization to statistical accuracy
Mokhtari, A., Daneshmand, H., Lucchi, A., Hofmann, T., and Ribeiro, A. (2016) · 2016
Later among the works it cites.
Semi-stochastic gradient descent methods
Konečnỳ, J. and Richtárik, P. (2017) · 2017
Later among the works it cites.
Underestimate sequences via quadratic averaging
Ma, C., Gudapati, N. V. C., Jahani, M., Tappenden, R., and Takáč, M. (2017) · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S. (2014) · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J. (2014) · 2014
Cited alongside, same era.
Global convergence of online limited memory bfgs
Mokhtari, A. and Ribeiro, A. (2015) · 2015
Cited alongside, same era.
Disco: Distributed optimization for self-concordant empirical loss
Zhang, Y. and Lin, X. (2015) · 2015
Cited alongside, same era.
A multi-batch l-bfgs method for machine learning
Berahas, A. S., Nocedal, J., and Takác, M. (2016) · 2016
Cited alongside, same era.
Sampled quasi-newton methods for deep learning
Berahas, A. S., Jahani, M., and Takác, M
Cited in the paper.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
Nguyen, L. M., Liu, J., Scheinberg, K., and Takáč, M. (2017) · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2017) · 2017
Later among the works it cites.
An optimal first order method based on optimal quadratic averaging
Drusvyatskiy, D., Fazel, M., and Roy, S. (2018) · 2018
Closest in time.
Large scale empirical risk minimization via truncated adaptive Newton method
Eisen, M., Mokhtari, A., and Ribeiro, A. (2018) · 2018
Closest in time.
Sub-sampled newton methods
Roosta-Khorasani, F. and Mahoney, M. W. (2018) · 2018
Closest in time.
First-order adaptive sample size methods to reduce complexity of empirical risk minimization
Mokhtari, A. and Ribeiro, A. (2017) · 2065
Closest in time.