Fetching the paper…
Reading the bibliography…
We introduce and analyze stochastic optimization methods where the input to each gradient update is perturbed by bounded noise.
Chaotic relaxation
Daniel Chazan and Willard Miranker · 1969
Earlier work this paper cites.
Distributed asynchronous deterministic and stochastic gradient optimization algorithms
John N Tsitsiklis, Dimitri P Bertsekas, and Michael Athans · 1986
Earlier work this paper cites.
Parallel and distributed computation: numerical methods
Dimitri P Bertsekas and John N Tsitsiklis · 1989
Earlier work this paper cites.
Rcv1: A new benchmark collection for text categorization research
David D Lewis, Yiming Yang, Tony G Rose, and Fan Li · 2004
Earlier work this paper cites.
The WebGraph framework I: Compression techniques
Paolo Boldi and Sebastiano Vigna · 2004
Earlier work this paper cites.
Ubicrawler: A scalable fully distributed web crawler
Paolo Boldi, Bruno Codenotti, Massimo Santini, and Sebastiano Vigna · 2004
Earlier work this paper cites.
Slow learners are fast
Martin Zinkevich, John Langford, and Alex J Smola · 2009
Earlier work this paper cites.
Identifying suspicious urls: an application of large-scale online learning
Justin Ma, Lawrence K Saul, Stefan Savage, and Geoffrey M Voelker · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Feng Niu, Benjamin Recht, Christopher Re, and Stephen Wright · 2011
Earlier work this paper cites.
Large-scale matrix factorization with distributed stochastic gradient descent
Rainer Gemulla, Erik Nijkamp, Peter J Haas, and Yannis Sismanis · 2011
Earlier work this paper cites.
Distributed delayed stochastic optimization
Alekh Agarwal and John C Duchi · 2011
Earlier work this paper cites.
Layered label propagation: A multiresolution coordinate-free ordering for compressing social networks
Paolo Boldi, Marco Rosa, Massimo Santini, and Sebastiano Vigna · 2011
Cited alongside, same era.
Factoring nonnegative matrices with linear programs
Ben Recht, Christopher Re, Joel Tropp, and Victor Bittorf · 2012
Cited alongside, same era.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Andrew Senior, Paul Tucker, Ke Yang, Quoc V Le, et al · 2012
Cited alongside, same era.
Parallel coordinate descent methods for big data optimization
Peter Richtárik and Martin Takáč · 2012
Cited alongside, same era.
A fast parallel sgd for matrix factorization in shared memory systems
Yong Zhuang, Wei-Sheng Chin, Yu-Chin Juan, and Chih-Jen Lin · 2013
Cited alongside, same era.
Mingyi Hong · 2014
Later among the works it cites.
An asynchronous parallel randomized kaczmarz algorithm
Ji Liu, Stephen J Wright, and Srikrishna Sridhar · 2014
Later among the works it cites.
Revisiting asynchronous linear solvers: Provable convergence rate through randomization
Haim Avron, Alex Druinsky, and Anshul Gupta · 2014
Later among the works it cites.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
Deanna Needell, Rachel Ward, and Nati Srebro · 2014
Later among the works it cites.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
Elad Hazan and Satyen Kale · 2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hyokun Yun, Hsiang-Fu Yu, Cho-Jui Hsieh, SVN Vishwanathan, and Inderjit Dhillon · 2013
Cited alongside, same era.
Estimation, optimization, and parallelism when data is sparse
John Duchi, Michael I Jordan, and Brendan McMahan · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
An asynchronous parallel stochastic coordinate descent algorithm
Ji Liu, Steve Wright, Christopher Re, Victor Bittorf, and Srikrishna Sridhar · 2014
Cited alongside, same era.
Asynchronous parallel block-coordinate frank-wolfe
Yu-Xiang Wang, Veeranjaneyulu Sadhanala, Wei Dai, Willie Neiswanger, Suvrit Sra, and Eric P Xing · 2014
Cited alongside, same era.
Project adam: Building an efficient and scalable deep learning training system
Trishul Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman · 2014
Cited alongside, same era.
Communication-efficient distributed dual coordinate ascent
Martin Jaggi, Virginia Smith, Martin Takác, Jonathan Terhorst, Sanjay Krishnan, Thomas Hofmann, and Michael I Jordan · 2014
Cited alongside, same era.
Theory of convex optimization for machine learning
Sébastien Bubeck · 2014
Later among the works it cites.
Passcode: Parallel asynchronous stochastic dual co-ordinate descent
Cho-Jui Hsieh, Hsiang-Fu Yu, and Inderjit S Dhillon · 2015
Closest in time.
Asynchronous stochastic coordinate descent: Parallelism and convergence properties
Ji Liu and Stephen J Wright · 2015
Closest in time.
An asynchronous mini-batch algorithm for regularized stochastic optimization
Hamid Reza Feyzmahdavian, Arda Aytekin, and Mikael Johansson · 2015
Closest in time.
Asynchronous parallel stochastic gradient for nonconvex optimization
Xiangru Lian, Yijun Huang, Yuncheng Li, and Ji Liu · 2015
Closest in time.
ARock: an Algorithmic Framework for Asynchronous Parallel Coordinate Updates
Zhimin Peng, Yangyang Xu, Ming Yan, and Wotao Yin · 2015
Closest in time.