Fetching the paper…
Reading the bibliography…
We propose an efficient protocol for decentralized training of deep neural networks from distributed data sources.
Hypercube Multiprocessors 86(181-195), 31 (1986)
Moler, C.: Matrix computation on distributed memory multiprocessors · 1986
Earlier work this paper cites.
In: Advances in Neural Information Processing Systems. pp. 305–313 (1989)
Pomerleau, D.A.: Alvinn: An autonomous land vehicle in a neural network · 1989
Earlier work this paper cites.
Proceedings of Neuro-Nımes 91(8), 0 (1991)
Bottou, L.: Stochastic gradient learning in neural networks · 1991
Earlier work this paper cites.
http://yann. lecun. com/exdb/mnist/ (1998)
LeCun, Y.: The mnist database of handwritten digits · 1998
Earlier work this paper cites.
Transactions on Database Systems 32(4) (2007)
Sharfman, I., Schuster, A., Keren, D.: A geometric approach to monitoring threshold functions over distributed data streams · 2007
Earlier work this paper cites.
In: Proceedings of the ACM SIGMOD-SIGACT-SIGART Symposium on Principles of Database Systems. pp. 301–310. ACM (2008)
Sharfman, I., Schuster, A., Keren, D.: Shape sensitive geometric monitoring · 2008
Earlier work this paper cites.
In: Advances in Neural Information Processing Systems. pp. 1231–1239 (2009)
Mcdonald, R., Mohri, M., Silberman, N., Walker, D., Mann, G.S.: Efficient large-scale distributed training of conditional maximum entropy models · 2009
Earlier work this paper cites.
Proceedings of the VLDB Endowment 4(2), 46–57 (2010)
Sagy, G., Keren, D., Sharfman, I., Schuster, A.: Distributed threshold querying of general functions by a difference of monotonic representation · 2010
Earlier work this paper cites.
In: Proceedings of the International Conference on Artificial Intelligence and Statistics (2010)
Xavier Glorot, Y.B.: Understanding the difficulty of training deep feedforward neural networks · 2010
Earlier work this paper cites.
In: Advances in Neural Information Processing Systems. pp. 2595–2603 (2010)
Zinkevich, M., Weimer, M., Smola, A.J., Li, L.: Parallelized stochastic gradient descent · 2010
Earlier work this paper cites.
In: Proceedings of the International Conference on Supercomputing. pp. 120–129. ACM (2011)
Verner, U., Schuster, A., Silberstein, M.: Processing data streams with hard real-time constraints on heterogeneous systems · 2011
Earlier work this paper cites.
Machine Learning 86(2), 209–231 (2012)
Bshouty, N.H., Long, P.M.: Linear classifiers are nearly optimal when hidden variables have diverse effects · 2012
Earlier work this paper cites.
In: Advances in Neural Information Processing Systems. pp. 1223–1231 (2012)
Dean, J., Corrado, G.S., Monga, R., Chen, K., Devin, M., Quoc, V.L., Mao, M.Z., Ranzato, M., Senior, A., Tucker, P., Yang, K., Ng, A.Y.: Large scale distributed deep networks · 2012
Earlier work this paper cites.
Journal of Machine Learning Research 13, 165–202 (2012)
Dekel, O., Gilad-Bachrach, R., Shamir, O., Xiao, L.: Optimal distributed online prediction using mini-batches · 2012
Earlier work this paper cites.
In: Proceedings of the 2012 ACM SIGMOD International Conference on Management of Data. pp. 265–276. ACM (2012)
Giatrakos, N., Deligiannakis, A., Garofalakis, M., Sharfman, I., Schuster, A.: Prediction-based geometric monitoring over distributed data streams · 2012
Earlier work this paper cites.
IEEE Transactions on Knowledge and Data Engineering 24(8), 1520–1535 (2012)
Keren, D., Sharfman, I., Schuster, A., Livne, A.: Shape sensitive geometric monitoring · 2012
Cited alongside, same era.
COURSERA: Neural networks for machine learning 4(2), 26–31 (2012)
Tieleman, T., Hinton, G.: Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude · 2012
Cited alongside, same era.
In: Advances in Neural Information Processing Systems. pp. 1502–1510 (2012)
Zhang, Y., Wainwright, M.J., Duchi, J.C.: Communication-efficient algorithms for statistical optimization · 2012
Cited alongside, same era.
In: BD3@ VLDB. pp. 13–18 (2013)
Boley, M., Kamp, M., Keren, D., Schuster, A., Sharfman, I.: Communication-efficient distributed online prediction using dynamic model synchronizations · 2013
Cited alongside, same era.
Springer Science & Business Media (2013)
Nesterov, Y.: Introductory lectures on convex optimization: A basic course, vol. 87 · 2013
Cited alongside, same era.
In: Machine Learning and Knowledge Discovery in Databases. pp. 805–819. Springer (2016)
Kamp, M., Bothe, S., Boley, M., Mock, M.: Communication-efficient distributed online learning with kernels · 2016
Later among the works it cites.
In: Advances in Neural Information Processing Systems. pp. 46–54 (2016)
Shamir, O.: Without-replacement sampling for stochastic gradient methods · 2016
Later among the works it cites.
Feng, J., Xu, H., Mannor, S.: Outlier robust online learning · 2017
Later among the works it cites.
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 214–221 (2017)
Fernando, T., Denman, S., Sridharan, S., Fookes, C.: Going deeper: Autonomous steering with neural memory networks · 2017
Later among the works it cites.
In: Advances in Neural Information Processing Systems. pp. 5904–5914 (2017)
Jiang, Z., Balu, A., Hegde, C., Sarkar, S.: Collaborative deep learning in fixed topology networks · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
In: Proceedings of the 28th International Parallel and Distributed Processing Symposium. pp. 37–47. IEEE (2014)
Gabel, M., Keren, D., Schuster, A.: Communication-efficient distributed variance monitoring and outlier detection for multivariate time series · 2014
Cited alongside, same era.
IEEE Transactions on Knowledge and Data Engineering 26(8), 1890–1903 (2014)
Keren, D., Sagy, G., Abboud, A., Ben-David, D., Schuster, A., Sharfman, I., Deligiannakis, A.: Geometric monitoring of heterogeneous streams · 2014
Cited alongside, same era.
In: Proceedings of the 3rd International Conference on Learning Representations (2014)
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization · 2014
Cited alongside, same era.
In: International Conference on Machine Learning. pp. 1000–1008 (2014)
Shamir, O., Srebro, N., Zhang, T.: Communication-efficient distributed optimization using an approximate newton-type method · 2014
Cited alongside, same era.
https://keras.io (2015)
Chollet, F., et al.: Keras · 2015
Cited alongside, same era.
Proceedings of the VLDB Endowment 8(5), 545–556 (2015)
Lazerson, A., Sharfman, I., Keren, D., Schuster, A., Garofalakis, M., Samoladas, V.: Monitoring distributed streams using convex decompositions · 2015
Cited alongside, same era.
In: Advances in Neural Information Processing Systems. pp. 685–693 (2015)
Zhang, S., Choromanska, A.E., LeCun, Y.: Deep learning with elastic averaging sgd · 2015
Cited alongside, same era.
Later among the works it cites.
In: Advances in Neural Information Processing Systems. pp. 6480–6491 (2017)
Kamp, M., Boley, M., Missura, O., Gärtner, T.: Effective parallelisation for machine learning · 2017
Later among the works it cites.
In: International Conference on Learning Representations (2017)
Keskar, N.S., Mudigere, D., Nocedal, J., Smelyanskiy, M., Tang, P.T.P.: On large-batch training for deep learning: Generalization gap and sharp minima · 2017
Later among the works it cites.
In: Artificial Intelligence and Statistics. pp. 1273–1282 (2017)
McMahan, B., Moore, E., Ramage, D., Hampson, S., y Arcas, B.A.: Communication-efficient learning of deep networks from decentralized data · 2017
Later among the works it cites.
In: International Conference on Machine Learning. pp. 2603–2612 (2017)
Nguyen, Q., Hein, M.: The loss surface of deep and wide neural networks · 2017
Later among the works it cites.
Results in Mathematics 71(3-4), 569–608 (2017)
Sanghavi, S., Ward, R., White, C.D.: The local convexity of solving systems of quadratic equations · 2017
Later among the works it cites.
In: Advances in Neural Information Processing Systems. pp. 4424–4434 (2017)
Smith, V., Chiang, C.K., Sanjabi, M., Talwalkar, A.S.: Federated multi-task learning · 2017
Later among the works it cites.
Wang, W., Srebro, N.: Stochastic nonconvex optimization with large minibatches · 2017
Later among the works it cites.
In: Proceedings of the International Conference on Learning Representations (2017)
Zhang, C., Bengio, S., Hardt, M., Recht, B., Vinyals, O.: Understanding deep learning requires rethinking generalization · 2017
Later among the works it cites.
In: International Conference on Learning Representations (2018)
McMahan, B., Ramage, D., Talwar, K., Zhang, L.: Learning differentially private recurrent language models · 2018
Closest in time.