Fetching the paper…
Reading the bibliography…
We propose a new algorithm called Parle for parallel training of deep networks that converges 2-4x faster than a data-parallel implementation of SGD, while achieving significantly improved error rates that are nearly state-of-the-art on several benchmarks including CIFAR-10 and CIFAR-100, without introducing any additional hyper-parameters.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014) · 1958
Earlier work this paper cites.
Flat minima
Hochreiter, S. and Schmidhuber, J. (1997) · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998) · 1998
Earlier work this paper cites.
A perception-driven autonomous urban vehicle
Leonard, J., How, J., Teller, S., Berger, M., Campbell, S., Fiore, G., Fletcher, L., Frazzoli, E., Huang, A., Karaman, S., et al. (2008) · 2008
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Nair, V. and Hinton, G. E. (2010) · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Recht, B., Re, C., Wright, S., and Niu, F. (2011) · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., et al. (2012) · 2012
Earlier work this paper cites.
Estimation, optimization, and parallelism when data is sparse
Duchi, J., Jordan, M. I., and McMahan, B. (2013) · 2013
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Krizhevsky, A. (2014) · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D. (2014) · 2014
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Springenberg, J., Dosovitskiy, A., Brox, T., and Riedmiller, M. (2014) · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., et al. (2015) · 2015
Cited alongside, same era.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Chen, T., Li, M., Li, Y., Lin, M., Wang, N., Wang, M., Xiao, T., Xu, B., Zhang, C., and Zhang, Z. (2015) · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C. (2015) · 2015
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
McMahan, H. B., Moore, E., Ramage, D., Hampson, S., et al. (2016) · 2016
Later among the works it cites.
Paleo: A performance model for deep neural networks
Qi, H., Sparks, E. R., and Talwalkar, A. (2016) · 2016
Later among the works it cites.
You only look once: Unified, real-time object detection
Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016) · 2016
Later among the works it cites.
Training neural networks without gradients: A scalable admm approach
Taylor, G., Burmeister, R., Xu, Z., Singh, B., Patel, A., and Goldstein, T. (2016) · 2016
Later among the works it cites.
Zagoruyko, S. and Komodakis, N. (2016) · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Moritz, P., Nishihara, R., Stoica, I., and Jordan, M. I. (2015) · 2015
Cited alongside, same era.
Deep image: Scaling up image recognition
Wu, R., Yan, S., Shan, Y., Dang, Q., and Sun, G. (2015) · 2015
Cited alongside, same era.
Entropy-SGD: Biasing Gradient Descent Into Wide Valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R. (2016) · 2016
Cited alongside, same era.
Densely connected convolutional networks
Huang, G., Liu, Z., Weinberger, K. Q., and van der Maaten, L. (2016) · 2016
Cited alongside, same era.
How to scale distributed deep learning?
Jin, P. H., Yuan, Q., Iandola, F., and Keutzer, K. (2016) · 2016
Cited alongside, same era.
Sgdr: stochastic gradient descent with restarts
Loshchilov, I. and Hutter, F. (2016) · 2016
Cited alongside, same era.
Unreasonable effectiveness of learning neural networks: From accessible states and robust ensembles to basic algorithmic schemes
Baldassi, C., Borgs, C., Chayes, J., Ingrosso, A., Lucibello, C., Saglietti, L., and Zecchina, R. (2016a)
Cited in the paper.
Learning may need only a few bits of synaptic precision
Baldassi, C., Gerace, F., Lucibello, C., Saglietti, L., and Zecchina, R. (2016b)
Cited in the paper.
Chaudhari, P., Oberman, A., Osher, S., Soatto, S., and Guillame, C. (2017) · 2017
Closest in time.
Gastaldi, X. (2017) · 2017
Closest in time.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K. (2017) · 2017
Closest in time.
Snapshot ensembles: Train 1, get m for free
Huang, G., Li, Y., Pleiss, G., Liu, Z., Hopcroft, J. E., and Weinberger, K. Q. (2017) · 2017
Closest in time.