Fetching the paper…
Reading the bibliography…
With an increasing demand for training powers for deep learning algorithms and the rapid growth of computation resources in data centers, it is desirable to dynamically schedule different distributed deep learning tasks to maximize resource utilization and reduce cost.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/k
Y. E. Nesterov · 1983
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
N. Qian · 1999
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, A. Senior, P. Tucker, K. Yang, Q. V. Le, et al · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
J. Long, E. Shelhamer, and T. Darrell · 2015
Earlier work this paper cites.
On variance reduction in stochastic gradient descent and its asynchronous variants
S. J. Reddi, A. Hefny, S. Sra, B. Poczos, and A. J. Smola · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2015
Earlier work this paper cites.
Staleness-aware async-sgd for distributed deep learning
W. Zhang, S. Gupta, X. Lian, and J. Liu · 2015
Earlier work this paper cites.
Revisiting distributed synchronous sgd
J. Chen, X. Pan, R. Monga, S. Bengio, and R. Jozefowicz · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Mini-batch semi-stochastic gradient descent in the proximal setting
J. Konečnỳ, J. Liu, P. Richtárik, and M. Takáč · 2016
Cited alongside, same era.
Ssd: Single shot multibox detector
W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C.-Y. Fu, and A. C. Berg · 2016
Cited alongside, same era.
SGDR: stochastic gradient descent with restarts
I. Loshchilov and F. Hutter · 2016
Cited alongside, same era.
Movieqa: Understanding stories in movies through question-answering
Asynchronous stochastic gradient descent with delay compensation
S. Zheng, Q. Meng, T. Wang, W. Chen, N. Yu, Z.-M. Ma, and T.-Y. Liu · 2017
Later among the works it cites.
Scene parsing through ade20k dataset
B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba · 2017
Later among the works it cites.
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs
L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille · 2018
Later among the works it cites.
Slowfast networks for video recognition
C. Feichtenhofer, H. Fan, J. Malik, and K. He · 2018
Later among the works it cites.
Preemptible vm, Jan 2018
Google · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Tapaswi, Y. Zhu, R. Stiefelhagen, A. Torralba, R. Urtasun, and S. Fidler · 2016
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
Z. Allen-Zhu · 2017
Cited alongside, same era.
Adabatch: Adaptive batch sizes for training deep neural networks
A. Devarakonda, M. Naumov, and M. Garland · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Cited alongside, same era.
Accurate, large minibatch SGD: training imagenet in 1 hour
P. Goyal, P. Dollár, R. B. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
E. Hoffer, I. Hubara, and D. Soudry · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Cited alongside, same era.
X. Jia, S. Song, W. He, Y. Wang, H. Rong, F. Zhou, L. Xie, Z. Guo, Y. Yang, L. Yu, et al · 2018
Later among the works it cites.
Tvqa: Localized, compositional video question answering
J. Lei, L. Yu, M. Bansal, and T. L. Berg · 2018
Later among the works it cites.
Low-latency video semantic segmentation
Y. Li, J. Shi, and D. Lin · 2018
Later among the works it cites.
An empirical model of large-batch training
S. McCandlish, J. Kaplan, D. Amodei, and O. D. Team · 2018
Later among the works it cites.
An empirical model of large-batch training
S. McCandlish, J. Kaplan, D. Amodei, and O. D. Team · 2018
Later among the works it cites.
Azure low priority vm, April 2018
Microsoft · 2018
Later among the works it cites.
Amazon ec2 spot instances, August 2018
A. W. Services · 2018
Later among the works it cites.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2018
Later among the works it cites.
Bag of tricks for image classification with convolutional neural networks
J. Xie, T. He, Z. Zhang, H. Zhang, Z. Zhang, and M. Li · 2018
Later among the works it cites.
Context encoding for semantic segmentation
H. Zhang, K. Dana, J. Shi, Z. Zhang, X. Wang, A. Tyagi, and A. Agrawal · 2018
Later among the works it cites.
Feelvos: Fast end-to-end embedding learning for video object segmentation
P. Voigtlaender, Y. Chai, F. Schroff, H. Adam, B. Leibe, and L.-C. Chen · 2019
Closest in time.
Bag of freebies for training object detection neural networks
Z. Zhang, T. He, H. Zhang, Z. Zhang, J. Xie, and M. Li · 2019
Closest in time.