Fetching the paper…
Reading the bibliography…
Mini-batch sub-sampling (MBSS) is favored in deep neural network training to reduce the computational cost.
Robbins H, Monro S (1951) A stochastic approximation method. The annals of mathematical statistics pp 400–407
1951
Earlier work this paper cites.
Wolfe P (1969) Convergence conditions for ascent methods. SIAM Review 11(2):226–235
1969
Earlier work this paper cites.
Wolfe P (1971) Convergence conditions for ascent methods. II: Some corrections. SIAM Review 13(2):185–188
1971
Earlier work this paper cites.
Lyapunov AM (1992) The general problem of the stability of motion. International journal of control 55(3):531–534
1992
Earlier work this paper cites.
LeCun Y, Bottou L, Bengio Y, Haffner P (1998) Gradient-based learning applied to document recognition. Proceedings of the IEEE 86(11):2278–2324
1998
Earlier work this paper cites.
2009
Earlier work this paper cites.
Glorot X, Bengio Y (2010) Understanding the difficulty of training deep feedforward neural networks. In: Proceedings of the thirteenth international conference on artificial intelligence and statistics, JMLR Workshop and Conference Proceedings, pp 249–256
2010
Earlier work this paper cites.
Byrd RH, Chin GM, Neveitt W, Nocedal J (2011) On the use of stochastic hessian information in optimization methods for machine learning. SIAM Journal on Optimization 21(3):977–995
2011
Earlier work this paper cites.
Bengio Y (2012) Practical recommendations for gradient-based training of deep architectures. In: Neural networks: Tricks of the trade, Springer, pp 437–478
2012
Earlier work this paper cites.
Byrd RH, Chin GM, Nocedal J, Wu Y (2012) Sample Size Selection in Optimization Methods for Machine Learning. Mathematical Programming 134(1):127–155, DOI 10.1007/s10107-012-0572-5
2012
Earlier work this paper cites.
Friedlander MP, Schmidt M (2012) Hybrid deterministic-stochastic methods for data fitting. SIAM Journal on Scientific Computing 34(3):A1380–A1405
2012
Cited alongside, same era.
Wilke DN, Kok S, Johannes, Snyman A, Groenwold AA (2013) Gradient-only approaches to avoid spurious local minima in unconstrained optimization. Optimization and Engineering 14(2):275–304
2013
Cited alongside, same era.
Abadi M, Agarwal A, Barham P, Brevdo E, Chen Z, Citro C, Corrado GS, Davis A, Dean J, Devin M, Ghemawat S, Goodfellow I, Harp A, Irving G, Isard M, Jia Y, Jozefowicz R, Kaiser L, Kudlur M, Levenberg J, Mané D, Monga R, Moore S, Murray D, Olah C, Schuster M, Shlens J, Steiner B, Sutskever I, Talwar K, Tucker P, Vanhoucke V, Vasudevan V, Viégas F, Vinyals O, Warden P, Wattenberg M, Wicke M, Yu Y, Zheng X (2015) TensorFlow: Large-scale machine learning on heterogeneous systems. URL https://www.tensorflow.org/, software available from tensorflow.org
2015
Cited alongside, same era.
Masters D, Luschi C (2018) Revisiting small batch training for deep neural networks. arXiv preprint arXiv:180407612
2018
Later among the works it cites.
Chae Y, Wilke DN (2019) Empirical study towards understanding line search approximations for training neural networks. arXiv preprint arXiv:190906893
2019
Later among the works it cites.
Gupta RK (2019) Numerical Methods: Fundamentals and Applications. Cambridge University Press
2019
Later among the works it cites.
Mutschler M, Zell A (2019) Parabolic approximation line search: An efficient and effective line search approach for dnns. arXiv preprint arXiv:190311991
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 770–778
2016
Cited alongside, same era.
Loshchilov I, Hutter F (2016) Sgdr: Stochastic gradient descent with warm restarts. arXiv preprint arXiv:160803983
2016
Cited alongside, same era.
Mahsereci M, Hennig P (2017) Probabilistic line searches for stochastic optimization. The Journal of Machine Learning Research 18(1):4262–4320
2017
Cited alongside, same era.
Bergou Eh, Diouane Y, Kunc V, Kungurtsev V, Royer CW (2018) A subsampling line-search method with second-order results. arXiv preprint arXiv:181007211
2018
Cited alongside, same era.
Bollapragada R, Byrd R, Nocedal J (2018) Adaptive sampling strategies for stochastic optimization. SIAM Journal on Optimization 28(4):3312–3343
2018
Cited alongside, same era.
Kungurtsev V, Pevny T (2018) Algorithms for solving optimization problems arising from deep neural net models: smooth problems. arXiv preprint arXiv:180700172
2018
Cited alongside, same era.
Kafka D, Wilke D (2019a) Gradient-only line searches: An alternative to probabilistic line searches. arXiv preprint arXiv:190309383
Cited in the paper.
Kafka D, Wilke DN (2019b) Resolving Learning Rates Adaptively by locating Stochastic Non-Negative Associated Gradient Projection Points using Line Searches, unpublished: In review at the Journal of Global Optimization
Cited in the paper.
Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, Killeen T, Lin Z, Gimelshein N, Antiga L, Desmaison A, Kopf A, Yang E, DeVito Z, Raison M, Tejani A, Chilamkurthy S, Steiner B, Fang L, Bai J, Chintala S (2019) Pytorch: An imperative style, high-performance deep learning library. In: Wallach H, Larochelle H, Beygelzimer A, d'Alché-Buc F, Fox E, Garnett R (eds) Advances in Neural Information Processing Systems 32, Curran Associates, Inc., pp 8024–8035, URL http://papers.neurips.cc/paper/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf
2019
Later among the works it cites.
Strubell E, Ganesh A, McCallum A (2019) Energy and policy considerations for deep learning in nlp. arXiv preprint arXiv:190602243
2019
Later among the works it cites.
Tan M, Le Q (2019) Efficientnet: Rethinking model scaling for convolutional neural networks. In: International Conference on Machine Learning, PMLR, pp 6105–6114
2019
Later among the works it cites.
Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A, Agarwal S, Herbert-Voss A, Krueger G, Henighan T, Child R, Ramesh A, Ziegler D, Wu J, Winter C, Hesse C, Chen M, Sigler E, Litwin M, Gray S, Chess B, Clark J, Berner C, McCandlish S, Radford A, Sutskever I, Amodei D (2020) Language models are few-shot learners. In: Larochelle H, Ranzato M, Hadsell R, Balcan MF, Lin H (eds) Advances in Neural Information Processing Systems, Curran Associates, Inc., vol 33, pp 1877–1901, URL https://proceedings.neurips.cc/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf
2020
Later among the works it cites.
Liu K (2020) 95.16% on CIFAR10 with PyTorch. https://github.com/kuangliu/pytorch-cifar
2020
Later among the works it cites.