Fetching the paper…
Reading the bibliography…
Many machine learning models require a training procedure based on running stochastic gradient descent.
International journal of control 41(2), 329–344 (1985)
Leontaritis, I., Billings, S.A.: Input-output parametric models for non-linear systems part ii: stochastic non-linear systems · 1985
Earlier work this paper cites.
Evolutionary computation 9(2), 159–195 (2001)
Hansen, N., Ostermeier, A.: Completely derandomized self-adaptation in evolution strategies · 2001
Earlier work this paper cites.
arXiv:2003.01115 (2020), https://arxiv.org/abs/2003.01115
van der Wilk, M., Dutordoir, V., John, S., Artemev, A., Adam, V., Hensman, J.: A framework for interdomain and multioutput Gaussian processes · 2003
Earlier work this paper cites.
https://keras.io/examples/cifar10_resnet/
Chollet, F.: Keras implementation of ResNet for CIFAR · 2009
Earlier work this paper cites.
Tech. rep., Citeseer (2009)
Krizhevsky, A., Hinton, G.: Learning multiple layers of features from tiny images · 2009
Earlier work this paper cites.
Artificial Intelligence and Statistics (2009)
Titsias, M.: Variational Learning of Inducing Variables in Sparse Gaussian Processes · 2009
Earlier work this paper cites.
In: Proceedings of the 27th International Conference on International Conference on Machine Learning. pp. 1015–1022. Omnipress (2010)
Srinivas, N., Krause, A., Kakade, S., Seeger, M.: Gaussian Process optimization in the bandit setting: no regret and experimental design · 2010
Earlier work this paper cites.
In: Advances in neural information processing systems. pp. 2546–2554 (2011)
Bergstra, J.S., Bardenet, R., Bengio, Y., Kégl, B.: Algorithms for hyper-parameter optimization · 2011
Earlier work this paper cites.
Journal of Machine Learning Research 12(Jul), 2121–2159 (2011)
Duchi, J., Hazan, E., Singer, Y.: Adaptive subgradient methods for online learning and stochastic optimization · 2011
Earlier work this paper cites.
In: Neural networks: Tricks of the trade, pp. 437–478. Springer (2012)
Bengio, Y.: Practical recommendations for gradient-based training of deep architectures · 2012
Earlier work this paper cites.
In: Artificial intelligence and statistics. pp. 592–600 (2012)
Kaufmann, E., Cappé, O., Garivier, A.: On Bayesian upper confidence bounds for bandit problems · 2012
Earlier work this paper cites.
In: Advances in neural information processing systems. pp. 2951–2959 (2012)
Snoek, J., Larochelle, H., Adams, R.P.: Practical Bayesian optimization of machine learning algorithms · 2012
Earlier work this paper cites.
COURSERA: Neural networks for machine learning 4(2), 26–31 (2012)
Tieleman, T., Hinton, G.: Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude · 2012
Earlier work this paper cites.
Journal of Machine Learning Research (2013)
Hoffman, M.D., Blei, D.M., Wang, C., Paisley, J.: Stochastic Variational Inference · 2013
Earlier work this paper cites.
SIAM/ASA Journal on Uncertainty Quantification 1(1), 57–78 (2013)
Picheny, V., Ginsbourger, D.: A nonstationary space-time Gaussian Process model for partially converged simulations · 2013
Cited alongside, same era.
In: Advances in neural information processing systems. pp. 2004–2012 (2013)
Swersky, K., Snoek, J., Adams, R.P.: Multi-task Bayesian optimization · 2013
Cited alongside, same era.
SIAM/ASA Journal on Uncertainty Quantification 2(1), 490–510 (2014)
Ginsbourger, D., Baccou, J., Chevalier, C., Perales, F., Garland, N., Monerie, Y.: Bayesian adaptive reconstruction of profile optima and optimizers · 2014
Cited alongside, same era.
arXiv preprint arXiv:1406.3896 (2014)
Swersky, K., Snoek, J., Adams, R.P.: Freeze-thaw Bayesian optimization · 2014
Cited alongside, same era.
In: Twenty-Fourth International Joint Conference on Artificial Intelligence (2015)
Domhan, T., Springenberg, J.T., Hutter, F.: Speeding up automatic hyperparameter optimization of deep neural networks by extrapolation of learning curves · 2015
In: International Conference on Artificial Intelligence and Statistics (AISTATS 2017). pp. 528–536. PMLR (2017a)
Klein, A., Falkner, S., Bartels, S., Hennig, P., Hutter, F.: Fast Bayesian optimization of machine learning hyperparameters on large datasets · 2017
Later among the works it cites.
The Journal of Machine Learning Research 18(1), 1299–1304 (2017)
Matthews, A.G.d.G., Van Der Wilk, M., Nickson, T., Fujii, K., Boukouvalas, A., León-Villagrá, P., Ghahramani, Z., Hensman, J.: Gpflow: A Gaussian Process library using tensorflow · 2017
Later among the works it cites.
In: 2017 IEEE Winter Conference on Applications of Computer Vision (WACV). pp. 464–472. IEEE (2017)
Smith, L.N.: Cyclical learning rates for training neural networks · 2017
Later among the works it cites.
In: Advances in Neural Information Processing Systems. pp. 5760–5770 (2018)
Bogunovic, I., Scarlett, J., Jegelka, S., Cevher, V.: Adversarially robust optimization with Gaussian processes · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
In: Proceedings of the Eighteenth International Conference on Artificial Intelligence and Statistics (2015)
Hensman, J., Matthews, A.G.d.G., Ghahramani, Z.: Scalable variational Gaussian process classification · 2015
Cited alongside, same era.
In: Advances in Neural Information Processing Systems. pp. 3981–3989 (2016)
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M.W., Pfau, D., Schaul, T., Shillingford, B., De Freitas, N.: Learning to learn by gradient descent by gradient descent · 2016
Cited alongside, same era.
In: Advances in neural information processing systems. pp. 1507–1515 (2016)
Bogunovic, I., Scarlett, J., Krause, A., Cevher, V.: Truncated variance reduction: A unified approach to Bayesian optimization and level-set estimation · 2016
Cited alongside, same era.
In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 770–778 (2016)
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition · 2016
Cited alongside, same era.
Journal of Machine Learning Research 51, 231–239 (2016)
Matthews, A.G.d.G., Hensman, J., Turner, R., Ghahramani, Z.: On sparse variational methods and the kullback-leibler divergence between stochastic Processes · 2016
Cited alongside, same era.
In: Proceedings of the 2016 Winter Simulation Conference. pp. 770–781. IEEE Press (2016)
Poloczek, M., Wang, J., Frazier, P.I.: Warm starting Bayesian optimization · 2016
Cited alongside, same era.
In: AISTATS. pp. 1431–1440 (2016)
Saul, A.D., Hensman, J., Vehtari, A., Lawrence, N.D., et al.: Chained Gaussian Processes · 2016
Cited alongside, same era.
Falkner, S., Klein, A., Hutter, F.: Bohb: Robust and efficient hyperparameter optimization at scale · 2018
Later among the works it cites.
Gugger, S., Howard, J.: Adamw and super-convergence is now the fastest way to train neural nets (Jul 2018), https://www.fast.ai/2018/07/02/adam-weight-decay/
2018
Later among the works it cites.
Journal of Machine Learning Research 18(185), 1–52 (2018)
Li, L., Jamieson, K., DeSalvo, G., Rostamizadeh, A., Talwalkar, A.: Hyperband: A novel bandit-based approach to hyperparameter optimization · 2018
Later among the works it cites.
In: Proceedings of the Genetic and Evolutionary Computation Conference (pp. 865-872) (2018)
Nishida, K., Akimoto, Y.: PSA-CMA-ES: CMA-ES with population size adaptation · 2018
Later among the works it cites.
European Journal of Operational Research 270(3), 1074–1085 (2018)
Pearce, M., Branke, J.: Continuous multi-task Bayesian optimisation with correlation · 2018
Later among the works it cites.
In: ICLR (2018)
Reddi, S.J., Kale, S., Kumar, S.: On the convergence of ADAM and beyond · 2018
Later among the works it cites.
In: Advances in Neural Information Processing Systems. pp. 9884–9895 (2018)
Wilson, J., Hutter, F., Deisenroth, M.: Maximizing acquisition functions for Bayesian optimization · 2018
Later among the works it cites.
In: ICLR (2019)
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization · 2019
Later among the works it cites.
In: Artificial Intelligence and Machine Learning for Multi-Domain Operations Applications. vol. 11006, p. 1100612. International Society for Optics and Photonics (2019)
Smith, L.N., Topin, N.: Super-convergence: Very fast training of neural networks using large learning rates · 2019
Later among the works it cites.