Fetching the paper…
Reading the bibliography…
Many contemporary machine learning models require extensive tuning of hyperparameters to perform well.
On Bayesian methods for seeking the extremum
Močkus, Jonas. 1975 · 1975
Earlier work this paper cites.
Mining geostatistics
Journel, Andre G, & Huijbregts, Charles J. 1978 · 1978
Earlier work this paper cites.
A statistical method for global optimization
Cox, Dennis D, & John, Susan. 1992 · 1992
Earlier work this paper cites.
SDO: A Statistical Method for Global Optimization
Cox, Dennis D., & John, Susan. 1997 · 1997
Earlier work this paper cites.
Efficient global optimization of expensive black-box functions
Jones, Donald R, Schonlau, Matthias, & Welch, William J. 1998 · 1998
Earlier work this paper cites.
The MNIST database of handwritten digits
LeCun, Yann. 1998 · 1998
Earlier work this paper cites.
Gaussian processes in machine learning
Rasmussen, Carl Edward. 2003 · 2003
Earlier work this paper cites.
Sequential kriging optimization using multiple-fidelity evaluations
Huang, Deng, Allen, T. T., Notz, W. I., & Miller, R. A. 2006 · 2006
Earlier work this paper cites.
Multi-task Gaussian process prediction
Bonilla, Edwin V, Chai, Kian M, & Williams, Christopher. 2008 · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, Alex, Hinton, Geoffrey, et al · 2009
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
Bergstra, James S, Bardenet, Rémi, Bengio, Yoshua, & Kégl, Balázs. 2011 · 2011
Earlier work this paper cites.
Entropy search for information-efficient global optimization
Hennig, Philipp, & Schuler, Christian J. 2012 · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Snoek, Jasper, Larochelle, Hugo, & Adams, Ryan P. 2012 · 2012
Earlier work this paper cites.
Multi-task bayesian optimization
Swersky, Kevin, Snoek, Jasper, & Adams, Ryan P. 2013 · 2012
Cited alongside, same era.
Predictive entropy search for efficient global optimization of black-box functions
Hernández-Lobato, José Miguel, Hoffman, Matthew W, & Ghahramani, Zoubin. 2014 · 2014
Cited alongside, same era.
Stochastic Gradient Descent, Weighted Sampling, and the Randomized Kaczmarz algorithm
Needell, Deanna, Ward, Rachel, & Srebro, Nati. 2014 · 2014
Cited alongside, same era.
Freeze-thaw Bayesian optimization
Swersky, Kevin, Snoek, Jasper, & Adams, Ryan Prescott. 2014 · 2014
Cited alongside, same era.
Speeding up automatic hyperparameter optimization of deep neural networks by extrapolation of learning curves
Domhan, Tobias, Springenberg, Jost Tobias, & Hutter, Frank. 2015 · 2015
Cited alongside, same era.
Hyperparameter optimization of deep neural networks: Combining hyperband with Bayesian model selection
Bertrand, Hadrien, Ardon, Roberto, Perrot, Matthieu, & Bloch, Isabelle. 2017 · 2017
Later among the works it cites.
CPSG-MCMC: Clustering-Based Preprocessing method for Stochastic Gradient MCMC
Fu, Tianfan, & Zhang, Zhihua. 2017 · 2017
Later among the works it cites.
Google Vizier: A Service for Black-Box Optimization
Golovin, Daniel, Solnik, Benjamin, Moitra, Subhodeep, Kochanski, Greg, Karro, John Elliot, & Sculley, D. (eds). 2017 · 2017
Later among the works it cites.
Bayesian optimization with gradients
Wu, Jian, Poloczek, Matthias, Wilson, Andrew G, & Frazier, Peter. 2017 · 2017
Later among the works it cites.
Determinantal Point Processes for Mini-Batch Diversification
Zhang, Cheng, Kjellstrom, Hedvig, & Mandt, Stephan. 2017 · 2017
Later among the works it cites.
BOHB: Robust and efficient hyperparameter optimization at scale
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Non-Uniform Stochastic Average Gradient Method for Training Conditional Random Fields
Schmidt, Mark, Babanezhad, Reza, Ahmed, Mohamed, Defazio, Aaron, Clifton, Ann, & Sarkar, Anoop. 2015 · 2015
Cited alongside, same era.
Stochastic Optimization with Importance Sampling for Regularized Loss Minimization
Zhao, Peilin, & Zhang, Tong. 2015 · 2015
Cited alongside, same era.
Predictive entropy search for multi-objective bayesian optimization
Hernández-Lobato, Daniel, Hernandez-Lobato, Jose, Shah, Amar, & Adams, Ryan. 2016 · 2016
Cited alongside, same era.
Gaussian process bandit optimisation with multi-fidelity evaluations
Kandasamy, Kirthevasan, Dasarathy, Gautam, Oliva, Junier B, Schneider, Jeff, & Póczos, Barnabás. 2016 · 2016
Cited alongside, same era.
Fast bayesian optimization of machine learning hyperparameters on large datasets
Klein, Aaron, Falkner, Stefan, Bartels, Simon, Hennig, Philipp, & Hutter, Frank. 2016 · 2016
Cited alongside, same era.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Li, Lisha, Jamieson, Kevin, DeSalvo, Giulia, Rostamizadeh, Afshin, & Talwalkar, Ameet. 2016 · 2016
Cited alongside, same era.
Zagoruyko, Sergey, & Komodakis, Nikos. 2016 · 2016
Cited alongside, same era.
Falkner, Stefan, Klein, Aaron, & Hutter, Frank. 2018 · 2018
Later among the works it cites.
Training Deep Models Faster with Robust, Approximate Importance Sampling
Johnson, Tyler B, & Guestrin, Carlos. 2018 · 2018
Later among the works it cites.
Not all samples are created equal: Deep learning with importance sampling
Katharopoulos, Angelos, & Fleuret, François. 2018 · 2018
Later among the works it cites.
Combination of hyperband and Bayesian optimization for hyperparameter optimization in deep learning
Wang, Jiazhuo, Xu, Jason, & Wang, Xuejun. 2018 · 2018
Later among the works it cites.
Bayesian Optimization Meets Bayesian Optimal Stopping
Dai, Zhongxiang, Yu, Haibin, Low, Bryan Kian Hsiang, & Jaillet, Patrick. 2019 · 2019
Later among the works it cites.
Energy and Policy Considerations for Deep Learning in NLP
Strubell, Emma, Ganesh, Ananya, & Mccallum, Andrew. 2019 · 2019
Later among the works it cites.
Multi-fidelity optimization via surrogate modelling
Forrester, Alexander I.J., Sóbester, András, & Keane, Andy J. 2007 · 2088
Closest in time.