Fetching the paper…
Reading the bibliography…
Since deep neural networks were developed, they have made huge contributions to everyday lives.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
A new method of locating the maximum point of an arbitrary multipeak curve in the presence of noise
Harold J Kushner · 1964
Earlier work this paper cites.
A dynamic allocation index for the sequential design of experiments
John Gittins · 1974
Earlier work this paper cites.
On bayesian methods for seeking the extremum
Jonas Močkus · 1975
Earlier work this paper cites.
The application of bayesian methods for seeking the extremum
Jonas Mockus, Vytautas Tiesis, and Antanas Zilinskas · 1978
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Tze Leung Lai and Herbert Robbins · 1985
Earlier work this paper cites.
Statistical aspects of neural networks
Brian D Ripley · 1993
Earlier work this paper cites.
A review of techniques for parameter sensitivity analysis of environmental models
DM Hamby · 1994
Earlier work this paper cites.
Statlog: comparison of classification algorithms on large real-world problems
Ross D. King, Cao Feng, and Alistair Sutherland · 1995
Earlier work this paper cites.
Automatic parameter selection by minimizing estimated error
Ron Kohavi and George H John · 1995
Earlier work this paper cites.
An introduction to sensitivity analysis. massachusetts institute of technology
L Breierova and M Choudhari · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
No free lunch theorems for optimization
David H Wolpert and William G Macready · 1997
Earlier work this paper cites.
Comparison between genetic algorithms and particle swarm optimization
Russell C Eberhart and Yuhui Shi · 1998
Earlier work this paper cites.
Efficient global optimization of expensive black-box functions
Donald R Jones, Matthias Schonlau, and William J Welch · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Efficient progressive sampling
Foster Provost, David Jensen, and Tim Oates · 1999
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
Improving the rprop learning algorithm
Christian Igel and Michael Hüsken · 2000
Earlier work this paper cites.
Estimating the predictive accuracy of a classifier
Hilan Bensusan and Alexandros Kalousis · 2001
Earlier work this paper cites.
Metamodels for computer-based engineering design: survey and recommendations
Timothy W Simpson, JD Poplinski, Patrick N Koch, and Janet K Allen · 2001
Earlier work this paper cites.
Gaussian processes in machine learning
Carl Edward Rasmussen · 2003
Earlier work this paper cites.
Feature selection, l 1 vs. l 2 regularization, and rotational invariance
Andrew Y Ng · 2004
Earlier work this paper cites.
Parallel implementation of a random search procedure: an experimental study
Nikolai K Krivulin, Dennis Guster, and Charles Hall · 2005
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Geoffrey E Hinton, Simon Osindero, and Yee-Whye Teh · 2006
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle · 2007
Earlier work this paper cites.
Gaussian processes
Chuong B Do · 2007
Earlier work this paper cites.
Typical steps for solving optimization problems
Martin J. Strauss · 2007
Earlier work this paper cites.
Predicting the performance of learning algorithms using support vector machines as meta-regressors
Silvio B Guerra, Ricardo BC Prudêncio, and Teresa B Ludermir · 2008
Earlier work this paper cites.
Conditional variable importance for random forests
Carolin Strobl, Anne-Laure Boulesteix, Thomas Kneib, Thomas Augustin, and Achim Zeileis · 2008
Earlier work this paper cites.
The knowledge-gradient policy for correlated normal beliefs
Peter Frazier, Warren Powell, and Savas Dayanik · 2009
Earlier work this paper cites.
Automated configuration of algorithms for solving hard computational problems
Frank Hutter · 2009
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang · 2009
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
Niranjan Srinivas, Andreas Krause, Sham M Kakade, and Matthias Seeger · 2009
Earlier work this paper cites.
Kriging is well-suited to parallelize optimization
David Ginsbourger, Rodolphe Le Riche, and Laurent Carraro · 2010
Earlier work this paper cites.
Introduction to monte carlo simulation
Robert L Harrison · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Deep learners benefit more from out-of-distribution examples
Yoshua Bengio, Frédéric Bastien, Arnaud Bergeron, Nicolas Boulanger-Lewandowski, Thomas Breuel, Youssouf Chherawala, Moustapha Cisse, Myriam Côté, Dumitru Erhan, Jeremy Eustache, et al · 2011
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
James S Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Dealing with asynchronicity in parallel gaussian process based global optimization
David Ginsbourger, Janis Janusevskis, and Rodolphe Le Riche · 2011
Earlier work this paper cites.
Portfolio allocation for bayesian optimization
Matthew D Hoffman, Eric Brochu, and Nando de Freitas · 2011
Earlier work this paper cites.
Sequential model-based optimization for general algorithm configuration
Frank Hutter, Holger H Hoos, and Kevin Leyton-Brown · 2011
Earlier work this paper cites.
Oracle inequalities for computationally adaptive model selection
Alekh Agarwal, Peter L Bartlett, and John C Duchi · 2012
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio · 2012
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio · 2012
Earlier work this paper cites.
Entropy search for information-efficient global optimization
Philipp Hennig and Christian J Schuler · 2012
Earlier work this paper cites.
Parallel algorithm configuration
Frank Hutter, Holger H Hoos, and Kevin Leyton-Brown · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Provably convergent multifidelity optimization algorithm not requiring high-fidelity derivatives
Andrew March and Karen Willcox · 2012
Earlier work this paper cites.
The nature of code
Daniel Shiffman, Shannon Fry, and Zannah Marsh · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures
James Bergstra, Daniel Yamins, and David Daniel Cox · 2013
Earlier work this paper cites.
Towards an empirical foundation for assessing bayesian optimization of hyperparameters
Katharina Eggensperger, Matthias Feurer, Frank Hutter, James Bergstra, Jasper Snoek, Holger Hoos, and Kevin Leyton-Brown · 2013
Earlier work this paper cites.
Ian J Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron Courville, and Yoshua Bengio · 2013
Earlier work this paper cites.
Almost optimal exploration in multi-armed bandits
Zohar Karnin, Tomer Koren, and Oren Somekh · 2013
Cited alongside, same era.
Evolutionary optimization algorithms
Dan Simon · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Parallelizing exploration-exploitation tradeoffs in gaussian process bandit optimization
Thomas Desautels, Andreas Krause, and Joel W Burdick · 2014
Cited alongside, same era.
Automatic model construction with Gaussian processes
David Duvenaud · 2014
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
The number of hidden layers, 2017
Jeff Heaton · 2017
Later among the works it cites.
Population based training of neural networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, et al · 2017
Later among the works it cites.
Learning the distribution with largest mean: two bandit frameworks
Emilie Kaufmann and Aurélien Garivier · 2017
Later among the works it cites.
Learning rate schedules and adaptive learning rate methods for deep learning
Suki Lau · 2017
Later among the works it cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Predictive entropy search for efficient global optimization of black-box functions
José Miguel Hernández-Lobato, Matthew W Hoffman, and Zoubin Ghahramani · 2014
Cited alongside, same era.
No free lunch theorems: Limitations and perspectives of metaheuristics
Christian Igel · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding
Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Hyperopt-sklearn: automatic hyperparameter configuration for scikit-learn
Brent Komer, James Bergstra, and Chris Eliasmith · 2014
Cited alongside, same era.
Efficient mini-batch training for stochastic optimization
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J Smola · 2014
Cited alongside, same era.
Nicolas Loizou and Peter Richtárik · 2017
Later among the works it cites.
Particle swarm optimization for hyper-parameter selection in deep neural networks
Pablo Ribalta Lorenzo, Jakub Nalepa, Michal Kawulok, Luciano Sanchez Ramos, and José Ranilla Pastor · 2017
Later among the works it cites.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom · 2017
Later among the works it cites.
Improving deep neural networks: Hyperparameter tuning, regularization and optimization
Andrew Ng · 2017
Later among the works it cites.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Activation functions and it’s types-which is better?
Anish Singh Walia · 2017
Later among the works it cites.
The effectiveness of data augmentation in image classification using deep learning
Jason Wang and Luis Perez · 2017
Later among the works it cites.
Large batch training of convolutional networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Later among the works it cites.
https:/https://github.com/tobegit3hub/advisor , 2018
Advisor · 2018
Later among the works it cites.
Deep learning using rectified linear units (relu)
Abien Fred Agarap · 2018
Later among the works it cites.
Hyperparameters in deep learning
Samarth Agrawal · 2018
Later among the works it cites.
Bohb: robust and efficient hyperparameter optimization at scale
André Biedenkapp · 2018
Later among the works it cites.
Hyper-parameter optimization algorithms: a short review
Aloïs Bissuel · 2018
Later among the works it cites.
The intuition behind bayesian optimization with gaussian processes
Charles Brecque · 2018
Later among the works it cites.
Understanding rmsprop –faster neural network learning
Vitaly Bushaev · 2018
Later among the works it cites.
Deep Learning mit Python und Keras: Das Praxis-Handbuch vom Entwickler der Keras-Bibliothek
Francois Chollet · 2018
Later among the works it cites.
Annotative experts for hyperparameter selection
C Davis and C Giraud-Carrier · 2018
Later among the works it cites.
Bohb: Robust and efficient hyperparameter optimization at scale
Stefan Falkner, Aaron Klein, and Frank Hutter · 2018
Later among the works it cites.
A tutorial on bayesian optimization
Peter I Frazier · 2018
Later among the works it cites.
Practical guide to hyperparameters optimization for deep learning models
Charlie Harrington · 2018
Later among the works it cites.
Efficient neural architecture search with network morphism
Haifeng Jin, Qingquan Song, and Xia Hu · 2018
Later among the works it cites.
Grid search for model tuning
Rohan Joseph · 2018
Later among the works it cites.
L1 and l2 regularization
R Khandelwal · 2018
Later among the works it cites.
Comparison of activation functions for deep neural networks
Will Koehrsen · 2018
Later among the works it cites.
Massively parallel hyperparameter tuning
Liam Li, Kevin Jamieson, Afshin Rostamizadeh, Ekaterina Gonina, Moritz Hardt, Benjamin Recht, and Ameet Talwalkar · 2018
Later among the works it cites.
Tune: A research platform for distributed model selection and training
Richard Liaw, Eric Liang, Robert Nishihara, Philipp Moritz, Joseph E Gonzalez, and Ion Stoica · 2018
Later among the works it cites.
Shufflenet v2: Practical guidelines for efficient cnn architecture design
Ningning Ma, Xiangyu Zhang, Hai-Tao Zheng, and Jian Sun · 2018
Later among the works it cites.
Why is random search better than grid search more machine learning
Kishan Maladkar · 2018
Later among the works it cites.
Revisiting small batch training for deep neural networks
Dominic Masters and Carlo Luschi · 2018
Later among the works it cites.
Neural network intelligence
Microsoft · 2018
Later among the works it cites.
Finding good learning rate and the one cycle policy
nachiket tanksale · 2018
Later among the works it cites.
Understanding hyperparameters optimization in deep learning models: Concepts and tools, 2018
Jesus Rodriguez · 2018
Later among the works it cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Later among the works it cites.
Joaquin Vanschoren · 2018
Later among the works it cites.
Combination of hyperband and bayesian optimization for hyperparameter optimization in deep learning
Jiazhuo Wang, Jason Xu, and Xuejun Wang · 2018
Later among the works it cites.
Taking human out of learning applications: A survey on automated machine learning
Quanming Yao, Mengshuo Wang, Yuqiang Chen, Wenyuan Dai, Hu Yi-Qi, Li Yu-Feng, Tu Wei-Wei, Yang Qiang, and Yu Yang · 2018
Later among the works it cites.
How to configure the learning rate when training deep learning neural networks
Jason Brownlee · 2019
Later among the works it cites.
Hyperparameter optimization
Matthias Feurer and Frank Hutter · 2019
Later among the works it cites.
Multi-fidelity automatic hyper-parameter tuning via transfer series expansion
Yi-Qi Hu, Yang Yu, Wei-Wei Tu, Qiang Yang, Yuqiang Chen, and Wenyuan Dai · 2019
Later among the works it cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Later among the works it cites.
Comparison of activation functions for deep neural networks
Ayyüce Kizrak · 2019
Later among the works it cites.
A generalized framework for population based training
Ang Li, Ola Spyra, Sagi Perel, Valentin Dalibard, Max Jaderberg, Chenjie Gu, David Budden, Tim Harley, and Pramod Gupta · 2019
Later among the works it cites.
Dying relu and initialization: Theory and numerical examples
Lu Lu, Yeonjong Shin, Yanhui Su, and George Em Karniadakis · 2019
Later among the works it cites.
Overfitting and underfitting in deep learning
Artem Oppermann · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Introduction to multi-armed bandits
Aleksandrs Slivkins et al · 2019
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V Le · 2019
Later among the works it cites.
Mnasnet: Platform-aware neural architecture search for mobile
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V Le · 2019
Later among the works it cites.
The vanishing gradient problem
CF Wang · 2019
Later among the works it cites.
The curse of dimensionality
Tony Yiu · 2019
Later among the works it cites.