Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 1901
Earlier work this paper cites.
Genetic algorithms as a tool for feature selection in machine learning. In International Conference on Tools with Artificial Intelligence . 200–203
H. Vafaie and K. De Jong. 1992 · 1992
Earlier work this paper cites.
Support-vector networks
C. Cortes and V. Vapnik. 1995 · 1995
Earlier work this paper cites.
Machine Learning
T. Mitchell. 1997 · 1997
Earlier work this paper cites.
Support vector clustering
A. Ben-Hur, D. Horn, H. T. Siegelmann, and V. Vapnik. 2001 · 2001
Earlier work this paper cites.
Foundations of bilevel programming
S. Dempe. 2002 · 2002
Earlier work this paper cites.
Genetic programming with a genetic algorithm for feature construction and selection
M. Smith and L. Bull. 2005 · 2005
Earlier work this paper cites.
Getting the most out of ensemble selection. In International Conference on Data Mining . 828–833
R. Caruana, A. Munson, and A. Niculescu-Mizil. 2006 · 2006
Earlier work this paper cites.
Compressed sensing
D. L. Donoho. 2006 · 2006
Earlier work this paper cites.
The tradeoffs of large scale learning. In Advances in Neural Information Processing Systems . 161–168
L. Bottou and O. Bousquet. 2008 · 2008
Earlier work this paper cites.
Introduction to derivative-free optimization
A. R. Conn, k. Scheinberg, and L. N. Vicente. 2009 · 2009
Earlier work this paper cites.
A Survey on Transfer Learning
S. Pan and Q. Yang. 2009 · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines. In International Conference on Machine Learning . 807–814
V. Nair and G. Hinton. 2010 · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization. In Advances in Neural Information Processing Systems . 2546–2554
J. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl. 2011 · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer. 2011 · 2011
Earlier work this paper cites.
Scikit-learn: Machine Learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011 · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
J. Bergstra and Y. Bengio. 2012 · 2012
Earlier work this paper cites.
Imagenet: classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems . 1097–1105
A. Krizhevsky, I. Sutskever, and G. Hinton. 2012 · 2012
Earlier work this paper cites.
Almost optimal exploration in multi-armed bandits. In International Conference on Machine Learning . 1238–1246
Z. Karnin, T. Koren, and O. Somekh. 2013 · 2013
Earlier work this paper cites.
Auto-WEKA: Combined selection and hyperparameter optimization of classification algorithms. In ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . 847–855
C. Thornton, F. Hutter, H. Hoos, and K. Leyton-Brown. 2013 · 2013
Earlier work this paper cites.
A survey on feature selection methods
G. Chandrashekar and F. Sahin. 2014 · 2014
Earlier work this paper cites.
Raiders of the lost architecture: Kernels for Bayesian optimization in conditional parameter spaces
Original
K. Swersky, D. Duvenaud, J. Snoek, F. Hutter, and M. Osborne. 2014 · 2014
Earlier work this paper cites.
Speeding up automatic hyperparameter optimization of deep neural networks by extrapolation of learning curves. In International Joint Conference on Artificial Intelligence
Tobias Domhan, Jost Tobias Springenberg, and Frank Hutter. 2015 · 2015
Earlier work this paper cites.
Efficient Benchmarking of Hyperparameter Optimizers via Surrogates. In AAAI Conference on Artificial Intelligence
K. Eggensperger, F. Hutter, H. Hoos, and K. Leyton-Brown. 2015 · 2015
Earlier work this paper cites.
Efficient and robust automated machine learning. In Advances in Neural Information Processing Systems . 2962–2970
M. Feurer, A. Klein, K. Eggensperger, J. Springenberg, M. Blum, and F. Hutter. 2015 · 2015
Earlier work this paper cites.
Unsupervised domain adaptation by backpropagation. In International Conference on Machine Learning . 1180–1189
Y. Ganin and V. Lempitsky. 2015 · 2015
Earlier work this paper cites.
Fast r-cnn. In International Conference on Computer Vision . 1440–1448
R. Girshick. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba. 2015 · 2015
Earlier work this paper cites.
Bi-level stochastic gradient for large scale support vector machine. In Neurocomputing . 300–308
C. Nicolas and W. Wang. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
S. Ren, K. He, R. Girshick, and J. Sun. 2015 · 2015
Earlier work this paper cites.
Optimizing deep learning hyper-parameters through an evolutionary algorithm. In Workshop on Machine Learning in High-performance Computing Environments
S. Young, D. Rose, T. Karnowski, S. Lim, and R. Patton. 2015 · 2015
Earlier work this paper cites.
Deep learning
I. Goodfellow, Y. Bengio, and A. Courville. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Conference on Computer Vision and Pattern Recognition . 770–778
K. He, X. Zhang, S. Ren, and J. Sun. 2016 · 2016
Earlier work this paper cites.
Non-stochastic best arm identification and hyperparameter optimization. In Artificial Intelligence and Statistics . 240–248
K. Jamieson and A. Talwalkar. 2016 · 2016
Earlier work this paper cites.
Learning curve prediction with Bayesian neural networks. In International Conference on Learning Representations
A. Klein, S. Falkner, J. T. Springenberg, and F. Hutter. 2016 · 2016
Earlier work this paper cites.
SGDR: Stochastic Gradient Descent with Warm Restarts. In International Conference on Learning Representations
I. Loshchilov and F. Hutter. 2016 · 2016
Earlier work this paper cites.
Evaluation of a tree-based pipeline optimization tool for automating data science. In Genetic and Evolutionary Computation Conference . 485–492
R. Olson, N. Bartley, R. Urbanowicz, and J. Moore. 2016 · 2016
Earlier work this paper cites.
TPOT: A tree-based pipeline optimization tool for automating machine learning. In Workshop on Automatic Machine Learning . 66–74
R. Olson and J. Moore. 2016 · 2016
Earlier work this paper cites.
Hyperparameter optimization with approximate gradient. In International Conference on Machine Learning . 737–746
F. Pedregosa. 2016 · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection. In Conference on Computer Vision and Pattern Recognition . 779–788
J. Redmon, S. Divvala, R. Girshick, and A. Farhadi. 2016 · 2016
Earlier work this paper cites.
Convolutional neural fabrics
S. Saxena and J. Verbeek. 2016 · 2016
Earlier work this paper cites.
Accelerating neural architecture search using performance prediction
Original
B. Baker, O. Gupta, R. Raskar, and N. Naik. 2017b · 2017
Earlier work this paper cites.
A downsampled variant of imagenet as an alternative to the cifar datasets
Original
P. Chrabaszcz, I. Loshchilov, and F. Hutter. 2017 · 2017
Earlier work this paper cites.
AdaNet: Adaptive Structural Learning of Artificial Neural Networks. In International Conference on Machine Learning . 874–883
C. Cortes, X. Gonzalvo, V. Kuznetsov, M. Mohri, and S. Yang. 2017 · 2017
Earlier work this paper cites.
Neural architecture search: A survey
T. Elsken, J. Metzen, and F. Hutter. 2019 · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning . 1126–1135
C. Finn, P. Abbeel, and S. Levine. 2017 · 2017
Earlier work this paper cites.
How machine learning could help to improve climate forecasts
N. Jones. 2017 · 2017
Earlier work this paper cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
L. Li, K. Jamieson, G. DeSalvo, A. Rostamizadeh, and A. Talwalkar. 2017 · 2017
Earlier work this paper cites.
SVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability
M. Raghu, J. Gilmer, J. Yosinski, and J. Sohl-Dickstein. 2017 · 2017
Earlier work this paper cites.
Large-scale evolution of image classifiers. In International Conference on Machine Learning
E. Real, S. Moore, A. Selle, S. Saxena, Y. Suematsu, J. Tan, Q. Le, and A. Kurakin. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, et al · 2017
Earlier work this paper cites.
Neural architecture search with reinforcement learning. In International Conference on Learning Representations
B. Zoph and Q. Le. 2017 · 2017
Earlier work this paper cites.
Learning Transferable Architectures for Scalable Image Recognition. In Conference on Computer Vision and Pattern Recognition
B. Zoph, V. Vasudevan, J. Shlens, and Q. Le. 2017 · 2017
Earlier work this paper cites.
Understanding and Simplifying One-Shot Architecture Search. In International Conference on Machine Learning . 549–558
G. Bender, P. J. Kindermans, B. Zoph, V. Vasudevan, and Q. Le. 2018 · 2018
Earlier work this paper cites.
BOHB: Robust and efficient hyperparameter optimization at scale. In International Conference on Machine Learning . 1437–1446
S. Falkner, A. Klein, and F. Hutter. 2018 · 2018
Earlier work this paper cites.