Fetching the paper…
Reading the bibliography…
When training deep learning models, the performance depends largely on the selected hyperparameters.
Transferability and hardness of supervised classification tasks
Anh Tuan Tran, Cuong V. Nguyen, and Tal Hassner · 1908
Earlier work this paper cites.
BANANAS: bayesian optimization with neural architectures for neural architecture search
Colin White, Willie Neiswanger, and Yash Savani · 1910
Earlier work this paper cites.
Multi-objective neural architecture search via predictive network performance optimization
Han Shi, Renjie Pi, Hang Xu, Zhenguo Li, James T. Kwok, and Tong Zhang · 1911
Earlier work this paper cites.
A new measure of rank correlation
Maurice G Kendall · 1938
Earlier work this paper cites.
On the algebraic structure of feedforward network weight spaces
Robert Hecht-Nielsen · 1990
Earlier work this paper cites.
Python reference manual
Guido Van Rossum and Fred L Drake Jr · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
LEEP: A new measure to evaluate transferability of learned representations
Cuong V. Nguyen, Tal Hassner, Cédric Archambeau, and Matthias W. Seeger · 2002
Earlier work this paper cites.
Information Theory, Inference and Learning Algorithms
David J.C. MacKay · 2003
Earlier work this paper cites.
Python for scientific computing
Travis E Oliphant · 2007
Earlier work this paper cites.
Matplotlib: A 2D graphics environment
John D Hunter · 2007
Earlier work this paper cites.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2010
Earlier work this paper cites.
Kriging is well-suited to parallelize optimization
David Ginsbourger, Rodolphe Le Riche, and Laurent Carraro · 2010
Earlier work this paper cites.
Contextual gaussian process bandit optimization
Andreas Krause and Cheng Ong · 2011
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
James Bergstra, Rémi Bardenet, Yoshua Bengio, and Balázs Kégl · 2011
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio · 2012
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
Auto-weka: Combined selection and hyperparameter optimization of classification algorithms
Chris Thornton, Frank Hutter, Holger H Hoos, and Kevin Leyton-Brown · 2013
Earlier work this paper cites.
Multi-task bayesian optimization
Kevin Swersky, Jasper Snoek, and Ryan P Adams · 2013
Earlier work this paper cites.
Collaborative hyperparameter tuning
Rémi Bardenet, Mátyás Brendel, Balázs Kégl, and Michele Sebag · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Efficient transfer learning method for automatic hyperparameter tuning
Dani Yogatama and Gideon Mann · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Earlier work this paper cites.
Initializing bayesian hyperparameter optimization via meta-learning
Matthias Feurer, Jost Springenberg, and Frank Hutter · 2015
Earlier work this paper cites.
Non-stochastic best arm identification and hyperparameter optimization
Kevin Jamieson and Ameet Talwalkar · 2016
Earlier work this paper cites.
Gaussian process bandit optimisation with multi-fidelity evaluations
Kirthevasan Kandasamy, Gautam Dasarathy, Junier B Oliva, Jeff Schneider, and Barnabás Póczos · 2016
Earlier work this paper cites.
Warm starting bayesian optimization
Matthias Poloczek, Jialei Wang, and Peter I Frazier · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V. Le · 2016
Earlier work this paper cites.
Torchvision: Pytorch’s computer vision library
TorchVision maintainers and contributors · 2016
Earlier work this paper cites.
Deep kernel learning
Andrew G Wilson, Zhiting Hu, Ruslan Salakhutdinov, and Eric P Xing · 2016
Earlier work this paper cites.
Deep sets
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola · 2017
Earlier work this paper cites.
Forward and reverse gradient-based hyperparameter optimization
Luca Franceschi, Michele Donini, Paolo Frasconi, and Massimiliano Pontil · 2017
Earlier work this paper cites.
Multi-information source optimization
Matthias Poloczek, Jialei Wang, and Peter Frazier · 2017
Earlier work this paper cites.
Multi-fidelity bayesian optimisation with continuous approximations
Kirthevasan Kandasamy, Gautam Dasarathy, Jeff Schneider, and Barnabás Póczos · 2017
Earlier work this paper cites.
Population based training of neural networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, et al · 2017
Earlier work this paper cites.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Deep information propagation
Samuel S. Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2017
Cited alongside, same era.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Cited alongside, same era.
Stochastic hyperparameter optimization through hypernetworks
Jonathan Lorraine and David Duvenaud · 2018
Cited alongside, same era.
Self-tuning networks: Bilevel optimization of hyperparameters using structured best-response functions
Pacoh: Bayes-optimal meta-learning with pac-guarantees
Jonas Rothfuss, Vincent Fortuin, Martin Josifoski, and Andreas Krause · 2021
Later among the works it cites.
Dehb: Evolutionary hyperband for scalable, robust and efficient hyperparameter optimization, 2021
Noor Awad, Neeratyoy Mallik, and Frank Hutter · 2021
Later among the works it cites.
St-nas: Efficient optimization of joint neural architecture and hyperparameter
Jinhang Cai, Yimin Ou, Xiu Li, and Haoqian Wang · 2021
Later among the works it cites.
How powerful are performance predictors in neural architecture search?
Colin White, Arber Zela, Binxin Ru, Yang Liu, and Frank Hutter · 2021
Later among the works it cites.
Understanding convolutions on graphs
Ameya Daigavane, Balaraman Ravindran, and Gaurav Aggarwal · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Matthew Mackay, Paul Vicol, Jonathan Lorraine, David Duvenaud, and Roger Grosse · 2018
Cited alongside, same era.
BOHB: Robust and efficient hyperparameter optimization at scale
Stefan Falkner, Aaron Klein, and Frank Hutter · 2018
Cited alongside, same era.
Scalable meta-learning for bayesian optimization using ranking-weighted gaussian process ensembles
Matthias Feurer, Benjamin Letham, and Eytan Bakshy · 2018
Cited alongside, same era.
Scalable hyperparameter transfer learning
Valerio Perrone, Rodolphe Jenatton, Matthias W Seeger, and Cédric Archambeau · 2018
Cited alongside, same era.
Gpytorch: Blackbox matrix-matrix gaussian process inference with gpu acceleration
Jacob Gardner, Geoff Pleiss, Kilian Q Weinberger, David Bindel, and Andrew G Wilson · 2018
Cited alongside, same era.
DARTS: differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2018
Cited alongside, same era.
Predicting the generalization gap in deep networks with margin distributions
Yiding Jiang, Dilip Krishnan, Hossein Mobahi, and Samy Bengio · 2018
Cited alongside, same era.
Benjamin Sanchez-Lengeling, Emily Reif, Adam Pearce, and Alexander B. Wiltschko · 2021
Later among the works it cites.
Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for learning
Charles H Martin and Michael W Mahoney · 2021
Later among the works it cites.
Methods and analysis of the first competition in predicting generalization of deep learning
Yiding Jiang, Parth Natekar, Manik Sharma, Sumukh K Aithal, Dhruva Kashyap, Natarajan Subramanyam, Carlos Lassance, Daniel M Roy, Gintare Karolina Dziugaite, Suriya Gunasekar, et al · 2021
Later among the works it cites.
Self-supervised representation learning on neural network weights for model characteristic prediction
Konstantin Schürholt, Dimche Kostadinov, and Damian Borth · 2021
Later among the works it cites.
Pytorch imagenet training example
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2021
Later among the works it cites.
Towards learning universal hyperparameter optimizers with transformers, 2022
Yutian Chen, Xingyou Song, Chansoo Lee, Zi Wang, Qiuyi Zhang, David Dohan, Kazuya Kawakami, Greg Kochanski, Arnaud Doucet, Marc’aurelio Ranzato, Sagi Perel, and Nando de Freitas · 2022
Later among the works it cites.
Scalable gaussian processes for data-driven design using big data with categorical factors
Liwei Wang, Suraj Yerramilli, Akshay Iyer, Daniel Apley, Ping Zhu, and Wei Chen · 2022
Later among the works it cites.
Task selection for automl system evaluation
Jonathan Lorraine, Nihesh Anderson, Chansoo Lee, Quentin De Laroussilhe, and Mehadi Hassen · 2022
Later among the works it cites.
Multi-rate vae: Train once, get the full rate-distortion curve
Juhan Bae, Michael R Zhang, Michael Ruan, Eric Wang, So Hasegawa, Jimmy Ba, and Roger Baker Grosse · 2022
Later among the works it cites.
Tutorial on amortized optimization for learning to optimize over continuous domains
Brandon Amos · 2022
Later among the works it cites.
Supervised training of conditional monge maps
Charlotte Bunne, Andreas Krause, and Marco Cuturi · 2022
Later among the works it cites.
Learning to learn with generative models of neural network checkpoints
William Peebles, Ilija Radosavovic, Tim Brooks, Alexei A Efros, and Jitendra Malik · 2022
Later among the works it cites.
Datamodels: Predicting predictions from training data, 2022
Andrew Ilyas, Sung Min Park, Logan Engstrom, Guillaume Leclerc, and Aleksander Madry · 2022
Later among the works it cites.
Zero-shot automl with pretrained models, 2022
Ekrem Öztürk, Fabio Ferreira, Hadi S. Jomaa, Lars Schmidt-Thieme, Josif Grabocka, and Frank Hutter · 2022
Later among the works it cites.
Model zoos: A dataset of diverse populations of neural network models
Konstantin Schürholt, Diyar Taskiran, Boris Knyazev, Xavier Giró-i Nieto, and Damian Borth · 2022
Later among the works it cites.
On implicit bias in overparameterized bilevel optimization
Paul Vicol, Jonathan P Lorraine, Fabian Pedregosa, David Duvenaud, and Roger B Grosse · 2022
Later among the works it cites.
Supervising the multi-fidelity race of hyperparameter configurations, 2023
Martin Wistuba, Arlind Kadra, and Josif Grabocka · 2023
Later among the works it cites.
Graph metanetworks for processing diverse neural architectures, 2023
Derek Lim, Haggai Maron, Marc T. Law, Jonathan Lorraine, and James Lucas · 2023
Later among the works it cites.
Permutation equivariant neural functionals, 2023
Allan Zhou, Kaien Yang, Kaylee Burns, Adriano Cardace, Yiding Jiang, Samuel Sokota, J. Zico Kolter, and Chelsea Finn · 2023
Later among the works it cites.
Hyperparameter optimization: Foundations, algorithms, best practices, and open challenges
Bernd Bischl, Martin Binder, Michel Lang, Tobias Pielok, Jakob Richter, Stefan Coors, Janek Thomas, Theresa Ullmann, Marc Becker, Anne-Laure Boulesteix, et al · 2023
Later among the works it cites.
On Bilevel Optimization without Full Unrolls: Methods and Applications
Paul Adrian Vicol · 2023
Later among the works it cites.
Neural architecture search: Insights from 1000 papers, 2023
Colin White, Mahmoud Safari, Rhea Sukthanker, Binxin Ru, Thomas Elsken, Arber Zela, Debadeepta Dey, and Frank Hutter · 2023
Later among the works it cites.
Att3d: Amortized text-to-3d object synthesis
Jonathan Lorraine, Kevin Xie, Xiaohui Zeng, Chen-Hsuan Lin, Towaki Takikawa, Nicholas Sharp, Tsung-Yi Lin, Ming-Yu Liu, Sanja Fidler, and James Lucas · 2023
Later among the works it cites.
Using large language models for hyperparameter optimization
Michael R Zhang, Nishkrit Desai, Juhan Bae, Jonathan Lorraine, and Jimmy Ba · 2023
Later among the works it cites.
Equivariant architectures for learning in deep weight spaces
Aviv Navon, Aviv Shamsian, Idan Achituve, Ethan Fetaya, Gal Chechik, and Haggai Maron · 2023
Later among the works it cites.
Data selection for language models via importance resampling, 2023
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy Liang · 2023
Later among the works it cites.
Quick-tune: Quickly learning which pretrained model to finetune and how, 2024
Sebastian Pineda Arango, Fabio Ferreira, Arlind Kadra, Frank Hutter, and Josif Grabocka · 2024
Closest in time.
Javier Antoran · 2024
Closest in time.
Scalable Nested Optimization for Deep Learning
Jonathan Lorraine · 2024
Closest in time.
Latte3d: Large-scale amortized text-to-enhanced3d synthesis
Kevin Xie, Jonathan Lorraine, Tianshi Cao, Jun Gao, James Lucas, Antonio Torralba, Sanja Fidler, and Xiaohui Zeng · 2024
Closest in time.
A general framework for user-guided bayesian optimization, 2024
Carl Hvarfner, Frank Hutter, and Luigi Nardi · 2024
Closest in time.
Dsdm: Model-aware dataset selection with datamodels, 2024
Logan Engstrom, Axel Feldmann, and Aleksander Madry · 2024
Closest in time.
Graph neural networks for learning equivariant representations of neural networks, 2024
Miltiadis Kofinas, Boris Knyazev, Yan Zhang, Yunlu Chen, Gertjan J. Burghouts, Efstratios Gavves, Cees G. M. Snoek, and David W. Zhang · 2024
Closest in time.
Universal neural functionals, 2024
Allan Zhou, Chelsea Finn, and James Harrison · 2024
Closest in time.