Fetching the paper…
Reading the bibliography…
This paper explores the use of foundational large language models (LLMs) in hyperparameter optimization (HPO).
An automatic method for finding the greatest or least value of a function
HoHo Rosenbrock · 1960
Earlier work this paper cites.
The global optimization problem: an introduction
Laurence Charles Ward Dixon · 1978
Earlier work this paper cites.
A connectionist machine for genetic hillclimbing, 1987
D Ackley · 1987
Earlier work this paper cites.
The application of bayesian methods for seeking the extremum
Jonas Mockus · 1998
Earlier work this paper cites.
Sequential model-based optimization for general algorithm configuration
Frank Hutter, Holger H Hoos, and Kevin Leyton-Brown · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Random search for hyper-parameter optimization
James Bergstra and Yoshua Bengio · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Jasper Snoek, Hugo Larochelle, and Ryan P Adams · 2012
Earlier work this paper cites.
Hyperopt: A python library for optimizing the hyperparameters of machine learning algorithms
James Bergstra, Dan Yamins, David D Cox, et al · 2013
Earlier work this paper cites.
Multi-task bayesian optimization
Kevin Swersky, Jasper Snoek, and Ryan P Adams · 2013
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Earlier work this paper cites.
Taking the human out of the loop: A review of bayesian optimization
Bobak Shahriari, Kevin Swersky, Ziyu Wang, Ryan P Adams, and Nando De Freitas · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Non-stochastic best arm identification and hyperparameter optimization
Kevin Jamieson and Ameet Talwalkar · 2016
Earlier work this paper cites.
Bayesian optimization with a finite budget: An approximate dynamic programming approach
Remi Lam, Karen Willcox, and David H Wolpert · 2016
Earlier work this paper cites.
Forward and reverse gradient-based hyperparameter optimization
Luca Franceschi, Michele Donini, Paolo Frasconi, and Massimiliano Pontil · 2017
Earlier work this paper cites.
Population based training of neural networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, et al · 2017
Earlier work this paper cites.
Fast bayesian optimization of machine learning hyperparameters on large datasets
Aaron Klein, Stefan Falkner, Simon Bartels, Philipp Hennig, and Frank Hutter · 2017
Earlier work this paper cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar · 2017
Earlier work this paper cites.
Ae: A domain-agnostic platform for adaptive experimentation
Eytan Bakshy, Lili Dworkin, Brian Karrer, Konstantin Kashin, Benjamin Letham, Ashwin Murthy, and Shaun Singh · 2018
Earlier work this paper cites.
Applied nonlinear programming
David M Himmelblau et al · 2018
Earlier work this paper cites.
Stochastic hyperparameter optimization through hypernetworks
Jonathan Lorraine and David Duvenaud · 2018
Cited alongside, same era.
Hyperparameter optimization
Matthias Feurer and Frank Hutter · 2019
Cited alongside, same era.
Towards assessing the impact of bayesian optimization’s own hyperparameters
Marius Lindauer, Matthias Feurer, Katharina Eggensperger, André Biedenkapp, and Frank Hutter · 2019
Cited alongside, same era.
Matthew MacKay, Paul Vicol, Jon Lorraine, David Duvenaud, and Roger Grosse · 2019
Cited alongside, same era.
Delta-stn: Efficient bilevel optimization for neural networks using structured response jacobians
Juhan Bae and Roger B Grosse · 2020
Task selection for automl system evaluation
Jonathan Lorraine, Nihesh Anderson, Chansoo Lee, Quentin De Laroussilhe, and Mehadi Hassen · 2022
Later among the works it cites.
On implicit bias in overparameterized bilevel optimization
Paul Vicol, Jonathan P Lorraine, Fabian Pedregosa, David Duvenaud, and Roger B Grosse · 2022
Later among the works it cites.
Large language models are human-level prompt engineers
Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba · 2022
Later among the works it cites.
Hyperparameter optimization: Foundations, algorithms, best practices, and open challenges
Bernd Bischl, Martin Binder, Michel Lang, Tobias Pielok, Jakob Richter, Stefan Coors, Janek Thomas, Theresa Ullmann, Marc Becker, Anne-Laure Boulesteix, et al · 2023
Closest in time.
Eight things to know about large language models
Samuel R Bowman · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Botorch: A framework for efficient monte-carlo bayesian optimization
Maximilian Balandat, Brian Karrer, Daniel Jiang, Samuel Daulton, Ben Letham, Andrew G Wilson, and Eytan Bakshy · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Tuning hyperparameters without grad students: Scalable and robust bayesian optimisation with dragonfly
Kirthevasan Kandasamy, Karun Raju Vysyaraju, Willie Neiswanger, Biswajit Paria, Christopher R Collins, Jeff Schneider, Barnabas Poczos, and Eric P Xing · 2020
Cited alongside, same era.
Optimizing millions of hyperparameters by implicit differentiation
Jonathan Lorraine, Paul Vicol, and David Duvenaud · 2020
Cited alongside, same era.
On hyperparameter optimization of machine learning algorithms: Theory and practice
Li Yang and Abdallah Shami · 2020
Cited alongside, same era.
Hpobench: A collection of reproducible multi-fidelity benchmark problems for hpo
Katharina Eggensperger, Philipp Müller, Neeratyoy Mallik, Matthias Feurer, René Sass, Aaron Klein, Noor Awad, Marius Lindauer, and Frank Hutter · 2021
Cited alongside, same era.
Evoprompting: Language models for code-level neural architecture search
Angelica Chen, David M Dohan, and David R So · 2023
Closest in time.
Walid Hariri · 2023
Closest in time.
Noah Hollmann, Samuel Müller, and Frank Hutter · 2023
Closest in time.
Survey of hallucination in natural language generation
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung · 2023
Closest in time.
Measuring faithfulness in chain-of-thought reasoning
Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, et al · 2023
Closest in time.
Summary of chatgpt-related research and perspective towards the future of large language models
Yiheng Liu, Tianle Han, Siyuan Ma, Jiayue Zhang, Yuanyuan Yang, Jiaming Tian, Hao He, Antong Li, Mengshen He, Zhengliang Liu, et al · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
NYC Taxi Trip Records from JAN 2023 to JUN 2023, 2023a
NYC Taxi and Limousine Commission · 2023
Closest in time.
Self-taught optimizer (stop): Recursively self-improving code generation
Eric Zelikman, Eliana Lorch, Lester Mackey, and Adam Tauman Kalai · 2023
Closest in time.
Automl-gpt: Automatic machine learning with gpt
Shujian Zhang, Chengyue Gong, Lemeng Wu, Xingchao Liu, and Mingyuan Zhou · 2023
Closest in time.
Can gpt-4 perform neural architecture search?
Mingkai Zheng, Xiu Su, Shan You, Fei Wang, Chen Qian, Chang Xu, and Samuel Albanie · 2023
Closest in time.
Scalable Nested Optimization for Deep Learning
Jonathan Lorraine · 2024
Closest in time.
Improving hyperparameter optimization with checkpointed model weights
Nikhil Mehta, Jonathan Lorraine, Steve Masson, Ramanathan Arunachalam, Zaid Pervaiz Bhat, James Lucas, and Arun George Zachariah · 2024
Closest in time.
Large language models as optimizers
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V Le, Denny Zhou, and Xinyun Chen · 2024
Closest in time.