Fetching the paper…
Reading the bibliography…
Undeniably, Large Language Models (LLMs) have stirred an extraordinary wave of innovation in the machine learning research domain, resulting in substantial impact across diverse fields such as reinforcement learning, robotics, and computer vision.
The traveling-salesman problem
Flood, M. M · 1956
Earlier work this paper cites.
Introduction to Automata Theory, Languages and Computation
Hopcroft, J. E. and Ullman, J. D · 1979
Earlier work this paper cites.
Upper and lower bounds for randomized search heuristics in black-box optimization
Droste, S., Jansen, T., and Wegener, I · 2006
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
Bergstra, J., Bardenet, R., Bengio, Y., and Kégl, B · 2011
Earlier work this paper cites.
Sequential model-based optimization for general algorithm configuration
Hutter, F., Hoos, H. H., and Leyton-Brown, K · 2011
Earlier work this paper cites.
A textbook of graph theory
Balakrishnan, R. and Ranganathan, K · 2012
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
Snoek, J., Larochelle, H., and Adams, R. P · 2012
Earlier work this paper cites.
Multi-task bayesian optimization
Swersky, K., Snoek, J., and Adams, R. P · 2013
Earlier work this paper cites.
Efficient transfer learning method for automatic hyperparameter tuning
Yogatama, D. and Mann, G · 2014
Earlier work this paper cites.
Efficient and robust automated machine learning
Feurer, M., Klein, A., Eggensperger, K., Springenberg, J., Blum, M., and Hutter, F · 2015
Earlier work this paper cites.
Taking the human out of the loop: A review of bayesian optimization
Shahriari, B., Swersky, K., Wang, Z., Adams, R. P., and de Freitas, N · 2015
Earlier work this paper cites.
The cma evolution strategy: A tutorial
Hansen, N · 2016
Earlier work this paper cites.
Li, K. and Malik, J · 2016
Earlier work this paper cites.
Neural combinatorial optimization with reinforcement learning
Bello, I., Pham, H., Le, Q. V., Norouzi, M., and Bengio, S · 2017
Earlier work this paper cites.
Learning to learn without gradient descent by gradient descent
Chen, Y., Hoffman, M. W., Colmenarejo, S. G., Denil, M., Lillicrap, T. P., Botvinick, M., and Freitas, N · 2017
Earlier work this paper cites.
Google vizier: A service for black-box optimization
Golovin, D., Solnik, B., Moitra, S., Kochanski, G., Karro, J., and Sculley, D · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Kudo, T. and Richardson, J · 2018
Earlier work this paper cites.
Scalable hyperparameter transfer learning
Perrone, V., Jenatton, R., Seeger, M. W., and Archambeau, C · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
Coco: the large scale black-box optimization benchmarking (bbob-largescale) test suite
Elhara, O., Varelas, K., Nguyen, D., Tusar, T., Brockhoff, D., Hansen, N., and Auger, A · 2019
Earlier work this paper cites.
Combinatorial bayesian optimization using the graph cartesian product
Oh, C., Tomczak, J. M., Gavves, E., and Welling, M · 2019
Earlier work this paper cites.
Meta-learning acquisition functions for transfer learning in bayesian optimization
Volpp, M., Fröhlich, L. P., Fischer, K., Doerr, A., Falkner, S., Hutter, F., and Daniel, C · 2019
Earlier work this paper cites.
Nas-bench-101: Towards reproducible neural architecture search
Ying, C., Klein, A., Christiansen, E., Real, E., Murphy, K., and Hutter, F · 2019
Earlier work this paper cites.
Population-based black-box optimization for biological sequence design
Angermueller, C., Belanger, D., Gane, A., Mariet, Z., Dohan, D., Murphy, K., Colwell, L., and Sculley, D · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Meta-learning requires meta-augmentation
Rajendran, J., Irpan, A., and Jang, E · 2020
Earlier work this paper cites.
Automl-zero: Evolving machine learning algorithms from scratch
Real, E., Liang, C., So, D., and Le, Q · 2020
Cited alongside, same era.
Predicting neural network accuracy from weights
Unterthiner, T., Keysers, D., Gelly, S., Bousquet, O., and Tolstikhin, I. O · 2020
Cited alongside, same era.
Openml: A benchmarking layer on top of openml to quickly create, download, and share systematic benchmarks
Bischl, B., Casalicchio, G., Feurer, M., Gijsbers, P., Hutter, F., Lang, M., Mantovani, R. G., van Rijn, J. N., and Vanschoren, J · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R. B., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N. S., Chen, A. S., Creel, K., Davis, J. Q., Demszky, D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etchemendy, J., Ethayarajh, K., Fei-Fei, L., Finn, C., Gale, T., Gillespie, L., Goel, K., Goodman, N. D., Grossman, S., Guha, N., Hashimoto, T., Henderson, P., Hewitt, J., Ho, D. E., Hong, J., Hsu, K., Huang, J., Icard, T., Jain, S., Jurafsky, D., Kalluri, P., Karamcheti, S., Keeling, G., Khani, F., Khattab, O., Koh, P. W., Krass, M. S., Krishna, R., Kuditipudi, R., and et al · 2021
Evolution through large models
Lehman, J., Gordon, J., Jain, S., Ndousse, K., Yeh, C., and Stanley, K. O · 2023
Later among the works it cites.
Eureka: Human-level reward design via coding large language models
Ma, Y. J., Liang, W., Wang, G., Huang, D.-A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Later among the works it cites.
Large language models generate functional protein sequences across diverse families
Madani, A., Krause, B., Greene, E. R., Subramanian, S., Mohr, B. P., Holton, J. M., Olmos, J. L., Xiong, C., Sun, Z. Z., Socher, R., Fraser, J. S., and Naik, N · 2023
Later among the works it cites.
Generative pretraining for black-box optimization
Mashkaria, S. M., Krishnamoorthy, S., and Grover, A · 2023
Later among the works it cites.
Language model crossover: Variation through few-shot prompting, 2023
Meyerson, E., Nelson, M. J., Bradley, H., Moradi, A., Hoover, A. K., and Lehman, J · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Cited alongside, same era.
Hpobench: A collection of reproducible multi-fidelity benchmark problems for hpo
Eggensperger, K., Müller, P., Mallik, N., Feurer, M., Sass, R., Klein, A., Awad, N., Lindauer, M., and Hutter, F · 2021
Cited alongside, same era.
Adaptive experimentation platform, 2021
Facebook · 2021
Cited alongside, same era.
Neural architecture search without training
Mellor, J., Turner, J., Storkey, A., and Crowley, E. J · 2021
Cited alongside, same era.
Transformers can do bayesian inference
Müller, S., Hollmann, N., Arango, S. P., Grabocka, J., and Hutter, F · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Interpretable neural architecture search via bayesian optimisation with weisfeiler-lehman kernels
Ru, B. X., Wan, X., Dong, X., and Osborne, M. A · 2021
Cited alongside, same era.
Bayesian optimization is superior to random search for machine learning hyperparameter tuning: Analysis of the black-box optimization challenge 2020
Turner, R., Eriksson, D., McCourt, M., Kiili, J., Laaksonen, E., Xu, Z., and Guyon, I · 2021
Cited alongside, same era.
Müller, S., Feurer, M., Hollmann, N., and Hutter, F · 2023
Later among the works it cites.
Llmatic: Neural architecture search via large language models and quality-diversity optimization
Nasir, M. U., Earle, S., Togelius, J., James, S., and Cleghorn, C. W · 2023
Later among the works it cites.
Expt: Synthetic pretraining for few-shot experimental design
Nguyen, T., Agrawal, S., and Grover, A · 2023
Later among the works it cites.
Importance of directional feedback for llm-based optimizers
Nie, A., Cheng, C.-A., Kolobov, A., and Swaminathan, A · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
OpenLLM: Operating LLMs in production, June 2023
Pham, A., Yang, C., Sheng, S., Zhao, S., Lee, S., Jiang, B., Dong, F., Guan, X., and Ming, F · 2023
Later among the works it cites.
Mathematical discoveries from program search with large language models
Romera-Paredes, B., Barekatain, M., Novikov, A., Balog, M., Kumar, M. P., Dupont, E., Ruiz, F. J., Ellenberg, J. S., Wang, P., Fawzi, O., et al · 2023
Later among the works it cites.
Languagempc: Large language models as decision makers for autonomous driving
Sha, H., Mu, Y., Jiang, Y., Chen, L., Xu, C., Luo, P., Li, S. E., Tomizuka, M., Zhan, W., and Ding, M · 2023
Later among the works it cites.
Mixture-of-experts meets instruction tuning: A winning combination for large language models
Shen, S., Hou, L., Zhou, Y., Du, N., Longpre, S., Wei, J., Chung, H. W., Zoph, B., Fedus, W., Chen, X., et al · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D., Blecher, L., Canton-Ferrer, C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I., Korenev, A., Koura, P. S., Lachaux, M., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X. E., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Later among the works it cites.
Review of large vision models and visual prompt engineering
Wang, J., Liu, Z., Zhao, L., Wu, Z., Ma, C., Yu, S., Dai, H., Yang, Q., Liu, Y., Zhang, S., Shi, E., Pan, Y., Zhang, T., Zhu, D., Li, X., Jiang, X., Ge, B., Yuan, Y., Shen, D., Liu, T., and Zhang, S · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2023
Later among the works it cites.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., and Hooi, B · 2023
Later among the works it cites.
Large language models as optimizers, 2023
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X · 2023
Later among the works it cites.
Language to rewards for robotic skill synthesis
Yu, W., Gileadi, N., Fu, C., Kirmani, S., Lee, K.-H., Arenas, M. G., Chiang, H.-T. L., Erez, T., Hasenclever, L., Humplik, J., et al · 2023
Later among the works it cites.
Using large language models for hyperparameter optimization
Zhang, M., Desai, N., Bae, J., Lorraine, J., and Ba, J · 2023
Later among the works it cites.
Llm augmented llms: Expanding capabilities through composition
Bansal, R., Samanta, B., Dalmia, S., Gupta, N., Vashishth, S., Ganapathy, S., Bapna, A., Jain, P., and Talukdar, P · 2024
Closest in time.
Llm maybe longlm: Self-extend llm context window without tuning
Jin, H., Han, X., Yang, J., Jiang, Z., Liu, Z., Chang, C.-Y., Chen, H., and Hu, X · 2024
Closest in time.
Can large language models explore in-context?, 2024
Krishnamurthy, A., Harris, K., Foster, D. J., Zhang, C., and Slivkins, A · 2024
Closest in time.
Large language models to enhance bayesian optimization
Liu, T., Astorga, N., Seedat, N., and van der Schaar, M · 2024
Closest in time.
From understanding to utilization: A survey on explainability for large language models
Luo, H. and Specia, L · 2024
Closest in time.
Omnipred: Language models as universal regressors
Song, X., Li, O., Lee, C., Yang, B., Peng, D., Perel, S., and Chen, Y · 2024
Closest in time.