Fetching the paper…
Reading the bibliography…
The remarkable instruction-following capability of large language models (LLMs) has sparked a growing interest in automatically finding good prompts, i.e., prompt optimization.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
A survey on practical applications of multi-armed and contextual bandits
Bouneffouf, D. and Rish, I. (2019) · 1904
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
Thompson, W. R. (1933) · 1933
Earlier work this paper cites.
Asymptotically efficient adaptive allocation rules
Lai, T. L. and Robbins, H. (1985) · 1985
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Auer, P., Cesa-Bianchi, N., and Fischer, P. (2002) · 2002
Earlier work this paper cites.
Stochastic neighbor embedding
Hinton, G. E. and Roweis, S. (2002) · 2002
Earlier work this paper cites.
Minimax policies for adversarial and stochastic bandits
Audibert, J.-Y., Bubeck, S., et al. (2009) · 2009
Earlier work this paper cites.
Best arm identification in multi-armed bandits
Audibert, J.-Y., Bubeck, S., and Munos, R. (2010) · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Li, L., Chu, W., Langford, J., and Schapire, R. E. (2010) · 2010
Earlier work this paper cites.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Shin, T., Razeghi, Y., Logan IV, R. L., Wallace, E., and Singh, S. (2020) · 2010
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y., Pál, D., and Szepesvári, C. (2011) · 2011
Earlier work this paper cites.
Multi-bandit best arm identification
Gabillon, V., Ghavamzadeh, M., Lazaric, A., and Bubeck, S. (2011) · 2011
Earlier work this paper cites.
The kl-ucb algorithm for bounded stochastic bandits and beyond
Garivier, A. and Cappé, O. (2011) · 2011
Earlier work this paper cites.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S., Cesa-Bianchi, N., et al. (2012) · 2012
Earlier work this paper cites.
Best arm identification: A unified approach to fixed budget and fixed confidence
Gabillon, V., Ghavamzadeh, M., and Lazaric, A. (2012) · 2012
Earlier work this paper cites.
Combinatorial network optimization with unknown variables: Multi-armed bandits with linear rewards and individual observations
Gai, Y., Krishnamachari, B., and Jain, R. (2012) · 2012
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
Gao, T., Fisch, A., and Chen, D. (2020) · 2012
Earlier work this paper cites.
Multiple identifications in multi-armed bandits
Bubeck, S., Wang, T., and Viswanathan, N. (2013) · 2013
Earlier work this paper cites.
Combinatorial multi-armed bandit: General framework and applications
Chen, W., Wang, Y., and Yuan, Y. (2013) · 2013
Earlier work this paper cites.
Almost optimal exploration in multi-armed bandits
Karnin, Z., Koren, T., and Somekh, O. (2013) · 2013
Earlier work this paper cites.
Combinatorial pure exploration of multi-armed bandits
Chen, S., Lin, T., King, I., Lyu, M. R., and Chen, W. (2014) · 2014
Earlier work this paper cites.
Best-arm identification algorithms for multi-armed bandits in the fixed confidence setting
Jamieson, K. and Nowak, R. (2014) · 2014
Earlier work this paper cites.
Best-arm identification in linear bandits
Soare, M., Lazaric, A., and Munos, R. (2014) · 2014
Earlier work this paper cites.
Combinatorial bandits revisited
Combes, R., Talebi Mazraeh Shahi, M. S., Proutiere, A., et al. (2015) · 2015
Earlier work this paper cites.
Taking the human out of the loop: A review of bayesian optimization
Shahriari, B., Swersky, K., Wang, Z., Adams, R. P., and De Freitas, N. (2015) · 2015
Earlier work this paper cites.
Combinatorial multi-armed bandit with general reward functions
Chen, W., Hu, W., Li, F., Li, J., Liu, Y., and Lu, P. (2016) · 2016
Earlier work this paper cites.
Optimal best arm identification with fixed confidence
Garivier, A. and Kaufmann, E. (2016) · 2016
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Cited alongside, same era.
Optimal best-arm identification in linear bandits
Jedra, Y. and Proutiere, A. (2020) · 2020
Cited alongside, same era.
How can we know what language models know?
Jiang, Z., Xu, F. F., Araki, J., and Neubig, G. (2020) · 2020
Cited alongside, same era.
Bandit algorithms
Lattimore, T. and Szepesvári, C. (2020) · 2020
Cited alongside, same era.
Learning for dose allocation in adaptive clinical trials with safety constraints
Shen, C., Wang, Z., Villar, S., and Van Der Schaar, M. (2020) · 2020
Gps: Genetic prompt search for efficient few-shot learning
Xu, H., Chen, Y., Du, Y., Shao, N., Wang, Y., Li, H., and Yang, Z. (2022) · 2022
Later among the works it cites.
Minimax optimal fixed-budget best arm identification in linear bandits
Yang, J. and Tan, V. (2022) · 2022
Later among the works it cites.
Automatic chain of thought prompting in large language models
Zhang, Z., Zhang, A., Li, M., and Smola, A. (2022) · 2022
Later among the works it cites.
Large language models are human-level prompt engineers
Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J. (2022) · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural contextual bandits with ucb-based exploration
Zhou, D., Li, L., and Gu, Q. (2020) · 2020
Cited alongside, same era.
Pure exploration in kernel and neural bandits
Zhu, Y., Zhou, D., Jiang, R., Gu, Q., Willett, R., and Nowak, R. (2021) · 2020
Cited alongside, same era.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N. (2021) · 2021
Cited alongside, same era.
Prefix-tuning: Optimizing continuous prompts for generation
Li, X. L. and Liang, P. (2021) · 2021
Cited alongside, same era.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P. (2021) · 2021
Cited alongside, same era.
Reframing instructional prompts to gptk’s language
Mishra, S., Khashabi, D., Baral, C., Choi, Y., and Hajishirzi, H. (2021) · 2021
Cited alongside, same era.
Azizi, M. J., Kveton, B., and Ghavamzadeh, M. (2023) · 2023
Later among the works it cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. (2023) · 2023
Later among the works it cites.
Instructzero: Efficient instruction optimization for black-box large language models
Chen, L., Chen, J., Goldstein, T., Huang, H., and Zhou, T. (2023) · 2023
Later among the works it cites.
Connecting large language models with evolutionary algorithms yields powerful prompt optimizers
Guo, Q., Wang, R., Guo, J., Li, B., Song, K., Tan, X., Liu, G., Bian, J., and Yang, Y. (2023) · 2023
Later among the works it cites.
Mistral 7b
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E. (2023) · 2023
Later among the works it cites.
Rate-optimal bayesian simple regret in best arm identification
Komiyama, J., Ariu, K., Kato, M., and Qin, C. (2023) · 2023
Later among the works it cites.
Diverse demonstrations improve in-context compositional generalization
Levy, I., Bogin, B., and Berant, J. (2023) · 2023
Later among the works it cites.
Unified demonstration retriever for in-context learning
Li, X., Lv, K., Yan, H., Lin, T., Zhu, W., Ni, Y., Xie, G., Wang, X., and Qiu, X. (2023) · 2023
Later among the works it cites.
Use your instinct: Instruction optimization using neural bandits coupled with transformers
Lin, X., Wu, Z., Dai, Z., Hu, W., Shu, Y., Ng, S.-K., Jaillet, P., and Low, B. K. H. (2023) · 2023
Later among the works it cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G. (2023) · 2023
Later among the works it cites.
Plum: Prompt learning using metaheuristic
Pan, R., Xing, S., Diao, S., Liu, X., Shum, K., Zhang, J., and Zhang, T. (2023) · 2023
Later among the works it cites.
Automatic prompt optimization with “gradient descent” and beam search
Pryzant, R., Iter, D., Li, J., Lee, Y. T., Zhu, C., and Zeng, M. (2023) · 2023
Later among the works it cites.
Query-dependent prompt evaluation and optimization with offline inverse rl
Sun, H., Hüyük, A., and van der Schaar, M. (2023) · 2023
Later among the works it cites.
Best arm identification with fixed budget: A large deviation perspective
Wang, P.-A., Tzeng, R.-C., and Proutiere, A. (2023) · 2023
Later among the works it cites.
Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery
Wen, Y., Jain, N., Kirchenbauer, J., Goldblum, M., Geiping, J., and Goldstein, T. (2023) · 2023
Later among the works it cites.
Large language models as optimizers
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X. (2023) · 2023
Later among the works it cites.
Fixed-budget best-arm identification in sparse linear bandits
Yavas, R. C. and Tan, V. Y. (2023) · 2023
Later among the works it cites.
Cost aware best arm identification
Kanarios, K., Zhang, Q., and Ying, L. (2024) · 2024
Closest in time.
Gpt-3.5-turbo
OpenAI (2023a) · 2024
Closest in time.
Text-embedding-ada-002
OpenAI (2023b) · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Team, G., Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al. (2024) · 2024
Closest in time.
Misconfidence-based demonstration selection for llm in-context learning
Xu, S. and Zhang, C. (2024) · 2024
Closest in time.