Fetching the paper…
Reading the bibliography…
Well-designed prompts are crucial for enhancing Large language models' (LLMs) reasoning capabilities while aligning their outputs with task requirements across diverse domains.
The winograd schema challenge
Levesque, H. J., Davis, E., and Morgenstern, L · 2012
Earlier work this paper cites.
”liar, liar pants on fire”: A new benchmark dataset for fake news detection
Wang, W. Y · 2017
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Earlier work this paper cites.
Rephrase and respond: Let large language models ask better questions for themselves
Deng, Y., Zhang, W., Chen, Z., and Gu, Q · 2023
Earlier work this paper cites.
Dspy: Compiling declarative language model calls into self-improving pipelines
Khattab, O., Singhvi, A., Maheshwari, P., Zhang, Z., Santhanam, K., Vardhamanan, S., Haq, S., Sharma, A., Joshi, T. T., Moazam, H., Miller, H., Zaharia, M., and Potts, C · 2023
Earlier work this paper cites.
Automatic prompt optimization with ”gradient descent” and beam search
Pryzant, R., Iter, D., Li, J., Lee, Y. T., Zhu, C., and Zeng, M · 2023
Earlier work this paper cites.
GPQA: A graduate-level google-proof q&a benchmark
Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R · 2023
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., and Wei, J · 2023
Earlier work this paper cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Earlier work this paper cites.
Large language models are human-level prompt engineers
Zhou, Y., Muresanu, A. I., Han, Z., Paster, K., Pitis, S., Chan, H., and Ba, J · 2023
Cited alongside, same era.
Efficient prompting methods for large language models: A survey
Chang, K., Xu, S., Wang, C., Luo, Y., Xiao, T., and Zhu, J · 2024
Cited alongside, same era.
Prompt optimization in multi-step tasks (PROMST): integrating human feedback and heuristic-based sampling
Chen, Y., Arkin, J., Hao, Y., Zhang, Y., Roy, N., and Fan, C · 2024
Cited alongside, same era.
Promptbreeder: Self-referential self-improvement via prompt evolution
Fernando, C., Banarse, D., Michalewski, H., Osindero, S., and Rocktäschel, T · 2024
Cited alongside, same era.
Does prompt formatting have any impact on LLM performance?
He, J., Rungta, M., Koleczek, D., Sekhon, A., Wang, F. X., and Hasan, S · 2024
Cited alongside, same era.
Voyager: An open-ended embodied agent with large language models
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A · 2024
Later among the works it cites.
Efficient and accurate prompt optimization: the benefit of memory in exemplar-guided reflection
Yan, C., Wang, J., Zhang, L., Zhao, R., Wu, X., Xiong, K., Liu, Q., Kang, G., and Kang, Y · 2024
Later among the works it cites.
Generative AI for visualization: State of the art and future directions
Ye, Y., Hao, J., Hou, Y., Wang, Z., Xiao, S., Luo, Y., and Zeng, W · 2024
Later among the works it cites.
Textgrad: Automatic ”differentiation” via text
Yüksekgönül, M., Bianchi, F., Boen, J., Liu, S., Huang, Z., Guestrin, C., and Zou, J · 2024
Later among the works it cites.
Take a step back: Evoking reasoning via abstraction in large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Prompt optimization with human feedback
Lin, X., Dai, Z., Verma, A., Ng, S., Jaillet, P., and Low, B. K. H · 2024
Cited alongside, same era.
Code generation with alphacodium: From prompt engineering to flow engineering
Ridnik, T., Kredo, D., and Friedman, I · 2024
Cited alongside, same era.
Archon: An architecture search framework for inference-time techniques
Saad-Falcon, J., Lafuente, A. G., Natarajan, S., Maru, N., Todorov, H., Guha, E., Buchanan, E. K., Chen, M., Guha, N., Ré, C., and Mirhoseini, A · 2024
Cited alongside, same era.
Data playwright: Authoring data videos with annotated narration
Shen, L., Li, H., Wang, Y., Luo, T., Luo, Y., and Qu, H · 2024
Cited alongside, same era.
Tam, Z. R., Wu, C., Tsai, Y., Lin, C., Lee, H., and Chen, Y · 2024
Cited alongside, same era.
Connecting large language models with evolutionary algorithms yields powerful prompt optimizers
Guo, Q., Wang, R., Guo, J., Li, B., Song, K., Tan, X., Liu, G., Bian, J., and Yang, Y
Cited in the paper.
Large language model based multi-agents: A survey of progress and challenges
Guo, T., Chen, X., Wang, Y., Chang, R., Pei, S., Chawla, N. V., Wiest, O., and Zhang, X
Cited in the paper.
Zheng, H. S., Mishra, S., Chen, X., Cheng, H., Chi, E. H., Le, Q. V., and Zhou, D · 2024
Later among the works it cites.
Agieval: A human-centric benchmark for evaluating foundation models
Zhong, W., Cui, R., Guo, Y., Liang, Y., Lu, S., Wang, Y., Saied, A., Chen, W., and Duan, N · 2024
Later among the works it cites.
Fairer preferences elicit improved human-aligned large language model judgments
Zhou, H., Wan, X., Liu, Y., Collier, N., Vulic, I., and Korhonen, A · 2024
Later among the works it cites.
Are large language models good statisticians?
Zhu, Y., Du, S., Li, B., Luo, Y., and Tang, N · 2024
Later among the works it cites.
Test-time preference optimization: On-the-fly alignment via iterative textual feedback
Li, Y., Hu, X., Qu, X., Li, L., and Cheng, Y · 2025
Closest in time.
Agent laboratory: Using llm agents as research assistants
Schmidgall, S., Su, Y., Wang, Z., Sun, X., Wu, J., Yu, X., Liu, J., Liu, Z., and Barsoum, E · 2025
Closest in time.