Fetching the paper…
Reading the bibliography…
Existing automatic prompt engineering methods are typically designed for discriminative tasks, where new task prompts are iteratively refined with limited feedback from a single metric reflecting a single aspect.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
Lin, C.-Y. 2004 · 2004
Earlier work this paper cites.
Visualizing data using t-SNE
Van der Maaten, L.; and Hinton, G. 2008 · 2008
Earlier work this paper cites.
Teaching machines to read and comprehend
Hermann, K. M.; Kocisky, T.; Grefenstette, E.; Espeholt, L.; Kay, W.; Suleyman, M.; and Blunsom, P. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P.; Zhang, J.; Lopyrev, K.; and Liang, P. 2016 · 2016
Earlier work this paper cites.
Creating training corpora for nlg micro-planning
Gardent, C.; Shimorina, A.; Narayan, S.; and Perez-Beltrachini, L. 2017 · 2017
Earlier work this paper cites.
TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Joshi, M.; Choi, E.; Weld, D.; and Zettlemoyer, L. 2017 · 2017
Earlier work this paper cites.
DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset
Li, Y.; Su, H.; Shen, X.; Li, W.; Cao, Z.; and Niu, S. 2017 · 2017
Earlier work this paper cites.
The narrativeqa reading comprehension challenge
Kočiskỳ, T.; Schwarz, J.; Blunsom, P.; Dyer, C.; Hermann, K. M.; Melis, G.; and Grefenstette, E. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Earlier work this paper cites.
SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization
Gliwa, B.; Mochol, I.; Biesek, M.; and Wawer, A. 2019 · 2019
Earlier work this paper cites.
Natural Questions: A Benchmark for Question Answering Research
Kwiatkowski, T.; Palomaki, J.; Redfield, O.; Collins, M.; Parikh, A.; Alberti, C.; Epstein, D.; Polosukhin, I.; Devlin, J.; Lee, K.; Toutanova, K.; Jones, L.; Kelcey, M.; Chang, M.-W.; Dai, A. M.; Uszkoreit, J.; Le, Q.; and Petrov, S. 2019 · 2019
Earlier work this paper cites.
Multi-style Generative Reading Comprehension
Nishida, K.; Saito, I.; Nishida, K.; Shinoda, K.; Otsuka, A.; Asano, H.; and Tomita, J. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reimers, N.; and Gurevych, I. 2019 · 2019
Earlier work this paper cites.
BERTScore: Evaluating Text Generation with BERT
Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2019 · 2019
Earlier work this paper cites.
Dice Loss for Data-imbalanced NLP Tasks
Li, X.; Sun, X.; Meng, Y.; Liang, J.; Wu, F.; and Li, J. 2020 · 2020
Earlier work this paper cites.
SummEval: Re-evaluating Summarization Evaluation
Fabbri, A. R.; Kryściński, W.; McCann, B.; Xiong, C.; Socher, R.; and Radev, D. 2021 · 2021
Earlier work this paper cites.
The Stem Cell Hypothesis: Dilemma behind Multi-Task Learning with Transformer Encoders
He, H.; and Choi, J. D. 2021 · 2021
Earlier work this paper cites.
Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering
Izacard, G.; and Grave, E. 2021 · 2021
Cited alongside, same era.
DialSummEval: Revisiting Summarization Evaluation for Dialogues
Gao, M.; and Wan, X. 2022 · 2022
Cited alongside, same era.
What Makes Good In-Context Examples for GPT-3?
Liu, J.; Shen, D.; Zhang, Y.; Dolan, B.; Carin, L.; and Chen, W. 2022 · 2022
Cited alongside, same era.
Universal Evasion Attacks on Summarization Scoring
Mu, W.; and Lim, K. H. 2022 · 2022
Cited alongside, same era.
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Pal, A.; Umapathi, L. K.; and Sankarasubbu, M. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Xia, F.; Chi, E.; Le, Q. V.; Zhou, D.; et al. 2022 · 2022
Pan, L.; Saxon, M.; Xu, W.; Nathani, D.; Wang, X.; and Wang, W. Y. 2023 · 2023
Later among the works it cites.
GrIPS: Gradient-free, Edit-based Instruction Search for Prompting Large Language Models
Prasad, A.; Hase, P.; Zhou, X.; and Bansal, M. 2023 · 2023
Later among the works it cites.
Automatic Prompt Optimization with “Gradient Descent” and Beam Search
Pryzant, R.; Iter, D.; Li, J.; Lee, Y.; Zhu, C.; and Zeng, M. 2023 · 2023
Later among the works it cites.
Reflexion: language agents with verbal reinforcement learning
Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; and Yao, S. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
GPS: Genetic Prompt Search for Efficient Few-Shot Learning
Xu, H.; Chen, Y.; Du, Y.; Shao, N.; Yanggang, W.; Li, H.; and Yang, Z. 2022 · 2022
Cited alongside, same era.
Large Language Models are Human-Level Prompt Engineers
Zhou, Y.; Muresanu, A. I.; Han, Z.; Paster, K.; Pitis, S.; Chan, H.; and Ba, J. 2022 · 2022
Cited alongside, same era.
Factuality challenges in the era of large language models
Augenstein, I.; Baldwin, T.; Cha, M.; Chakraborty, T.; Ciampaglia, G. L.; Corney, D.; DiResta, R.; Ferrara, E.; Hale, S.; Halevy, A.; et al. 2023 · 2023
Cited alongside, same era.
Promptbreeder: Self-referential self-improvement via prompt evolution
Fernando, C.; Banarse, D.; Michalewski, H.; Osindero, S.; and Rocktäschel, T. 2023 · 2023
Cited alongside, same era.
Self-verification improves few-shot clinical information extraction
Gero, Z.; Singh, C.; Cheng, H.; Naumann, T.; Galley, M.; Gao, J.; and Poon, H. 2023 · 2023
Cited alongside, same era.
Connecting Large Language Models with Evolutionary Algorithms Yields Powerful Prompt Optimizers
Guo, Q.; Wang, R.; Guo, J.; Li, B.; Song, K.; Tan, X.; Liu, G.; Bian, J.; and Yang, Y. 2023 · 2023
Cited alongside, same era.
Later among the works it cites.
Instructive Dialogue Summarization with Query Aggregations
Wang, B.; Liu, Z.; and Chen, N. 2023 · 2023
Later among the works it cites.
PromptAgent: Strategic Planning with Language Models Enables Expert-level Prompt Optimization
Wang, X.; Li, C.; Wang, Z.; Bai, F.; Luo, H.; Zhang, J.; Jojic, N.; Xing, E.; and Hu, Z. 2023 · 2023
Later among the works it cites.
Large Language Models as Optimizers
Yang, C.; Wang, X.; Lu, Y.; Liu, H.; Le, Q. V.; Zhou, D.; and Chen, X. 2023 · 2023
Later among the works it cites.
Aci-bench: a novel ambient clinical intelligence dataset for benchmarking automatic visit note generation
Yim, W.-w.; Fu, Y.; Ben Abacha, A.; Snider, N.; Lin, T.; and Yetisgen, M. 2023 · 2023
Later among the works it cites.
AlignScore: Evaluating Factual Consistency with A Unified Alignment Function
Zha, Y.; Yang, Y.; Li, R.; and Hu, Z. 2023 · 2023
Later among the works it cites.
Claude instant model 1.2
Anthropic. 2023 · 2024
Closest in time.
The claude 3 model family: Opus, sonnet, haiku
Anthropic, A. 2024 · 2024
Closest in time.
Elangovan, A.; Liu, L.; Xu, L.; Bodapati, S.; and Roth, D. 2024 · 2024
Closest in time.
Self-refine: Iterative refinement with self-feedback
Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. 2024 · 2024
Closest in time.
LLaMA 3 Model
MetaAI. 2024 · 2024
Closest in time.
Joint prompt optimization of stacked llms using variational inference
Sordoni, A.; Yuan, E.; Côté, M.-A.; Pereira, M.; Trischler, A.; Xiao, Z.; Hosseini, A.; Niedtner, F.; and Le Roux, N. 2024 · 2024
Closest in time.
Predicting Text Preference Via Structured Comparative Reasoning
Yan, J. N.; Liu, T.; Chiu, J.; Shen, J.; Qin, Z.; Yu, Y.; Lakshmanan, C.; Kurzion, Y.; Rush, A.; Liu, J.; and Bendersky, M. 2024 · 2024
Closest in time.
TextGrad: Automatic ”Differentiation” via Text
Yuksekgonul, M.; Bianchi, F.; Boen, J.; Liu, S.; Huang, Z.; Guestrin, C.; and Zou, J. 2024 · 2024
Closest in time.