Fetching the paper…
Reading the bibliography…
Frontier models can be prompted or conditioned to do many tasks, but finding good prompts is not always easy, nor is understanding some performant prompts.
“On the likelihood that one unknown probability exceeds another in view of the evidence of two samples”
William Thompson · 1933
Earlier work this paper cites.
“On perceptual readiness.”
Jerome Bruner · 1957
Earlier work this paper cites.
“Finding structure in time”
Jeffrey Elman · 1990
Earlier work this paper cites.
“Quadratic programming is in NP”
Stephen Vavasis · 1990
Earlier work this paper cites.
“A strategy of win-stay, lose-shift that outperforms tit-for-tat in the Prisoner’s Dilemma game”
Martin Nowak and Karl Sigmund · 1993
Earlier work this paper cites.
“Long short-term memory.”
S Hochreiter and J Schmidhuber · 1997
Earlier work this paper cites.
“Latent dirichlet allocation”
David Blei, Andrew Ng and Michael Jordan · 2003
Earlier work this paper cites.
“Learning Word Vectors for Sentiment Analysis”
Andrew. Maas et al · 2011
Earlier work this paper cites.
“Analysis of thompson sampling for the multi-armed bandit problem”
Shipra Agrawal and Navin Goyal · 2012
Earlier work this paper cites.
“A Note on Sequence Prediction over Large Alphabets”
Travis Gagie · 2012
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“Attention is all you need”
A Vaswani · 2017
Earlier work this paper cites.
“Neural network trained with supervision represents uncertainty by nonlinear moments”
Li. Wenliang and Maneesh Sahani · 2018
Earlier work this paper cites.
“JAX: composable transformations of Python+NumPy programs”, 2018
James Bradbury et al · 2018
Earlier work this paper cites.
“Meta-learning of sequential strategies”
Pedro Ortega et al · 2019
Earlier work this paper cites.
“Language models are unsupervised multitask learners”
Alec Radford et al · 2019
Earlier work this paper cites.
“Language models are few-shot learners”
Tom Brown et al · 2020
Earlier work this paper cites.
“Meta-trained agents implement bayes-optimal agents”
Vladimir Mikulik et al · 2020
Earlier work this paper cites.
“On the ability and limitations of transformers to recognize formal languages”
Satwik Bhattamishra, Kabir Ahuja and Navin Goyal · 2020
Earlier work this paper cites.
“The DeepMind JAX Ecosystem”, 2020
DeepMind et al · 2020
Earlier work this paper cites.
“Haiku: Sonnet for JAX”, 2020
Tom Hennigan et al · 2020
Earlier work this paper cites.
“Transformers are rnns: Fast autoregressive transformers with linear attention”
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas and François Fleuret · 2020
Earlier work this paper cites.
“Linear attention mechanism: An efficient attention for semantic segmentation”
Rui Li, Jianlin Su, Chenxi Duan and Shunyi Zheng · 2020
Earlier work this paper cites.
“Array programming with NumPy”
Charles. Harris et al · 2020
Earlier work this paper cites.
“An explanation of in-context learning as implicit Bayesian inference”
Sang Xie, Aditi Raghunathan, Percy Liang and Tengyu Ma · 2021
Earlier work this paper cites.
“The power of scale for parameter-efficient prompt tuning”
Brian Lester, Rami Al-Rfou and Noah Constant · 2021
Earlier work this paper cites.
“Why do pretrained language models help in downstream tasks? an analysis of head and prompt tuning”
Colin Wei, Sang Xie and Tengyu Ma · 2021
Earlier work this paper cites.
“SynthBio: A Case Study in Faster Curation of Text Datasets”
Ann Yuan et al · 2021
Earlier work this paper cites.
“Efficient attention: Attention with linear complexities”
Zhuoran Shen et al · 2021
Earlier work this paper cites.
“Fast approximation of Beta inequalities”, 2021
John. Cook · 2021
Earlier work this paper cites.
“Frequency Effects on Syntactic Rule Learning in Transformers”
Jason Wei, Dan Garrette, Tal Linzen and Ellie Pavlick · 2021
Earlier work this paper cites.
“Statistically meaningful approximation: a case study on approximating turing machines with transformers”
Colin Wei, Yining Chen and Tengyu Ma · 2022
Earlier work this paper cites.
“Neural Networks and the Chomsky Hierarchy”
Gregoire Deletang et al · 2022
Earlier work this paper cites.
“Chain-of-thought prompting elicits reasoning in large language models”
Jason Wei et al · 2022
Earlier work this paper cites.
“A Survey on In-context Learning”
Qingxiu Dong et al · 2022
Earlier work this paper cites.
“Do prompt-based models really understand the meaning of their prompts?”
Albert Webson and Ellie Pavlick · 2022
Earlier work this paper cites.
“Discovering the hidden vocabulary of DALLE-2”
Giannis Daras and Alex Dimakis · 2022
Earlier work this paper cites.
“Large language models are zero-shot reasoners”
Takeshi Kojima et al · 2022
Earlier work this paper cites.
“RLPrompt: Optimizing Discrete Text Prompts with Reinforcement Learning”
Mingkai Deng et al · 2022
Earlier work this paper cites.
“Fairdistillation: mitigating stereotyping in language models”
Pieter Delobelle and Bettina Berendt · 2022
Earlier work this paper cites.
“Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents”
Wenlong Huang, Pieter Abbeel, Deepak Pathak and Igor Mordatch · 2022
Cited alongside, same era.
“Rethinking the role of demonstrations: What makes in-context learning work?”
Sewon Min et al · 2022
Cited alongside, same era.
“Exploring length generalization in large language models”
Cem Anil et al · 2022
Cited alongside, same era.
“Impact of Pretraining Term Frequencies on Few-Shot Numerical Reasoning”
Yasaman Razeghi, Robert Logan, Matt Gardner and Sameer Singh · 2022
Cited alongside, same era.
“Transformers generalize differently from information stored in context vs in weights”
Stephanie Chan et al · 2022
Cited alongside, same era.
“Large Language Models Are Implicitly Topic Models: Explaining and Finding Good Demonstrations for In-Context Learning”
“Prompting large language models for recommender systems: A comprehensive framework and empirical analysis”
Lanling Xu et al · 2024
Later among the works it cites.
“Prompt Optimization with EASE? Efficient Ordering-aware Automated Selection of Exemplars”
Zhaoxuan Wu et al · 2024
Later among the works it cites.
“Optimizing prompts for text-to-image generation”
Yaru Hao, Zewen Chi, Li Dong and Furu Wei · 2024
Later among the works it cites.
“LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations”
Anian Ruoss et al · 2024
Later among the works it cites.
“Many-shot in-context learning”
Rishabh Agarwal et al · 2024
Later among the works it cites.
“Does Prompt Formatting Have Any Impact on LLM Performance?”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xinyi Wang et al · 2023
Cited alongside, same era.
“A latent space theory for emergent abilities in large language models”
Hui Jiang · 2023
Cited alongside, same era.
“The learnability of in-context learning”
Noam Wies, Yoav Levine and Amnon Shashua · 2023
Cited alongside, same era.
“Memory-based meta-learning on non-stationary distributions”
Tim Genewein et al · 2023
Cited alongside, same era.
“Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing”
Pengfei Liu et al · 2023
Cited alongside, same era.
“Unleashing the potential of prompt engineering in Large Language Models: a comprehensive review”
Banghao Chen, Zhaofeng Zhang, Nicolas Langrené and Shengxin Zhu · 2023
Cited alongside, same era.
“Prompt engineering in large language models”
Ggaliwango Marvin, Nakayiza Hellen, Daudi Jjingo and Joyce Nakatumba-Nabende · 2023
Cited alongside, same era.
Jia He et al · 2024
Later among the works it cites.
“State of what art? a call for multi-prompt llm evaluation”
Moran Mizrahi et al · 2024
Later among the works it cites.
“Why and when llm-based assistants can go wrong: Investigating the effectiveness of prompt-based interactions for software help-seeking”
Anjali Khurana, Hariharan Subramonyam and Parmit Chilana · 2024
Later among the works it cites.
“Jailbroken: How does llm safety training fail?”
Alexander Wei, Nika Haghtalab and Jacob Steinhardt · 2024
Later among the works it cites.
“A comprehensive study of jailbreak attack versus defense for large language models”
Zihao Xu et al · 2024
Later among the works it cites.
“Talking Nonsense: Probing Large Language Models’ Understanding of Adversarial Gibberish Inputs”
Valeriia Cherepanova and James Zou · 2024
Later among the works it cites.
“AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models”
Xiaogeng Liu, Nan Xu, Muhao Chen and Chaowei Xiao · 2024
Later among the works it cites.
“Learning Universal Predictors”
Jordi Grau-Moya et al · 2024
Later among the works it cites.
“Does learning the right latent variables necessarily improve in-context learning?”
Sarthak Mittal et al · 2024
Later among the works it cites.
“Function Vectors in Large Language Models”
Eric Todd et al · 2024
Later among the works it cites.
“The benefits of a concise chain of thought on problem-solving in large language models”
Matthew Renze and Erhan Guven · 2024
Later among the works it cites.
“Are Longer Prompts Always Better? Prompt Selection in Large Language Models for Recommendation Systems”
Genki Kusano, Kosuke Akimoto and Kunihiro Takeoka · 2024
Later among the works it cites.
“Robustness-aware Automatic Prompt Optimization”
Zeru Shi et al · 2024
Later among the works it cites.
“When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations”
Aleksandar Petrov, Philip Torr and Adel Bibi · 2024
Later among the works it cites.
“Bias Vector: Mitigating Biases in Language Models with Task Arithmetic Approach”
Daiki Shirafuji, Makoto Takenaka and Shinya Taguchi · 2024
Later among the works it cites.
“In-context learning with retrieved demonstrations for language models: A survey”
Man Luo et al · 2024
Later among the works it cites.
“RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models”
Noah Wang et al · 2024
Later among the works it cites.
“In-context Exploration-Exploitation for Reinforcement Learning”
Zhenwen Dai, Federico Tomasi and Sina Ghiassian · 2024
Later among the works it cites.
“GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents”
Anthony Costarelli et al · 2024
Later among the works it cites.
“Do Prompts Really Prompt? Exploring the Prompt Understanding Capability of Whisper”
Chih-Kai Yang, Kuan-Po Huang and Hung-yi Lee · 2024
Later among the works it cites.
“xLSTM: Extended Long Short-Term Memory”
Maximilian Beck et al · 2024
Later among the works it cites.
“Efficient Prompting Methods for Large Language Models: A Survey”
Kaiyan Chang et al · 2024
Later among the works it cites.
“Toward a Theory of Tokenization in LLMs”
Nived Rajaraman, Jiantao Jiao and Kannan Ramchandran · 2024
Later among the works it cites.
“Protein language models meet reduced amino acid alphabets”
Ioan Ieremie, Rob Ewing and Mahesan Niranjan · 2024
Later among the works it cites.
“Compression via pre-trained transformers: A study on byte-level multimodal data”
David Heurtel-Depeiges, Anian Ruoss, Joel Veness and Tim Genewein · 2024
Later among the works it cites.
“Towards Understanding the Characteristics of Code Generation Errors Made by Large Language Models”
Zhijie Wang et al · 2025
Closest in time.
“Effects of prompt length on domain-specific tasks for large language models”
Qibang Liu, Wenzhe Wang and Jeffrey Willard · 2025
Closest in time.
“Towards LLMs Robustness to Changes in Prompt Format Styles”
Lilian Ngweta, Kiran Kate, Jason Tsay and Yara Rizk · 2025
Closest in time.
“DLPO: Towards a Robust, Efficient, and Generalizable Prompt Optimization Framework from a Deep-Learning Perspective”
Dengyun Peng et al · 2025
Closest in time.
“Demystifying optimized prompts in language models”
Rimon Melamed, Lucas McCabe and H Huang · 2025
Closest in time.
“Sweeping heterogeneity with smart mops: Mixture of prompts for LLM task adaptation”
Chen Dun et al · 2025
Closest in time.
“Gemma 3 technical report”
Gemma Team et al · 2025
Closest in time.
“Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities”
Gheorghe Comanici et al · 2025
Closest in time.
“LLMs are Bayesian, in Expectation, not in Realization”
Leon Chlon, Sarah Rashidi, Zein Khamis and MarcAntonio Awada · 2025
Closest in time.