Fetching the paper…
Reading the bibliography…
Chain-of-Thought (CoT) prompting and its variants have gained popularity as effective methods for solving multi-step reasoning problems using pretrained large language models (LLMs).
Language models are few-shot learners
Brown, T · 1901
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R · 1905
Earlier work this paper cites.
Joint measures and cross-covariance operators
Baker, C. R · 1973
Earlier work this paper cites.
An introduction to hidden markov models
Rabiner, L · 1986
Earlier work this paper cites.
Some pac-bayesian theorems
McAllester, D. A · 1998
Earlier work this paper cites.
Bayesian model averaging: a tutorial (with comments by m. clyde, david draper and ei george, and a rejoinder by the authors
Hoeting, J. A · 1999
Earlier work this paper cites.
Information theory, inference and learning algorithms
MacKay, D. J · 2003
Earlier work this paper cites.
Optimal rates for the regularized least-squares algorithm
Caponnetto, A · 2007
Earlier work this paper cites.
The elements of statistical learning: data mining, inference, and prediction
Hastie, T · 2009
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
Erhan, D · 2010
Earlier work this paper cites.
Wang, Y.-A · 2010
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P · 2014
Earlier work this paper cites.
Ba, J. L · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Deep sets
Zaheer, M · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J · 2018
Earlier work this paper cites.
Using pre-training can improve model robustness and uncertainty
Hendrycks, D · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A · 2019
Earlier work this paper cites.
On layer normalization in the transformer architecture
Xiong, R · 2020
Earlier work this paper cites.
Rethinking pre-training and self-training
Zoph, B · 2020
Earlier work this paper cites.
Unified pre-training for program understanding and generation
Ahmad, W. U · 2021
Earlier work this paper cites.
User-friendly introduction to pac-bayes bounds
Alquier, P · 2021
Earlier work this paper cites.
Deep neural network approximation theory
Elbrächter, D · 2021
Earlier work this paper cites.
The statistical complexity of interactive decision making
Foster, D. J · 2021
Earlier work this paper cites.
Learning to retrieve prompts for in-context learning
Rubin, O · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M · 2021
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Akyürek, E · 2022
Cited alongside, same era.
Chen, W · 2022
Cited alongside, same era.
Selection-inference: Exploiting large language models for interpretable logical reasoning
Towards revealing the mystery behind chain of thought: a theoretical perspective
Feng, G · 2023
Later among the works it cites.
Fu, D · 2023
Later among the works it cites.
A theory of emergent in-context learning as implicit structure induction
Hahn, M · 2023
Later among the works it cites.
Towards a mechanistic interpretation of multi-step reasoning capabilities of language models
Hou, Y · 2023
Later among the works it cites.
A latent space theory for emergent abilities in large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Creswell, A · 2022
Cited alongside, same era.
A survey on in-context learning
Dong, Q · 2022
Cited alongside, same era.
Inductive biases and variable creation in self-attention mechanisms
Edelman, B. L · 2022
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes
Garg, S · 2022
Cited alongside, same era.
Kim, H. J · 2022
Cited alongside, same era.
Text and patterns: For effective chain of thought, it takes two to tango
Madaan, A · 2022
Cited alongside, same era.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S · 2022
Cited alongside, same era.
In-context learning and induction heads
Olsson, C · 2022
Cited alongside, same era.
Jiang, H · 2023
Later among the works it cites.
Measuring faithfulness in chain-of-thought reasoning
Lanham, T · 2023
Later among the works it cites.
Mahankali, A · 2023
Later among the works it cites.
The expresssive power of transformers with chain of thought
Merrill, W · 2023
Later among the works it cites.
Gpt-4 technical report. arxiv 2303.08774
OpenAI, R · 2023
Later among the works it cites.
Refiner: Reasoning feedback on intermediate representations
Paul, D · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
Su, J · 2023
Later among the works it cites.
Large language models are in-context semantic reasoners rather than symbolic reasoners
Tang, X · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, G · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H · 2023
Later among the works it cites.
Why can large language models generate correct chain-of-thoughts?
Tutunov, R · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
Von Oswald, J · 2023
Later among the works it cites.
Towards understanding chain-of-thought prompting: An empirical study of what matters. arxiv 2023
Wang, B · 2023
Later among the works it cites.
The learnability of in-context learning
Wies, N · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S · 2023
Later among the works it cites.
Graph of thoughts: Solving elaborate problems with large language models
Besta, M · 2024
Closest in time.
Faith and fate: Limits of transformers on compositionality
Dziri, N · 2024
Closest in time.
From words to actions: Unveiling the theoretical underpinnings of llm-driven autonomous systems
He, J · 2024
Closest in time.
Can large language models explore in-context?
Krishnamurthy, A · 2024
Closest in time.
Why think step by step? reasoning emerges from the locality of experience
Prystawski, B · 2024
Closest in time.
A systematic survey of prompt engineering in large language models: Techniques and applications
Sahoo, P · 2024
Closest in time.
A comprehensive survey of hallucination mitigation techniques in large language models
Tonmoy, S · 2024
Closest in time.