Fetching the paper…
Reading the bibliography…
This paper introduces a novel Bayesian learning model to explain the behavior of Large Language Models (LLMs), focusing on their core optimization metric of next token prediction.
Theory of Majorization and Its Applications
Albert W. Marshall and Ingram Olkin · 1979
Earlier work this paper cites.
On approximating parametric bayes models by nonparametric bayes models
S. R. Dalal and G. J. Hall · 1980
Earlier work this paper cites.
Approximating priors by mixtures of natural conjugate priors
S. R. Dalal and W. J. Hall · 1983
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
On the marginal likelihood and cross-validation
E Fong and C C Holmes · 2020
Cited alongside, same era.
WT5?! Training text-to-text models to explain their predictions
Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan · 2020
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2021
Cited alongside, same era.
Generalization in deep learning
K. Kawaguchi, Y. Bengio, and L. Kaelbling · 2022
Cited alongside, same era.
Can language models learn from explanations in context?
Andrew K. Lampinen, Ishita Dasgupta, Stephanie C.Y. Chan, Kory Matthewson, Michael Henry Tessler, Antonia Creswell, James L. McClelland, Jane X. Wang, and Felix Hill · 2022
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces, 2023
Albert Gu and Tri Dao · 2023
Later among the works it cites.
Symbolic chain-of-thought distillation: Small models can also "think" step-by-step
Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, and Yejin Choi · 2023
Later among the works it cites.
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Later among the works it cites.
Chatgpt: Optimizing language models for dialogue
OpenAI · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2022
Cited alongside, same era.
Lipschitz continuity of probability kernels in the optimal transport framework, 2023
Emanuele Dolera and Edoardo Mainini · 2023
Cited alongside, same era.
Larger language models do in-context learning differently, 2023
Jerry Wei, Jason Wei, Yi Tay, Dustin Tran, Albert Webson, Yifeng Lu, Xinyun Chen, Hanxiao Liu, Da Huang, Denny Zhou, and Tengyu Ma · 2023
Later among the works it cites.
On masked pre-training and the marginal likelihood
Pablo Moreno-Muñoz, Pol G. Recasens, and Søren Hauberg · 2024
Closest in time.