Fetching the paper…
Reading the bibliography…
Intelligent agents must be able to articulate its own uncertainty.
Classi di numeri aleatori equivalenti. la legge dei grandi numeri nel caso dei numeri aleatori equivalenti. sulla legge di distribuzione dei valori in una successione di numeri aleatori equivalenti
B. De Finetti · 1933
Earlier work this paper cites.
La prévision: ses lois logiques, ses sources subjectives
B. De Finetti · 1937
Earlier work this paper cites.
Symmetric measures on cartesian products
E. Hewitt and L. J. Savage · 1955
Earlier work this paper cites.
Probabilistic prediction
H. V. Roberts · 1965
Earlier work this paper cites.
Subjective bayesian models in sampling finite populations
W. A. Ericson · 1969
Earlier work this paper cites.
Information-theoretic asymptotics of bayes methods
B. S. Clarke and A. R. Barron · 1990
Earlier work this paper cites.
Prequential analysis, stochastic complexity and Bayesian inference
A. P. Dawid · 1992
Earlier work this paper cites.
An Introduction to the Bootstrap
B. Efron and R. J. Tibshirani · 1993
Earlier work this paper cites.
De Finetti’s contribution to probability and statistics
D. M. Cifarelli and E. Regazzini · 1996
Earlier work this paper cites.
Prequential probability: principles and properties
A. Dawid and V. Vovk · 1997
Earlier work this paper cites.
Well calibrated, coherent forecasting systems
P. Berti, E. Regazzini, and P. Rigo · 1998
Earlier work this paper cites.
Strong laws for martingale differences and independent random variables
H. Teicher · 1998
Earlier work this paper cites.
Asymptotic Statistics
A. W. van der Vaart · 1998
Earlier work this paper cites.
Exchangeability, predictive distributions and parametric models
S. Fortini, L. Ladelli, and E. Regazzini · 2000
Earlier work this paper cites.
Latent dirichlet allocation
D. Blei, A. Ng, and M. Jordan · 2003
Earlier work this paper cites.
Limit theorems for a class of identically distributed random variables
P. Berti, L. Pratelli, and P. Rigo · 2004
Earlier work this paper cites.
The bernstein-von-mises theorem under misspecification
B. Kleijn and A. Van der Vaart · 2012
Earlier work this paper cites.
Bayesian Data Analysis
A. Gelman, J. Carlin, H. Stern, D. Dunson, A. Vehtari, and D. Rubin · 2013
Cited alongside, same era.
Predictive distribution (de f inetti’s view)
S. Fortini and S. Petrone · 2014
Cited alongside, same era.
Weight uncertainty in neural network
C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra · 2015
Cited alongside, same era.
Theory of probability: A critical introductory treatment , volume 6
B. De Finetti · 2017
Cited alongside, same era.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Cited alongside, same era.
Conditional neural processes
M. Garnelo, D. Rosenbaum, C. Maddison, T. Ramalho, D. Saxton, M. Shanahan, Y. W. Teh, D. Rezende, and S. A. Eslami · 2018
Cited alongside, same era.
Transformers can do bayesian inference
S. Müller, N. Hollmann, S. P. Arango, J. Grabocka, and F. Hutter · 2022
Later among the works it cites.
Transformer neural processes: Uncertainty-aware meta learning via sequence modeling
T. Nguyen and A. Grover · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, b. ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou · 2022
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models
E. Akyürek, D. Schuurmans, J. Andreas, T. Ma, and D. Zhou · 2023
Later among the works it cites.
Transformers as statisticians: Provable in-context learning with in-context algorithm selection
Y. Bai, F. Chen, H. Wang, C. Xiong, and S. Mei · 2023
Later among the works it cites.
Why can GPT learn in-context? language models secretly perform gradient descent as meta-optimizers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On recursive bayesian predictive distributions
P. R. Hahn, R. Martin, and S. G. Walker · 2018
Cited alongside, same era.
Randomized prior functions for deep reinforcement learning
I. Osband, J. Aslanides, and A. Cassirer · 2018
Cited alongside, same era.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, and A. Askell · 2020
Cited alongside, same era.
A survey on bayesian deep learning
H. Wang and D.-Y. Yeung · 2020
Cited alongside, same era.
A class of models for bayesian predictive inference
P. Berti, E. Dreassi, L. Pratelli, and P. Rigo · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
S. M. Xie, A. Raghunathan, P. Liang, and T. Ma · 2021
Cited alongside, same era.
D. Dai, Y. Sun, L. Dong, Y. Hao, S. Ma, Z. Sui, and F. Wei · 2023
Later among the works it cites.
Martingale posterior distributions
E. Fong, C. Holmes, and S. G. Walker · 2023
Later among the works it cites.
Tabpfn: A transformer that solves small tabular classification problems in a second
N. Hollmann, S. Müller, K. Eggensperger, and F. Hutter · 2023
Later among the works it cites.
Transformers as algorithms: Generalization and stability in in-context learning
Y. Li, M. E. Ildiz, D. Papailiopoulos, and S. Oymak · 2023
Later among the works it cites.
K. Tian, E. Mitchell, A. Zhou, A. Sharma, R. Rafailov, H. Yao, C. Finn, and C. D. Manning · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
J. Von Oswald, E. Niklasson, E. Randazzo, J. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov · 2023
Later among the works it cites.
Is in-context learning in large language models bayesian? a martingale perspective, 2024
F. Falck, Z. Wang, and C. Holmes · 2024
Closest in time.
An information-theoretic analysis of in-context learning
H. J. Jeon, J. D. Lee, Q. Lei, and B. Van Roy · 2024
Closest in time.
Calibrating large language models with sample consistency
Q. Lyu, K. Shridhar, C. Malaviya, L. Zhang, Y. Elazar, N. Tandon, M. Apidianaki, M. Sachan, and C. Callison-Burch · 2024
Closest in time.
How many pretraining tasks are needed for in-context learning of linear regression?
J. Wu, D. Zou, Z. Chen, V. Braverman, Q. Gu, and P. Bartlett · 2024
Closest in time.
Posterior sampling via autoregressive generation
K. W. Zhang, H. Namkoong, D. Russo, et al · 2024
Closest in time.