Fetching the paper…
Reading the bibliography…
In-context learning (ICL) has emerged as a particularly remarkable characteristic of Large Language Models (LLM): given a pretrained LLM and an observed dataset, LLMs can make predictions for new data points from the same distribution without fine-tuning.
Funzione caratteristica di un fenomeno aleatorio
De Finetti, B · 1929
Earlier work this paper cites.
Application of the theory of martingales
Doob, J. L · 1949
Earlier work this paper cites.
Optimal information processing and bayes’s theorem
Zellner, A · 1988
Earlier work this paper cites.
Foundations of modern probability , volume 2
Kallenberg, O · 1997
Earlier work this paper cites.
Asymptotic statistics , volume 3
Van der Vaart, A. W · 2000
Earlier work this paper cites.
Limit theorems for a class of identically distributed random variables
Berti, P., Pratelli, L., and Rigo, P · 2004
Earlier work this paper cites.
Optimal predictions in everyday cognition
Griffiths, T. L. and Tenenbaum, J. B · 2006
Earlier work this paper cites.
Matplotlib: A 2D graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
Data Structures for Statistical Computing in Python
Wes McKinney · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E · 2011
Earlier work this paper cites.
Fundamentals of nonparametric Bayesian inference , volume 44
Ghosal, S. and Van der Vaart, A · 2017
Earlier work this paper cites.
Bayesian fractional posteriors
Bhattacharya, A., Pati, D., and Yang, Y · 2019
Earlier work this paper cites.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Earlier work this paper cites.
Experiment tracking with weights and biases, 2020
Biewald, L · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Array programming with NumPy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., et al · 2020
Earlier work this paper cites.
The Python Library Reference, release 3.8.2
Van Rossum, G · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., Davison, J., Shleifer, S., von Platen, P., Ma, C., Jernite, Y., Plu, J., Xu, C., Scao, T. L., Gugger, S., Drame, M., Lhoest, Q., and Rush, A. M · 2020
Earlier work this paper cites.
Martingale posterior distributions
Fong, E., Holmes, C., and Walker, S. G · 2021
Cited alongside, same era.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P · 2021
Cited alongside, same era.
Capturing semantics for imputation with pre-trained language models
Mei, Y., Song, S., Fang, C., Yang, H., Fang, J., and Long, J · 2021
Cited alongside, same era.
Metaicl: Learning to learn in context
Min, S., Lewis, M., Zettlemoyer, L., and Hajishirzi, H · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2021
A latent space theory for emergent abilities in large language models
Jiang, H · 2023
Later among the works it cites.
Time-llm: Time series forecasting by reprogramming large language models
Jin, M., Wang, S., Ma, L., Chu, Z., Zhang, J. Y., Shi, X., Chen, P.-Y., Liang, Y., Li, Y.-F., Pan, S., et al · 2023
Later among the works it cites.
Calibrated language models must hallucinate
Kalai, A. T. and Vempala, S. S · 2023
Later among the works it cites.
Li, Z., Zhu, H., Lu, Z., and Yin, M · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Zhao, Z., Wallace, E., Feng, S., Klein, D., and Singh, S · 2021
Cited alongside, same era.
What learning algorithm is in-context learning? investigations with linear models
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D · 2022
Cited alongside, same era.
Language models are realistic tabular data generators
Borisov, V., Seßler, K., Leemann, T., Pawelczyk, M., and Kasneci, G · 2022
Cited alongside, same era.
A survey for in-context learning
Dong, Q., Li, L., Dai, D., Zheng, C., Wu, Z., Chang, B., Sun, X., Xu, J., and Sui, Z · 2022
Cited alongside, same era.
Implicit bayesian inference in large language models
Huszár, F · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Manikandan, H., Jiang, Y., and Kolter, J. Z · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Pretraining task diversity and the emergence of non-bayesian in-context learning for regression
Raventós, A., Paul, M., Chen, F., and Ganguli, S · 2023
Later among the works it cites.
Model dementia: Generated data makes models forget
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., and Anderson, R · 2023
Later among the works it cites.
The transient nature of emergent in-context learning in transformers
Singh, A. K., Chan, S. C., Moskovitz, T., Grant, E., Saxe, A. M., and Hill, F · 2023
Later among the works it cites.
Does synthetic data generation of llms help clinical text mining?
Tang, R., Han, X., Jiang, X., and Hu, X · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Veselovsky, V., Ribeiro, M. H., Arora, A., Josifoski, M., Anderson, A., and West, R · 2023
Later among the works it cites.
Large language models are latent variable models: Explaining and finding good demonstrations for in-context learning
Wang, X., Zhu, W., Saxon, M., Steyvers, M., and Wang, W. Y · 2023
Later among the works it cites.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al · 2023
Later among the works it cites.
In-context learning through the bayesian prism
Panwar, M., Ahuja, K., and Goyal, N · 2024
Closest in time.
Making pre-trained language models great on tabular prediction
Yan, J., Zheng, B., Xu, H., Zhu, Y., Chen, D., Sun, J., Wu, J., and Chen, J · 2024
Closest in time.
Pre-training and in-context learning IS bayesian inference a la de finetti
Ye, N., Yang, H., Siah, A., and Namkoong, H · 2024
Closest in time.