Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are remarkably efficient across a wide range of natural language processing tasks and well beyond them.
Finite state Markov chains
Gallager, R. G · 1996
Earlier work this paper cites.
Real analysis: modern techniques and their applications , volume 40
Folland, G. B · 1999
Earlier work this paper cites.
An overview of statistical learning theory
Vapnik, V. N · 1999
Earlier work this paper cites.
Measure concentration for Euclidean distance in the case of dependent random variables
Marton, K · 2004
Earlier work this paper cites.
General state space Markov chains and MCMC algorithms
Roberts, G. O. and Rosenthal, J. S · 2004
Earlier work this paper cites.
Introduction to Nonparametric Estimation
Tsybakov, A. B · 2008
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Hinton, G · 2015
Earlier work this paper cites.
Concentration inequalities for Markov chains by Marton couplings and spectral methods
Paulin, D · 2015
Earlier work this paper cites.
Neural Machine Translation of Rare Words with Subword Units
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
On learning markov chains
Hao, Y., Orlitsky, A., and Pichapati, V · 2018
Earlier work this paper cites.
Learning classifiers with fenchel-young losses: Generalized entropies, margins, and algorithms
Blondel, M., Martins, A., and Niculae, V · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Pytorch: an imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Earlier work this paper cites.
State of the Art of Statistical Learning Theory
Redko, I., Habrard, A., Morvant, E., Sebban, M., and Bennani, Y · 2019
Earlier work this paper cites.
Minimax learning of ergodic Markov chains
Wolfer, G. and Kontorovich, A · 2019
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2019
Earlier work this paper cites.
Flambe: Structural complexity and representation learning of low rank mdps
Agarwal, A., Kakade, S., Krishnamurthy, A., and Sun, W · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., et al · 2020
Earlier work this paper cites.
How much knowledge can you pack into the parameters of a language model?
Roberts, A., Raffel, C., and Shazeer, N · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Cited alongside, same era.
Inductive biases and variable creation in self-attention mechanisms
Edelman, B. L., Goel, S., Kakade, S., and Zhang, C · 2022
Cited alongside, same era.
An Explanation of In-context Learning as Implicit Bayesian Inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Gemini: a family of highly capable multimodal models
Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., and Hauth, A. e. a · 2023
Cited alongside, same era.
The evolution of statistical induction heads: In-context learning markov chains
Edelman, B. L., Edelman, E., Goel, S., Malach, E., and Tsilivis, N · 2024
Closest in time.
Transformers are universal in-context learners
Furuya, T., de Hoop, M. V., and Peyré, G · 2024
Closest in time.
Unveiling the Statistical Foundations of Chain-of-Thought Prompting Methods
Hu, X., Zhang, F., Chen, S., and Yang, Z · 2024
Closest in time.
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
Ildiz, M. E., Huang, Y., Li, Y., Rawat, A. S., and Oymak, S · 2024
Closest in time.
From loops to oops: Fallback behaviors of language models under uncertainty, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Birth of a Transformer: A Memory Viewpoint
Bietti, A., Cabannes, V., Bouchacourt, D., Jegou, H., and Bottou, L · 2023
Cited alongside, same era.
Probing the “creativity” of large language models: Can models produce divergent semantic association?
Chen, H. and Ding, N · 2023
Cited alongside, same era.
xval: A continuous number encoding for large language models
Golkar, S., Pettee, M., Eickenberg, M., Bietti, A., Cranmer, M., Krawezik, G., Lanusse, F., McCabe, M., Ohana, R., Parker, L., et al · 2023
Cited alongside, same era.
Large language models are zero-shot time series forecasters
Gruver, N., Finzi, M., Qiu, S., and Wilson, A. G · 2023
Cited alongside, same era.
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. d. l., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., et al · 2023
Cited alongside, same era.
minGPT: A minimal PyTorch re-implementation of the GPT (Generative Pretrained Transformer)
Karpathy, A · 2023
Cited alongside, same era.
Transformers as algorithms: Generalization and stability in in-context learning
Li, Y., Ildiz, M. E., Papailiopoulos, D., and Oymak, S · 2023
Cited alongside, same era.
Ivgi, M., Yoran, O., Berant, J., and Geva, M · 2024
Closest in time.
An Information-Theoretic Analysis of In-Context Learning
Jeon, H. J., Lee, J. D., Lei, Q., and Van Roy, B · 2024
Closest in time.
Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models
Jurafsky, D. and Martin, J. H · 2024
Closest in time.
Transformers are Minimax Optimal Nonparametric In-Context Learners
Kim, J., Nakamaki, T., and Suzuki, T · 2024
Closest in time.
LLMs learn governing principles of dynamical systems, revealing an in-context neural scaling law
Liu, T. J., Boullé, N., Sarfati, R., and Earls, C. J · 2024
Closest in time.
Unlocking tokens as data points for generalization bounds on larger language models
Lotfi, S., Kuang, Y., Amos, B., Goldblum, M., Finzi, M., and Wilson, A. G · 2024
Closest in time.
Attention with Markov: A framework for principled analysis of transformers via markov chains
Makkuva, A. V., Bondaschi, M., Girish, A., Nagle, A., Jaggi, M., Kim, H., and Gastpar, M · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., and Love, J. e. a · 2024
Closest in time.
Is temperature the creativity parameter of large language models?, 2024
Peeperkorn, M., Kouwenhoven, T., Brown, D., and Jordanous, A · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size, 2024
Riviere, M., Pathak, S., Sessa, P. G., Hardin, C., Bhupatiraju, S., Hussenot, L., Mesnard, T., Shahriari, B., Ramé, A., and et al., J. F · 2024
Closest in time.
Tokenization counts: the impact of tokenization on arithmetic in frontier llms
Singh, A. K. and Strouse, D · 2024
Closest in time.
softmax is not enough (for sharp out-of-distribution), 2024
Veličković, P., Perivolaropoulos, C., Barbero, F., and Pascanu, R · 2024
Closest in time.
The learnability of in-context learning
Wies, N., Levine, Y., and Shashua, A · 2024
Closest in time.
Pre-tokenization of numbers for large language models
Wu, Z., Qi, Q., Zhuang, Z., Sun, H., and Wang, J · 2024
Closest in time.
Mano: Exploiting matrix norm for unsupervised accuracy estimation under distribution shifts, 2024
Xie, R., Odonnat, A., Feofanov, V., Deng, W., Zhang, J., and An, B · 2024
Closest in time.