Fetching the paper…
Reading the bibliography…
Large language models (LLMs) with billions of parameters excel at predicting the next token in a sequence.
On tables of random numbers
A. N. Kolmogorov · 1963
Earlier work this paper cites.
A formal theory of inductive inference. part i
R. J. Solomonoff · 1964
Earlier work this paper cites.
Weighted sums of certain dependent random variables
K. Azuma · 1967
Earlier work this paper cites.
An introduction to arithmetic coding
G. G. Langdon · 1984
Earlier work this paper cites.
Rfc1952: Gzip file format specification version 4.3, 1996
P. Deutsch · 1996
Earlier work this paper cites.
Markov chains
J. R. Norris · 1998
Earlier work this paper cites.
Pac-bayesian supervised classification: the thermodynamics of statistical learning
O. Catoni · 2007
Earlier work this paper cites.
Stability bounds for non-iid processes
M. Mohri and A. Rostamizadeh · 2007
Earlier work this paper cites.
Pearson correlation coefficient
J. Benesty, J. Chen, Y. Huang, and I. Cohen · 2009
Earlier work this paper cites.
Chromatic pac-bayes bounds for non-iid data: Applications to ranking and stationary β \beta -mixing processes
L. Ralaivola, M. Szafranski, and G. Stempfel · 2010
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Earlier work this paper cites.
G. K. Dziugaite and D. M. Roy · 2017
Earlier work this paper cites.
Generalization bounds for non-stationary mixing processes
V. Kuznetsov and M. Mohri · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
On equivalence of martingale tail bounds and deterministic regret inequalities
A. Rakhlin and K. Sridharan · 2017
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Non-vacuous generalization bounds at the imagenet scale: a pac-bayesian compression approach
W. Zhou, V. Veitch, M. Austern, R. P. Adams, and P. Orbanz · 2019
Earlier work this paper cites.
On the role of data in pac-bayes bounds
G. K. Dziugaite, K. Hsu, W. Gharbieh, G. Arpino, and D. Roy · 2021
Cited alongside, same era.
Probabilistic fine-tuning of pruning masks and pac-bayes self-bounded learning
S. Hayou, B. He, and G. K. Dziugaite · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Cited alongside, same era.
Tighter risk certificates for neural networks
M. Pérez-Ortiz, O. Rivasplata, J. Shawe-Taylor, and C. Szepesvári · 2021
Cited alongside, same era.
Generative language modeling for antibody design
R. W. Shuai, J. A. Ruffolo, and J. J. Gray · 2021
Cited alongside, same era.
Redpajama: an open dataset for training large language models, 2023
T. Computer · 2023
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
T. Dettmers, S. Shmitchell, A. Roberts, K. Lee, T. B. Brown, D. Song, and C. Raffel · 2023
Later among the works it cites.
M. Goldblum, M. Finzi, K. Rowan, and A. G. Wilson · 2023
Later among the works it cites.
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Later among the works it cites.
The cost of down-scaling language models: Fact recall deteriorates before in-context learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2021
Cited alongside, same era.
Monarch: Expressive structured matrices for efficient and accurate training
T. Dao, B. Chen, N. S. Sohoni, A. Desai, M. Poli, J. Grogan, A. Liu, A. Rao, A. Rudra, and C. Ré · 2022
Cited alongside, same era.
8-bit optimizers via block-wise quantization, 2022
T. Dettmers, M. Lewis, S. Shleifer, and L. Zettlemoyer · 2022
Cited alongside, same era.
Krona: Parameter efficient tuning with kronecker adapter
A. Edalati, M. Tahaei, I. Kobyzev, V. P. Nia, J. J. Clark, and M. Rezagholizadeh · 2022
Cited alongside, same era.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Z. Frantal, A. Gruslys, and D. Kiela · 2022
Cited alongside, same era.
Pac-bayes compression bounds so tight that they can explain generalization
S. Lotfi, M. Finzi, S. Kapoor, A. Potapczynski, M. Goldblum, and A. G. Wilson · 2022
Cited alongside, same era.
Observed antibody space: A diverse database of cleaned, annotated, and translated unpaired and paired antibody sequences
T. H. Olsen, F. Boyles, and C. M. Deane · 2022
Cited alongside, same era.
T. Jin, N. Clement, X. Dong, V. Nagarajan, M. Carbin, J. Ragan-Kelley, and G. K. Dziugaite · 2023
Later among the works it cites.
Memory-efficient fine-tuning of compressed large language models via sub-4-bit integer quantization
J. Kim, J. H. Lee, S. Kim, J. Park, K. M. Yoo, S. J. Kwon, and D. Lee · 2023
Later among the works it cites.
Non-vacuous generalization bounds for large language models
S. Lotfi, M. Finzi, Y. Kuang, T. G. Rudner, M. Goldblum, and A. G. Wilson · 2023
Later among the works it cites.
The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only, 2023
G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei, and J. Launay · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Later among the works it cites.
Fast clonal family inference from large-scale B cell repertoire sequencing data
K. Wang, X. Hu, and J. Zhang · 2023
Later among the works it cites.
Q. Xu, W. Xu, and J. Zhu · 2023
Later among the works it cites.
Bayesian optimization of antibodies informed by a generative model of evolving sequences
Anonymous · 2024
Closest in time.
Quip: 2-bit quantization of large language models with guarantees, 2024
J. Chee, Y. Cai, V. Kuleshov, and C. D. Sa · 2024
Closest in time.
Quip#: Quip with lattice codebooks
C. RelaxML · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu · 2024
Closest in time.
Quip#: Even better llm quantization with hadamard incoherence and lattice codebooks
A. Tseng, J. Chee, Q. Sun, V. Kuleshov, and C. De Sa · 2024
Closest in time.