Fetching the paper…
Reading the bibliography…
Modern language models can contain billions of parameters, raising the question of whether they can generalize beyond the training data or simply parrot their training corpora.
On tables of random numbers
Kolmogorov, A. N · 1963
Earlier work this paper cites.
A formal theory of inductive inference. part i
Solomonoff, R. J · 1964
Earlier work this paper cites.
An introduction to arithmetic coding
Langdon, G. G · 1984
Earlier work this paper cites.
Principles of risk minimization for learning theory
Vapnik, V · 1991
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Hoeffding, W · 1994
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2002
Earlier work this paper cites.
Pac-bayesian supervised classification: the thermodynamics of statistical learning
Catoni, O · 2007
Earlier work this paper cites.
Generalization error bounds for stationary autoregressive models
McDonald, D. J., Shalizi, C. R., and Schervish, M · 2011
Earlier work this paper cites.
A PAC-Bayesian approach to minimum perplexity language modeling
Bharadwaj, S. and Hasegawa-Johnson, M · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Earlier work this paper cites.
On the properties of variational approximations of gibbs posteriors
Alquier, P., Ridgway, J., and Chopin, N · 2016
Earlier work this paper cites.
Pac-bayesian theory meets bayesian inference
Germain, P., Bach, F., Lacoste, A., and Lacoste-Julien, S · 2016
Earlier work this paper cites.
Dziugaite, G. K. and Roy, D. M · 2017
Earlier work this paper cites.
Implicit regularization in matrix factorization
Gunasekar, S., Woodworth, B. E., Bhojanapalli, S., Neyshabur, B., and Srebro, N · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Cited alongside, same era.
Measuring the intrinsic dimension of objective landscapes
Li, C., Farkhoor, H., Liu, R., and Yosinski, J · 2018
Cited alongside, same era.
Pac-bayes under potentially heavy tails
Holland, M · 2019
Cited alongside, same era.
Efron-stein pac-bayesian inequalities
Kuzborskij, I. and Szepesvári, C · 2019
Cited alongside, same era.
Uniform convergence may be unable to explain generalization in deep learning
Nagarajan, V. and Kolter, J. Z · 2019
Cited alongside, same era.
Understanding deep learning (still) requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways, 2022
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
Later among the works it cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers
Frantal, Z., Gruslys, A., and Kiela, D · 2022
Later among the works it cites.
Pac-bayes compression bounds so tight that they can explain generalization
Lotfi, S., Finzi, M., Kapoor, S., Potapczynski, A., Goldblum, M., and Wilson, A. G · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Glue: A multi-task benchmark and analysis platform for natural language understanding, 2019
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Cited alongside, same era.
Non-vacuous generalization bounds at the imagenet scale: a pac-bayesian compression approach
Zhou, W., Veitch, V., Austern, M., Adams, R. P., and Orbanz, P · 2019
Cited alongside, same era.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Aghajanyan, A., Zettlemoyer, L., and Gupta, S · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Extracting training data from large language models
Carlini, N., Tramèr, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T. B., Song, D., Erlingsson, U., Oprea, A., and Raffel, C · 2020
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Pac-bayes unleashed: Generalisation bounds with unbounded losses
Haddouche, M., Guedj, B., Rivasplata, O., and Shawe-Taylor, J · 2021
Cited alongside, same era.
Park, G., Kim, J., Kim, J., Choi, E., Kim, S., Kim, S., Lee, M., Shin, H., and Lee, J · 2022
Later among the works it cites.
Causal forecasting: generalization bounds for autoregressive models
Vankadara, L. C., Faller, P. M., Hardt, M., Minorics, L., Ghoshdastidar, D., and Janzing, D · 2022
Later among the works it cites.
Quantifying memorization across neural language models
Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., and Zhang, C · 2023
Closest in time.
Language modeling is compression
Delétang, G., Ruoss, A., Duquenne, P.-A., Catt, E., Genewein, T., Mattern, C., Grau-Moya, J., Wenliang, L. K., Aitchison, M., Orseau, L., et al · 2023
Closest in time.
Goldblum, M., Finzi, M., Rowan, K., and Wilson, A. G · 2023
Closest in time.
Tabllm: Few-shot classification of tabular data with large language models
Hegselmann, S., Buendia, A., Lang, H., Agrawal, M., Jiang, X., and Sontag, D · 2023
Closest in time.
Memory-efficient fine-tuning of compressed large language models via sub-4-bit integer quantization
Kim, J., Lee, J. H., Kim, S., Park, J., Yoo, K. M., Kwon, S. J., and Lee, D · 2023
Closest in time.
Llm-qat: Data-free quantization aware training for large language models
Liu, Y., Xu, Q., Xu, W., and Zhu, J · 2023
Closest in time.
Xu, Q., Xu, W., and Zhu, J · 2023
Closest in time.