Fetching the paper…
Reading the bibliography…
Transformer based large-language models (LLMs) display extreme proficiency with language yet a precise understanding of how they work remains elusive.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 1901
Earlier work this paper cites.
The connectionist scientist game: Rule extraction and refinement in a neural network
Mcmillan, C., Mozer, M., and Smolensky, P. (1991) · 1991
Earlier work this paper cites.
Distributional clustering of English words
Pereira, F., Tishby, N., and Lee, L. (1993) · 1993
Earlier work this paper cites.
Improved backing-off for m-gram language modeling
Kneser, R. and Ney, H. (1995) · 1995
Earlier work this paper cites.
A bit of progress in language modeling
Goodman, J. (2001) · 2001
Earlier work this paper cites.
Rule extraction from recurrent neural networks: Ataxonomy and review
Jacobsson, H. (2005) · 2005
Earlier work this paper cites.
Large language models in machine translation
Brants, T., Popat, A. C., Xu, P., Och, F. J., and Dean, J. (2007) · 2007
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2017) · 2017
Earlier work this paper cites.
The grammar-learning trajectories of neural language models
Choshen, L., Hacohen, G., Weinshall, D., and Abend, O. (2022) · 2022
Earlier work this paper cites.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C. (2022) · 2022
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P.-S., Glaese, A., Welbl, J., Dathathri, S., Huang, S., Uesato, J., Mellor, J., Higgins, I., Creswell, A., McAleese, N., Wu, A., Elsen, E., Jayakumar, S., Buchatskaya, E., Budden, D., Sutherland, E., Simonyan, K., Paganini, M., Sifre, L., Martens, L., Li, X. L., Kuncoro, A., Nematzadeh, A., Gribovskaya, E., Donato, D., Lazaridou, A., Mensch, A., Lespiau, J.-B., Tsimpoukelli, M., Grigorev, N., Fritz, D., Sottiaux, T., Pajarskas, M., Pohlen, T., Gong, Z., Toyama, D., de Masson d’Autume, C., Li, Y., Terzi, T., Mikulik, V., Babuschkin, I., Clark, A., de Las Casas, D., Guy, A., Jones, C., Bradbury, J., Johnson, M., Hechtman, B., Weidinger, L., Gabriel, I., Isaac, W., Lockhart, E., Osindero, S., Rimell, L., Dyer, C., Vinyals, O., Ayoub, K., Stanway, J., Bennett, L., Hassabis, D., Kavukcuoglu, K., and Irving, G. (2022) · 2022
Cited alongside, same era.
Impact of pretraining term frequencies on few-shot numerical reasoning
Razeghi, Y., RobertL.Logan, I., Gardner, M., and Singh, S. (2022) · 2022
Cited alongside, same era.
Zoology: Measuring and improving recall in efficient language models
Arora, S., Eyuboglu, S., Timalsina, A., Johnson, I., Poli, M., Zou, J., Rudra, A., and Ré, C. (2023) · 2023
Cited alongside, same era.
Quantifying memorization across neural language models
Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramèr, F., and Zhang, C. (2023) · 2023
Neurons in large language models: Dead, n-gram, positional
Voita, E., Ferrando, J., and Nalmpantis, C. (2023) · 2023
Later among the works it cites.
In-context language learning: Architectures and algorithms
Akyürek, E., Wang, B., Kim, Y., and Andreas, J. (2024) · 2024
Closest in time.
The reversal curse: Llms trained on "a is b" fail to learn "b is a"
Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., and Evans, O. (2024) · 2024
Closest in time.
Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs
Chen, A., Shwartz-Ziv, R., Cho, K., Leavitt, M. L., and Saphra, N. (2024) · 2024
Closest in time.
The evolution of statistical induction heads: In-context learning markov chains
Edelman, B. L., Edelman, E., Goel, S., Malach, E., and Tsilivis, N. (2024) · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Measuring causal effects of data statistics on language model’s ‘factual’ predictions
Elazar, Y., Kassner, N., Ravfogel, S., Feder, A., Ravichander, A., Mosbach, M., Belinkov, Y., Schütze, H., and Goldberg, Y. (2023) · 2023
Cited alongside, same era.
Tinystories: How small can language models be and still speak coherent english?
Eldan, R. and Li, Y. (2023) · 2023
Cited alongside, same era.
Large language models struggle to learn long-tail knowledge
Kandpal, N., Deng, H., Roberts, A., Wallace, E., and Raffel, C. (2023) · 2023
Cited alongside, same era.
Impact of co-occurrence on factual knowledge of large language models
Kang, C. and Choi, J. (2023) · 2023
Cited alongside, same era.
Scalable extraction of training data from (production) language models
Nasr, M., Carlini, N., Hayase, J., Jagielski, M., Cooper, A. F., Ippolito, D., Choquette-Choo, C. A., Wallace, E., Tramèr, F., and Lee, K. (2023) · 2023
Cited alongside, same era.
Bias and fairness in large language models: A survey
Gallegos, I. O., Rossi, R. A., Barrow, J., Tanjim, M. M., Kim, S., Dernoncourt, F., Yu, T., Zhang, R., and Ahmed, N. K. (2024) · 2024
Closest in time.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J. W., and Sifre, L. (2024) · 2024
Closest in time.
Infini-gram: Scaling unbounded n-gram language models to a trillion tokens
Liu, J., Min, S., Zettlemoyer, L., Choi, Y., and Hajishirzi, H. (2024) · 2024
Closest in time.
Transformers can represent n n -gram language models
Svete, A. and Cotterell, R. (2024) · 2024
Closest in time.