Fetching the paper…
Reading the bibliography…
We demonstrate that LLMs may learn indicators of document usefulness and modulate their updates accordingly.
The mnist database of handwritten digit images for machine learning research [best of the web]
Deng, L · 2012
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
T-rex: A large scale alignment of natural language with knowledge base triples
Elsahar, H., Vougiouklis, P., Remaci, A., Gravier, C., Hare, J., Laforest, F., and Simperl, E · 2018
Earlier work this paper cites.
Learning to generalize: Meta-learning for domain generalization
Li, D., Yang, Y., Song, Y.-Z., and Hospedales, T · 2018
Earlier work this paper cites.
On first-order meta-learning algorithms
Nichol, A., Achiam, J., and Schulman, J · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Earlier work this paper cites.
Stiffness: A new perspective on generalization in neural networks
Fort, S., Nowak, P. K., Jastrzebski, S., and Narayanan, S · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Petroni, F., Rocktäschel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A. H., and Riedel, S · 2019
Earlier work this paper cites.
Lookahead optimizer: k steps forward, 1 step back
Zhang, M., Lucas, J., Ba, J., and Hinton, G. E · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Hidden incentives for auto-induced distributional shift
Krueger, D., Maharaj, T., and Leike, J · 2020
Earlier work this paper cites.
Cheating death in damascus
Levinstein, B. A. and Soares, N · 2020
Earlier work this paper cites.
Learning explanations that are hard to vary
Parascandolo, G., Neitz, A., Orvieto, A., Gresele, L., and Schölkopf, B · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Sinitsin, A., Plokhotnyuk, V., Pyrkin, D., Popov, S., and Babenko, A · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al · 2020
Earlier work this paper cites.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow
Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S · 2021
Cited alongside, same era.
Cross-document coreference resolution over predicted mentions
Cattan, A., Eirew, A., Stanovsky, G., Joshi, M., and Dagan, I · 2021
Cited alongside, same era.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Elazar, Y., Ravfogel, S., Jacovi, A., and Goldberg, Y · 2021
Cited alongside, same era.
Sequential reptile: Inter-task gradient alignment for multilingual learning
Lee, S., Lee, H. B., Lee, J., and Hwang, S. J · 2021
Cited alongside, same era.
Implicit representations of meaning in neural language models
Li, B. Z., Nye, M., and Andreas, J · 2021
Cited alongside, same era.
Implicit gradient alignment in distributed and federated learning
Dandi, Y., Barba, L., and Jaggi, M · 2022
Later among the works it cites.
A cross-verified database of notable people, 3500bc-2018ad
Laouenan, M., Bhargava, P., Eyméoud, J.-B., Gergaud, O., Plique, G., and Wasmer, E · 2022
Later among the works it cites.
A convnet for the 2020s
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S · 2022
Later among the works it cites.
Locating and editing factual knowledge in gpt
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Later among the works it cites.
The alignment problem from a deep learning perspective
Ngo, R., Chan, L., and Mindermann, S · 2022
Later among the works it cites.
Taken out of context: On measuring situational awareness in llms
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P · 2021
Cited alongside, same era.
Mitchell, E., Lin, C., Bosselut, A., Finn, C., and Manning, C. D · 2021
Cited alongside, same era.
Sgd implicitly regularizes generalization error
Roberts, D. A · 2021
Cited alongside, same era.
Gradient matching for domain generalization
Shi, Y., Seely, J., Torr, P. H., Siddharth, N., Hannun, A., Usunier, N., and Synnaeve, G · 2021
Cited alongside, same era.
On the origin of implicit regularization in stochastic gradient descent
Smith, S. L., Dherin, B., Barrett, D. G., and De, S · 2021
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2021
Cited alongside, same era.
Calibrate before use: Improving few-shot performance of language models
Zhao, Z., Wallace, E., Feng, S., Klein, D., and Singh, S · 2021
Cited alongside, same era.
Berglund, L., Stickland, A. C., Balesni, M., Kaufmann, M., Tong, M., Korbak, T., Kokotajlo, D., and Evans, O · 2023
Closest in time.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Closest in time.
Challenges with unsupervised llm knowledge discovery
Farquhar, S., Varma, V., Kenton, Z., Gasteiger, J., Mikulik, V., and Shah, R · 2023
Closest in time.
Memory-based meta-learning on non-stationary distributions
Genewein, T., Delétang, G., Ruoss, A., Wenliang, L. K., Catt, E., Dutordoir, V., Grau-Moya, J., Orseau, L., Hutter, M., and Veness, J · 2023
Closest in time.
Studying large language model generalization with influence functions
Grosse, R., Bae, J., Anil, C., Elhage, N., Tamkin, A., Tajdini, A., Steiner, B., Li, D., Durmus, E., Perez, E., et al · 2023
Closest in time.
Tell, don’t show: Declarative facts influence how llms generalize
Meinke, A. and Evans, O · 2023
Closest in time.
The debate over understanding in ai’s large language models
Mitchell, M. and Krakauer, D. C · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Convnext v2: Co-designing and scaling convnets with masked autoencoders
Woo, S., Debnath, S., Hu, R., Chen, X., Liu, Z., Kweon, I. S., and Xie, S · 2023
Closest in time.
Physics of language models: Part 3.3, knowledge capacity scaling laws
Allen-Zhu, Z. and Li, Y · 2024
Closest in time.
The reversal curse: Llms trained on" a is b" fail to learn" b is a"
Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., and Evans, O · 2024
Closest in time.