Fetching the paper…
Reading the bibliography…
Large language models are pre-trained on ever-growing token budgets under the assumption that better pre-training performance translates to improved downstream models.
On tiny episodic memories in continual learning, 2019b
Chaudhry, A., Rohrbach, M., Elhoseiny, M., Ajanthan, T., Dokania, P. K., Torr, P. H. S., and Ranzato, M · 1902
Earlier work this paper cites.
Implicit regularization of discrete gradient dynamics in linear neural networks, 2019
Gidel, G., Bach, F., and Lacoste-Julien, S · 1904
Earlier work this paper cites.
Uncertainty-based continual learning with adaptive regularization, 2019
Ahn, H., Cha, S., Lee, D., and Moon, T · 1905
Earlier work this paper cites.
Episodic memory in lifelong language learning, 2019
de Masson d’Autume, C., Ruder, S., Kong, L., and Yogatama, D · 1906
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
French, R. M · 1999
Earlier work this paper cites.
Building a question answering test collection
Voorhees, E. M. and Tice, D. M · 2000
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts
Pang, B. and Lee, L · 2004
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Dagan, I., Glickman, O., and Magnini, B · 2005
Earlier work this paper cites.
Understanding and improving information transfer in multi-task learning, 2020
Wu, S., Zhang, H. R., and Ré, C · 2005
Earlier work this paper cites.
On the theory of transfer learning: The importance of task diversity, 2020
Tripuraneni, N., Jordan, M. I., and Jin, C · 2006
Earlier work this paper cites.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Goodfellow, I. J., Mirza, M., Xiao, D., Courville, A., and Bengio, Y · 2013
Earlier work this paper cites.
A diagram is worth a dozen images, 2016
Kembhavi, A., Salvato, M., Kolve, E., Seo, M., Hajishirzi, H., and Farhadi, A · 2016
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., and Hadsell, R · 2017
Earlier work this paper cites.
icarl: Incremental classifier and representation learning
Rebuffi, S.-A., Kolesnikov, A., Sperl, G., and Lampert, C. H · 2017
Earlier work this paper cites.
Continual learning with deep generative replay, 2017
Shin, H., Lee, J. K., Kim, J., and Kim, J · 2017
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy, 2018
Arora, S., Li, Y., Liang, Y., Ma, T., and Risteski, A · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Senteval: An evaluation toolkit for universal sentence representations
Conneau, A. and Kiela, D · 2018
Earlier work this paper cites.
Learning to learn around a common mean
Denevi, G., Ciliberto, C., Stamos, D., and Pontil, M · 2018
Earlier work this paper cites.
Lifelong learning via progressive distillation and retrospection
Hou, S., Pan, X., Loy, C. C., Wang, Z., and Lin, D · 2018
Earlier work this paper cites.
Measuring catastrophic forgetting in neural networks
Kemker, R., McClure, M., Abitino, A., Hayes, T., and Kanan, C · 2018
Earlier work this paper cites.
A mathematical theory of semantic development in deep neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2018
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science , volume 47
Vershynin, R · 2018
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A. and Manning, C. D · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Earlier work this paper cites.
Socialiqa: Commonsense reasoning about social interactions
Sap, M., Rashkin, H., Chen, D., LeBras, R., and Choi, Y · 2019
Earlier work this paper cites.
Towards vqa models that can read
Singh, A., Natarjan, V., Shah, M., Jiang, Y., Chen, X., Parikh, D., and Rohrbach, M · 2019
Earlier work this paper cites.
On warm-starting neural network training
Ash, J. and Adams, R. P · 2020
Earlier work this paper cites.
Rewriting a deep generative model
Bau, D., Liu, S., Wang, T., Zhu, J.-Y., and Torralba, A · 2020
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2020
Cited alongside, same era.
Implicit under-parameterization inhibits data-efficient deep reinforcement learning
Kumar, A., Agarwal, R., Ghosh, D., and Levine, S · 2020
Cited alongside, same era.
Tweet sentiment extraction
Maggie, Culliton, P., and Chen, W · 2020
Cited alongside, same era.
How fine-tuning allows for effective meta-learning, 2021
Chua, K., Lei, Q., and Lee, J. D · 2021
Cited alongside, same era.
Understanding catastrophic forgetting in language models via implicit inference
Kotha, S., Springer, J. M., and Raghunathan, A · 2023
Later among the works it cites.
Maintaining plasticity via regenerative regularization
Kumar, S., Marklund, H., and Van Roy, B · 2023
Later among the works it cites.
Directions of curvature as an explanation for loss of plasticity
Lewandowski, A., Tanaka, H., Schuurmans, D., and Machado, M. C · 2023
Later among the works it cites.
Understanding plasticity in neural networks
Lyle, C., Zheng, Z., Nikishin, E., Pires, B. A., Pascanu, R., and Dabney, W · 2023
Later among the works it cites.
Revisiting plasticity in visual reinforcement learning: Data, modules and training stages
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
Cohen, J. M., Kaur, S., Li, Y., Kolter, J. Z., and Talwalkar, A · 2021
Cited alongside, same era.
Continual backprop: Stochastic gradient descent with persistent randomness
Dohare, S., Sutton, R. S., and Mahmood, A. R · 2021
Cited alongside, same era.
Scaling laws for transfer, 2021
Hernandez, D., Kaplan, J., Henighan, T., and McCandlish, S · 2021
Cited alongside, same era.
Winogrande: An adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Cited alongside, same era.
A mathematical exploration of why language models help solve downstream tasks
Saunshi, N., Malladi, S., and Arora, S · 2021
Cited alongside, same era.
A theoretical analysis of fine-tuning with linear teachers, 2021
Shachaf, G., Brutzkus, A., and Globerson, A · 2021
Cited alongside, same era.
Ma, G., Li, L., Zhang, S., Liu, Z., Wang, Z., Chen, Y., Shen, L., Wang, X., and Tao, D · 2023
Later among the works it cites.
Inverse scaling prize: Second round winners, 2023
McKenzie, I., Lyzhov, A., Parrish, A., Prabhu, A., Mueller, A., Kim, N., Bowman, S., and Perez, E · 2023
Later among the works it cites.
Scaling data-constrained language models, 2023
Muennighoff, N., Rush, A. M., Barak, B., Scao, T. L., Piktus, A., Tazi, N., Pyysalo, S., Wolf, T., and Raffel, C · 2023
Later among the works it cites.
Are emergent abilities of large language models a mirage?
Schaeffer, R., Miranda, B., and Koyejo, S · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
How far can camels go? exploring the state of instruction tuning on open resources
Wang, Y., Ivison, H., Dasigi, P., Hessel, J., Khot, T., Chandu, K., Wadden, D., MacMillan, K., Smith, N. A., Beltagy, I., et al · 2023
Later among the works it cites.
Unveiling transformers with lego: a synthetic reasoning task, 2023
Zhang, Y., Backurs, A., Bubeck, S., Eldan, R., Gunasekar, S., and Wagner, T · 2023
Later among the works it cites.
Establishing task scaling laws via compute-efficient model ladders
Bhagia, A., Liu, J., Wettig, A., Heineman, D., Tafjord, O., Jha, A. H., Soldaini, L., Smith, N. A., Groeneveld, D., Koh, P. W., et al · 2024
Later among the works it cites.
Scaling laws do not scale
Diaz, F. and Madaio, M · 2024
Later among the works it cites.
Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., Wu, Y., and Ji, R · 2024
Later among the works it cites.
Language models scale reliably with over-training and on downstream tasks
Gadre, S. Y., Smyrnis, G., Shankar, V., Gururangan, S., Wortsman, M., Shao, R., Mercat, J., Fang, A., Li, J., Keh, S., et al · 2024
Later among the works it cites.
Scaling laws for data filtering – data curation cannot be compute agnostic, 2024
Goyal, S., Maini, P., Lipton, Z. C., Raghunathan, A., and Kolter, J. Z · 2024
Later among the works it cites.
The llama 3 herd of models, 2024
Grattafiori, A. et al · 2024
Later among the works it cites.
Model editing with canonical examples
Hewitt, J., Chen, S., Xie, L. L., Adams, E., Liang, P., and Manning, C. D · 2024
Later among the works it cites.
Minicpm: Unveiling the potential of small language models with scalable training strategies
Hu, S., Tu, Y., Han, X., He, C., Cui, G., Long, X., Zheng, Z., Fang, Y., Huang, Y., Zhao, W., et al · 2024
Later among the works it cites.
Scaling laws for downstream task performance of large language models
Isik, B., Ponomareva, N., Hazimeh, H., Paparas, D., Vassilvitskii, S., and Koyejo, S · 2024
Later among the works it cites.
Tofu: A task of fictitious unlearning for llms
Maini, P., Feng, Z., Schwarzschild, A., Lipton, Z. C., and Kolter, J. Z · 2024
Later among the works it cites.
Deep reinforcement learning with plasticity injection
Nikishin, E., Oh, J., Ostrovski, G., Lyle, C., Pascanu, R., Dabney, W., and Barreto, A · 2024
Later among the works it cites.
OLMo, T., Walsh, P., Soldaini, L., Groeneveld, D., Lo, K., Arora, S., Bhagia, A., Gu, Y., Huang, S., Jordan, M., Lambert, N., Schwenk, D., Tafjord, O., Anderson, T., Atkinson, D., Brahman, F., Clark, C., Dasigi, P., Dziri, N., Guerquin, M., Ivison, H., Koh, P. W., Liu, J., Malik, S., Merrill, W., Miranda, L. J. V., Morrison, J., Murray, T., Nam, C., Pyatkin, V., Rangapur, A., Schmitz, M., Skjonsberg, S., Wadden, D., Wilhelm, C., Wilson, M., Zettlemoyer, L., Farhadi, A., Smith, N. A., and Hajishirzi, H · 2024
Later among the works it cites.
Beyond chinchilla-optimal: Accounting for inference in language model scaling laws, 2024
Sardana, N., Portes, J., Doubov, S., and Frankle, J · 2024
Later among the works it cites.
Why has predicting downstream capabilities of frontier ai models with scale remained elusive?
Schaeffer, R., Schoelkopf, H., Miranda, B., Mukobi, G., Madan, V., Ibrahim, A., Bradley, H., Biderman, S., and Koyejo, S · 2024
Later among the works it cites.
Decomposing and editing predictions by modeling model computation
Shah, H., Ilyas, A., and Madry, A · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
Provable unlearning in topic modeling and downstream tasks
Wei, S., Malladi, S., Arora, S., and Sanyal, A · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI et al · 2025
Closest in time.
Liu, E., Bertsch, A., Sutawika, L., Tjuatja, L., Fernandes, P., Marinov, L., Chen, M., Singhal, S., Lawrence, C., Raghunathan, A., et al · 2025
Closest in time.