Fetching the paper…
Reading the bibliography…
The coverage and composition of the pretraining data significantly impacts the generalization ability of Large Language Models (LLMs).
Problem complexity and Method Efficiency in Optimization , volume 1
Nemirovski, A. and Yudin, D · 1983
Earlier work this paper cites.
Mirror descent and nonlinear projected subgradient methods for convex optimization
Beck, A. and Teboulle, M · 2003
Earlier work this paper cites.
SemEval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Gordon, A., Kozareva, Z., and Roemmele, M · 2012
Earlier work this paper cites.
Crowdsourcing multiple choice science questions
Welbl, J., Liu, N. F., and Gardner, M · 2017
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language, 2019
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
On the power of curriculum learning in training deep networks, 2019
Hacohen, G. and Weinshall, D · 2019
Earlier work this paper cites.
Wic: the word-in-context dataset for evaluating context-sensitive meaning representations, 2019
Pilehvar, M. T. and Camacho-Collados, J · 2019
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale, 2019
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2019
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling, 2020
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
Wiki-40B: Multilingual language model dataset
Guo, M., Dai, Z., Vrandečić, D., and Al-Rfou, R · 2020
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Earlier work this paper cites.
Logiqa: A challenge dataset for machine reading comprehension with logical reasoning, 2020
Liu, J., Cui, L., Liu, H., Huang, D., Wang, Y., and Zhang, Y · 2020
Cited alongside, same era.
Estimating training data influence by tracing gradient descent, 2020
Pruthi, G., Liu, F., Sundararajan, M., and Kale, S · 2020
Cited alongside, same era.
Curriculum learning for natural language understanding
Xu, B., Zhang, L., Mao, Z., Wang, Q., Xie, H., and Zhang, Y · 2020
Cited alongside, same era.
A framework for few-shot language model evaluation, September 2021
Gao, L., Tow, J., Biderman, S., Black, S., DiPofi, A., Foster, C., Golding, L., Hsu, J., McDonell, K., Muennighoff, N., Phang, J., Reynolds, L., Tang, E., Thite, A., Wang, B., Wang, K., and Zou, A · 2021
Cited alongside, same era.
Model performance scaling with multiple data sources
Hashimoto, T · 2021
Cited alongside, same era.
Palm: Scaling language modeling with pathways, 2022
Adaptive training distributions with scalable online bilevel optimization, 2023
Grangier, D., Ablin, P., and Hannun, A · 2023
Closest in time.
Textbooks are all you need, 2023
Gunasekar, S., Zhang, Y., Aneja, J., Mendes, C. C. T., Giorno, A. D., Gopi, S., Javaheripi, M., Kauffmann, P., de Rosa, G., Saarikivi, O., Salim, A., Shah, S., Behl, H. S., Wang, X., Bubeck, S., Eldan, R., Kalai, A. T., Lee, Y. T., and Li, Y · 2023
Closest in time.
Beyond scale: the diversity coefficient as a data quality metric demonstrates llms are pre-trained on formally diverse data, 2023
Lee, A., Miranda, B., and Koyejo, S · 2023
Closest in time.
Textbooks are all you need ii: phi-1.5 technical report, 2023
Li, Y., Bubeck, S., Eldan, R., Giorno, A. D., Gunasekar, S., and Lee, Y. T · 2023
Closest in time.
A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, toxicity, 2023
Longpre, S., Yauney, G., Reif, E., Lee, K., Roberts, A., Zoph, B., Zhou, D., Wei, J., Robinson, K., Mimno, D., and Ippolito, D · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
Cited alongside, same era.
Family tree of languages – part i: Indo-european (2022), 03 2022
Cole, T. and Siebert-Cole, E · 2022
Cited alongside, same era.
Glam: Efficient scaling of language models with mixture-of-experts, 2022
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., Zoph, B., Fedus, L., Bosma, M., Zhou, Z., Wang, T., Wang, Y. E., Webster, K., Pellat, M., Robinson, K., Meier-Hellstern, K., Duke, T., Dixon, L., Zhang, K., Le, Q. V., Wu, Y., Chen, Z., and Cui, C · 2022
Cited alongside, same era.
Training compute-optimal large language models, 2022
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Cited alongside, same era.
First is better than last for language data influence, 2022
Yeh, C.-K., Taly, A., Sundararajan, M., Liu, F., and Ravikumar, P · 2022
Cited alongside, same era.
Skill-it! a data-driven skills framework for understanding and training language models, 2023
Chen, M. F., Roberts, N., Bhatia, K., Wang, J., Zhang, C., Sala, F., and Ré, C · 2023
Cited alongside, same era.
Gio: Gradient information optimization for training dataset selection, 2023
Everaert, D. and Potts, C · 2023
Cited alongside, same era.
Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E., and Launay, J · 2023
Closest in time.
SlimPajama: A 627B token cleaned and deduplicated version of RedPajama
Soboleva, D., Al-Khateeb, F., Myers, R., Steeves, J. R., Hestness, J., and Dey, N · 2023
Closest in time.
Self-influence guided data reweighting for language model pre-training, 2023
Thakkar, M., Bolukbasi, T., Ganapathy, S., Vashishth, S., Chandar, S., and Talukdar, P · 2023
Closest in time.
Redpajama: An open source recipe to reproduce llama training dataset, 2023
Together Computer · 2023
Closest in time.
Attention is all you need, 2023
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2023
Closest in time.
Probabilistic bilevel coreset selection, 2023
Zhou, X., Pi, R., Zhang, W., Lin, Y., and Zhang, T · 2023
Closest in time.
Dsdm: Model-aware dataset selection with datamodels, 2024
Engstrom, L., Feldmann, A., and Madry, A · 2024
Closest in time.