Fetching the paper…
Reading the bibliography…
In this technical report, we present Zyda-2: a five trillion token dataset for language model pretraining.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 2005
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al. (2020) · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2020) · 2020
Earlier work this paper cites.
Deduplicating training data makes language models better
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N. (2021) · 2021
Earlier work this paper cites.
Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E., and Launay, J. (2023) · 2023
Earlier work this paper cites.
Phi-3 technical report: A highly capable language model locally on your phone
Abdin, M., Jacobs, S. A., Awan, A. A., Aneja, J., Awadallah, A., Awadalla, H., Bach, N., Bahree, A., Bakhtiari, A., Behl, H., et al. (2024) · 2024
Earlier work this paper cites.
Perplexed by perplexity: Perplexity-based data pruning with small reference models
Ankner, Z., Blakeney, C., Sreenivasan, K., Marion, M., Leavitt, M. L., and Paul, M. (2024) · 2024
Earlier work this paper cites.
Blackmamba: Mixture of experts for state-space models
Anthony, Q., Tokpanov, Y., Glorioso, P., and Millidge, B. (2024) · 2024
Cited alongside, same era.
Does your data spark joy? performance gains from domain upsampling at the end of training
Blakeney, C., Paul, M., Larsen, B. W., Owen, S., and Frankle, J. (2024) · 2024
Cited alongside, same era.
Gemma: Open models based on gemini research and technology
Gemma Team, Mesnard, T., Hardin, C., Dadashi, R., Bhupatiraju, S., Pathak, S., Sifre, L., Rivière, M., Kale, M. S., Love, J., et al. (2024) · 2024
Cited alongside, same era.
Zamba: A compact 7b ssm hybrid model
Glorioso, P., Anthony, Q., Tokpanov, Y., Whittington, J., Pilault, J., Ibrahim, A., and Millidge, B. (2024) · 2024
Cited alongside, same era.
Minicpm: Unveiling the potential of small language models with scalable training strategies
Nemo-curator: a toolkit for data curation
Nvidia (2024) · 2024
Closest in time.
Nemotron-4 15b technical report
Parmar, J., Prabhumoye, S., Jennings, J., Patwary, M., Subramanian, S., Su, D., Zhu, C., Narayanan, D., Jhunjhunwala, A., Dattagupta, A., et al. (2024) · 2024
Closest in time.
The fineweb datasets: Decanting the web for the finest text data at scale
Penedo, G., Kydlíček, H., allal, L. B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L. V., and Wolf, T. (2024) · 2024
Closest in time.
Dolma: an open corpus of three trillion tokens for language model pretraining research
Soldaini, L., Kinney, R., Bhagia, A., Schwenk, D., Atkinson, D., Authur, R., Bogin, B., Chandu, K., Dumas, J., Elazar, Y., Hofmann, V., Jha, A. H., Kumar, S., Lucy, L., Lyu, X., Lambert, N., Magnusson, I., Morrison, J., Muennighoff, N., Naik, A., Nam, C., Peters, M. E., Ravichander, A., Richardson, K., Shen, Z., Strubell, E., Subramani, N., Tafjord, O., Walsh, P., Zettlemoyer, L., Smith, N. A., Hajishirzi, H., Beltagy, I., Groeneveld, D., Dodge, J., and Lo, K. (2024) · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hu, S., Tu, Y., Han, X., He, C., Cui, G., Long, X., Zheng, Z., Fang, Y., Huang, Y., Zhao, W., et al. (2024) · 2024
Cited alongside, same era.
Datacomp-lm: In search of the next generation of training sets for language models
Li, J., Fang, A., Smyrnis, G., Ivgi, M., Jordan, M., Gadre, S., Bansal, H., Guha, E., Keh, S., Arora, K., Garg, S., Xin, R., Muennighoff, N., Heckel, R., Mercat, J., Chen, M., Gururangan, S., Wortsman, M., Albalak, A., Bitton, Y., Nezhurina, M., Abbas, A., Hsieh, C.-Y., Ghosh, D., Gardner, J., Kilian, M., Zhang, H., Shao, R., Pratt, S., Sanyal, S., Ilharco, G., Daras, G., Marathe, K., Gokaslan, A., Zhang, J., Chandu, K., Nguyen, T., Vasiljevic, I., Kakade, S., Song, S., Sanghavi, S., Faghri, F., Oh, S., Zettlemoyer, L., Lo, K., El-Nouby, A., Pouransari, H., Toshev, A., Wang, S., Groeneveld, D., Soldaini, L., Koh, P. W., Jitsev, J., Kollar, T., Dimakis, A. G., Carmon, Y., Dave, A., Schmidt, L., and Shankar, V. (2024) · 2024
Cited alongside, same era.
Rephrasing the web: A recipe for compute and data-efficient language modeling
Maini, P., Seto, S., Bai, H., Grangier, D., Zhang, Y., and Jaitly, N. (2024) · 2024
Cited alongside, same era.
Tokpanov, Y., Millidge, B., Glorioso, P., Pilault, J., Ibrahim, A., Whittington, J., and Anthony, Q. (2024) · 2024
Closest in time.
Introducing Meta Llama 3: The most capable openly available LLM to date
Meta (2024) · 2026
Closest in time.