Fetching the paper…
Reading the bibliography…
OpenDiLoCo is an open-source implementation and replication of the Distributed Low-Communication (DiLoCo) training method for large language models.
A method for solving the convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Fixing weight decay regularization in adam
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Micikevicius, P., Narang, S., Alben, J., Diamos, G. F., Elsen, E., García, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., and Wu, H · 2017
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N. M., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Earlier work this paper cites.
Local sgd converges fast and communicates little, 2019
Stich, S. U · 2019
Earlier work this paper cites.
Glu variants improve transformer, 2020
Shazeer, N · 2020
Cited alongside, same era.
Hivemind: a Library for Decentralized Deep Learning
team, L · 2020
Cited alongside, same era.
Training compute-optimal large language models, 2022
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Cited alongside, same era.
Diloco: Distributed low-communication training of language models, 2023
Douillard, A., Feng, Q., Rusu, A. A., Chhaparia, R., Donchev, Y., Kuncoro, A., Ranzato, M., Szlam, A., and Shen, J · 2023
Cited alongside, same era.
Efficient parallelization layouts for large-scale distributed model training, 2023
Hagemann, J., Weinbach, S., Dobler, K., Schall, M., and de Melo, G · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
Pytorch fsdp: Experiences on scaling fully sharded data parallel, 2023
Zhao, Y., Gu, A., Varma, R., Luo, L., Huang, C.-C., Xu, M., Wright, L., Shojanazeri, H., Ott, M., Shleifer, S., Desmaison, A., Balioglu, C., Damania, P., Nguyen, B., Chauhan, G., Hao, Y., Mathews, A., and Li, S · 2023
Later among the works it cites.
Asynchronous local-sgd training for language modeling, 2024
Liu, B., Chhaparia, R., Douillard, A., Kale, S., Rusu, A. A., Shen, J., Szlam, A., and Ranzato, M · 2024
Closest in time.
Tinyllama: An open-source small language model, 2024
Zhang, P., Zeng, G., Wang, T., and Lu, W · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…