Fetching the paper…
Reading the bibliography…
Large language models are typically trained densely: all parameters are updated with respect to all inputs.
fairseq: A fast, extensible toolkit for sequence modeling, 2019
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 1904
Earlier work this paper cites.
Sparse networks from scratch: Faster training without losing performance
Dettmers, T. and Zettlemoyer, L · 1907
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2019
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 1910
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models, 2019
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M · 1911
Earlier work this paper cites.
S2orc: The semantic scholar open research corpus, 2019
Lo, K., Wang, L. L., Neumann, M., Kinney, R., and Weld, D. S · 1911
Earlier work this paper cites.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Auction algorithms for network flow problems: A tutorial introduction
Bertsekas, D. P · 1992
Earlier work this paper cites.
A density-based algorithm for discovering clusters in large spatial databases with noise
Ester, M., Kriegel, H.-P., Sander, J., and Xu, X · 1996
Earlier work this paper cites.
K-means++: The advantages of careful seeding
Arthur, D. and Vassilvitskii, S · 2007
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Balanced k-means and min-cut clustering, 2014
Chang, X., Nie, F., Ma, Z., and Yang, Y · 2014
Earlier work this paper cites.
Dbpedia - a large-scale, multilingual knowledge base extracted from wikipedia
Lehmann, J., Isele, R., Jakob, M., Jentzsch, A., Kontokostas, D., Mendes, P., Hellmann, S., Morsey, M., Van Kleef, P., Auer, S., and Bizer, C · 2014
Earlier work this paper cites.
Balanced k-means for clustering
Malinen, M. I. and Fränti, P · 2014
Earlier work this paper cites.
Good debt or bad debt: Detecting semantic orientations in economic texts
Malo, P., Sinha, A., Korhonen, P., Wallenius, J., and Takala, P · 2014
Earlier work this paper cites.
Character-level convolutional networks for text classification, 2016
Zhang, X., Zhao, J., and LeCun, Y · 2016
Earlier work this paper cites.
Hard mixtures of experts for large scale weakly supervised vision
Gross, S., Ranzato, M., and Szlam, A · 2017
Earlier work this paper cites.
Hash embeddings for efficient word representations, 2017
Svenstrup, D., Hansen, J. M., and Winther, O · 2017
Earlier work this paper cites.
An empirical model of large-batch training, 2018
McCandlish, S., Kaplan, J., Amodei, D., and Team, O. D · 2018
Earlier work this paper cites.
Measuring the effects of data parallelism on neural network training
Shallue, C. J., Lee, J., Antognini, J., Sohl-Dickstein, J., Frostig, R., and Dahl, G. E · 2018
Earlier work this paper cites.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Mostafa, H. and Wang, X · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
Unsupervised domain clusters in pretrained language models
Aharoni, R. and Goldberg, Y · 2020
Cited alongside, same era.
TweetEval: Unified benchmark and comparative evaluation for tweet classification
Barbieri, F., Camacho-Collados, J., Espinosa Anke, L., and Neves, L · 2020
Cited alongside, same era.
Rigging the lottery: Making all tickets winners
Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E · 2020
Cited alongside, same era.
Efficient large scale language modeling with mixtures of experts, 2021
Artetxe, M., Bhosale, S., Goyal, N., Mihaylov, T., Ott, M., Shleifer, S., Lin, X. V., Du, J., Iyer, S., Pasunuru, R., Anantharaman, G., Li, X., Chen, S., Akin, H., Baines, M., Martin, L., Zhou, X., Koura, P. S., O’Horo, B., Wang, J., Zettlemoyer, L., Diab, M., Kozareva, Z., and Stoyanov, V · 2021
Cited alongside, same era.
Unified scaling laws for routed language models, 2022
Clark, A., Casas, D. d. l., Guy, A., Mensch, A., Paganini, M., Hoffmann, J., Damoc, B., Hechtman, B., Cai, T., Borgeaud, S., Driessche, G. v. d., Rutherford, E., Hennigan, T., Johnson, M., Millican, K., Cassirer, A., Jones, C., Buchatskaya, E., Budden, D., Sifre, L., Osindero, S., Vinyals, O., Rae, J., Elsen, E., Kavukcuoglu, K., and Simonyan, K · 2022
Later among the works it cites.
A review of sparse expert models in deep learning, 2022
Fedus, W., Dean, J., and Zoph, B · 2022
Later among the works it cites.
DEMix layers: Disentangling domains for modular language modeling
Gururangan, S., Lewis, M., Holtzman, A., Smith, N. A., and Zettlemoyer, L · 2022
Later among the works it cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Later among the works it cites.
Sparse upcycling: Training mixture-of-experts from dense checkpoints, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dehghani, M., Arnab, A., Beyer, L., Vaswani, A., and Tay, Y · 2021
Cited alongside, same era.
Enslm: Ensemble language model for data diversity by semantic clustering
Duan, Z., Zhang, H., Wang, C., Wang, Z., Chen, B., and Zhou, M · 2021
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2021
Cited alongside, same era.
Bagua: Scaling up distributed learning with system relaxations, 2021
Gan, S., Lian, X., Wang, R., Chang, J., Liu, C., Shi, H., Zhang, S., Li, X., Sun, T., Jiang, J., Yuan, B., Yang, S., Liu, J., and Zhang, C · 2021
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling, 2021
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2021
Cited alongside, same era.
Beyond distillation: Task-level mixture-of-experts for efficient inference, 2021
Kudugunta, S., Huang, Y., Bapna, A., Krikun, M., Lepikhin, D., Luong, M.-T., and Firat, O · 2021
Cited alongside, same era.
{GS}hard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2021
Cited alongside, same era.
Komatsuzaki, A., Puigcerver, J., Lee-Thorp, J., Ruiz, C. R., Mustafa, B., Ainslie, J., Tay, Y., Dehghani, M., and Houlsby, N · 2022
Later among the works it cites.
Branch-train-merge: Embarrassingly parallel training of expert language models, 2022
Li, M., Gururangan, S., Dettmers, T., Lewis, M., Althoff, T., Smith, N. A., and Zettlemoyer, L · 2022
Later among the works it cites.
Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Lu, Y., Bartolo, M., Moore, A., Riedel, S., and Stenetorp, P · 2022
Later among the works it cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Later among the works it cites.
Lifting the curse of multilinguality by pre-training modular transformers
Pfeiffer, J., Goyal, N., Lin, X., Li, X., Cross, J., Riedel, S., and Artetxe, M · 2022
Later among the works it cites.
knn-prompt: Nearest neighbor zero-shot inference, 2022
Shi, W., Michael, J., Gururangan, S., and Zettlemoyer, L · 2022
Later among the works it cites.
Galactica: A large language model for science, 2022
Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., Kerkez, V., and Stojnic, R · 2022
Later among the works it cites.
Fine-tuning language models over slow networks using activation compression with guarantees, 2022
Wang, J., Yuan, B., Rimanic, L., He, Y., Dao, T., Chen, B., Re, C., and Zhang, C · 2022
Later among the works it cites.
lo-fi: distributed fine-tuning without communication, 2022
Wortsman, M., Gururangan, S., Li, S., Farhadi, A., Schmidt, L., Rabbat, M., and Morcos, A. S · 2022
Later among the works it cites.
Decentralized training of foundation models in heterogeneous environments, 2022
Yuan, B., He, Y., Davis, J. Q., Zhang, T., Dao, T., Chen, B., Liang, P., Re, C., and Zhang, C · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models, 2022
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing, 2022
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A., Chen, Z., Le, Q., and Laudon, J · 2022
Later among the works it cites.
Adaptersoup: Weight averaging to improve generalization of pretrained language models
Chronopoulou, A., Peters, M. E., Fraser, A. M., and Dodge, J · 2023
Closest in time.
Exploring the benefits of training expert language models over instruction tuning, 2023
Jang, J., Kim, S., Ye, S., Kim, D., Logeswaran, L., Lee, M., Lee, K., and Seo, M · 2023
Closest in time.
Swarm parallelism: Training large models can be surprisingly communication-efficient, 2023
Ryabinin, M., Dettmers, T., Diskin, M., and Borzunov, A · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Closest in time.