Fetching the paper…
Reading the bibliography…
Overparameterized large-scale language models have impressive generalization performance of in-context few-shot learning.
Individual differences in reasoning: Implications for the rationality debate?
Stanovich, K. E. and West, R. F · 2000
Earlier work this paper cites.
Expectation-based syntactic comprehension
Levy, R · 2007
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A. C · 2013
Earlier work this paper cites.
Conditional computation in neural networks for faster models
Bengio, E., Bacon, P., Pineau, J., and Precup, D · 2015
Earlier work this paper cites.
Srivastava, R. K., Greff, K., and Schmidhuber, J · 2015
Earlier work this paper cites.
Efficient object localization using convolutional networks
Tompson, J., Goroshin, R., Jain, A., LeCun, Y., and Bregler, C · 2015
Earlier work this paper cites.
Hard mixtures of experts for large scale weakly supervised vision
Gross, S., Ranzato, M., and Szlam, A · 2017
Earlier work this paper cites.
In-datacenter performance analysis of a tensor processing unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
IDK cascades: Fast deep learning by learning not to overthink
Wang, X., Luo, Y., Crankshaw, D., Tumanov, A., and Gonzalez, J. E · 2017
Earlier work this paper cites.
Dropblock: A regularization method for convolutional networks
Ghiasi, G., Lin, T.-Y., and Le, Q. V · 2018
Earlier work this paper cites.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Kudo, T. and Richardson, J · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2018
Earlier work this paper cites.
Skipnet: Learning dynamic routing in convolutional networks
Wang, X., Yu, F., Dou, Z., Darrell, T., and Gonzalez, J. E · 2018
Cited alongside, same era.
Batch dropblock network for person re-identification and beyond
Dai, Z., Chen, M., Gu, X., Zhu, S., and Tan, P · 2019
Cited alongside, same era.
Reducing transformer depth on demand with structured dropout
Fan, A., Grave, E., and Joulin, A · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using gpu model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Cited alongside, same era.
Controlling computation versus quality for neural sequence models, 2020
Bapna, A., Arivazhagan, N., and Firat, O · 2020
Cited alongside, same era.
Base layers: Simplifying training of large, sparse models
Lewis, M., Bhosale, S., Dettmers, T., Goyal, N., and Zettlemoyer, L · 2021
Later among the works it cites.
Carbon emissions and large neural network training
Patterson, D., Gonzalez, J., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D., Texier, M., and Dean, J · 2021
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher
Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, H. F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P., Glaese, A., Welbl, J., Dathathri, S., Huang, S., Uesato, J., Mellor, J., Higgins, I., Creswell, A., McAleese, N., Wu, A., Elsen, E., Jayakumar, S. M., Buchatskaya, E., Budden, D., Sutherland, E., Simonyan, K., Paganini, M., Sifre, L., Martens, L., Li, X. L., Kuncoro, A., Nematzadeh, A., Gribovskaya, E., Donato, D., Lazaridou, A., Mensch, A., Lespiau, J., Tsimpoukelli, M., Grigorev, N., Fritz, D., Sottiaux, T., Pajarskas, M., Pohlen, T., Gong, Z., Toyama, D., de Masson d’Autume, C., Li, Y., Terzi, T., Mikulik, V., Babuschkin, I., Clark, A., de Las Casas, D., Guy, A., Jones, C., Bradbury, J., Johnson, M., Hechtman, B. A., Weidinger, L., Gabriel, I., Isaac, W. S., Lockhart, E., Osindero, S., Rimell, L., Dyer, C., Vinyals, O., Ayoub, K., Stanway, J., Bennett, L., Hassabis, D., Kavukcuoglu, K., and Irving, G · 2021
Later among the works it cites.
Hash layers for large sparse models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Balancing cost and benefit with tied-multi transformers
Dabre, R., Rubino, R., and Fujita, A · 2020
Cited alongside, same era.
Depth-adaptive transformer
Elbayad, M., Gu, J., Grave, E., and Auli, M · 2020
Cited alongside, same era.
FastBERT: a self-distilling BERT with adaptive inference time
Liu, W., Zhou, P., Wang, Z., Zhao, Z., Deng, H., and Ju, Q · 2020
Cited alongside, same era.
The right tool for the job: Matching model and instance complexities
Schwartz, R., Stanovsky, G., Swayamdipta, S., Dodge, J., and Smith, N. A · 2020
Cited alongside, same era.
Deebert: Dynamic early exiting for accelerating bert inference
Xin, J., Tang, R., Lee, J., Yu, Y., and Lin, J · 2020
Cited alongside, same era.
Efficient large scale language modeling with mixtures of experts
Artetxe, M., Bhosale, S., Goyal, N., Mihaylov, T., Ott, M., Shleifer, S., Lin, X. V., Du, J., Iyer, S., Pasunuru, R., Anantharaman, G., Li, X., Chen, S., Akin, H., Baines, M., Martin, L., Zhou, X., Koura, P. S., O’Horo, B., Wang, J., Zettlemoyer, L., Diab, M., Kozareva, Z., and Stoyanov, V · 2021
Cited alongside, same era.
Roller, S., Sukhbaatar, S., szlam, a., and Weston, J · 2021
Later among the works it cites.
Correlation-based structural dropout for convolutional neural networks
Zeng, Y., Dai, T., Chen, B., Xia, S.-T., and Lu, J · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., Reif, E., Du, N., Hutchinson, B., Pope, R., Bradbury, J., Austin, J., Isard, M., Gur-Ari, G., Yin, P., Duke, T., Levskaya, A., Ghemawat, S., Dev, S., Michalewski, H., Garcia, X., Misra, V., Robinson, K., Fedus, L., Zhou, D., Ippolito, D., Luan, D., Lim, H., Zoph, B., Spiridonov, A., Sepassi, R., Dohan, D., Agrawal, S., Omernick, M., Dai, A. M., Pillai, T. S., Pellat, M., Lewkowycz, A., Moreira, E., Child, R., Polozov, O., Lee, K., Zhou, Z., Wang, X., Saeta, B., Diaz, M., Firat, O., Catasta, M., Wei, J., Meier-Hellstern, K., Eck, D., Dean, J., Petrov, S., and Fiedel, N · 2022
Later among the works it cites.
GLaM: Efficient scaling of language models with mixture-of-experts
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., Zoph, B., Fedus, L., Bosma, M. P., Zhou, Z., Wang, T., Wang, E., Webster, K., Pellat, M., Robinson, K., Meier-Hellstern, K., Duke, T., Dixon, L., Zhang, K., Le, Q., Wu, Y., Chen, Z., and Cui, C · 2022
Later among the works it cites.
An empirical analysis of compute-optimal large language model training
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Vinyals, O., Rae, J. W., and Sifre, L · 2022
Later among the works it cites.
The unreasonable effectiveness of fully-connected layers for low-data regimes
Kocsis, P., Súkeník, P., Brasó, G., Nießner, M., Leal-Taixé, L., and Elezi, I · 2022
Later among the works it cites.
DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation AI scale
Rajbhandari, S., Li, C., Yao, Z., Zhang, M., Aminabadi, R. Y., Awan, A. A., Rasley, J., and He, Y · 2022
Later among the works it cites.
Confident adaptive language modeling
Schuster, T., Fisch, A., Gupta, J. P., Dehghani, M., Bahri, D., Tran, V. Q., Tay, Y., and Metzler, D · 2022
Later among the works it cites.
Mixture-of-experts with expert choice routing
Zhou, Y., Lei, T., Liu, H., Du, N., Huang, Y., Zhao, V., Dai, A., Chen, Z., Le, Q., and Laudon, J · 2022
Later among the works it cites.
Colt5: Faster long-range transformers with conditional computation
Ainslie, J., Lei, T., de Jong, M., Ontañón, S., Brahma, S., Zemlyanskiy, Y., Uthus, D., Guo, M., Lee-Thorp, J., Tay, Y., et al · 2023
Closest in time.
Conditional adapters: Parameter-efficient transfer learning with fast inference
Lei, T., Bai, J., Brahma, S., Ainslie, J., Lee, K., Zhou, Y., Du, N., Zhao, V. Y., Wu, Y., Li, B., Zhang, Y., and Chang, M.-W · 2023
Closest in time.