Fetching the paper…
Reading the bibliography…
The computation necessary for training Transformer-based language models has skyrocketed in recent years.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 1901
Earlier work this paper cites.
Zhang, C., Bengio, S., and Singer, Y · 1902
Earlier work this paper cites.
Lookahead Optimizer: k steps forward, 1 step back, December 2019b
Zhang, M. R., Lucas, J., Hinton, G., and Ba, J · 1907
Earlier work this paper cites.
Accelerating Deep Learning by Focusing on the Biggest Losers, October 2019
Jiang, A. H., Wong, D. L.-K., Zhou, G., Andersen, D. G., Dean, J., Ganger, G. R., Joshi, G., Kaminksy, M., Kozuch, M., Lipton, Z. C., and Pillai, P · 1910
Earlier work this paper cites.
Big Transfer (BiT): General Visual Representation Learning, May 2020
Kolesnikov, A., Beyer, L., Zhai, X., Puigcerver, J., Yung, J., Gelly, S., and Houlsby, N · 1912
Earlier work this paper cites.
The cascade-correlation learning architecture
Fahlman, S. and Lebiere, C · 1989
Earlier work this paper cites.
Training mlps layer by layer using an objective function for internal representations
Lengellé, R. and Denoeux, T · 1996
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Variance reduction in SGD by distributed importance sampling
Alain, G., Lamb, A., Sankar, C., Courville, A. C., and Bengio, Y · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R. S., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Incorporating nesterov momentum into adam
Dozat, T · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q · 2016
Earlier work this paper cites.
FreezeOut: Accelerate Training by Progressively Freezing Layers, June 2017
Brock, A., Lim, T., Ritchie, J. M., and Weston, N · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N. M., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q. V., Hinton, G. E., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
On the optimization of deep networks: Implicit acceleration by overparameterization
Arora, S., Cohen, N., and Hazan, E · 2018
Earlier work this paper cites.
Not all samples are created equal: Deep learning with importance sampling
Katharopoulos, A. and Fleuret, F · 2018
Earlier work this paper cites.
An alternative view: When does sgd escape local minima?
Kleinberg, B., Li, Y., and Yuan, Y · 2018
Earlier work this paper cites.
Kudo, T. and Richardson, J · 2018
Earlier work this paper cites.
Mixed precision training
Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., and Wu, H · 2018
Earlier work this paper cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost, 2018
Shazeer, N. and Stern, M · 2018
Earlier work this paper cites.
Super-convergence: Very fast training of neural networks using large learning rates, 2018
Smith, L. N. and Topin, N · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
Xing, C., Arpit, D., Tsirigotis, C., and Bengio, Y · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
Sparse networks from scratch: Faster training without losing performance
Dettmers, T. and Zettlemoyer, L · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout, 2019
Fan, A., Grave, E., and Joulin, A · 2019
Earlier work this paper cites.
Efficient Training of BERT by Progressively Stacking
Gong, L., He, D., Li, Z., Qin, T., Wang, L., and Liu, T · 2019
Earlier work this paper cites.
Asymmetric valleys: Beyond sharp and flat local minima
He, H., Huang, G., and Yuan, Y · 2019
Earlier work this paper cites.
Komatsuzaki, A · 2019
Earlier work this paper cites.
Green ai, 2019
Schwartz, R., Dodge, J., Smith, N. A., and Etzioni, O · 2019
Earlier work this paper cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2019
Earlier work this paper cites.
Scalable second order optimization for deep learning
Anil, R., Gupta, V., Koren, T., Regan, K., and Singer, Y · 2020
Earlier work this paper cites.
Selection via proxy: Efficient data selection for deep learning
Coleman, C., Yeh, C., Mussmann, S., Mirzasoleiman, B., Bailis, P., Liang, P., Leskovec, J., and Zaharia, M · 2020
Cited alongside, same era.
Rigging the lottery: Making all tickets winners
Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E · 2020
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Cited alongside, same era.
Realm: Retrieval-augmented language model pre-training, 2020
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M.-W · 2020
Cited alongside, same era.
Deduplicating Training Data Makes Language Models Better, March 2022
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N · 2022
Later among the works it cites.
Automated Progressive Learning for Efficient Training of Vision Transformers
Li, C., Zhuang, B., Wang, G., Liang, X., Chang, X., and Yang, Y · 2022
Later among the works it cites.
Prioritized training on points that are learnable, worth learning, and not yet learnt
Mindermann, S., Brauner, J. M., Razzak, M. T., Sharma, M., Kirsch, A., Xu, W., Höltgen, B., Gomez, A. N., Morisot, A., Farquhar, S., et al · 2022
Later among the works it cites.
Efficient transformers with dynamic token pooling
Nawrot, P., Chorowski, J., La’ncucki, A., and Ponti, E · 2022
Later among the works it cites.
Active learning is a strong baseline for data subset selection
Park, D., Papailiopoulos, D., and Lee, K · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
Ordered sgd: A new stochastic optimization framework for empirical risk minimization
Kawaguchi, K. and Lu, H · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Glu variants improve transformer, 2020
Shazeer, N · 2020
Cited alongside, same era.
Wu, X., Dyer, E., and Neyshabur, B · 2020
Cited alongside, same era.
Accelerating training of transformer-based language models with progressive layer dropping
Zhang, M. and He, Y · 2020
Cited alongside, same era.
AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed Gradients
Zhuang, J., Tang, T., Ding, Y., Tatikonda, S. C., Dvornek, N., Papademetris, X., and Duncan, J · 2020
Cited alongside, same era.
Later among the works it cites.
The carbon footprint of machine learning training will plateau, then shrink
Patterson, D., Gonzalez, J., Hölzle, U., Le, Q., Liang, C., Munguia, L.-M., Rothchild, D., So, D. R., Texier, M., and Dean, J · 2022
Later among the works it cites.
Fast Benchmarking of Accuracy vs. Training Time with Cyclic Learning Rates, November 2022
Portes, J., Blalock, D., Stephenson, C., and Frankle, J · 2022
Later among the works it cites.
Staged Training for Transformer Language Models
Shen, S., Walsh, P., Keutzer, K., Dodge, J., Peters, M., and Beltagy, I · 2022
Later among the works it cites.
Beyond neural scaling laws: beating power law scaling via data pruning
Sorscher, B., Geirhos, R., Shekhar, S., Ganguli, S., and Morcos, A. S · 2022
Later among the works it cites.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks, 2022
Wang, Y., Mishra, S., Alipoormolabashi, P., Kordi, Y., Mirzaei, A., Arunkumar, A., Ashok, A., Dhanasekaran, A. S., Naik, A., Stap, D., Pathak, E., Karamanolakis, G., Lai, H. G., Purohit, I., Mondal, I., Anderson, J., Kuznia, K., Doshi, K., Patel, M., Pal, K. K., Moradshahi, M., Parmar, M., Purohit, M., Varshney, N., Kaza, P. R., Verma, P., Puri, R. S., Karia, R., Sampat, S. K., Doshi, S., Mishra, S., Reddy, S., Patro, S., Dixit, T., Shen, X., Baral, C., Choi, Y., Smith, N. A., Hajishirzi, H., and Khashabi, D · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E., Le, Q., and Zhou, D · 2022
Later among the works it cites.
Nlp from scratch without large-scale pretraining: A simple and efficient framework
Yao, X., Zheng, Y., Yang, X., and Yang, Z · 2022
Later among the works it cites.
OPT: Open Pre-trained Transformer Language Models, June 2022
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L · 2022
Later among the works it cites.
Semdedup: Data-efficient learning at web-scale through semantic deduplication
Abbas, A., Tirumala, K., Simig, D., Ganguli, S., and Morcos, A. S · 2023
Closest in time.
Colt5: Faster long-range transformers with conditional computation
Ainslie, J., Lei, T., de Jong, M., Ontan’on, S., Brahma, S., Zemlyanskiy, Y., Uthus, D. C., Guo, M., Lee-Thorp, J., Tay, Y., Sung, Y.-H., and Sanghai, S. K · 2023
Closest in time.
Compute-efficient deep learning: Algorithmic trends and opportunities
Bartoldson, B. R., Kailkhura, B., and Blalock, D · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Closest in time.
Symbolic discovery of optimization algorithms
Chen, X., Liang, C., Huang, D., Real, E., Wang, K., Liu, Y., Pham, H., Dong, X., Luong, T., Hsieh, C.-J., et al · 2023
Closest in time.
Benchmarking neural network training algorithms
Dahl, G. E., Schneider, F., Nado, Z., Agarwal, N., Sastry, C. S., Hennig, P., Medapati, S., Eschenhagen, R., Kasimbeg, P., Suo, D., et al · 2023
Closest in time.
The minipile challenge for data-efficient language models
Kaddour, J · 2023
Closest in time.
Challenges and Applications of Large Language Models
Kaddour, J., Harris, J., Mozes, M., Bradley, H., Raileanu, R., and McHardy, R · 2023
Closest in time.
Does this work with 16-mixed precision
Liu, H · 2023
Closest in time.
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training, May 2023
Liu, H., Li, Z., Hall, D., Liang, P., and Ma, T · 2023
Closest in time.
Liu, S. and Wang, Z · 2023
Closest in time.
Longpre, S., Yauney, G., Reif, E., Lee, K., Roberts, A., Zoph, B., Zhou, D., Wei, J., Robinson, K., Mimno, D., and Ippolito, D · 2023
Closest in time.
Efficient deep learning: A survey on making deep learning models smaller, faster, and better
Menghani, G · 2023
Closest in time.
nanot5: A pytorch framework for pre-training and fine-tuning t5-style models with limited resources
Nawrot, P · 2023
Closest in time.
Training trajectories, mini-batch losses and the curious role of the learning rate
Sandler, M., Zhmoginov, A., Vladymyrov, M., and Miller, N · 2023
Closest in time.
Understanding the effectiveness of early weight averaging for training large language models
Sanyal, S., Kaddour, J., Kumar, A., and Sanghavi, S · 2023
Closest in time.
On efficient training of large-scale deep learning models: A literature review
Shen, L., Sun, Y., Yu, Z., Ding, L., Tian, X., and Tao, D · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Closest in time.
Learning to grow pretrained models for efficient transformer training
Wang, P., Panda, R., Hennigen, L. T., Greengard, P., Karlinsky, L., Feris, R., Cox, D. D., Wang, Z., and Kim, Y · 2023
Closest in time.
To Repeat or Not To Repeat: Insights from Scaling LLM under Token-Crisis, May 2023
Xue, F., Fu, Y., Zhou, W., Zheng, Z., and You, Y · 2023
Closest in time.
A survey on efficient training of transformers
Zhuang, B., Liu, J., Pan, Z., He, H., Weng, Y., and Shen, C · 2023
Closest in time.
Budgeted training for vision transformer
zhuofan xia, Pan, X., Jin, X., He, Y., Xue’, H., Song, S., and Huang, G · 2023
Closest in time.