Fetching the paper…
Reading the bibliography…
Since hardware resources are limited, the objective of training deep learning models is typically to maximize accuracy subject to the time and memory constraints of training and inference.
RoBERTa: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 1907
Earlier work this paper cites.
Attentive student meets multi-task teacher: Improved knowledge distillation for pretrained models
Liu, L., Wang, H., Lin, J., Socher, R., and Xiong, C · 1911
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A · 1990
Earlier work this paper cites.
Sparse connection and pruning in large dynamic artificial neural networks
Ström, N · 1997
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Efficient and robust automated machine learning
Feurer, M., Klein, A., Eggensperger, K., Springenberg, J., Blum, M., and Hutter, F · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Earlier work this paper cites.
Layer normalization
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Memory-efficient backpropagation through time
Gruslys, A., Munos, R., Danihelka, I., Lanctot, M., and Graves, A · 2016
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
Han, S., Mao, H., and Dally, W. J · 2016
Earlier work this paper cites.
Compression of neural machine translation models via pruning
See, A., Luong, M.-T., and Manning, C. D · 2016
Earlier work this paper cites.
DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs
Chen, L.-C., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A. L · 2017
Earlier work this paper cites.
Clipper: A low-latency online prediction serving system
Crankshaw, D., Wang, X., Zhou, G., Franklin, M. J., Gonzalez, J. E., and Stoica, I · 2017
Earlier work this paper cites.
The reversible residual network: Backpropagation without storing activations
Gomez, A. N., Ren, M., Urtasun, R., and Grosse, R. B · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
Population based training of neural networks
Jaderberg, M., Dalibard, V., Osindero, S., Czarnecki, W. M., Donahue, J., Razavi, A., Vinyals, O., Green, T., Dunning, I., Simonyan, K., et al · 2017
Earlier work this paper cites.
In-datacenter performance analysis of a tensor processing unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al · 2017
Earlier work this paper cites.
Pruning filters for efficient convnets
Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H. P · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
The expressive power of neural networks: A view from the width
Lu, Z., Pu, H., Wang, F., Hu, Z., and Wang, L · 2017
Earlier work this paper cites.
ThiNet: A filter level pruning method for deep neural network compression
Luo, J.-H., Wu, J., and Lin, W · 2017
Cited alongside, same era.
Building an AI chip saved Google from building a dozen new data centers
Metz, C · 2017
Cited alongside, same era.
On the expressive power of deep neural networks
Raghu, M., Poole, B., Kleinberg, J., Ganguli, S., and Dickstein, J. S · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2017
Cited alongside, same era.
Neural architecture search with reinforcement learning
Zoph, B. and Le, Q. V · 2017
Cited alongside, same era.
Efficient training of BERT by progressively stacking
Gong, L., He, D., Li, Z., Qin, T., Wang, L., and Liu, T · 2019
Later among the works it cites.
Data-efficient image recognition with contrastive predictive coding
Hénaff, O. J., Srinivas, A., Fauw, J. D., Razavi, A., Doersch, C., Eslami, S. M. A., and van den Oord, A · 2019
Later among the works it cites.
Are sixteen heads really better than one?
Michel, P., Levy, O., and Neubig, G · 2019
Later among the works it cites.
Fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reconciling modern machine learning and the bias-variance trade-off
Belkin, M., Hsu, D., Ma, S., and Mandal, S · 2018
Cited alongside, same era.
Efficient neural audio synthesis
Kalchbrenner, N., Elsen, E., Simonyan, K., Noury, S., Casagrande, N., Lockhart, E., Stimberg, F., Oord, A. v. d., Dieleman, S., and Kavukcuoglu, K · 2018
Cited alongside, same era.
Learning sparse neural networks through L 0 {L}_{0} regularization
Louizos, C., Welling, M., and Kingma, D. P · 2018
Cited alongside, same era.
An empirical model of large-batch training
McCandlish, S., Kaplan, J., Amodei, D., and Team, O. D · 2018
Cited alongside, same era.
Scaling neural machine translation
Ott, M., Edunov, S., Grangier, D., and Auli, M · 2018
Cited alongside, same era.
Mesh-TensorFlow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., et al · 2018
Cited alongside, same era.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Later among the works it cites.
Schwartz, R., Dodge, J., Smith, N. A., and Etzioni, O · 2019
Later among the works it cites.
Megatron-LM: Training multi-billion parameter language models using GPU model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Later among the works it cites.
Patient knowledge distillation for BERT model compression
Sun, S., Cheng, Y., Gan, Z., and Liu, J · 2019
Later among the works it cites.
EfficientNet: Rethinking model scaling for convolutional neural networks
Tan, M. and Le, Q. V · 2019
Later among the works it cites.
Compressing RNNs for IOT devices by 15-38x using kronecker products
Thakker, U., Beu, J., Gope, D., Zhou, C., Fedorov, I., Dasika, G., and Mattina, M · 2019
Later among the works it cites.
Well-read students learn better: The impact of student initialization on knowledge distillation
Turc, I., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Later among the works it cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Voita, E., Talbot, D., Moiseev, F., Sennrich, R., and Titov, I · 2019
Later among the works it cites.
ELECTRA: Pre-training text encoders as discriminators rather than generators
Clark, K., Luong, M.-T., Le, Q. V., and Manning, C. D · 2020
Closest in time.
Rigging the lottery: Making all tickets winners
Evci, U., Gale, T., Menick, J., Castro, P. S., and Elsen, E · 2020
Closest in time.
Checkmate: Breaking the memory wall with optimal tensor rematerialization
Jain, P., Jain, A., Nrusimha, A., Gholami, A., Abbeel, P., Keutzer, K., Stoica, I., and Gonzalez, J. E · 2020
Closest in time.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Closest in time.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2020
Closest in time.
Soft threshold weight reparameterization for learnable sparsity
Kusupati, A., Ramanujan, V., Somani, R., Wortsman, M., Jain, P., Kakade, S., and Farhadi, A · 2020
Closest in time.
ALBERT: A lite BERT for self-supervised learning of language representations
Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R · 2020
Closest in time.
Budgeted training: Rethinking deep neural network training under resource constraints
Li, M., Yumer, E., and Ramanan, D · 2020
Closest in time.
Q-BERT: Hessian based ultra low precision quantization of BERT
Shen, S., Dong, Z., Ye, J., Ma, L., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2020
Closest in time.
Large batch optimization for deep learning: Training BERT in 76 minutes
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.-J · 2020
Closest in time.