Fetching the paper…
Reading the bibliography…
Methods for improving the efficiency of deep network training (i.e.
Training Data Subset Search with Ensemble Active Learning, November 2020
Chitta, K., Alvarez, J. M., Haussmann, E., and Farabet, C · 1905
Earlier work this paper cites.
Ensemble Distribution Distillation, November 2019
Malinin, A., Mlodozeniec, B., and Gales, M · 1905
Earlier work this paper cites.
Selection via Proxy: Efficient Data Selection for Deep Learning, October 2020
Coleman, C., Yeh, C., Mussmann, S., Mirzasoleiman, B., Bailis, P., Liang, P., Leskovec, J., and Zaharia, M · 1906
Earlier work this paper cites.
Patient Knowledge Distillation for BERT Model Compression, August 2019
Sun, S., Cheng, Y., Gan, Z., and Liu, J · 1908
Earlier work this paper cites.
Ensemble Knowledge Distillation for Learning Improved and Efficient Networks, April 2020
Asif, U., Tang, J., and Harrer, S · 1909
Earlier work this paper cites.
TinyBERT: Distilling BERT for Natural Language Understanding, October 2020
Jiao, X., Yin, Y., Shang, L., Jiang, X., Chen, X., Li, L., Wang, F., and Liu, Q · 1909
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter, February 2020
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 1910
Earlier work this paper cites.
The Early Phase of Neural Network Training
Frankle, J., Schwab, D. J., and Morcos, A. S · 2002
Earlier work this paper cites.
BERT-of-Theseus: Compressing BERT by Progressive Module Replacing, October 2020a
Xu, C., Zhou, W., Ge, T., Wei, F., and Zhou, M · 2002
Earlier work this paper cites.
Improving BERT Fine-Tuning via Self-Ensemble and Self-Distillation, February 2020b
Xu, Y., Qiu, X., Zhou, L., and Huang, X · 2002
Earlier work this paper cites.
FastBERT: a Self-distilling BERT with Adaptive Inference Time, April 2020
Liu, W., Zhou, P., Zhao, Z., Wang, Z., Deng, H., and Ju, Q · 2004
Earlier work this paper cites.
Knowledge Distillation: A Survey
Gou, J., Yu, B., Maybank, S. J., and Tao, D · 2006
Earlier work this paper cites.
Feldman, V. and Zhang, C · 2008
Earlier work this paper cites.
Dataset Cartography: Mapping and Diagnosing Datasets with Training Dynamics
Swayamdipta, S., Schwartz, R., Lourie, N., Wang, Y., Hajishirzi, H., Smith, N. A., and Choi, Y · 2009
Earlier work this paper cites.
TernaryBERT: Distillation-aware Ultra-low Bit BERT, October 2020
Zhang, W., Hou, L., Yin, Y., Shang, L., Chen, X., Jiang, X., and Liu, Q · 2009
Earlier work this paper cites.
Allen-Zhu, Z. and Li, Y · 2012
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2012
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., Dean, J., et al · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Cited alongside, same era.
SGDR: stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Cited alongside, same era.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Yim, J., Joo, D., Bae, J., and Kim, J · 2017
Cited alongside, same era.
Critical learning periods in deep networks
Achille, A., Rovere, M., and Soatto, S · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Improved knowledge distillation via teacher assistant
Mirzadeh, S. I., Farajtabar, M., Li, A., Levine, N., Matsukawa, A., and Ghasemzadeh, H · 2020
Later among the works it cites.
Self-training with noisy student improves imagenet classification
Xie, Q., Luong, M.-T., Hovy, E., and Le, Q. V · 2020
Later among the works it cites.
Knowledge distillation: A good teacher is patient and consistent
Beyer, L., Zhai, X., Royer, A., Markeeva, L., Anil, R., and Kolesnikov, A · 2021
Later among the works it cites.
On Evaluating and Improving the Efficiency of Deep Networks
Blalock, D., Carbin, M., Florescu, L., Frankle, J., Leavitt, M. L., Lee, T., Nadeem, M., Portes, J., Rao, N., Seguin, L., Stephenson, C., Tang, H., and Venigalla, A · 2021
Later among the works it cites.
No One Representation to Rule Them All: Overlapping Features of Training Methods
Gontijo-Lopes, R., Dauphin, Y., and Cubuk, E. D · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Born again neural networks
Furlanello, T., Lipton, Z., Tschannen, M., Itti, L., and Anandkumar, A · 2018
Cited alongside, same era.
Gradient Descent Happens in a Tiny Subspace, December 2018
Gur-Ari, G., Roberts, D. A., and Dyer, E · 2018
Cited alongside, same era.
Empirical Analysis of the Hessian of Over-Parametrized Neural Networks, May 2018
Sagun, L., Evci, U., Guney, V. U., Dauphin, Y., and Bottou, L · 2018
Cited alongside, same era.
Are All Training Examples Created Equal? An Empirical Study, November 2018
Vodrahalli, K., Li, K., and Malik, J · 2018
Cited alongside, same era.
Snapshot Distillation: Teacher-Student Optimization in One Generation
Yang, C., Xie, L., Su, C., and Yuille, A. L · 2018
Cited alongside, same era.
Time Matters in Regularizing Deep Networks: Weight Decay and Data Augmentation Affect Early Learning Dynamics, Matter Little Near Convergence
Golatkar, A. S., Achille, A., and Soatto, S · 2019
Cited alongside, same era.
Deep Learning on a Data Diet: Finding Important Examples Early in Training
Paul, M., Ganguli, S., and Dziugaite, G. K · 2021
Later among the works it cites.
A Survey of Deep Active Learning
Ren, P., Xiao, Y., Chang, X., Huang, P.-Y., Li, Z., Gupta, B. B., Chen, X., and Wang, X · 2021
Later among the works it cites.
seaborn: statistical data visualization
Waskom, M. L · 2021
Later among the works it cites.
Robust fine-tuning of zero-shot models
Wortsman, M., Ilharco, G., Kim, J. W., Li, M., Kornblith, S., Roelofs, R., Gontijo-Lopes, R., Hajishirzi, H., Farhadi, A., Namkoong, H., and Schmidt, L · 2021
Later among the works it cites.
Knowledge Distillation: Bad Models Can Be Good Role Models, March 2022
Kaplun, G., Malach, E., Nakkiran, P., and Shalev-Shwartz, S · 2022
Closest in time.
Knowledge Distillation for Efficient Sequences of Training Runs
Liu, X., Leonardi, A., Yu, L., Gilmer-Hill, C., Leavitt, M. L., and Frankle, J · 2022
Closest in time.
Merging Models with Fisher-Weighted Averaging, August 2022
Matena, M. and Raffel, C · 2022
Closest in time.
Mindermann, S., Brauner, J., Razzak, M., Sharma, M., Kirsch, A., Xu, W., Höltgen, B., Gomez, A. N., Morisot, A., Farquhar, S., and Gal, Y · 2022
Closest in time.
Beyond neural scaling laws: beating power law scaling via data pruning, August 2022
Sorscher, B., Geirhos, R., Shekhar, S., Ganguli, S., and Morcos, A. S · 2022
Closest in time.
Composer: A PyTorch Library for Efficient Neural Network Training, 2022
Tang, H., Rahman, R., Patel, M., Nadeem, M., Venigalla, A., Seguin, L., Khudia, D. S., Blalock, D., Leavitt, M. L., Shah, B., Bloxham, J., Racah, E., Jacobson, A., Stephenson, C., Saini, A., King, D., Knighton, J., Ehsani, A., Jariwala, K., Niklas, N., Lamp, A., Shastri, I., Trott, A., Cress, M., Lee, T., Cui, B., Portes, J., Florescu, L., Li, L., Zosa-Forde, J., Ivanchuk, V., Sardana, N., Blakeney, C., Carbin, M., Lupesko, H., Frankle, J., and Rao, N · 2022
Closest in time.
Efficient Knowledge Distillation from Model Checkpoints, October 2022
Wang, C., Yang, Q., Huang, R., Song, S., and Huang, G · 2022
Closest in time.
Wortsman, M., Ilharco, G., Gadre, S. Y., Roelofs, R., Gontijo-Lopes, R., Morcos, A. S., Namkoong, H., Farhadi, A., Carmon, Y., Kornblith, S., and Schmidt, L · 2022
Closest in time.