Fetching the paper…
Reading the bibliography…
Scaling up the convolutional neural network (CNN) size (e.g., width, depth, etc.) is known to effectively improve model accuracy.
Iterative solution of nonlinear equations in several variables , vol. 30
Ortega, J. M., W. C. Rheinboldt · 1970
Earlier work this paper cites.
Parallel and distributed computation: numerical methods , vol. 23
Bertsekas, D. P., J. N. Tsitsiklis · 1989
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Qian, N · 1999
Earlier work this paper cites.
Model compression
Buciluǎ, C., R. Caruana, A. Niculescu-Mizil · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A., G. Hinton, et al · 2009
Earlier work this paper cites.
Proximal alternating minimization and projection methods for nonconvex problems: An approach based on the kurdyka-łojasiewicz inequality
Attouch, H., J. Bolte, P. Redont, et al · 2010
Earlier work this paper cites.
A unified convergence analysis of block successive minimization methods for nonsmooth optimization
Razaviyayn, M., M. Hong, Z.-Q. Luo · 2013
Earlier work this paper cites.
Proximal alternating linearized minimization for nonconvex and nonsmooth problems
Bolte, J., S. Sabach, M. Teboulle · 2014
Earlier work this paper cites.
Do deep nets really need to be deep?
Ba, J., R. Caruana · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Romero, A., N. Ballas, S. E. Kahou, et al · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P., J. Ba · 2014
Earlier work this paper cites.
Coordinate descent algorithms
Wright, S. J · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., O. Vinyals, J. Dean · 2015
Earlier work this paper cites.
Han, S., H. Mao, W. J. Dally · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., X. Zhang, S. Ren, et al · 2016
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
McMahan, H. B., E. Moore, D. Ramage, et al · 2016
Earlier work this paper cites.
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and< 0.5 mb model size
Iandola, F. N., S. Han, M. W. Moskewicz, et al · 2016
Earlier work this paper cites.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
Alistarh, D., D. Grubic, J. Li, et al · 2017
Earlier work this paper cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Lin, Y., S. Han, H. Mao, et al · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G., M. Zhu, B. Chen, et al · 2017
Earlier work this paper cites.
Darts: Differentiable architecture search
Liu, H., K. Simonyan, Y. Yang · 2018
Earlier work this paper cites.
Distributed learning of deep neural network over multiple agents
Gupta, O., R. Raskar · 2018
Earlier work this paper cites.
Split learning for health: Distributed deep learning without sharing raw patient data
Vepakomma, P., O. Gupta, T. Swedish, et al · 2018
Earlier work this paper cites.
Cinic-10 is not imagenet or cifar-10
Darlow, L. N., E. J. Crowley, A. Antoniou, et al · 2018
Earlier work this paper cites.
signsgd: Compressed optimisation for non-convex problems
Bernstein, J., Y.-X. Wang, K. Azizzadenesheli, et al · 2018
Cited alongside, same era.
Gradient sparsification for communication-efficient distributed optimization
Wangni, J., J. Wang, J. Liu, et al · 2018
Cited alongside, same era.
Communication compression for decentralized training
Tang, H., S. Gan, C. Zhang, et al · 2018
Cited alongside, same era.
Atomo: Communication-efficient learning via atomic sparsification
Wang, H., S. Sievert, S. Liu, et al · 2018
Cited alongside, same era.
Deep mutual learning
Zhang, Y., T. Xiang, T. M. Hospedales, et al · 2018
Cited alongside, same era.
Large scale distributed neural network training through online distillation
Fbnet: Hardware-aware efficient convnet design via differentiable neural architecture search
Wu, B., X. Dai, P. Zhang, et al · 2019
Later among the works it cites.
A sufficient condition for convergences of adam and rmsprop
Zou, F., L. Shen, Z. Jie, et al · 2019
Later among the works it cites.
An exponential learning rate schedule for deep learning
Li, Z., S. Arora · 2019
Later among the works it cites.
Milenas: Efficient neural architecture search via mixed-level reformulation
He, C., H. Ye, L. Shen, et al · 2020
Closest in time.
Federated learning with matched averaging
Wang, H., M. Yurochkin, Y. Sun, et al · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anil, R., G. Pereyra, A. Passos, et al · 2018
Cited alongside, same era.
Collaborative learning for deep neural networks
Song, G., W. Chai · 2018
Cited alongside, same era.
Jeong, E., S. Oh, H. Kim, et al · 2018
Cited alongside, same era.
Knowledge distillation by on-the-fly native ensemble
Zhu, X., S. Gong, et al · 2018
Cited alongside, same era.
Amc: Automl for model compression and acceleration on mobile devices
He, Y., J. Lin, Z. Liu, et al · 2018
Cited alongside, same era.
Netadapt: Platform-aware neural network adaptation for mobile applications
Yang, T.-J., A. Howard, B. Chen, et al · 2018
Cited alongside, same era.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
Zhang, X., X. Zhou, M. Lin, et al · 2018
Cited alongside, same era.
He, C., S. Li, J. So, et al · 2020
Closest in time.
Adaptive federated optimization
Reddi, S., Z. Charles, M. Zaheer, et al · 2020
Closest in time.
Fednas: Federated deep learning via neural architecture search
He, C., M. Annavaram, S. Avestimehr · 2020
Closest in time.
Turbo-aggregate: Breaking the quadratic aggregation barrier in secure federated learning
So, J., B. Guler, A. S. Avestimehr · 2020
Closest in time.
Byzantine-resilient secure federated learning
So, J., B. Guler, A. S. Avestimehr · 2020
Closest in time.
Secure aggregation with heterogeneous quantization in federated learning
Elkordy, A. R., A. S. Avestimehr · 2020
Closest in time.
Ensemble distillation for robust model fusion in federated learning
Lin, T., L. Kong, S. U. Stich, et al · 2020
Closest in time.
Distributed distillation for on-device learning
Bistritz, I., A. Mann, N. Bambos · 2020
Closest in time.
Asynchronous edge learning using cloned knowledge distillation
Lee, S.-h., K. Yoo, N. Kwak · 2020
Closest in time.
Distilled one-shot federated learning
Zhou, Y., G. Pu, X. Ma, et al · 2020
Closest in time.
Federated model distillation with noise-free differential privacy
Sun, L., L. Lyu · 2020
Closest in time.
Adaptive distillation for decentralized learning from heterogeneous clients
Ma, J., R. Yonetani, Z. Iqbal · 2020
Closest in time.
Itahara, S., T. Nishio, Y. Koda, et al · 2020
Closest in time.
Model-agnostic round-optimal federated learning via knowledge transfer
Li, Q., B. He, D. Song · 2020
Closest in time.
Hydra: Preserving ensemble diversity for model distillation
Tran, L., B. S. Veeling, K. Roth, et al · 2020
Closest in time.
Tackling the objective inconsistency problem in heterogeneous federated optimization
Wang, J., Q. Liu, H. Liang, et al · 2020
Closest in time.
Shi, B., W. J. Su, M. I. Jordan · 2020
Closest in time.
Measuring the algorithmic efficiency of neural networks
Hernandez, D., T. B. Brown · 2020
Closest in time.
Attack of the tails: Yes, you really can backdoor federated learning
Wang, H., K. Sreenivasan, S. Rajput, et al · 2020
Closest in time.