Fetching the paper…
Reading the bibliography…
State-of-the-art training algorithms for deep learning models are based on stochastic gradient descent (SGD).
Autoslim: Towards one-shot architecture search for channel numbers
Jiahui Yu and Thomas Huang · 1903
Earlier work this paper cites.
The extragradient method for finding saddle points and other problems
Galina M Korpelevich · 1976
Earlier work this paper cites.
Almost sure convergence of dropout algorithms for neural networks
Albert Senen-Cerda and Jaron Sanders · 2002
Earlier work this paper cites.
Introductory Lectures on Convex Optimization , volume 87 of Springer Science & Business Media
Yurii Nesterov · 2004
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Yoshua Bengio, Pascal Lamblin, Dan Popovici, Hugo Larochelle, et al · 2007
Earlier work this paper cites.
Asymptotic convergence rate of dropout on shallow linear neural networks
Albert Senen-Cerda and Jaron Sanders · 2012
Earlier work this paper cites.
Report on the 11th iwslt evaluation campaign
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, and Marcello Federico · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Wide & deep learning for recommender systems
Heng-Tze Cheng, Levent Koc, Jeremiah Harmsen, Tal Shaked, Tushar Chandra, Hrishi Aradhye, Glen Anderson, Greg Corrado, Wei Chai, Mustafa Ispir, et al · 2016
Earlier work this paper cites.
Communication quantization for data-parallel training of deep neural networks
Nikoli Dryden, Tim Moon, Sam Ade Jacobs, and Brian Van Essen · 2016
Earlier work this paper cites.
Training skinny deep neural networks with iterative hard thresholding methods
Xiaojie Jin, Xiaotong Yuan, Jiashi Feng, and Shuicheng Yan · 2016
Earlier work this paper cites.
Sparse communication for distributed gradient descent
Alham Fikri Aji and Kenneth Heafield · 2017
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic · 2017
Earlier work this paper cites.
Findings of the 2017 conference on machine translation (wmt17)
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, et al · 2017
Earlier work this paper cites.
DeepFM: A factorization-machine based neural network for CTR prediction
Huifeng Guo, Ruiming TANG, Yunming Ye, Zhenguo Li, and Xiuqiang He · 2017
Cited alongside, same era.
meprop: Sparsified back propagation for accelerated deep learning with reduced overfitting
Xu Sun, Xuancheng Ren, Shuming Ma, and Houfeng Wang · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Coordinating filters for faster deep neural networks
Wei Wen, Cong Xu, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2017
Cited alongside, same era.
To prune, or not to prune: exploring the efficacy of pruning for model compression
Michael Zhu and Suyog Gupta · 2017
Cited alongside, same era.
The convergence of sparsified gradient methods
Distributed learning of deep neural networks using independent subnet training
Binhang Yuan, Cameron R. Wolfe, Chen Dun, Yuxin Tang, Anastasios Kyrillidis, and Christopher M. Jermaine · 2019
Later among the works it cites.
On the convergence of SGD with biased gradients
Ahmad Ajalloeian and Sebastian U. Stich · 2020
Later among the works it cites.
A tight convergence analysis for stochastic gradient descent with delayed updates
Yossi Arjevani, Ohad Shamir, and Nathan Srebro · 2020
Later among the works it cites.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin · 2020
Later among the works it cites.
Characterising bias in compressed models
Sara Hooker, Nyalleng Moorosi, Gregory Clark, Samy Bengio, and Emily Denton · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dan Alistarh, Torsten Hoefler, Mikael Johansson, Nikola Konstantinov, Sarit Khirirat, and Cedric Renggli · 2018
Cited alongside, same era.
Optimization methods for large-scale machine learning
L. Bottou, F. Curtis, and J. Nocedal · 2018
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Cited alongside, same era.
On first-order meta-learning algorithms
Alex Nichol, Joshua Achiam, and John Schulman · 2018
Cited alongside, same era.
Sparsified sgd with memory
Sebastian U Stich, Jean-Baptiste Cordonnier, and Martin Jaggi · 2018
Cited alongside, same era.
Once-for-all: Train one network and specialize it for efficient deployment
Han Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang, and Song Han · 2019
Cited alongside, same era.
Dynamic model pruning with feedback
Tao Lin, Sebastian U Stich, Luis Barba, Daniil Dmitriev, and Martin Jaggi · 2019
Cited alongside, same era.
Later among the works it cites.
Ensemble distillation for robust model fusion in federated learning
Tao Lin, Lingjing Kong, Sebastian U Stich, and Martin Jaggi · 2020
Later among the works it cites.
The error-feedback framework: Better rates for sgd with delayed gradients and compressed updates
Sebastian U Stich and Sai Praneeth Karimireddy · 2020
Later among the works it cites.
BigNAS: Scaling up neural architecture search with big single-stage models
Jiahui Yu, Pengchong Jin, Hanxiao Liu, Gabriel Bender, Pieter-Jan Kindermans, Mingxing Tan, Thomas Huang, Xiaodan Song, Ruoming Pang, and Quoc Le · 2020
Later among the works it cites.
Implicit gradient alignment in distributed and federated learning
Yatin Dandi, Luis Barba, and Martin Jaggi · 2021
Closest in time.
Neural tangent kernel: convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2021
Closest in time.
On the convergence of shallow neural network training with randomly masked neurons
Fangshuo Liao and Anastasios Kyrillidis · 2021
Closest in time.
Ac/dc: Alternating compressed/decompressed training of deep neural networks
Alexandra Peste, Eugenia Iofinova, Adrian Vladu, and Dan Alistarh · 2021
Closest in time.
Critical parameters for scalable distributed learning with large batches and asynchronous updates
Sebastian Stich, Amirkeivan Mohtashami, and Martin Jaggi · 2021
Closest in time.
Pufferfish: Communication-efficient models at no extra cost
Hongyi Wang, Saurabh Agarwal, and Dimitris Papailiopoulos · 2021
Closest in time.
Dynamic sparsity neural networks for automatic speech recognition
Zhaofeng Wu, Ding Zhao, Qiao Liang, Jiahui Yu, Anmol Gulati, and Ruoming Pang · 2021
Closest in time.