Fetching the paper…
Reading the bibliography…
We observe that the traditional use of DP with the Adam optimizer introduces a bias in the second moment estimation, due to the addition of independent noise in the gradient computation.
On the convergence of adam and beyond, 2019
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 1904
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2018
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Cifar-10 (canadian institute for advanced research)
Alex Krizhevsky, Vinod Nair, and Geoffrey Hinton
Cited in the paper.
Large language models can be strong differentially private learners, 2021
Xuechen Li, Florian Tramèr, Percy Liang, and Tatsunori Hashimoto · 2021
Later among the works it cites.
Heavy-tailed noise does not explain the gap between SGD and adam, but sign descent might
Frederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, and Mark Schmidt · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…