Fetching the paper…
Reading the bibliography…
In this work, we study the large-scale pretraining of BERT-Large with differentially private SGD (DP-SGD).
Our data, ourselves: Privacy via distributed noise generation
Cynthia Dwork, Krishnaram Kenthapadi, Frank McSherry, Ilya Mironov, and Moni Naor · 2006
Earlier work this paper cites.
Calibrating noise to sensitivity in private data analysis
Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith · 2006
Earlier work this paper cites.
The Algorithmic Foundations of Differential Privacy
Cynthia Dwork and Aaron Roth · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Deep learning with differential privacy
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang · 2016
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Earlier work this paper cites.
Riemannian approach to batch normalization
Minhyung Cho and Jaehyung Lee · 2017
Earlier work this paper cites.
Fixing weight decay regularization in Adam
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Rényi differential privacy
Ilya Mironov · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
XLA: Optimizing compiler for machine learning
XLA team and collaborators · 2017
Earlier work this paper cites.
Large scale distributed neural network training through online distillation
Rohan Anil, Gabriel Pereyra, Alexandre Passos, Róbert Ormándi, George E. Dahl, and Geoffrey E. Hinton · 2018
Earlier work this paper cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Earlier work this paper cites.
Compiling machine learning programs via high-level tracing
Roy Frostig, Matthew Johnson, and Chris Leary · 2018
Earlier work this paper cites.
Shampoo: Preconditioned stochastic tensor optimization
Vineet Gupta, Tomer Koren, and Yoram Singer · 2018
Cited alongside, same era.
Learning differentially private language models without losing accuracy
H. Brendan McMahan, Daniel Ramage, Kunal Talwar, and Li Zhang · 2018
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2018
Cited alongside, same era.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Cited alongside, same era.
Memory efficient adaptive optimization
Rohan Anil, Vineet Gupta, Tomer Koren, and Yoram Singer · 2019
Cited alongside, same era.
Differentially private learning with adaptive clipping
Galen Andrew, Om Thakkar, H Brendan McMahan, and Swaroop Ramaswamy · 2019
Stochastic optimization with laggard data pipelines
Naman Agarwal, Rohan Anil, Tomer Koren, Kunal Talwar, and Cyril Zhang · 2020
Later among the works it cites.
Scalable Second Order Optimization for Deep Learning
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer · 2020
Later among the works it cites.
Robust differentially private training of deep neural networks
Ali Davody, David Ifeoluwa Adelani, Thomas Kleinbauer, and Dietrich Klakow · 2020
Later among the works it cites.
Does learning require memorization? A short tale about a long tail
Vitaly Feldman · 2020
Later among the works it cites.
Flax: A neural network library and ecosystem for JAX, 2020
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, and Marc van Zee · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Robust bi-tempered logistic loss based on Bregman divergences
Ehsan Amid, Manfred K. Warmuth, Rohan Anil, and Tomer Koren · 2019
Cited alongside, same era.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song · 2019
Cited alongside, same era.
Faster neural network training with data echoing
Dami Choi, Alexandre Passos, Christopher J Shallue, and George E Dahl · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Norm matters: efficient and accurate normalization schemes in deep networks
Elad Hoffer, Ron Banner, Itay Golan, and Daniel Soudry · 2019
Cited alongside, same era.
The Ethical Algorithm: The Science of Socially Aware Algorithm Design
Michael Kearns and Aaron Roth · 2019
Cited alongside, same era.
Learning rate adaptation for differentially private learning
Antti Koskela and Antti Honkela · 2020
Later among the works it cites.
Enabling fast differentially private SGD via just-in-time compilation and vectorization
Pranav Subramani, Nicholas Vadivelu, and Gautam Kamath · 2020
Later among the works it cites.
Locoprop: Enhancing backprop via local loss optimization
Ehsan Amid, Rohan Anil, and Manfred K Warmuth · 2021
Closest in time.
When is memorization of irrelevant training data necessary for high-accuracy learning?
Gavin Brown, Mark Bun, Vitaly Feldman, Adam Smith, and Kunal Talwar · 2021
Closest in time.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Closest in time.
Adamp: Slowing down the slowdown for momentum optimizers on scale-invariant weights
Byeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han, Sangdoo Yun, Gyuwan Kim, Youngjung Uh, and Jung-Woo Ha · 2021
Closest in time.
Learning and evaluating a differentially private pre-trained language model
Shlomo Hoory, Amir Feder, Avichai Tendler, Alon Cohen, Sofia Erell, Itay Laish, Hootan Nakhost, Uri Stemmer, Ayelet Benjamini, Avinatan Hassidim, and Yossi Matias · 2021
Closest in time.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2021
Closest in time.
Fnet: Mixing tokens with Fourier transforms
James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, and Santiago Ontanon · 2021
Closest in time.
A large batch optimizer reality check: Traditional, generic optimizers suffice across batch sizes
Zachary Nado, Justin Gilmer, Christopher J. Shallue, Rohan Anil, and George E. Dahl · 2021
Closest in time.
MLP-Mixer: An all-MLP architecture for vision
Ilya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic, and Alexey Dosovitskiy · 2021
Closest in time.
Gspmd: General and scalable parallelization for ML computation graphs
Yuanzhong Xu, HyoukJoong Lee, Dehao Chen, Blake Hechtman, Yanping Huang, Rahul Joshi, Maxim Krikun, Dmitry Lepikhin, Andy Ly, Marcello Maggioni, et al · 2021
Closest in time.