Security policies and security models
J. A. Goguen and J. Meseguer · 1982
Earlier work this paper cites.
Input partitioning to mixture of experts
B. Tang, M. I. Heywood, and M. Shepherd · 2002
Earlier work this paper cites.
Calibrating noise to sensitivity in private data analysis
C. Dwork, F. McSherry, K. Nissim, and A. Smith · 2006
Earlier work this paper cites.
No free lunch in data privacy
D. Kifer and A. Machanavajjhala · 2011
Earlier work this paper cites.
The algorithmic foundations of differential privacy
C. Dwork, A. Roth, et al · 2014
Earlier work this paper cites.
Mixture of experts: a literature survey
S. Masoudnia and R. Ebrahimpour · 2014
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Y. Cao and J. Yang · 2015
Earlier work this paper cites.
Deep learning with differential privacy
M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang · 2016
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas · 2017
Earlier work this paper cites.
Semi-supervised knowledge transfer for deep learning from private training data
N. Papernot, M. Abadi, Ú. Erlingsson, I. Goodfellow, and K. Talwar · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Original
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Privacy-preserving prediction
C. Dwork and V. Feldman · 2018
Earlier work this paper cites.
Averaging weights leads to wider optima and better generalization
P. Izmailov, D. Podoprikhin, T. Garipov, D. Vetrov, and A. G. Wilson · 2018
Earlier work this paper cites.
Learning differentially private recurrent language models
H. B. McMahan, D. Ramage, K. Talwar, and L. Zhang · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Eli5: Long form question answering
Original
A. Fan, Y. Jernite, E. Perez, D. Grangier, J. Weston, and M. Auli · 2019
Earlier work this paper cites.
Making ai forget you: Data deletion in machine learning
A. Ginart, M. Guan, G. Valiant, and J. Y. Zou · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly · 2019
Earlier work this paper cites.
CTRL: A conditional transformer language model for controllable generation
Original
N. S. Keskar, B. McCann, L. R. Varshney, C. Xiong, and R. Socher · 2019
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Original
U. Khandelwal, O. Levy, D. Jurafsky, L. Zettlemoyer, and M. Lewis · 2019
Earlier work this paper cites.