Fetching the paper…
Reading the bibliography…
We train a recurrent neural network language model using a distributed, on-device learning framework called federated learning for the purpose of next-word prediction in a virtual keyboard for smartphones.
“A method for solving the convex programming problem with convergence rate o ( 1 / k 2 ) o(1/{k}^{2}) ,”
Yurii Nesterov, · 1983
Earlier work this paper cites.
“A learning algorithm for continually running fully recurrent neural networks,”
R. J. Williams and D. Zipser, · 1989
Earlier work this paper cites.
“Finite-state transducers in language and speech processing,”
Mehryar Mohri, · 1997
Earlier work this paper cites.
“Long short-term memory,”
S. Hochreiter and J. Schmidhuber, · 1997
Earlier work this paper cites.
“A neural probabilistic language model,”
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin, · 2003
Earlier work this paper cites.
“Differential privacy,”
Cynthia Dwork, · 2006
Earlier work this paper cites.
“Bayesian language model interpolation for mobile speech input,”
Cyril Allauzen and Michael Riley, · 2011
Earlier work this paper cites.
“Adaptive subgradient methods for online learning and stochastic optimization,”
John Duchi, Elad Hazan, and Yoram Singer, · 2011
Earlier work this paper cites.
“Consumer data privacy in a networked world: A framework for protecting privacy and promoting innovation in the global digital economy,” 01 2013
The White House, · 2013
Earlier work this paper cites.
“One billion word benchmark for measuring progress in statistical language modeling,”
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson, · 2014
Earlier work this paper cites.
“Learning phrase representations using RNN encoder-decoder for statistical machine translation,”
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, · 2014
Cited alongside, same era.
“Adam: A Method for Stochastic Optimization,”
D. P. Kingma and J. Ba, · 2014
Cited alongside, same era.
“Long short term memory neural network for keyboard gesture decoding,”
Ouais Alsharif, Tom Ouyang, Françoise Beaufays, Shumin Zhai, Thomas Breuel, and Johan Schalkwyk, · 2015
Cited alongside, same era.
“Privacy-preserving deep learning,”
Reza Shokri and Vitaly Shmatikov, · 2015
Cited alongside, same era.
“Character-aware neural language models,”
Yoon Kim, Yacine Jernite, David Sontag, and Alexander M. Rush, · 2016
Cited alongside, same era.
“Exploring the limits of language modeling,” 2016
Rafal Jozefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and Yonghui Wu, · 2016
“Outrageously large neural networks: The sparsely-gated mixture-of-experts layer,”
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean, · 2017
Later among the works it cites.
“Communication-efficient learning of deep networks from decentralized data,”
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas, · 2017
Later among the works it cites.
“Neural networks for text correction and completion in keyboard decoding,”
Shaona Ghosh and Per Ola Kristensson, · 2017
Later among the works it cites.
“Learning differentially private language models without losing accuracy,”
H. Brendan McMahan, Daniel Ramage, Kunal Talwar, and Li Zhang, · 2017
Later among the works it cites.
“LSTM: A search space odyssey,”
Klaus Greff, Rupesh Kumar Srivastava, Jan Koutník, Bas R. Steunebrink, and Jürgen Schmidhuber, · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Tensorflow: A system for large-scale machine learning,”
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek Gordon Murray, Benoit Steiner, Paul A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng, · 2016
Cited alongside, same era.
“Tying word vectors and word classifiers: A loss framework for language modeling,”
Hakan Inan, Khashayar Khosravi, and Richard Socher, · 2016
Cited alongside, same era.
“Practical secure aggregation for federated learning on user-held data,”
Keith Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone, H. Brendan McMahan, Sarvar Patel, Daniel Ramage, Aaron Segal, and Karn Seth, · 2016
Cited alongside, same era.
“Mobile keyboard input decoding with finite-state transducers,”
Tom Ouyang, David Rybach, Françoise Beaufays, and Michael Riley, · 2017
Cited alongside, same era.
Later among the works it cites.
“Using the output embedding to improve language models,”
Ofir Press and Lior Wolf, · 2017
Later among the works it cites.
“Technology device ownership: 2015,” http://www.pewinternet.org/2015/10/29/technology-device-ownership-2015/
Monica Anderson, · 2018
Closest in time.
“On-device neural language model based word prediction,”
Seunghak Yu, Nilesh Kulkarni, Haejun Lee, and Jihie Kim, · 2018
Closest in time.
“Federated learning: Collaborative machine learning without centralized training data,” https://ai.googleblog.com/2017/04/federated-learning-collaborative.html
Brendan McMahan and Daniel Ramage, · 2018
Closest in time.