Fetching the paper…
Reading the bibliography…
We propose algorithms to train production-quality n-gram language models using federated learning.
Approximating probabilistic models as weighted finite automata
Ananda Theertha Suresh, Michael Riley, Brian Roark, and Vlad Schogol. 2019a · 1905
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
Stanley F Chen and Joshua Goodman. 1999 · 1999
Earlier work this paper cites.
Byte pair encoding: A text compression scheme that accelerates pattern matching
Yusuxke Shibata, Takuya Kida, Shuichi Fukamachi, Masayuki Takeda, Ayumi Shinohara, Takeshi Shinohara, and Setsuo Arikawa. 1999 · 1999
Earlier work this paper cites.
Intelligent selection of language model training data
Robert C. Moore and William Lewis. 2010 · 2010
Earlier work this paper cites.
A filter-based algorithm for efficient composition of finite-state transducers
Cyril Allauzen, Michael Riley, and Johan Schalkwyk. 2011 · 2011
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima. 2012 · 2012
Earlier work this paper cites.
Language model capitalization
Françoise Beaufays and Brian Strope. 2013 · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
The algorithmic foundations of differential privacy
Cynthia Dwork, Aaron Roth, et al. 2014 · 2014
Earlier work this paper cites.
Long short-term memory recurrent neural network architectures for large scale acoustic modeling
Haşim Sak, Andrew Senior, and Françoise Beaufays. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Cited alongside, same era.
Testing closeness with unequal sized samples
Bhaswar Bhattacharya and Gregory Valiant. 2015 · 2015
Cited alongside, same era.
Deep learning with differential privacy
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. 2016 · 2016
Cited alongside, same era.
Federated optimization: Distributed machine learning for on-device intelligence
Jakub Konečnỳ, H Brendan McMahan, Daniel Ramage, and Peter Richtárik. 2016 · 2016
Cited alongside, same era.
Federated learning: Strategies for improving communication efficiency
Jakub Konečný, H. Brendan McMahan, Felix X. Yu, Peter Richtarik, Ananda Theertha Suresh, and Dave Bacon. 2016 · 2016
Cited alongside, same era.
Factorization tricks for LSTM networks
Oleksii Kuchaiev and Boris Ginsburg. 2017 · 2017
Later among the works it cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Agüera y Arcas. 2017 · 2017
Later among the works it cites.
Federated learning: Collaborative machine learning without centralized training data
Brendan McMahan and Daniel Ramage. 2017 · 2017
Later among the works it cites.
Mobile keyboard input decoding with finite-state transducers
Tom Ouyang, David Rybach, Françoise Beaufays, and Michael Riley. 2017 · 2017
Later among the works it cites.
cpsgd: Communication-efficient and differentially-private distributed sgd
Naman Agarwal, Ananda Theertha Suresh, Felix Yu, Sanjiv Kumar, and Brendan McMahan. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
Practical secure aggregation for privacy-preserving machine learning
Keith Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone, H. Brendan McMahan, Sarvar Patel, Daniel Ramage, Aaron Segal, and Karn Seth. 2017 · 2017
Cited alongside, same era.
Lstm: A search space odyssey
Klaus Greff, Rupesh K Srivastava, Jan Koutník, Bas R Steunebrink, and Jürgen Schmidhuber. 2017 · 2017
Cited alongside, same era.
Transliterated mobile keyboard input via weighted finite-state transducers
Lars Hellsten, Brian Roark, Prasoon Goyal, Cyril Allauzen, Francoise Beaufays, Tom Ouyang, Michael Riley, and David Rybach. 2017 · 2017
Cited alongside, same era.
Residual LSTM: design of a deep recurrent architecture for distant speech recognition
Jaeyoung Kim, Mostafa El-Khamy, and Jungwon Lee. 2017 · 2017
Cited alongside, same era.
State-of-the-art speech recognition with sequence-to-sequence models
Chung-Cheng Chiu, Tara N Sainath, Yonghui Wu, Rohit Prabhavalkar, Patrick Nguyen, Zhifeng Chen, Anjuli Kannan, Ron J Weiss, Kanishka Rao, Ekaterina Gonina, et al. 2018 · 2018
Later among the works it cites.
Federated learning for mobile keyboard prediction
Andrew Hard, Kanishka Rao, Rajiv Mathews, Françoise Beaufays, Sean Augenstein, Hubert Eichner, Chloé Kiddon, and Daniel Ramage. 2018 · 2018
Later among the works it cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Later among the works it cites.
Learning differentially private recurrent language models
Brendan McMahan, Daniel Ramage, Kunal Talwar, and Li Zhang. 2018 · 2018
Later among the works it cites.
Distilling weighted finite automata from arbitrary probabilistic models
Ananda Theertha Suresh, Brian Roark, Michael Riley, and Vlad Schogol. 2019b · 2019
Closest in time.
Long short term memory neural network for keyboard gesture decoding
Ouais Alsharif, Tom Ouyang, Françoise Beaufays, Shumin Zhai, Thomas Breuel, and Johan Schalkwyk. 2015 · 2080
Closest in time.