Fetching the paper…
Reading the bibliography…
We show how Zipf's Law can be used to scale up language modeling (LM) to take advantage of more training data and more GPUs.
G. K. Zipf, “The Psycho-Biology of Language,” Linguistic Society of America , vol. 12, no. 3, pp. 196–210, 1936
1936
Earlier work this paper cites.
G. K. Zipf, Human Behaviour and the Principle of Least Effort: an Introduction to Human Ecology . Addison-Wesley, 1949
1949
Earlier work this paper cites.
K. W. Church and R. L. Mercer, “Introduction to the Special Issue on Computational Linguistics Using Large Corpora,” Computational Linguistics , vol. 19, no. 1, pp. 1–24, Mar. 1993. [Online]. Available: http://dl.acm.org/citation.cfm?id=972450.972452
1993
Earlier work this paper cites.
M. Banko and E. Brill, “Scaling to very very large corpora for natural language disambiguation,” in 39th annual meeting on association for computational linguistics , 2001, pp. 26–33
2001
Earlier work this paper cites.
Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin, “A neural probabilistic language model,” Journal of machine learning research , vol. 3, no. Feb, pp. 1137–1155, 2003
2003
Earlier work this paper cites.
S. Bird, E. Klein, and E. Loper, Natural Language Processing with Python , 1st ed. O’Reilly Media, Inc., 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “ImageNet: A Large-Scale Hierarchical Image Database,” in CVPR09 , 2009
2009
Earlier work this paper cites.
2013
Earlier work this paper cites.
C. Buck, K. Heafield, and B. van Ooyen, “N-gram Counts and Language Models from the Common Crawl,” in Proceedings of the Language Resources and Evaluation Conference , Reykjavik, Iceland, May 2014. [Online]. Available: http://commoncrawl.org/
2014
Earlier work this paper cites.
J. McAuley, C. Targett, Q. Shi, and A. van den Hengel, “Image-Based Recommendations on Styles and Substitutes,” ser. SIGIR ’15. New York, NY, USA: ACM, 2015, pp. 43–52
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Cited alongside, same era.
M. Abadi et al. , “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, software available from tensorflow.org. [Online]. Available: https://www.tensorflow.org/
2015
Cited alongside, same era.
M.-S. Isabel, F.-C. Francesc, and C. Alvaro, “Large-Scale Analysis of Zipf’s Law in English Texts,” PLOS ONE , 2016
2016
Cited alongside, same era.
2016
Cited alongside, same era.
R. Kim, “Flashblade – Now 5X Bigger, 5X Faster,” https://blog.purestorage.com/flashblade-now-5x-bigger-5x-faster , 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Wikipedia, “Zipfś law,” https://en.wikipedia.org/wiki/Zipf%27s_law , (Accessed on 10/12/2018)
2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
I. Mitliagkas, C. Zhang, S. Hadjis, and C. Ré, “Asynchrony begets momentum, with an application to deep learning,” in Communication, Control, and Computing (Allerton) . IEEE, 2016, pp. 997–1004
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2018
Closest in time.
P. di Miceli, “Project Gutenberg,” https://www.gutenberg.org/ , 2018
2018
Closest in time.
BAIDU, “Baidu Tieba Log File,” https://tieba.baidu.com/index.html , 2018
2018
Closest in time.
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in Proc. of NAACL , 2018
2018
Closest in time.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” https://blog.openai.com/language-unsupervised/, 2018, accessed: 2018/10/15
2018
Closest in time.
R. Puri, R. Kirby, N. Yakovenko, and B. Catanzaro, “Large Scale Language Modeling: Converging on 40GB of Text in Four Hours,” ArXiv e-prints , Aug. 2018
2018
Closest in time.
2018
Closest in time.
2018
Closest in time.
Y. Ueno and K. Fukuda, “Technologies behind Distributed Deep Learning: AllReduce,” https://preferredresearch.jp/2018/07/10/technologies-behind-distributed-deep-learning-allreduce/ , 2018
2018
Closest in time.