Fetching the paper…
Reading the bibliography…
Many recent breakthroughs in deep learning were achieved by training increasingly larger models on massive datasets.
A bridging model for parallel computation
Leslie G Valiant · 1990
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Michael I Jordan and Robert A Jacobs · 1994
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Algorithm 799: revolve: an implementation of checkpointing for the reverse or adjoint mode of computational differentiation
Andreas Griewank and Andrea Walther · 2000
Earlier work this paper cites.
A scalable content-addressable network
Sylvia Ratnasamy, Paul Francis, Mark Handley, Richard Karp, and Scott Shenker · 2001
Earlier work this paper cites.
Pastry: Scalable, decentralized object location, and routing for large-scale peer-to-peer systems
Antony Rowstron and Peter Druschel · 2001
Earlier work this paper cites.
Infinite mixtures of gaussian process experts
Carl E Rasmussen and Zoubin Ghahramani · 2002
Earlier work this paper cites.
A parallel mixture of svms for very large scale problems
Ronan Collobert, Samy Bengio, and Yoshua Bengio · 2002
Earlier work this paper cites.
Kademlia: A peer-to-peer information system based on the xor metric
Petar Maymounkov and David Mazieres · 2002
Earlier work this paper cites.
Looking up data in p2p systems
Hari Balakrishnan, M Frans Kaashoek, David Karger, Robert Morris, and Ion Stoica · 2003
Earlier work this paper cites.
Tapestry: A resilient global-scale overlay for service deployment
Ben Zhao, Ling Huang, Jeremy Stribling, Sean Rhea, Anthony Joseph, and John Kubiatowicz · 2003
Earlier work this paper cites.
Koorde: A simple degree-optimal distributed hash table
M Frans Kaashoek and David R Karger · 2003
Earlier work this paper cites.
Boinc: A system for public-resource computing and storage
David P Anderson · 2004
Earlier work this paper cites.
A denial-of-service resistant dht
Baruch Awerbuch and Christian Scheideler · 2007
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Folding@home and genome@home: Using distributed computing to tackle previously intractable problems in computational biology
Stefan Larson, Christopher Snow, Michael Shirts, and Vijay Pande · 2009
Earlier work this paper cites.
Hierarchical mixture of classification experts uncovers interactions between brain regions
Bangpeng Yao, Dirk Walther, Diane Beck, and Li Fei-Fei · 2009
Earlier work this paper cites.
Nonlinear models using dirichlet process mixtures
Babak Shahbaba and Radford Neal · 2009
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
A survey of dht security techniques
Guido Urdaneta, Guillaume Pierre, and Maarten Van Steen · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Folding research recruits unconventional help
Michael Gross · 2012
Earlier work this paper cites.
Real-world sybil attacks in bittorrent mainline dht
Liang Wang and Jussi Kangasharju · 2012
Earlier work this paper cites.
Learning factored representations in a deep mixture of experts
David Eigen, Marc’Aurelio Ranzato, and Ilya Sutskever · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Sybil nodes as a mitigation strategy against sybil attack
Zied Trifa and Maher Khemakhem · 2014
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Atlas@ home: harnessing volunteer computing for hep
C Adam-Bourdarios, D Cameron, A Filipčič, E Lancon, Wenjing Wu, et al · 2015
Cited alongside, same era.
8-bit approximations for parallelism in deep learning
Tim Dettmers · 2015
Cited alongside, same era.
Staleness-aware async-sgd for distributed deep learning
Wei Zhang, Suyog Gupta, Xiangru Lian, and Ji Liu · 2015
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Later among the works it cites.
Megatron-lm: Training multi-billion parameter language models using gpu model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Later among the works it cites.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Cited alongside, same era.
Generating a function for network delay
Andrei M Sukhov, MA Astrakhantseva, AK Pervitsky, SS Boldyrev, and AA Bukatov · 2016
Cited alongside, same era.
Gaussian error linear units (gelus), 2016
Dan Hendrycks and Kevin Gimpel · 2016
Cited alongside, same era.
Wikitext-2
2016 Stephen Merity et al · 2016
Cited alongside, same era.
A case study of ipv6 network performance: Packet delay, loss, and reordering
Fuliang Li, Xingwei Wang, Tian Pan, and Jiahai Yang · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour, 2017
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Cited alongside, same era.
Peng Sun, Wansen Feng, Ruobing Han, Shengen Yan, and Yonggang Wen · 2019
Later among the works it cites.
Zero: Memory optimization towards training a trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2019
Later among the works it cites.
Pipemare: Asynchronous pipeline parallel dnn training
Bowen Yang, Jian Zhang, Jonathan Li, Christopher Ré, Christopher R. Aberger, and Christopher De Sa · 2019
Later among the works it cites.
Pipedream: Generalized pipeline parallelism for dnn training
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R. Devanur, Gregory R. Ganger, Phillip B. Gibbons, and Matei Zaharia · 2019
Later among the works it cites.
Leela chess zero
Pascutto, Gian-Carlo and Linscott, Gary · 2019
Later among the works it cites.
Large memory layers with product keys
Guillaume Lample, Alexandre Sablayrolles, Marc´Aurelio Ranzato, Ludovic Denoyer, and Herve Jegou · 2019
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Natural compression for distributed deep learning
Samuel Horvath, Chen-Yu Ho, Ludovit Horvath, Atal Narayan Sahu, Marco Canini, and Peter Richtárik · 2019
Later among the works it cites.
Hybrid 8-bit floating point (hfp8) training and inference for deep neural networks
Xiao Sun, Jungwook Choi, Chia-Yu Chen, Naigang Wang, Swagath Venkataramani, Vijayalakshmi (Viji) Srinivasan, Xiaodong Cui, Wei Zhang, and Kailash Gopalakrishnan · 2019
Later among the works it cites.
The hsic bottleneck: Deep learning without back-propagation, 2019
Wan-Duo Kurt Ma, J. P. Lewis, and W. Bastiaan Kleijn · 2019
Later among the works it cites.
Scaling laws for neural language models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Closest in time.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Closest in time.
Estimate of GPT-3 training cost based on public cloud GPU/TPU cost models, from Elliot Turner’s personal page (accessed on May 29, 2020)
Elliot Turner · 2020
Closest in time.
https://foldingathome.org/project-timeline (accessed on May 30, 2020)
Folding@home project timeline · 2020
Closest in time.
https://www.speedtest.net/global-index (accessed on 11.08.2020, bandwidth for top countries and general trend)
Speedtest global index for fixed broadband · 2020
Closest in time.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2020
Closest in time.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, H. Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Y. Huang, M. Krikun, Noam Shazeer, and Z. Chen · 2020
Closest in time.
Scalable transfer learning with expert models
Joan Puigcerver, Carlos Riquelme, Basil Mustafa, Cedric Renggli, André Susano Pinto, Sylvain Gelly, Daniel Keysers, and Neil Houlsby · 2020
Closest in time.
https://foldingathome.org/covid19/ (accessed on June 4, 2020)
2020
Closest in time.
Large-scale evolution of image classifiers
Esteban Real, Sherry Moore, Andrew Selle, Saurabh Saxena, Yutaka Leon Suematsu, Jie Tan, Quoc V Le, and Alexey Kurakin · 2080
Closest in time.