Fetching the paper…
Reading the bibliography…
Modern deep learning applications require increasingly more compute to train state-of-the-art models.
Application of programs with maximin objective functions to problems of optimal resource allocation
Seymour Kaplan · 1974
Earlier work this paper cites.
The MOSEK interior point optimizer for linear programming: An implementation of the homogeneous algorithm
Erling D. Andersen and Knud D. Andersen · 2000
Earlier work this paper cites.
SETI@home: An experiment in public-resource computing
David Anderson, Jeff Cobb, Eric Korpela, Matt Lebofsky, and Dan Werthimer · 2002
Earlier work this paper cites.
Kademlia: A peer-to-peer information system based on the XOR metric
Petar Maymounkov and David Mazieres · 2002
Earlier work this paper cites.
Koorde: A simple degree-optimal distributed hash table
M Frans Kaashoek and David R Karger · 2003
Earlier work this paper cites.
STUN - simple traversal of user datagram protocol (UDP) through network address translators (NATs)
J. Rosenberg, J. Weinberger, C. Huitema, and R. Mahy · 2003
Earlier work this paper cites.
BOINC: A system for public-resource computing and storage
David P Anderson · 2004
Earlier work this paper cites.
Optimization of collective communication operations in MPICH
Rajeev Thakur, Rolf Rabenseifner, and William Gropp · 2005
Earlier work this paper cites.
NATBLASTER: Establishing TCP connections between hosts behind NATs
Andrew Biggadike, Daniel Ferullo, Geoffrey Wilson, and Adrian Perrig · 2005
Earlier work this paper cites.
Peer-to-peer communication across network address translators
Bryan Ford, Pyda Srisuresh, and Dan Kegel · 2005
Earlier work this paper cites.
Accelerating molecular dynamics simulations on playstation 3 platform using virtual-grape programming model
Tetsu Narumi, Shun Kameoka, Makoto Taiji, and Kenji Yasuoka · 2008
Earlier work this paper cites.
The BitTorrent Protocol Specification
Bram Cohen · 2008
Earlier work this paper cites.
ImageNet: a large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Folding@home: Lessons from eight years of volunteer distributed computing
A. L. Beberg, D. Ensign, G. Jayachandran, S. Khaliq, and V. Pande · 2009
Earlier work this paper cites.
Bandwidth optimal all-reduce algorithms for clusters of workstations
Pitch Patarasuk and Xin Yuan · 2009
Earlier work this paper cites.
Stefan M. Larson, Christopher D. Snow, Michael Shirts, and Vijay S. Pande · 2009
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex Smola · 2010
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Re, Stephen Wright, and Feng Niu · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc' aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, Quoc Le, and Andrew Ng · 2012
Earlier work this paper cites.
Folding research recruits unconventional help
Michael Gross · 2012
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
Earlier work this paper cites.
DeCAF: A deep convolutional activation feature for generic visual recognition
J. Donahue, Y. Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, and Trevor Darrell · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Mu Li, D. Andersen, J. Park, Alex Smola, Amr Ahmed, V. Josifovski, J. Long, E. Shekita, and Bor-Yiing Su · 2014
Earlier work this paper cites.
Project adam: Building an efficient and scalable deep learning training system
Trishul Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Fully convolutional networks for semantic segmentation
Jonathan Long, Evan Shelhamer, and Trevor Darrell · 2015
Earlier work this paper cites.
Collective algorithms for multiported torus networks
Paul Sack and William Gropp · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Net2Net: Accelerating learning via knowledge transfer
Tianqi Chen, Ian Goodfellow, and Jonathon Shlens · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei · 2016
Earlier work this paper cites.
Volunteer computing on mobile devices: State of the art and future research directions
C. Tapparello, Colin Funai, Shurouq Hijazi, Abner Aquino, Bora Karaoglu, H. Ba, J. Shi, and W. Heinzelman · 2016
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous distributed systems
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mane, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viegas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Holistically evaluating the environmental impacts in modern computing systems
Donald Kline, Nikolas Parshook, Xiaoyu Ge, E. Brunvand, R. Melhem, Panos K. Chrysanthis, and A. Jones · 2016
Earlier work this paper cites.
Can decentralized algorithms outperform centralized algorithms? A case study for decentralized parallel stochastic gradient descent
Xiangru Lian, Ce Zhang, Huan Zhang, Cho-Jui Hsieh, Wei Zhang, and Ji Liu · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: Training ImageNet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Earlier work this paper cites.
LHC@Home: a BOINC-based volunteer computing infrastructure for physics studies at CERN
Javier Barranco, Yunhi Cai, David Cameron, Matthew Crouch, Riccardo De Maria, Laurence Field, M. Giovannozzi, Pascal Hermes, Nils Høimyr, Dobrin Kaltchev, Nikos Karastathis, Cinzia Luzzi, Ewen Maclean, Eric Mcintosh, Alessio Mereghetti, James Molson, Yuri Nosochkov, Tatiana Pieloni, Ivan Reid, and Igor Zacharov · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Earlier work this paper cites.
An efficient task-based all-reduce for machine learning applications
Zhenyu Li, James Davis, and Stephen Jarvis · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and R. Socher · 2017
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji · 2017
Earlier work this paper cites.
Practical secure aggregation for privacy-preserving machine learning
Keith Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone, H Brendan McMahan, Sarvar Patel, Daniel Ramage, Aaron Segal, and Karn Seth · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Performance modeling and evaluation of distributed deep learning frameworks on GPUs
Shaohuai Shi, Qiang Wang, and Xiaowen Chu · 2018
Cited alongside, same era.
Horovod: fast and easy distributed deep learning in TensorFlow
Alexander Sergeev and Mike Del Balso · 2018
Cited alongside, same era.
Deep Gradient Compression: Reducing the communication bandwidth for distributed training
Yujun Lin, Song Han, Huizi Mao, Yu Wang, and William J Dally · 2018
Cited alongside, same era.
Decentralized federated learning preserves model and data privacy
Thorsten Wittkopp and Alexander Acker · 2020
Later among the works it cites.
Towards crowdsourced training of large neural networks using decentralized mixture-of-experts
Max Ryabinin and Anton Gusev · 2020
Later among the works it cites.
Understanding the difficulty of training transformers
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han · 2020
Later among the works it cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Pytorch distributed: Experiences on accelerating data parallel training
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, and Soumith Chintala · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
QMCPACK: an open sourceab initioquantum monte carlo package for the electronic structure of atoms, molecules and solids
Jeongnim Kim, Andrew D Baczewski, Todd D Beaudet, Anouar Benali, M Chandler Bennett, Mark A Berrill, Nick S Blunt, Edgar Josué Landinez Borda, Michele Casula, David M Ceperley, Simone Chiesa, Bryan K Clark, Raymond C Clay, Kris T Delaney, Mark Dewing, Kenneth P Esler, Hongxia Hao, Olle Heinonen, Paul R C Kent, Jaron T Krogel, Ilkka Kylänpää, Ying Wai Li, M Graham Lopez, Ye Luo, Fionn D Malone, Richard M Martin, Amrita Mathuriya, Jeremy McMinis, Cody A Melton, Lubos Mitas, Miguel A Morales, Eric Neuscamman, William D Parker, Sergio D Pineda Flores, Nichols A Romero, Brenda M Rubenstein, Jacqueline A R Shea, Hyeondeok Shin, Luke Shulenburger, Andreas F Tillack, Joshua P Townsend, Norm M Tubman, Brett Van Der Goetz, Jordan E Vincent, D ChangMo Yang, Yubo Yang, Shuai Zhang, and Luning Zhao · 2018
Cited alongside, same era.
A hybrid gpu cluster and volunteer computing platform for scalable deep learning
Ekasit Kijsipongse, Apivadee Piyatumrong, and Suriya U-ruekolan · 2018
Cited alongside, same era.
Training tips for the transformer model
Martin Popel and Ondřej Bojar · 2018
Cited alongside, same era.
Scaling neural machine translation
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli · 2018
Cited alongside, same era.
Mixed precision training
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu · 2018
Cited alongside, same era.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo · 2018
Cited alongside, same era.
A comprehensive reasoning framework for hardware refresh in data centers
Rabih Bashroush · 2018
Cited alongside, same era.
Traversal using relays around NAT (TURN): Relay extensions to session traversal utilities for NAT (STUN)
T. Reddy, A. Johnston, P. Matthews, and J. Rosenberg · 2020
Later among the works it cites.
Unsupervised learning of visual features by contrasting cluster assignments
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Later among the works it cites.
Hivemind: a Library for Decentralized Deep Learning
Learning@home team · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2020
Later among the works it cites.
A monolingual approach to contextualized word embeddings for mid-resource languages
Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot · 2020
Later among the works it cites.
Experiment tracking with Weights and Biases, 2020
Lukas Biewald · 2020
Later among the works it cites.
IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages
Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar · 2020
Later among the works it cites.
Indic-transformers: An analysis of transformer language models for Indian languages
Kushal Jain, Adwait Deshpande, Kumar Shridhar, Felix Laumann, and Ayushman Dash · 2020
Later among the works it cites.
Carbontracker: Tracking and predicting the carbon footprint of training deep learning models
Lasse F. Wolff Anthony, Benjamin Kanding, and Raghavendra Selvan · 2020
Later among the works it cites.
Federated learning with differential privacy: Algorithms and performance analysis
Kang Wei, Jun Li, Ming Ding, Chuan Ma, Howard H. Yang, Farhad Farokhi, Shi Jin, Tony Q. S. Quek, and H. Vincent Poor · 2020
Later among the works it cites.
Unified analysis of stochastic gradient methods for composite convex and smooth optimization, 2020
Ahmed Khaled, Othmane Sebbouh, Nicolas Loizou, Robert M. Gower, and Peter Richtárik · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Later among the works it cites.
GLU variants improve transformer
Noam Shazeer · 2020
Later among the works it cites.
Green AI
Roy Schwartz, Jesse Dodge, Noah Smith, and Oren Etzioni · 2020
Later among the works it cites.
Towards the systematic reporting of the energy and carbon footprints of machine learning
Peter Henderson, Jieru Hu, Joshua Romoff, Emma Brunskill, Dan Jurafsky, and Joelle Pineau · 2020
Later among the works it cites.
Can federated learning save the planet?
Xinchi Qiu, Titouan Parcollet, Daniel J. Beutel, Taner Topal, Akhil Mathur, and Nicholas D. Lane · 2020
Later among the works it cites.
Efficient large-scale language model training on GPU clusters using Megatron-LM
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Anand Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia · 2021
Closest in time.
https://www.tensorflow.org/hub
TensorFlow Hub · 2021
Closest in time.
https://pytorch.org/hub/
PyTorch Hub · 2021
Closest in time.
https://huggingface.co/models
Hugging Face Hub · 2021
Closest in time.
https://blogs.nvidia.com/blog/2020/04/01/foldingathome-exaflop-coronavirus/
Folding@home gets 1.5+ exaflops to fight covid-19 · 2021
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
https://pytorch.org/elastic
PyTorch Elastic · 2021
Closest in time.
https://horovod.rtfd.io/en/stable/elastic_include.html
Elastic Horovod · 2021
Closest in time.
Moshpit SGD: Communication-efficient decentralized training on heterogeneous unreliable devices
Max Ryabinin, Eduard Gorbunov, Vsevolod Plokhotnyuk, and Gennady Pekhimenko · 2021
Closest in time.
foldingathome.org/2020/03/10/covid19-update
Folding@home update on SARS-CoV-2 (10 mar 2020) · 2021
Closest in time.
https://foldingathome.org/project-timeline
Folding@home project timeline · 2021
Closest in time.
Einstein@Home all-sky search for continuous gravitational waves in LIGO O2 public data
B. Steltner, M. A. Papa, H. B. Eggenstein, B. Allen, V. Dergachev, R. Prix, B. Machenschalk, S. Walsh, S. J. Zhu, and S. Kwang · 2021
Closest in time.
MLDS: A dataset for weight-space analysis of neural networks
John Clemens · 2021
Closest in time.
Leela chess zero
Gian-Carlo Pascutto and Gary Linscott · 2021
Closest in time.
Distributed deep learning using volunteer computing-like paradigm
Medha Atre, Birendra Jha, and Ashwini Rao · 2021
Closest in time.
ZeRO-offload: Democratizing billion-scale model training
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He · 2021
Closest in time.
Invisible internet project (i2p) project overview
jrandom (Pseudonym) · 2021
Closest in time.
https://support.google.com/stadia/answer/9607891
Google Stadia data usage · 2021
Closest in time.
https://help.netflix.com/en/node/87
Netflix data usage · 2021
Closest in time.
https://libp2p.io/
libp2p · 2021
Closest in time.
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander M. Rush, and Thomas Wolf · 2021
Closest in time.
Do transformer modifications transfer across implementations and applications?
Sharan Narang, Hyung Won Chung, Yi Tay, William Fedus, Thibault Févry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam M. Shazeer, Zhenzhong Lan, Yanqi Zhou, Wen hong Li, Nan Ding, Jake Marcus, Adam Roberts, and Colin Raffel · 2021
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Closest in time.
Rotary embeddings: A relative revolution
Stella Biderman, Sid Black, Charles Foster, Leo Gao, Eric Hallahan, Horace He, Ben Wang, and Phil Wang · 2021
Closest in time.