Fetching the paper…
Reading the bibliography…
Deep learning researchers and practitioners usually leverage GPUs to help train their deep neural networks (DNNs) faster.
Systolic VLSI Arrays for Polynomial GCD Computation
Richard P Brent and Hsiang-Tsung Kung · 1984
Earlier work this paper cites.
Parallel Distributed Processing: Explorations in the Microstructure of Cognition, Vol. 1
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
NVIDIA’s Fermi: The First Complete GPU Computing Architecture
Peter N. Glaskowsky · 2009
Earlier work this paper cites.
Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures
Samuel Williams, Andrew Waterman, and David Patterson · 2009
Earlier work this paper cites.
Large-Scale Machine Learning with Stochastic Gradient Descent
Léon Bottou · 2010
Earlier work this paper cites.
Large Scale Distributed Deep Networks
Jeffrey Dean, Greg S. Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Quoc V. Le, Mark Z. Mao, Marc’Aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, and Andrew Y. Ng · 2012
Earlier work this paper cites.
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
cuDNN: Efficient Primitives for Deep Learning
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer · 2014
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
HBM2 - High Bandwidth Memory-2, 2015
Advanced Micro Devices, Inc · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
GDDR5, 2015
Micron Technology, Inc · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Rethinking the Inception Architecture for Computer Vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna · 2015
Earlier work this paper cites.
LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop
Fisher Yu, Yinda Zhang, Shuran Song, Ari Seff, and Jianxiong Xiao · 2015
Earlier work this paper cites.
TensorFlow: A System for Large-Scale Machine Learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Earlier work this paper cites.
Fathom: Reference Workloads for Modern Deep Learning Methods
Robert Adolf, Saketh Rama, Brandon Reagen, Gu-Yeon Wei, and David Brooks · 2016
Earlier work this paper cites.
Findings of the 2016 Conference on Machine Translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurelie Neveol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri · 2016
Earlier work this paper cites.
ESTIMA: Extrapolating Scalability of in-Memory Applications
Georgios Chatzopoulos, Aleksandar Dragojević, and Rachid Guerraoui · 2016
Earlier work this paper cites.
MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2016
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Asynchrony Begets Momentum, with an Application to Deep Learning
Ioannis Mitliagkas, Ce Zhang, Stefan Hadjis, and Christopher Ré · 2016
Earlier work this paper cites.
NVIDIA Pascal P100, 2016
NVIDIA Corporation · 2016
Earlier work this paper cites.
NVIDIA Tesla P100
NVIDIA Corporation · 2016
Earlier work this paper cites.
Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
Alec Radford, Luke Metz, and Soumith Chintala · 2016
Earlier work this paper cites.
Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Earlier work this paper cites.
AMD Ryzen Threadripper 1950X Processor, 2017
Advanced Micro Devices, Inc · 2017
Earlier work this paper cites.
DAWNBench: An End-to-End Deep Learning Benchmark and Competition
Cody Coleman, Deepak Narayanan, Daniel Kang, Tian Zhao, Jian Zhang, Luigi Nardi, Peter Bailis, Kunle Olukotun, Chris Ré, and Matei Zaharia · 2017
Earlier work this paper cites.
Densely Connected Convolutional Networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
GDDR6, 2017
Micron Technology, Inc · 2017
Earlier work this paper cites.
NVIDIA GeForce GTX 1080Ti, 2017
NVIDIA Corporation · 2017
Earlier work this paper cites.
NVIDIA Quadro P4000, 2017
NVIDIA Corporation · 2017
Earlier work this paper cites.
NVIDIA Tesla V100, 2017
NVIDIA Corporation · 2017
Earlier work this paper cites.
NVIDIA Tesla V100
NVIDIA Corporation · 2017
Cited alongside, same era.
NVIDIA TITAN Xp, 2017
NVIDIA Corporation · 2017
Cited alongside, same era.
Paleo: A Performance Model for Deep Neural Networks
Hang Qi, Evan R. Sparks, and Ameet Talwalkar · 2017
Cited alongside, same era.
Attention is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
JAX: Composable Transformations of Python+NumPy Programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang · 2018
Cited alongside, same era.
Ubuntu 18.04 LTS (Bionic Beaver), 2018
Canonical Ltd · 2018
Cited alongside, same era.
Google Cloud N1 Machine Types, 2020
Google, Inc · 2020
Later among the works it cites.
GPUs on Compute Engine, 2020
Google, Inc · 2020
Later among the works it cites.
Graphcore, 2020
Graphcore · 2020
Later among the works it cites.
Habana Labs, 2020
Habana Labs · 2020
Later among the works it cites.
Intel Xeon Processor E5-2680, 2020
Intel Corporation · 2020
Later among the works it cites.
Google breaks AI performance records in MLPerf with world’s fastest training supercomputer, 2020
Naveen Kumar · 2020
Later among the works it cites.
Lambda: Deep Learning Workstations, Servers, Laptops, 2020
Lambda Labs Inc · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to Optimize Tensor Programs
Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy · 2018
Cited alongside, same era.
Predicting the Computational Cost of Deep Learning Models
Daniel Justus, John Brennan, Stephen Bonner, and Andrew Stephen McGough · 2018
Cited alongside, same era.
Mixed Precision Training
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory F. Diamos, Erich Elsen, David García, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu · 2018
Cited alongside, same era.
NVIDIA Tesla T4, 2018
NVIDIA Corporation · 2018
Cited alongside, same era.
NVIDIA Turing Architecture
NVIDIA Corporation · 2018
Cited alongside, same era.
Gandiva: Introspective Cluster Scheduling for Deep Learning
Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Xuan Peng, Hanyu Zhao, Quanlu Zhang, Fan Yang, and Lidong Zhou · 2018
Cited alongside, same era.
Later among the works it cites.
MLPerf Training Benchmark
Peter Mattson, Christine Cheng, Cody Coleman, Greg Diamos, Paulius Micikevicius, David Patterson, Hanlin Tang, Gu-Yeon Wei, Peter Bailis, Victor Bittorf, David Brooks, Dehao Chen, Debojyoti Dutta, Udit Gupta, Kim Hazelwood, Andrew Hock, Xinyuan Huang, Bill Jia, Daniel Kang, David Kanter, Naveen Kumar, Jeffery Liao, Guokai Ma, Deepak Narayanan, Tayo Oguntebi, Gennady Pekhimenko, Lillian Pentecost, Vijay Janapa Reddi, Taylor Robie, Tom St. John, Carole-Jean Wu, Lingjie Xu, Cliff Young, and Matei Zaharia · 2020
Later among the works it cites.
Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads
Deepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee, and Matei Zaharia · 2020
Later among the works it cites.
cuBLAS: Dense Linear Algebra on GPUs, 2020
NVIDIA Corporation · 2020
Later among the works it cites.
NVIDIA A100, 2020
NVIDIA Corporation · 2020
Later among the works it cites.
NVIDIA Ampere Architecture In-Depth, 2020
NVIDIA Corporation · 2020
Later among the works it cites.
NVIDIA CUDA Toolkit, 2020
NVIDIA Corporation · 2020
Later among the works it cites.
NVIDIA Data Center Deep Learning Product Performance, 2020
NVIDIA Corporation · 2020
Later among the works it cites.
NVIDIA GeForce RTX 3090, 2020
NVIDIA Corporation · 2020
Later among the works it cites.
NVIDIA GPU Cloud Virtual Machine Image Release Notes, 2020
NVIDIA Corporation · 2020
Later among the works it cites.
Quadro RTX 6000 Graphics Card, 2020
NVIDIA Corporation · 2020
Later among the works it cites.
Announcing availability of Inf1 instances in Amazon SageMaker for high performance and cost-effective machine learning inference, 2020
Julien Simon · 2020
Later among the works it cites.
TechPowerUp GPU Database (P4000 and 2070), 2020
TechPowerUp · 2020
Later among the works it cites.
Skyline: Interactive In-Editor Computational Performance Profiling for Deep Neural Network Training
Geoffrey X. Yu, Tovi Grossman, and Gennady Pekhimenko · 2020
Later among the works it cites.
Optimizing Memory-Access Patterns for Deep Learning Accelerators
Hongbin Zheng, Sejong Oh, Huiqing Wang, Preston Briggs, Jiading Gai, Animesh Jain, Yizhi Liu, Rich Heaton, Randy Huang, and Yida Wang · 2020
Later among the works it cites.
Daydream: Accurately Estimating the Efficacy of Optimizations for DNN Training
Hongyu Zhu, Amar Phanishayee, and Gennady Pekhimenko · 2020
Later among the works it cites.
Amazon SageMaker, 2021
Amazon, Inc · 2021
Closest in time.
AWS Inferentia, 2021
Amazon, Inc · 2021
Closest in time.
AWS Trainium, 2021
Amazon, Inc · 2021
Closest in time.
torchvision, 2021
PyTorch Contributors · 2021
Closest in time.
Google Cloud Vertex AI, 2021
Google, Inc · 2021
Closest in time.
Supported TPU Versions, 2021
Google, Inc · 2021
Closest in time.
XLA: Optimizing Compiler for Machine Learning, 2021
Google, Inc · 2021
Closest in time.
A Learned Performance Model for Tensor Processing Units
Samuel J. Kaufman, Phitchaya Mangpo Phothilimthana, Yanqi Zhou, Charith Mendis, Sudip Roy, Amit Sabne, and Mike Burrows · 2021
Closest in time.
Azure Machine Learning, 2021
Microsoft Corporation · 2021
Closest in time.
CUDA Programming Guide, 2021
NVIDIA Corporation · 2021
Closest in time.
Habitat: A Runtime-Based Computational Performance Predictor for Deep Neural Network Training (Code), 2021
Geoffrey X. Yu, Yubo Gao, Pavel Golikov, and Gennady Pekhimenko · 2021
Closest in time.
Habitat Pre-Trained Models and Kernel Metadata, 2021
Geoffrey X. Yu, Yubo Gao, Pavel Golikov, and Gennady Pekhimenko · 2021
Closest in time.
NVIDIA GeForce RTX 2070, 2018
NVIDIA Corporation · 2070
Closest in time.
NVIDIA GeForce RTX 2080Ti, 2018
NVIDIA Corporation · 2080
Closest in time.