Fetching the paper…
Reading the bibliography…
Deep learning has become widely used in complex AI applications.
NVIDIA · 1909
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
Felix A Gers, Jürgen Schmidhuber, and Fred Cummins · 1999
Earlier work this paper cites.
ETA: Experience with an Intel Xeon processor as a packet processing engine
Greg Regnier, Dave Minturn, Gary McAlpine, Vikram A Saletore, and Annie Foong · 2004
Earlier work this paper cites.
GPGPU: general-purpose computation on graphics hardware
David Luebke, Mark Harris, Naga Govindaraju, Aaron Lefohn, Mike Houston, John Owens, Mark Segal, Matthew Papakipos, and Ian Buck · 2006
Earlier work this paper cites.
Roofline: An insightful visual performance model for floating-point programs and multicore architectures*
Samuel Williams, Andrew Waterman, and David A. Patterson · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng et al · 2009
Earlier work this paper cites.
Large scale distributed deep networks
etc Jeffrey Dean · 2012
Earlier work this paper cites.
cuDNN: Efficient primitives for deep learning
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Efficient mini-batch training for stochastic optimization
Mu Li, Tong Zhang, Yuqiang Chen, and Alexander J Smola · 2014
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Martın Abadi et al · 2015
Cited alongside, same era.
Harmonia: Balancing compute and memory power in high-performance gpus
I. Paul et al · 2015
Cited alongside, same era.
You only look once: Unified, real-time object detection
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi · 2016
Cited alongside, same era.
Deep speech 2: End-to-end speech recognition in english and mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Cited alongside, same era.
Fathom: Reference workloads for modern deep learning methods
Robert Adolf, Saketh Rama, Brandon Reagen, Gu-Yeon Wei, and David M. Brooks · 2016
Cited alongside, same era.
A survey and measurement study of GPU DVFS on energy conservation
Xinxin Mei, Qiang Wang, and Xiaowen Chu · 2017
Later among the works it cites.
Energy efficient job scheduling with DVFS for CPU-GPU heterogeneous systems
Vincent Chau, Xiaowen Chu, Hai Liu, and Yiu-Wing Leung · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Energy efficient real-time task scheduling on CPU-GPU hybrid clusters
Xinxin Mei, Xiaowen Chu, Hai Liu, Yiu-Wing Leung, and Zongpeng Li · 2017
Later among the works it cites.
Image classification at supercomputer scale
Chris Ying, Sameer Kumar, Dehao Chen, Tao Wang, and Youlong Cheng · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shaohuai Shi, Qiang Wang, Pengfei Xu, and Xiaowen Chu · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy et al · 2016
Cited alongside, same era.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2016
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit
Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, et al · 2017
Cited alongside, same era.
Dawnbench: An end-to-end deep learning benchmark and competition
Cody Coleman, Deepak Narayanan, Daniel Kang, Tian Zhao, Jian Zhang, Luigi Nardi, Peter Bailis, Kunle Olukotun, Chris Ré, and Matei Zaharia · 2017
Cited alongside, same era.
EPPMiner: An extended benchmark suite for energy, power and performance characterization of heterogeneous architecture
Qiang Wang et al · 2017
Cited alongside, same era.
ROCm System Management Library
AMD
Cited in the paper.
Scott Cyphers et al · 2018
Later among the works it cites.
TBD: benchmarking and analyzing deep neural network training
Hongyu Zhu et al · 2018
Later among the works it cites.
GPGPU performance estimation with core and memory frequency scaling
Qiang Wang and Xiaowen Chu · 2018
Later among the works it cites.
Benchmarking TPU, GPU, and CPU platforms for deep learning
Yu Emma Wang, Gu-Yeon Wei, David Brooks, et al · 2019
Closest in time.
Aibench: An industry standard internet service ai benchmark suite, 2019
Wanling Gao, Fei Tang, Lei Wang, Jianfeng Zhan, Chunxin Lan, Chunjie Luo, Yunyou Huang, Chen Zheng, Jiahui Dai, Zheng Cao, Daoyi Zheng, Haoning Tang, Kunlin Zhan, Biao Wang, Defei Kong, Tong Wu, Minghe Yu, Chongkang Tan, Huan Li, Xinhui Tian, Yatao Li, Junchao Shao, Zhenyu Wang, Xiaoyu Wang, and Hainan Ye · 2019
Closest in time.
The impact of GPU DVFS on the energy and performance of deep learning: an empirical study
Zhenheng Tang et al · 2019
Closest in time.