Fetching the paper…
Reading the bibliography…
Modern deep learning models have been exploited in various domains, including computer vision (CV), natural language processing (NLP), search and recommendation.
Analyzing cuda workloads using a detailed gpu simulator
Ali Bakhoda, George L Yuan, Wilson WL Fung, Henry Wong, and Tor M Aamodt · 2009
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Nonlinear latent factorization by embedding multiple user interests
Jason Weston, Ron J Weiss, and Hector Yee · 2013
Earlier work this paper cites.
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell · 2014
Earlier work this paper cites.
An introduction to computational networks and the computational network toolkit
Dong Yu, Adam Eversole, Mike Seltzer, Kaisheng Yao, Zhiheng Huang, Brian Guenter, Oleksii Kuchaiev, Yu Zhang, Frank Seide, Huaming Wang, et al · 2014
Earlier work this paper cites.
Attention-based models for speech recognition
Jan K Chorowski, Dzmitry Bahdanau, Dmitriy Serdyuk, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Deepdriving: Learning affordance for direct perception in autonomous driving
Chenyi Chen, Ari Seff, Alain Kornhauser, and Jianxiong Xiao · 2015
Earlier work this paper cites.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Earlier work this paper cites.
Identity mappings in deep residual networks
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Earlier work this paper cites.
Benchmarking state-of-the-art deep learning software tools
Shaohuai Shi, Qiang Wang, Pengfei Xu, and Xiaowen Chu · 2016
Earlier work this paper cites.
Fathom: Reference workloads for modern deep learning methods
Robert Adolf, Saketh Rama, Brandon Reagen, Gu-Yeon Wei, and David Brooks · 2016
Earlier work this paper cites.
Peter Goldsborough · 2016
Cited alongside, same era.
Paleo: A performance model for deep neural networks
Hang Qi, Evan R Sparks, and Ameet Talwalkar · 2016
Cited alongside, same era.
R-fcn: Object detection via region-based fully convolutional networks
Jifeng Dai, Yi Li, Kaiming He, and Jian Sun · 2016
Cited alongside, same era.
Deep neural networks for youtube recommendations
Paul Covington, Jay Adams, and Emre Sargin · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Cited alongside, same era.
Deepprof: Performance analysis for deep learning applications via mining gpu execution patterns
Graph convolutional neural networks for web-scale recommender systems
Rex Ying, Ruining He, Kaifeng Chen, Pong Eksombatchai, William L Hamilton, and Jure Leskovec · 2018
Later among the works it cites.
Tartan: Evaluating modern gpu interconnect via a multi-gpu benchmark suite
Ang Li, Shuaiwen Leon Song, Jieyang Chen, Xu Liu, Nathan Tallent, and Kevin Barker · 2018
Later among the works it cites.
Data motif-based proxy benchmarks for big data and ai workloads
Wanling Gao, Jianfeng Zhan, Lei Wang, Chunjie Luo, Zhen Jia, Daoyi Zheng, Chen Zheng, Xiwen He, Hainan Ye, Haibin Wang, et al · 2018
Later among the works it cites.
Applied machine learning at facebook: a datacenter infrastructure perspective
Kim Hazelwood, Sarah Bird, David Brooks, Soumith Chintala, Utku Diril, Dmytro Dzhulgakov, Mohamed Fawzy, Bill Jia, Yangqing Jia, Aditya Kalro, et al · 2018
Later among the works it cites.
Jongsoo Park, Maxim Naumov, Protonu Basu, Summer Deng, Aravind Kalaiah, Daya Khudia, James Law, Parth Malani, Andrey Malevich, Satish Nadathur, et al · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiazhen Gu, Huan Liu, Yangfan Zhou, and Xin Wang · 2017
Cited alongside, same era.
Nvidia tesla v100 gpu architecture
NVIDIA · 2017
Cited alongside, same era.
Xla – tensorflow compiled. post in the google developers blog
The XLA Team · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Poseidon: An efficient communication architecture for distributed deep learning on { \{ GPU } \} clusters
Hao Zhang, Zeyu Zheng, Shizhen Xu, Wei Dai, Qirong Ho, Xiaodan Liang, Zhiting Hu, Jinliang Wei, Pengtao Xie, and Eric P Xing · 2017
Cited alongside, same era.
Evaluating on-node gpu interconnects for deep learning workloads
Nathan R Tallent, Nitin A Gawande, Charles Siegel, Abhinav Vishnu, and Adolfy Hoisie · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Later among the works it cites.
Performance modeling and evaluation of distributed deep learning frameworks on gpus
Shaohuai Shi, Qiang Wang, and Xiaowen Chu · 2018
Later among the works it cites.
Benchmarking and analyzing deep neural network training
Hongyu Zhu, Mohamed Akrout, Bojian Zheng, Andrew Pelegris, Anand Jayarajan, Amar Phanishayee, Bianca Schroeder, and Gennady Pekhimenko · 2018
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
Christopher J Shallue, Jaehoon Lee, Joe Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl · 2018
Later among the works it cites.
Nvidia collective communications library
NVIDIA · 2018
Later among the works it cites.
Tictac: Accelerating distributed deep learning with communication scheduling
Sayed Hadi Hashemi, Sangeetha Abdu Jyothi, and Roy H Campbell · 2018
Later among the works it cites.
Fusionstitching: Deep fusion and code generation for tensorflow computations on gpus
Guoping Long, Jun Yang, Kai Zhu, and Wei Lin · 2018
Later among the works it cites.
Multi-tenant gpu clusters for deep learning workloads: Analysis and implications
Myeongjae Jeon, Shivaram Venkataraman, Junjie Qian, Amar Phanishayee, Wencong Xiao, and Fan Yang · 2018
Later among the works it cites.
Data motifs: a lens towards fully understanding big data and ai workloads
Wanling Gao, Jianfeng Zhan, Lei Wang, Chunjie Luo, Daoyi Zheng, Fei Tang, Biwei Xie, Chen Zheng, Xu Wen, Xiwen He, et al · 2018
Later among the works it cites.
Performance characterization of state-of-the-art deep learning workloads on an ibm” minsky” platform
Mauricio Guignard, Marcelo Schild, Carlos S Bederián, Nicolás Wolovick, and Augusto J Vega · 2018
Later among the works it cites.
Characterizing deep-learning i/o workloads in tensorflow
Steven WD Chien, Stefano Markidis, Chaitanya Prasad Sishtla, Luis Santos, Pawel Herman, Sai Narasimhamurthy, and Erwin Laure · 2018
Later among the works it cites.
Predicting the computational cost of deep learning models
Daniel Justus, John Brennan, Stephen Bonner, and Andrew Stephen McGough · 2018
Later among the works it cites.
Scalable deep learning on distributed infrastructures: Challenges, techniques and tools
Ruben Mayer and Hans-Arno Jacobsen · 2019
Closest in time.