Fetching the paper…
Reading the bibliography…
The strong demand for efficient and performant deployment of Deep Learning (DL) applications prompts the rapid development of a rich DL ecosystem.
DEAP: Evolutionary algorithms made easy
Félix-Antoine Fortin, François-Michel De Rainville, Marc-André Gardner Gardner, Marc Parizeau, and Christian Gagné. 2012 · 2012
Earlier work this paper cites.
Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe. 2013 · 2013
Earlier work this paper cites.
Opentuner: An extensible framework for program autotuning. In Proceedings of the 23rd international conference on Parallel architectures and compilation . 303–316
Jason Ansel, Shoaib Kamil, Kalyan Veeramachaneni, Jonathan Ragan-Kelley, Jeffrey Bosboom, Una-May O’Reilly, and Saman Amarasinghe. 2014 · 2014
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014 · 2014
Earlier work this paper cites.
Intel math kernel library
Endong Wang, Qing Zhang, Bo Shen, Guangyong Zhang, Xiaowei Lu, Qing Wu, and Yajuan Wang. 2014 · 2014
Earlier work this paper cites.
On optimizing machine learning workloads via kernel fusion
Arash Ashari, Shirish Tatikonda, Matthias Boehm, Berthold Reinwald, Keith Campbell, John Keenleyside, and P Sadayappan. 2015 · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala. 2015 · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning. In 12th { \{ USENIX } \} symposium on operating systems design and implementation ( { \{ OSDI } \} 16) . 265–283
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Latte: A language, compiler, and runtime for elegant and efficient deep neural networks. In Proceedings of the 37th ACM SIGPLAN Conference on Programming Language Design and Implementation . 209–223
Leonard Truong, Rajkishore Barik, Ehsan Totoni, Hai Liu, Chick Markley, Armando Fox, and Tatiana Shpeisman. 2016 · 2016
Earlier work this paper cites.
SPOOF: Sum-Product Optimization and Operator Fusion for Large-Scale Machine Learning.. In CIDR
Tarek Elgamal, Shangyu Luo, Matthias Boehm, Alexandre V Evfimievski, Shirish Tatikonda, Berthold Reinwald, and Prithviraj Sen. 2017 · 2017
Earlier work this paper cites.
In-datacenter performance analysis of a tensor processing unit. In Proceedings of the 44th annual international symposium on computer architecture . 1–12
Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, et al · 2017
Earlier work this paper cites.
The tensor algebra compiler
Fredrik Kjolstad, Shoaib Kamil, Stephen Chou, David Lugato, and Saman Amarasinghe. 2017 · 2017
Earlier work this paper cites.
Device placement optimization with reinforcement learning. In International Conference on Machine Learning . PMLR, 2430–2439
Azalia Mirhoseini, Hieu Pham, Quoc V Le, Benoit Steiner, Rasmus Larsen, Yuefeng Zhou, Naveen Kumar, Mohammad Norouzi, Samy Bengio, and Jeff Dean. 2017 · 2017
Earlier work this paper cites.
Aggregated Residual Transformations for Deep Neural Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Saining Xie, Ross Girshick, Piotr Dollar, Zhuowen Tu, and Kaiming He. 2017 · 2017
Earlier work this paper cites.
Placeto: Efficient progressive device placement optimization. In NIPS Machine Learning for Systems Workshop
Ravichandra Addanki, Shaileshh Bojja Venkatakrishnan, Shreyan Gupta, Hongzi Mao, and Mohammad Alizadeh. 2018 · 2018
Earlier work this paper cites.
On optimizing operator fusion plans for large-scale machine learning in systemml
Matthias Boehm, Berthold Reinwald, Dylan Hutchison, Alexandre V Evfimievski, and Prithviraj Sen. 2018 · 2018
Earlier work this paper cites.
Learning to optimize tensor programs
Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018b · 2018
Earlier work this paper cites.
Intel ngraph: An intermediate representation, compiler, and executor for deep learning
Scott Cyphers, Arjun K Bansal, Anahita Bhiwandiwalla, Jayaram Bobba, Matthew Brookhart, Avijit Chakraborty, Will Constable, Christian Convey, Leona Cook, Omar Kanawi, et al · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Spotlight: Optimizing device placement for training deep neural networks. In International Conference on Machine Learning . PMLR, 1676–1684
Yuanxiang Gao, Li Chen, and Baochun Li. 2018 · 2018
Earlier work this paper cites.
Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 6546–6555
Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh. 2018 · 2018
Earlier work this paper cites.
Exploring Hidden Dimensions in Parallelizing Convolutional Neural Networks.. In ICML . 2279–2288
Zhihao Jia, Sina Lin, Charles R Qi, and Alex Aiken. 2018 · 2018
Earlier work this paper cites.
A hierarchical model for device placement. In International Conference on Learning Representations
Azalia Mirhoseini, Anna Goldie, Hieu Pham, Benoit Steiner, Quoc V Le, and Jeff Dean. 2018 · 2018
Cited alongside, same era.
Relay: A new ir for machine learning frameworks. In Proceedings of the 2nd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages . 58–68
Jared Roesch, Steven Lyubomirsky, Logan Weber, Josh Pollock, Marisa Kirisame, Tianqi Chen, and Zachary Tatlock. 2018 · 2018
Cited alongside, same era.
Glow: Graph lowering compiler techniques for neural networks
Nadav Rotem, Jordan Fix, Saleem Abdulrasool, Garret Catron, Summer Deng, Roman Dzhabarov, Nick Gibson, James Hegeman, Meghan Lele, Roman Levenstein, et al · 2018
Cited alongside, same era.
Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions
Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen. 2018 · 2018
Cited alongside, same era.
Fusionstitching: boosting memory intensive computations for deep learning workloads
Zhen Zheng, Pengzhan Zhao, Guoping Long, Feiwen Zhu, Kai Zhu, Wenyi Zhao, Lansong Diao, Jun Yang, and Wei Lin. 2020c · 2020
Later among the works it cites.
Apple Neural Engine (ANE)
[n.d.] · 2021
Closest in time.
Intel OneDNN
[n.d.] · 2021
Closest in time.
Intel OpenVINO
[n.d.] · 2021
Closest in time.
NVIDIA cuBLAS
[n.d.] · 2021
Closest in time.
NVIDIA Deep Learning Accelerator (NVDLA)
[n.d.] · 2021
Closest in time.
NVIDIA TensorRT
[n.d.]a · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V. Le. 2018 · 2018
Cited alongside, same era.
Learning to fuse. In NeurIPS ML for Systems Workshop
Amirali Abdolrashidi, Qiumin Xu, Shibo Wang, Sudip Roy, and Yanqi Zhou. 2019 · 2019
Cited alongside, same era.
Learning to optimize Halide with tree search and random programs
Andrew Adams, Karima Ma, Luke Anderson, Riyadh Baghdadi, Tzu-Mao Li, Michaël Gharbi, Benoit Steiner, Steven Johnson, Kayvon Fatahalian, Frédo Durand, and Jonathan Ragan-Kelley. 2019 · 2019
Cited alongside, same era.
Tiramisu: A polyhedral compiler for expressing fast and portable code. In 2019 IEEE/ACM International Symposium on Code Generation and Optimization (CGO) . IEEE, 193–205
Riyadh Baghdadi, Jessica Ray, Malek Ben Romdhane, Emanuele Del Sozzo, Abdurrahman Akkas, Yunming Zhang, Patricia Suriana, Shoaib Kamil, and Saman Amarasinghe. 2019 · 2019
Cited alongside, same era.
Igc: The open source intel graphics compiler. In 2019 IEEE/ACM International Symposium on Code Generation and Optimization (CGO) . IEEE, 254–265
Anupama Chandrasekhar, Gang Chen, Po-Yu Chen, Wei-Yu Chen, Junjie Gu, Peng Guo, Shruthi Hebbur Prasanna Kumar, Guei-Yuan Lueh, Pankaj Mistry, Wei Pan, et al · 2019
Cited alongside, same era.
Optimizing dnn computation with relaxed graph substitutions
Zhihao Jia, James Thomas, Tod Warszawski, Mingyu Gao, Matei Zaharia, and Alex Aiken. 2019b · 2019
Cited alongside, same era.
Beyond Data and Model Parallelism for Deep Neural Networks.. In Proceedings of Machine Learning and Systems , A. Talwalkar, V. Smith, and M. Zaharia (Eds.), Vol. 1. 1–13
Zhihao Jia, Matei Zaharia, and Alex Aiken. 2019c · 2019
Cited alongside, same era.
Learned TPU cost model for XLA tensor programs. In Proc. Workshop ML Syst. NeurIPS . 1–6
Samuel Kaufman, Phitchaya Mangpo Phothilimthana, and Mike Burrows. 2019 · 2019
Cited alongside, same era.
Tensorflow XLA
[n.d.] · 2021
Closest in time.
Cortex: A Compiler for Recursive Deep Learning Models
Pratik Fegade, Tianqi Chen, Phillip Gibbons, and Todd Mowry. 2021 · 2021
Closest in time.
Union: A unified HW-SW Co-Design ecosystem in MLIR for evaluating tensor operations on spatial accelerators. In 2021 30th International Conference on Parallel Architectures and Compilation Techniques (PACT) . IEEE, 30–44
Geonhwa Jeong, Gokcen Kestor, Prasanth Chatarasi, Angshuman Parashar, Po-An Tsai, Sivasankaran Rajamanickam, Roberto Gioiosa, and Tushar Krishna. 2021 · 2021
Closest in time.
DeepCuts: a deep learning optimization framework for versatile GPU workloads. In Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation . 190–205
Wookeun Jung, Thanh Tuan Dao, and Jaejin Lee. 2021 · 2021
Closest in time.
Mlir: Scaling compiler infrastructure for domain specific computation. In 2021 IEEE/ACM International Symposium on Code Generation and Optimization (CGO) . IEEE, 2–14
Chris Lattner, Mehdi Amini, Uday Bondhugula, Albert Cohen, Andy Davis, Jacques Pienaar, River Riddle, Tatiana Shpeisman, Nicolas Vasilache, and Oleksandr Zinenko. 2021 · 2021
Closest in time.
C-for-metal: high performance SIMD programming on intel GPUs. In 2021 IEEE/ACM International Symposium on Code Generation and Optimization (CGO) . IEEE, 289–300
Guei-Yuan Lueh, Kaiyu Chen, Gang Chen, Joel Fuentes, Wei-Yu Chen, Fangwen Fu, Hong Jiang, Hongzheng Li, and Daniel Rhee. 2021 · 2021
Closest in time.
DNNFusion: accelerating deep neural networks execution with advanced operator fusion. In Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation . 883–898
Wei Niu, Jiexiong Guan, Yanzhi Wang, Gagan Agrawal, and Bin Ren. 2021 · 2021
Closest in time.
A Flexible Approach to Autotuning Multi-Pass Machine Learning Compilers. In 2021 30th International Conference on Parallel Architectures and Compilation Techniques (PACT) . IEEE, 1–16
Phitchaya Mangpo Phothilimthana, Amit Sabne, Nikhil Sarda, Karthik Srinivasa Murthy, Yanqi Zhou, Christof Angermueller, Mike Burrows, Sudip Roy, Ketan Mandke, Rezsa Farahani, et al · 2021
Closest in time.
Demystifying TensorRT: Characterizing Neural Network Inference Engine on Nvidia Edge Devices. In 2021 IEEE International Symposium on Workload Characterization (IISWC) . 226–237
Omais Shafi, Chinmay Rai, Rijurekha Sen, and Gayathri Ananthanarayanan. 2021 · 2021
Closest in time.
Equality Saturation for Tensor Graph Superoptimization
Yichen Yang, Phitchaya Mangpo Phothilimtha, Yisu Remy Wang, Max Willsey, Sudip Roy, and Jacques Pienaar. 2021 · 2021
Closest in time.
DLPack: Open In Memory Tensor Structure
[n.d.] · 2022
Closest in time.
NVIDIA TensorRT Deserialization
[n.d.]b · 2022
Closest in time.
Optimizing GPU Deep Learning Operators with Polyhedral Scheduling Constraint Injection. In 2022 IEEE/ACM International Symposium on Code Generation and Optimization (CGO) . IEEE, 313–324
Cedric Bastoul, Zhen Zhang, Harenome Razanajato, Nelson Lossing, Adilla Susungi, Javier de Juan, Etienne Filhol, Baptiste Jarry, Gianpietro Consolaro, and Renwei Zhang. 2022 · 2022
Closest in time.
Automatic horizontal fusion for GPU kernels. In 2022 IEEE/ACM International Symposium on Code Generation and Optimization (CGO) . IEEE, 14–27
Ao Li, Bojian Zheng, Gennady Pekhimenko, and Fan Long. 2022 · 2022
Closest in time.
Alpa: Automating Inter-and Intra-Operator Parallelism for Distributed Deep Learning
Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Joseph E Gonzalez, et al · 2022
Closest in time.
Persistent rnns: Stashing recurrent weights on-chip. In International Conference on Machine Learning . PMLR, 2024–2033
Greg Diamos, Shubho Sengupta, Bryan Catanzaro, Mike Chrzanowski, Adam Coates, Erich Elsen, Jesse Engel, Awni Hannun, and Sanjeev Satheesh. 2016 · 2033
Closest in time.