Fetching the paper…
Reading the bibliography…
The difficulty of deploying various deep learning (DL) models on diverse DL hardware has boosted the research and development of DL compilers in the community.
Machine learning in compiler optimization
Zheng Wang and Michael O’Boyle. 2018 · 1901
Earlier work this paper cites.
Relay: A High-Level Compiler for Deep Learning
Jared Roesch, Steven Lyubomirsky, Marisa Kirisame, Logan Weber, Josh Pollock, Luis Vega, Ziheng Jiang, Tianqi Chen, Thierry Moreau, and Zachary Tatlock. 2019 · 1904
Earlier work this paper cites.
A Study of BFLOAT16 for Deep Learning Training
Dhiraj Kalamkar, Dheevatsa Mudigere, Naveen Mellempudi, Dipankar Das, Kunal Banerjee, Sasikanth Avancha, Dharma Teja Vooturi, Nataraj Jammalamadaka, Jianyu Huang, Hector Yuen, Jiyan Yang, Jongsoo Park, Alexander Heinecke, Evangelos Georganas, Sudarshan Srinivasan, Abhisek Kundu, Misha Smelyanskiy, Bharat Kaul, and Pradeep Dubey. 2019 · 1905
Earlier work this paper cites.
FusionStitching: Boosting Execution Efficiency of Memory Intensive Computations for DL Workloads
Guoping Long, Jun Yang, and Wei Lin. 2019 · 1911
Earlier work this paper cites.
LISP 1.5 programmer’s manual
John McCarthy and Michael I Levin. 1965 · 1965
Earlier work this paper cites.
Dependence Graphs and Compiler Optimizations. In Proceedings of the 8th ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages (POPL ’81) . Association for Computing Machinery, New York, NY, USA, 207–218
D. J. Kuck, R. H. Kuhn, D. A. Padua, B. Leasure, and M. Wolfe. 1981 · 1981
Earlier work this paper cites.
Compilers, principles, techniques
Alfred V Aho, Ravi Sethi, and Jeffrey D Ullman. 1986 · 1986
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams. 1986 · 1986
Earlier work this paper cites.
Parametric integer programming
P. Feautrier. 1988 · 1988
Earlier work this paper cites.
Genetic Algorithms in Search, Optimization and Machine Learning (1st ed.)
David E. Goldberg. 1989 · 1989
Earlier work this paper cites.
Efficiently computing static single assignment form and the control dependence graph
Ron Cytron, Jeanne Ferrante, Barry K Rosen, Mark N Wegman, and F Kenneth Zadeck. 1991 · 1991
Earlier work this paper cites.
Simulated annealing
Dimitris Bertsimas, John Tsitsiklis, et al · 1993
Earlier work this paper cites.
The Omega calculator and library, version 1.1. 0
Wayne Kelly, Vadim Maslov, William Pugh, Evan Rosser, Tatiana Shpeisman, and Dave Wonnacott. 1996 · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Revised 5 report on the algorithmic language Scheme
Harold Abelson, R. Kent Dybvig, Christopher T. Haynes, Guillermo Juan Rozas, NI Adams, Daniel P. Friedman, E Kohlbecker, GL Steele, David H Bartley, Robert Halstead, et al · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner. 1998 · 1998
Earlier work this paper cites.
PolyLib: A library for manipulating parameterized polyhedra
Vincent Loechner. 1999 · 1999
Earlier work this paper cites.
Foundations of statistical natural language processing
Christopher D Manning, Christopher D Manning, and Hinrich Schütze. 1999 · 1999
Earlier work this paper cites.
Chameleon: Adaptive Code Optimization for Expedited Deep Neural Network Compilation
Byung Hoon Ahn, Prannoy Pilligundla, Amir Yazdanbakhsh, and Hadi Esmaeilzadeh. 2020b · 2001
Earlier work this paper cites.
Computer vision: a modern approach
David A Forsyth and Jean Ponce. 2002 · 2002
Earlier work this paper cites.
MLIR: A Compiler Infrastructure for the End of Moore’s Law
Chris Lattner, Mehdi Amini, Uday Bondhugula, Albert Cohen, Andy Davis, Jacques Pienaar, River Riddle, Tatiana Shpeisman, Nicolas Vasilache, and Oleksandr Zinenko. 2020 · 2002
Earlier work this paper cites.
Ordering Chaos: Memory-Aware Scheduling of Irregularly Wired Neural Networks for Edge Devices
Byung Hoon Ahn, Jinwon Lee, Jamie Menjay Lin, Hsin-Pai Cheng, Jilei Hou, and Hadi Esmaeilzadeh. 2020a · 2003
Earlier work this paper cites.
LLVM: A compilation framework for lifelong program analysis & transformation. In Proceedings of the international symposium on Code generation and optimization: feedback-directed and runtime optimization . IEEE Computer Society, IEEE Computer Society, San Jose, CA, USA, 75
Chris Lattner and Vikram Adve. 2004 · 2004
Earlier work this paper cites.
The Parma Polyhedra Library: Toward a complete set of numerical abstractions for the analysis and verification of hardware and software systems
Roberto Bagnara, Patricia M Hill, and Enea Zaffanella. 2006 · 2006
Earlier work this paper cites.
Polyhedral code generation in the real world. In International Conference on Compiler Construction . Springer, Springer, Vienna, Austria, 185–201
Nicolas Vasilache, Cédric Bastoul, and Albert Cohen. 2006 · 2006
Earlier work this paper cites.
JavaScript bible
Danny Goodman. 2007 · 2007
Earlier work this paper cites.
isl: An integer set library for the polyhedral model. In International Congress on Mathematical Software . Springer, Springer, Kobe, Japan, 299–302
Sven Verdoolaege. 2010 · 2010
Earlier work this paper cites.
Torch7: A Matlab-like Environment for Machine Learning. In BigLearn, NIPS Workshop . Curran Associates, Granada, Spain, 6
R. Collobert, K. Kavukcuoglu, and C. Farabet. 2011 · 2011
Earlier work this paper cites.
Chainer: A deep learning framework for accelerating the research cycle. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . ACM, Anchorage, AK, USA, 2002–2011
Seiya Tokui, Ryosuke Okuta, Takuya Akiba, Yusuke Niitani, Toru Ogawa, Shunta Saito, Shuji Suzuki, Kota Uenishi, Brian Vogel, and Hiroyuki Yamazaki Vincent. 2019 · 2011
Earlier work this paper cites.
Polyhedra scanning revisited. In Proceedings of the 33rd ACM SIGPLAN conference on Programming Language Design and Implementation . ACM, Beijing, China, 499–508
Chun Chen. 2012 · 2012
Earlier work this paper cites.
Syntax Matters: Writing abstract computations in F#
Tomas Petricek and Don Syme. 2012 · 2012
Earlier work this paper cites.
Halide: A Language and Compiler for Optimizing Parallelism, Locality, and Recomputation in Image Processing Pipelines. In Proceedings of the 34th ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI ’13) . Association for Computing Machinery, New York, NY, USA, 519–530
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe. 2013 · 2013
Earlier work this paper cites.
Polyhedral parallel code generation for CUDA
Sven Verdoolaege, Juan Carlos Juega, Albert Cohen, José Ignacio Gómez, Christian Tenllado, and Francky Catthoor. 2013 · 2013
Earlier work this paper cites.
cuDNN: Efficient Primitives for Deep Learning
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer. 2014 · 2014
Cited alongside, same era.
Generative Adversarial Networks
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Caffe: Convolutional architecture for fast feature embedding. In Proceedings of the 22nd ACM international conference on Multimedia . ACM, ACM, Orlando, FL, USA, 675–678
Yangqing Jia, Evan Shelhamer, Jeff Donahue, Sergey Karayev, Jonathan Long, Ross Girshick, Sergio Guadarrama, and Trevor Darrell. 2014 · 2014
Cited alongside, same era.
Non-affine extensions to polyhedral code generation. In Proceedings of Annual IEEE/ACM International Symposium on Code Generation and Optimization . ACM, Orlando, FL, USA, 185–194
Anand Venkat, Manu Shantharam, Mary Hall, and Michelle Mills Strout. 2014 · 2014
Cited alongside, same era.
Comparative Study of Deep Learning Software Frameworks
Soheil Bahrampour, Naveen Ramakrishnan, Lukas Schott, and Mohak Shah. 2015 · 2015
Fashionable Modelling with Flux
Michael Innes, Elliot Saba, Keno Fischer, Dhairya Gandhi, Marco Concetto Rudilosso, Neethu Mariya Joy, Tejan Karmali, Avik Pal Singh, and Viral Shah. 2018 · 2018
Later among the works it cites.
C-GOOD: C-code generation framework for optimized on-device deep learning. In Proceedings of the International Conference on Computer-Aided Design . ACM, ACM, San Diego, CA, USA, 105
Duseok Kang, Euiseok Kim, Inpyo Bae, Bernhard Egger, and Soonhoi Ha. 2018 · 2018
Later among the works it cites.
FusionStitching: Deep Fusion and Code Generation for Tensorflow Computations on GPUs
Guoping Long, Jun Yang, Kai Zhu, and Wei Lin. 2018 · 2018
Later among the works it cites.
ALAMO: FPGA acceleration of deep learning algorithms with a modularized RTL compiler
Yufei Ma, Naveen Suda, Yu Cao, Sarma Vrudhula, and Jae-sun Seo. 2018 · 2018
Later among the works it cites.
Deep private-feature extraction
Seyed Ali Osia, Ali Taheri, Ali Shahin Shamsabadi, Kleomenis Katevas, Hamed Haddadi, and Hamid R Rabiee. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang. 2015 · 2015
Cited alongside, same era.
Tensor-matrix products with a compressed sparse tensor. In Proceedings of the 5th Workshop on Irregular Applications: Architectures and Algorithms . ACM, Austin, Texas, USA, 1–7
Shaden Smith and George Karypis. 2015 · 2015
Cited alongside, same era.
Loop and Data Transformations for Sparse Matrix Code. In Proceedings of the 36th ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI ’15) . ACM, Portland, OR, USA, 521–532
Anand Venkat, Mary Hall, and Michelle Strout. 2015 · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning. In 12th { \{ USENIX } \} Symposium on Operating Systems Design and Implementation ( { \{ OSDI } \} 16) . USENIX Association, Savannah, GA, USA, 265–283
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Cited alongside, same era.
Large-scale item categorization in e-commerce using multiple recurrent neural networks. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . ACM, San Francisco, CA, USA, 107–115
Jung-Woo Ha, Hyuna Pyo, and Jeonghee Kim. 2016 · 2016
Cited alongside, same era.
Cambricon: An Instruction Set Architecture for Neural Networks. In 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) . IEEE Computer Society, Seoul, South Korea, 393–405
S. Liu, Z. Du, J. Tao, D. Han, T. Luo, Y. Xie, Y. Chen, and T. Chen. 2016 · 2016
Cited alongside, same era.
Automatic code generation of convolutional neural networks in FPGA implementation. In 2016 International Conference on Field-Programmable Technology (FPT) . IEEE, IEEE, Xi’an, China, 61–68
Zhiqiang Liu, Yong Dou, Jingfei Jiang, and Jinwei Xu. 2016 · 2016
Cited alongside, same era.
Glow: Graph Lowering Compiler Techniques for Neural Networks
Nadav Rotem, Jordan Fix, Saleem Abdulrasool, Garret Catron, Summer Deng, Roman Dzhabarov, Nick Gibson, James Hegeman, Meghan Lele, Roman Levenstein, Jack Montgomery, Bert Maher, Satish Nadathur, Jakob Olesen, Jongsoo Park, Artem Rakhov, Misha Smelyanskiy, and Man Wang. 2018 · 2018
Later among the works it cites.
Automatic differentiation in ML: Where we are and where we should be going. In Advances in neural information processing systems . Curran Associates, Montréal, Canada, 8757–8767
Bart Van Merriënboer, Olivier Breuleux, Arnaud Bergeron, and Pascal Lamblin. 2018 · 2018
Later among the works it cites.
Automatic differentiation in ML: Where we are and where we should be going
Bart van Merriënboer, Olivier Breuleux, Arnaud Bergeron, and Pascal Lamblin. 2018 · 2018
Later among the works it cites.
Tensor comprehensions: Framework-agnostic high-performance machine learning abstractions
Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen. 2018 · 2018
Later among the works it cites.
Toolflows for Mapping Convolutional Neural Networks on FPGAs
Stylianos I. Venieris, Alexandros Kouris, and Christos-Savvas Bouganis. 2018 · 2018
Later among the works it cites.
Dynamic Control Flow in Large-Scale Machine Learning. In Proceedings of the Thirteenth EuroSys Conference (EuroSys ’18) . Association for Computing Machinery, New York, NY, USA, Article Article 18, 15 pages
Yuan Yu, Martín Abadi, Paul Barham, Eugene Brevdo, Mike Burrows, Andy Davis, Jeff Dean, Sanjay Ghemawat, Tim Harley, Peter Hawkins, Michael Isard, Manjunath Kudlur, Rajat Monga, Derek Murray, and Xiaoqiang Zheng. 2018 · 2018
Later among the works it cites.
Hardware Compilation of Deep Neural Networks: An Overview. In 2018 IEEE 29th International Conference on Application-specific Systems, Architectures and Processors (ASAP) . IEEE Computer Society, Milano, Italy, 1–8
R. Zhao, S. Liu, H. Ng, E. Wang, J. J. Davis, X. Niu, X. Wang, H. Shi, G. A. Constantinides, P. Y. K. Cheung, and W. Luk. 2018 · 2018
Later among the works it cites.
Privacy for Rescue: A New Testimony Why Privacy is Vulnerable In Deep Models
Ruiyuan Gao, Ming Dun, Hailong Yang, Zhongzhi Luan, and Depei Qian. 2019 · 2019
Later among the works it cites.
Dissecting the Graphcore IPU Architecture via Microbenchmarking
Zhe Jia, Blake Tillman, Marco Maggioni, and Daniele Paolo Scarpazza. 2019 · 2019
Later among the works it cites.
Learned TPU Cost Model for XLA Tensor Programs. In Proceedings of the Workshop on ML for Systems at NeurIPS 2019 . Curran Associates, Vancouver, Canada, 1–6
Samuel Kaufman, Phitchaya Mangpo Phothilimthana, and Mike Burrows. 2019 · 2019
Later among the works it cites.
DaVinci: A Scalable Architecture for Neural Network Computing. In 2019 IEEE Hot Chips 31 Symposium (HCS) . IEEE, IEEE, Cupertino, CA, USA, 1–44
Heng Liao, Jiajin Tu, Jing Xia, and Xiping Zhou. 2019 · 2019
Later among the works it cites.
Optimizing { \{ CNN } \} Model Inference on CPUs. In 2019 { \{ USENIX } \} Annual Technical Conference ( { \{ USENIX } \} { \{ ATC } \} 19) . USENIX Association, Renton, WA, USA, 1025–1040
Yizhi Liu, Yao Wang, Ruofei Yu, Mu Li, Vin Sharma, and Yida Wang. 2019 · 2019
Later among the works it cites.
Performance Evaluation of Deep Learning frameworks on Computer Vision problems. In 2019 3rd International Conference on Trends in Electronics and Informatics (ICOEI) . IEEE, Tirunelveli, India, India, 670–674
Madhumitha Nara, BR Mukesh, Preethi Padala, and Bharath Kinnal. 2019 · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems . Curran Associates, Vancouver, BC, Canada, 8024–8035
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Benchmarking tpu, gpu, and cpu platforms for deep learning
Gu-Yeon Wei, David Brooks, et al · 2019
Later among the works it cites.
An In-depth Comparison of Compilers for Deep Neural Networks on Hardware. In 2019 IEEE International Conference on Embedded Software and Systems (ICESS) . IEEE, IEEE, Las Vegas, NV, USA, 1–8
Yu Xing, Jian Weng, Yushun Wang, Lingzhi Sui, Yi Shan, and Yu Wang. 2019 · 2019
Later among the works it cites.
Announcing Hanguang 800: Alibaba’s First AI-Inference Chip
Alibaba. 2019 · 2020
Closest in time.
AWS Inferentia
Amazon. 2018 · 2020
Closest in time.
PaddlePaddle Github repository
Baidu. 2016 · 2020
Closest in time.
Polyhedral Compilation
Tobias Grosser. 2000 · 2020
Closest in time.
Nervana Neural Network Processor
Intel. 2019 · 2020
Closest in time.
The Deep Learning Compiler: A Comprehensive Survey
Mingzhen Li, Yi Liu, Xiaoyan Liu, Qingxiao Sun, Xin You, Hailong Yang, Zhongzhi Luan, and Depei Qian. 2020 · 2020
Closest in time.
ONNX Github repository
Microsoft. 2017 · 2020
Closest in time.
Shredder: Learning noise distributions to protect inference privacy. In Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems . ACM, Lausanne, Switzerland, 3–18
Fatemehsadat Mireshghallah, Mohammadkazem Taram, Prakash Ramrakhyani, Ali Jalali, Dean Tullsen, and Hadi Esmaeilzadeh. 2020 · 2020
Closest in time.
DLProf User-guide
NVIDIA. 2019a · 2020
Closest in time.
Nvidia Turing Architecture
NVIDIA. 2019b · 2020
Closest in time.
TensorRT Github repository
NVIDIA. 2019c · 2020
Closest in time.