Fetching the paper…
Reading the bibliography…
High-performance tensor programs are crucial to guarantee efficient execution of deep neural networks.
Speech understanding systems: report of a steering committee
Mark F. Medress, Franklin S Cooper, Jim W. Forgie, CC Green, Dennis H. Klatt, Michael H. O’Malley, Edward P Neuburg, Allen Newell, DR Reddy, B Ritea, et al · 1977
Earlier work this paper cites.
Fftw: an adaptive software architecture for the fft
Matteo Frigo and Steven G Johnson · 1998
Earlier work this paper cites.
Automatically tuned linear algebra software
R Clinton Whaley and Jack J Dongarra · 1998
Earlier work this paper cites.
A practical automatic polyhedral parallelizer and locality optimizer
Uday Bondhugula, Albert Hartono, Jagannathan Ramanujam, and Ponnuswamy Sadayappan · 2008
Earlier work this paper cites.
Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe · 2013
Earlier work this paper cites.
Stochastic superoptimization
Eric Schkufza, Rahul Sharma, and Alex Aiken · 2013
Earlier work this paper cites.
Polyhedral parallel code generation for cuda
Sven Verdoolaege, Juan Carlos Juega, Albert Cohen, Jose Ignacio Gomez, Christian Tenllado, and Francky Catthoor · 2013
Earlier work this paper cites.
Opentuner: an extensible framework for program autotuning
Jason Ansel, Shoaib Kamil, Kalyan Veeramachaneni, Jonathan Ragan-Kelley, Jeffrey Bosboom, Una-May O’Reilly, and Saman Amarasinghe · 2014
Earlier work this paper cites.
cudnn: efficient primitives for deep learning
Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, and Evan Shelhamer · 2014
Earlier work this paper cites.
Mxnet: a flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Earlier work this paper cites.
Batch normalization: accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Alec Radford, Luke Metz, and Soumith Chintala · 2015
Earlier work this paper cites.
Multi-scale context aggregation by dilated convolutions
Fisher Yu and Vladlen Koltun · 2015
Earlier work this paper cites.
Tensorflow: a system for large-scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Earlier work this paper cites.
Xgboost: a scalable tree boosting system
Tianqi Chen and Carlos Guestrin · 2016
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benenson, Uwe Franke, Stefan Roth, and Bernt Schiele · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Fast algorithms for convolutional neural networks
Andrew Lavin and Scott Gray · 2016
Earlier work this paper cites.
Automatically scheduling halide image processing pipelines
Ravi Teja Mullapudi, Andrew Adams, Dillon Sharlet, Jonathan Ragan-Kelley, and Kayvon Fatahalian · 2016
Earlier work this paper cites.
Presburger formulas and polyhedral compilation
Sven Verdoolaege · 2016
Earlier work this paper cites.
Evolutionary algorithms: a critical review and its future prospects
Pradnya A Vikhar · 2016
Cited alongside, same era.
Augmented reality meets deep learning for car instance segmentation in urban scenes
Hassan Abu Alhaija, Siva Karthik Mustikovela, Lars Mescheder, Andreas Geiger, and Carsten Rother · 2017
Cited alongside, same era.
Mobilenets: efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Cited alongside, same era.
Intel® math kernel library for deep learning networks, 2017
Intel · 2017
Cited alongside, same era.
Deep learning with dynamic computation graphs
Moshe Looks, Marcello Herreshoff, DeLesley Hutchins, and Peter Norvig · 2017
Cited alongside, same era.
Onnx: open neural network exchange, 2019
Junjie Bai, Fang Lu, Ke Zhang, et al · 2019
Later among the works it cites.
Machine learning systems are stuck in a rut
Paul Barham and Michael Isard · 2019
Later among the works it cites.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker · 2019
Later among the works it cites.
Autophase: compiler phase-ordering for hls with deep reinforcement learning
Qijing Huang, Ameer Haj-Ali, William Moses, John Xiang, Ion Stoica, Krste Asanovic, and John Wawrzynek · 2019
Later among the works it cites.
Taso: optimizing deep learning computation with automatic generation of graph substitutions
Zhihao Jia, Oded Padon, James Thomas, Todd Warszawski, Matei Zaharia, and Alex Aiken · 2019
Later among the works it cites.
Optimizing cnn model inference on cpus
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nvidia tensor cores, 2017
Nvidia · 2017
Cited alongside, same era.
Nvidia tensorrt: programmable inference accelerator, 2017
Nvidia · 2017
Cited alongside, same era.
Parallel associative reductions in halide
Patricia Suriana, Andrew Adams, and Shoaib Kamil · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Tvm: an automated end-to-end optimizing compiler for deep learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, et al · 2018
Cited alongside, same era.
Learning to optimize tensor programs
Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy · 2018
Cited alongside, same era.
Bert: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Yizhi Liu, Yao Wang, Ruofei Yu, Mu Li, Vin Sharma, and Yida Wang · 2019
Later among the works it cites.
A hardware–software blueprint for flexible deep learning specialization
Thierry Moreau, Tianqi Chen, Luis Vega, Jared Roesch, Eddie Yan, Lianmin Zheng, Josh Fromm, Ziheng Jiang, Luis Ceze, Carlos Guestrin, et al · 2019
Later among the works it cites.
Pytorch: an imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
Relay: a high-level compiler for deep learning
Jared Roesch, Steven Lyubomirsky, Marisa Kirisame, Josh Pollock, Logan Weber, Ziheng Jiang, Tianqi Chen, Thierry Moreau, and Zachary Tatlock · 2019
Later among the works it cites.
Triton: an intermediate language and compiler for tiled neural network computations
Philippe Tillet, HT Kung, and David Cox · 2019
Later among the works it cites.
The next 700 accelerated layers: from mathematical expressions of network computation graphs to accelerated gpu kernels, automatically
Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary Devito, William S Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen · 2019
Later among the works it cites.
A unified optimization approach for cnn model inference on integrated gpus
Leyuan Wang, Zhi Chen, Yizhi Liu, Yao Wang, Lianmin Zheng, Mu Li, and Yida Wang · 2019
Later among the works it cites.
Neurovectorizer: end-to-end vectorization with deep reinforcement learning
Ameer Haj-Ali, Nesreen K Ahmed, Ted Willke, Yakun Sophia Shao, Krste Asanovic, and Ion Stoica · 2020
Closest in time.
Protuner: tuning programs with monte carlo tree search
Ameer Haj-Ali, Hasan Genc, Qijing Huang, William Moses, John Wawrzynek, Krste Asanović, and Ion Stoica · 2020
Closest in time.
Autophase: juggling hls phase orderings in random forests with deep reinforcement learning
Ameer Haj-Ali, Qijing Huang, William Moses, John Xiang, John Wawrzynek, Krste Asanovic, and Ion Stoica · 2020
Closest in time.
Featgraph: A flexible and efficient backend for graph neural network systems
Yuwei Hu, Zihao Ye, Minjie Wang, Jiali Yu, Da Zheng, Mu Li, Zheng Zhang, Zhiru Zhang, and Yida Wang · 2020
Closest in time.
Nimble: Efficiently compiling dynamic neural networks for model inference
Haichen Shen, Jared Roesch, Zhi Chen, Wei Chen, Yong Wu, Mu Li, Vin Sharma, Zachary Tatlock, and Yida Wang · 2020
Closest in time.
Flextensor: an automatic schedule exploration and optimization framework for tensor computation on heterogeneous system
Size Zheng, Yun Liang, Shuo Wang, Renze Chen, and Kaiwen Sheng · 2020
Closest in time.
Fusionstitching: boosting memory intensive computations for deep learning workloads
Zhen Zheng, Pengzhan Zhao, Guoping Long, Feiwen Zhu, Kai Zhu, Wenyi Zhao, Lansong Diao, Jun Yang, and Wei Lin · 2020
Closest in time.