Fetching the paper…
Reading the bibliography…
Deep learning models with convolutional and recurrent networks are now ubiquitous and analyze massive amounts of audio, image, video, text and graph data, with applications in automatic translation, speech-to-text, scene understanding, ranking user preferences, ad placement, etc.
Perceptrons: an introduction to computational geometry; 1st ed
M. L. Minsky and S. Papert · 1969
Earlier work this paper cites.
Genetic Algorithms in Search, Optimization and Machine Learning
D. E. Goldberg · 1989
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
Y. LeCun, B. E. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. E. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
Dataflow Analysis of Array and Scalar References
P. Feautrier · 1991
Earlier work this paper cites.
The ALPHA language and its use for the design of systolic arrays
H. Le Verge, C. Mauras, and P. Quinton · 1991
Earlier work this paper cites.
Some Efficient Solutions to the Affine Scheduling Problem. Part II. Multidimensional Time
P. Feautrier · 1992
Earlier work this paper cites.
Static Analysis of Upper and Lower Bounds on Dependences and Parallelism
W. Pugh and D. Wonnacott · 1994
Earlier work this paper cites.
FFTW: An adaptive software architecture for the FFT
M. Frigo and S. G. Johnson · 1998
Earlier work this paper cites.
Active libraries: Rethinking the roles of compilers and libraries
T. Veldhuizen and E. Gannon · 1998
Earlier work this paper cites.
Automatically tuned linear algebra software
R. C. Whaley and J. J. Dongarra · 1998
Earlier work this paper cites.
Principles of Program Analysis
F. Nielson, H. Nielson, and C. Hankin · 1999
Earlier work this paper cites.
Oolala: An object oriented analysis and design of numerical linear algebra
M. Luján, T. L. Freeman, and J. R. Gurd · 2000
Earlier work this paper cites.
Overcoming the challenges to feedback-directed optimization (keynote talk)
M. D. Smith · 2000
Earlier work this paper cites.
Lush reference manual
L. Bottou and Y. LeCun · 2002
Earlier work this paper cites.
Optimizing Compilers for Modern Architectures: A Dependence-Based Approach
K. Kennedy and J. R. Allen · 2002
Earlier work this paper cites.
Run-time code generation in C++ as a foundation for domain-specific optimisation
O. Beckmann, A. Houghton, P. H. J. Kelly, and M. Mellor · 2003
Earlier work this paper cites.
Code Generation in the Polyhedral Model Is Easier Than You Think
C. Bastoul · 2004
Earlier work this paper cites.
In search of a program generator to implement generic transformations for high-performance computing
A. Cohen, S. Donadio, M. J. Garzarán, C. Herrmann, O. Kiselyov, and D. Padua · 2004
Earlier work this paper cites.
Spiral: A generator for platform-adapted libraries of signal processing alogorithms
M. Püschel, J. M. F. Moura, B. Singer, J. Xiong, J. Johnson, D. Padua, M. Veloso, and R. W. Johnson · 2004
Earlier work this paper cites.
Facilitating the search for compositions of program transformations
A. Cohen, S. Girbal, D. Parello, M. Sigler, O. Temam, and N. Vasilache · 2005
Earlier work this paper cites.
A language for the compact representation of multiple program versions
S. Donadio, J. Brodman, T. Roeder, K. Yotov, D. Barthou, A. Cohen, M. J. Garzarán, D. Padua, and K. Pingali · 2005
Earlier work this paper cites.
Semi-Automatic Composition of Loop Transformations for Deep Parallelism and Memory Hierarchies
S. Girbal, N. Vasilache, C. Bastoul, A. Cohen, D. Parello, M. Sigler, and O. Temam · 2006
Earlier work this paper cites.
A compiler framework for optimization of affine loop nests for GPGPUs
M. M. Baskaran, U. Bondhugula, S. Krishnamoorthy, J. Ramanujam, A. Rountev, and P. Sadayappan · 2008
Earlier work this paper cites.
A Practical Automatic Polyhedral Parallelizer and Locality Optimizer
U. Bondhugula, A. Hartono, J. Ramanujam, and P. Sadayappan · 2008
Earlier work this paper cites.
Chill: A framework for composing high-level loop transformations
C. Chen, J. Chame, and M. Hall · 2008
Earlier work this paper cites.
Automating the generation of composed linear algebra kernels
G. Belter, E. R. Jessup, I. Karlin, and J. G. Siek · 2009
Earlier work this paper cites.
Large-scale deep unsupervised learning using graphics processors
R. Raina, A. Madhavan, and A. Y. Ng · 2009
Earlier work this paper cites.
The Polyhedral Model Is More Widely Applicable Than You Think
M.-W. Benabderrahmane, L.-N. Pouchet, A. Cohen, and C. Bastoul · 2010
Earlier work this paper cites.
Deep big simple neural nets excel on handwritten digit recognition
D. C. Ciresan, U. Meier, L. M. Gambardella, and J. Schmidhuber · 2010
Earlier work this paper cites.
Factorization machines
S. Rendle · 2010
Earlier work this paper cites.
Lightweight modular staging: A pragmatic approach to runtime code generation and compiled dsls
T. Rompf and M. Odersky · 2010
Earlier work this paper cites.
Isl: An integer set library for the polyhedral model
S. Verdoolaege · 2010
Cited alongside, same era.
Polyhedron Model
P. Feautrier and C. Lengauer · 2011
Cited alongside, same era.
R-Stream Compiler
B. Meister, N. Vasilache, D. Wohlford, M. M. Baskaran, A. Leung, and R. Lethin · 2011
Cited alongside, same era.
Loop transformations: Convexity, pruning and optimization
L.-N. Pouchet, U. Bondhugula, C. Bastoul, A. Cohen, J. Ramanujam, P. Sadayappan, and N. Vasilache · 2011
Cited alongside, same era.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Cited alongside, same era.
Model-driven SIMD code generation for a multi-resolution tensor kernel
K. Stock, T. Henretty, I. Murugandi, P. Sadayappan, and R. Harrison · 2011
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Later among the works it cites.
Pytorch examples
J. Johnson · 2015
Later among the works it cites.
Polymage: Automatic optimization for image processing pipelines
R. T. Mullapudi, V. Vasista, and U. Bondhugula · 2015
Later among the works it cites.
TensorFlow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, et al · 2016
Later among the works it cites.
Opening Polyhedral Compiler’s Black Box
L. Bagnères, O. Zinenko, S. Huot, and C. Bastoul · 2016
Later among the works it cites.
The Pluto+ Algorithm: A Practical Approach for Parallelization and Locality Optimization of Affine Loop Nests
U. Bondhugula, A. Acharya, and A. Cohen · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Counting Affine Calculator and Applications
S. Verdoolaege · 2011
Cited alongside, same era.
Polyhedra Scanning Revisited
C. Chen · 2012
Cited alongside, same era.
Implementing neural networks efficiently
R. Collobert, K. Kavukcuoglu, and C. Farabet · 2012
Cited alongside, same era.
Optimization techniques for efficient hta programs
B. B. Fraguela, G. Bikshandi, J. Guo, M. J. Garzarán, D. Padua, and C. von Praun · 2012
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2012
Cited alongside, same era.
ImageNet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
Later among the works it cites.
Wide & deep learning for recommender systems
H. Cheng, L. Koc, J. Harmsen, T. Shaked, T. Chandra, H. Aradhye, G. Anderson, G. Corrado, W. Chai, M. Ispir, R. Anil, Z. Haque, L. Hong, V. Jain, X. Liu, and H. Shah · 2016
Later among the works it cites.
Simit: A language for physical simulation
F. Kjolstad, S. Kamil, J. Ragan-Kelley, D. I. W. Levin, S. Sueda, D. Chen, E. Vouga, D. M. Kaufman, G. Kanwar, W. Matusik, and S. Amarasinghe · 2016
Later among the works it cites.
Automatically scheduling halide image processing pipelines
R. T. Mullapudi, A. Adams, D. Sharlet, J. Ragan-Kelley, and K. Fatahalian · 2016
Later among the works it cites.
Automatically scheduling halide image processing pipelines
R. T. Mullapudi, A. Adams, D. Sharlet, J. Ragan-Kelley, and K. Fatahalian · 2016
Later among the works it cites.
A review of relational machine learning for knowledge graphs
M. Nickel, K. Murphy, V. Tresp, and E. Gabrilovich · 2016
Later among the works it cites.
A basic linear algebra compiler for structured matrices
D. G. Spampinato and M. Püschel · 2016
Later among the works it cites.
Theano: A Python framework for fast computation of mathematical expressions
Theano Development Team · 2016
Later among the works it cites.
Latte: A language, compiler, and runtime for elegant and efficient deep neural networks
L. Truong, R. Barik, E. Totoni, H. Liu, C. Markley, A. Fox, and T. Shpeisman · 2016
Later among the works it cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. B. Girshick, P. Dollár, Z. Tu, and K. He · 2016
Later among the works it cites.
Google Vizier: A Service for Black-Box Optimization
D. Golovin, B. Solnik, S. Moitra, G. Kochanski, J. E. Karro, and D. Sculley, editors · 2017
Later among the works it cites.
https://developers.google.com/protocol-buffers/docs/overview
Protocol buffers developer guide · 2017
Later among the works it cites.
https://www.tensorflow.org/performance/xla
XLA: Domain-specific compiler for linear algebra to optimizes tensorflow computations · 2017
Later among the works it cites.
Accurate, large minibatch SGD: training ImageNet in 1 hour
P. Goyal, P. Dollár, R. B. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Later among the works it cites.
personal communication, 2017
V. Grover · 2017
Later among the works it cites.
The tensor algebra compiler
F. Kjolstad, S. Kamil, S. Chou, D. Lugato, and S. Amarasinghe · 2017
Later among the works it cites.
https://www.microsoft.com/en-us/research/blog/microsoft-unveils-project-brainwave
Microsoft unveils project brainwave for real-time ai · 2017
Later among the works it cites.
https://devblogs.nvidia.com/parallelforall/deploying-deep-learning-nvidia-tensorrt
Deploying deep neural networks with Nvidia TensorRT · 2017
Later among the works it cites.
https://devblogs.nvidia.com/parallelforall/inside-volta
Inside Volta: The world’s most advanced data center GPU · 2017
Later among the works it cites.
Polyhedral optimization of tensorflow computation graphs
B. Pradelle, B. Meister, M. Baskaran, J. Springer, and R. Lethin · 2017
Later among the works it cites.
https://pytorch.org
PyTorch: Tensors and dynamic neural networks in python with strong GPU acceleration · 2017
Later among the works it cites.
Scheduling for PPCG
S. Verdoolaege and G. Janssens · 2017
Later among the works it cites.
Deep Interest Network for Click-Through Rate Prediction
G. Zhou, C. Song, X. Zhu, X. Ma, Y. Yan, X. Dai, H. Zhu, J. Jin, H. Li, and K. Gai · 2017
Later among the works it cites.
Unified Polyhedral Modeling of Temporal and Spatial Locality
O. Zinenko, S. Verdoolaege, C. Reddy, J. Shirako, T. Grosser, V. Sarkar, and A. Cohen · 2017
Later among the works it cites.
TVM: End-to-end optimization stack for deep learning, 2018, arXiv:1802.04799
T. Chen, T. Moreau, Z. Jiang, H. Shen, E. Yan, L. Wang, Y. Hu, L. Ceze, C. Guestrin, and A. Krishnamurthy · 2018
Closest in time.