Fetching the paper…
Reading the bibliography…
There is an increasing need to bring machine learning to a wide diversity of hardware devices.
Decoupled access/execute computer architectures
1982
Earlier work this paper cites.
Optimization by simulated annealing
1983
Earlier work this paper cites.
Improving direct-mapped cache performance by the addition of a small fully-associative cache and prefetch buffers
1990
Earlier work this paper cites.
Effective hardware-based data prefetching for high-performance processors
1995
Earlier work this paper cites.
Simultaneous multithreading: a platform for next-generation processors
1997
Earlier work this paper cites.
Fftw: an adaptive software architecture for the fft
1998
Earlier work this paper cites.
Automatically tuned linear algebra software
1998
Earlier work this paper cites.
Roofline: An insightful visual performance model for multicore architectures
2009
Earlier work this paper cites.
Optiml: An implicitly parallel domain-specific language for machine learning
2011
Earlier work this paper cites.
Theano: new features and speed improvements
2012
Earlier work this paper cites.
Halide: A language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines
2013
Earlier work this paper cites.
Polyhedral parallel code generation for cuda
2013
Earlier work this paper cites.
An introduction to computational networks and the computational network toolkit
2014
Earlier work this paper cites.
Opentuner: An extensible framework for program autotuning
2014
Earlier work this paper cites.
Dadiannao: A machine-learning supercomputer
2014
Earlier work this paper cites.
Darkroom: Compiling high-level image processing code into hardware pipelines
2014
Earlier work this paper cites.
Loo.py: transformation-based code generation for GPUs and CPUs
2014
Cited alongside, same era.
Recurrent neural network regularization
2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
2015
Cited alongside, same era.
Pencil: A platform-neutral compute intermediate language for accelerator programming
2015
Cited alongside, same era.
MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems
2015
Cited alongside, same era.
Binaryconnect: Training deep neural networks with binary weights during propagations
Efficient hyperparameter optimization and infinitely many armed bandits
2016
Later among the works it cites.
Automatically scheduling halide image processing pipelines
2016
Later among the works it cites.
Xnor-net: Imagenet classification using binary convolutional neural networks
2016
Later among the works it cites.
From high-level deep neural models to fpgas
2016
Later among the works it cites.
FINN: A framework for fast, scalable binarized neural network inference
2016
Later among the works it cites.
Understanding Latency Hiding on GPUs
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
Pudiannao: A polyvalent machine learning accelerator
2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
2015
Cited alongside, same era.
Unsupervised representation learning with deep convolutional generative adversarial networks
2015
Cited alongside, same era.
Improved semantic representations from tree-structured long short-term memory networks
2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
2016
Cited alongside, same era.
Xgboost: A scalable tree boosting system
2016
Cited alongside, same era.
NVIDIA Tesla V100 GPU Architecture: The World’s Most Advanced Data Center GPU, 2017
2017
Later among the works it cites.
Futhark: Purely functional gpu-programming with nested parallelism and in-place array updates
2017
Later among the works it cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
2017
Later among the works it cites.
The tensor algebra compiler
2017
Later among the works it cites.
Weld: Rethinking the interface between data-intensive applications
2017
Later among the works it cites.
Lift: A functional data-parallel ir for high-performance gpu code generation
2017
Later among the works it cites.
High performance ultra-low-precision convolutions on mobile devices
2017
Later among the works it cites.
Dlvm: A modern compiler infrastructure for deep learning systems
2017
Later among the works it cites.