Fetching the paper…
Reading the bibliography…
In recent years, there is a surge on machine learning applications in industry.
Programming parallel algorithms
Guy, B · 1996
Earlier work this paper cites.
Optimizing Compilers for Modern Architectures: A Dependence-based Approach
Kennedy, K. and Allen, J. R · 2002
Earlier work this paper cites.
“improving effective bandwidth through compiler enhancement of global cache reuse
Ding, C. and Kennedy, K · 2004
Earlier work this paper cites.
Kernel weaver: Automatically fusing database primitives for efficient gpu computation
Wu, H. C., Diamos, G., Cadambi, S., and Yalamanchili, S · 2012
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Jia, Y. Q., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, G., and Darrell, T · 2014
Earlier work this paper cites.
Loo.py: transformation-based code generation for gpus and cpus
Klöckner, A · 2014
Earlier work this paper cites.
Scalable kernel fusion for memory-bound gpu applications
Wahib, M. and Maruyama, N · 2014
Earlier work this paper cites.
An introduction to computational networks and the computational network toolkit
Yu, D., Eversole, A., Seltzer, M., Yao, K., Kuchaiev, O., Zhang, Y., Seide, F., Huang, Z. H., Guenter, B., Wang, H. M., Droppo, J., Zweig, G., Rossbach, C., Gao, J., Stolcke, A., Currey, J., Slaney, M., Chen, G. G., Agarwal, A., Basoglu, C., Padmilac, M., Kamenev, A., Ivanov, V., Cypher, S., Parthasarathi, M., Mitra, B., Peng, B. L., and Huang, X. D · 2014
Cited alongside, same era.
URL https://github.com/torch/nn
Torch nn, 2015 · 2015
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Cited alongside, same era.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
Chen, T., Li, M., Li, Y., Lin, M., Wang, N., Wang, M., Xiao, T., Xu, B., Zhang, C., and Zhang, Z · 2015
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Later among the works it cites.
Optimal dnn primitive selection with partitioned boolean quadratic programming
Anderson, A. and Gregg, D · 2018
Closest in time.
Tvm: An automated end-to-end optimizing compiler for deep learning
Chen, T., Moreau, T., Jiang, Z., Zheng, L., Yan, E., Shen, H., Cowan, M., Wang, L., Hu, Y., Ceze, L., Guestrin, C., and Krishnamurthy, A · 2018
Closest in time.
Halide: decoupling algorithms from schedules for high-performance image processing
Kelley, J. R., Adams, A., Sharlet, D., Barnes, C., Paris, S., Levoy, M., Amarasinghe, S., and Durand, F · 2018
Closest in time.
Bringing tvm into tensorflow for optimizing neural machine translation on gpu
PAI · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Highly optimized code generation for stencil codes with computation reuse for gpus
Ma, W. J., Gao, K., and Long, G. P · 2016
Cited alongside, same era.
Latte: A language, compiler, and runtime for elegant and efficient deep neural networks
Truong, L., Barik, R., Totoni, E., Liu, H., Markley, C., Fox, A., and Shpeisman, T · 2016
Cited alongside, same era.
Boda: A holistic approach for implementing neural network computations
Moskewicz, M. W., Jannesari, A., and Keutzer, K · 2017
Cited alongside, same era.
URL https://github.com/dmlc/tvm
Tvm: Open deep learning compiler stack
Cited in the paper.
Tensorflow-examples
aymericdamien
Cited in the paper.
A simple, fast dominance algorithm
Cooper, K. D., Harvey, T. J., and Kennedy, K
Cited in the paper.
Program generation for small-scale linear algebra applications
Spampinato, D. G., Traver, D. F., Bientinesi, P., and Püschel, M · 2018
Closest in time.
Attention focusing for neural machine translation by bridging source and target embeddings
Xiong, D. Y., Li, J. H., Branco, A., Kuang, S. H., and Luo, W. H · 2018
Closest in time.