Fetching the paper…
Reading the bibliography…
Accurate hardware performance models are critical to efficient code generation.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Learning to rank using gradient descent
Burges, C., Shaked, T., Renshaw, E., Lazier, A., Deeds, M., Hamilton, N., and Hullender, G · 2005
Earlier work this paper cites.
Chaos in Computer Performance
Berry, H., Gracia Pérez, D., and Temam, O · 2006
Earlier work this paper cites.
Fast Compiler Optimisation Evaluation Using Code-Feature Based Performance Prediction
Dubach, C., Cavazos, J., Franke, B., Fursin, G., O’Boyle, M. F., and Temam, O · 2007
Earlier work this paper cites.
End-to-end deep learning of optimization heuristics
Cummins, C., Petoumenos, P., Wang, Z., and Leather, H · 2017
Earlier work this paper cites.
Peephole: Predicting network performance before training
Deng, B., Yan, J., and Lin, D · 2017
Earlier work this paper cites.
Inductive Representation Learning on Large Graphs
Hamilton, W. L., Ying, Z., and Leskovec, J · 2017
Earlier work this paper cites.
In-Datacenter Performance Analysis of a Tensor Processing Unit
Jouppi, N. P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., Boyle, R., Cantin, P.-l., Chao, C., Clark, C., Coriell, J., Daley, M., Dau, M., Dean, J., Gelb, B., Ghaemmaghami, T. V., Gottipati, R., Gulland, W., Hagmann, R., Ho, C. R., Hogberg, D., Hu, J., Hundt, R., Hurt, D., Ibarz, J., Jaffey, A., Jaworski, A., Kaplan, A., Khaitan, H., Killebrew, D., Koch, A., Kumar, N., Lacy, S., Laudon, J., Law, J., Le, D., Leary, C., Liu, Z., Lucke, K., Lundin, A., MacKean, G., Maggiore, A., Mahony, M., Miller, K., Nagarajan, R., Narayanaswami, R., Ni, R., Nix, K., Norrie, T., Omernick, M., Penukonda, N., Phelps, A., Ross, J., Ross, M., Salek, A., Samadiani, E., Severn, C., Sizikov, G., Snelham, M., Souter, J., Steinberg, D., Swing, A., Tan, M., Thorson, G., Tian, B., Toma, H., Tuttle, E., Vasudevan, V., Walter, R., Wang, W., Wilcox, E., and Yoon, D. H · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, u., and Polosukhin, I · 2017
Earlier work this paper cites.
Learning to Optimize Tensor Programs
Chen, T., Zheng, L., Yan, E., Jiang, Z., Moreau, T., Ceze, L., Guestrin, C., and Krishnamurthy, A · 2018
Cited alongside, same era.
Tapas: Train-less accuracy predictor for architecture search, 2018
Istrate, R., Scheidegger, F., Mariani, G., Nikolopoulos, D., Bekas, C., and Malossi, A. C. I · 2018
Cited alongside, same era.
Predicting the computational cost of deep learning models
Justus, D., Brennan, J., Bonner, S., and McGough, A. S · 2018
Cited alongside, same era.
Learning to Optimize Halide with Tree Search and Random Programs
Adams, A., Ma, K., Anderson, L., Baghdadi, R., Li, T.-M., Gharbi, M., Steiner, B., Johnson, S., Fatahalian, K., Durand, F., and Ragan-Kelley, J · 2019
Cited alongside, same era.
Learning generalizable device placement algorithms for distributed machine learning
Addanki, R., Bojja Venkatakrishnan, S., Gupta, S., Mao, H., and Alizadeh, M · 2019
Cited alongside, same era.
Auto-Vectorization in GCC
Renas:relativistic evaluation of neural architecture search, 2019
Xu, Y., Wang, Y., Han, K., Jui, S., Xu, C., Tian, Q., and Xu, C · 2019
Later among the works it cites.
Graph transformer networks
Yun, S., Jeong, M., Kim, R., Kang, J., and Kim, H. J · 2019
Later among the works it cites.
Chameleon: Adaptive code optimization for expedited deep neural network compilation
Ahn, B. H., Pilligundla, P., Yazdanbakhsh, A., and Esmaeilzadeh, H · 2020
Closest in time.
Programl: Graph-based deep learning for program optimization and analysis, 2020
Cummins, C., Fisches, Z. V., Ben-Nun, T., Hoefler, T., and Leather, H · 2020
Closest in time.
Improving the accuracy, scalability, and performance of graph neural networks with roc
Jia, Z., Lin, S., Gao, M., Zaharia, M., and Aiken, A · 2020
Closest in time.
A domain-specific supercomputer for training deep neural networks
Jouppi, N. P., Yoon, D. H., Kurian, G., Li, S., Patil, N., Laudon, J., Young, C., and Patterson, D · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GCC · 2019
Cited alongside, same era.
Pipedream: Generalized pipeline parallelism for dnn training
Narayanan, D., Harlap, A., Phanishayee, A., Seshadri, V., Devanur, N. R., Ganger, G. R., Gibbons, P. B., and Zaharia, M · 2019
Cited alongside, same era.
XLA: Optimizing Compiler for TensorFlow
TensorFlow · 2019
Cited alongside, same era.
Neural predictor for neural architecture search, 2019
Wen, W., Liu, H., Li, H., Chen, Y., Bender, G., and Kindermans, P.-J · 2019
Cited alongside, same era.
Optimizing dnn computation with relaxed graph substitutions
Jia, Z., Thomas, J., Warszawski, T., Gao, M., Zaharia, M., and Aiken, A
Cited in the paper.
Beyond data and model parallelism for deep neural networks
Jia, Z., Zaharia, M., and Aiken, A
Cited in the paper.
Ithemal: Accurate, Portable and Fast Basic Block Throughput Estimation using Deep Neural Networks
Mendis, C., Renda, A., Amarasinghe, S. P., and Carbin, M
Cited in the paper.
Closest in time.
Auto-Vectorization in LLVM
LLVM · 2020
Closest in time.
Reinforced genetic algorithm learning for optimizing computation graphs
Paliwal, A., Gimeno, F., Nair, V., Li, Y., Lubin, M., Kohli, P., and Vinyals, O · 2020
Closest in time.
Transferable graph optimizers for ml compilers
Zhou, Y., Roy, S., Abdolrashidi, A., Wong, D., Ma, P. C., Xu, Q., Liu, H., Phothilimtha, P., Wang, S., Goldie, A., Mirhoseini, A., and Laudon, J · 2020
Closest in time.