Fetching the paper…
Reading the bibliography…
Sparse tensors are rapidly becoming critical components of modern deep learning workloads.
Generating Long Sequences with Sparse Transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, and Jamie Brew. 2019 · 1910
Earlier work this paper cites.
Direct Methods for Sparse Matrices
Iain S Duff, Albert M Erisman, and John K Reid. 1986 · 1986
Earlier work this paper cites.
The Use of Vector and Parallel Computers in the Solution of Large Sparse Linear Equations
Iain S. Duff. 1987 · 1987
Earlier work this paper cites.
The performance of ITPACK on vector computers for solving large sparse linear systems arising in sample oil reseervoir simulation problems
Thomas C. Oppe and David R. Kincaid. 1987 · 1987
Earlier work this paper cites.
Krylov Subspace Methods on Supercomputers
Youcef Saad. 1989 · 1989
Earlier work this paper cites.
//ELLPACK: A Numerical Simulation Programming Environment for Parallel MIMD Machines
E. N. Houstis, J. R. Rice, N. P. Chrisochoides, H. C. Karathanasis, P. N. Papachiou, M. K. Samartzis, E. A. Vavalis, Ko Yang Wang, and S. Weerawarana. 1990 · 1990
Earlier work this paper cites.
SPARSKIT: A basic tool kit for sparse matrix computations
Youcef Saad. 1990 · 1990
Earlier work this paper cites.
Compilation Techniques for Sparse Matrix Computations. In Proceedings of the 7th international conference on Supercomputing, ICS 1993, Tokyo, Japan, July 20-22, 1993 , Yoichi Muraoka (Ed.). ACM, 416–424
Aart J. C. Bik and Harry A. G. Wijshoff. 1993 · 1993
Earlier work this paper cites.
Nonzero Structure Analysis. In Proceedings of the 8th International Conference on Supercomputing (Manchester, England) (ICS ’94) . Association for Computing Machinery, New York, NY, USA, 226–235
Aart J. C. Bik and Harry A. G. Wijshoff. 1994 · 1994
Earlier work this paper cites.
Advanced Compiler Optimizations for Sparse Computations
A.J.C. Bik and H.A.G. Wijshoff. 1995 · 1995
Earlier work this paper cites.
Random butterfly transformations with applications in computational linear algebra
Douglass Stott Parker. 1995 · 1995
Earlier work this paper cites.
Compiler Support for Sparse Matrix Computations
Aart J. C. Bik. 1996 · 1996
Earlier work this paper cites.
Automatic Data Structure Selection and Transformation for Sparse Matrix Computations
Aart J. C. Bik and Harry A. G. Wijshoff. 1996 · 1996
Earlier work this paper cites.
SIPR: A New Framework for Generating Efficient Code for Sparse Matrix Computations. In Languages and Compilers for Parallel Computing, 11th International Workshop, LCPC’98, Chapel Hill, NC, USA, August 7-9, 1998, Proceedings (Lecture Notes in Computer Science, Vol. 1656) , Siddhartha Chatterjee, Jan F. Prins, Larry Carter, Jeanne Ferrante, Zhiyuan Li, David C. Sehr, and Pen-Chung Yew (Eds.). Springer, 213–229
William W. Pugh and Tatiana Shpeisman. 1998 · 1998
Earlier work this paper cites.
A Framework for Sparse Matrix Code Synthesis from High-level Specifications. In Proceedings Supercomputing 2000, November 4-10, 2000, Dallas, Texas, USA. IEEE Computer Society, CD-ROM , Jed Donnelley (Ed.). IEEE Computer Society, 58
Nawaaz Ahmed, Nikolay Mateev, Keshav Pingali, and Paul Stodghill. 2000 · 2000
Earlier work this paper cites.
Next-generation generic programming and its application to sparse matrix computations. In Proceedings of the 14th international conference on Supercomputing, ICS 2000, Santa Fe, NM, USA, May 8-11, 2000 , John Reynders and Alexander V. Veidenbaum (Eds.). ACM, 88–99
Nikolay Mateev, Keshav Pingali, Paul Stodghill, and Vladimir Kotlyar. 2000 · 2000
Earlier work this paper cites.
Longformer: The Long-Document Transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
TIRAMISU: A Polyhedral Compiler for Dense and Sparse Deep Learning
Riyadh Baghdadi, Abdelkader Nadir Debbagh, Kamel Abdous, Fatima-Zohra Benhamida, Alex Renda, Jonathan Elliott Frankle, Michael Carbin, and Saman P. Amarasinghe. 2020 · 2005
Earlier work this paper cites.
OSKI: A library of automatically tuned sparse matrix kernels
Richard Vuduc, James W Demmel, and Katherine A Yelick. 2005 · 2005
Earlier work this paper cites.
On the representation and multiplication of hypersparse matrices. In 22nd IEEE International Symposium on Parallel and Distributed Processing, IPDPS 2008, Miami, Florida USA, April 14-18, 2008 . IEEE, 1–11
Aydin Buluç and John R. Gilbert. 2008 · 2008
Earlier work this paper cites.
Collective Classification in Network Data
Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Gallagher, and Tina Eliassi-Rad. 2008 · 2008
Earlier work this paper cites.
Halide: A Language and Compiler for Optimizing Parallelism, Locality, and Recomputation in Image Processing Pipelines. In Proceedings of the 34th ACM SIGPLAN Conference on Programming Language Design and Implementation (Seattle, Washington, USA) (PLDI ’13) . Association for Computing Machinery, New York, NY, USA, 519–530
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe. 2013 · 2013
Earlier work this paper cites.
Intel math kernel library
Endong Wang, Qing Zhang, Bo Shen, Guangyong Zhang, Xiaowei Lu, Qing Wu, and Yajuan Wang. 2014 · 2014
Earlier work this paper cites.
Tensor-Matrix Products with a Compressed Sparse Tensor. In Proceedings of the 5th Workshop on Irregular Applications: Architectures and Algorithms (Austin, Texas) (IA<sup>3</sup> ’15) . Association for Computing Machinery, New York, NY, USA, Article 5, 7 pages
Shaden Smith and George Karypis. 2015 · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J. Dally. 2016 · 2016
Earlier work this paper cites.
A Collection of Benchmark Datasets for Systematic Evaluations of Machine Learning on the Semantic Web. In The Semantic Web – ISWC 2016 , Paul Groth, Elena Simperl, Alasdair Gray, Marta Sabou, Markus Krötzsch, Freddy Lecue, Fabian Flöck, and Yolanda Gil (Eds.). Springer International Publishing, Cham, 186–194
Petar Ristoski, Gerben Klaas Dirk de Vries, and Heiko Paulheim. 2016 · 2016
Earlier work this paper cites.
Sympiler: transforming sparse matrix codes by decoupling symbolic analysis. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC 2017, Denver, CO, USA, November 12 - 17, 2017 , Bernd Mohr and Padma Raghavan (Eds.). ACM, 13
Kazem Cheshmi, Shoaib Kamil, Michelle Mills Strout, and Maryam Mehri Dehnavi. 2017 · 2017
Earlier work this paper cites.
Neural Message Passing for Quantum Chemistry. In Proceedings of the 34th International Conference on Machine Learning - Volume 70 (Sydney, NSW, Australia) (ICML’17) . JMLR.org, 1263–1272
Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. 2017 · 2017
Earlier work this paper cites.
Inductive Representation Learning on Large Graphs. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17) . Curran Associates Inc., Red Hook, NY, USA, 1025–1035
William L. Hamilton, Rex Ying, and Jure Leskovec. 2017 · 2017
Earlier work this paper cites.
In-Datacenter Performance Analysis of a Tensor Processing Unit. In Proceedings of the 44th Annual International Symposium on Computer Architecture, ISCA 2017, Toronto, ON, Canada, June 24-28, 2017 . ACM, 1–12
Norman P. Jouppi, Cliff Young, Nishant Patil, David A. Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazir Ghaemmaghami, Rajendra Gottipati, William Gulland, Robert Hagmann, C. Richard Ho, Doug Hogberg, John Hu, Robert Hundt, Dan Hurt, Julian Ibarz, Aaron Jaffey, Alek Jaworski, Alexander Kaplan, Harshit Khaitan, Daniel Killebrew, Andy Koch, Naveen Kumar, Steve Lacy, James Laudon, James Law, Diemthu Le, Chris Leary, Zhuyuan Liu, Kyle Lucke, Alan Lundin, Gordon MacKean, Adriana Maggiore, Maire Mahony, Kieran Miller, Rahul Nagarajan, Ravi Narayanaswami, Ray Ni, Kathy Nix, Thomas Norrie, Mark Omernick, Narayana Penukonda, Andy Phelps, Jonathan Ross, Matt Ross, Amir Salek, Emad Samadiani, Chris Severn, Gregory Sizikov, Matthew Snelham, Jed Souter, Dan Steinberg, Andy Swing, Mercedes Tan, Gregory Thorson, Bo Tian, Horia Toma, Erick Tuttle, Vijay Vasudevan, Richard Walter, Walter Wang, Eric Wilcox, and Doe Hyun Yoon. 2017 · 2017
Earlier work this paper cites.
Semi-Supervised Classification with Graph Convolutional Networks. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings . OpenReview.net
Thomas N. Kipf and Max Welling. 2017 · 2017
Earlier work this paper cites.
The Tensor Algebra Compiler
Fredrik Kjolstad, Shoaib Kamil, Stephen Chou, David Lugato, and Saman Amarasinghe. 2017 · 2017
Earlier work this paper cites.
Reordering Strategy for Blocking Optimization in Sparse Linear Solvers
Gregoire Pichon, Mathieu Faverge, Pierre Ramet, and Jean Roman. 2017 · 2017
Earlier work this paper cites.
Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Conference Track Proceedings . OpenReview.net
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean. 2017 · 2017
Earlier work this paper cites.
Parallel Associative Reductions in Halide. In Proceedings of the 2017 International Symposium on Code Generation and Optimization (Austin, USA) (CGO ’17) . IEEE Press, 281–291
Patricia Suriana, Andrew Adams, and Shoaib Kamil. 2017 · 2017
Cited alongside, same era.
Attention is All You Need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17) . Curran Associates Inc., Red Hook, NY, USA, 6000–6010
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
ParSy: inspection and transformation of sparse matrix computations for parallelism. In Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis, SC 2018, Dallas, TX, USA, November 11-16, 2018 . IEEE / ACM, 62:1–62:15
Kazem Cheshmi, Shoaib Kamil, Michelle Mills Strout, and Maryam Mehri Dehnavi. 2018 · 2018
Cited alongside, same era.
Format Abstraction for Sparse Tensor Algebra Compilers
Stephen Chou, Fredrik Kjolstad, and Saman Amarasinghe. 2018 · 2018
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python
Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J. van der Walt, Matthew Brett, Joshua Wilson, K. Jarrod Millman, Nikolay Mayorov, Andrew R. J. Nelson, Eric Jones, Robert Kern, Eric Larson, C J Carey, İlhan Polat, Yu Feng, Eric W. Moore, Jake VanderPlas, Denis Laxalde, Josef Perktold, Robert Cimrman, Ian Henriksen, E. A. Quintero, Charles R. Harris, Anne M. Archibald, Antônio H. Ribeiro, Fabian Pedregosa, Paul van Mulbregt, and SciPy 1.0 Contributors. 2020 · 2020
Later among the works it cites.
FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous System. In ASPLOS ’20: Architectural Support for Programming Languages and Operating Systems, Lausanne, Switzerland, March 16-20, 2020 , James R. Larus, Luis Ceze, and Karin Strauss (Eds.). ACM, 859–873
Size Zheng, Yun Liang, Shuo Wang, Renze Chen, and Kaiwen Sheng. 2020b · 2020
Later among the works it cites.
dgSPARSE Library
dgSPARSE team. 2021 · 2021
Later among the works it cites.
Compilation of Sparse Array Programming Models
Rawn Henry, Olivia Hsu, Rohan Yadav, Stephen Chou, Kunle Olukotun, Saman Amarasinghe, and Fredrik Kjolstad. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Sampled Dense Matrix Multiplication for High-Performance Machine Learning. In 2018 IEEE 25th International Conference on High Performance Computing (HiPC) . 32–41
Israt Nisa, Aravind Sukumaran-Rajam, Sureyya Emre Kurt, Changwan Hong, and P. Sadayappan. 2018 · 2018
Cited alongside, same era.
Relay: A New IR for Machine Learning Frameworks. In Proceedings of the 2nd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages (Philadelphia, PA, USA) (MAPL 2018) . Association for Computing Machinery, New York, NY, USA, 58–68
Jared Roesch, Steven Lyubomirsky, Logan Weber, Josh Pollock, Marisa Kirisame, Tianqi Chen, and Zachary Tatlock. 2018 · 2018
Cited alongside, same era.
Modeling Relational Data with Graph Convolutional Networks. In The Semantic Web - 15th International Conference, ESWC 2018, Heraklion, Crete, Greece, June 3-7, 2018, Proceedings (Lecture Notes in Computer Science, Vol. 10843) , Aldo Gangemi, Roberto Navigli, Maria-Esther Vidal, Pascal Hitzler, Raphaël Troncy, Laura Hollink, Anna Tordai, and Mehwish Alam (Eds.). Springer, 593–607
Michael Sejr Schlichtkrull, Thomas N. Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov, and Max Welling. 2018 · 2018
Cited alongside, same era.
The Sparse Polyhedral Framework: Composing Compiler-Generated Inspector-Executor Code
Michelle Mills Strout, Mary W. Hall, and Catherine Olschanowsky. 2018 · 2018
Cited alongside, same era.
Graph Attention Networks. In International Conference on Learning Representations
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. 2018 · 2018
Cited alongside, same era.
Design Principles for Sparse Matrix Multiplication on the GPU. In Euro-Par 2018: Parallel Processing - 24th International Conference on Parallel and Distributed Computing, Turin, Italy, August 27-31, 2018, Proceedings (Lecture Notes in Computer Science, Vol. 11014) , Marco Aldinucci, Luca Padovani, and Massimo Torquati (Eds.). Springer, 672–687
Carl Yang, Aydin Buluç, and John D. Owens. 2018 · 2018
Cited alongside, same era.
Learning to optimize halide with tree search and random programs
Andrew Adams, Karima Ma, Luke Anderson, Riyadh Baghdadi, Tzu-Mao Li, Michaël Gharbi, Benoit Steiner, Steven Johnson, Kayvon Fatahalian, Frédo Durand, and Jonathan Ragan-Kelley. 2019 · 2019
Cited alongside, same era.
SemanticKITTI: A Dataset for Semantic Scene Understanding of LiDAR Sequences. In Proc. of the IEEE/CVF International Conf. on Computer Vision (ICCV)
J. Behley, M. Garbade, A. Milioto, J. Quenzel, S. Behnke, C. Stachniss, and J. Gall. 2019 · 2019
Cited alongside, same era.
Sparsity in Deep Learning: Pruning and growth for efficient inference and training in neural networks
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste. 2021 · 2021
Later among the works it cites.
Block Pruning For Faster Transformers. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, 10619–10629
François Lagunas, Ella Charlaix, Victor Sanh, and Alexander Rush. 2021 · 2021
Later among the works it cites.
Accelerating SpMM Kernel with Cache-First Edge Sampling for Graph Neural Networks
Chien-Yu Lin, Liang Luo, and Luis Ceze. 2021 · 2021
Later among the works it cites.
Learning Sparse Matrix Row Permutations for Efficient SpMM on GPU Architectures. In IEEE International Symposium on Performance Analysis of Systems and Software, ISPASS 2021, Stony Brook, NY, USA, March 28-30, 2021 . IEEE, 48–58
Atefeh Mehrabi, Donghyuk Lee, Niladrish Chatterjee, Daniel J. Sorin, Benjamin C. Lee, and Mike O’Connor. 2021 · 2021
Later among the works it cites.
FusedMM: A Unified SDDMM-SpMM Kernel for Graph Embedding and Graph Neural Networks. In 2021 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . 256–266
Md. Khaledur Rahman, Majedul Haque Sujon, and Ariful Azad. 2021 · 2021
Later among the works it cites.
A High Performance Sparse Tensor Algebra Compiler in MLIR. In 2021 IEEE/ACM 7th Workshop on the LLVM Compiler Infrastructure in HPC (LLVM-HPC) . 27–38
Ruiqin Tian, Luanzheng Guo, Jiajia Li, Bin Ren, and Gokcen Kestor. 2021 · 2021
Later among the works it cites.
TC-GNN: Accelerating Sparse Graph Neural Network Computation Via Dense Tensor Core on GPUs
Yuke Wang, Boyuan Feng, and Yufei Ding. 2021a · 2021
Later among the works it cites.
Seastar: vertex-centric programming for graph neural networks. In EuroSys ’21: Sixteenth European Conference on Computer Systems, Online Event, United Kingdom, April 26-28, 2021 , Antonio Barbalace, Pramod Bhatotia, Lorenzo Alvisi, and Cristian Cadar (Eds.). ACM, 359–375
Yidi Wu, Kaihao Ma, Zhenkun Cai, Tatiana Jin, Boyang Li, Chengguang Zheng, James Cheng, and Fan Yu. 2021 · 2021
Later among the works it cites.
Exploiting Online Locality and Reduction Parallelism for Sampled Dense Matrix Multiplication on GPUs. In 39th IEEE International Conference on Computer Design, ICCD 2021, Storrs, CT, USA, October 24-27, 2021 . IEEE, 567–574
Zhongming Yu, Guohao Dai, Guyue Huang, Yu Wang, and Huazhong Yang. 2021 · 2021
Later among the works it cites.
AKG: Automatic Kernel Generation for Neural Processing Units Using Polyhedral Transformations. In Proceedings of the 42nd ACM SIGPLAN International Conference on Programming Language Design and Implementation (Virtual, Canada) (PLDI 2021) . Association for Computing Machinery, New York, NY, USA, 1233–1248
Jie Zhao, Bojie Li, Wang Nie, Zhen Geng, Renwei Zhang, Xiong Gao, Bin Cheng, Chen Wu, Yun Cheng, Zheng Li, Peng Di, Kun Zhang, and Xuefeng Jin. 2021 · 2021
Later among the works it cites.
Learning N: M Fine-grained Structured Sparse Neural Networks From Scratch. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021 . OpenReview.net
Aojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu, Zhijie Zhang, Kun Yuan, Wenxiu Sun, and Hongsheng Li. 2021 · 2021
Later among the works it cites.
Autoscheduling for Sparse Tensor Algebra with an Asymptotic Cost Model. In Proceedings of the 43rd ACM SIGPLAN International Conference on Programming Language Design and Implementation (San Diego, CA, USA) (PLDI 2022) . Association for Computing Machinery, New York, NY, USA, 269–285
Peter Ahrens, Fredrik Kjolstad, and Saman Amarasinghe. 2022 · 2022
Closest in time.
Compiler Support for Sparse Tensor Computations in MLIR
Aart Bik, Penporn Koanantakool, Tatiana Shpeisman, Nicolas Vasilache, Bixia Zheng, and Fredrik Kjolstad. 2022 · 2022
Closest in time.
Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net
Beidi Chen, Tri Dao, Kaizhao Liang, Jiaming Yang, Zhao Song, Atri Rudra, and Christopher Ré. 2022 · 2022
Closest in time.
cuSPARSE :: CUDA Toolkit Documentation v11.7.1
NVIDIA Corporation. 2022 · 2022
Closest in time.
Heuristic Adaptability to Input Dynamics for SpMM on GPUs. In Proceedings of the 59th ACM/IEEE Design Automation Conference (San Francisco, California) (DAC ’22) . Association for Computing Machinery, New York, NY, USA, 595–600
Guohao Dai, Guyue Huang, Shang Yang, Zhongming Yu, Hengrui Zhang, Yufei Ding, Yuan Xie, Huazhong Yang, and Yu Wang. 2022 · 2022
Closest in time.
SparseLNR: Accelerating Sparse Tensor Computations Using Loop Nest Restructuring. In Proceedings of the 36th ACM International Conference on Supercomputing (Virtual Event) (ICS ’22) . Association for Computing Machinery, New York, NY, USA, Article 15, 14 pages
Adhitha Dias, Kirshanthan Sundararajah, Charitha Saumya, and Milind Kulkarni. 2022 · 2022
Closest in time.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
William Fedus, Barret Zoph, and Noam Shazeer. 2022 · 2022
Closest in time.
The CoRa Tensor Compiler: Compilation for Ragged Tensors with Minimal Padding. In Proceedings of Machine Learning and Systems , A. Smola, A. Dimakis, and I. Stoica (Eds.)
Pratik Fegade, Tianqi Chen, Phillip B. Gibbons, and Todd C. Mowry. 2022 · 2022
Closest in time.
TensorIR: An Abstraction for Automatic Tensorized Program Optimization
Siyuan Feng, Bohan Hou, Hongyi Jin, Wuwei Lin, Junru Shao, Ruihang Lai, Zihao Ye, Lianmin Zheng, Cody Hao Yu, Yong Yu, and Tianqi Chen. 2022 · 2022
Closest in time.
Automatic Horizontal Fusion for GPU Kernels. In Proceedings of the 20th IEEE/ACM International Symposium on Code Generation and Optimization (Virtual Event, Republic of Korea) (CGO ’22) . IEEE Press, 14–27
Ao Li, Bojian Zheng, Gennady Pekhimenko, and Fan Long. 2022b · 2022
Closest in time.
Efficient Quantized Sparse Matrix Operations on Tensor Cores
Shigang Li, Kazuki Osawa, and Torsten Hoefler. 2022a · 2022
Closest in time.
Tensor Program Optimization with Probabilistic Programs
Junru Shao, Xiyou Zhou, Siyuan Feng, Bohan Hou, Ruihang Lai, Hongyi Jin, Wuwei Lin, Masahiro Masuda, Cody Hao Yu, and Tianqi Chen. 2022 · 2022
Closest in time.
TorchSparse: Efficient Point Cloud Inference Engine. In Proceedings of Machine Learning and Systems , D. Marculescu, Y. Chi, and C. Wu (Eds.), Vol. 4. 302–315
Haotian Tang, Zhijian Liu, Xiuyu Li, Yujun Lin, and Song Han. 2022a · 2022
Closest in time.
FreeTensor: A Free-Form DSL with Holistic Optimizations for Irregular Tensor Programs. In Proceedings of the 43rd ACM SIGPLAN International Conference on Programming Language Design and Implementation (San Diego, CA, USA) (PLDI 2022) . Association for Computing Machinery, New York, NY, USA, 872–887
Shizhi Tang, Jidong Zhai, Haojie Wang, Lin Jiang, Liyan Zheng, Zhenhao Yuan, and Chen Zhang. 2022b · 2022
Closest in time.
Nicolas Vasilache, Oleksandr Zinenko, Aart J. C. Bik, Mahesh Ravishankar, Thomas Raoux, Alexander Belyaev, Matthias Springer, Tobias Gysi, Diego Caballero, Stephan Herhut, Stella Laurenzo, and Albert Cohen. 2022 · 2022
Closest in time.
QGTC: Accelerating Quantized Graph Neural Networks via GPU Tensor Core. In Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (Seoul, Republic of Korea) (PPoPP ’22) . Association for Computing Machinery, New York, NY, USA, 107–119
Yuke Wang, Boyuan Feng, and Yufei Ding. 2022 · 2022
Closest in time.
Graphiler: Optimizing Graph Neural Networks with Message Passing Data Flow Graph. In Proceedings of Machine Learning and Systems , D. Marculescu, Y. Chi, and C. Wu (Eds.), Vol. 4. 515–528
Zhiqiang Xie, Minjie Wang, Zihao Ye, Zheng Zhang, and Rui Fan. 2022 · 2022
Closest in time.
DietCode: Automatic Optimization for Dynamic Tensor Programs. In Proceedings of Machine Learning and Systems , D. Marculescu, Y. Chi, and C. Wu (Eds.), Vol. 4. 848–863
Bojian Zheng, Ziheng Jiang, Cody Hao Yu, Haichen Shen, Joshua Fromm, Yizhi Liu, Yida Wang, Luis Ceze, Tianqi Chen, and Gennady Pekhimenko. 2022a · 2022
Closest in time.
Heterogeneous Graph Attention Network. In The World Wide Web Conference, WWW 2019, San Francisco, CA, USA, May 13-17, 2019 , Ling Liu, Ryen W. White, Amin Mantrach, Fabrizio Silvestri, Julian J. McAuley, Ricardo Baeza-Yates, and Leila Zia (Eds.). ACM, 2022–2032
Xiao Wang, Houye Ji, Chuan Shi, Bai Wang, Yanfang Ye, Peng Cui, and Philip S. Yu. 2019a · 2032
Closest in time.