Fetching the paper…
Reading the bibliography…
Special-purpose hardware accelerators are increasingly pivotal for sustaining performance improvements in emerging applications, especially as the benefits of technology scaling continue to diminish.
A Lattice-Theoretical Fixpoint Theorem and Its Applications
Alfred Tarski. 1955 · 1955
Earlier work this paper cites.
The Principal Type-Scheme of an Object in Combinatory Logic
R. Hindley. 1969 · 1969
Earlier work this paper cites.
A Unified Approach to Global Program Optimization. In Proceedings of the 1st Annual ACM SIGACT-SIGPLAN Symposium on Principles of Programming Languages (Boston, Massachusetts) (POPL’73) . Association for Computing Machinery, New York, NY, USA, 194–206
Gary A. Kildall. 1973 · 1973
Earlier work this paper cites.
Systolic Arrays for (VLSI)
H. T. Kung and Charles E. Leiserson. 1978 · 1978
Earlier work this paper cites.
Principal Type-Schemes for Functional Programs. In Proceedings of the 9th ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages (Albuquerque, New Mexico) (POPL’82) . Association for Computing Machinery, New York, NY, USA, 207–212
Luis Damas and Robin Milner. 1982 · 1982
Earlier work this paper cites.
Type assignment in programming languages
Luis Damas. 1984 · 1984
Earlier work this paper cites.
Type Inference in the Presence of Subtyping: From Theory to Practice
François Pottier. 1998 · 1998
Earlier work this paper cites.
LLVM: A Compilation Framework for Lifelong Program Analysis & Transformation. In International symposium on code generation and optimization, 2004. CGO 2004. IEEE, 75–86
Chris Lattner and Vikram Adve. 2004 · 2004
Earlier work this paper cites.
High-Level Synthesis for FPGAs: From Prototyping to Deployment
Jason Cong, Bin Liu, Stephen Neuendorffer, Juanjo Noguera, Kees Vissers, and Zhiru Zhang. 2011 · 2011
Earlier work this paper cites.
Chisel: Constructing Hardware in a Scala Embedded Language. In Proceedings of the 49th Annual Design Automation Conference . Association for Computing Machinery, New York, NY, USA, 1216–1225
Jonathan Bachrach, Huy Vo, Brian Richards, Yunsup Lee, Andrew Waterman, Rimas Avižienis, John Wawrzynek, and Krste Asanović. 2012 · 2012
Earlier work this paper cites.
Computer Generation of Hardware for Linear Digital Signal Processing Transforms
Peter Milder, Franz Franchetti, James C. Hoe, and Markus Püschel. 2012 · 2012
Earlier work this paper cites.
Polybench: The polyhedral benchmark suite
Louis-Noël Pouchet et al · 2012
Earlier work this paper cites.
Polyhedral-Based Data Reuse Optimization for Configurable Computing. In Proceedings of the ACM/SIGDA International Symposium on Field Programmable Gate Arrays (FPGA’13) . Association for Computing Machinery, New York, NY, USA, 29–38
Louis-Noël Pouchet, Peng Zhang, P. Sadayappan, and Jason Cong. 2013 · 2013
Earlier work this paper cites.
Halide: A Language and Compiler for Optimizing Parallelism, Locality, and Recomputation in Image Processing Pipelines
Jonathan Ragan-Kelley, Connelly Barnes, Andrew Adams, Sylvain Paris, Frédo Durand, and Saman Amarasinghe. 2013 · 2013
Earlier work this paper cites.
Darkroom: Compiling High-Level Image Processing Code into Hardware Pipelines
James Hegarty, John Brunhaver, Zachary DeVito, Jonathan Ragan-Kelley, Noy Cohen, Steven Bell, Artem Vasilyev, Mark Horowitz, and Pat Hanrahan. 2014 · 2014
Earlier work this paper cites.
Chlorophyll: Synthesis-Aided Compiler for Low-Power Spatial Architectures. In Proceedings of the 35th ACM SIGPLAN Conference on Programming Language Design and Implementation (Edinburgh, United Kingdom) (PLDI ’14) . Association for Computing Machinery, New York, NY, USA, 396–407
Phitchaya Mangpo Phothilimthana, Tikhon Jelvis, Rohin Shah, Nishant Totla, Sarah Chasins, and Rastislav Bodik. 2014 · 2014
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Rigel: Flexible Multi-Rate Image Processing Hardware
James Hegarty, Ross Daly, Zachary DeVito, Jonathan Ragan-Kelley, Mark Horowitz, and Pat Hanrahan. 2016 · 2016
Earlier work this paper cites.
Automatic Generation of Efficient Accelerators for Reconfigurable Hardware. In Proceedings of the 43rd International Symposium on Computer Architecture (Seoul, Republic of Korea) (ISCA’16) . IEEE Press, 115–127
David Koeplinger, Christina Delimitrou, Raghu Prabhakar, Christos Kozyrakis, Yaqi Zhang, and Kunle Olukotun. 2016 · 2016
Earlier work this paper cites.
TorchVision: PyTorch’s Computer Vision library
TorchVision maintainers and contributors. 2016 · 2016
Earlier work this paper cites.
From High-Level Deep Neural Models to FPGAs. In 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . Institute of Electrical and Electronic Engineers, Taipei, Taiwan, 1–12
Hardik Sharma, Jongse Park, Divya Mahajan, Emmanuel Amaro, Joon Kyung Kim, Chenkai Shao, Asit Mishra, and Hadi Esmaeilzadeh. 2016 · 2016
Earlier work this paper cites.
Polymorphism, Subtyping, and Type Inference in MLsub
Stephen Dolan and Alan Mycroft. 2017 · 2017
Earlier work this paper cites.
Mobilenets: Efficient Convolutional Neural Networks for Mobile Vision Applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017 · 2017
Earlier work this paper cites.
In-Datacenter Performance Analysis of a Tensor Processing Unit. In Proceedings of the 44th Annual International Symposium on Computer Architecture (Toronto, ON, Canada) (ISCA’17) . Association for Computing Machinery, New York, NY, USA, 1–12
Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, et al · 2017
Earlier work this paper cites.
The Tensor Algebra Compiler
Fredrik Kjolstad, Shoaib Kamil, Stephen Chou, David Lugato, and Saman Amarasinghe. 2017 · 2017
Earlier work this paper cites.
Generating FPGA-based image processing accelerators with Hipacc: (Invited paper). In 2017 IEEE/ACM International Conference on Computer-Aided Design (ICCAD) . 1026–1033
Oliver Reiche, M. Akif Özkan, Richard Membarth, Jürgen Teich, and Frank Hannig. 2017 · 2017
Earlier work this paper cites.
Attention is All You Need. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NIPS’17) . Curran Associates Inc., Red Hook, NY, USA, 6000–6010
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
P4FPGA: A Rapid Prototyping Framework for P4. In Proceedings of the Symposium on SDN Research (Santa Clara, CA, USA) (SOSR’17) . Association for Computing Machinery, New York, NY, USA, 122–135
Han Wang, Robert Soulé, Huynh Tu Dang, Ki Suh Lee, Vishal Shrivastav, Nate Foster, and Hakim Weatherspoon. 2017 · 2017
Earlier work this paper cites.
COMBA: A Comprehensive Model-Based Analysis Framework for High Level Synthesis of Real Applications. In 2017 IEEE/ACM International Conference on Computer-Aided Design (ICCAD) . Insititute of Electrical and Electronic Engineers, Irvine, CA, USA, 430–437
Jieru Zhao, Liang Feng, Sharad Sinha, Wei Zhang, Yun Liang, and Bingsheng He. 2017 · 2017
Earlier work this paper cites.
SODA: Stencil with Optimized Dataflow Architecture. In 2018 IEEE/ACM International Conference on Computer-Aided Design (ICCAD) . Institute of Electrical and Electronic Engineers, New York, NY, USA, 1–8
Yuze Chi, Jason Cong, Peng Wei, and Peipei Zhou. 2018 · 2018
Earlier work this paper cites.
PolySA: Polyhedral-Based Systolic Array Auto-Compilation. In 2018 IEEE/ACM International Conference on Computer-Aided Design (ICCAD) . Institute of Electrical and Electronic Engineers, New York, NY, USA, 1–8
Jason Cong and Jie Wang. 2018 · 2018
Earlier work this paper cites.
Bert: Pre-Training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Fast Inference of Deep Neural Networks in FPGAs for Particle Physics
Javier Duarte, Song Han, Philip Harris, Sergo Jindariani, Edward Kreinar, Benjamin Kreis, Jennifer Ngadiuba, Maurizio Pierini, Ryan Rivera, Nhan Tran, et al · 2018
Earlier work this paper cites.
Spatial: A Language and Compiler for Application Accelerators. In Proceedings of the 39th ACM SIGPLAN Conference on Programming Language Design and Implementation (Philadelphia, PA, USA) (PLDI 2018) . Association for Computing Machinery, New York, NY, USA, 296–311
David Koeplinger, Matthew Feldman, Raghu Prabhakar, Yaqi Zhang, Stefan Hadjis, Ruben Fiszel, Tian Zhao, Luigi Nardi, Ardavan Pedram, Christos Kozyrakis, and Kunle Olukotun. 2018 · 2018
Cited alongside, same era.
NVIDIA Tensor Core Programmability, Performance & Precision. In 2018 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW) . 522–531
Stefano Markidis, Steven Wei Der Chien, Erwin Laure, Ivy Bo Peng, and Jeffrey S. Vetter. 2018 · 2018
Cited alongside, same era.
RIPL: A Parallel Image Processing Language for FPGAs
Robert Stewart, Kirsty Duncan, Greg Michaelson, Paulo Garcia, Deepayan Bhowmik, and Andrew Wallace. 2018 · 2018
Cited alongside, same era.
Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions
Nicolas Vasilache, Oleksandr Zinenko, Theodoros Theodoridis, Priya Goyal, Zachary DeVito, William S Moses, Sven Verdoolaege, Andrew Adams, and Albert Cohen. 2018 · 2018
Alveo U280 Data Center Accelerator Card
AMD Xilinx. 2021 · 2021
Later among the works it cites.
Bridging Python to Silicon: The SODA Toolchain
Nicolas Bohm Agostini, Serena Curzel, Jeff Jun Zhang, Ankur Limaye, Cheng Tan, Vinay Amatya, Marco Minutoli, Vito Giovanni Castellana, Joseph Manzano, David Brooks, Gu-Yeon Wei, and Antonino Tumeo. 2022 · 2022
Later among the works it cites.
FPGA HLS Today: Successes, Challenges, and Opportunities
Jason Cong, Jason Lau, Gai Liu, Stephen Neuendorffer, Peichen Pan, Kees Vissers, and Zhiru Zhang. 2022 · 2022
Later among the works it cites.
LLM.int8 (): 8-bit Matrix Multiplication for Transformers at Scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer. 2022 · 2022
Later among the works it cites.
IREE (Intermediate Representation Execution Environment
IREE Developers. 2022 · 2022
Later among the works it cites.
GPTQ: Accurate Post-training Compression for Generative Pretrained Transformers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
GraphIt: A High-Performance Graph DSL
Yunming Zhang, Mengjiao Yang, Riyadh Baghdadi, Shoaib Kamil, Julian Shun, and Saman Amarasinghe. 2018 · 2018
Cited alongside, same era.
Tiramisu: A Polyhedral Compiler for Expressing Fast and Portable Code. In Proceedings of the 2019 IEEE/ACM International Symposium on Code Generation and Optimization (Washington, DC, USA) (CGO’19) . IEEE Press, 193–205
Riyadh Baghdadi, Jessica Ray, Malek Ben Romdhane, Emanuele Del Sozzo, Abdurrahman Akkas, Yunming Zhang, Patricia Suriana, Shoaib Kamil, and Saman Amarasinghe. 2019 · 2019
Cited alongside, same era.
Stateful Dataflow Multigraphs: A Data-Centric Model for Performance Portability on Heterogeneous Architectures. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (Denver, Colorado) (SC’19) . Association for Computing Machinery, New York, NY, USA, Article 81, 14 pages
Tal Ben-Nun, Johannes de Fine Licht, Alexandros N. Ziogas, Timo Schneider, and Torsten Hoefler. 2019 · 2019
Cited alongside, same era.
HeteroCL: A Multi-Paradigm Programming Infrastructure for Software-Defined Reconfigurable Computing. In Proceedings of the 2019 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (Seaside, CA, USA) (FPGA’19) . Association for Computing Machinery, New York, NY, USA, 242–251
Yi-Hsiang Lai, Yuze Chi, Yuwei Hu, Jie Wang, Cody Hao Yu, Yuan Zhou, Jason Cong, and Zhiru Zhang. 2019 · 2019
Cited alongside, same era.
PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Proceedings of the 33rd International Conference on Neural Information Processing Systems . IEEE Press, New York, NY, USA, 172–198
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, et al · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
T2S-Tensor: Productively Generating High-Performance Spatial Hardware for Dense Tensor Computations. In 2019 IEEE 27th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM) . 181–189
Nitish Srivastava, Hongbo Rong, Prithayan Barua, Guanyu Feng, Huanqi Cao, Zhiru Zhang, et al · 2019
Cited alongside, same era.
Huggingface’s Transformers: State-of-the-Art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al · 2019
Cited alongside, same era.
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022 · 2022
Later among the works it cites.
DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation. In 2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO) . 616–630
Seongmin Hong, Seungjae Moon, Junsoo Kim, Sungjae Lee, Minsub Kim, Dongsoo Lee, and Joo-Young Kim. 2022 · 2022
Later among the works it cites.
Exocompilation for Productive Programming of Hardware Accelerators. In Proceedings of the 43rd ACM SIGPLAN International Conference on Programming Language Design and Implementation (San Diego, CA, USA) (PLDI 2022) . Association for Computing Machinery, New York, NY, USA, 703–718
Yuka Ikarashi, Gilbert Louis Bernstein, Alex Reinking, Hasan Genc, and Jonathan Ragan-Kelley. 2022 · 2022
Later among the works it cites.
Accelerator Design with Decoupled Hardware Customizations: Benefits and Challenges: Invited. In Proceedings of the 59th ACM/IEEE Design Automation Conference (San Francisco, California) (DAC ’22) . Association for Computing Machinery, New York, NY, USA, 1351–1354
Debjit Pal, Yi-Hsiang Lai, Shaojie Xiang, Niansong Zhang, Hongzheng Chen, Jeremy Casas, Pasquale Cocchini, Zhenkun Yang, Jin Yang, Louis-Noël Pouchet, and Zhiru Zhang. 2022 · 2022
Later among the works it cites.
TorchDynamo Overview
PyTorch. 2022 · 2022
Later among the works it cites.
torch.fx: Practical Program Capture and Transformation for Deep Learning in Python. In Proceedings of Machine Learning and Systems , Vol. 4
James Reed, Zachary DeVito, Horace He, Ansley Ussery, and Jason Ansel. 2022 · 2022
Later among the works it cites.
Tensor Program Optimization with Probabilistic Programs. In Advances in Neural Information Processing Systems
Junru Shao, Xiyou Zhou, Siyuan Feng, Bohan Hou, Ruihang Lai, Hongyi Jin, Wuwei Lin, Masahiro Masuda, Cody Hao Yu, and Tianqi Chen. 2022 · 2022
Later among the works it cites.
Automated Accelerator Optimization Aided by Graph Neural Networks. In 2022 59th ACM/IEEE Design Automation Conference (DAC) . Association for Computing Machinery, New York, NY, USA, 55–60
Atefeh Sohrabizadeh, Yunsheng Bai, Yizhou Sun, and Jason Cong. 2022 · 2022
Later among the works it cites.
Nicolas Vasilache, Oleksandr Zinenko, Aart JC Bik, Mahesh Ravishankar, Thomas Raoux, Alexander Belyaev, et al · 2022
Later among the works it cites.
HeteroFlow: An Accelerator Programming Model with Decoupled Data Placement for Software-Defined FPGAs. In Proceedings of the 2022 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays (Virtual Event, USA) (FPGA’22) . Association for Computing Machinery, New York, NY, USA, 78–88
Shaojie Xiang, Yi-Hsiang Lai, Yuan Zhou, Hongzheng Chen, Niansong Zhang, Debjit Pal, and Zhiru Zhang. 2022 · 2022
Later among the works it cites.
ScaleHLS: A New Scalable High-Level Synthesis Framework on Multi-Level Intermediate Representation. In 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
Hanchen Ye, Cong Hao, Jianyi Cheng, Hyunmin Jeong, Jack Huang, Stephen Neuendorffer, and Deming Chen. 2022 · 2022
Later among the works it cites.
POLSCA: Polyhedral High-Level Synthesis with Compiler Transformations. In 2022 32nd International Conference on Field-Programmable Logic and Applications (FPL) . Institute of Electrical and Electronic Engineers, Belfast, United Kingdom, 235–242
Ruizhe Zhao, Jianyi Cheng, Wayne Luk, and George A. Constantinides. 2022 · 2022
Later among the works it cites.
[RFC] Interfaces and Dialects for Precise IR Transformation Control
Alex Zinenko. 2022 · 2022
Later among the works it cites.
Inferentia Architecture
AWS. 2023 · 2023
Later among the works it cites.
TensorIR: An Abstraction for Automatic Tensorized Program Optimization. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (Vancouver, BC, Canada) (ASPLOS 2023) . Association for Computing Machinery, New York, NY, USA, 804–817
Siyuan Feng, Bohan Hou, Hongyi Jin, Wuwei Lin, Junru Shao, Ruihang Lai, Zihao Ye, Lianmin Zheng, Cody Hao Yu, Yong Yu, and Tianqi Chen. 2023 · 2023
Later among the works it cites.
TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings. In Proceedings of the 50th Annual International Symposium on Computer Architecture (Orlando, FL, USA) (ISCA’23) . Association for Computing Machinery, New York, NY, USA, Article 82, 14 pages
Norm Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, et al · 2023
Later among the works it cites.
Modular Hardware Design with Timeline Types
Rachit Nigam, Pedro Henrique Azevedo de Amorim, and Adrian Sampson. 2023 · 2023
Later among the works it cites.
OpenAI. 2023 · 2023
Later among the works it cites.
Efficiently Scaling Transformer Inference. In Proceedings of Machine Learning and Systems , Vol. 5
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. 2023 · 2023
Later among the works it cites.
PyBind11
PyBind. 2023 · 2023
Later among the works it cites.
Llama: Open and Efficient Foundation Language Models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. In International Conference on Machine Learning . PMLR, 38087–38099
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023 · 2023
Later among the works it cites.
Merlin Compiler
AMD Xilinx. 2023 · 2023
Later among the works it cites.
Atom: Low-Bit Quantization for Efficient and Accurate LLM Serving
Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. 2023 · 2023
Later among the works it cites.
PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS’24) . Association for Computing Machinery, New York, NY, USA, 317–335
Jason Ansel, Edward Yang, Horace He, Natalia Gimelshein, Animesh Jain, Michael Voznesensky, et al · 2024
Closest in time.
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference
Hongzheng Chen, Jiahao Zhang, Yixiao Du, Shaojie Xiang, Zichao Yue, Niansong Zhang, Yaohui Cai, and Zhiru Zhang. 2024b · 2024
Closest in time.
CIRCT: Circuit IR Compilers and Tools
CIRCT. 2024 · 2024
Closest in time.
Formal Verification of Source-to-Source Transformations for HLS. In Proceedings of the 2024 ACM/SIGDA International Symposium on Field Programmable Gate Arrays (Monterey, CA, USA) (FPGA’24) . Association for Computing Machinery, New York, NY, USA, 97–107
Louis-Noël Pouchet, Emily Tucker, Niansong Zhang, Hongzheng Chen, Debjit Pal, Gabriel Rodríguez, and Zhiru Zhang. 2024 · 2024
Closest in time.