Fetching the paper…
Reading the bibliography…
Spatial computing architectures promise a major stride in performance and energy efficiency over the traditional load/store devices currently employed in large scale computing systems.
A. P. Yershov, “ALPHA – an automatic programming system of high efficiency,” J. ACM , 1966
1966
Earlier work this paper cites.
F. E. Allen and J. Cocke, A catalogue of optimizing transformations , 1971
1971
Earlier work this paper cites.
D. J. Kuck, “A survey of parallel machine organization and programming,” CSUR , Mar. 1977
1977
Earlier work this paper cites.
J. Cocke and K. Kennedy, “An algorithm for reduction of operator strength,” CACM , 1977
1977
Earlier work this paper cites.
G. L. Steele, “Arithmetic shifting considered harmful,” ACM SIGPLAN Notices , 1977
1977
Earlier work this paper cites.
H. Kung and C. E. Leiserson, “Systolic arrays (for VLSI),” in Sparse Matrix Proceedings , 1978
1978
Earlier work this paper cites.
J. J. Dongarra and A. R. Hinds, “Unrolling loops in Fortran,” Software: Practice and Experience , 1979
1979
Earlier work this paper cites.
D. J. Kuck et al. , “Dependence graphs and compiler optimizations,” in POPL , 1981
1981
Earlier work this paper cites.
D. D. Gajski et al. , “A second opinion on data flow machines and languages,” Computer , 1982
1982
Earlier work this paper cites.
M. J. Wolfe, “Optimizing supercompilers for supercomputers,” Ph.D. dissertation, 1982
1982
Earlier work this paper cites.
J. R. Allen and K. Kennedy, “Automatic loop interchange,” in SIGPLAN , 1984
1984
Earlier work this paper cites.
G. D. Smith, Numerical solution of partial differential equations: finite difference methods , 1985
1985
Earlier work this paper cites.
R. Bernstein, “Multiplication by integer constants,” Softw. Pract. Exper. , 1986
1986
Earlier work this paper cites.
A. V. Aho et al. , “Compilers, principles, techniques,” Addison Wesley , 1986
1986
Earlier work this paper cites.
C. D. Polychronopoulos, “Loop coalescing: A compiler transformation for parallel machines,” Tech. Rep., 1987
1987
Earlier work this paper cites.
C. A. Fletcher, Computational Techniques for Fluid Dynamics 2 , 1988
1988
Earlier work this paper cites.
C. D. Polychronopoulos, “Advanced loop optimizations for parallel computers,” in ICS , 1988
1988
Earlier work this paper cites.
M. Lam, “Software pipelining: An effective scheduling technique for VLIW machines,” in PLDI , 1988
1988
Earlier work this paper cites.
M. Weiss, “Strip mining on SIMD architectures,” in ICS , 1991
1991
Earlier work this paper cites.
M. D. Lam et al. , “The cache performance and optimizations of blocked algorithms,” 1991
1991
Earlier work this paper cites.
D. F. Bacon et al. , “Compiler transformations for high-performance computing,” CSUR , 1994
1994
Earlier work this paper cites.
W. A. Wulf and S. A. McKee, “Hitting the memory wall: implications of the obvious,” SIGARCH , 1995
1995
Earlier work this paper cites.
A. Taflove and S. C. Hagness, “Computational electrodynamics: The finite-difference time-domain method,” 1995
1995
Earlier work this paper cites.
M. B. Gokhale et al. , “Stream-oriented FPGA computing in the Streams-C high level language,” in FCCM , 2000
2000
Earlier work this paper cites.
J. Hammarberg and S. Nadjm-Tehrani, “Development of safety-critical reconfigurable hardware with Esterel,” FMICS , 2003
2003
Earlier work this paper cites.
S. Gupta et al. , “SPARK: a high-level synthesis framework for applying parallelizing compiler transformations,” in VLSID , 2003
2003
Earlier work this paper cites.
R. Nikhil, “Bluespec system Verilog: efficient, correct RTL from high level specifications,” in MEMOCODE , 2004
2004
Earlier work this paper cites.
——, “Coordinated parallelizing compiler optimizations and high-level synthesis,” TODAES , 2004
2004
Earlier work this paper cites.
Y. Y. Leow et al. , “Generating hardware from OpenMP programs,” in FPT , 2006
2006
Earlier work this paper cites.
S. Sirowy and A. Forin, “Where’s the beef? why FPGAs are so fast,” MS Research , 2008
2008
Earlier work this paper cites.
Z. Zhang et al. , “AutoPilot: A platform-based ESL synthesis system,” in High-Level Synthesis , 2008
2008
Earlier work this paper cites.
S. Ryoo et al. , “Optimization principles and application performance evaluation of a multithreaded GPU using CUDA,” in PPoPP , 2008
2008
Earlier work this paper cites.
D. B. Thomas et al. , “A comparison of CPUs, GPUs, FPGAs, and massively parallel processor arrays for random number generation,” in FPGA , 2009
2009
Earlier work this paper cites.
G. Martin and G. Smith, “High-level synthesis: Past, present, and future,” D&T , 2009
2009
Earlier work this paper cites.
A. Papakonstantinou et al. , “FCUDA: Enabling efficient compilation of CUDA kernels onto FPGAs,” in SASP , 2009
2009
Earlier work this paper cites.
A. R. Brodtkorb et al. , “State-of-the-art in heterogeneous computing,” Scientific Programming , 2010
2010
Cited alongside, same era.
J. Auerbach et al. , “Lime: A Java-compatible and synthesizable language for heterogeneous architectures,” in OOPSLA , 2010
2010
Cited alongside, same era.
J. Cong et al. , “High-level synthesis for FPGAs: From prototyping to deployment,” TCAD , 2011
2011
Cited alongside, same era.
A. Canis et al. , “LegUp: High-level synthesis for FPGA-based processor/accelerator systems,” in FPGA , 2011
2011
Cited alongside, same era.
M. Owaida et al. , “Synthesis of platform architectures from OpenCL programs,” in FCCM , 2011
2011
Cited alongside, same era.
S. Lee et al. , “OpenACC to FPGA: A framework for directive-based high-performance reconfigurable computing,” in IPDPS , 2016
2016
Later among the works it cites.
J. Cong et al. , “Source-to-source optimization for HLS,” in FPGAs for Software Programmers , 2016
2016
Later among the works it cites.
D. Koeplinger et al. , “Automatic generation of efficient accelerators for reconfigurable hardware,” in ISCA , 2016
2016
Later among the works it cites.
Y. Umuroglu et al. , “FINN: A framework for fast, scalable binarized neural network inference,” in FPGA , 2017
2017
Later among the works it cites.
H. M. Waidyasooriya et al. , “OpenCL-based FPGA-platform for stencil computation and its optimization methodology,” TPDS , May 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011
Cited alongside, same era.
W. Meeus et al. , “An overview of today’s high-level synthesis tools,” DAEM , 2012
2012
Cited alongside, same era.
R. Nane et al. , “DWARV 2.0: A CoSy-based C-to-VHDL hardware compiler,” in FPL , 2012
2012
Cited alongside, same era.
T. Czajkowski et al. , “From OpenCL to high-performance hardware on FPGAs,” in FPL , 2012
2012
Cited alongside, same era.
——, “A compiler and runtime for heterogeneous computing,” in DAC , 2012
2012
Cited alongside, same era.
X. Niu et al. , “Exploiting run-time reconfiguration in stencil computation,” in FPL , 2012
2012
Cited alongside, same era.
J. Fowers et al. , “A performance and energy comparison of FPGAs, GPUs, and multicores for sliding-window applications,” in FPGA , 2012
2012
Cited alongside, same era.
D. Weller et al. , “Energy efficient scientific computing on FPGAs using OpenCL,” in FPGA , 2017
2017
Later among the works it cites.
J. Zhang and J. Li, “Improving the performance of OpenCL-based FPGA accelerator for convolutional neural network,” in FPGA , 2017
2017
Later among the works it cites.
E. H. D’Hollander, “High-level synthesis optimization for blocked floating-point matrix multiplication,” SIGARCH , 2017
2017
Later among the works it cites.
T. Lloyd et al. , “A case for better integration of host and target compilation when using OpenCL for FPGAs,” in FSP , 2017
2017
Later among the works it cites.
J. Pu et al. , “Programming heterogeneous systems from an image processing DSL,” TACO , 2017
2017
Later among the works it cites.
E. D. Sozzo et al. , “A common backend for hardware acceleration on FPGA,” in ICCD , 2017
2017
Later among the works it cites.
A. Izraelevitz et al. , “Reusability is FIRRTL ground: Hardware construction languages, compiler frameworks, and transformations,” in ICCAD , 2017
2017
Later among the works it cites.
M. Blott et al. , “FINN-R: An end-to-end deep-learning framework for fast exploration of quantized neural networks,” TRETS , 2018
2018
Closest in time.
H. R. Zohouri et al. , “Combined spatial and temporal blocking for high-performance stencil computation on FPGAs using OpenCL,” in FPGA , 2018
2018
Closest in time.
T. Kenter et al. , “OpenCL-based FPGA design to accelerate the nodal discontinuous Galerkin method for unstructured meshes,” in FCCM , 2018
2018
Closest in time.
L. Josipović et al. , “Dynamically scheduled high-level synthesis,” in FPGA , 2018
2018
Closest in time.
R. Kastner et al. , “Parallel programming for FPGAs,” arXiv:1805.03648 , 2018
2018
Closest in time.
T. Ben-Nun and T. Hoefler, “Demystifying parallel and distributed deep learning: An in-depth concurrency analysis,” CSUR , 2019
2019
Closest in time.
X. Chen et al. , “On-the-fly parallel data shuffling for graph processing on OpenCL-based FPGAs,” in FPL , 2019
2019
Closest in time.
2019
Closest in time.
T. De Matteis et al. , “Streaming message interface: High-performance distributed memory programming on reconfigurable hardware,” in SC , 2019
2019
Closest in time.
P. Gorlani et al. , “OpenCL implementation of Cannon’s matrix multiplication algorithm on Intel Stratix 10 FPGAs,” in ICFPT , 2019
2019
Closest in time.
M. Besta et al. , “Graph processing on FPGAs: Taxonomy, survey, challenges,” arXiv:1903.06697 , 2019
2019
Closest in time.
——, “Substream-centric maximum matchings on FPGA,” in FPGA , 2019
2019
Closest in time.
H. Eran et al. , “Design patterns for code reuse in HLS packet processing pipelines,” in FCCM , 2019
2019
Closest in time.
T. Ben-Nun et al. , “Stateful dataflow multigraphs: A data-centric model for performance portability on heterogeneous architectures,” in SC , 2019
2019
Closest in time.
J. S. da Silva et al. , “Module-per-Object: a human-driven methodology for C++-based high-level synthesis design,” in FCCM , 2019
2019
Closest in time.
Intel High-Level Synthesis (HLS) Compiler. https://www.intel.com/content/www/us/en/software/programmable/quartus-prime/hls-compiler.html . Accessed May 15, 2020
2020
Closest in time.
Mentor Graphics. Catapult high-level synthesis. https://www.mentor.com/hls-lp/catapult-high-level-synthesis/c-systemc-hls . Accessed May 15, 2020
2020
Closest in time.
T. Young-Schultz et al. , “Using OpenCL to enable software-like development of an FPGA-accelerated biophotonic cancer treatment simulator,” in FPGA , 2020
2020
Closest in time.
J. de Fine Licht et al. , “Flexible communication avoiding matrix multiplication on FPGA with high-level synthesis,” in FPGA , 2020
2020
Closest in time.
J. Li et al. , “HeteroHalide: From image processing DSL to efficient FPGA acceleration,” in FPGA , 2020
2020
Closest in time.
2020
Closest in time.
J. de Fine Licht et al. , “StencilFlow: Mapping large stencil programs to distributed spatial computing systems,” in CGO , 2021
2021
Closest in time.