Fetching the paper…
Reading the bibliography…
Stencil computation is one of the most used kernels in a wide variety of scientific applications, ranging from large-scale weather prediction to solving partial differential equations.
H. S. Stone, “A Logic-in-Memory Computer,” IEEE Trans. Comput. , 1970
1970
Earlier work this paper cites.
G. A. McMechan, “Migration by Extrapolation of Time-Dependent Boundary Values,” Geophysical Prospecting , 1983
1983
Earlier work this paper cites.
J. Canny, “A Computational Approach to Edge Detection,” TPAMI , 1986
1986
Earlier work this paper cites.
A. Taflove, “Review of the Formulation and Applications of the Finite-Difference Time-Domain Method for Numerical Modeling of Electromagnetic Wave Interactions With Arbitrary Structures,” Wave Motion , 1988
1988
Earlier work this paper cites.
A. Saini, “Design of the Intel Pentium Processor,” in ICCD , 1993
1993
Earlier work this paper cites.
J. D. Anderson and J. Wendt, Computational Fluid Dynamics . Springer, 1995
1995
Earlier work this paper cites.
M. Gokhale, B. Holmes, and K. Iobst, “Processing in Memory: The Terasys Massively Parallel PIM Array,” Computer , 1995
1995
Earlier work this paper cites.
D. Patterson, T. Anderson, N. Cardwell et al. , “A Case for Intelligent RAM,” IEEE Micro , 1997
1997
Earlier work this paper cites.
J. W. Demmel, Applied Numerical Linear Algebra . SIAM, 1997
1997
Earlier work this paper cites.
D. Jaggar, “ARM Architecture and Systems,” IEEE Annals of the History of Computing , 1997
1997
Earlier work this paper cites.
M. Oskin, F. T. Chong, and T. Sherwood, “Active Pages: A Computation Model for Intelligent Memory,” in ISCA , 1998
1998
Earlier work this paper cites.
K. Keeton, D. A. Patterson, and J. M. Hellerstein, “A Case for Intelligent Disks (IDISKs),” SIGMOD Rec. , 1998
1998
Earlier work this paper cites.
A. Acharya, M. Uysal, and J. Saltz, “Active Disks: Programming Model, Algorithms and Evaluation,” in ASPLOS , 1998
1998
Earlier work this paper cites.
G. Doms and U. Schättler, “The Nonhydrostatic Limited-Area Model LM (Lokalmodel) of the DWD. Part I: Scientific Documentation,” DWD, GB Forschung und Entwicklung , 1999
1999
Earlier work this paper cites.
Y. Kang, W. Huang, S.-M. Yoo, D. Keen, Z. Ge, V. Lam, P. Pattnaik, and J. Torrellas, “FlexRAM: Toward an Advanced Intelligent Memory System,” in ICCD , 1999
1999
Earlier work this paper cites.
D. G. Elliott, M. Stumm, W. M. Snelgrove et al. , “Computational RAM: Implementing Processors in Memory,” IEEE Design & Test , 1999
1999
Earlier work this paper cites.
G. Allen, T. Goodale, G. Lanfermann, T. Radke, E. Seidel, W. Benger, H.-C. Hege, A. Merzky, J. Masso, and J. Shalf, “Solving Einstein’s Equations on Supercomputers,” Computer , 1999
1999
Earlier work this paper cites.
B. Khailany, W. J. Dally, U. J. Kapasi, P. Mattson, J. Namkoong, J. D. Owens, B. Towles, A. Chang, and S. Rixner, “Imagine: Media Processing with Streams,” IEEE Micro , 2001
2001
Earlier work this paper cites.
J. Draper, J. Chame, M. Hall, C. Steele, T. Barrett, J. LaCoss, J. Granacki, J. Shin, C. Chen, C. W. Kang, I. Kim, and G. Daglikoca, “The Architecture of the DIVA Processing-in-Memory Chip,” in SC , 2002
2002
Earlier work this paper cites.
S. Ciricescu, R. Essick, B. Lucas, P. May, K. Moat, J. Norris, M. Schuette, and A. Saidi, “The Reconfigurable Streaming Vector Processor (RSVP),” in MICRO , 2003
2003
Earlier work this paper cites.
O. Mutlu, J. Stark, C. Wilkerson, and Y. N. Patt, “Runahead execution: An alternative to very large instruction windows for out-of-order processors,” in The Ninth International Symposium on High-Performance Computer Architecture, 2003. HPCA-9 2003. Proceedings. IEEE, 2003, pp. 129–140
2003
Earlier work this paper cites.
——, “Runahead Execution: An Effective Alternative to Large Instruction Windows,” in IEEE Micro , 2003
2003
Earlier work this paper cites.
P. Colella, “Defining Software Requirements for Scientific Computing,” https://www.krellinst.org/doecsgf/conf/2013/pres/pcolella.pdf , 2004
2004
Earlier work this paper cites.
J. B. Brockman, S. Thoziyoor, S. K. Kuntz, and P. M. Kogge, “A Low Cost, Multithreaded Processing-in-Memory System,” in WMPI , 2004
2004
Earlier work this paper cites.
S. Kamil, P. Husbands, L. Oliker, J. Shalf, and K. Yelick, “Impact of Modern Memory Subsystems on Cache Optimizations for Stencil Computations,” in MSP , 2005
2005
Earlier work this paper cites.
O. Mutlu, H. Kim, and Y. N. Patt, “Techniques for Efficient Processing in Runahead Execution Engines,” in ISCA , 2005
2005
Earlier work this paper cites.
M. Frigo and V. Strumpen, “The Memory Behavior of Cache Oblivious Stencil Computations,” The Journal of Supercomputing , 2007
2007
Earlier work this paper cites.
K. Datta, M. Murphy, V. Volkov, S. Williams, J. Carter, L. Oliker, D. Patterson, J. Shalf, and K. Yelick, “Stencil Computation Optimization and Auto-Tuning on State-Of-The-Art Multicore Architectures ,” in SC , 2008
2008
Earlier work this paper cites.
K. Datta, S. Kamil, S. Williams, L. Oliker, J. Shalf, and K. Yelick, “Optimization and Performance Modeling of Stencil Computations on Modern Microprocessors,” SIAM Review , 2009
2009
Earlier work this paper cites.
S. Williams, A. Waterman, and D. Patterson, “Roofline: An Insightful Visual Performance Model for Multicore Architectures,” CACM , 2009
2009
Earlier work this paper cites.
W. Augustin, V. Heuveline, and J.-P. Weiss, “Optimized Stencil Computation using In-Place Calculation on Modern Multicore Systems,” in ECPP , 2009
2009
Earlier work this paper cites.
K. Datta and K. A. Yelick, Auto-Tuning Stencil Codes for Cache-Based Multicore Platforms . University of California, Berkeley, 2009
2009
Earlier work this paper cites.
R. Strzodka, M. Shaheen, D. Pajak, and H.-P. Seidel, “Cache Oblivious Parallelograms in Iterative Stencil Computations,” in SC , 2010
2010
Earlier work this paper cites.
A. Nguyen, N. Satish, J. Chhugani, C. Kim, and P. Dubey, “3.5-D Blocking Optimization for Stencil Computations on Modern CPUs And GPUs,” in SC , 2010
2010
Earlier work this paper cites.
T. Brandvik and G. Pullan, “SBLOCK: A Framework for Efficient Stencil-Based PDE Solvers on Multi-Core Platforms,” in ICCIT , 2010
2010
Earlier work this paper cites.
E. H. Phillips and M. Fatica, “Implementing the Himeno Benchmark with CUDA on GPU Clusters,” in IPDPS , 2010
2010
Earlier work this paper cites.
M. Christen, O. Schenk, and H. Burkhart, “Patus: A Code Generation and Autotuning Framework for Parallel Iterative Stencil Computations on Modern Microarchitectures,” in IPDPS , 2011
2011
Earlier work this paper cites.
Y. Tang, R. A. Chowdhury, B. C. Kuszmaul, C.-K. Luk, and C. E. Leiserson, “The Pochoir Stencil Compiler,” in SPAA , 2011
2011
Earlier work this paper cites.
J. Meng and K. Skadron, “A Performance Study for Iterative Stencil Loops on GPUs with Ghost Zone Optimizations,” IJPP , 2011
2011
Earlier work this paper cites.
T. Henretty, K. Stock, L.-N. Pouchet, F. Franchetti, J. Ramanujam, and P. Sadayappan, “Data Layout Transformation for Stencil Computations on Short-Vector SIMD Architectures,” in CC , 2011
2011
Earlier work this paper cites.
N. Binkert, B. Beckmann, G. Black, S. K. Reinhardt, A. Saidi, A. Basu, J. Hestness, D. R. Hower, T. Krishna, S. Sardashti et al. , “The gem5 Simulator,” Comp. Arch. News , 2011
2011
Earlier work this paper cites.
M. Schönherr, K. Kucher, M. Geier, M. Stiebler, S. Freudiger, and M. Krafczyk, “Multi-Thread Implementations of the Lattice Boltzmann Method on Non-Uniform Grids for CPUs and GPUs,” Computers & Mathematics with Applications , 2011
2011
Earlier work this paper cites.
J. Jaeger and D. Barthou, “Automatic Efficient Data Layout for Multithreaded Stencil Codes on CPU Sand GPUs,” in HiPC , 2012
2012
Earlier work this paper cites.
J. Holewinski, L.-N. Pouchet, and P. Sadayappan, “High-Performance Code Generation for Stencil Computations on GPU Architectures,” in SC , 2012
2012
Earlier work this paper cites.
H. Ltaief, P. Luszczek, and J. Dongarra, “Profiling High Performance Dense Linear Algebra Algorithms on Multicore Architectures for Power and Energy Efficiency,” Computer Science-Research and Development , 2012
2012
Earlier work this paper cites.
O. Fuhrer, C. Osuna, X. Lapillonne, T. Gysi, M. Bianco, and T. Schulthess, “Towards GPU-Accelerated Operational Weather Forecasting,” in GTC , 2013
2013
Earlier work this paper cites.
J. Ragan-Kelley, C. Barnes, A. Adams, S. Paris, F. Durand, and S. Amarasinghe, “Halide: A Language and Compiler for Optimizing Parallelism, Locality, and Recomputation in Image Processing Pipelines,” in PLDI , 2013
2013
Earlier work this paper cites.
V. Seshadri, Y. Kim, C. Fallin, D. Lee, R. Ausavarungnirun, G. Pekhimenko, Y. Luo, O. Mutlu, P. B. Gibbons, M. A. Kozuch et al. , “Rowclone: fast and energy-efficient in-dram bulk data copy and initialization,” in Proceedings of the 46th Annual IEEE/ACM International Symposium on Microarchitecture . ACM, 2013, pp. 185–197
2013
Earlier work this paper cites.
Q. Zhu, T. Graf, H. E. Sumbul, L. Pileggi, and F. Franchetti, “Accelerating Sparse Matrix-Matrix Multiplication with 3D-Stacked Logic-in-Memory Hardware,” in HPEC , 2013
2013
Earlier work this paper cites.
A. Basu, J. Gandhi, J. Chang, M. D. Hill, and M. M. Swift, “Efficient Virtual Memory for Big Memory Servers,” in ISCA , 2013
2013
Earlier work this paper cites.
L. Szustak, K. Rojek, and P. Gepner, “Using Intel Xeon Phi Coprocessor to Accelerate Computations in MPDATA Algorithm,” in PPAM , 2013
2013
Earlier work this paper cites.
N. Maruyama and T. Aoki, “Optimizing Stencil Computations for NVIDIA Kepler GPUs,” in HiStencils , 2014
2014
Earlier work this paper cites.
C. Olschanowsky, M. M. Strout, S. Guzik, J. Loffeld, and J. Hittinger, “A Study on Balancing Parallelism, Data Locality, and Recomputation in Existing PDE Solvers,” in SC , 2014
2014
Earlier work this paper cites.
K. Sano, Y. Hatsuda, and S. Yamamoto, “Multi-FPGA Accelerator for Scalable Stencil Computation With Constant Memory Bandwidth,” TPDS , 2014
2014
Earlier work this paper cites.
R. Wester and J. Kuper, “Deriving Stencil Hardware Accelerators from a Single Higher-Order Function,” CPA , 2014
2014
Earlier work this paper cites.
D. Zhang, N. Jayasena, A. Lyashevsky, J. L. Greathouse, L. Xu, and M. Ignatowski, “TOP-PIM: Throughput-Oriented Programmable Processing in Memory,” in HPDC , 2014
2014
Earlier work this paper cites.
S. H. Pugsley, J. Jestes, H. Zhang, R. Balasubramonian, V. Srinivasan, A. Buyuktosunoglu, A. Davis, and F. Li, “NDC: Analyzing the Impact of 3D-Stacked Memory+Logic Devices on MapReduce Workloads,” in ISPASS , 2014
2014
Earlier work this paper cites.
M. Kang, M.-S. Keel, N. R. Shanbhag, S. Eilert, and K. Curewitz, “An Energy-Efficient VLSI Architecture for Pattern Recognition via Deep Embedding of Computation in SRAM,” in ICASSP , 2014
2014
Earlier work this paper cites.
P. Rosenfeld, “Performance Exploration of the Hybrid Memory Cube,” Ph.D. dissertation, University of Maryland, 2014
2014
Earlier work this paper cites.
Y. S. Shao, B. Reagen, G.-Y. Wei, and D. Brooks, “Aladdin: A Pre-RTL, Power-Performance Accelerator Simulator Enabling Large Design Space Exploration of Customized Architectures,” in ISCA , 2014
2014
Earlier work this paper cites.
D. U. Lee, K. W. Kim, K. W. Kim, H. Kim, J. Y. Kim, Y. J. Park, J. H. Kim, D. S. Kim, H. B. Park, J. W. Shin et al. , “A 1.2V 8Gb 8-Channel 128GB/s High-Bandwidth Memory (HBM) Stacked DRAM with Effective Microbump I/O Test Methods Using 29nm Process and TSV,” in ISSCC , 2014
2014
Earlier work this paper cites.
M. Wahib and N. Maruyama, “Scalable Kernel Fusion for Memory-Bound GPU Applications,” in SC , 2014
2014
Earlier work this paper cites.
T. Gysi, T. Grosser, and T. Hoefler, “Modesto: Data-Centric Analytic Optimization of Complex Stencil Programs on Heterogeneous Architectures,” in SC , 2015
2015
Earlier work this paper cites.
H. Stengel, J. Treibig, G. Hager, and G. Wellein, “Quantifying Performance Bottlenecks of Stencil Computations Using the Execution-Cache-Memory Model,” in ICS , 2015
2015
Cited alongside, same era.
R. Cattaneo, G. Natale, C. Sicignano, D. Sciuto, and M. D. Santambrogio, “On How to Accelerate Iterative Stencil Loops: A Scalable Streaming-Based Approach,” TACO , 2015
2015
Cited alongside, same era.
J. Ahn, S. Yoo, O. Mutlu, and K. Choi, “PIM-Enabled Instructions: A Low-Overhead, Locality-Aware Processing-in-Memory Architecture,” in ISCA , 2015
2015
Cited alongside, same era.
J. Ahn, S. Hong, S. Yoo, O. Mutlu, and K. Choi, “A Scalable Processing-in-Memory Accelerator for Parallel Graph Processing,” in ISCA , 2015
2015
Cited alongside, same era.
R. Nair, S. F. Antao, C. Bertolli, P. Bose et al. , “Active Memory Cube: A Processing-in-Memory Architecture for Exascale Systems,” IBM JRD , 2015
A. Boroumand, S. Ghose, Y. Kim, R. Ausavarungnirun, E. Shiu, R. Thakur, D. Kim, A. Kuusela, A. Knies, P. Ranganathan et al. , “Google Workloads for Consumer Devices: Mitigating Data Movement Bottlenecks,” in ASPLOS , 2018
2018
Later among the works it cites.
Q. Deng, L. Jiang, Y. Zhang, M. Zhang, and J. Yang, “DrAcc: A DRAM Based Accelerator for Accurate CNN Inference,” in DAC , 2018
2018
Later among the works it cites.
D. Fujiki, S. Mahlke, and R. Das, “In-Memory Data Parallel Processor,” in ASPLOS , 2018
2018
Later among the works it cites.
H. R. Zohouri, A. Podobas, and S. Matsuoka, “Combined spatial and temporal blocking for high-performance stencil computation on FPGAs using OpenCL,” in Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays , 2018, pp. 153–162
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
A. Farmahini-Farahani, J. H. Ahn, K. Morrow, and N. S. Kim, “NDA: Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules,” in HPCA , 2015
2015
Cited alongside, same era.
V. Seshadri, K. Hsieh, A. Boroumabd, D. Lee, M. A. Kozuch, O. Mutlu, P. B. Gibbons, and T. C. Mowry, “Fast Bulk Bitwise AND and OR in DRAM,” CAL , 2015
2015
Cited alongside, same era.
B. Akin, F. Franchetti, and J. C. Hoe, “Data Reorganization in Memory Using 3D-Stacked DRAM,” in ISCA , 2015
2015
Cited alongside, same era.
J. H. Lee, J. Sim, and H. Kim, “BSSync: Processing Near Memory for Machine Learning Workloads with Bounded Staleness Consistency Models,” in PACT , 2015
2015
Cited alongside, same era.
V. Seshadri, T. Mullins, A. Boroumand, O. Mutlu, P. B. Gibbons, M. A. Kozuch, and T. C. Mowry, “Gather-Scatter DRAM: in-DRAM Address Translation to Improve the Spatial Locality of Non-unit Strided Accesses,” in MICRO , 2015
2015
Cited alongside, same era.
M. Gao, G. Ayers, and C. Kozyrakis, “Practical Near-Data Processing for In-Memory Analytics Frameworks,” in PACT , 2015
2015
Cited alongside, same era.
A. Morad, L. Yavits, and R. Ginosar, “Gp-simd processing-in-memory,” ACM Trans. Archit. Code Optim. , vol. 11, no. 4, pp. 53:1–53:26, Jan. 2015. [Online]. Available: http://doi.acm.org/10.1145/2686875
2015
Cited alongside, same era.
G. Dai, T. Huang, Y. Chi, J. Zhao, G. Sun, Y. Liu, Y. Wang, Y. Xie, and H. Yang, “GraphH: A Processing-in-Memory Architecture for Large-Scale Graph Processing,” in IEEE TCAD , 2018
2018
Later among the works it cites.
P.-A. Tsai, C. Chen, and D. Sanchez, “Adaptive Scheduling for Systems with Asymmetric Memory Hierarchies,” in MICRO , 2018
2018
Later among the works it cites.
J. van Lunteren, R. Luijten, D. Diamantopoulos, F. Auernhammer, C. Hagleitner, L. Chelini, S. Corda, and G. Singh, “Coherently Attached Programmable Near-Memory Acceleration Platform and its Application to Stencil Processing,” in DATE , 2019
2019
Later among the works it cites.
G. Singh, D. Diamantopoulos, C. Hagleitner, S. Stuijk, and H. Corporaal, “NARMADA: Near-Memory Horizontal Diffusion Accelerator for Scalable Stencil Computations,” in FPL , 2019
2019
Later among the works it cites.
J. Li, X. Wang, A. Tumeo, B. Williams, J. D. Leidel, and Y. Chen, “PIMS: A Lightweight Processing-in-Memory Accelerator for Stencil Computations,” in ISMS , 2019
2019
Later among the works it cites.
H. M. Waidyasooriya and M. Hariyama, “Multi-FPGA Accelerator Architecture for Stencil Computation Exploiting Spacial and Temporal Scalability,” IEEE Access , 2019
2019
Later among the works it cites.
G. Singh, D. Diamantopoulos, S. Stuijk, C. Hagleitner, and H. Corporaal, “Low Precision Processing for High Order Stencil Computations,” in Springer LNCS , 2019
2019
Later among the works it cites.
L. Szustak and P. Bratek, “Performance Portable Parallel Programming of Heterogeneous Stencils across Shared-memory Platforms with Modern Intel Processors,” IJHPCA , 2019
2019
Later among the works it cites.
M. Imani, S. Gupta, Y. Kim, and T. Rosing, “FloatPIM: In-memory Acceleration of Deep Neural Network Training with High Precision,” in ISCA , 2019
2019
Later among the works it cites.
S. Ghose, A. Boroumand, J. S. Kim, J. Gómez-Luna, and O. Mutlu, “Processing-in-Memory: A Workload-Driven Perspective,” IBM JRD , 2019
2019
Later among the works it cites.
O. Mutlu, S. Ghose, J. Gómez-Luna, and R. Ausavarungnirun, “Processing Data Where It Makes Sense: Enabling In-Memory Computation,” MicPro , 2019
2019
Later among the works it cites.
D. Fujiki, S. Mahlke, and R. Das, “Duality Cache for Data Parallel Acceleration,” in ISCA , 2019
2019
Later among the works it cites.
A. Nag, C. Ramachandra, R. Balasubramonian, R. Stutsman, E. Giacomin, H. Kambalasubramanyam, and P.-E. Gaillardon, “GenCache: Leveraging In-Cache Operators for Efficient Sequence Alignment,” in MICRO , 2019
2019
Later among the works it cites.
A. Pattnaik, X. Tang, O. Kayiran, A. Jog, A. Mishra, M. T. Kandemir, A. Sivasubramaniam, and C. R. Das, “Opportunistic computing in GPU architectures,” in ISCA , 2019
2019
Later among the works it cites.
A. Boroumand, S. Ghose, M. Patel, H. Hassan, B. Lucia, R. Ausavarungnirun, K. Hsieh, N. Hajinazar, K. T. Malladi, H. Zheng et al. , “CoNDA: Efficient Cache Coherence Support for Near-Data Accelerators,” in ISCA , 2019
2019
Later among the works it cites.
G. Singh, G. , G. F. Oliveira, S. Corda, S. Stuijk, O. Mutlu, and H. Corporaal, “NAPEL: Near-Memory Computing Application Performance Prediction via Ensemble Learning,” in DAC , 2019
2019
Later among the works it cites.
O. Mutlu, S. Ghose, J. Gómez-Luna, and R. Ausavarungnirun, “Enabling Practical Processing in and Near Memory for Data-Intensive Computing,” in DAC , 2019
2019
Later among the works it cites.
S. Angizi and D. Fan, “GraphiDe: A Graph Processing Accelerator leveraging In-DRAM-Computing,” in GLSVLSI , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
F. Thaler, S. Moosbrugger, C. Osuna, M. Bianco, H. Vogt, A. Afanasyev, L. Mosimann, O. Fuhrer, T. C. Schulthess, and T. Hoefler, “Porting the COSMO Weather Model to Manycore CPUs,” in PASC , 2019
2019
Later among the works it cites.
Z. Wang and T. Nowatzki, “Stream-Based Memory Access Specialization for General Purpose Processors,” in ISCA , 2019
2019
Later among the works it cites.
J. S. Kim, M. Patel, H. Hassan, L. Orosa, and O. Mutlu, “D-RaNGe: Using Commodity DRAM Devices to Generate True Random Numbers with Low Latency and High Throughput,” in HPCA , 2019
2019
Later among the works it cites.
Y. Zhuo, M. Zhang, R. Wang, D. Niu, Y. Wang, and a. Qian, “GraphQ: Scalable PIM-Based Graph Processing,” in MICRO , 2019
2019
Later among the works it cites.
O. Anjum, S. Garcia de Gonzalo, M. Hidayetoglu, and W.-M. Hwu, “An Efficient GPU Implementation Technique for Higher-Order 3D Stencils,” in HPCC , 2019
2019
Later among the works it cites.
M. Koraei, O. Fatemi, and M. Jahre, “DCMI: A Scalable Strategy for Accelerating Iterative Stencil Loops on FPGAs,” TACO , 2019
2019
Later among the works it cites.
S. Huang, L.-W. Chang, I. El Hajj, S. Garcia de Gonzalo, J. Gómez-Luna, S. R. Chalamalasetti, M. El-Hadedy, D. Milojicic, O. Mutlu, D. Chen et al. , “Analysis and Modeling of Collaborative Execution Strategies for Heterogeneous CPU-FPGA Architectures,” in ICPE , 2019
2019
Later among the works it cites.
G. Singh, D. Diamantopoulos, C. Hagleitner, J. Gómez-Luna, S. Stuijk, O. Mutlu, and H. Corporaal, “NERO: A Near High-Bandwidth Memory Stencil Accelerator for Weather Prediction Modeling,” in FPL , 2020
2020
Later among the works it cites.
H. E. Yantır, A. M. Eltawil, and K. N. Salama, “Efficient Acceleration of Stencil Applications through In-Memory Computing,” Micromachines , 2020
2020
Later among the works it cites.
E. Lockerman, A. Feldmann, M. Bakhshalipour, A. Stanescu, S. Gupta, D. Sanchez, and N. Beckmann, “Livia: Data-Centric Computing Throughout the Memory Hierarchy,” in ASPLOS , 2020
2020
Later among the works it cites.
I. Fernandez, R. Quislant, E. Gutiérrez, O. Plata, C. Giannoula, M. Alser, J. Gómez-Luna, and O. Mutlu, “NATSA: A Near-Data Processing Accelerator for Time Series Analysis,” in ICCD , 2020
2020
Later among the works it cites.
D. S. Cali, G. S. Kalsi, Z. Bingöl, C. Firtina, L. Subramanian, J. S. Kim, R. Ausavarungnirun, M. Alser, J. Gomez-Luna, A. Boroumand et al. , “GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis,” in MICRO , 2020
2020
Later among the works it cites.
P. Gu, X. Xie, Y. Ding, G. Chen, W. Zhang, D. Niu, and Y. Xie, “iPIM: Programmable In-Memory Image Processing Accelerator Using Near-Bank Architecture,” in ISCA , 2020
2020
Later among the works it cites.
Y. Huang, L. Zheng, P. Yao, J. Zhao, X. Liao, H. Jin, and J. Xue, “A Heterogeneous PIM Hardware-Software Co-Design for Energy-Efficient Graph Processing,” in IPDPS , 2020
2020
Later among the works it cites.
M. Alser, Z. Bingol, D. Senol Cali, S. Kim, Jeremie andGhose, C. Alkan, and O. Mutlu, “Accelerating Genome Analysis: A Primer on an Ongoing Journey,” in IEEE MICRO , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
J. Jiang, Z. Wang, X. Liu, J. Gómez-Luna, N. Guan, Q. Deng, W. Zhang, and O. Mutlu, “Boyi: A Systematic Framework for Automatically Deciding the Right Execution Model of OpenCL Applications on FPGAs,” in FPGA , 2020
2020
Later among the works it cites.
G. Singh, D. Diamantopoulos, J. Gómez-Luna, C. Hagleitner, S. Stuijk, H. Corporaal, and O. Mutlu, “Accelerating Weather Prediction using Near-Memory Reconfigurable Fabric,” in ACM Transactions on Reconfigurable Technology and Systems , 2021
2021
Closest in time.
O. Mutlu, S. Ghose, J. Gómez-Luna, and R. Ausavarungnirun, “A Modern Primer on Processing in Memory,” in Emerging Computing: From Devices to Systems — Looking Beyond Moore and Von Neumann . Springer, 2021
2021
Closest in time.
2021
Closest in time.
G. F. Oliveira, J. Gómez-Luna, L. Orosa, S. Ghose, N. Vijaykumar, I. Fernandez, M. Sadrosadati, and O. Mutlu, “DAMOV: A New Methodology and Benchmark Suite for Evaluating Data Movement Bottlenecks,” IEEE Access , 2021
2021
Closest in time.
A. V. Nori, R. Bera, S. Balachandran, J. Rakshit, O. J. Omer, A. Abuhatzera, B. Kuttanna, and S. Subramoney, “REDUCT: Keep it Close, Keep it Cool! : Efficient Scaling of DNN Inference on Multi-core CPUs with Near-Cache Compute,” in ISCA , 2021
2021
Closest in time.
N. Hajinazar, G. F. Oliveira, S. Gregorio, J. D. Ferreira, N. M. Ghiasi, M. Patel, M. Alser, S. Ghose, J. Gómez-Luna, and O. Mutlu, “SIMDRAM: A Framework for Bit-Serial SIMD Processing Using DRAM,” in ASPLOS , 2021
2021
Closest in time.
M. Besta, R. Kanakagiri, G. Kwasniewski, R. Ausavarungnirun, J. Beránek, K. Kanellopoulos, K. Janda, Z. Vonarburg-Shmaria, L. Gianinazzi, I. Stefan et al. , “SISA: Set-Centric Instruction Set Architecture for Graph Mining on Processing-in-Memory Systems,” in MICRO , 2021
2021
Closest in time.
G. Singh, M. Alser, D. S. Cali, D. Diamantopoulos, J. Gómez-Luna, H. Corporaal, and O. Mutlu, “FPGA-Based Near-Memory Acceleration of Modern Data-Intensive Applications,” IEEE Micro , 2021
2021
Closest in time.
C. Giannoula, N. Vijaykumar, N. Papadopoulou, V. Karakostas, I. Fernandez, J. Gómez-Luna, L. Orosa, N. Koziris, G. Goumas, and O. Mutlu, “SynCron: Efficient Synchronization Support for Near-Data-Processing Architectures,” in HPCA , 2021
2021
Closest in time.
A. Boroumand, S. Ghose, B. Akin, R. Narayanaswami, G. F. Oliveira, X. Ma, E. Shiu, and O. Mutlu, “Google Neural Network Models for Edge Devices: Analyzing and Mitigating Machine Learning Inference Bottlenecks,” in PACT , 2021
2021
Closest in time.
L. Yavits, R. Kaplan, and R. Ginosar, “GIRAF: General Purpose In-Storage Resistive Associative Framework,” TPDS , 2021
2021
Closest in time.
S. Kang, J. An, J. Kim, and S.-W. Jun, “Near-Storage Accelerator for High-Performance Log Analytics,” in MICRO , 2021
2021
Closest in time.
2021
Closest in time.
O. Ataberk, M. Patel, A. G. Yaglikci, H. Lu, J. S. Kim, F. N. Bostanci, N. Vijaykumar, O. Ergin, and O. Mutlu, “QUAC-TRNG: High-Throughput True Random Number Generation Using Quadruple Row Activation in Commodity DRAM Chips,” in ISCA , 2021
2021
Closest in time.
O. Mutlu, “Intelligent Architectures for Intelligent Computing Systems,” in DATE , 2021
2021
Closest in time.
Q. Sun, Y. Liu, H. Yang, Z. Jiang, X. Liu, M. Dun, Z. Luan, and D. Qian, “csTuner: Scalable Auto-tuning Framework for Complex Stencil Computation on GPUs,” in CLUSTER , 2021
2021
Closest in time.
J. de Fine Licht, A. Kuster, T. De Matteis, T. Ben-Nun, D. Hofer, and T. Hoefler, “StencilFlow: Mapping Large Stencil Programs to Distributed Spatial Computing Systems,” in CGO , 2021
2021
Closest in time.
E. Reggiani, E. Del Sozzo, D. Conficconi, G. Natale, C. Moroni, and M. D. Santambrogio, “Enhancing the Scalability of Multi-FPGA Stencil Computations via Highly Optimized HDL Components,” TRETS , 2021
2021
Closest in time.
——, “Benchmarking a New Paradigm: Experimental Analysis and Characterization of a Real Processing-in-Memory System,” IEEE Access , 2022
2022
Closest in time.
N. M. Ghiasi, J. Park, H. Mustafa, J. Kim, A. Olgun, A. Gollwitzer, D. S. Cali, C. Firtina, H. Mao, N. A. Alserr, R. Ausavarungnirun, N. Vijaykumar, M. Alser, and O. Mutlu, “GenStore: A High-Performance In-Storage Processing System for Genome Sequence Analysis,” in ASPLOS , 2022
2022
Closest in time.
A. Sohrabizadeh, C. H. Yu, M. Gao, and J. Cong, “AutoDSE: Enabling Software Programmers to Design Efficient FPGA Accelerators,” TODAES , 2022
2022
Closest in time.