Fetching the paper…
Reading the bibliography…
Many modern workloads, such as neural networks, databases, and graph processing, are fundamentally memory-bound.
M. J. Flynn, “Very High-speed Computing Systems,” Proc. IEEE , 1966
1966
Earlier work this paper cites.
G. M. Amdahl, “Validity of the Single Processor Approach to Achieving Large Scale,” in AFIPS , 1967
1967
Earlier work this paper cites.
W. H. Kautz, “Cellular Logic-in-Memory Arrays,” IEEE TC , 1969
1969
Earlier work this paper cites.
H. S. Stone, “A Logic-in-Memory Computer,” IEEE TC , 1970
1970
Earlier work this paper cites.
J. E. Thornton, CDC 6600: Design of a Computer , 1970
1970
Earlier work this paper cites.
S. Needleman and C. Wunsch, “A General Method Applicable to the Search for Similarities in the Amino Acid Sequence of Two Proteins,” Journal of Molecular Biology , 1970
1970
Earlier work this paper cites.
D. E. Knuth, “Optimum Binary Search Trees,” Acta informatica , 1971
1971
Earlier work this paper cites.
B. J. Smith, “A Pipelined, Shared Resource MIMD Computer,” in ICPP , 1978
1978
Earlier work this paper cites.
D. E. Shaw et al. , “The NON-VON Database Machine: A Brief Overview,” IEEE Database Eng. Bull. , 1981
1981
Earlier work this paper cites.
B. J. Smith, “Architecture and Applications of the HEP Multiprocessor Computer System,” in SPIE, Real-Time signal processing IV , 1981
1981
Earlier work this paper cites.
A. Bundy and L. Wallen, “Breadth-first Search,” in Catalogue of Artificial Intelligence Tools . Springer, 1984
1984
Earlier work this paper cites.
S. Ceri and G. Gottlob, “Translating SQL into Relational Algebra: Optimization, Semantics, and Equivalence of SQL Queries,” IEEE TSE , 1985
1985
Earlier work this paper cites.
G. E. Hinton, “Learning Translation Invariant Recognition in a Massively Parallel Networks,” in PARLE , 1987
1987
Earlier work this paper cites.
J. L. Gustafson, “Reevaluating Amdahl’s Law,” CACM , 1988
1988
Earlier work this paper cites.
G. E. Blelloch, “Scans as Primitive Parallel Operations,” IEEE TC , 1989
1989
Earlier work this paper cites.
A. Saini, “Design of the Intel Pentium Processor,” in ICCD , 1993
1993
Earlier work this paper cites.
P. M. Kogge, “EXECUBE - A New Architecture for Scaleable MPPs,” in ICPP , 1994
1994
Earlier work this paper cites.
M. Gokhale et al. , “Processing in Memory: The Terasys Massively Parallel PIM Array,” IEEE Computer , 1995
1995
Earlier work this paper cites.
J. D. McCalpin, “Memory Bandwidth and Machine Balance in Current High Performance Computers,” IEEE TCCA newsletter , 1995
1995
Earlier work this paper cites.
R. F. Boisvert et al. , “Matrix Market: A Web Resource for Test Matrix Collections,” in Quality of Numerical Software , 1996
1996
Earlier work this paper cites.
D. Patterson et al. , “A Case for Intelligent RAM,” IEEE Micro , 1997
1997
Earlier work this paper cites.
D. Jaggar, “ARM Architecture and Systems,” IEEE Annals of the History of Computing , 1997
1997
Earlier work this paper cites.
M. Oskin et al. , “Active Pages: A Computation Model for Intelligent Memory,” in ISCA , 1998
1998
Earlier work this paper cites.
J. H. van Hateren and A. van der Schaaf, “Independent Component Filters of Natural Images Compared with Simple Cells in Primary Visual Cortex,” Proceedings of the Royal Society of London. Series B: Biological Sciences , 1998
1998
Earlier work this paper cites.
Y. Kang et al. , “FlexRAM: Toward an Advanced Intelligent Memory System,” in ICCD , 1999
1999
Earlier work this paper cites.
D. G. Elliott et al. , “Computational RAM: Implementing Processors in Memory,” IEEE Design & Test of Computers , 1999
1999
Earlier work this paper cites.
K. Mai et al. , “Smart Memories: A Modular Reconfigurable Architecture,” in ISCA , 2000
2000
Earlier work this paper cites.
J. Draper et al. , “The Architecture of the DIVA Processing-in-Memory Chip,” in SC , 2002
2002
Earlier work this paper cites.
J. A. Mandelman et al. , “Challenges and Future Directions for the Scaling of Dynamic Random-Access Memory (DRAM),” IBM JRD , 2002
2002
Earlier work this paper cites.
L. S. Blackford et al. , “An Updated Set of Basic Linear Algebra Subprograms (BLAS),” ACM TOMS , 2002
2002
Earlier work this paper cites.
Y. Saad, Iterative Methods for Sparse Linear Systems, 2nd Edition, Chapter 1: Parallel Implementations . SIAM, 2003
2003
Earlier work this paper cites.
Y. Saad, Iterative Methods for Sparse Linear Systems, 2nd Edition, Chapter 3: Sparse Matrices . SIAM, 2003
2003
Earlier work this paper cites.
C. Lattner and V. Adve, “LLVM: A Compilation Framework for Lifelong Program Analysis & Transformation,” in CGO , 2004
2004
Earlier work this paper cites.
R. Rabenseifner, “Optimization of Collective Reduction Operations,” in ICCS , 2004
2004
Earlier work this paper cites.
D. Chakrabarti et al. , “R-MAT: A Recursive Model for Graph Mining,” in SDM , 2004
2004
Earlier work this paper cites.
D. Weber et al. , “Current and Future Challenges of DRAM Metallization,” in IITC , 2005
2005
Earlier work this paper cites.
P. R. Luszczek et al. , “The HPC Challenge (HPCC) Benchmark Suite,” in SC , 2006
2006
Earlier work this paper cites.
M. Harris, “Optimizing Parallel Reduction in CUDA,” Nvidia Developer Technology , 2007
2007
Earlier work this paper cites.
D. B. Strukov et al. , “The Missing Memristor Found,” Nature , 2008
2008
Earlier work this paper cites.
S. Sengupta et al. , “Efficient Parallel Scan Algorithms for GPUs,” NVIDIA Technical Report NVR-2008-003 , 2008
2008
Earlier work this paper cites.
Y. Dotsenko et al. , “Fast Scan Algorithms on Graphics Processors,” in ICS , 2008
2008
Earlier work this paper cites.
B. C. Lee et al. , “Architecting Phase Change Memory as a Scalable DRAM Alternative,” in ISCA , 2009
2009
Earlier work this paper cites.
M. K. Qureshi et al. , “Scalable High Performance Main Memory System Using Phase-Change Memory Technology,” in ISCA , 2009
2009
Earlier work this paper cites.
P. Zhou et al. , “A Durable and Energy Efficient Main Memory Using Phase Change Memory Technology,” in ISCA , 2009
2009
Earlier work this paper cites.
S. Williams et al. , “Roofline: An Insightful Visual Performance Model for Multicore Architectures,” CACM , 2009
2009
Earlier work this paper cites.
S. Che et al. , “Rodinia: A Benchmark Suite for Heterogeneous Computing,” in IISWC , 2009
2009
Earlier work this paper cites.
S. Hong, “Memory Technology Trend and Future Challenges,” in IEDM , 2010
2010
Earlier work this paper cites.
B. C. Lee et al. , “Phase Change Memory Architecture and the Quest for Scalability,” CACM , 2010
2010
Earlier work this paper cites.
B. C. Lee et al. , “Phase-Change Technology and the Future of Main Memory,” IEEE Micro , 2010
2010
Earlier work this paper cites.
H.-S. P. Wong et al. , “Phase Change Memory,” Proc. IEEE , 2010
2010
Earlier work this paper cites.
G. Hager and G. Wellein, Introduction to High Performance Computing for Scientists and Engineers, Chapter 5: Basics of Parallelization . CRC Press, 2010
2010
Earlier work this paper cites.
L. Luo et al. , “An Effective GPU Implementation of Breadth-first Search,” in DAC , 2010
2010
Earlier work this paper cites.
S. Kvatinsky et al. , “Memristor-Based IMPLY Logic Design Procedure,” in ICCD , 2011
2011
Earlier work this paper cites.
H. David et al. , “Memory Power Management via Dynamic Voltage/Frequency Scaling,” in ICAC , 2011
2011
Earlier work this paper cites.
Q. Deng et al. , “Memscale: Active Low-power Modes for Main Memory,” in ASPLOS , 2011
2011
Earlier work this paper cites.
M. Yuffe et al. , “A Fully Integrated Multi-CPU, GPU and Memory Controller 32nm processor,” in ISSCC , 2011
2011
Earlier work this paper cites.
Intel, “Intel Advanced Vector Extensions Programming Reference,” 2011
2011
Earlier work this paper cites.
E. Cho et al. , “Friendship and Mobility: User Movement in Location-Based Social Networks,” in KDD , 2011
2011
Earlier work this paper cites.
Intel, “Intel 64 and IA-32 Architectures Software Developer’s Manual,” Volume 3B: System Programming Guide , 2011
2011
Earlier work this paper cites.
2012
Earlier work this paper cites.
J. Liu et al. , “RAIDR: Retention-Aware Intelligent DRAM Refresh,” in ISCA , 2012
2012
Earlier work this paper cites.
H.-S. P. Wong et al. , “Metal-Oxide RRAM,” Proc. IEEE , 2012
2012
Earlier work this paper cites.
H. Yoon et al. , “Row Buffer Locality Aware Caching Policies for Hybrid Memories,” in ICCD , 2012
2012
Earlier work this paper cites.
Y. Kim et al. , “A Case for Exploiting Subarray-Level Parallelism (SALP) in DRAM,” in ISCA , 2012
2012
Earlier work this paper cites.
J. L. Hennessy and D. A. Patterson, Computer Architecture - A Quantitative Approach, 5th Edition, Chapter 3: Instruction-level Parallelism and Its Exploitation . Morgan Kaufmann, 2012
2012
Earlier work this paper cites.
J. L. Hennessy and D. A. Patterson, Computer Architecture - A Quantitative Approach, 5th Edition, Chapter 4: Data-level Parallelism in Vector, SIMD, and GPU Architectures . Morgan Kaufmann, 2012
2012
Earlier work this paper cites.
JEDEC, “JESD79-4 DDR4 SDRAM standard,” 2012
2012
Earlier work this paper cites.
T. W. Hungerford, Abstract Algebra: An Introduction, 3rd Edition . Cengage Learning, 2012
2012
Earlier work this paper cites.
I.-J. Sung et al. , “DL: A Data Layout Transformation System for Heterogeneous Computing,” in Innovative Parallel Computing , 2012
2012
Earlier work this paper cites.
N. Bell and J. Hoberock, “Thrust: A Productivity-Oriented Library for CUDA,” in GPU Computing Gems, Jade Edition , 2012
2012
Earlier work this paper cites.
G. Kestor et al. , “Quantifying the Energy Cost of Data Movement in Scientific Applications,” in IISWC , 2013
2013
Earlier work this paper cites.
V. Seshadri et al. , “RowClone: Fast and Energy-Efficient In-DRAM Bulk Data Copy and Initialization,” in MICRO , 2013
2013
Earlier work this paper cites.
Q. Zhu et al. , “Accelerating Sparse Matrix-Matrix Multiplication with 3D-Stacked Logic-in-Memory Hardware,” in HPEC , 2013
2013
Earlier work this paper cites.
J. Liu et al. , “An Experimental Study of Data Retention Behavior in Modern DRAM Devices: Implications for Retention Time Profiling Mechanisms,” in ISCA , 2013
2013
Earlier work this paper cites.
O. Mutlu, “Memory Scaling: A Systems Architecture Perspective,” IMW , 2013
2013
Earlier work this paper cites.
JEDEC, “High Bandwidth Memory (HBM) DRAM,” Standard No. JESD235, 2013
2013
Earlier work this paper cites.
E. Kültürsay et al. , “Evaluating STT-RAM as an Energy-Efficient Main Memory Alternative,” in ISPASS , 2013
2013
Earlier work this paper cites.
T. Rauber and G. Rünger, Parallel Programming, 2nd Edition, Chapter 3: Parallel Programming Models . Springer, 2013
2013
Earlier work this paper cites.
S. Yan et al. , “StreamScan: Fast Scan Algorithms for GPUs without Global Barrier Synchronization,” in PPoPP , 2013
2013
Earlier work this paper cites.
J. Gómez-Luna et al. , “An Optimized Approach to Histogram Computation on GPU,” MVAP , 2013
2013
Earlier work this paper cites.
J. Gómez-Luna et al. , “Performance Modeling of Atomic Additions on GPU Scratchpad Memory,” IEEE TPDS , 2013
2013
Earlier work this paper cites.
G.-J. van den Braak et al. , “Simulation and Architecture Improvements of Atomic Operations on GPU Scratchpad Memory,” in ICCD , 2013
2013
Earlier work this paper cites.
D. Pandiyan and C.-J. Wu, “Quantifying the Energy Cost of Data Movement for Emerging Smart Phone Workloads on Mobile Platforms,” in IISWC , 2014
2014
Earlier work this paper cites.
M. Kang et al. , “An Energy-Efficient VLSI Architecture for Pattern Recognition via Deep Embedding of Computation in SRAM,” in ICASSP , 2014
2014
Earlier work this paper cites.
Y. Levy et al. , “Logic Operations in Memory Using a Memristive Akers Array,” Microelectronics Journal , 2014
2014
Earlier work this paper cites.
S. Kvatinsky et al. , “MAGIC—Memristor-Aided Logic,” IEEE TCAS II: Express Briefs , 2014
2014
Earlier work this paper cites.
S. Kvatinsky et al. , “Memristor-Based Material Implication (IMPLY) Logic: Design Principles and Methodologies,” TVLSI , 2014
2014
Earlier work this paper cites.
Q. Guo et al. , “3D-Stacked Memory-Side Acceleration: Accelerator and System Design,” in WoNDP , 2014
2014
Earlier work this paper cites.
S. H. Pugsley et al. , “NDC: Analyzing the Impact of 3D-Stacked Memory+Logic Devices on MapReduce Workloads,” in ISPASS , 2014
2014
Earlier work this paper cites.
D. P. Zhang et al. , “TOP-PIM: Throughput-Oriented Programmable Processing in Memory,” in HPDC , 2014
2014
Earlier work this paper cites.
R. Balasubramonian et al. , “Near-Data Processing: Insights from a MICRO-46 Workshop,” IEEE Micro , 2014
2014
Earlier work this paper cites.
U. Kang et al. , “Co-Architecting Controllers and DRAM to Enhance DRAM Process Scaling,” in The Memory Forum , 2014
2014
Earlier work this paper cites.
Y. Kim et al. , “Flipping Bits in Memory Without Accessing Them: An Experimental Study of DRAM Disturbance Errors,” in ISCA , 2014
2014
Earlier work this paper cites.
O. Mutlu and L. Subramanian, “Research Problems and Opportunities in Memory Systems,” SUPERFRI , 2014
2014
Earlier work this paper cites.
S. Khan et al. , “The Efficacy of Error Mitigation Techniques for DRAM Retention Failures: A Comparative Experimental Study,” in SIGMETRICS , 2014
2014
Earlier work this paper cites.
K. K. Chang et al. , “Improving DRAM Performance by Parallelizing Refreshes with Accesses,” in HPCA , 2014
2014
Earlier work this paper cites.
Hybrid Memory Cube Consortium, “HMC Specification 2.0,” 2014
2014
Earlier work this paper cites.
H. Yoon et al. , “Efficient Data Mapping and Buffering Techniques for Multilevel Cell Phase-Change Memories,” ACM TACO , 2014
2014
Earlier work this paper cites.
Y. Kim and O. Mutlu, Computing Handbook: Computer Science and Software Engineering, 3rd Edition, Chapter 1: Memory Systems . CRC Press, 2014
2014
Earlier work this paper cites.
I.-J. Sung et al. , “In-place Transposition of Rectangular Matrices on Accelerators,” in PPoPP , 2014
2014
Earlier work this paper cites.
Intel Open Source, “RAPL Power Meter,” https://01.org/rapl-power-meter , 2014
2014
Earlier work this paper cites.
B. Catanzaro et al. , “A Decomposition for In-place Matrix Transposition,” in PPoPP , 2014
2014
Cited alongside, same era.
S. Hamdioui et al. , “Memristor Based Computation-in-Memory Architecture for Data-intensive Applications,” in DATE , 2015
2015
Cited alongside, same era.
L. Xie et al. , “Fast Boolean Logic Papped on Memristor Crossbar,” in ICCD , 2015
2015
Cited alongside, same era.
J. Ahn et al. , “PIM-Enabled Instructions: A Low-Overhead, Locality-Aware Processing-in-Memory Architecture,” in ISCA , 2015
2015
Cited alongside, same era.
J. Ahn et al. , “A Scalable Processing-in-Memory Accelerator for Parallel Graph Processing,” in ISCA , 2015
2015
Cited alongside, same era.
O. O. Babarinsa and S. Idreos, “JAFAR: Near-Data Processing for Databases,” in SIGMOD , 2015
G. Dai et al. , “GraphH: A Processing-in-Memory Architecture for Large-scale Graph Processing,” IEEE TCAD , 2018
2018
Later among the works it cites.
M. Zhang et al. , “GraphP: Reducing Communication for PIM-based Graph Processing with Efficient Data Partition,” in HPCA , 2018
2018
Later among the works it cites.
S. Lloyd and M. Gokhale, “Design Space Exploration of Near Memory Accelerators,” in MEMSYS , 2018
2018
Later among the works it cites.
S. Ghose et al. , “What Your DRAM Power Models Are Not Telling You: Lessons from a Detailed Experimental Study,” in SIGMETRICS , 2018
2018
Later among the works it cites.
J. Kim et al. , “Solar-DRAM: Reducing DRAM Access Latency by Exploiting the Variation in Local Bitlines,” in ICCD , 2018
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
A. Farmahini-Farahani et al. , “NDA: Near-DRAM acceleration architecture leveraging commodity DRAM devices and standard memory modules,” in HPCA , 2015
2015
Cited alongside, same era.
M. Gao et al. , “Practical Near-Data Processing for In-Memory Analytics Frameworks,” in PACT , 2015
2015
Cited alongside, same era.
J. H. Lee et al. , “BSSync: Processing Near Memory for Machine Learning Workloads with Bounded Staleness Consistency Models,” in PACT , 2015
2015
Cited alongside, same era.
A. Morad et al. , “GP-SIMD Processing-in-Memory,” ACM TACO , 2015
2015
Cited alongside, same era.
B. Akin et al. , “Data Reorganization in Memory Using 3D-Stacked DRAM,” in ISCA , 2015
2015
Cited alongside, same era.
S. Lloyd and M. Gokhale, “In-memory Data Rearrangement for Irregular, Data-intensive Computing,” Computer , 2015
2015
Cited alongside, same era.
J. Ambrosi et al. , “Hardware-software Co-design for an Analog-digital Accelerator for Machine Learning,” in ICRC , 2018
2018
Later among the works it cites.
UPMEM, “Introduction to UPMEM PIM. Processing-in-memory (PIM) on DRAM Accelerator (White Paper),” 2018
2018
Later among the works it cites.
Y. Zhu et al. , “Matrix Profile XI: SCRIMP++: Time Series Motif Discovery at Interactive Speeds,” in ICDM , 2018
2018
Later among the works it cites.
L. Song et al. , “GraphR: Accelerating Graph Processing using ReRAM,” in HPCA , 2018
2018
Later among the works it cites.
V. Zois et al. , “Massively Parallel Skyline Computation for Processing-in-Memory Architectures,” in PACT , 2018
2018
Later among the works it cites.
H. Shin et al. , “McDRAM: Low latency and energy-efficient matrix computations in DRAM,” IEEE TCADICS , 2018
2018
Later among the works it cites.
O. Mutlu et al. , “Processing Data Where It Makes Sense: Enabling In-Memory Computation,” MicPro , 2019
2019
Later among the works it cites.
S. Ghose et al. , “Processing-in-Memory: A Workload-Driven Perspective,” IBM JRD , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
O. Mutlu et al. , “Enabling Practical Processing in and near Memory for Data-Intensive Computing,” in DAC , 2019
2019
Later among the works it cites.
D. Fujiki et al. , “Duality Cache for Data Parallel Acceleration,” in ISCA , 2019
2019
Later among the works it cites.
S. Angizi and D. Fan, “Graphide: A Graph Processing Accelerator Leveraging In-dram-computing,” in GLSVLSI , 2019
2019
Later among the works it cites.
J. Kim et al. , “D-RaNGe: Using Commodity DRAM Devices to Generate True Random Numbers with Low Latency and High Throughput,” in HPCA , 2019
2019
Later among the works it cites.
F. Gao et al. , “ComputeDRAM: In-Memory Compute Using Off-the-Shelf DRAMs,” in MICRO , 2019
2019
Later among the works it cites.
M. F. Ali et al. , “In-Memory Low-Cost Bit-Serial Addition Using Commodity DRAM Technology,” in TCAS-I , 2019
2019
Later among the works it cites.
S. Angizi et al. , “AlignS: A Processing-in-Memory Accelerator for DNA Short Read Alignment Leveraging SOT-MRAM,” in DAC , 2019
2019
Later among the works it cites.
A. Boroumand et al. , “CoNDA: Efficient Cache Coherence Support for near-Data Accelerators,” in ISCA , 2019
2019
Later among the works it cites.
G. Singh et al. , “NAPEL: Near-memory Computing Application Performance Prediction via Ensemble Learning,” in DAC , 2019
2019
Later among the works it cites.
Y. Zhuo et al. , “GraphQ: Scalable PIM-based Graph Processing,” in MICRO , 2019
2019
Later among the works it cites.
A. Rodrigues et al. , “Towards a Scatter-Gather Architecture: Hardware and Software Issues,” in MEMSYS , 2019
2019
Later among the works it cites.
O. Mutlu and J. S. Kim, “RowHammer: A Retrospective,” IEEE TCAD , 2019
2019
Later among the works it cites.
S. Ghose et al. , “Demystifying Complex Workload-DRAM Interactions: An Experimental Study,” in SIGMETRICS , 2019
2019
Later among the works it cites.
A. Ankit et al. , “PUMA: A Programmable Ultra-Efficient Memristor-Based Accelerator for Machine Learning Inference,” in ASPLOS , 2019
2019
Later among the works it cites.
F. Devaux, “The True Processing In Memory Accelerator,” in Hot Chips , 2019
2019
Later among the works it cites.
Intel, “Intel Xeon Silver 4215 Processor,” https://ark.intel.com/content/www/us/en/ark/products/193389/intel-xeon-silver-4215-processor-11m-cache-2-50-ghz.html , 2019
2019
Later among the works it cites.
K. Kanellopoulos et al. , “SMASH: Co-designing Software Compression and Hardware-Accelerated Indexing for Efficient Sparse Matrix Operations,” in MICRO , 2019
2019
Later among the works it cites.
S. G. De Gonzalo et al. , “Automatic Generation of Warp-level Primitives and Atomic Instructions for Fast and Portable Parallel Reduction on GPUs,” in CGO , 2019
2019
Later among the works it cites.
W. Huangfu et al. , “MEDAL: Scalable DIMM Based Near Data Processing Accelerator for DNA Seeding Algorithm,” in MICRO , 2019
2019
Later among the works it cites.
X. Xin et al. , “ELP2IM: Efficient and Low Power Bitwise Operation Processing in DRAM,” in HPCA , 2020
2020
Later among the works it cites.
S. H. S. Rezaei et al. , “NoM: Network-on-Memory for Inter-Bank Data Transfer in Highly-Banked Memories,” CAL , 2020
2020
Later among the works it cites.
Y. Wang et al. , “FIGARO: Improving System Performance via Fine-Grained In-DRAM Data Relocation and Caching,” in MICRO , 2020
2020
Later among the works it cites.
I. Fernandez et al. , “NATSA: A Near-Data Processing Accelerator for Time Series Analysis,” in ICCD , 2020
2020
Later among the works it cites.
M. Alser et al. , “Accelerating Genome Analysis: A Primer on an Ongoing Journey,” IEEE Micro , 2020
2020
Later among the works it cites.
D. S. Cali et al. , “GenASM: A High-Performance, Low-Power Approximate String Matching Acceleration Framework for Genome Sequence Analysis,” in MICRO , 2020
2020
Later among the works it cites.
Y. Huang et al. , “A Heterogeneous PIM Hardware-Software Co-Design for Energy-Efficient Graph Processing,” in IPDPS , 2020
2020
Later among the works it cites.
Y. Xi et al. , “In-Memory Learning With Analog Resistive Switching Memory: A Review and Perspective,” Proceedings of the IEEE , 2020
2020
Later among the works it cites.
J. S. Kim et al. , “Revisiting RowHammer: An Experimental Analysis of Modern DRAM Devices and Mitigation Techniques,” in ISCA , 2020
2020
Later among the works it cites.
P. Frigo et al. , “TRRespass: Exploiting the Many Sides of Target Row Refresh,” in S&P , 2020
2020
Later among the works it cites.
L. Cojocar et al. , “Are We Susceptible to Rowhammer? An End-to-End Methodology for Cloud Providers,” in S&P , 2020
2020
Later among the works it cites.
P. Girard et al. , “A Survey of Test and Reliability Solutions for Magnetic Random Access Memories,” Proceedings of the IEEE , 2020
2020
Later among the works it cites.
G. Singh et al. , “NERO: A Near High-Bandwidth Memory Stencil Accelerator for Weather Prediction Modeling,” in FPL , 2020
2020
Later among the works it cites.
V. Seshadri and O. Mutlu, “In-DRAM Bulk Bitwise Execution Engine,” arXiv:1905.09822 [cs.AR], 2020
2020
Later among the works it cites.
A. Ankit et al. , “PANTHER: A Programmable Architecture for Neural Network Training Harnessing Energy-efficient ReRAM,” IEEE TC , 2020
2020
Later among the works it cites.
R. Christy et al. , “8.3 A 3GHz ARM Neoverse N1 CPU in 7nm FinFET for Infrastructure Applications,” in ISSCC , 2020
2020
Later among the works it cites.
O. Mutlu, “Lecture 18c: Fine-Grained Multithreading,” https://safari.ethz.ch/digitaltechnik/spring2020/lib/exe/fetch.php?media=onur-digitaldesign-2020-lecture18c-fgmt-beforelecture.pptx , video available at http://www.youtube.com/watch?v=bu5dxKTvQVs , 2020, Digital Design and Computer Architecture. Spring 2020
2020
Later among the works it cites.
O. Mutlu, “Lecture 19: SIMD Processors,” https://safari.ethz.ch/digitaltechnik/spring2020/lib/exe/fetch.php?media=onur-digitaldesign-2020-lecture19-simd-beforelecture.pptx , video available at http://www.youtube.com/watch?v=2XZ3ik6xSzM , 2020, Digital Design and Computer Architecture. Spring 2020
2020
Later among the works it cites.
O. Mutlu, “Lecture 2b: Data Retention and Memory Refresh,” https://bit.ly/3KhVgAq , video available at http://www.youtube.com/watch?v=v702wUnaWGE , 2020, Computer Architecture. Fall 2020
2020
Later among the works it cites.
O. Mutlu, “Lecture 3b: Memory Systems: Challenges and Opportunities,” https://bit.ly/3sQIfYu , video available at http://www.youtube.com/watch?v=Q2FbUxD7GHs , 2020, Computer Architecture. Fall 2020
2020
Later among the works it cites.
O. Mutlu, “Lecture 4a: Memory Systems: Solution Directions,” https://safari.ethz.ch/architecture/fall2020/lib/exe/fetch.php?media=onur-comparch-fall2020-lecture4a-memory-solutions-afterlecture.pptx , video available at http://www.youtube.com/watch?v=PANTCVTYe8M , 2020, Computer Architecture. Fall 2020
2020
Later among the works it cites.
UPMEM, “UPMEM Website,” https://www.upmem.com , 2020
2020
Later among the works it cites.
R. Jodin and R. Cimadomo, “UPMEM. Personal Communication,” October 2020
2020
Later among the works it cites.
O. Mutlu, “Lecture 20: Graphics Processing Units,” https://safari.ethz.ch/digitaltechnik/spring2020/lib/exe/fetch.php?media=onur-digitaldesign-2020-lecture20-gpu-beforelecture.pptx , video available at http://www.youtube.com/watch?v=dg0VN-XCGKQ , 2020, Digital Design and Computer Architecture. Spring 2020
2020
Later among the works it cites.
UPMEM, “Compiler Explorer,” https://dpu.dev , 2020
2020
Later among the works it cites.
Intel, “Intel Advisor,” 2020
2020
Later among the works it cites.
N. Hajinazar et al. , “The Virtual Block Interface: A Flexible Alternative to the Conventional Virtual Memory Framework,” in ISCA , 2020
2020
Later among the works it cites.
D. Lavenier et al. , “Variant Calling Parallelization on Processor-in-Memory Architecture,” in BIBM , 2020
2020
Later among the works it cites.
J. Nider et al. , “Processing in Storage Class Memory,” in HotStorage , 2020
2020
Later among the works it cites.
S. Cho et al. , “McDRAM v2: In-Dynamic Random Access Memory Systolic Array Accelerator to Address the Large Model Problem in Deep Neural Networks on the Edge,” IEEE Access , 2020
2020
Later among the works it cites.
N. Hajinazar et al. , “SIMDRAM: A Framework for Bit-Serial SIMD Processing Using DRAM,” in ASPLOS , 2021
2021
Closest in time.
C. Giannoula et al. , “SynCron: Efficient Synchronization Support for Near-Data-Processing Architectures,” in HPCA , 2021
2021
Closest in time.
M. Besta et al. , “SISA: Set-Centric Instruction Set Architecture for Graph Mining on Processing-in-Memory Systems,” in MICRO , 2021
2021
Closest in time.
2021
Closest in time.
A. Olgun et al. , “QUAC-TRNG: High-Throughput True Random Number Generation Using Quadruple Row Activation in Commodity DRAMs,” in ISCA , 2021
2021
Closest in time.
J. Landgraf et al. , “Combining Emulation and Simulation to Evaluate a Near Memory Key/Value Lookup Accelerator,” 2021
2021
Closest in time.
Onur Mutlu and Juan Gómez-Luna, “Exploring the Processing-in-Memory Paradigm for Future Computing Systems (Fall 2021),” https://safari.ethz.ch/projects_and_seminars/fall2021/doku.php?id=processing_in_memory
2021
Closest in time.
Onur Mutlu, “Computer Architecture (Fall 2021),” https://safari.ethz.ch/architecture/fall2021/doku.php?id=start
2021
Closest in time.
Onur Mutlu and Mohammed Alser and Juan Gómez-Luna, “Seminar in Computer Architecture (Fall 2021),” https://safari.ethz.ch/architecture_seminar/spring2022/doku.php?id=start
2021
Closest in time.
L. Yavits et al. , “GIRAF: General Purpose In-Storage Resistive Associative Framework,” IEEE TPDS , 2021
2021
Closest in time.
2021
Closest in time.
L. Ke et al. , “Near-Memory Processing in Action: Accelerating Personalized Recommendation with AxDIMM,” IEEE Micro , 2021
2021
Closest in time.
Y.-C. Kwon et al. , “25.4 A 20nm 6GB Function-In-Memory DRAM, Based on HBM2 with a 1.2 TFLOPS Programmable Computing Unit Using Bank-Level Parallelism, for Machine Learning Applications,” in ISSCC , 2021
2021
Closest in time.
S. Lee et al. , “Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology: Industrial Product,” in ISCA , 2021
2021
Closest in time.
B. Asgari et al. , “FAFNIR: Accelerating Sparse Gathering by Using Efficient Near-Memory Intelligent Reduction,” in HPCA , 2021
2021
Closest in time.
J. M. Herruzo et al. , “Enabling Fast and Energy-Efficient FM-Index Exact Matching Using Processing-Near-Memory,” The Journal of Supercomputing , 2021
2021
Closest in time.
G. Singh et al. , “Fpga-based Near-memory Acceleration of Modern Data-intensive Applications,” IEEE Micro , 2021
2021
Closest in time.
G. Singh et al. , “Accelerating Weather Prediction using Near-Memory Reconfigurable Fabric,” ACM TRETS , 2021
2021
Closest in time.
G. F. Oliveira et al. , “DAMOV: A New Methodology and Benchmark Suite for Evaluating Data Movement Bottlenecks,” IEEE Access , 2021
2021
Closest in time.
A. Boroumand et al. , “Google Neural Network Models for Edge Devices: Analyzing and Mitigating Machine Learning Inference Bottlenecks,” in PACT , 2021
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
A. G. Yağlikçi et al. , “BlockHammer: Preventing RowHammer at Low Cost by Blacklisting Rapidly-Accessed DRAM Rows,” in HPCA , 2021
2021
Closest in time.
L. Orosa et al. , “A Deeper Look into RowHammer’s Sensitivities: Experimental Analysis of Real DRAM Chips and Implications on Future Attacks and Defenses,” in MICRO , 2021
2021
Closest in time.
H. Hassan et al. , “Uncovering In-DRAM RowHammer Protection Mechanisms: A New Methodology, Custom RowHammer Patterns, and Implications,” in MICRO , 2021
2021
Closest in time.
M. Patel et al. , “Harp: Practically and effectively identifying uncorrectable errors in memory chips that use on-die error-correcting codes,” in MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture , 2021, pp. 623–640
2021
Closest in time.
2021
Closest in time.
S. Huang et al. , “Mixed Precision Quantization for ReRAM-based DNN Inference Accelerators,” in ASP-DAC , 2021
2021
Closest in time.
UPMEM, “UPMEM User Manual. Version 2021.1.0,” 2021
2021
Closest in time.
UPMEM, “UPMEM Software Development Kit (SDK).” https://sdk.upmem.com , 2021
2021
Closest in time.
LLVM, “Compiler-RT, LLVM project,” https://github.com/llvm/llvm-project/tree/main/compiler-rt/lib/builtins , 2021
2021
Closest in time.
NVIDIA, “CUDA Samples v. 11.2,” 2021
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
J. Gómez-Luna et al. , “Benchmarking Memory-Centric Computing Systems: Analysis of Real Processing-in-Memory Hardware,” in IGSC , 2021
2021
Closest in time.
S. Lee et al. , “A 1ynm 1.25V 8Gb, 16Gb/s/pin GDDR6-based Accelerator-in-Memory supporting 1TFLOPS MAC Operation and Various Activation Functions for Deep-Learning Applications,” in ISSCC , 2022
2022
Closest in time.