Fetching the paper…
Reading the bibliography…
Personalized recommendation systems leverage deep learning models and account for the majority of data center AI cycles.
1906
Earlier work this paper cites.
1906
Earlier work this paper cites.
M.Gorman, “Understanding the Linux virtual memory manager.” Prentice Hall Upper Saddle River , 2004
2004
Earlier work this paper cites.
Onur Mutlu, Thomas Moscibroda, “Stall-Time Fair Memory Access Scheduling for Chip Multiprocessors,” MICRO , pp. 146–160, 2007
2007
Earlier work this paper cites.
N. Muralimanohar, R. Balasubramonian, and N. P. Jouppi, “Cacti 6.0: A tool to model large caches,” HP laboratories , pp. 22–31, 2009
2009
Earlier work this paper cites.
Samuel Williams, Andrew Waterman, and David Patterson, “Roofline: An Insightful Visual Performance Model for Floating-Point Programs and Multicore Architectures,” Communications of the ACM , 2009
2009
Earlier work this paper cites.
Xiao Zhang, Sandhya Dwarkadas, Kai Shen, “Towards practical page coloring-based multicore cache management,” EuroSys , pp. 89–102, 2009
2009
Earlier work this paper cites.
J. E. Stone, D. Gohara, and G. Shi, “Opencl: A parallel programming standard for heterogeneous computing systems,” IEEE Computing in Science and Engineering , vol. 12, no. 3, 2010
2010
Earlier work this paper cites.
J. Ngiam, Z. Chen, D. Chia, P. W. Koh, Q. V. Le, and A. Y. Ng, “Tiled convolutional neural networks,” NIPS , pp. 1279–1287, 2010
2010
Earlier work this paper cites.
Jaekyu Lee, Hyesoon Kim, and Richard Vuduc, “When Prefetching Works, When It Doesn’t, and Why,” ACM TACO , vol. 9, no. 1, 2012
2012
Earlier work this paper cites.
K. Chen, S. Li, N. Muralimanohar, J. H. Ahn, J. B. Brockman, and N. P. Jouppi, “Cacti-3dd: Architecture-level modeling for 3d die-stacked dram main memory,” VLSI , pp. 33–38, 2012
2012
Earlier work this paper cites.
Q. Guo, N. Alachiotis, B. Akin, F. Sadi, G. Xu, T. M. Low, L. Pileggi, J. C. Hoe, and F. Franchetti, “3d-stacked memory-side acceleration: Accelerator and system design,” WoNDP , 2014
2014
Earlier work this paper cites.
J. Ahn, S. Hong, S. Yoo, O. Mutlu, and K. Choi, “A scalable processing-in-memory accelerator for parallel graph processing,” ISCA , pp. 105–117, 2015
2015
Earlier work this paper cites.
Amin Farmahini-Farahani, Jung Ho Ahn, Katherine Morrow, and Nam Sung Kim, “NDA: Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules,” HPCA , 2015
2015
Earlier work this paper cites.
M. Gao, G. Ayers, and C. Kozyrakis, “Practical near-data processing for in-memory analytics frameworks,” PACT , pp. 113–124, 2015
2015
Earlier work this paper cites.
Junwhan Ahn, Sungjoo Yoo, Onur Mutlu, Kiyoung Choi, “Pim-enabled instructions: A low-overhead, locality-aware processing-in-memory architecture,” ISCA , pp. 336–348, 2015
2015
Earlier work this paper cites.
Y. Kim, W. Yang, and O. Mutlu, “Ramulator: A fast and extensible dram simulator,” IEEE Computer architecture letters , vol. 15, no. 1, pp. 45–49, 2015
2015
Earlier work this paper cites.
N. P. Jouppi, A. B. Kahng, N. Muralimanohar, and V. Srinivas, “Cacti-io: Cacti with off-chip power-area-timing models,” VLSI , pp. 1254–1267, 2015
2015
Earlier work this paper cites.
P. Meaney, L. Curley, G. Gilda, M. Hodges, D. Buerkle, R. Siegl, and R. Dong, “The IBM z13 Memory Subsystem for Big Data,” IBM Journal of Research and Development , 2015
2015
Earlier work this paper cites.
R. Adolf, S. Rama, B. Reagen, G.-Y. Wei, and D. Brooks, “Fathom: Reference workloads for modern deep learning methods,” in Proceedings of the IEEE International Symposium on Workload Characterization (IISWC) . IEEE, 2016, pp. 1–10
2016
Earlier work this paper cites.
J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, “Cnvlutin: Ineffectual-neuron-free deep neural network computing,” in Proc. of the Intl. Symp. on Computer Architecture , 2016
2016
Earlier work this paper cites.
H. Cheng, L. Koc, J. Harmsen, T. Shaked, T. Chandra, H. Aradhye, G. Anderson, G. Corrado, W. Chai, M. Ispir, R. Anil, Z. Haque, L. Hong, V. Jain, X. Liu, and H. Shah, “Wide & deep learning for recommender systems,” in Proceedings of the 1st Workshop on Deep Learning for Recommender Systems, DLRS RecSys 2016, Boston, MA, USA, September 15, 2016 , 2016, pp. 7–10
2016
Earlier work this paper cites.
H.-T. Cheng, L. Koc, J. Harmsen, T. Shaked, T. Chandra, H. Aradhye, G. Anderson, G. Corrado, W. Chai, M. Ispir et al. , “Wide & deep learning for recommender systems,” in Proceedings of the 1st Workshop on Deep Learning for Recommender Systems . ACM, 2016, pp. 7–10
2016
Earlier work this paper cites.
P. Covington, J. Adams, and E. Sargin, “Deep neural networks for youtube recommendations,” in Proceedings of the 10th ACM Conference on Recommender Systems , ser. RecSys ’16. New York, NY, USA: ACM, 2016, pp. 191–198. [Online]. Available: http://doi.acm.org/10.1145/2959100.2959190
2016
Cited alongside, same era.
P. Covington, J. Adams, and E. Sargin, “Deep neural networks for youtube recommendations,” in Proceedings of the 10th ACM conference on recommender systems . ACM, 2016, pp. 191–198
2016
Cited alongside, same era.
Hadi Asghari-Moghaddam, Young Hoon Son, Jung Ho Ahn, Nam Sung Kim, “Chameleon: Versatile and Practical Near-DRAM Acceleration Architecture for Large Memory Systems,” MICRO , 2016
2016
Cited alongside, same era.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “Eie: efficient inference engine on compressed deep neural network,” in Proceedings of the ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2016, pp. 243–254
2018
Later among the works it cites.
B. Feinberg, S. Wang, and E. Ipek, “Making memristive neural network accelerators reliable,” in Proc. of the Intl. Symp. on High Performance Computer Architecture , 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
K. Hazelwood, S. Bird, D. Brooks, S. Chintala, U. Diril, D. Dzhulgakov, M. Fawzy, B. Jia, Y. Jia, A. Kalro et al. , “Applied machine learning at Facebook: a datacenter infrastructure perspective,” in Proceedings of the IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2018, pp. 620–629
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
B. Hong, G. Kim, J. H. Ahn, Y. Kwon, H. Kim, and J. Kim, “Accelerating linked-list traversal through near-data processing,” PACT , pp. 113–124, 2016
2016
Cited alongside, same era.
K. Hsieh, E. Ebrahim, G. Kim, N. Chatterjee, M. O’Connor, N. Vijaykumar, O. Mutlu, and S. W. Keckler, “Transparent offloading and mapping (tom): Enabling programmer-transparent near-data processing in gpu systems,” ISCA , pp. 204–216, 2016
2016
Cited alongside, same era.
K. Hsieh, S. Khan, N. Vijaykumar, K. K. Chang, A. Boroumand, S. Ghose, and O. Mutlu, “Accelerating pointer chasing in 3d-stacked memory: Challenges, mechanisms, evaluation,” ICCD , pp. 25–32, 2016
2016
Cited alongside, same era.
D. Kim, J. Kung, S. Chai, S. Yalamanchili, and S. Mukhopadhyay, “Neurocube: A programmable digital neuromorphic architecture with high-density 3D memory,” in Proc. of the Intl. Symp. on Computer Architecture , 2016
2016
Cited alongside, same era.
S. Liu, Z. Du, J. Tao, D. Han, T. Luo, Y. Xie, Y. Chen, and T. Chen, “Cambricon: An instruction set architecture for neural networks,” in Proc. of the Intl. Symp. on Computer Architecture , 2016
2016
Cited alongside, same era.
P. Pessl, D. Gruss, C. Maurice, M. Schwarz, and S. Mangard, “Drama: Exploiting dram addressing for cross-cpu attack,” USENIX Security Symposium , vol. pp.565-581, 2016
2016
Cited alongside, same era.
B. Reagen, P. Whatmough, R. Adolf, S. Rama, H. Lee, S. K. Lee, J. M. Hernández-Lobato, G.-Y. Wei, and D. Brooks, “Minerva: Enabling low-power, highly-accurate deep neural network accelerators,” in Proceedings of the ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2016, pp. 267–278
2016
Cited alongside, same era.
A. Shafiee, A. Nag, N. Muralimanohar, R. Balasubramonian, J. P. Strachan, M. Hu, R. S. Williams, and V. Srikumar, “ISAAC: a convolutional neural network accelerator with in-situ analog arithmetic in crossbars,” in Proc. of the Intl. Symp. on Computer Architecture , 2016, pp. 14–26
2016
Cited alongside, same era.
2018
Later among the works it cites.
K. Hegde, J. Yu, R. Agrawal, M. Yan, M. Pellauer, and C. W. Fletcher, “UCNN: Exploiting computational reuse in deep neural networks via weight repetition,” Proc. of the Intl. Symp. on Computer Architecture , 2018
2018
Later among the works it cites.
A. Jain, A. Phanishayee, J. Mars, L. Tang, and G. Pekhimenko, “Gist: Efficient data encoding for deep neural network training,” Proc. of the Intl. Symp. on Computer Architecture , 2018
2018
Later among the works it cites.
Jiawen Liu, Hengyu Zhao, Matheus Almeida Ogleari, Dong Li, Jishen Zhao, “Processing-in-memory for energy-efficient neural network training: A heterogeneous approach,” MICRO , pp. 655–668, 2018
2018
Later among the works it cites.
2018
Later among the works it cites.
2018
Later among the works it cites.
E. Park, D. Kim, and S. Yoo, “Energy-efficient neural network accelerator based on outlier-aware low-precision computation,” Proc. of the Intl. Symp. on Computer Architecture , 2018
2018
Later among the works it cites.
M. Rhu, M. O’Connor, N. Chatterjee, J. Pool, Y. Kwon, and S. W. Keckler, “Compressing DMA engine: Leveraging activation sparsity for training deep neural networks,” in Proc. of the Intl. Symp. on High Performance Computer Architecture , 2018
2018
Later among the works it cites.
M. Riera, J. M. Arnau, and A. Gonzalez, “Computation reuse in DNNs by exploiting input similarity,” Proc. of the Intl. Symp. on Computer Architecture , 2018
2018
Later among the works it cites.
H. Sharma, J. Park, N. Suda, L. Lai, B. Chau, V. Chandra, and H. Esmaeilzadeh, “Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural network,” Proc. of the Intl. Symp. on Computer Architecture , 2018
2018
Later among the works it cites.
M. Song, J. Zhang, H. Chen, and T. Li, “Towards efficient microarchitectural design for accelerating unsupervised GAN-based deep learning,” in Proc. of the Intl. Symp. on High Performance Computer Architecture , 2018
2018
Later among the works it cites.
M. Song, K. Zhong, J. Zhang, Y. Hu, D. Liu, W. Zhang, J. Wang, and T. Li, “In-situ AI: towards autonomous and incremental deep learning for IoT systems,” in Proc. of the Intl. Symp. on High Performance Computer Architecture , 2018
2018
Later among the works it cites.
M. Song, J. Zhao, Y. Hu, J. Zhang, and T. Li, “Prediction based execution on deep neural networks,” Proc. of the Intl. Symp. on Computer Architecture , 2018
2018
Later among the works it cites.
A. Yazdanbakhsh, K. Samadi, H. Esmaeilzadeh, and N. S. Kim, “GANAX: a unified SIMD-MIMD acceleration for generative adversarial network,” in Proc. of the Intl. Symp. on Computer Architecture , 2018
2018
Later among the works it cites.
R. Yazdani, M. Riera, J.-M. Arnau, and A. Gonzalez, “The dark side of DNN pruning,” Proc. of the Intl. Symp. on Computer Architecture , 2018
2018
Later among the works it cites.
G. Zhou, X. Zhu, C. Song, Y. Fan, H. Zhu, X. Ma, Y. Yan, J. Jin, H. Li, and K. Gai, “Deep interest network for click-through rate prediction,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2018, London, UK, August 19-23, 2018 , 2018, pp. 1059–1068
2018
Later among the works it cites.
Fortune, https://fortune.com/2019/04/30/artificial-intelligence-walmart-stores/
2019
Closest in time.
Hanhwi Jang, Joonsung Kim, Jae-Eon Jo, Jaewon Lee, Jangwoo Kim, “MnnFast: a fast and scalable system architecture for memory-augmented neural networks,” ISCA , 2019
2019
Closest in time.
Y. L. Youngeun Kwon and M. Rhu, “Tensordimm: A practical near-memory processing architecture for embeddings and tensor operations in deep learning,” MICRO , pp. 740–753, 2019
2019
Closest in time.
Z. Zhao, L. Hong, L. Wei, J. Chen, A. Nath, S. Andrews, A. Kumthekar, M. Sathiamoorthy, X. Yi, and E. Chi, “Recommending what video to watch next: A multitask ranking system,” in Proceedings of the 13th ACM Conference on Recommender Systems , ser. RecSys ’19. New York, NY, USA: ACM, 2019, pp. 43–51. [Online]. Available: http://doi.acm.org/10.1145/3298689.3346997
2019
Closest in time.