Fetching the paper…
Reading the bibliography…
As Deep Learning continues to drive a variety of applications in edge and cloud data centers, there is a growing trend towards building large accelerators with several sub-accelerator cores/chiplets.
K. Pearson, “Liii. on lines and planes of closest fit to systems of points in space,” The London, Edinburgh, and Dublin philosophical magazine and journal of science , vol. 2, no. 11, pp. 559–572, 1901
1901
Earlier work this paper cites.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International conference on machine learning , 2016, pp. 1928–1937
1937
Earlier work this paper cites.
J. Kennedy and R. Eberhart, “Particle swarm optimization,” in Proceedings of ICNN’95-International Conference on Neural Networks , vol. 4. IEEE, 1995, pp. 1942–1948
1948
Earlier work this paper cites.
J. H. Holland, “Genetic algorithms,” Scientific american , vol. 267, no. 1, pp. 66–73, 1992
1992
Earlier work this paper cites.
E. S. Hou, N. Ansari, and H. Ren, “A genetic algorithm for multiprocessor scheduling,” IEEE Transactions on Parallel and Distributed systems , vol. 5, no. 2, pp. 113–120, 1994
1994
Earlier work this paper cites.
I. Rechenberg, “Evolutionsstrategie: Optimierung technischer systeme nach prinzipien der biologischen evolution. frommann-holzbog, stuttgart, 1973,” Step-Size Adaptation Based on Non-Local Use of Selection Information. In PPSN3 , 1994
1994
Earlier work this paper cites.
P. Shroff, D. W. Watson, N. S. Flann, and R. F. Freund, “Genetic simulated annealing for scheduling data-dependent tasks in heterogeneous environments,” in 5th Heterogeneous Computing Workshop (HCW’96) , 1996, pp. 98–117
1996
Earlier work this paper cites.
H. Singh and A. Youssef, “Mapping and scheduling heterogeneous task graphs using genetic algorithms,” in 5th IEEE heterogeneous computing workshop (HCW’96) , 1996, pp. 86–97
1996
Earlier work this paper cites.
L. Wang, H. J. Siegel, and V. P. Roychowdhury, “A genetic-algorithm-based approach for task matching and scheduling in heterogeneous computing environments,” in Proc. Heterogeneous Computing Workshop , 1996, pp. 72–85
1996
Earlier work this paper cites.
R. C. Corrêa, A. Ferreira, and P. Rebreyend, “Scheduling multiprocessor tasks with genetic algorithms,” IEEE Transactions on Parallel and Distributed systems , vol. 10, no. 8, pp. 825–837, 1999
1999
Earlier work this paper cites.
J. Sherwani, N. Ali, N. Lotia, Z. Hayat, and R. Buyya, “Libra: a computational economy-based job scheduling system for clusters,” Software: Practice and Experience , vol. 34, no. 6, pp. 573–590, 2004
2004
Earlier work this paper cites.
N. Hansen, “The cma evolution strategy: a comparing review,” in Towards a new evolutionary computation . Springer, 2006, pp. 75–102
2006
Earlier work this paper cites.
M. Zaharia, D. Borthakur, J. S. Sarma, K. Elmeleegy, S. Shenker, and I. Stoica, “Job scheduling for multi-user mapreduce clusters,” Technical Report UCB/EECS-2009-55, EECS Department, University of California …, Tech. Rep., 2009
2009
Earlier work this paper cites.
P. Bajpai and M. Kumar, “Genetic algorithm–an approach to solve global optimization problems,” Indian Journal of computer science and engineering , vol. 1, no. 3, pp. 199–206, 2010
2010
Earlier work this paper cites.
Z. Wu, Z. Ni, L. Gu, and X. Liu, “A revised discrete particle swarm optimization for cloud workflow scheduling,” in 2010 International Conference on Computational Intelligence and Security . IEEE, 2010, pp. 184–188
2010
Earlier work this paper cites.
T. Beisel, T. Wiersema, C. Plessl, and A. Brinkmann, “Cooperative multitasking for heterogeneous accelerators in the linux completely fair scheduler,” in ASAP 2011-22nd IEEE International Conference on Application-specific Systems, Architectures and Processors . IEEE, 2011, pp. 223–226
2011
Earlier work this paper cites.
C. Li, S. Yang, and T. T. Nguyen, “A self-learning particle swarm optimizer for global optimization problems,” IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics) , vol. 42, no. 3, pp. 627–646, 2011
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in neural information processing systems , 2012, pp. 1097–1105
2012
Earlier work this paper cites.
J. Bergstra, D. Yamins, and D. D. Cox, “Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures,” 2013
2013
Earlier work this paper cites.
K. V. Price, “Differential evolution,” in Handbook of Optimization . Springer, 2013, pp. 187–214
2013
Earlier work this paper cites.
E. Schkufza, R. Sharma, and A. Aiken, “Stochastic superoptimization,” ACM SIGARCH Computer Architecture News , vol. 41, no. 1, pp. 305–316, 2013
2013
Earlier work this paper cites.
W. Joo and D. Shin, “Resource-constrained spatial multi-tasking for embedded gpu,” in 2014 IEEE International Conference on Consumer Electronics (ICCE) . IEEE, 2014, pp. 339–340
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2015, pp. 1–9
2015
Earlier work this paper cites.
C. Zhang, P. Li, G. Sun, Y. Guan, B. Xiao, and J. Cong, “Optimizing fpga-based accelerator design for deep convolutional neural networks,” in Proceedings of the 2015 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays , 2015, pp. 161–170
2015
Earlier work this paper cites.
Y.-H. Chen, J. Emer, and V. Sze, “Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks,” in International Symposium on Computer Architecture (ISCA) , 2016
2016
Earlier work this paper cites.
Y.-H. Chen, T. Krishna, J. S. Emer, and V. Sze, “Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks,” IEEE journal of solid-state circuits , vol. 52, no. 1, pp. 127–138, 2016
2016
Earlier work this paper cites.
H.-T. Cheng, L. Koc, J. Harmsen, T. Shaked, T. Chandra, H. Aradhye, G. Anderson, G. Corrado, W. Chai, M. Ispir et al. , “Wide & deep learning for recommender systems,” in Proceedings of the 1st workshop on deep learning for recommender systems , 2016, pp. 7–10
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
M. Hellwig and H.-G. Beyer, “Evolution under strong noise: A self-adaptive evolution strategy can reach the lower performance bound-the pccmsa-es,” in International Conference on Parallel Problem Solving from Nature . Springer, 2016, pp. 26–36
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
H. Mao, M. Alizadeh, I. Menache, and S. Kandula, “Resource management with deep reinforcement learning,” in Proceedings of the 15th ACM Workshop on Hot Topics in Networks , 2016, pp. 50–56
2016
Earlier work this paper cites.
N. Suda, V. Chandra, G. Dasika, A. Mohanty, Y. Ma, S. Vrudhula, J.-s. Seo, and Y. Cao, “Throughput-optimized opencl-based fpga accelerator for large-scale convolutional neural networks,” in Proceedings of the 2016 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays , 2016, pp. 16–25
2016
Earlier work this paper cites.
S. Bang, J. Wang, Z. Li, C. Gao, Y. Kim, Q. Dong, Y.-P. Chen, L. Fick, X. Sun, R. Dreslinski et al. , “14.7 a 288 μ \mu w programmable deep-learning processor with 270kb on-chip weight storage using non-uniform memory hierarchy for mobile intelligence,” in 2017 IEEE International Solid-State Circuits Conference (ISSCC) . IEEE, 2017, pp. 250–251
2017
Earlier work this paper cites.
Q. Chen, H. Yang, M. Guo, R. S. Kannan, J. Mars, and L. Tang, “Prophet: Precise qos prediction on non-preemptive accelerators to improve utilization in warehouse-scale computers,” in Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems , 2017, pp. 17–32
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
G. Emadi, A. M. Rahmani, and H. Shahhoseini, “Task scheduling algorithm using covariance matrix adaptation evolution strategy (cma-es) in cloud computing,” Journal of Advances in Computer Engineering and Technology , vol. 3, no. 3, pp. 135–144, 2017
2017
Earlier work this paper cites.
X. He, L. Liao, H. Zhang, L. Nie, X. Hu, and T.-S. Chua, “Neural collaborative filtering,” in Proceedings of the 26th international conference on world wide web , 2017, pp. 173–182
2017
Cited alongside, same era.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers et al. , “In-datacenter performance analysis of a tensor processing unit,” in Proceedings of the 44th Annual International Symposium on Computer Architecture , 2017, pp. 1–12
2017
Cited alongside, same era.
H. Jun, J. Cho, K. Lee, H.-Y. Son, K. Kim, H. Jin, and K. Kim, “Hbm (high bandwidth memory) dram technology and architecture,” in 2017 IEEE International Memory Workshop (IMW) . IEEE, 2017, pp. 1–4
2017
Cited alongside, same era.
W. Lu, G. Yan, J. Li, S. Gong, Y. Han, and X. Li, “Flexflow: A flexible dataflow accelerator architecture for convolutional neural networks,” in International Symposium on High Performance Computer Architecture (HPCA) , 2017
2017
2019
Later among the works it cites.
A. Parashar, P. Raina, Y. S. Shao, Y.-H. Chen, V. A. Ying, A. Mukkara, R. Venkatesan, B. Khailany, S. W. Keckler, and J. Emer, “Timeloop: A systematic approach to dnn accelerator evaluation,” in 2019 IEEE international symposium on performance analysis of systems and software (ISPASS) . IEEE, 2019, pp. 304–315
2019
Later among the works it cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” OpenAI Blog , vol. 1, no. 8, p. 9, 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Y. Shen, M. Ferdman, and P. Milder, “Maximizing cnn accelerator efficiency through resource partitioning,” in 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2017, pp. 535–547
2017
Cited alongside, same era.
2017
Cited alongside, same era.
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi, “Inception-v4, inception-resnet and the impact of residual connections on learning,” in Thirty-first AAAI conference on artificial intelligence , 2017
2017
Cited alongside, same era.
L. Thamsen, B. Rabier, F. Schmidt, T. Renner, and O. Kao, “Scheduling recurring distributed dataflow jobs based on resource utilization and interference,” in 2017 IEEE International Congress on Big Data (BigData Congress) . IEEE, 2017, pp. 145–152
2017
Cited alongside, same era.
X. Wei et al. , “Automated systolic array architecture synthesis for high throughput cnn inference on fpgas,” in DAC , 2017, pp. 1–6
2017
Cited alongside, same era.
S. Chang, J. Yang, J. Choi, and N. Kwak, “Genetic-gated networks for deep reinforcement learning.” in NeurIPS , 2018, pp. 1754–1763
2018
Cited alongside, same era.
2019
Later among the works it cites.
Y. S. Shao et al. , “Simba: Scaling deep-learning inference with multi-chip-module-based architecture,” in MICRO , 2019, pp. 14–27
2019
Later among the works it cites.
H. Shen, L. Chen, Y. Jin, L. Zhao, B. Kong, M. Philipose, A. Krishnamurthy, and R. Sundaram, “Nexus: a gpu cluster engine for accelerating dnn-based video analysis,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles , 2019, pp. 322–337
2019
Later among the works it cites.
2019
Later among the works it cites.
M. Tan, B. Chen, R. Pang, V. Vasudevan, M. Sandler, A. Howard, and Q. V. Le, “Mnasnet: Platform-aware neural architecture search for mobile,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 2820–2828
2019
Later among the works it cites.
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le, “Xlnet: Generalized autoregressive pretraining for language understanding,” in Advances in neural information processing systems , 2019, pp. 5753–5763
2019
Later among the works it cites.
S. Zheng, P. Ouyang, D. Song, X. Li, L. Liu, S. Wei, and S. Yin, “An ultra-low power binarized convolutional neural network-based speech recognition processor with on-chip self-learning,” IEEE Transactions on Circuits and Systems I: Regular Papers , vol. 66, no. 12, pp. 4648–4661, 2019
2019
Later among the works it cites.
G. Zhou, N. Mou, Y. Fan, Q. Pi, W. Bian, C. Zhou, X. Zhu, and K. Gai, “Deep interest evolution network for click-through rate prediction,” in Proceedings of the AAAI conference on artificial intelligence , vol. 33, 2019, pp. 5941–5948
2019
Later among the works it cites.
2019
Later among the works it cites.
“Maestro tool,” http://maestro.ece.gatech.edu/ , 2020
2020
Later among the works it cites.
E. Baek, D. Kwon, and J. Kim, “A multi-neural network acceleration architecture,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2020, pp. 940–953
2020
Later among the works it cites.
Cerebras, “Cerebras cs-1,” 2020. [Online]. Available: https://www.cerebras.net/
2020
Later among the works it cites.
Y. Choi and M. Rhu, “Prema: A predictive multi-task scheduling algorithm for preemptible neural processing units,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2020, pp. 220–233
2020
Later among the works it cites.
2020
Later among the works it cites.
H. Du, S. Leng, F. Wu, X. Chen, and S. Mao, “A new vehicular fog computing architecture for cooperative sensing of autonomous driving,” IEEE Access , vol. 8, pp. 10 997–11 006, 2020
2020
Later among the works it cites.
F. Fu, Y. Kang, Z. Zhang, F. R. Yu, and T. Wu, “Soft actor-critic drl for live transcoding and streaming in vehicular fog computing-enabled iov,” IEEE Internet of Things Journal , 2020
2020
Later among the works it cites.
Google, “Cloud tpu,” 2020. [Online]. Available: https://cloud.google.com/tpu/
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
N. P. Jouppi, D. H. Yoon, G. Kurian, S. Li, N. Patil, J. Laudon, C. Young, and D. Patterson, “A domain-specific supercomputer for training deep neural networks,” Commun. ACM , vol. 63, no. 7, p. 67–78, Jun. 2020. [Online]. Available: https://doi.org/10.1145/3360307
2020
Later among the works it cites.
S.-C. Kao and T. Krishna, “Confuciux: Autonomous hardware resource assignment for dnn accelerators using reinforcement learning,” in 2020 53th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2020, pp. 1–14
2020
Later among the works it cites.
S.-C. Kao and T. Krishna, “Gamma: Automating the hw mapping of dnn models on accelerators via genetic algorithm,” in 2020 IEEE/ACM International Conference On Computer Aided Design (ICCAD) . IEEE, 2020, pp. 1–9
2020
Later among the works it cites.
2020
Later among the works it cites.
X. Li, Y. Qin, H. Zhou, D. Chen, S. Yang, and Z. Zhang, “An intelligent adaptive algorithm for servers balancing and tasks scheduling over mobile fog computing networks,” Wireless Communications and Mobile Computing , vol. 2020, 2020
2020
Later among the works it cites.
Micron, “Ddr5 sdram,” 2020. [Online]. Available: https://www.micron.com/products/dram/ddr5-sdram
2020
Later among the works it cites.
NVIDIA, “Nvidia volta, tensor core gpu architecture,” 2020. [Online]. Available: https://www.nvidia.com/en-us/data-center/volta-gpu-architecture/
2020
Later among the works it cites.
PCI-SIG, “Pce spec,” 2020. [Online]. Available: https://pcisig.com/newsroom
2020
Later among the works it cites.
D. Richins, D. Doshi, M. Blackmore, A. T. Nair, N. Pathapati, A. Patel, B. Daguman, D. Dobrijalowski, R. Illikkal, K. Long et al. , “Missing the forest for the trees: End-to-end ai application performance in edge data centers,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2020, pp. 515–528
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
Transcend, “The tranfer rate of ddr1, ddr2, ddr3, and ddr4,” 2020. [Online]. Available: https://www.transcend-info.com/Support/FAQ-292
2020
Later among the works it cites.
2021
Closest in time.