Fetching the paper…
Reading the bibliography…
Transformer-based language models such as BERT provide significant accuracy improvement for a multitude of natural language processing (NLP) tasks.
2005
Earlier work this paper cites.
X. Dong, C. Xu, Y. Xie, and N. P. Jouppi, “Nvsim: A circuit-level performance, energy, and area model for emerging nonvolatile memory,” IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems , vol. 31, no. 7, pp. 994–1007, 2012
2012
Earlier work this paper cites.
G. F. Close, U. Frey, J. Morrish, R. Jordan, S. C. Lewis, T. Maffitt, M. J. BrightSky, C. Hagleitner, C. H. Lam, and E. Eleftheriou, “A 256-mcell phase-change memory chip operating at 2 + 2{+} bit/cell,” IEEE Transactions on Circuits and Systems I: Regular Papers , vol. 60, no. 6, pp. 1521–1533, 2013
2013
Earlier work this paper cites.
Cong Xu, Dimin Niu, N. Muralimanohar, N. P. Jouppi, and Yuan Xie, “Understanding the trade-offs in multi-level cell reram memory design,” in 2013 50th ACM/EDAC/IEEE Design Automation Conference (DAC) , 2013, pp. 1–6
2013
Earlier work this paper cites.
T. Liu, T. H. Yan, R. Scheuerlein, Y. Chen, J. K. Lee, G. Balakrishnan, G. Yee, H. Zhang, A. Yap, J. Ouyang, T. Sasaki, S. Addepalli, A. Al-Shamma, C. Chen, M. Gupta, G. Hilton, S. Joshi, A. Kathuria, V. Lai, D. Masiwal, M. Matsumoto, A. Nigam, A. Pai, J. Pakhale, C. H. Siau, X. Wu, R. Yin, L. Peng, J. Y. Kang, S. Huynh, H. Wang, N. Nagel, Y. Tanaka, M. Higashitani, T. Minvielle, C. Gorla, T. Tsukamoto, T. Yamaguchi, M. Okajima, T. Okamura, S. Takase, T. Hara, H. Inoue, L. Fasoli, M. Mofidi, R. Shrivastava, and K. Quader, “A 130.7mm2 2-layer 32gb reram memory device in 24nm technology,” in 2013 IEEE International Solid-State Circuits Conference Digest of Technical Papers , 2013
2013
Earlier work this paper cites.
M. Chang, J. Wu, T. Chien, Y. Liu, T. Yang, W. Shen, Y. King, C. Lin, K. Lin, Y. Chih, S. Natarajan, and J. Chang, “19.4 embedded 1mb reram in 28nm cmos with 0.27-to-1v read using swing-sample-and-couple sense amplifier and self-boost-write-termination scheme,” in 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC) , 2014
2014
Earlier work this paper cites.
T. Chen, Z. Du, N. Sun, J. Wang, C. Wu, Y. Chen, and O. Temam, “Diannao: A small-footprint high-throughput accelerator for ubiquitous machine-learning,” in Proceedings of the 19th International Conference on Architectural Support for Programming Languages and Operating Systems , ser. ASPLOS ’14. New York, NY, USA: ACM, 2014, pp. 269–284
2014
Earlier work this paper cites.
Z. Toprak-Deniz, M. Sperling, J. Bulzacchelli, G. Still, R. Kruse, S. Kim, D. Boerstler, T. Gloekler, R. Robertazzi, K. Stawiasz, T. Diemoz, G. English, D. Hui, P. Muench, and J. Friedrich, “5.2 distributed system of digitally controlled microregulators enabling per-core dvfs for the power8 tm microprocessor,” in 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC) , 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
D. Liu, T. Chen, S. Liu, J. Zhou, S. Zhou, O. Teman, X. Feng, X. Zhou, and Y. Chen, “Pudiannao: A polyvalent machine learning accelerator,” SIGPLAN Not. , vol. 50, no. 4, p. 369–381, Mar. 2015. [Online]. Available: https://doi.org/10.1145/2775054.2694358
2015
Earlier work this paper cites.
PubNub. (2015) How fast is real-time? human perception and technology. [Online]. Available: https://www.pubnub.com/blog/how-fast-is-realtime-human-perception-and-technology/
2015
Earlier work this paper cites.
J. Albericio, P. Judd, T. Hetherington, T. Aamodt, N. E. Jerger, and A. Moshovos, “Cnvlutin: Ineffectual-neuron-free deep neural network computing,” in Proceedings of the 43rd International Symposium on Computer Architecture , 2016
2016
Earlier work this paper cites.
L. J. Ba et al. , “Layer normalization,” ArXiv , vol. abs/1607.06450, 2016
2016
Earlier work this paper cites.
Y. Chen, J. Emer, and V. Sze, “Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks,” in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) , June 2016, pp. 367–379
2016
Earlier work this paper cites.
R. Eisele. (2016) The log-sum-exp trick in machine learning. [Online]. Available: https://www.xarg.org/2016/06/the-log-sum-exp-trick-in-machine-learning/
2016
Earlier work this paper cites.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “Eie: Efficient inference engine on compressed deep neural network,” SIGARCH Comput. Archit. News , vol. 44, no. 3, Jun. 2016
2016
Earlier work this paper cites.
S. Liu, Z. Du, J. Tao, D. Han, T. Luo, Y. Xie, Y. Chen, and T. Chen, “Cambricon: An instruction set architecture for neural networks,” in Proceedings of the 43rd International Symposium on Computer Architecture , ser. ISCA ’16, 2016, p. 393–405
2016
Earlier work this paper cites.
D. Mahajan, J. Park, E. Amaro, H. Sharma, A. Yazdanbakhsh, J. K. Kim, and H. Esmaeilzadeh, “Tabla: A unified template-based framework for accelerating statistical machine learning,” in 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2016, pp. 14–26
2016
Earlier work this paper cites.
J. McCaffrey. (2016) The max trick when computing softmax. [Online]. Available: https://jamesmccaffrey.wordpress.com/2016/03/04/the-max-trick-when-computing-softmax/
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
B. Reagen, P. Whatmough, R. Adolf, S. Rama, H. Lee, S. K. Lee, J. M. Hernández-Lobato, G. Wei, and D. Brooks, “Minerva: Enabling low-power, highly-accurate deep neural network accelerators,” in 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA) , June 2016, pp. 267–278
2016
Earlier work this paper cites.
H. Sharma, J. Park, D. Mahajan, E. Amaro, J. K. Kim, C. Shao, A. Mishra, and H. Esmaeilzadeh, “From high-level deep neural models to fpgas,” in 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2016, pp. 1–12
2016
Earlier work this paper cites.
S. Teerapittayanon, B. McDanel, and H. T. Kung, “Branchynet: Fast inference via early exiting from deep neural networks,” in 2016 23rd International Conference on Pattern Recognition (ICPR) , 2016, pp. 2464–2469
2016
Earlier work this paper cites.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V. Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D. H. Yoon, “In-datacenter performance analysis of a tensor processing unit,” in 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA) , June 2017, pp. 1–12
2017
Earlier work this paper cites.
A. Parashar, M. Rhu, A. Mukkara, A. Puglielli, R. Venkatesan, B. Khailany, J. Emer, S. W. Keckler, and W. J. Dally, “Scnn: An accelerator for compressed-sparse convolutional neural networks,” in Proceedings of the 44th Annual International Symposium on Computer Architecture , ser. ISCA ’17. New York, NY, USA: Association for Computing Machinery, 2017, p. 27–40
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Venkataramani, A. Ranjan, S. Banerjee, D. Das, S. Avancha, A. Jagannathan, A. Durg, D. Nagaraj, B. Kaul, P. Dubey, and A. Raghunathan, “Scaledeep: A scalable compute architecture for learning and evaluating deep networks,” SIGARCH Comput. Archit. News , 2017
2017
Earlier work this paper cites.
V. Akhlaghi, A. Yazdanbakhsh, K. Samadi, R. K. Gupta, and H. Esmaeilzadeh, “Snapea: Predictive early activation for reducing computation in deep convolutional neural networks,” in Proceedings of the 45th Annual International Symposium on Computer Architecture , 2018, p. 662–673
2018
Earlier work this paper cites.
2018
Cited alongside, same era.
K. Hegde, J. Yu, R. Agrawal, M. Yan, M. Pellauer, and C. W. Fletcher, “Ucnn: Exploiting computational reuse in deep neural networks via weight repetition,” in Proceedings of the 45th Annual International Symposium on Computer Architecture , 2018, p. 674–687
2018
Cited alongside, same era.
A. Jain, A. Phanishayee, J. Mars, L. Tang, and G. Pekhimenko, “Gist: Efficient data encoding for deep neural network training,” in Proceedings of the 45th Annual International Symposium on Computer Architecture , ser. ISCA ’18, 2018, p. 776–789
2018
Cited alongside, same era.
2018
O. Zafrir, G. Boudoukh, P. Izsak, and M. Wasserblat, “Q8bert: Quantized 8bit bert,” in 33rd Conference on Neural Information Processing Systems (NeurIPS) , 2019
2019
Later among the works it cites.
Catapult High-Level Synthesis , accessed Oct 1, 2020. [Online]. Available: https://www.mentor.com/hls-lp/catapult-high-level-synthesis
2020
Closest in time.
Jetson TX2 Module , accessed Oct 1, 2020. [Online]. Available: https://developer.nvidia.com/embedded/jetson-tx2
2020
Closest in time.
T. Ajayi, S. Kamineni, Y. Cherivirala, M. Fayazi, K. Kwon, M. Saligane, S. Gupta, C. Chen, D. Sylvester, D. Dreslinski, B. Calhoun, and D. Wentzloff, “An open-source framework for autonomous soc design with analog block generation,” in 020 IFIP/IEEE 28th International Conference on Very Large Scale Integration (VLSI-SoC) , 2020
2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
B. Khailany, E. Khmer, R. Venkatesan, J. Clemons, J. S. Emer, M. Fojtik, A. Klinefelter, M. Pellauer, N. Pinckney, Y. S. Shao, S. Srinath, C. Torng, S. L. Xi, Y. Zhang, and B. Zimmer, “A modular digital vlsi flow for high-productivity soc design,” in Proceedings of the 55th Annual Design Automation Conference , ser. DAC ’18. New York, NY, USA: ACM, 2018, pp. 72:1–72:6. [Online]. Available: http://doi.acm.org/10.1145/3195970.3199846
2018
Cited alongside, same era.
H. Kwon, A. Samajdar, and T. Krishna, “Maeri: Enabling flexible dataflow mapping over dnn accelerators via reconfigurable interconnects,” SIGPLAN Not. , vol. 53, no. 2, p. 461–475, Mar. 2018. [Online]. Available: https://doi.org/10.1145/3296957.3173176
2018
Cited alongside, same era.
P. Meinerzhagen, C. Tokunaga, A. Malavasi, V. Vaidya, A. Mendon, D. Mathaikutty, J. Kulkarni, C. Augustine, M. Cho, S. Kim, G. Matthew, R. Jain, J. Ryan, C. Peng, S. Paul, S. Vangal, B. Esparza, L. Cuellar, M. Woodman, B. Iyer, S. Maiyuran, G. Chinya, C. Zou, Y. Liao, K. Ravichandran, H. Wang, M. Khellah, J. Tschanz, and V. De, “2.3 an energy-efficient graphics processor featuring fine-grain dvfs with integrated voltage regulators, execution-unit turbo, and retentive sleep in 14nm tri-gate cmos,” in 2018 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC) , 2018
2018
Cited alongside, same era.
E. Park, D. Kim, and S. Yoo, “Energy-efficient neural network accelerator based on outlier-aware low-precision computation,” in Proceedings of the 45th Annual International Symposium on Computer Architecture , 2018, p. 688–698
2018
Cited alongside, same era.
B. Reagen, U. Gupta, L. Pentecost, P. Whatmough, S. K. Lee, N. Mulholland, D. Brooks, and G. Wei, “Ares: A framework for quantifying the resilience of deep neural networks,” in 2018 55th ACM/ESDA/IEEE Design Automation Conference (DAC) , 2018, pp. 1–6
2018
Cited alongside, same era.
M. Riera, J.-M. Arnau, and A. González, “Computation reuse in dnns by exploiting input similarity,” in Proceedings of the 45th Annual International Symposium on Computer Architecture , 2018, p. 57–68
2018
Cited alongside, same era.
2018
Cited alongside, same era.
P. N. Whatmough, S. K. Lee, D. Brooks, and G.-Y. Wei, “DNN Engine: A 28-nm Timing-Error Tolerant Sparse Deep Neural Network Processor for IoT Applications,” IEEE Journal of Solid-State Circuits , vol. 53, no. 9, pp. 2722–2731, 2018
2018
Cited alongside, same era.
S. Bang, W. Lim, C. Augustine, A. Malavasi, M. Khellah, J. Tschanz, and V. De, “25.1 a fully synthesizable distributed and scalable all-digital ldo in 10nm cmos,” in 2020 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC) , 2020
2020
Closest in time.
I. Fedorov, M. Stamenovic, C. Jensen, L.-C. Yang, A. Mandell, Y. Gan, M. Mattina, and P. N. Whatmough, “TinyLSTMs: Efficient Neural Speech Enhancement for Hearing Aids,” in Proc. Interspeech 2020 , 2020, pp. 4054–4058. [Online]. Available: http://dx.doi.org/10.21437/Interspeech.2020-1864
2020
Closest in time.
2020
Closest in time.
T. J. Ham, S. J. Jung, S. Kim, Y. H. Oh, Y. Park, Y. Song, J. Park, S.-H. Lee, K. Park, J. Lee, and D.-K. Jeong, “A 3 : Accelerating attention mechanisms in neural networks with approximation,” 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) , pp. 328–341, 2020
2020
Closest in time.
2020
Closest in time.
G. G. Ko, Y. Chai, M. Donato, P. N. Whatmough, T. Tambe, R. A. Rutenbar, D. Brooks, and G.-Y. Wei, “A 3mm<sup>2</sup> programmable bayesian inference accelerator for unsupervised machine perception using parallel gibbs sampling in 16nm,” in 2020 IEEE Symposium on VLSI Circuits , 2020, pp. 1–2
2020
Closest in time.
2020
Closest in time.
S. Li, Z. Yang, D. Reddy, A. Srivastava, and B. Jacob, “Dramsim3: A cycle-accurate, thermal-capable dram simulator,” IEEE Computer Architecture Letters , vol. 19, no. 2, pp. 106–109, 2020
2020
Closest in time.
Z.-G. Liu, P. N. Whatmough, and M. Mattina, “Systolic tensor array: An efficient structured-sparse gemm accelerator for mobile cnn inference,” IEEE Computer Architecture Letters , vol. 19, no. 1, pp. 34–37, 2020
2020
Closest in time.
2020
Closest in time.
J. Park, H. Yoon, D. Ahn, J. Choi, and J.-J. Kim, “Optimus: Optimized matrix multiplication structure for transformer neural network accelerator,” in Proceedings of Machine Learning and Systems , I. Dhillon, D. Papailiopoulos, and V. Sze, Eds., 2020, vol. 2, pp. 363–378
2020
Closest in time.
A. Samajdar, J. M. Joseph, Y. Zhu, P. Whatmough, M. Mattina, and T. Krishna, “A Systematic Methodology for Characterizing Scalability of DNN Accelerators using SCALE-Sim,” in 2020 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) , 2020, pp. 58–68
2020
Closest in time.
R. Schwartz, G. Stanovsky, S. Swayamdipta, J. Dodge, and N. A. Smith, “The right tool for the job: Matching model and instance complexities,” in ACL , 2020
2020
Closest in time.
S. Shen, Z. Dong, J. Ye, L. Ma, Z. Yao, A. Gholami, M. Mahoney, and K. Keutzer, “Q-bert: Hessian based ultra low precision quantization of bert,” in AAAI , 2020
2020
Closest in time.
Z. Sun, H. Yu, X. Song, R. Liu, Y. Yang, and D. Zhou, “Mobilebert: a compact task-agnostic bert for resource-limited devices,” in ACL , 2020
2020
Closest in time.
J. Weng, S. Liu, V. Dadu, Z. Wang, P. Shah, and T. Nowatzki, “Dsagen: Synthesizing programmable spatial accelerators,” in Proceedings of the ACM/IEEE 47th Annual International Symposium on Computer Architecture , ser. ISCA ’20, 2020
2020
Closest in time.
P. N. Whatmough, M. Donato, G. G. Ko, S. K. Lee, D. Brooks, and G.-Y. Wei, “CHIPKIT: An Agile, Reusable Open-Source Framework for Rapid Test Chip Development,” IEEE Micro , vol. 40, no. 4, pp. 32–40, 2020
2020
Closest in time.
S. L. Xi, Y. Yao, K. Bhardwaj, P. Whatmough, G.-Y. Wei, and D. Brooks, “SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning Workloads,” ACM Trans. Archit. Code Optim. , vol. 17, no. 4, Nov. 2020. [Online]. Available: https://doi.org/10.1145/3424669
2020
Closest in time.
2020
Closest in time.
A. H. Zadeh and A. Moshovos, “Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference,” in 53rd IEEE/ACM International Symposium on Microarchitecture (MICRO) , 2020
2020
Closest in time.
2020
Closest in time.
A. Agrawal, S. Lee, J. Silberman, M. Ziegler, M. Kang, S. Venkataramani, N. Cao, B. Fleischer, M. Guillorn, M. Cohen, S. Mueller, J. Oh, M. Lutz, J. Jung, S. Koswatta, C. Zhou, V. Zalani, J. Bonanno, R. Casatuta, C. Chen, J. Choi, H. Haynie, A. Herbert, R. Jain, M. Kar, K. Kim, Y. Li, Z. Ren, S. Rider, M. Schaal, K. Schelm, M. Scheuermann, X. Sun, H. Tran, N. Wang, W. Wang, X. Zhang, V. Shah, B. Curran, V. Srinivasan, P. Lu, S. Shukla, L. Chang, and K. Gopalakrishnan, “9.1 a 7nm 4-core ai chip with 25.6 tflops hybrid fp8 training, 102.4 tops int4 inference and workload-aware throttling,” in 2021 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC) , 2021
2021
Closest in time.
B. Huang, E. Fang, S. Hsueh, R. Huang, A. Lin, C. Chiang, Y. Lin, W. Hsieh, B. Chen, Y. Zhuang, C. Wu, J. Chen, Y. Chen, C. Wan, E. Wang, A. Chiou, P. Kao, Y. Tsai, H. Chen, and S. Hwang, “35.1 an octa-core 2.8/2ghz dual-gear sensor-assisted high-speed and power-efficient cpu in 7nm finfet 5g smartphone soc,” in 2021 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC) , 2021
2021
Closest in time.
T. Tambe, E.-Y. Yang, G. G. Ko, Y. Chai, C. Hooper, M. Donato, P. N. Whatmough, A. M. Rush, D. Brooks, and G.-Y. Wei, “9.8 A 25mm2 SoC for IoT Devices with 18ms Noise-Robust Speech-to-Text Latency via Bayesian Speech Denoising and Attention-Based Sequence-to-Sequence DNN Speech Recognition in 16nm FinFET,” in 2021 IEEE International Solid- State Circuits Conference (ISSCC) , vol. 64, 2021, pp. 158–160
2021
Closest in time.
H. Wang, Z. Zhang, and S. Han, “Spatten: Efficient sparse attention architecture with cascade token and head pruning,” 2021 IEEE International Symposium on High Performance Computer Architecture (HPCA) , 2021
2021
Closest in time.