Fetching the paper…
Reading the bibliography…
Attention-based models have demonstrated remarkable success in various natural language understanding tasks.
1909
Earlier work this paper cites.
1909
Earlier work this paper cites.
2004
Earlier work this paper cites.
N. Muralimanohar, R. Balasubramonian, and N. Jouppi, “Cacti 6.0: A tool to model large caches,” HP Laboratories , 01 2009
2009
Earlier work this paper cites.
P. Rosenfeld, E. Cooper-Balis, and B. Jacob, “Dramsim2: A cycle accurate memory system simulator,” IEEE Computer Architecture Letters , vol. 10, no. 1, pp. 16–19, Jan 2011
2011
Earlier work this paper cites.
2016
Earlier work this paper cites.
S. Han, X. Liu, H. Mao, J. Pu, A. Pedram, M. A. Horowitz, and W. J. Dally, “Eie: Efficient inference engine on compressed deep neural network,” in Proceedings of the 43rd International Symposium on Computer Architecture , ser. ISCA ’16. IEEE Press, 2016, p. 243–254. [Online]. Available: https://doi.org/10.1109/ISCA.2016.30
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
Y.-H. Chen, J. Emer, and V. Sze, “Eyeriss: A spatial architecture for energy-efficient dataflow for convolutional neural networks,” in Intl’ Symp. on Computer Architecture , 2016
2016
Earlier work this paper cites.
Y. Cheng, D. Wang, P. Zhou, and T. Zhang, “A survey of model compression and acceleration for deep neural networks,” 2017
2017
Earlier work this paper cites.
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P.-l. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V. Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C. R. Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D. H. Yoon, “In-datacenter performance analysis of a tensor processing unit,” in International Symposium on Computer Architecture , 2017
2017
Earlier work this paper cites.
L. Durant, O. Giroux, M. Harris, and N. Stam, “Nvidia developer blog,” May 2017. [Online]. Available: https://devblogs.nvidia.com/inside-volta/
2017
Earlier work this paper cites.
E. Park, S. Yoo, and P. Vajda, “Value-aware quantization for training and inference of neural networks,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 580–595
2018
Earlier work this paper cites.
E. Park, D. Kim, and S. Yoo, Energy-Efficient Neural Network Accelerator Based on Outlier-Aware Low-Precision Computation . IEEE Press, 2018, p. 688–698. [Online]. Available: https://doi.org/10.1109/ISCA.2018.00063
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT , 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
R. Tang, Y. Lu, and J. Lin, “Natural language generation for effective knowledge distillation,” in Proceedings of the 2nd Workshop on Deep Learning Approaches for Low-Resource NLP (DeepLo 2019) , 2019, pp. 202–208
2019
Later among the works it cites.
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
V. Sanh, L. Debut, J. Chaumond, and T. Wolf, “Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter,” 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing - NeurIPS , 2019
2019
Cited alongside, same era.
O. Zafrir, G. Boudoukh, P. Izsak, and M. Wasserblat, “Q8bert: Quantized 8bit bert,” 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing - NeurIPS , 2019
2019
Cited alongside, same era.
2019
Cited alongside, same era.
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman, “GLUE: A multi-task benchmark and analysis platform for natural language understanding,” in 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019 . OpenReview.net, 2019. [Online]. Available: https://openreview.net/forum?id=rJ4km2R5t7
2019
Cited alongside, same era.
S. Narasimhan, “Nvidia developer blog,” Aug 2019. [Online]. Available: https://devblogs.nvidia.com/training-bert-with-gpus]/
2019
Cited alongside, same era.
E. Wang, J. J. Davis, R. Zhao, H.-C. Ng, X. Niu, W. Luk, P. Y. K. Cheung, and G. A. Constantinides, “Deep neural network approximation for custom hardware,” ACM Computing Surveys , vol. 52, no. 2, p. 1–39, May 2019. [Online]. Available: http://dx.doi.org/10.1145/3309551
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Cited alongside, same era.
2019
Later among the works it cites.
M. A. Raihan, N. Goli, and T. M. Aamodt, “Modeling deep learning accelerator enabled gpus,” in 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS) , 2019, pp. 79–92
2019
Later among the works it cites.
S. Shen, Z. Dong, J. Ye, L. Ma, Z. Yao, A. Gholami, M. W. Mahoney, and K. Keutzer, “Q-bert: Hessian based ultra low precision quantization of bert.” in AAAI , 2020, pp. 8815–8821
2020
Closest in time.
J. Fang, A. Shafiee, H. Abdel-Aziz, D. Thorsley, G. Georgiadis, and J. Hassoun, “Post-training piecewise linear quantization for deep neural networks,” 2020
2020
Closest in time.
H. Wang, Z. Wu, Z. Liu, H. Cai, L. Zhu, C. Gan, and S. Han, “HAT: Hardware-aware transformers for efficient natural language processing,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics . Online: Association for Computational Linguistics, Jul. 2020, pp. 7675–7688. [Online]. Available: https://www.aclweb.org/anthology/2020.acl-main.686
2020
Closest in time.
M. Joshi, D. Chen, Y. Liu, D. S. Weld, L. Zettlemoyer, and O. Levy, “Spanbert: Improving pre-training by representing and predicting spans,” Transactions of the Association for Computational Linguistics , vol. 8, pp. 64–77, 2020
2020
Closest in time.
2020
Closest in time.
2020
Closest in time.