Fetching the paper…
Reading the bibliography…
On-device learning allows AI models to adapt to user data, thereby enhancing service quality on edge platforms.
J. H. Wilkinson, Rounding Errors in Algebraic Processes . Dover Publications, 1964
1964
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on . IEEE, 2009, pp. 248–255
2009
Earlier work this paper cites.
K. C. Chun, P. Jain, J. H. Lee, and C. H. Kim, “A 3t gain cell embedded dram utilizing preferential boosting for high density and low power on-die caches,” IEEE Journal of Solid-State Circuits , vol. 46, no. 6, pp. 1495–1505, 2011
2011
Earlier work this paper cites.
Y. Chen, T. Luo, S. Liu, S. Zhang, L. He, J. Wang, L. Li, T. Chen, Z. Xu, N. Sun, and O. Temam, “Dadiannao: A machine-learning supercomputer,” in Proceedings of the 47th Annual IEEE/ACM International Symposium on Microarchitecture . IEEE Computer Society, 2014, pp. 609–622
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
A. Krizhevsky, V. Nair, and G. Hinton, “The cifar-10 dataset,” 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
I. Bhati, M.-T. Chang, Z. Chishti, S.-L. Lu, and B. Jacob, “Dram refresh mechanisms, penalties, and trade-offs,” IEEE Transactions on Computers , vol. 65, no. 1, pp. 108–121, 2015
2015
Earlier work this paper cites.
M. Courbariaux, Y. Bengio, and J.-P. David, “Binaryconnect: Training deep neural networks with binary weights during propagations,” in Advances in neural information processing systems , 2015, pp. 3123–3131
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in International conference on machine learning . PMLR, 2015, pp. 448–456
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
P. Judd, J. Albericio, T. Hetherington, T. M. Aamodt, and A. Moshovos, “Stripes: Bit-serial deep neural network computing,” in Microarchitecture (MICRO), 2016 49th Annual IEEE/ACM International Symposium on . IEEE, 2016, pp. 1–12
2016
Earlier work this paper cites.
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 2818–2826
2016
Earlier work this paper cites.
J. Zhang, Z. Wang, and N. Verma, “A machine-learning classifier implemented in a standard 6t sram array,” in 2016 ieee symposium on vlsi circuits (vlsi-circuits) . IEEE, 2016, pp. 1–2
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
C. Coleman, D. Narayanan, D. Kang, T. Zhao, J. Zhang, L. Nardi, P. Bailis, K. Olukotun, C. Ré, and M. Zaharia, “Dawnbench: An end-to-end deep learning benchmark and competition,” Training , vol. 100, no. 101, p. 102, 2017
2017
Earlier work this paper cites.
A. N. Gomez, M. Ren, R. Urtasun, and R. B. Grosse, “The reversible residual network: Backpropagation without storing activations,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio, “Quantized neural networks: Training neural networks with low precision weights and activations,” The Journal of Machine Learning Research , vol. 18, no. 1, pp. 6869–6898, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
E. Park, J. Ahn, and S. Yoo, “Weighted-entropy-based quantization for deep neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017, pp. 5456–5464
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
R. Banner, I. Hubara, E. Hoffer, and D. Soudry, “Scalable methods for 8-bit training of neural networks,” in NeurIPS , 2018, pp. 5151–5159
2018
Cited alongside, same era.
A. Biswas and A. P. Chandrakasan, “Conv-ram: An energy-efficient sram with embedded convolution computation for low-power cnn-based machine learning applications,” in 2018 IEEE International Solid-State Circuits Conference-(ISSCC) . IEEE, 2018, pp. 488–490
2018
Cited alongside, same era.
2018
Cited alongside, same era.
M. Mahmoud, I. Edo, A. H. Zadeh, O. M. Awad, G. Pekhimenko, J. Albericio, and A. Moshovos, “Tensordash: Exploiting sparsity to accelerate deep neural network training,” in 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2020, pp. 781–795
2020
Later among the works it cites.
E. Qin, A. Samajdar, H. Kwon, V. Nadella, S. Srinivasan, D. Das, B. Kaul, and T. Krishna, “Sigma: A sparse and irregular gemm accelerator with flexible interconnects for dnn training,” in 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA) . IEEE, 2020, pp. 58–70
2020
Later among the works it cites.
T. Tambe, E.-Y. Yang, Z. Wan, Y. Deng, V. Janapa Reddi, A. Rush, D. Brooks, and G.-Y. Wei, “Algorithm-hardware co-design of adaptive floating-point encodings for resilient deep learning inference,” in 2020 57th ACM/IEEE Design Automation Conference (DAC) , 2020, pp. 1–6
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko, “Quantization and training of neural networks for efficient integer-arithmetic-only inference,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 2704–2713
2018
Cited alongside, same era.
P. Micikevicius, S. Narang, J. Alben, G. Diamos, E. Elsen, D. Garcia, B. Ginsburg, M. Houston, O. Kuchaiev, G. Venkatesh, and H. Wu, “Mixed precision training,” in International Conference on Learning Representations , 2018. [Online]. Available: https://openreview.net/forum?id=r1gs9JgRZ
2018
Cited alongside, same era.
2018
Cited alongside, same era.
F. Tu, W. Wu, S. Yin, L. Liu, and S. Wei, “Rana: Towards efficient neural acceleration with refresh-optimized embedded dram,” in 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA) . IEEE, 2018, pp. 340–352
2018
Cited alongside, same era.
S. Koppula, L. Orosa, A. G. Yağlıkçı, R. Azizi, T. Shahroodi, K. Kanellopoulos, and O. Mutlu, “Eden: Enabling energy-efficient, high-performance deep neural network inference using approximate dram,” in Proceedings of the 52nd Annual IEEE/ACM International Symposium on Microarchitecture , ser. MICRO ’52. New York, NY, USA: Association for Computing Machinery, 2019, p. 166–181. [Online]. Available: https://doi.org/10.1145/3352460.3358280
2019
Cited alongside, same era.
J. Lee, J. Lee, D. Han, J. Lee, G. Park, and H.-J. Yoo, “7.7 lnpu: A 25.3 tflops/w sparse deep-neural-network learning processor with fine-grained mixed precision of fp8-fp16,” in 2019 IEEE International Solid-State Circuits Conference-(ISSCC) . IEEE, 2019, pp. 142–144
2019
Cited alongside, same era.
B. McDanel, S. Q. Zhang, H. Kung, and X. Dong, “Full-stack optimization for accelerating cnns using powers-of-two weights with fpga validation,” in Proceedings of the ACM International Conference on Supercomputing , 2019, pp. 449–460
2019
Cited alongside, same era.
J. Narinx, R. Giterman, A. Bonetti, N. Frigerio, C. Aprile, A. Burg, and Y. Leblebici, “A 24 kb single-well mixed 3t gain-cell edram with body-bias in 28 nm fd-soi for refresh-free dsp applications,” in 2019 IEEE Asian Solid-State Circuits Conference (A-SSCC) , 2019, pp. 219–222
2019
Cited alongside, same era.
D. Yang, A. Ghasemazar, X. Ren, M. Golub, G. Lemieux, and M. Lis, “Procrustes: a dataflow and accelerator for sparse deep neural network training,” in 2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) . IEEE, 2020, pp. 711–724
2020
Later among the works it cites.
C. Yu, T. Yoo, H. Kim, T. T.-H. Kim, K. C. T. Chuan, and B. Kim, “A logic-compatible edram compute-in-memory with embedded adcs for processing neural networks,” IEEE Transactions on Circuits and Systems I: Regular Papers , vol. 68, no. 2, pp. 667–679, 2020
2020
Later among the works it cites.
“Bfloat16: The secret to high performance on cloud tpus,” https://cloud.google.com/blog/products/ai-machine-learning/bfloat16-the-secret-to-high-performance-on-cloud-tpus , Google, accessed: 2021-03-29
2021
Later among the works it cites.
D. Han, D. Im, G. Park, Y. Kim, S. Song, J. Lee, and H.-J. Yoo, “Hnpu: An adaptive dnn training processor utilizing stochastic dynamic fixed-point and active bit-precision searching,” IEEE Journal of Solid-State Circuits , vol. 56, no. 9, pp. 2858–2869, 2021
2021
Later among the works it cites.
S. M. A. H. Jafri, H. Hassan, A. Hemani, and O. Mutlu, “Refresh triggered computation: Improving the energy efficiency of convolutional neural network accelerators,” ACM Trans. Archit. Code Optim. , vol. 18, no. 1, dec 2021. [Online]. Available: https://doi.org/10.1145/3417708
2021
Later among the works it cites.
“Accelerating ai training with nvidia tf32 tensor cores,” https://developer.nvidia.com/blog/accelerating-ai-training-with-tf32-tensor-cores/ , Nvidia, accessed: 2021-03-29
2021
Later among the works it cites.
S. Q. Zhang, B. McDanel, H. Kung, and X. Dong, “Training for multi-resolution inference using reusable quantization terms,” in Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , 2021, pp. 845–860
2021
Later among the works it cites.
Catapult High-Level Synthesis , accessed Nov 1, 2022. [Online]. Available: https://www.mentor.com/hls-lp/catapult-high-level-synthesis
2022
Later among the works it cites.
A. Hatamizadeh, Y. Tang, V. Nath, D. Yang, A. Myronenko, B. Landman, H. R. Roth, and D. Xu, “Unetr: Transformers for 3d medical image segmentation,” in Proceedings of the IEEE/CVF winter conference on applications of computer vision , 2022, pp. 574–584
2022
Later among the works it cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 16 000–16 009
2022
Later among the works it cites.
Semiconductor Research Corporation, “The decadal plan for semiconductors,” https://www.src.org/about/decadal-plan/ , accessed: 2022-11-04
2022
Later among the works it cites.
2022
Later among the works it cites.
S. Q. Zhang, B. McDanel, and H. Kung, “Fast: Dnn training under variable precision block floating point with stochastic rounding,” in 2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2022, pp. 846–860
2022
Later among the works it cites.
Y. Zheng, H. Yang, Y. Shu, Y. Jia, and Z. Huang, “Optimizing off-chip memory access for deep neural network accelerator,” IEEE Transactions on Circuits and Systems II: Express Briefs , vol. 69, no. 4, pp. 2316–2320, 2022
2022
Later among the works it cites.
2023
Closest in time.
2023
Closest in time.