2018

CMSIS-NN: Efficient Neural Network Kernels for Arm Cortex-M CPUs

Lai, Liangzhen, Suda, Naveen, Chandra, Vikas

Understand

Deep Neural Networks are becoming increasingly popular in always-on IoT edge devices performing data analytics right at the source, reducing latency as well as energy consumption for data communication.

  • This paper presents CMSIS-NN, efficient kernels developed to maximize the performance and minimize the memory footprint of neural network (NN) applications on Arm Cortex-M processors targeted for intelligent IoT edge devices.
  • Neural network inference based on CMSIS-NN kernels achieves 4.6X improvement in runtime/throughput and 4.9X improvement in energy efficiency.

Built on

Similar

  • Fixed point quantization of deep convolutional networks

    Darryl Lin, Sachin Talathi, and Sreekanth Annapureddy · 2016

    Cited alongside, same era.

  • Throughput-optimized opencl-based fpga accelerator for large-scale convolutional neural networks

    Naveen Suda, Vikas Chandra, Ganesh Dasika, Abinash Mohanty, Yufei Ma, Sarma Vrudhula, Jae-sun Seo, and Yu Cao · 2016

    Cited alongside, same era.

  • Optimizing memory efficiency for deep convolutional neural networks on gpus

    Chao Li, Yi Yang, Min Feng, Srimat Chakradhar, and Huiyang Zhou · 2016

    Cited alongside, same era.

  • The route to a trillion devices

    Philip Sparks

    Cited in the paper.

  • https://github.com/ARM-software/CMSIS_5

    Cited in the paper.

  • http://www.arm.com/products/processors/cortex-m

    Cited in the paper.

  • http://www.st.com/en/evaluation-tools/nucleo-f746zg.html

    Original

    Nucleo-f746zg development board

    Cited in the paper.

Then

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…