Low Latency Implementations of CNN for Resource-Constrained IoT Devices

Citations

WEB OF SCIENCE

7
Citations

SCOPUS

10

초록

Convolutional Neural Network (CNN) inference on a resource-constrained Internet-of-Things (IoT) device (i.e., ARM Cortex-M microcontroller) requires careful optimization to reduce the timing overhead. We propose two novel techniques to improve the computational efficiency of CNNs by targeting low-cost microcontrollers. Our techniques utilize on-chip memory and minimize redundant operations, yielding low-latency inference results on complex quantized models such as MobileNetV1. On the ImageNet dataset for per-layer quantization, we reduce inference latency and Multiply-and-Accumulate (MAC) per cycle by 22.4% and 22.9%, respectively, compared to the state-of-theart mixed-precision CMix-NN library. On the CIFAR-10 dataset for per-channel quantization, we reduce inference latency and MAC per cycle by 31.7% and 31.3%, respectively. The achieved low-latency inference results can improve the user experience and save power budget in resource-constrained IoT devices.

키워드

Convolutional neural networksInternet-of-Thingsmicrocontrollerstiny ML
제목
Low Latency Implementations of CNN for Resource-Constrained IoT Devices
저자
Mujtaba, AhmedLee, Wai-KongHwang, Seong Oun
DOI
10.1109/TCSII.2022.3205029
발행일
2022-12
유형
Article
저널명
IEEE Transactions on Circuits and Systems II: Express Briefs
69
12
페이지
5124 ~ 5128