상세 보기
Low Latency Implementations of CNN for Resource-Constrained IoT Devices
- Mujtaba, Ahmed;
- Lee, Wai-Kong;
- Hwang, Seong Oun
WEB OF SCIENCE
7SCOPUS
10초록
Convolutional Neural Network (CNN) inference on a resource-constrained Internet-of-Things (IoT) device (i.e., ARM Cortex-M microcontroller) requires careful optimization to reduce the timing overhead. We propose two novel techniques to improve the computational efficiency of CNNs by targeting low-cost microcontrollers. Our techniques utilize on-chip memory and minimize redundant operations, yielding low-latency inference results on complex quantized models such as MobileNetV1. On the ImageNet dataset for per-layer quantization, we reduce inference latency and Multiply-and-Accumulate (MAC) per cycle by 22.4% and 22.9%, respectively, compared to the state-of-theart mixed-precision CMix-NN library. On the CIFAR-10 dataset for per-channel quantization, we reduce inference latency and MAC per cycle by 31.7% and 31.3%, respectively. The achieved low-latency inference results can improve the user experience and save power budget in resource-constrained IoT devices.
키워드
- 제목
- Low Latency Implementations of CNN for Resource-Constrained IoT Devices
- 저자
- Mujtaba, Ahmed; Lee, Wai-Kong; Hwang, Seong Oun
- 발행일
- 2022-12
- 유형
- Article
- 권
- 69
- 호
- 12
- 페이지
- 5124 ~ 5128