Deep Neural Network (DNN) is the state-of-the-art neural network computing model that successfully achieves close-to or better than human performance in many large scale cognitive applications, like computer vision, speech recognition, nature language processing, object recognition, etc. The most successful DNN is deep convolutional neural network consisting of multiple types of layers including convolution, activation, pooling and fully-connected layers. Typically, a DNN may have tens to thousands of layers to achieve optimized inference accuracy for practical applications, which makes it heavily memory (tens of GB working memory) and computing intensive (needs powerful CPU, GPU, FPGA, ASIC, etc.). Our research focus on:
- Explore automated and general methodologies to simultaneously reduce DNN model size and computing complexity, while maintaining state-of-the-art accuracy
- Explore how to design and deploy hardware-efficient DNN model in low power and resource limited mobile system, embedded system, IoT, edge devices for various applications, such as pattern recognition, object tracking/detection, etc.
- Explore memory- and computing-efficient on-device continual learning algorithm and system design