← 返回 JSSC 论文列表
📄 下载 JSSC 原文 PDF
JSSC 2024第6期Digital Circuits40nm

A 28.8-mW Accelerator IC for Dark Channel Prior-Based Blind Image Deblurring Po-Shao Chen

本文提出了一种基于暗通道先验的盲图像去模糊加速器,显著降低了计算延迟和功耗。
40nm CMOS, 0.65V, 65MHz, 28.8mW
暗通道先验盲图像去模糊加速器FFT卷积引擎
2-D Laplace方程求解器减少56%边界包裹延迟
梯度数据局部性减少57%潜在图像估计延迟
并行架构2-D卷积引擎减少79%模糊核估计延迟
Abstract
This work presents an accelerator that performs blind deblurring based on the dark channel prior. The alternating minimization algorithm is leveraged for latent image and blur kernel estimation. A 2-D Laplace equation solver is embedded to reduce the latency by 56% for boundary wrapping. For latent image estimation, gradient data locality is employed to reduce the latency by 57%. A sorting engine is designed to reduce the latency in data access by 96% for calculating the dark channel. A pipelined mixed-radix 1-D fast Fourier transform (FFT) engine is used for efficient latent image estimation and blur kernel estimation. By employing image size approximation, 85% of additions and 97% of multiplications for FFT can further be saved. In the blur kernel estimator, a 2-D convolution engine with a parallel architecture is implemented, reducing the latency by 79%. The accelerator supports blur kernels of 25 × 25 and 49 × 49 pixels for blurred images of 129 × 129 and 257 × 257 pixels, respectively. Fabricated in 40-nm CMOS, the accelerator’s core area is 3.98 mm2. The chip dissipates 28.8 mW at 65 MHz from a 0.65-V supply. It can estimate a blur kernel of 25 × 25 pixels for an image patch with 129 × 129 pixels for deblurring a full-HD image in 1.7 s, achieving a 2562× shorter latency than a high-end CPU. Compared with the state-of-the- art design, the chip achieves a four times higher normalized area efficiency and a 7.5× higher normalized energy efficiency.