← 返回论文列表 📄 下载原文 PDF  ISSCC 2025 · 23.4
ISSCC 2025Session 23 · AI-ACCELERATORSOther28nm CMOS

Nebula: A 28nm 109.8TOPS/W 3D PNN Accelerator Featuring

⚡ 本页包含 AI 生成的分析内容,仅供参考

📋 论文概要

该论文提出了一款名为Nebula的28nm工艺3D点云神经网络(PNN)加速器,通过自适应分区、多跳过和块聚合技术,实现了109.8TOPS/W的高能效。解决了点云处理中计算复杂度和冗余距离计算的问题。

💡 主要创新点

核心指标
109.8TOPS/W
工艺节点
28nm CMOS
重要性
发表年份
ISSCC 2025

🏷 关键词

3D点云神经网络加速器自适应分区多跳过块聚合

📄 原文摘要

Changchun Zhou1, Tianling Huang1, Yanzhe Ma1, Yuzhe Fu1, Xiangjie Song1, Siyuan Qiu1, Jiacong Sun1, Min Liu1, Ge Li1, Yifan He2, Yuchao Yang1,3, Hailong Jiao1 Peking University, Shenzhen, China Reconova Technologies, Xiamen, China 3 Peking University, Beijing, China exceeds Pth, the address of the block is recorded, and the corresponding points are further partitioned at the next level. In the hardware implementation of sampling, 16 cores handle 16 blocks, respectively. A mask (see Fig. 23.4.3), in which “0” indicates RP, is used to skip redundant distance calculations between the already-generated RPs, thereby reducing the number of cycles in sampling. When evaluated on PointNeXt-Cls@ModelNet40, with the same number of blocks (e.g. 16 blocks), APS achieves a more balanced partition than [13], reducing the mean square error by 83.6%, improving the accuracy by 1.5%, and reducing the latency by 76.8%.

👥 作者与机构

Adaptive Partition, Multi-Skipping, and Block-Wise Aggregation

分类:Other · 年份:ISSCC 2025