⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文提出了一款名为Nebula的28nm工艺3D点云神经网络(PNN)加速器,通过自适应分区、多跳过和块聚合技术,实现了109.8TOPS/W的高能效。解决了点云处理中计算复杂度和冗余距离计算的问题。
Changchun Zhou1, Tianling Huang1, Yanzhe Ma1, Yuzhe Fu1, Xiangjie Song1, Siyuan Qiu1, Jiacong Sun1, Min Liu1, Ge Li1, Yifan He2, Yuchao Yang1,3, Hailong Jiao1 Peking University, Shenzhen, China Reconova Technologies, Xiamen, China 3 Peking University, Beijing, China exceeds Pth, the address of the block is recorded, and the corresponding points are further partitioned at the next level. In the hardware implementation of sampling, 16 cores handle 16 blocks, respectively. A mask (see Fig. 23.4.3), in which “0” indicates RP, is used to skip redundant distance calculations between the already-generated RPs, thereby reducing the number of cycles in sampling. When evaluated on PointNeXt-Cls@ModelNet40, with the same number of blocks (e.g. 16 blocks), APS achieves a more balanced partition than [13], reducing the mean square error by 83.6%, improving the accuracy by 1.5%, and reducing the latency by 76.8%.
Adaptive Partition, Multi-Skipping, and Block-Wise Aggregation