← 返回 JSSC 论文列表JSSC 2012第2期Digital Circuits0.13μm
MRTP: Mobile Ray Tracing Processor With Reconfigurable Stream Multi-Processors for High Datapath Utilization Hong-Y un Kim,S t u d e n tM e m b e r ,I E E E, Y oung-Jun Kim , Student Member , IEEE ,a n d
提出一种可重构流多处理器移动光线追踪处理器,提高数据路径利用率。
0.13μm CMOS, 100MHz, 673K rays/s, 156mW
光线追踪可重构处理器SIMT架构数据路径利用率表加载器
▸可重构流多处理器(RSMP)支持MIMD和SIMT模式
▸通过表加载器(TBLD)减少LUT访问冲突
▸部分特殊功能单元(PSFU)降低硬件开销
Abstract
This paper presents a mobile ray tracing processor (MRTP) with recon figurable stream multi-processors (RSMPs) for high datapath utilization. The MRTP includes three RSMPs that operate in multiple instruction multiple d ata (MIMD) mode asynchronously to exploit instr uction-level parallelism. Each RSMP is based on single instruction multiple thread (SIMT) architecture to exploit thread-level paral lelism. An RSMP consists of twelve scalar processing elements (SPEs) that run multiple threads in parallel synchronously: twelve scalar threads or four vector threads depending on an operati ng mode. A low datapath utilization caused by a branch di vergence in SIMT architecture is improved by 19.9% on average by recon figuring twelve SPEs between scalar SIMT and vector SIMT w ith 0.1% area overheads. Special function instructions occupy only 2% 8% of kernel instructions so that a partial spec ial function unit (PSFU) is imple- mented instead of a large dedic ated SFU. The access con flicts with a look-up table (LUT) caused by c oncurrent accesses of twelve SPEs are reduced by a table l oader (TBLD). The TBLD monitors concurrent requests from tw elve SPEs and red uces an access count to LUT by distributing a coef ficient to multiple SPEs with only one read-access to LUT. MRTP with area of mm has been fabricated in 0.13 mC MOS technology. MRTP achieves a peak performance of 673 K rays per second while consuming 156 mW at 100 MHz with .