⚡ 本页包含 AI 生成的分析内容,仅供参考
本文介绍了AMD Bulldozer 32nm SOI 8核CPU中2核模块的设计方案,重点解决了整数数据通路的0周期结果旁路、关键路径线延迟消除等问题。通过复制物理寄存器文件阵列和地址生成器增量器,以及采用双轨读取关键位等方法,提升了处理器性能。
Michael Golden2, Scott Hilker2, Aaron Horiuchi1, Kevin A. Hurd1, Dave Johnson1, Hugh McIntyre2, Samuel Naffziger1, James Vinh2, Jonathan White4, Kathryn Wilcox4 instead of a traditional mismatch CAM. The integer datapath supports 0-cycle result bypass to dependent instructions. To remove critical-path wire delay, the physical register file arrays and address generator (AGEN) incrementor are replicated. Most register file read bits are single-ended, but dual-rail reads for a few critical bits supply clocked data to the dynamic shifter. 2 The floating-point (FP) unit adds many features within the same basic pipeline as previous AMD x86-64 processors: a) new multiply-accumulate functionality and instructions, including AVX, AES, SSSE3, SSE4.1, SSE4.2, XOP, and PCLMULQDQ; b) dual-threaded support; c) 4-wide instruction issue; and, d) increased size of performance-critical structures [2, 4]. AMD’s 2-core “Bulldozer” module contains 213 million transistors in an 11metal layer 32nm HKMG SOI C
Tim Fischer1, Srikanth Arekapudi2, Eric Busta1, Carl Dietz3,