⚡ 本页包含 AI 生成的分析内容,仅供参考
本文介绍了AMD Bulldozer x86-64核心中一个40条目的统一乱序调度器和整数执行单元,该设计在32nm SOI HKMG工艺下实现了每周期最多发出4个操作并支持单周期唤醒,同时通过减少FO4延迟超过20%来提升频率和性能。
AMD’s two-core Bulldozer module [1,2] implements the AMD x86-64 microarchitecture in an 11-layer 32-nm SOI HKMG technology. The 40-instruction outof-order unified integer scheduler issues up to four operations per cycle and supports single-cycle wake-up of dependent operations. The 2.37mm2 integer execution unit supports single-cycle data bypass among four independent functional units. Compared to previous AMD x86-64 cores [3-6], project goals reduce the number of FO4 inverter delays per cycle by more than 20%, while maintaining constant IPC, to achieve higher frequency and performance in the same power envelope, even with increased core counts. Critical paths (Fig 4.6.1) are implemented without exotic circuit techniques or heavy reliance on full-custom design. Dynamic logic appears only when required for density. Inputs and outputs of dynamic gates tolerate duty-cycle variation by flowing through timing test points if clock arrives early. In lieu of dynamic logic,
Michael Golden, Srikanth Arekapudi, James Vinh
AMD, Sunnyvale, CA