⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文介绍了特斯拉DOJO超算系统中的D1机器学习训练处理器,采用7nm工艺和波时钟分布技术,实现了354个计算节点和440MB分布式SRAM的高性能、可扩展架构,旨在解决自动驾驶AI训练的高算力需求。
Ram Bharti1, Derek Carson1, Anton Lawrendra1, Vineet Mudgal1, Vivek Santhosh1, Sunil Shukla2, Te-Chen Tsai1 Tesla, Palo Alto, CA Tesla, Austin, TX 1 2 D1 is the ML training processor in the DOJO exa-scale computer system [1,2]. DOJO applications include training ML networks that run on the Tesla Full Self-Driving (FSD) computer in Tesla vehicles [3,4]. D1 chip design goals were high performance, wide operating range, and scalable and modular construction featuring reduced design mismatch and high block-reuse. The D1 processor has 354 compute nodes, 440 MB of distributed SRAM, a 2-dimensional high-bandwidth on-die network-on-chip (NoC), and 576 lanes of high-speed SerDes, providing a total of 362 TFlops of BFP16/CFP8 performance. D1 is implemented in TSMC 7nm process with 13 layers of copper interconnect and 1 layer of aluminum RDL. The D1 die contains 50B transistors in 645mm2 area, has a 400W TDP, and its core clock
Tim C. Fischer1, Anantha Kumar Nivarti1, Raghuvir Ramachandran1,