← 返回论文列表 📄 下载原文 PDF  ISSCC 2021 · 3.2
ISSCC 2021Session 3 · HIGHLIGHTED CHIP RELEASES: MODERN DIGITAL SoCsDigital Processors

The A100 Datacenter GPU and Ampere Architecture

⚡ 本页包含 AI 生成的分析内容,仅供参考

📋 论文概要

该论文介绍了Nvidia A100数据中心GPU及其Ampere架构,针对现代云数据中心中AI深度学习、数据分析、科学计算等多样化计算密集型应用进行优化。通过引入第三代Tensor Core支持细粒度稀疏性、新数据类型(BF16、TF32、FP64)以及多实例GPU(MIG)实现高效scale-out,显著提升了计算性能和资源利用率。

💡 主要创新点

重要性
发表年份
ISSCC 2021

🏷 关键词

GPU数据中心Tensor Core稀疏性多实例GPU

📄 原文摘要

Nvidia, Santa Clara, CA The diversity of compute-intensive applications in modern cloud data centers has driven the explosion of GPU-accelerated cloud computing. Such applications include AI deep learning training and inference, data analytics, scientific computing, genomics, edge video analytics and 5G services, graphics rendering, and cloud gaming. The A100 GPU introduces several features targeting these workloads: a 3rd-generation Tensor Core with support for fine-grained sparsity, new BFloat16 (BF16), TensorFloat-32 (TF32), and FP64 datatypes, scale-out support with multi-instance GPU (MIG) virtualization, and scale-up support with a 3rd-generation 50Gbps NVLink I/O interface (NVLink3) and NVSwitch inter-GPU communication. As shown in Fig. 3.2.1, A100 contains 108 Streaming Multiprocessors (SMs) and 6912 CUDA cores. The SMs are fed by a 40MB L2 cache and 1.56TB/s of HBM2 memory bandwidth (BW). At 1.41GHz, A100 provides an effective

👥 作者与机构

Jack Choquette, Edward Lee, Ronny Krashinsky, Vishnu Balan, Brucek Khailany

分类:Digital Processors · 年份:ISSCC 2021