⚡ 本页包含 AI 生成的分析内容,仅供参考
该论文提出了一款基于28nm工艺的稀疏Transformer加速器,通过内存内蝶形零跳过单元实现非结构化剪枝神经网络的高效计算,解决了Transformer中自注意力机制计算量大且稀疏性利用不足的问题。
Transformer networks, from BERT, GPT to Alphafold, have demonstrated unprecedented advances in a variety of AI tasks. Fig. 16.2.1 shows the computing flow of self-attention – the fundamental operation in transformers. Queries (Q), keys (K) and values (V) are first obtained by multiplying inputs with 3 weight matrices. Afterward, scores that evaluate Q-K relevance are computed as scaled dot products and converted to probabilities through the softmax function. The probabilities are then multiplied by V, generating the final self-attention results. Transformer networks have led to an explosion in parameter counts, for example, 175B parameters for GPT-3. This demands significant growth in computing hardware and memory. Owing to expanding network sizes and
Shiwei Liu1, Peizhe Li1, Jinshan Zhang1, Yunzhengmao Wang1, Haozhe Zhu1,
Wenning Jiang1, Shan Tang2, Chixiao Chen1,3, Qi Liu1, Ming Liu1 Fudan University, Shanghai, China BIRENTECH, Shanghai, China 3 Peng Cheng Laboratory, Shenzhen, China 1 2