跳到正文
原文
Hugging Face·· 2024-05-16AI 评分54

Hugging Face Transformers 正式支持 KV Cache 量化功能

Unlocking Longer Generation with Key-Value Cache Quantization

AI 导读

Hugging Face 在 Transformers 库中集成了 KV Cache 量化功能,可大幅降低大语言模型长文本生成时的显存占用。该功能基于 quanto 与 HQQ 后端支持 int2、int4 和 int8 精度,其中 int4 精度可节省约 2.5 倍显存且模型困惑度与原版 fp16 相当。

来源:Hugging Face · huggingface.co