<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>E5M2 on PlumePHP</title><link>https://plumephp.com/tags/e5m2/</link><description>Recent content in E5M2 on PlumePHP</description><generator>Hugo</generator><language>zh-CN</language><lastBuildDate>Mon, 28 Sep 2026 12:00:00 +0800</lastBuildDate><atom:link href="https://plumephp.com/tags/e5m2/index.xml" rel="self" type="application/rss+xml"/><item><title>FP8 推理：8 位浮点量化、E4M3/E5M2 与 NVIDIA FP8 生态</title><link>https://plumephp.com/ai-fp8-inference/</link><pubDate>Mon, 28 Sep 2026 12:00:00 +0800</pubDate><guid>https://plumephp.com/ai-fp8-inference/</guid><description>&lt;p&gt;FP16 存不下的大模型，很多人第一反应是「降到 INT8」。但 INT8 动态范围窄，激活（activation）的数值分布宽，量化后精度损失明显。FP8 提供了一个「中间地带」：&lt;strong&gt;8 位浮点，动态范围接近 FP16，显存带宽比 FP16 减半&lt;/strong&gt;，而且在 H100/GH200 上有原生硬件支持。本文把 FP8 讲透：格式长什么样、和 INT8 怎么选、量化流程怎么做、NVIDIA 生态怎么用，以及生产里哪些模型该用 FP8、哪些不该。&lt;/p&gt;</description></item></channel></rss>