<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Draft Model on PlumePHP</title><link>https://plumephp.com/tags/draft-model/</link><description>Recent content in Draft Model on PlumePHP</description><generator>Hugo</generator><language>zh-CN</language><lastBuildDate>Mon, 28 Sep 2026 10:00:00 +0800</lastBuildDate><atom:link href="https://plumephp.com/tags/draft-model/index.xml" rel="self" type="application/rss+xml"/><item><title>投机采样（Speculative Decoding）：草案模型、并行验证与加速原理</title><link>https://plumephp.com/ai-speculative-decoding/</link><pubDate>Mon, 28 Sep 2026 10:00:00 +0800</pubDate><guid>https://plumephp.com/ai-speculative-decoding/</guid><description>&lt;p&gt;LLM 自回归解码的根本瓶颈是&lt;strong&gt;串行&lt;/strong&gt;：每一步只能生成一个 token，而 Decode 阶段的计算密度极低（带宽受限）。投机采样（Speculative Decoding）换了个思路：&lt;strong&gt;让一个小草案模型替你「猜」后面的几个 token，大模型一次并行验证这一串&lt;/strong&gt;——猜对了就一次接受多个 token，用「一次并行验证」换「多步串行生成」。本文把投机采样的原理、实现路径（草拟-验证、Medusa、EAGLE）、正确性保证、工程取舍与踩坑讲透。&lt;/p&gt;</description></item></channel></rss>