<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>显存优化 on PlumePHP</title><link>https://plumephp.com/tags/%E6%98%BE%E5%AD%98%E4%BC%98%E5%8C%96/</link><description>Recent content in 显存优化 on PlumePHP</description><generator>Hugo</generator><language>zh-CN</language><lastBuildDate>Mon, 28 Sep 2026 14:00:00 +0800</lastBuildDate><atom:link href="https://plumephp.com/tags/%E6%98%BE%E5%AD%98%E4%BC%98%E5%8C%96/index.xml" rel="self" type="application/rss+xml"/><item><title>LLM 推理性能调优实战：从瓶颈定位到端到端优化的检查清单</title><link>https://plumephp.com/ai-performance-tuning-checklist/</link><pubDate>Mon, 28 Sep 2026 14:00:00 +0800</pubDate><guid>https://plumephp.com/ai-performance-tuning-checklist/</guid><description>&lt;p&gt;模型是同一份，部署的人不同，性能可以差一个数量级。LLM 推理性能调优最忌讳「盲人摸象」——一会儿调这个参数、一会儿换个引擎，没有定位就乱调。本文给出一份&lt;strong&gt;可执行的调优检查清单&lt;/strong&gt;：先明确优化目标（吞吐还是延迟），再按「显存→计算→带宽→调度」逐层定位瓶颈，最后按优先级上优化手段，并用基准测试闭环验证。照着清单走，性能问题能快速收敛。&lt;/p&gt;</description></item></channel></rss>