本节目标:用 hypothesis 覆盖想不到的输入、用 pytest-benchmark 盯住性能退化、用 pytest-cov 把「测试够不够」变成一条可执行的红线。
适用版本:Python 3.12+(实测 3.14.6);hypothesis 6.168.5;pytest-benchmark 5.3.0;pytest-cov 7.1.0
4.3 属性测试、性能回归与覆盖率门禁
前两节把测试「搭起来了」:4.1 解决结构,4.2 解决集成。但一套测试写完,还有三个问题没答:输入覆盖够不够(手写用例永远只覆盖想到的)、性能有没有退化(功能全绿但慢了三倍)、底线在哪(谁来判断测试是不是太少了)。本节把这三块补齐,全部输出实测。
4.3.1 手写用例的天花板
手写测试的本质是「枚举你认为重要的输入」。问题在于你枚举的正是你已经想到的——真正的 bug 往往藏在想不到的组合里:空列表、恰好等于阈值、极大整数、重复元素、0 和负数。属性测试换了个思路:不写具体输入,而是描述对任意输入都应成立的「性质」,让机器去生成输入并尝试证伪。
4.3.2 hypothesis:用「性质」代替「用例」
以本节的 clamp 为例——把值夹到 [low, high] 区间。与其列举几组输入,不如声明三条性质:
import pytest
from hypothesis import assume, given
from hypothesis import strategies as st
from app.pricing import clamp
pytestmark = pytest.mark.unit
@given(st.integers(), st.integers(), st.integers())
def test_clamp_within_bounds(value, low, high):
assume(low <= high)
assert low <= clamp(value, low, high) <= high
@given(st.integers(), st.integers(), st.integers())
def test_clamp_idempotent(value, low, high):
assume(low <= high)
once = clamp(value, low, high)
assert clamp(once, low, high) == once
st.integers()生成任意整数(含负数、零、极大值)。assume(low <= high)告诉 hypothesis:只测合法区间,不满足的输入直接丢弃。- 第一条性质是值域约束,第二条是幂等性——两者都是「对一切输入成立」的陈述,而不是「对这几组输入成立」。
复杂对象用 @st.composite 把生成逻辑封装成一个可复用的策略:
@st.composite
def carts(draw):
items = draw(st.lists(
st.tuples(st.integers(min_value=0, max_value=10**6),
st.integers(min_value=0, max_value=100)),
max_size=30,
))
cart = Cart()
for price, qty in items:
cart.add("sku", price, qty)
return cart
@given(carts())
def test_subtotal_non_negative(cart):
assert cart.subtotal() >= 0
@given(carts(),
st.floats(min_value=0, max_value=100, allow_nan=False, allow_infinity=False),
st.floats(min_value=0, max_value=100, allow_nan=False, allow_infinity=False))
def test_discount_monotonic(cart, a, b):
lo, hi = min(a, b), max(a, b)
assert cart.discount(lo) >= cart.discount(hi) # 折扣越大,折后价越低
默认每个用例生成 100 组输入,实测六个属性测试全绿:
tests/property/test_properties.py ...... [ 34%]
加 --hypothesis-show-statistics 还能看到生成了多少组、丢弃了多少:
tests/property/test_properties.py::test_clamp_within_bounds:
- during generate phase (0.06 seconds):
- 100 passing, 0 failing, and 37 invalid test cases
- Events:
* 27.01%, gave up because: failed to satisfy assume() ... (line 12)
注意 37 invalid test cases:assume 丢弃了 37 组不满足 low <= high 的输入。assume 用太多会让测试变慢甚至报 FailedHealthCheck——能改策略(如 st.integers().map(...))就别靠 assume 过滤。
4.3.3 反例缩小:一个真实缺陷的定位
属性测试最漂亮的能力是 shrinking(反例缩小):发现反例后,自动把它裁剪到最小的失败输入。看一个真实的边界缺陷——「满 5000 分免运费」的实现用了 > 而非 >=:
def shipping_fee(subtotal: int, threshold: int = 5000) -> int:
return 0 if subtotal > threshold else 800 # 缺陷:应为 >=
@given(st.integers(min_value=0, max_value=20000),
st.integers(min_value=1, max_value=20000))
def test_exactly_at_threshold_is_free(subtotal, threshold):
if subtotal == threshold:
assert shipping_fee(subtotal, threshold) == 0
实测失败输出,hypothesis 把反例缩到了 (1, 1):
subtotal = 1, threshold = 1
def test_exactly_at_threshold_is_free(subtotal, threshold):
if subtotal == threshold:
> assert shipping_fee(subtotal, threshold) == 0
E assert 800 == 0
E + where 800 = shipping_fee(1, 1)
E Failing test case: test_exactly_at_threshold_is_free(
E subtotal=1,
E threshold=1,
E )
如果没有 shrinking,报出来的可能是 subtotal=7842, threshold=7842 这种大数字,你还得自己回推边界。(1, 1) 一眼就指向「相等时未免运费」这个边界。这就是属性测试相比随机 fuzz 的核心价值:不只是找到 bug,还把它缩到你能读懂。
4.3.4 性能回归:pytest-benchmark
功能测试全绿不代表没有退化——一个「顺手」的重构可能让关键路径慢三倍,而没有任何测试会报警。pytest-benchmark 把「这段代码多快」变成可断言的基准:
def _make_cart(n: int) -> Cart:
cart = Cart()
for i in range(n):
cart.add(f"sku-{i}", 100, 2)
return cart
def test_bench_subtotal(benchmark):
cart = _make_cart(1000)
result = benchmark(cart.subtotal)
assert result == 200_000 # 基准里同样可以断言正确性
def test_bench_discount(benchmark):
cart = _make_cart(1000)
result = benchmark(cart.discount, 15.0)
assert result == 170_000
benchmark fixture 会自动决定轮数——先跑几轮校准,再跑足够多轮让统计稳定。实测输出(--benchmark-columns=min,mean,max,rounds):
Name (time in us) Min Mean Max Rounds
----------------------------------------------------------------------------------------
test_bench_subtotal 54.6250 (1.0) 57.3869 (1.0) 538.5830 (1.92) 12258
test_bench_discount 54.7080 (1.00) 57.4606 (1.00) 280.2500 (1.0) 15239
两件事值得注意:轮数上万(Mean 才可信),以及 Max 远高于 Mean(GC 或调度抖动造成离群)。所以永远不要用 Max 当阈值,判断退化看 Mean 或 Median。
4.3.5 把基准接进 CI
基准只有「和谁比」才有意义。pytest-benchmark 的流程是先存基线、再对比:
# 1. 在稳定分支上存基线(写入 .benchmarks/)
pytest tests/perf --benchmark-autosave
# 2. 在改动分支上对比,Mean 退化超过 10% 即失败
pytest tests/perf --benchmark-compare --benchmark-compare-fail=mean:10%
对比结果会把基线与当前并排显示:
Name (time in us) Min Mean Rounds
test_bench_subtotal (0001_unversi) 55.0420 (1.0) 58.0744 (1.01) 16140
test_bench_subtotal (NOW) 55.0830 (1.00) 58.6422 (1.02) 16217
本机实测这套流程 exit=0(没有超过 10% 的退化)。--benchmark-compare-fail 的阈值要留出噪声余量——设成 mean:2% 会因为 CI 机器抖动频繁误报,mean:10%~20% 是更务实的起点。基线的 .benchmarks/ 目录应纳入缓存或 artifact,避免每次从零开始。
4.3.6 覆盖率:pytest-cov
覆盖率回答的是「哪些代码从来没被执行过」。pytest-cov 7.1.0 把 coverage 7.16.2 接进 pytest:
pytest --cov=app --cov-report=term-missing
实测输出(Missing 列精确到行号):
Name Stmts Miss Cover Missing
-----------------------------------------------
app/__init__.py 0 0 100%
app/db.py 24 0 100%
app/pricing.py 29 2 93% 40, 46
-----------------------------------------------
TOTAL 53 2 96%
Missing 40, 46 直接指到未覆盖的两行(Cart.total 与 clamp 的非法区间分支)。覆盖率配置写进 pyproject.toml:
[tool.coverage.run]
branch = true
source = ["app"]
[tool.coverage.report]
show_missing = true
fail_under = 90
exclude_lines = [
"pragma: no cover",
"if __name__ == .__main__.:",
"raise NotImplementedError",
]
4.3.7 覆盖率门禁:–cov-fail-under
覆盖率最大的价值不是那个百分比,而是它能当门禁。--cov-fail-under=90 让低于阈值时 pytest 返回非零退出码:
pytest --cov=app --cov-fail-under=90
TOTAL 53 2 96%
Required test coverage of 90% reached. Total coverage: 96.23%
当前 96% 通过。把阈值抬到 98%,门禁立刻亮红灯——实测退出码为 1:
TOTAL 53 2 96%
FAIL Required test coverage of 98% not reached. Total coverage: 96.23%
这正是门禁该有的行为:CI 里这一条命令就能挡住「新增代码但没加测试」的 PR。要点:
- 阈值要可达成且渐进。一上来设 95% 会让老项目寸步难行;先按现状定,再逐步抬高。
- 门禁卡的是总覆盖率,容易掩盖「新代码 0 覆盖、老代码高覆盖」的稀释。更严的做法是只对改动行设阈值(
diff-cover一类工具,本机未实测)。 - 门禁是底线不是目标。为了凑数字写无断言的测试,只会制造维护负担。
4.3.8 分支覆盖率与合理的排除
行覆盖有盲区:if x > 0 只要执行过这一行就算覆盖,真假两个分支是否都走过它不管。branch = true 补上这一课,实测:
Name Stmts Miss Branch BrPart Cover Missing
-------------------------------------------------------------
app/pricing.py 29 2 6 1 91% 40, 46
TOTAL 53 2 6 1 95%
多出的 Branch/BrPart 两列专门盯「分支未走全」。有些代码合理地不该被覆盖,用 exclude_lines 显式排除,比为了数字硬写测试更诚实:类型检查用的 if TYPE_CHECKING:、raise NotImplementedError、if __name__ == "__main__":。排除必须写进配置、可审查,而不是悄悄加 # pragma: no cover 蒙混过关。
4.3.9 三层门禁的取舍
三种手段各管一段,组合起来才完整:
| 手段 | 拦什么 | 成本 | 误报风险 |
|---|---|---|---|
| hypothesis | 想不到的输入 / 边界 | 中(生成耗时) | 低(性质写错才会) |
| pytest-benchmark | 性能退化 | 高(上万轮) | 中(机器抖动) |
| pytest-cov 门禁 | 测试太少 | 低 | 低(但会诱发凑数) |
务实组合:属性测试用于纯函数与不变量(定价、解析、状态机),基准只测关键路径(别给每个函数都上 benchmark),覆盖率门禁守住总底线并逐步抬高。三者都不该喧宾夺主——测试的目的是让你敢改,不是让数字好看。
小结
- 手写用例只能覆盖「已想到的输入」;hypothesis 用性质描述对一切输入成立的条件,机器负责找反例。
assume过滤会丢弃输入、拖慢测试,能用策略表达就别用它。- shrinking 把反例缩到最小(实测
(1, 1)),让边界缺陷一眼可读。 - pytest-benchmark 自动校准轮数;判断退化看
Mean/Median,绝不用Max;CI 对比用--benchmark-compare-fail,阈值留噪声余量。 - 覆盖率门禁
--cov-fail-under是底线不是目标;开branch = true补行覆盖盲区,排除项要写进配置可审查。
测试体系到这里就完整了:结构(4.1)、集成(4.2)、深度与门禁(4.3)。接下来进入第二部分 Web 后端工程——下一章把这些测试能力用在真实服务上,从 FastAPI 应用结构与依赖注入 开始搭第一个可测、可部署的 HTTP 服务。
延伸阅读:Python 测试与质量工程 、《Python编程入门》14.3 覆盖率、ruff/mypy 与 CI 门禁 。
阅读导航:上一节:集成测试与容器化测试环境 · 下一节:FastAPI 应用结构与依赖注入 。
继续阅读
探索更多技术文章
浏览归档,发现更多关于系统设计、工具链和工程实践的内容。