《Python编程实战》4.3 属性测试、性能回归与覆盖率门禁

给测试体系补上三块短板:用 hypothesis 的性质测试突破手写用例的输入盲区并靠反例缩小定位真实缺陷,用 pytest-benchmark 把性能退化成一道可对比的门禁,用 pytest-cov 的覆盖率与 --cov-fail-under 卡住底线。

本节目标:用 hypothesis 覆盖想不到的输入、用 pytest-benchmark 盯住性能退化、用 pytest-cov 把「测试够不够」变成一条可执行的红线。
适用版本:Python 3.12+(实测 3.14.6);hypothesis 6.168.5;pytest-benchmark 5.3.0;pytest-cov 7.1.0

4.3 属性测试、性能回归与覆盖率门禁

前两节把测试「搭起来了」:4.1 解决结构,4.2 解决集成。但一套测试写完,还有三个问题没答:输入覆盖够不够(手写用例永远只覆盖想到的)、性能有没有退化(功能全绿但慢了三倍)、底线在哪(谁来判断测试是不是太少了)。本节把这三块补齐,全部输出实测。

4.3.1 手写用例的天花板

手写测试的本质是「枚举你认为重要的输入」。问题在于你枚举的正是你已经想到的——真正的 bug 往往藏在想不到的组合里:空列表、恰好等于阈值、极大整数、重复元素、0 和负数。属性测试换了个思路:不写具体输入,而是描述对任意输入都应成立的「性质」,让机器去生成输入并尝试证伪。

4.3.2 hypothesis:用「性质」代替「用例」

以本节的 clamp 为例——把值夹到 [low, high] 区间。与其列举几组输入,不如声明三条性质:

import pytest
from hypothesis import assume, given
from hypothesis import strategies as st

from app.pricing import clamp

pytestmark = pytest.mark.unit


@given(st.integers(), st.integers(), st.integers())
def test_clamp_within_bounds(value, low, high):
    assume(low <= high)
    assert low <= clamp(value, low, high) <= high


@given(st.integers(), st.integers(), st.integers())
def test_clamp_idempotent(value, low, high):
    assume(low <= high)
    once = clamp(value, low, high)
    assert clamp(once, low, high) == once
  • st.integers() 生成任意整数(含负数、零、极大值)。
  • assume(low <= high) 告诉 hypothesis:只测合法区间,不满足的输入直接丢弃。
  • 第一条性质是值域约束,第二条是幂等性——两者都是「对一切输入成立」的陈述,而不是「对这几组输入成立」。

复杂对象用 @st.composite 把生成逻辑封装成一个可复用的策略:

@st.composite
def carts(draw):
    items = draw(st.lists(
        st.tuples(st.integers(min_value=0, max_value=10**6),
                  st.integers(min_value=0, max_value=100)),
        max_size=30,
    ))
    cart = Cart()
    for price, qty in items:
        cart.add("sku", price, qty)
    return cart


@given(carts())
def test_subtotal_non_negative(cart):
    assert cart.subtotal() >= 0


@given(carts(),
       st.floats(min_value=0, max_value=100, allow_nan=False, allow_infinity=False),
       st.floats(min_value=0, max_value=100, allow_nan=False, allow_infinity=False))
def test_discount_monotonic(cart, a, b):
    lo, hi = min(a, b), max(a, b)
    assert cart.discount(lo) >= cart.discount(hi)     # 折扣越大,折后价越低

默认每个用例生成 100 组输入,实测六个属性测试全绿:

tests/property/test_properties.py ......                                 [ 34%]

加 --hypothesis-show-statistics 还能看到生成了多少组、丢弃了多少:

tests/property/test_properties.py::test_clamp_within_bounds:
  - during generate phase (0.06 seconds):
    - 100 passing, 0 failing, and 37 invalid test cases
    - Events:
      * 27.01%, gave up because: failed to satisfy assume() ... (line 12)

注意 37 invalid test cases:assume 丢弃了 37 组不满足 low <= high 的输入。assume 用太多会让测试变慢甚至报 FailedHealthCheck——能改策略(如 st.integers().map(...))就别靠 assume 过滤。

4.3.3 反例缩小:一个真实缺陷的定位

属性测试最漂亮的能力是 shrinking(反例缩小):发现反例后,自动把它裁剪到最小的失败输入。看一个真实的边界缺陷——「满 5000 分免运费」的实现用了 > 而非 >=:

def shipping_fee(subtotal: int, threshold: int = 5000) -> int:
    return 0 if subtotal > threshold else 800     # 缺陷:应为 >=


@given(st.integers(min_value=0, max_value=20000),
       st.integers(min_value=1, max_value=20000))
def test_exactly_at_threshold_is_free(subtotal, threshold):
    if subtotal == threshold:
        assert shipping_fee(subtotal, threshold) == 0

实测失败输出,hypothesis 把反例缩到了 (1, 1):

subtotal = 1, threshold = 1

    def test_exactly_at_threshold_is_free(subtotal, threshold):
        if subtotal == threshold:
>           assert shipping_fee(subtotal, threshold) == 0
E           assert 800 == 0
E            +  where 800 = shipping_fee(1, 1)
E           Failing test case: test_exactly_at_threshold_is_free(
E               subtotal=1,
E               threshold=1,
E           )

如果没有 shrinking,报出来的可能是 subtotal=7842, threshold=7842 这种大数字,你还得自己回推边界。(1, 1) 一眼就指向「相等时未免运费」这个边界。这就是属性测试相比随机 fuzz 的核心价值:不只是找到 bug,还把它缩到你能读懂。

4.3.4 性能回归:pytest-benchmark

功能测试全绿不代表没有退化——一个「顺手」的重构可能让关键路径慢三倍,而没有任何测试会报警。pytest-benchmark 把「这段代码多快」变成可断言的基准:

def _make_cart(n: int) -> Cart:
    cart = Cart()
    for i in range(n):
        cart.add(f"sku-{i}", 100, 2)
    return cart


def test_bench_subtotal(benchmark):
    cart = _make_cart(1000)
    result = benchmark(cart.subtotal)
    assert result == 200_000      # 基准里同样可以断言正确性


def test_bench_discount(benchmark):
    cart = _make_cart(1000)
    result = benchmark(cart.discount, 15.0)
    assert result == 170_000

benchmark fixture 会自动决定轮数——先跑几轮校准,再跑足够多轮让统计稳定。实测输出(--benchmark-columns=min,mean,max,rounds):

Name (time in us)           Min               Mean                 Max            Rounds
----------------------------------------------------------------------------------------
test_bench_subtotal     54.6250 (1.0)      57.3869 (1.0)      538.5830 (1.92)      12258
test_bench_discount     54.7080 (1.00)     57.4606 (1.00)     280.2500 (1.0)       15239

两件事值得注意:轮数上万(Mean 才可信),以及 Max 远高于 Mean(GC 或调度抖动造成离群)。所以永远不要用 Max 当阈值,判断退化看 Mean 或 Median。

4.3.5 把基准接进 CI

基准只有「和谁比」才有意义。pytest-benchmark 的流程是先存基线、再对比:

# 1. 在稳定分支上存基线(写入 .benchmarks/)
pytest tests/perf --benchmark-autosave

# 2. 在改动分支上对比,Mean 退化超过 10% 即失败
pytest tests/perf --benchmark-compare --benchmark-compare-fail=mean:10%

对比结果会把基线与当前并排显示:

Name (time in us)                    Min            Mean           Rounds
test_bench_subtotal (0001_unversi)  55.0420 (1.0)   58.0744 (1.01)    16140
test_bench_subtotal (NOW)           55.0830 (1.00)  58.6422 (1.02)    16217

本机实测这套流程 exit=0(没有超过 10% 的退化)。--benchmark-compare-fail 的阈值要留出噪声余量——设成 mean:2% 会因为 CI 机器抖动频繁误报,mean:10%~20% 是更务实的起点。基线的 .benchmarks/ 目录应纳入缓存或 artifact,避免每次从零开始。

4.3.6 覆盖率:pytest-cov

覆盖率回答的是「哪些代码从来没被执行过」。pytest-cov 7.1.0 把 coverage 7.16.2 接进 pytest:

pytest --cov=app --cov-report=term-missing

实测输出(Missing 列精确到行号):

Name              Stmts   Miss  Cover   Missing
-----------------------------------------------
app/__init__.py       0      0   100%
app/db.py            24      0   100%
app/pricing.py       29      2    93%   40, 46
-----------------------------------------------
TOTAL                53      2   96%

Missing 40, 46 直接指到未覆盖的两行(Cart.total 与 clamp 的非法区间分支)。覆盖率配置写进 pyproject.toml:

[tool.coverage.run]
branch = true
source = ["app"]

[tool.coverage.report]
show_missing = true
fail_under = 90
exclude_lines = [
    "pragma: no cover",
    "if __name__ == .__main__.:",
    "raise NotImplementedError",
]

4.3.7 覆盖率门禁:–cov-fail-under

覆盖率最大的价值不是那个百分比,而是它能当门禁。--cov-fail-under=90 让低于阈值时 pytest 返回非零退出码:

pytest --cov=app --cov-fail-under=90
TOTAL                53      2   96%
Required test coverage of 90% reached. Total coverage: 96.23%

当前 96% 通过。把阈值抬到 98%,门禁立刻亮红灯——实测退出码为 1:

TOTAL                53      2   96%
FAIL Required test coverage of 98% not reached. Total coverage: 96.23%

这正是门禁该有的行为:CI 里这一条命令就能挡住「新增代码但没加测试」的 PR。要点:

  • 阈值要可达成且渐进。一上来设 95% 会让老项目寸步难行;先按现状定,再逐步抬高。
  • 门禁卡的是总覆盖率,容易掩盖「新代码 0 覆盖、老代码高覆盖」的稀释。更严的做法是只对改动行设阈值(diff-cover 一类工具,本机未实测)。
  • 门禁是底线不是目标。为了凑数字写无断言的测试,只会制造维护负担。

4.3.8 分支覆盖率与合理的排除

行覆盖有盲区:if x > 0 只要执行过这一行就算覆盖,真假两个分支是否都走过它不管。branch = true 补上这一课,实测:

Name              Stmts   Miss Branch BrPart  Cover   Missing
-------------------------------------------------------------
app/pricing.py       29      2      6      1    91%   40, 46
TOTAL                53      2      6      1    95%

多出的 Branch/BrPart 两列专门盯「分支未走全」。有些代码合理地不该被覆盖,用 exclude_lines 显式排除,比为了数字硬写测试更诚实:类型检查用的 if TYPE_CHECKING:、raise NotImplementedError、if __name__ == "__main__":。排除必须写进配置、可审查,而不是悄悄加 # pragma: no cover 蒙混过关。

4.3.9 三层门禁的取舍

三种手段各管一段,组合起来才完整:

手段拦什么成本误报风险
hypothesis想不到的输入 / 边界中(生成耗时)低(性质写错才会)
pytest-benchmark性能退化高(上万轮)中(机器抖动)
pytest-cov 门禁测试太少低低(但会诱发凑数)

务实组合:属性测试用于纯函数与不变量(定价、解析、状态机),基准只测关键路径(别给每个函数都上 benchmark),覆盖率门禁守住总底线并逐步抬高。三者都不该喧宾夺主——测试的目的是让你敢改,不是让数字好看。

小结

  • 手写用例只能覆盖「已想到的输入」;hypothesis 用性质描述对一切输入成立的条件,机器负责找反例。
  • assume 过滤会丢弃输入、拖慢测试,能用策略表达就别用它。
  • shrinking 把反例缩到最小(实测 (1, 1)),让边界缺陷一眼可读。
  • pytest-benchmark 自动校准轮数;判断退化看 Mean/Median,绝不用 Max;CI 对比用 --benchmark-compare-fail,阈值留噪声余量。
  • 覆盖率门禁 --cov-fail-under 是底线不是目标;开 branch = true 补行覆盖盲区,排除项要写进配置可审查。

测试体系到这里就完整了:结构(4.1)、集成(4.2)、深度与门禁(4.3)。接下来进入第二部分 Web 后端工程——下一章把这些测试能力用在真实服务上,从 FastAPI 应用结构与依赖注入 开始搭第一个可测、可部署的 HTTP 服务。

延伸阅读:Python 测试与质量工程 、《Python编程入门》14.3 覆盖率、ruff/mypy 与 CI 门禁 。

阅读导航:上一节:集成测试与容器化测试环境 · 下一节:FastAPI 应用结构与依赖注入 。

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「python」更多文章

  1. 《Python高级编程》目录
  2. 《Python高级编程》11.3 PEP 流程与版本迁移策略
  3. 《Python高级编程》11.2 嵌入式与自由线程运行时