《Python编程入门》17.1 文件批处理与办公自动化

用 pathlib 与 shutil 把散落的文件批量重命名、按扩展名归类、安全删除,掌握 zipfile/tarfile 的解压炸弹与路径穿越防护,再用 openpyxl 读写带公式与样式的 Excel,最后用 subprocess 与轮询式监听接进日常工作流。

本节目标:把 10.2 节学过的 pathlib 用起来,写出「先 dry-run 再执行」的文件批处理脚本,安全地压缩解压,并用 openpyxl 读写 Excel。
适用版本:Python 3.12+(实测 3.14.6)

17.1 文件批处理与办公自动化

10.2 节我们学过 Path 的读写与遍历,但那是「操作单个文件」的视角。真实工作里更常见的是成批的琐事:几百张图片要按扩展名分文件夹、每月报表要重命名、一批压缩包要解开、一张 Excel 要更新公式。这些活手点鼠标能耗掉一下午,写成脚本却能一劳永逸。本节就把它做成可复用的套路。

17.1.1 pathlib 是批处理的起点

批处理的第一步永远是「列出目录里有什么」。Path 提供了三个层次的方法:

方法作用
iterdir()直接子项(文件 + 目录),不递归
glob(pat)按模式匹配当前目录
rglob(pat)按模式匹配整棵子树(递归)

先造一批杂乱的测试文件,再看它们的区别:

from pathlib import Path
import shutil

ROOT = Path("/tmp/python_book/scratch/ch17/lab1")
shutil.rmtree(ROOT, ignore_errors=True)
ROOT.mkdir(parents=True)
for n in ["notes.txt", "todo.txt", "data.csv", "logo.png", "banner.PNG"]:
    (ROOT / n).write_text("x", encoding="utf-8")
(ROOT / "downloads" / "deep").mkdir(parents=True)
(ROOT / "downloads" / "photo1.jpg").write_text("x", encoding="utf-8")
(ROOT / "downloads" / "deep" / "photo2.jpg").write_text("x", encoding="utf-8")

17.1.2 glob 与 rglob:模式匹配的边界

print(sorted(p.name for p in ROOT.glob("*.txt")))          # 只顶层
# ['notes.txt', 'todo.txt']
print(sorted(str(p.relative_to(ROOT)) for p in ROOT.rglob("*.jpg")))  # 递归
# ['downloads/deep/photo2.jpg', 'downloads/photo1.jpg']

三个容易踩的边界:

  • glob 只扫当前目录,子目录里的 .txt 不会出现;要递归必须用 rglob。
  • glob("*") 不含点开头的隐藏文件,要匹配它们得写 glob(".*")。
  • glob("*/*") 只下钻一层,模式里的 / 就是目录层级,不是「任意深度」。

rglob("*.jpg") 等价于 glob("**/*.jpg")。** 表示「任意层级」,所以 ROOT.glob("**/*.txt") 与 ROOT.rglob("*.txt") 结果一致。

17.1.3 实战:批量归类(先 dry-run 再执行)

批量操作的第一条铁律:先把「打算做什么」打印出来,人眼确认后再真做。这个模式叫 dry-run。下面按扩展名把文件收进同名子目录:

def plan_organize(root: Path):
    moves = []
    for p in sorted(root.iterdir()):
        if not p.is_file():
            continue
        ext = p.suffix.lower().lstrip(".") or "noext"
        moves.append((p, root / ext / p.name))
    return moves

def apply_moves(moves, dry_run=True):
    for src, dst in moves:
        if dry_run:
            print(f"[dry-run] {src.name} -> {dst.relative_to(src.parent)}")
        else:
            dst.parent.mkdir(parents=True, exist_ok=True)
            src.replace(dst)   # Path.replace 就是原子重命名/移动

moves = plan_organize(ROOT)
apply_moves(moves, dry_run=True)     # 第一步:只看不动
# [dry-run] a.txt -> txt/a.txt   ...   [dry-run] f.py -> py/f.py
apply_moves(moves, dry_run=False)    # 第二步:确认后执行

注意 e.PNG 与 d.png 都归进了 png/:.suffix.lower() 先把大小写抹平,否则 .PNG 和 .png 会分成两个目录。plan_organize 只返回「计划」,apply_moves 负责执行——把「决策」和「副作用」拆开,是安全批处理的核心设计。

17.1.4 批量重命名与安全删除

重命名同理,先算出新名字再统一 rename:

def new_name(p: Path) -> str:
    # report_2026_01.xlsx -> 2026-01-report.xlsx
    parts = p.stem.split("_")
    if len(parts) == 3 and parts[1].isdigit():
        return f"{parts[1]}-{parts[2]}-{parts[0]}{p.suffix}"
    return p.name

for p in sorted(rename_dir.iterdir()):
    print(f"{p.name} -> {new_name(p)}")
    p.rename(p.with_name(new_name(p)))
    # report_2026_01.xlsx -> 2026-01-report.xlsx

删除比改名更不可逆,所以更要 dry-run。先扫出待删清单,打印大小,再 unlink:

def find_temp_files(root: Path):
    return sorted(p for p in root.rglob("*.tmp") if p.is_file())

for p in find_temp_files(ROOT):          # dry-run
    print("[dry-run] 将删除:", p.relative_to(ROOT), f"({p.stat().st_size} 字节)")
    # [dry-run] 将删除: tmp/junk1.tmp (1 字节)
for p in find_temp_files(ROOT):          # 确认后
    p.unlink()

17.1.5 shutil:复制、移动、归档与 rmtree 的危险

Path 只处理「一个文件」的增删改,整棵树的操作交给 shutil:

函数作用注意
shutil.copy(src, dst)复制文件(内容 + 权限)dst 目录必须已存在
shutil.copytree(src, dst)递归复制整棵树dst 必须不存在
shutil.move(src, dst)移动 / 重命名跨设备时退化为复制 + 删除
shutil.rmtree(path)递归删除整棵树无法撤销
shutil.make_archive(base, fmt, root_dir)打包成 zip/tar返回归档文件路径

rmtree 是最危险的一个函数,因为它删除的是路径参数指向的一切。危险往往不在函数本身,而在「路径是怎么拼出来的」:

evil = "../../../etc"
naive = Path("/tmp/python_book/scratch/ch17/lab2/dst/" + evil)
print(naive.resolve())     # 解析到 /private/tmp/python_book/scratch/etc

字符串拼接不会阻止 .. 向上跳。永远不要用字符串拼路径,用 Path 组合,并在删除前校验它仍在允许的根之内:

def safe_rmtree(target: Path, allowed_root: Path) -> None:
    resolved, root = target.resolve(), allowed_root.resolve()
    if not resolved.is_relative_to(root) or resolved == root:
        raise ValueError(f"拒绝删除根之外的路径: {resolved}")
    shutil.rmtree(resolved)

Path.is_relative_to() 判断一个路径是否在另一个之下,正是这类守卫的标准写法。

17.1.6 zipfile / tarfile:读写与路径穿越防护

压缩包读写用标准库即可:

import zipfile
with zipfile.ZipFile("out.zip", "w", zipfile.ZIP_DEFLATED) as zf:
    zf.writestr("readme.txt", "打包内容\n")
    zf.writestr("data/nums.csv", "1,2,3\n")
with zipfile.ZipFile("out.zip") as zf:
    print(zf.namelist(), repr(zf.read("readme.txt").decode()))
    # ['readme.txt', 'data/nums.csv'] '打包内容\n'

真正的坑在解压。一个恶意压缩包可以把条目名写成 ../../etc/passwd,诱导解压程序把文件写到目标目录之外——这就是「路径穿越」(zip slip,tarfile 侧对应 CVE-2007-4559)。本机实测(Python 3.14.6)三种情况:

# 构造条目名为 ../pwned.txt 的 tar
info = tarfile.TarInfo(name="../escaped.txt")
info.size = len(data)
tf.addfile(info, io.BytesIO(data))

with tarfile.open(evil) as tf:
    tf.extractall(dest)                       # 默认 filter='data'
默认 extractall 抛错: OutsideDestinationError '../escaped.txt' would be extracted to '...'
BASE 下是否有越界的 escaped.txt: False
filter='fully_trusted' 后,BASE 下出现 escaped.txt: True
[拦截] ../escaped.txt -> /private/tmp/python_book/scratch/ch17/lab3/escaped.txt

结论要记牢:

  • Python 3.14 起 tarfile.extractall 默认 filter='data',会自动拦掉 .. 与绝对路径,抛 OutsideDestinationError。
  • zipfile.extractall 早已会剥掉条目名里的 ..(内部净化),所以现代 Python 上这两类 API 默认是安全的。
  • 但只要你写了 filter="fully_trusted"(为兼容旧行为而设),防护就没了——实测它会把文件写到目标目录之外。
  • 最稳的姿势仍是自己逐个 member 校验目标路径,再调用 extract:
def safe_extract(zip_path: Path, dest: Path) -> None:
    dest = dest.resolve()
    with zipfile.ZipFile(zip_path) as zf:
        for member in zf.infolist():
            target = (dest / member.filename).resolve()
            if not target.is_relative_to(dest):
                print("[拦截] 跳过越界条目:", member.filename)
                continue
            zf.extract(member, dest)

另外,「解压炸弹」(一个几十 KB 的 zip 解压出几十 GB)靠的是限制总大小:遍历 infolist() 累加 file_size,超过阈值就拒绝,别直接 extractall。

17.1.7 openpyxl:读写 Excel

openpyxl 3.1.5 是纯 Python 的 xlsx 读写库。写一个带公式和样式的表:

import openpyxl
from openpyxl.styles import Font, PatternFill

wb = openpyxl.Workbook()
ws = wb.active
ws.title = "销售"
ws.append(["商品", "单价", "数量", "金额"])
for r in [("键盘", 199.0, 3), ("鼠标", 89.5, 10), ("显示器", 1299.0, 2)]:
    ws.append(list(r))
for i in range(2, 5):
    ws.cell(row=i, column=4).value = f"=B{i}*C{i}"     # 公式写成字符串
ws["A5"], ws["D5"] = "合计", "=SUM(D2:D4)"
for cell in ws[1]:                                       # 表头加粗
    cell.font = Font(bold=True, color="FFFFFF")
    cell.fill = PatternFill("solid", fgColor="4472C4")
wb.save("sales.xlsx")   # 5175 字节

读回来时,公式默认按原样读出(不是算好的值):

wb2 = openpyxl.load_workbook("sales.xlsx")
for row in wb2["销售"].iter_rows(min_row=1, max_row=5, max_col=4):
    print([c.value for c in row])
['商品', '单价', '数量', '金额']
['键盘', 199, 3, '=B2*C2']
['鼠标', 89.5, 10, '=B3*C3']
['显示器', 1299, 2, '=B4*C4']
['合计', None, None, '=SUM(D2:D4)']

这里有个新手必踩的坑:load_workbook(path, data_only=True) 看起来能拿公式结果,但只有当文件被 Excel/LibreOffice 打开保存过、缓存了计算值才有数。纯 openpyxl 生成的表里没有缓存,实测 data_only=True 读到的 D5 是 None:

data_only=True 读 D5: None

所以「写公式」和「算结果」要分开:要结果就自己用 Python 算好再写值,或者交给 Excel 打开后另存。

17.1.8 csv 还是 xlsx?docx / pptx 的定位

场景推荐
程序间交换、体量大、纯数据CSV(csv / pandas)
给人看、要公式样式多 sheetxlsx(openpyxl)
需要保留计算逻辑、别人要改xlsx(写公式而非结果)
版本控制、可读 diffCSV 或 JSON,别用 xlsx(二进制)

CSV 是纯文本、无格式、无类型,程序友好;xlsx 有格式和公式,人友好。不要让程序去读写「人用的报表」再回头解析,两边的假设会打架。

python-docx(Word)与 python-pptx(PowerPoint)分别读写 .docx / .pptx,套路与 openpyxl 类似:打开文档、定位段落/幻灯片、改文本、保存。它们未预装,本节不贴输出,知道「有这两个库、适用场景是批量生成合同/周报/演示」即可。

17.1.9 subprocess.run:安全地调用外部命令

脚本常要调外部程序(git、ffmpeg、ImageMagick)。用 subprocess.run 时守住三条:列表传参、shell=False、check=True:

import subprocess, sys

r = subprocess.run(
    [sys.executable, "-c", "import sys; print('argv =', sys.argv[1:])", "hello", "world"],
    capture_output=True, text=True, check=True,
)
print(r.stdout.strip(), "| returncode:", r.returncode)
# argv = ['hello', 'world'] | returncode: 0

try:   # check=True 让非零退出码抛 CalledProcessError
    subprocess.run([sys.executable, "-c", "import sys; sys.exit(3)"], check=True)
except subprocess.CalledProcessError as e:
    print("捕获 CalledProcessError, returncode =", e.returncode)   # ...= 3

为什么绝不写 shell=True 拼字符串:subprocess.run(f"convert {user_input} out.png", shell=True) 里,如果 user_input 是 a.png; rm -rf ~,整条命令会被 shell 解释,等于把命令注入的口子敞开。列表传参把每个参数当作独立值,shell 不介入,注入就无从谈起。

17.1.10 文件监听:pathlib 轮询 + mtime

有些任务要「文件一变就处理」。watchdog 能监听文件系统事件,但依赖额外安装;不装库也能用轮询 + mtime 对比做最小实现:

def snapshot(root: Path) -> dict[str, float]:
    return {str(p): p.stat().st_mtime for p in root.rglob("*") if p.is_file()}

def watch_once(prev: dict[str, float]) -> dict[str, float]:
    cur = snapshot(WATCH)
    for path, mtime in cur.items():
        if path not in prev:
            print("  [新增]", Path(path).name)
        elif prev[path] != mtime:
            print("  [修改]", Path(path).name)
    for path in set(prev) - set(cur):
        print("  [删除]", Path(path).name)
    return cur
=== 文件监听(轮询)===
  [新增] a.txt
  [修改] a.txt
  [新增] b.txt

轮询的代价是延迟与开销:间隔太短空耗 CPU,太长漏掉瞬时变化。适合「几十秒扫一次」的批处理触发;要毫秒级响应、要监听成千上万文件,才值得上 watchdog。

小结

  • 批处理三件套:iterdir / glob / rglob 列文件,Path 组合路径,shutil 管整棵树。glob 只扫顶层、rglob 递归。
  • 一切破坏性操作先 dry-run:把「算计划」和「执行」拆成两个函数,打印确认后再动。
  • 绝不用字符串拼路径;删除前用 Path.resolve() + is_relative_to() 校验仍在允许的根内,rmtree 尤其如此。
  • 3.14 起 tarfile.extractall 默认 filter='data'、zipfile 会剥 ..,但 filter="fully_trusted" 会重开路径穿越;安全解压仍应逐个校验条目。解压炸弹靠限制总 file_size。
  • openpyxl 读写 xlsx,公式按原文读出,data_only=True 只有在 Excel 保存过缓存后才有值;CSV 给程序、xlsx 给人。
  • subprocess.run 用列表传参 + shell=False + check=True;轻量监听用轮询 + mtime,重场景才上 watchdog。

这一节把「本地文件与办公文档」这条线收口了。下一节我们走出本机,去网络上抓数据——同样要守住「先合规、再动手」的纪律。

阅读导航:上一节:数据校验、依赖注入与数据库访问 · 下一节:网络爬虫基础与合规边界 。

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「python」更多文章

  1. 《Python高级编程》目录
  2. 《Python高级编程》11.3 PEP 流程与版本迁移策略
  3. 《Python高级编程》11.2 嵌入式与自由线程运行时