Actions Runner Controller 与 Kubernetes 自动扩缩

系统讲解 Actions Runner Controller(ARC)在 Kubernetes 上托管自托管 Runner 的架构与调优,涵盖 AutoscalingRunnerSet 配置、基于 KEDA 的队列驱动扩缩、scale-to-zero、临时 Runner、Spot 节点成本优化以及可观测性与故障排查。

当 CI 负载出现"白天高峰、夜间归零"的潮汐特征时,固定数量的自托管 Runner 要么在高峰排长队,要么在夜间白白烧钱。Actions Runner Controller(ARC)解决的正是这个问题:把 Runner 变成 Kubernetes 里的临时 Pod,由队列深度驱动扩缩,负载来了就扩、没活干就缩到零。

本文聚焦 ARC 的第三代架构(gha-runner-scale-set),从部署、扩缩机制、Runner 镜像、成本优化一路讲到排障。读者需要具备基本的 Kubernetes 与 Helm 知识。

一、为什么需要 ARC

1.1 托管 Runner 的天花板

GitHub 托管的 Runner(ubuntu-latest 等)开箱即用,但有几条硬约束:

约束影响
规格固定无法用 32 核机器跑编译
无法访问内网私有依赖、内网服务不可达
分钟数计费大规模流水线成本不可控
环境不可定制无法预装大型工具链

1.2 自托管 Runner 的新问题

自托管解决了上面几条,却引入运维负担:谁来保证有机器可用?高峰排队怎么办?机器空闲时怎么回收?ARC 用 Kubernetes 的声明式能力回答了这三个问题。

1.3 ARC 的定位

ARC 不是"又一个 Runner",而是Runner 的编排器。它在集群里维护一个 listener,监听 GitHub 的作业队列,按需创建 Pod 执行作业,作业结束后销毁 Pod。这与 Kubernetes 自动扩缩 中的 HPA/KEDA 思路一脉相承,只是扩缩的指标从 CPU 换成了"待执行作业数"。

二、ARC 架构与部署

2.1 两个控制器

第三代 ARC 由两部分组成:

┌─────────────────────────────────────────────┐
│  gha-runner-scale-set-controller(集群级)    │
│  ── 监听 AutoscalingRunnerSet 资源           │
│  ── 为每个 scale set 创建 listener Pod        │
└─────────────────────────────────────────────┘
                     │
                     ▼
┌─────────────────────────────────────────────┐
│  gha-runner-scale-set(命名空间级)           │
│  ── listener 轮询 GitHub 作业队列             │
│  ── 按需创建 ephemeral runner Pod            │
└─────────────────────────────────────────────┘

2.2 安装控制器

# 安装集群级控制器
helm install arc \
  --namespace arc-systems \
  --create-namespace \
  oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set-controller

2.3 配置认证

ARC 需要一个能访问 GitHub API 的凭证。两种方式:

  • GitHub App(推荐):创建 App,授予 administration:write 与 actions:read 权限,安装到目标仓库/组织;
  • PAT(Personal Access Token):简单但权限过宽、会过期,仅适合试验。

创建 secret:

kubectl create secret generic pre-defined-secret \
  --namespace arc-runners \
  --from-literal=github_app_id=123456 \
  --from-literal=github_app_installation_id=654321 \
  --from-literal=github_app_private_key='-----BEGIN RSA PRIVATE KEY-----
...'

2.4 部署 scale set

helm install arc-runner-set \
  --namespace arc-runners \
  --create-namespace \
  --set githubConfigUrl="https://github.com/my-org" \
  --set githubConfigSecret=pre-defined-secret \
  oci://ghcr.io/actions/actions-runner-controller-charts/gha-runner-scale-set

部署完成后,工作流里用 runs-on 指向 scale set 名称即可:

jobs:
  build:
    runs-on: arc-runner-set
    steps:
      - uses: actions/checkout@v4

三、AutoscalingRunnerSet 配置详解

通过 Helm values 或直接声明 AutoscalingRunnerSet 资源来调参:

apiVersion: actions.github.com/v1alpha1
kind: AutoscalingRunnerSet
metadata:
  name: arc-runner-set
  namespace: arc-runners
spec:
  githubConfigUrl: "https://github.com/my-org"
  githubConfigSecret: pre-defined-secret
  minRunners: 0
  maxRunners: 30
  runnerScaleSetName: arc-runner-set
  template:
    spec:
      containers:
        - name: runner
          image: ghcr.io/actions/actions-runner:latest
          resources:
            requests:
              cpu: "2"
              memory: "4Gi"
            limits:
              cpu: "4"
              memory: "8Gi"

关键参数的含义与取值建议:

参数含义建议
minRunners常驻 Runner 数0 表示可 scale-to-zero
maxRunners峰值上限受集群容量与配额约束
runnerScaleSetName工作流 runs-on 引用的名字按用途区分,如 arc-linux-x64
template.spec.resourcesPod 资源请求/限制请求决定调度,限制决定上限

3.1 minRunners 的取舍

minRunners: 0 最省钱,但第一个作业要等 Pod 调度+启动(通常 20~60 秒)。对延迟敏感的仓库可以设 minRunners: 1~2 作为热池。

3.2 maxRunners 与集群容量

maxRunners 是逻辑上限,真正的瓶颈是集群的物理容量。如果集群只有 40 核,maxRunners: 30 且每个 Runner 请求 2 核,最多只能同时跑 20 个,其余会处于 Pending。因此 maxRunners 要与节点自动扩缩能力配合。

3.3 多 scale set 分池

不同作业对规格的需求差异很大。合理的做法是按用途分池:

arc-small    → 2C4G,跑 lint / 单元测试
arc-large    → 8C32G,跑编译 / 打包
arc-gpu      → 带 GPU 的节点池,跑模型推理测试

每个池独立设置 minRunners 与 maxRunners,互不抢占。容器化的作业定义可参考 /github-actions-container-jobs/。

四、扩缩容机制

4.1 队列驱动而非指标驱动

传统 HPA 依据 CPU/内存扩缩,ARC 依据待执行作业数扩缩——这是更直接的信号:队列里有 5 个作业,就准备 5 个 Runner。

GitHub 作业队列
      │  listener 轮询
      ▼
 待执行作业数 = N
      │
      ▼
 期望 Runner 数 = clamp(N, minRunners, maxRunners)
      │
      ▼
 创建/销毁 ephemeral runner Pod

4.2 scale-to-zero 的代价

scale-to-zero 意味着没有作业时集群里一个 Runner Pod 都没有。从零到第一个作业就绪的延迟由三部分构成:

阶段典型耗时优化手段
检测作业入队秒级listener 轮询间隔
Pod 调度秒级预热节点、减少亲和性约束
镜像拉取10~60 秒预拉取镜像、用镜像缓存

4.3 临时 Runner(ephemeral)

ARC 创建的 Runner 默认是一次性的:执行完一个作业就销毁。这带来强隔离——作业之间不会互相污染工作区、缓存、凭证。代价是每个作业都要重新拉代码、重建缓存。

对缓存敏感的场景,可以把缓存放在集群外部(S3、NFS)而不是 Runner 本地磁盘。

4.4 扩缩的观测指标

关注这几个指标判断扩缩是否健康:

# 待执行作业数(listener 暴露)
kubectl logs -n arc-runners deploy/arc-runner-set-...-listener | grep -i "job"

# Pod 状态分布
kubectl get pods -n arc-runners -l app.kubernetes.io/component=runner

# Pending 的 Pod(说明集群容量不足)
kubectl get pods -n arc-runners --field-selector status.phase=Pending

如果 Pending 长期不为零,说明 maxRunners 超过了集群实际容量,需要扩节点或调低上限。

4.5 第二代与第三代 ARC 的区别

如果你在网上看到大量 RunnerDeployment、Runner、HorizontalRunnerAutoscaler 的配置,那是第二代 ARC。第三代做了架构级重构:

维度第二代第三代
核心 CRDRunnerDeployment / HRAAutoscalingRunnerSet
扩缩机制自研 webhook 扩缩官方 listener,队列驱动
Runner 生命周期可复用强制 ephemeral
部署方式混合控制器与 scale set 分离
状态维护模式推荐使用

新项目应直接用第三代。迁移时主要工作量在于把 runs-on 名称与 Helm values 重写,逻辑本身不变。

五、Runner 镜像与容器模式

5.1 自定义镜像

默认镜像只有最基本的工具。生产环境通常要自建:

FROM ghcr.io/actions/actions-runner:latest

USER root
RUN apt-get update && apt-get install -y --no-install-recommends \
      build-essential git-lfs jq unzip \
    && rm -rf /var/lib/apt/lists/*

USER runner

构建后推送到镜像仓库,在 template.spec.containers[0].image 中引用。镜像越小,拉取越快,冷启动越短。

5.2 Docker in Docker

需要在 Runner 里跑 docker build 时,有两种模式:

模式做法安全性
DinD挂载 docker:dind sidecar需 privileged,风险高
挂载宿主 socket挂 /var/run/docker.sock等于给容器宿主权限

两者都有安全代价。更安全的选择是用 BuildKit 的 rootless 模式或 Kaniko/Buildah 这类无需守护进程的构建工具。

5.3 生命周期钩子

ARC 支持容器生命周期钩子,可以在作业开始前做预热:

template:
  spec:
    containers:
      - name: runner
        image: my-registry/actions-runner:1.0
        lifecycle:
          postStart:
            exec:
              command: ["/bin/sh", "-c", "preload-cache.sh &"]

六、成本与性能优化

6.1 Spot / 抢占式节点

CI 作业大多可中断可重试,非常适合跑在 Spot 实例上。用 Karpenter 或 Cluster Autoscaler 为 Runner 池配置 Spot 优先:

# Karpenter NodePool 片段
spec:
  disruption:
    consolidationPolicy: WhenEmpty
    consolidateAfter: 30s
  template:
    spec:
      requirements:
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["spot", "on-demand"]

WhenEmpty 让空节点尽快回收,配合 minRunners: 0 可以把夜间成本压到接近零。成本优化的整体思路可参考 /github-actions-cost-optimization/。

6.2 节点自动扩缩的配合

ARC 负责"创建 Runner Pod",节点自动扩缩负责"给 Pod 找机器"。两者必须协同:

  • Pod 因容量不足 Pending → 触发节点扩容;
  • 节点空闲 → 缩容回收。

如果只装 ARC 不装节点自动扩缩,maxRunners 就成了摆设。

6.3 缓存局部性

临时 Runner 每次重建工作区,缓存命中率低。三个缓解手段:

  • 把依赖缓存放到对象存储(S3 兼容),跨 Runner 共享;
  • 用 PVC 挂载一个共享缓存卷(注意并发写冲突);
  • 镜像里预装不常变的大依赖(工具链、SDK)。

6.4 资源请求的精确性

requests 设得过大浪费节点,过小会导致节点超卖、作业互相争抢。建议先用默认值观察一段时间,用 kubectl top pod 记录实际峰值,再回填请求值。

6.5 网络出口与代理

内网 Runner 访问公网(拉依赖、推镜像)通常要经过代理。两种配置方式:

template:
  spec:
    containers:
      - name: runner
        image: my-registry/actions-runner:1.0
        env:
          - name: HTTP_PROXY
            value: "http://proxy.internal:3128"
          - name: HTTPS_PROXY
            value: "http://proxy.internal:3128"
          - name: NO_PROXY
            value: "localhost,127.0.0.1,.cluster.local"

NO_PROXY 一定要包含集群内部域名后缀,否则访问 API Server、DNS 都会被代理劫持,导致 Runner 无法注册。

6.6 命名空间与配额

把 Runner 放在独立命名空间,并用 ResourceQuota 限制其总量,防止 CI 突发流量挤占生产工作负载:

apiVersion: v1
kind: ResourceQuota
metadata:
  name: ci-quota
  namespace: arc-runners
spec:
  hard:
    requests.cpu: "64"
    requests.memory: "256Gi"
    pods: "40"

配额是"安全阀":即使 maxRunners 设得很大,也不会越过配额上限。

七、可观测性与故障排查

7.1 常见故障对照表

现象可能原因排查方向
作业一直排队maxRunners 已满或集群无容量看 Pod 是否 Pending
Pod 启动后立即失败认证 secret 失效看 listener 日志
runs-on 找不到 Runnerscale set 名不匹配核对 runnerScaleSetName
作业卡在 “Waiting for a runner”listener 未连上 GitHub检查网络与 App 权限

7.2 listener 日志

listener 是 ARC 的"大脑",绝大多数问题都能从它的日志找到线索:

kubectl logs -n arc-runners \
  -l app.kubernetes.io/component=runner-scale-set-listener \
  --tail=200

关注 Failed to acquire job、Unauthorized 这类关键字。

7.3 监控指标

ARC 暴露 Prometheus 指标,建议接入 Grafana 观测:

  • gha_runner_scale_set_job_queue_duration_seconds:作业排队时长;
  • gha_runner_scale_set_desired_runners:期望 Runner 数;
  • gha_runner_scale_set_assigned_jobs:已分配作业数。

排队时长的 P95 是最有业务意义的指标——它直接等于"开发者等了多久"。

7.4 升级与回滚

ARC 的控制器与 scale set 是分离的,可以独立升级。升级控制器前先确认新版本与现有 scale set 的 API 兼容;回滚用 Helm:

helm rollback arc-runner-set 1 -n arc-runners

7.5 一个完整的排障流程

当"作业不执行"时,按下面顺序逐层排查:

# 1. listener 是否在运行
kubectl get pods -n arc-runners -l app.kubernetes.io/component=runner-scale-set-listener

# 2. listener 是否成功注册到 GitHub
kubectl logs -n arc-runners deploy/arc-runner-set-listener --tail=50

# 3. 是否有 Runner Pod 被创建
kubectl get pods -n arc-runners -w

# 4. Pod 是否被调度
kubectl describe pod -n arc-runners <pod-name> | tail -20

# 5. 事件层是否有配额/调度失败
kubectl get events -n arc-runners --sort-by=.lastTimestamp | tail -20

绝大多数问题会在这五步中的某一步暴露:权限(第 2 步)、容量(第 4 步)、配额(第 5 步)。

八、落地清单

  • 认证使用 GitHub App 而非 PAT;
  • 按用途分池,各池独立 minRunners/maxRunners;
  • 自定义精简镜像,预装常用工具;
  • 节点自动扩缩已配置且支持 Spot;
  • 队列等待时长、Pending Pod 数已纳入监控;
  • 缓存放在集群外部以缓解临时 Runner 的重建成本;
  • 网络策略限制 Runner 的出站访问。

总结

ARC 把自托管 Runner 从"一堆长期运行的机器"变成"Kubernetes 上按需生灭的 Pod",用队列深度驱动扩缩,用 scale-to-zero 换取成本弹性。它真正的难点不在安装,而在三个协同:maxRunners 与集群容量的协同、Pod 创建与节点扩缩的协同、以及临时 Runner 与缓存策略的协同。把这三处调顺,ARC 才能在潮汐负载下既快又省。它与托管 Runner 并非互斥——常见做法是默认跑托管、重负载与内网作业跑 ARC,两者在 /github-actions-self-hosted-runners/ 中有更全面的对比。

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「github-actions」更多文章

  1. 多云部署编排与基础设施漂移检测
  2. 文档站与静态站点发布流水线
  3. AI 代码审查与 PR 助手集成