《Go 语言运行时原理》4.2 逃逸分析与分配决策

把 -gcflags=-m 的每一行拆开读:从 go1.27.0 的 escape 包(escape.go、graph.go、solve.go、leaks.go)看数据流图怎么建、leaks 位图怎么编码、不动点怎么解,并用同一段代码在开/关内联下的不同输出,解释「moved to heap」「escapes to heap」「flow:」三行分别在说什么。

4.2 逃逸分析与分配决策

go build -gcflags='-m' 打印的每一行,都是一次编译期推断的结论。大多数人对它的用法停留在「看到 escapes to heap 就去改代码」——但如果不知道这行是谁写的、依据是什么,就很容易改错方向:把本该逃逸的对象硬塞回栈,或者为了「不逃逸」写出更难维护的代码。

本节不重复「什么写法会逃逸」的清单(那是卷三 6.1 的内容),而是从编译器的数据流分析实现入手,回答「这一行是谁打印的、它凭什么这么判断」。

本节要回答:-m 的输出每一行对应编译器里哪一步、-m -m 多出来的 flow: 轨迹是什么?结论是:逃逸分析在 src/cmd/compile/internal/escape/ 里先把函数体建成「位置—洞(location/hole)」数据流图,再用 leaks 位图记录每个位置到堆/到返回值的最短解引用距离,最后在 solve.go 里迭代到不动点;-m 打印结论,-m -m 额外打印 explainFlow 生成的路径。 与卷三《Go 语言高级编程》6.1 的分工:卷三写的是应用侧决策表(什么写法会逃逸、怎么改代码),本节写的是编译器实现(数据流图与位图编码),只写增量。同样,/posts/golang/ 下讲逃逸的专题文章偏概念,本节只补「输出逐行怎么读」。

4.2.1 实验:三种 -m 的输出怎么读

复现基线:

  • Go 工具链 go version go1.27.0 darwin/arm64(GOTOOLCHAIN=go1.27.0)
  • 机器:Apple M1 Pro,10 核,32 GiB
  • 分析对象是纯编译期行为,不受 GOGC/GOMAXPROCS 影响,也不需要运行程序
  • 命令均为 go build -gcflags=... -o /dev/null .;-m 是单级,-m -m 是两级,-l 关闭内联

被测代码刻意覆盖四类典型情形:返回结构体值、返回局部变量地址、装箱到 any、闭包捕获变量。

type Point struct{ X, Y int }

func stackAlloc() Point {
	p := Point{1, 2}
	return p
}

func heapAlloc() *Point {
	p := Point{3, 4}
	return &p
}

func leakToInterface() any {
	v := 42
	return v
}

func closureCapture() func() int {
	x := 0
	return func() int { x++; return x }
}

第一遍,默认内联,单级 -m。注意输出里混着内联决策,这是最容易被误读的地方:

$ GOTOOLCHAIN=go1.27.0 go build -gcflags='-m' -o /dev/null .
# escape
./main.go:7:6: can inline stackAlloc
./main.go:12:6: can inline heapAlloc
./main.go:17:6: can inline leakToInterface
./main.go:22:6: can inline closureCapture
./main.go:24:9: can inline closureCapture.func1
./main.go:27:6: can inline sliceLiteral
./main.go:32:24: inlining call to stackAlloc
./main.go:32:38: inlining call to heapAlloc
...
./main.go:13:2: moved to heap: p
./main.go:19:9: 42 escapes to heap
./main.go:23:2: moved to heap: x
./main.go:24:9: func literal escapes to heap
./main.go:28:14: []int{...} escapes to heap
./main.go:32:13: ... argument does not escape
./main.go:32:24: ~r0 escapes to heap
./main.go:32:28: *(~r0) escapes to heap

can inline / inlining call to 来自内联阶段,moved to heap / escapes to heap / does not escape 来自逃逸分析。两者共享同一个 -m 开关,这是第一件要知道的事。

第二遍,加 -l 关闭内联,让逃逸结论「裸奔」——这样输出里只剩逃逸分析自己的判断:

$ GOTOOLCHAIN=go1.27.0 go build -gcflags='-m -l' -o /dev/null .
# escape
./main.go:13:2: moved to heap: p
./main.go:19:9: 42 escapes to heap
./main.go:23:2: moved to heap: x
./main.go:24:9: func literal escapes to heap
./main.go:28:14: []int{...} escapes to heap
./main.go:32:13: ... argument does not escape
./main.go:32:24: stackAlloc() escapes to heap
./main.go:32:28: *heapAlloc() escapes to heap
./main.go:32:77: .autotmp_0() escapes to heap
./main.go:32:93: sliceLiteral(1) escapes to heap

对比两遍输出可以看到一个关键事实:同一段 main,开内联时打印的是 ~r0 escapes to heap(内联后 stackAlloc 的返回值没有名字,用临时 ~r0 表示),关内联时打印的是 stackAlloc() escapes to heap。内联会把「跨函数的逃逸」变成「函数内的逃逸」,因此 -m 的输出会随内联决策而变。想看清函数内部的原始判断,必须先 -l。

第三遍,两级 -m -m,多出 flow: 轨迹。这是本节最有用的一档,它把「为什么」摊开:

$ GOTOOLCHAIN=go1.27.0 go build -gcflags='-m -m' -o /dev/null .
./main.go:31:6: cannot inline main: function too complex: cost 218 exceeds budget 80
./main.go:13:2: p escapes to heap in heapAlloc:
./main.go:13:2:   flow: ~r0 ← &p:
./main.go:13:2:     from &p (address-of) at ./main.go:14:9
./main.go:13:2:     from return &p (return) at ./main.go:14:2
./main.go:13:2: moved to heap: p
./main.go:19:9: 42 escapes to heap in leakToInterface:
./main.go:19:9:   flow: ~r0 ← &{storage for 42}:
./main.go:19:9:     from 42 (spill) at ./main.go:19:9
./main.go:19:9:     from return 42 (return) at ./main.go:19:2
./main.go:23:2: closureCapture capturing by ref: x (addr=false assign=true width=8)
./main.go:23:2: x escapes to heap in closureCapture:
./main.go:23:2:   flow: {storage for func literal} ← &x:
./main.go:23:2:     from x (captured by a closure) at ./main.go:24:22
./main.go:23:2:     from x (reference) at ./main.go:24:22

flow: 一行是路径的终点(~r0 ← &p 读作「返回值 ~r0 收到了 p 的地址」),下面缩进的 from ... (原因) at 位置 是逐跳的边。整段合起来就是一条从「取地址」到「泄漏到返回值」的证据链。(address-of)、(spill)、(captured by a closure)、(slice-literal-element) 这些括号里的词是编译器给每条边打的标签。

4.2.2 源码:数据流图、leaks 位图、不动点求解

逃逸分析的入口在 src/cmd/compile/internal/escape/escape.go。Funcs 自底向上遍历函数,Batch 对一批函数做分析:

func Funcs(all []*ir.Func) {
	reassignOracles := make(map[*ir.Func]*ir.ReassignOracle)
	ir.VisitFuncsBottomUp(all, func(list []*ir.Func, recursive bool) {
		Batch(list, reassignOracles)
	})
}

Batch 的注释直接点明了它做三件事:建图、流闭包、解不动点:

func Batch(fns []*ir.Func, reassignOracles map[*ir.Func]*ir.ReassignOracle) {
	var b batch
	b.heapLoc.attrs = attrEscapes | attrPersists | attrMutates | attrCalls
	...
	// Construct data-flow graph from syntax trees.
	for _, fn := range fns {
		b.initFunc(fn)
	}
	for _, fn := range fns {
		if !fn.IsClosure() {
			b.walkFunc(fn)
		}
	}
	...
	b.walkAll()
	b.finish(fns)
}

图的基本单元在 src/cmd/compile/internal/escape/graph.go:location 是「一个可能有地址的位置」(变量、临时值、返回值……),hole 是「位置上的一个洞」(带解引用偏移和备注),edge 连接两者。hole 上的 addr/deref 操作会移动偏移:

func (k hole) shift(delta int) hole {
	n := k
	n.derefs += delta
	...
	return n
}

func (k hole) deref(where ir.Node, why string) hole { return k.shift(1).note(where, why) }
func (k hole) addr(where ir.Node, why string) hole  { return k.shift(-1).note(where, why) }

why 就是 4.2.1 输出里括号里的 (address-of)、(spill) 这些标签——它们是边在建立时被记下来的,不是打印时猜的。

每个位置的「逃逸程度」用 src/cmd/compile/internal/escape/leaks.go 的 leaks 位图编码。它是一个 [8]uint8,下标含义固定:

const (
	leakHeap = iota
	leakMutator
	leakCallee
	leakResult0
)

// Heap returns the minimum deref count of any assignment flow from l
// to the heap. If no such flows exist, Heap returns -1.
func (l leaks) Heap() int { return l.get(leakHeap) }

注意 Heap() 的语义是**「到堆的最短解引用次数」**,不是布尔值。get 用「减一」编码,所以「没有路径」是 -1:

func (l leaks) get(i int) int { return int(l[i]) - 1 }

func (l *leaks) add(i int, derefs int) {
	if old := l.get(i); old < 0 || derefs < old {
		l.set(i, derefs)
	}
}

add 只在「新的路径更短」时更新——这就是为什么结论是「最短距离」。Optimize 再砍掉比堆路径更长的结果路径:

func (l *leaks) Optimize() {
	if x := l.Heap(); x >= 0 {
		for i := 1; i < len(*l); i++ {
			if l.get(i) >= x {
				l.set(i, -1)
			}
		}
	}
}

求解在 src/cmd/compile/internal/escape/solve.go:walkAll。它把 allLocs、heapLoc、mutatorLoc、calleeLoc 全部推入队列,反复传播直到不动点:

func (b *batch) walkAll() {
	todo := newQueue(len(b.allLocs) + 3)
	...
	for _, loc := range b.allLocs {
		todo.pushFront(loc)
		loc.queuedWalkAll = true
	}
	todo.pushFront(&b.mutatorLoc)
	todo.pushFront(&b.calleeLoc)
	todo.pushFront(&b.heapLoc)
	...
	for todo.len() > 0 {
		root := todo.popFront()
		root.queuedWalkAll = false
		walkgen++
		b.walkOne(root, walkgen, enqueue)
	}
}

-m -m 的 flow: 行就来自这里。solve.go:explainFlow 沿 src.dst 一路回溯到根,按 base.Flag.LowerM >= 2 决定是否打印:

func (b *batch) explainFlow(pos string, dst, srcloc *location, derefs int, notes *note, explanation []*logopt.LoggedOpt) []*logopt.LoggedOpt {
	ops := "&"
	if derefs >= 0 {
		ops = strings.Repeat("*", derefs)
	}
	print := base.Flag.LowerM >= 2
	flow := fmt.Sprintf("   flow: %s ← %s%v:", b.explainLoc(dst), ops, b.explainLoc(srcloc))
	if print {
		fmt.Printf("%s:%s\n", pos, flow)
	}
	...
}

ops 的构造逻辑说明了一件事:flow: 里的 & 或 * 数量就是那条边的解引用增量。~r0 ← &p 里只有一个 &,代表「取一层地址」。

最终打印「逃逸/不逃逸」结论的是 escape.go:reportLeaks:

func (b *batch) reportLeaks(pos src.XPos, name string, esc leaks, sig *types.Type) {
	warned := false
	if x := esc.Heap(); x >= 0 {
		if x == 0 {
			base.WarnfAt(pos, "leaking param: %v", name)
		} else {
			base.WarnfAt(pos, "leaking param content: %v", name)
		}
		warned = true
	}
	...
	if !warned {
		base.WarnfAt(pos, "%v does not escape", name)
	}
}

esc.Heap() == 0 打印 leaking param,> 0 打印 leaking param content——这两个词的差别就是解引用层数:前者是把指针本身漏出去,后者是把指针指向的内容漏出去。

最后,还有一类逃逸与数据流无关,由 HeapAllocReason 单独判定(src/cmd/compile/internal/escape/utils.go),例如「过大的局部数组」「make 出来的切片长度非常量」。这类在 Batch 里被直接接到 heapLoc:

		if why := HeapAllocReason(loc.n); why != "" {
			b.flow(b.heapHole().addr(loc.n, why), loc)
		}

4.2.3 决策:从 -m 输出反推改法

-m 输出里的字样编译器里的来源该怎么读 / 怎么改
moved to heap: preportLeaks + esc.Heap(),变量本身要堆化该变量的地址被存进了堆对象或返回;想留栈上就得断开这条 flow: 链
42 escapes to heap常量/临时值被装箱(leaks 的 leakHeap)典型是装进 any 或闭包;改法是避免接口装箱(如预分配、泛型)
... argument does not escape变参切片未逃逸说明 fmt 这类调用没把变参留在堆上;不用改
~r0 escapes to heap内联后的匿名返回值开内联才出现的形态;用 -l 看原始判断
flow: ~r0 ← &pexplainFlow 打印的路径终点在左、起点在右;顺着缩进的 from 找到「谁取了这个地址」
leaking param: xesc.Heap() == 0指针本身漏出去;改函数签名(返回副本而非指针)
leaking param content: xesc.Heap() > 0指针指向的内容漏出去;通常无害,别急着改
x does not escapereportLeaks 的兜底分支理想状态;这是结论,不是建议
cannot inline main: cost 218 exceeds budget 80内联成本模型与逃逸无直接关系,但内联会改变逃逸结论;分析前先看有没有 cannot inline
closureCapture capturing by ref: xflowClosure(escape.go)闭包按引用捕获;by ref 意味着 x 会逃逸

三条实操建议:

  1. 分析前先加 -l。内联会把跨函数逃逸折叠成函数内逃逸,输出形态完全不同;先用 -m -l 看原始判断,再用 -m -m 看路径。
  2. 优先看 flow:,而不是结论行。结论行只告诉你「逃逸了」,flow: 告诉你「从哪一步漏出去的」——后者才是能改的地方。
  3. 别为了消灭 escapes to heap 而牺牲可读性。leaking param content 这类「内容逃逸」在很多场景下是良性的;把它当成 bug 去修,往往是把清晰代码改成晦涩代码。真正的收益点是热路径上的高频分配——先用 4.3 的 profile 找到热点,再回来用本节的方法逐行读。

阅读导航:上一节:4.1 size class 与 mcache/mcentral/mheap · 下一节:4.3 分配热点定位与对象复用 。

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「golang」更多文章

  1. 《Go 语言编程实战》目录
  2. 《Go 语言编程实战》18.3 上线、观测与迭代
  3. 《Go 语言编程实战》18.2 故障演练