5.3 可见性陷阱与 -race 实测
前两节是理论:happens-before 怎么定义、由哪些操作建立。这一节全部是反例——把「忘了建立边」的几种典型写法跑出来,看 -race 怎么把它们一个个揪出来。理论只有配上能复现的失败,才真的进脑子。
本节要回答:哪些看起来「应该没问题」的并发写法其实有 data race,
-race会给出什么报告。结论是:goroutine 退出、缓冲 channel 方向、无同步共享变量、手写双重检查锁这四类是最高频的坑;-race能抓到它们,但抓不到就等于没竞争——它的能力边界也要说清。
5.3.1 -race 怎么用
-race 是 Go 内置的竞争检测器,底层是 ThreadSanitizer。两种用法:
# 单文件 / 单命令
$ GOTOOLCHAIN=go1.27.0 go run -race ./prog
# 测试
$ GOTOOLCHAIN=go1.27.0 go test -race ./...
检测到竞争时的行为:打印 WARNING: DATA RACE 报告,然后以退出码 66 结束。这一点在 5.1 节引用的官方文档里有明确出处:
Any implementation can, upon detecting a data race,
report the race and halt execution of the program.
Implementations using ThreadSanitizer
(accessed with “go build -race”)
do exactly this.
本机 CGO_ENABLED=1(go env CGO_ENABLED 输出 1),-race 可用——-race 依赖 cgo 运行时。
5.3.2 陷阱一:goroutine 退出不建立边
最容易被「直觉」坑到的一类。写一个 goroutine 去赋值,主 goroutine 直接读:
var a string
func main() {
done := make(chan struct{})
go func() {
a = "hello"
close(done)
}()
fmt.Println("len(a) =", len(a))
<-done
fmt.Println("after join len(a) =", len(a))
}
-race 的真实输出(节选):
len(a) = 0
==================
WARNING: DATA RACE
Write at 0x000104383a10 by goroutine 7:
main.main.func1()
/tmp/gbadv2/mem/goroutine_destruction/main.go:10 +0x30
Previous read at 0x000104383a10 by main goroutine:
main.main()
/tmp/gbadv2/mem/goroutine_destruction/main.go:13 +0xa0
...
after join len(a) = 5
Found 1 data race(s)
exit status 66
两件事同时被印证:第一次读拿到的是 len(a) = 0(写还没发生),第二次读在 <-done 之后拿到 5。差别就在 <-done——它建立了边。这正是 5.2.4 那条规则的实证:goroutine 的退出本身不建立边,是 close(done) 与 <-done 这对同步操作建立的。
5.3.3 陷阱二:缓冲 channel 的方向搞反
5.2.5 的规则(3)只对无缓冲 channel 成立。把官方反例跑出来:
var c = make(chan int, 1) // 注意:带缓冲
var s string
func f() {
s = "hello, world"
<-c // 先接收
}
func main() {
go f()
c <- 0 // 后发送
fmt.Println(s)
}
-race 真实输出(节选):
==================
WARNING: DATA RACE
Write at 0x0001026bfa40 by goroutine 7:
main.f()
/tmp/gbadv2/mem/buffered_counter/main.go:9 +0x28
Previous read at 0x0001026bfa40 by main goroutine:
main.main()
/tmp/gbadv2/mem/buffered_counter/main.go:16 +0x58
...
Found 1 data race(s)
exit status 66
同样的代码,把缓冲去掉(make(chan int))就干净了——因为无缓冲 channel 上「接收 → 发送完成」这条边成立。这就是文档说的:If the channel were buffered (e.g., c = make(chan int, 1)) then the program would not be guaranteed to print "hello, world". (It might print the empty string, crash, or do something else.)
| channel 形态 | 语句顺序 | 是否建立边 |
|---|---|---|
| 无缓冲 | 先 <-c 后 c<-0 | 是(接收 → 发送完成) |
| 带缓冲(cap≥1) | 先 <-c 后 c<-0 | 否(可能竞争) |
| 任意 | 先 c<-0 后 <-c | 是(发送 → 接收完成) |
5.3.4 陷阱三:无同步计数器
四个 goroutine 各做 1000 次 counter++。counter++ 是「读-改-写」三步,不是原子操作:
var counter int
func main() {
var wg sync.WaitGroup
for i := 0; i < 4; i++ {
wg.Add(1)
go func() {
defer wg.Done()
for j := 0; j < 1000; j++ {
counter++
}
}()
}
wg.Wait()
fmt.Println("counter =", counter)
}
不带 -race 连跑三次的真实结果:
counter = 2962
counter = 2592
counter = 3411
期望 4000,实际每次都不同——这就是 data race 的「不确定性」。带 -race:
==================
WARNING: DATA RACE
Read at 0x000101102450 by goroutine 9:
main.main.func1()
/tmp/gbadv2/mem/race_demo/main.go:17 +0x88
Previous write at 0x000101102450 by goroutine 8:
main.main.func1()
/tmp/gbadv2/mem/race_demo/main.go:17 +0xa0
...
Found 2 data race(s)
exit status 66
注意报告里 Read at / Previous write at 的地址相同(0x000101102450)——检测器认的就是「同一内存位置上的无序读写」。修复方式按优先级:
| 修法 | 写法 | 适用 |
|---|---|---|
| 原子操作 | atomic.AddInt64(&counter, 1) | 单一计数器 |
| 互斥锁 | mu.Lock(); counter++; mu.Unlock() | 临界区含多步 |
| channel 归并 | 每个 goroutine 发结果,单 goroutine 汇总 | 结构性聚合 |
5.3.5 陷阱四:双重检查锁与 busy waiting
官方文档在「Incorrect synchronization」一节点名了两种错误惯用法。双重检查锁:
var a string
var done bool
func setup() { a = "hello, world"; done = true }
func doprint() {
if !done {
once.Do(setup)
}
print(a)
}
文档结论:there is no guarantee that, in doprint, observing the write to done implies observing the write to a. This version can (incorrectly) print an empty string. 因为 done 是普通 bool,读写它不建立任何边——「看到 done=true」不蕴含「看到 a 的写」。
busy waiting:
var a string
var done bool
func setup() { a = "hello, world"; done = true }
func main() {
go setup()
for !done {
}
print(a)
}
文档的两句结论一句比一句狠:could print an empty string too,以及 there is no guarantee that the write to done will ever be observed by main ... The loop in main is not guaranteed to finish. 这个 for 循环有可能永不退出——因为没有任何同步事件,编译器甚至可以把 done 的读提到循环外。
这两类错误的统一修法:把普通 bool 换成 atomic.Bool(5.2.8 已实测干净),或直接用 sync.Once(5.2.7)。区别是 atomic.Bool 提供「观察到即同步」,Once 提供「一次性执行 + 全体可见」。
5.3.6 -race 的能力边界
-race 很强,但必须知道它不是证明:
| 能力 | 说明 |
|---|---|
| 能报什么 | 实际发生的竞争访问(动态检测,非静态分析) |
| 报不出什么 | 本次运行没被执行到的竞争路径 |
| 假阴性 | 常见:竞争窗口没被调度到 |
| 假阳性 | 极少:理论上可能误报,实践中按报告改 |
| 代价 | 运行时内存与 CPU 开销显著(数倍),不适合常开在生产 |
| 退出码 | 检测到竞争 → 66 |
关键结论:**-race 干净 ≠ 无竞争,-race 报错 = 一定有竞争。**要覆盖更多路径,靠的是提高测试覆盖率与并发强度(多跑几轮、加 -count),而不是相信一次运行。
5.3.7 修复清单
| 症状 | 根因 | 修法 |
|---|---|---|
| 读到的值忽新忽旧 | 读写之间没有同步操作 | 加锁 / channel / 原子操作 |
| goroutine 结果丢失 | 等「它应该跑完了」 | WaitGroup、close(done) |
| 缓冲 channel 不保证顺序 | 方向与缓冲不匹配 | 用无缓冲,或改发送在前 |
for !flag {} 不退出 | 普通 bool 无同步 | atomic.Bool 或 Once |
| 计数器丢增量 | x++ 非原子 | atomic.Add / 锁 |
| 结构体字段撕裂 | 多字对象竞争 | 别共享可变多字对象(见 5.1.7) |
一句话收束:data race 不是「性能问题」,是「正确性问题」——它的后果从「读到旧值」到「循环不退出」再到「任意内存损坏」都有。-race 是发现它的最好工具,但真正的解法只有一条:在每一对跨 goroutine 的读写之间,主动建立一个 happens-before 边。
阅读导航:上一节:5.2 happens-before 的建立 · 下一节:6.1 逃逸分析决策表 。
继续阅读
探索更多技术文章
浏览归档,发现更多关于系统设计、工具链和工程实践的内容。