《Go 语言高级编程》4.2 math/rand/v2 与 iter.Pull 组合子

math/rand/v2 与 iter.Pull 常被记错版本归属。本节用 api 清单核实:前者 Go 1.22 引入、1.24 补 AppendBinary,后者是 Go 1.23 的 range-over-func 配套;再实测 PCG 与 ChaCha8 的可复现与 MarshalBinary 往返,给出 iter.Pull 比 range 慢约 132 倍的真实数字与使用判据。

4.2 math/rand/v2 与 iter.Pull 组合子

math/rand/v2 和 iter.Pull 是两件看起来毫不相干的东西:一个是随机数生成器,一个把「推送式」迭代器翻成「拉取式」。它们被放进同一节,是因为它们共享同一个陷阱——版本归属容易被记错。前者容易被记成「Go 1.20 就有」,后者容易被记成「1.24 才有」,两种印象都不准。

本节要回答:math/rand/v2 与 iter.Pull 各自属于哪个 Go 版本,它们真正解决了什么问题,以及 iter.Pull 的代价有多高。结论是:math/rand/v2 是 Go 1.22 引入(1.24 补 AppendBinary),iter.Pull 是 Go 1.23 引入;而 iter.Pull 每调用一次要付一个 goroutine 与 channel 握手的钱,本机实测比 range 迭代慢约 132 倍。

4.2.1 版本归属:api 清单给答案

本机 /usr/local/go 是 go1.26.0,api/ 下有 go1.1.txt … go1.26.txt。用包锚定的方式查首次出现版本(别用裸子串,rand 会命中一堆无关条目):

$ grep -ln "^pkg math/rand/v2," /usr/local/go/api/go1.*.txt | sort -V | head -1
/usr/local/go/api/go1.22.txt

$ grep -ln "^pkg iter," /usr/local/go/api/go1.*.txt
/usr/local/go/api/go1.23.txt

$ grep -h "func Pull" /usr/local/go/api/go1.*.txt
pkg iter, func Pull2[$0 interface{}, $1 interface{}](Seq2[$0, $1]) (func() ($0, $1, bool), func()) #61897
pkg iter, func Pull[$0 interface{}](Seq[$0]) (func() ($0, bool), func()) #61897

而 math/rand/v2 在 1.24 又补了两个方法:

$ grep "^pkg math/rand/v2" /usr/local/go/api/go1.24.txt
pkg math/rand/v2, method (*ChaCha8) AppendBinary([]uint8) ([]uint8, error) #62384
pkg math/rand/v2, method (*PCG) AppendBinary([]uint8) ([]uint8, error) #62384

整理成版本归属表:

符号 / 包引入版本核实命令
math/rand/v2(整包)Go 1.22grep -ln "^pkg math/rand/v2," api/go1.*.txt
math/rand/v2 的 AppendBinaryGo 1.24grep "^pkg math/rand/v2" api/go1.24.txt
iter(整包)Go 1.23grep -ln "^pkg iter," api/go1.*.txt
iter.Pull / iter.Pull2Go 1.23grep "func Pull" api/go1.*.txt

iter 与 range over func 同属 Go 1.23 的同一批特性(1.22 时还是 GOEXPERIMENT=rangefunc,1.23 转正)。

4.2.2 math/rand/v2:可复现性实测

v2 相比 v1 最大的结构性变化是去掉了全局可播种源,把随机源显式化:rand.NewPCG(seed1, seed2) 与 rand.NewChaCha8([32]byte) 返回可序列化的源,交给 rand.New(source) 使用。同一种子必须给出同一序列——这是「可复现」的全部意义。实测:

p1 := rand.NewPCG(1, 2)
p2 := rand.NewPCG(1, 2)
r1 := rand.New(p1)
r2 := rand.New(p2)
fmt.Printf("PCG same seed: %d %d\n", r1.IntN(1000), r2.IntN(1000))

var seed [32]byte
for i := range seed {
	seed[i] = byte(i)
}
c1 := rand.NewChaCha8(seed)
c2 := rand.NewChaCha8(seed)
rc1, rc2 := rand.New(c1), rand.New(c2)
fmt.Printf("ChaCha8 same seed: %d %d\n", rc1.Uint64(), rc2.Uint64())

b, _ := c1.MarshalBinary()
var c3 rand.ChaCha8
_ = c3.UnmarshalBinary(b)
fmt.Printf("ChaCha8 marshal len=%d roundtrip ok=%v\n", len(b), c3 == *c2)

真实输出:

PCG same seed: 769 769
ChaCha8 same seed: 12537355132343524571 12537355132343524571
ChaCha8 marshal len=48 roundtrip ok=true

三点值得注意:

  • PCG 与 ChaCha8 都严格可复现:相同种子给出相同序列,这是把随机数写进测试断言的前提。
  • ChaCha8 可序列化:MarshalBinary 产出 48 字节(32 字节种子 + 16 字节状态),UnmarshalBinary 还原后与原对象相等,说明状态被完整保存。
  • v2 没有 Seed 方法:rand.NewPCG / rand.NewChaCha8 就是唯一的播种入口。grep -i seed 在 go doc math/rand/v2 里只能找到这两个构造函数。

4.2.3 math/rand.Seed 的时间线

v1 的 math/rand.Seed 有一个被反复误传的生命周期。本机 1.27 的文档直接写明了它:

$ GOTOOLCHAIN=go1.27.0 go doc -all math/rand.Seed | grep -iE "deprecat|no-op"
    Deprecated: As of Go 1.20 there is no reason to call Seed with a random ...
    As of Go 1.24 Seed is a no-op. To restore the previous behavior set ...

时间线因此是:

时间点math/rand.Seed 状态
Go 1.20 之前需要手动播种才能得到确定序列
Go 1.20 起全局源自动随机播种,Seed 被标记 Deprecated
Go 1.24 起Seed 是 no-op,调用它不再改变全局序列

对依赖「播种复现」的老代码,v2 的正解是换成 rand.New(rand.NewPCG(...)) 并显式传递 *Rand,而不是继续依赖全局函数。全局 rand.IntN 之类的函数在 v2 里仍在,但底层源不可播种,因此不适合做可复现测试。

4.2.4 iter.Pull 的代价:本机基准

iter.Pull 把 func(yield func(V) bool) 这种「推送式」迭代器翻成 next() (V, bool) 的「拉取式」接口。它的实现代价是:Pull 在独立 goroutine里跑原迭代器,用 channel 与调用方交接每次产出。这意味着每个元素都过一次同步握手。

写三份基准对比同一段逻辑(对 0..999 求和)的三种写法:

func BenchmarkDirect(b *testing.B) { /* 裸 for 循环 */ }
func BenchmarkRangeFunc(b *testing.B) { for v := range seq { s += v } }
func BenchmarkIterPull(b *testing.B) {
	next, stop := iter.Pull(seq)
	for { v, ok := next(); if !ok { break }; s += v }
	stop()
}

真实输出(Apple M1 Pro,GOTOOLCHAIN=go1.27.0):

BenchmarkDirect-10          	     200	       589.2 ns/op	       0 B/op	       0 allocs/op
BenchmarkRangeFunc-10       	     200	       518.8 ns/op	       0 B/op	       0 allocs/op
BenchmarkIterPull-10        	     200	     68348 ns/op	     208 B/op	       6 allocs/op
BenchmarkIterPullSeq2-10    	     200	         1.250 ns/op	       0 B/op	       0 allocs/op

换算成倍率:

写法时间相对 range分配
裸 for 循环589.2 ns/op1.14×0
range over func518.8 ns/op1×(基准)0
iter.Pull68348 ns/op约 132×208 B / 6 allocs

(BenchmarkIterPullSeq2 只 break 一次就退出,只量到「创建 + 立刻停止」的成本,不代表完整迭代。)

结论很直接:iter.Pull 的开销不是常数级的小负担,而是两个数量级。它换来的是「可以在迭代中途做任意控制流」——提前 stop、把迭代状态存进结构体、在两次 next 之间插入别的逻辑。range over func 做不到这些,因为控制权交给了 yield。

4.2.5 什么时候该用 Pull

判据只有一条:你能不能改用 range 表达?

场景推荐写法理由
单纯遍历求和 / 过滤range over func零分配,编译器可内联
迭代中途要 break 后继续处理range + breakrange 本身支持 break
需要把迭代器存进字段、跨函数传递状态iter.Pull拉取式才可持有
要在两次取值之间交错做别的事iter.Pullyield 模型下无法插入
热路径上的逐元素处理别用 Pull每元素一次 goroutine 握手

iter.Pull 还有一个语义细节:stop 必须被调用(用 defer 最稳),否则内部 goroutine 会泄漏。stop 之后继续调 next 是合法的,会持续返回零值与 false——这一点在 go doc iter.Pull 里写得很清楚:

It is valid to call next after reaching the end of the sequence
or after calling stop. These calls will continue to return the zero V and false.

最后提醒一句:math/rand/v2 的 ChaCha8 是密码学强度不足的伪随机源(文档明确写「For random numbers suitable for security-sensitive work, see crypto/rand」)。做可复现测试用 v2,做密钥 / token 用 crypto/rand,别混。

4.2.6 把 Pull 用对:可持有的有状态迭代器

range over func 的控制权在迭代器手里——yield 一返回 false 就结束,调用方拿不到「迭代到一半」这个状态。iter.Pull 的真正用途就是把这个状态变成可持有的对象。下面把无限斐波那契序列包成一个可持有、可关闭的 Scanner:

type Scanner struct {
	next func() (int, bool)
	stop func()
}

func NewScanner(seq iter.Seq[int]) *Scanner {
	n, s := iter.Pull(seq)
	return &Scanner{next: n, stop: s}
}

func (s *Scanner) Next() (int, bool) { return s.next() }
func (s *Scanner) Close()            { s.stop() }

func fib() iter.Seq[int] {
	return func(yield func(int) bool) {
		a, b := 0, 1
		for {
			if !yield(a) {
				return
			}
			a, b = b, a+b
		}
	}
}

fib 是无限序列——用 range 遍历它会永远不返回。而 Pull 让调用方可以按需取前 8 个再 Close。真实输出:

fib[0]=0
fib[1]=1
fib[2]=1
fib[3]=2
fib[4]=3
fib[5]=5
fib[6]=8
fib[7]=13
after stop: v=0 ok=false

注意最后一行:Close(即 stop)之后再调 Next,返回 v=0 ok=false,与文档承诺一致。这正是「无限序列 + 手动关闭」的模型:迭代器 goroutine 只在 stop 时才被回收。

4.2.7 组合子:Pull2 与泛型 N

iter.Pull2 是二元版本,签名对称:

pairs := func(yield func(string, int) bool) {
	yield("a", 1)
	yield("b", 2)
}
next, stop := iter.Pull2(pairs)
defer stop()
k, v, ok := next()

真实输出:

Pull2: a=1 ok=true

math/rand/v2 一侧的「组合子」是泛型辅助函数 rand.N:rand.N[T intType](n T) T 让你对任意整数类型直接用上界 N,避免手写 IntN(int(...)) 的来回转换。它的版本归属与整包一致(1.22),grep "^pkg math/rand/v2, func N" api/go1.22.txt 可核实。

把两个「组合子」放在一起看,会发现它们的取舍刚好相反:rand.N 是零成本抽象(编译期特化,运行时就是一条取模),iter.Pull 是高成本抽象(每元素一次 goroutine 握手)。Go 里「组合子」这个词不承诺廉价——用什么代价换什么表达力,得逐个看实现。

阅读导航:上一节:4.1 encoding/json/v2 实跑与迁移 · 下一节:4.3 Go 1.26/1.27 变更逐项与迁移检查表 。

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「golang」更多文章

  1. 《Go 语言编程实战》目录
  2. 《Go 语言编程实战》18.3 上线、观测与迭代
  3. 《Go 语言编程实战》18.2 故障演练