ES 不是 NoSQL 数据库的简单替代,它把同一份数据拆成倒排索引、Doc Values、BKD 树等多种结构,Mapping 一旦创建就很难改。因此数据建模必须前置思考:字段用 text 还是 keyword,嵌套怎么建模,聚合怎么存,这些都直接决定检索能力与性能上限。
1. Mapping 基础
1.1 Mapping 是什么
Mapping 是索引的字段结构定义,等价于关系库的表结构,描述每个字段的类型、分析器、是否索引、是否存储 Doc Values。
| 概念 | ES 对应 | 关系库类比 |
|---|---|---|
| 索引 | Index | 数据库/表 |
| 文档 | Document | 行 |
| 字段 | Field | 列 |
| Mapping | Mapping | 表结构定义 |
| 分析器 | Analyzer | 无(全文索引) |
1.2 显式创建 Mapping
curl -X PUT 'http://localhost:9200/products?pretty' \
-H 'Content-Type: application/json' \
-d '{
"mappings": {
"properties": {
"name": { "type": "text", "analyzer": "ik_max_word" },
"brand": { "type": "keyword" },
"price": { "type": "float" },
"stock": { "type": "integer" },
"on_sale": { "type": "boolean" },
"tags": { "type": "keyword" },
"created_at": { "type": "date", "format": "yyyy-MM-dd HH:mm:ss" }
}
}
}'
1.3 字段类型总览
| 分类 | 类型 | 用途 |
|---|---|---|
| 全文 | text | 分词检索 |
| 精确 | keyword | 过滤/排序/聚合 |
| 数值 | long/integer/float/half_float | 数值范围 |
| 时间 | date | 日期范围/排序 |
| 结构 | object/nested | 嵌套对象 |
| 特殊 | geo_point/ip/join | 地理/网络/父子 |
2. 核心字段类型
2.1 text 与 keyword 的选择
| 对比 | text | keyword |
|---|---|---|
| 索引方式 | 分词 | 整体单值 |
| 匹配 | match 全文 | term 精确 |
| 聚合 | 不支持 | 支持 |
| 排序 | 不支持 | 支持 |
| 典型值 | 标题、正文 | 状态、品牌、ID |
{
"properties": {
"status": { "type": "keyword" },
"title": { "type": "text", "analyzer": "ik_max_word" }
}
}
最常见的建模错误是把所有字段设成 text,导致排序聚合全不可用,或把所有字段设成 keyword 导致无法全文检索。规则是:需要分词匹配的用 text,需要精确过滤与聚合的用 keyword。
2.2 数值与浮点精度
ES 数值类型遵循 Lucene BKD 树,但求和与统计依赖 Doc Values。金额类数据建议用 scaled_float 或存分为单位的 integer,避免 float 精度误差。
{
"properties": {
"amount_cents": { "type": "integer" },
"rating": { "type": "half_float" }
}
}
2.3 date 与 format
{
"properties": {
"created_at": { "type": "date", "format": "strict_date_optional_time||epoch_millis" }
}
}
date 字段底层存为自 epoch 的毫秒数,format 只影响读写展示。范围查询、排序、日期直方图聚合都依赖正确的 date 类型,字符串存 date 是常见反模式。
3. dynamic 策略与显式映射
3.1 三种 dynamic 行为
| 策略 | 行为 | 适用 |
|---|---|---|
| true | 自动推断并添加字段 | 开发期 |
| false | 忽略新字段,不索引 | 字段可扩展但不用 |
| strict | 遇到新字段直接拒绝 | 严格 schema |
{
"mappings": {
"dynamic": "strict",
"properties": {
"name": { "type": "text" }
}
}
}
3.2 自动推断的坑
dynamic 为 true 时,ES 会为第一个文档推断类型,同名字段后续类型不一致会被拒绝或产生索引错误。
{
"mappings": {
"dynamic_templates": [
{
"strings_as_keyword": {
"match_mapping_type": "string",
"mapping": { "type": "keyword" }
}
},
{
"longs_as_long": {
"match_mapping_type": "long",
"mapping": { "type": "long" }
}
}
]
}
}
生产环境建议 dynamic 设为 false 或 strict,用 dynamic_templates 兜底常见字段形态,避免字段类型漂移。
3.3 字段命名规范
字段名使用 snake_case 或 camelCase 统一风格,避免字段名过长。ES 不限制字段数量,但字段越多 mapping 与查询开销越大,动辄上千字段的文档应评估是否扁平化。
4. Analyzer 与多字段设计
4.1 为每个字段配分析器
{
"properties": {
"title": {
"type": "text",
"fields": {
"keyword": { "type": "keyword", "ignore_above": 256 },
"ik": { "type": "text", "analyzer": "ik_max_word" },
"en": { "type": "text", "analyzer": "english" }
}
}
}
}
multi-fields 让同一字段以多种方式索引:title 默认 standard 分词,title.keyword 支持精确排序,title.ik 支持中文切分,title.en 支持英文词干。查询时按需选择子字段。
4.2 中文场景的多字段
{
"mappings": {
"properties": {
"content": {
"type": "text",
"analyzer": "ik_max_word",
"search_analyzer": "ik_smart"
}
}
}
}
索引侧用 ik_max_word 全切分保召回,查询侧用 ik_smart 精切分保精度,是中文搜索的标准做法。分析器配置方法详见《倒排索引与分词原理》。
4.3 ignore_above 与空值
keyword 字段设置 ignore_above 后,超过长度的值不索引,避免超长字符串撑爆词典。空字符串、null 数组在聚合与过滤中的行为不同,建模时要明确默认值语义。
5. 嵌套与父子文档
5.1 object 的扁平化陷阱
{
"properties": {
"reviews": {
"type": "nested",
"properties": {
"author": { "type": "keyword" },
"rating": { "type": "integer" },
"content": { "type": "text" }
}
}
}
}
默认 object 类型把嵌套对象扁平化为独立字段,数组对象会丢失对象边界,例如两个评价的 author 与 rating 会交叉匹配。需要保持对象独立性时用 nested。
5.2 nested 与 join 对比
| 维度 | nested | join(父子) |
|---|---|---|
| 数据存储 | 同文档隐藏块 | 父子独立文档 |
| 更新粒度 | 整篇重写 | 子文档独立更新 |
| 查询能力 | nested_query | has_child/has_parent |
| 性能 | 快,单分片内 | 较慢,跨文档 |
| 父子数量 | 无限制 | 父文档扇出有限 |
{
"mappings": {
"properties": {
"question": { "type": "join", "relations": { "question": "answer" } }
}
}
}
5.3 建模选型建议
一对多且不频繁更新的内嵌列表用 nested;子文档大量独立更新(如订单明细状态)用 join。join 查询 has_child 性能开销大,千级扇出即可能劣化,能扁平化就扁平化。
6. 数值与聚合建模
6.1 Doc Values 与聚合
{
"properties": {
"price": {
"type": "float",
"doc_values": true
}
}
}
聚合与排序依赖 Doc Values 列存储。默认开启,若确认某字段只做检索不做聚合排序,可关闭 doc_values 省磁盘;反之需聚合的字段必须开启。
6.2 稀疏字段与成本
文档间字段差异大会形成稀疏 Doc Values。大量稀疏字段会浪费存储与遍历开销。建模上应尽量让同索引文档字段结构一致,不同业务的异构数据拆到独立索引。
6.3 聚合建模技巧
| 需求 | 建模建议 |
|---|---|
| 按品牌统计 | brand 用 keyword |
| 按时间统计 | date 字段,calendar_interval |
| 多字段组合聚合 | 预聚合宽表 |
| 大量唯一值聚合 | 控制基数,必要时拆索引 |
{
"aggs": {
"by_brand": { "terms": { "field": "brand", "size": 20 } }
}
}
7. Mapping 演进与重构
7.1 字段不可变与 reindex
字段类型创建后不可修改,调整分析器或类型必须重建索引。标准流程是新建索引、reindex、切换别名。
# 新建新版本索引
curl -X PUT 'http://localhost:9200/products_v2?pretty' \
-H 'Content-Type: application/json' \
-d '{"mappings": {"properties": {"name": {"type": "text", "analyzer": "ik_max_word"}}}}'
# 迁移数据
curl -X POST 'http://localhost:9200/_reindex?pretty' \
-H 'Content-Type: application/json' \
-d '{"source": {"index": "products_v1"}, "dest": {"index": "products_v2"}}'
7.2 别名切换实现零停机
curl -X POST 'http://localhost:9200/_aliases?pretty' \
-H 'Content-Type: application/json' \
-d '{
"actions": [
{ "remove": { "index": "products_v1", "alias": "products" } },
{ "add": { "index": "products_v2", "alias": "products" } }
]
}'
应用层只访问别名 products,reindex 完成后原子切换,实现零停机迁移。alias 还支持按日期过滤与写索引分离。
7.3 reindex 的性能与注意
curl -X POST 'http://localhost:9200/_reindex?pretty' \
-H 'Content-Type: application/json' \
-d '{
"source": { "index": "products_v1", "size": 5000 },
"dest": { "index": "products_v2" }
}'
reindex 本质是滚动读取加批量写入,可在 source 用 query 过滤只迁移部分数据。大量数据迁移建议在低峰执行,并临时拉长 refresh 与关闭副本以提速。
7.4 Mapping 版本与文档规范
维护一份 mapping 变更记录,每个索引带版本号后缀(products_v1/v2)。删除旧索引前确认数据已备份且无查询引用,避免线上误删。
8. 总结
| 环节 | 要点 |
|---|---|
| Mapping 即表结构 | 创建后字段难改,建模前置 |
| 类型选择 | text 全文、keyword 精确、date 时间、nested 保边界 |
| dynamic | 生产用 false/strict,配合 dynamic_templates |
| 多字段 | index/search analyzer 分离,multi-fields 按场景 |
| 嵌套选择 | 频繁更新用 join,否则 nested 优先 |
| 聚合建模 | Doc Values 列存,字段结构尽量一致 |
| 演进重构 | reindex + 别名切换零停机 |
| 命名规范 | 统一字段风格,控制字段数量 |
数据建模决定了 ES 的能力边界:检索靠倒排、排序靠 Doc Values、时间靠 date,mapping 一旦定型,后续优化成本极高。建模时应反复问自己每个字段的读法:要全文匹配、精确过滤、排序还是聚合。底部分词逻辑见《倒排索引与分词原理》,查询组合见《Query DSL 与相关性打分》。
延伸阅读
继续阅读
探索更多技术文章
浏览归档,发现更多关于系统设计、工具链和工程实践的内容。