自定义分析器与分词:字符过滤器、分词器、过滤器链与拼音同义词

系统讲解 Elasticsearch 自定义分析器:字符过滤器、分词器与 token 过滤器链的组装方式,同义词与拼音插件配置,analyze API 调试手段,以及索引期与查询期分析器的分离设计。

全文检索的质量在写入那一刻就已经决定了大半:文档被切成哪些词、是否做了归一化、同义词有没有展开,直接决定了之后能召回什么。Elasticsearch 把这一过程抽象成分析器,由字符过滤器、分词器、token 过滤器三段串成一条流水线。内置分析器能覆盖多数场景,但中文分词、拼音检索、业务同义词这些需求必须自己组装。本文从三段式结构讲起,覆盖各段的可选组件、拼音与同义词配置、analyze 调试,以及索引期与查询期分析器分离的设计原则。

1. 分析器的三段式结构

一句话总结: 分析器由字符过滤器、分词器、token 过滤器依次串联,前一段的输出是后一段的输入。

1.1 三段的职责

  • 字符过滤器(character filter):在切词之前处理原始字符串,做整体替换,例如去掉 HTML 标签、把 & 换成 and。
  • 分词器(tokenizer):把字符串切成一个个 token,同时记录每个 token 的起止偏移量。
  • token 过滤器(token filter):对 token 流做逐个处理,可以小写化、去停用词、加同义词、做词干还原。

三段中只有分词器是必需的,字符过滤器与 token 过滤器都可以没有。一个分析器最多一个分词器、可以零到多个字符过滤器与 token 过滤器。

1.2 内置分析器的对照

# 用 analyze API 观察不同内置分析器的切词差异
curl -X POST "localhost:9200/_analyze?pretty" -H 'Content-Type: application/json' -d '
{
  "analyzer": "standard",
  "text": "The Quick Brown-Fox 2026"
}'

standard 会输出 the、quick、brown、fox、2026,连字符被当分隔符。若换成 keyword,整句作为一个 token。若换成 simple,数字 2026 会被丢掉。理解这些差异是选型的前提。

1.3 索引期与查询期的分离

analyzer 决定文档怎么切,search_analyzer 决定查询词怎么切。默认查询沿用索引分析器,但在需要「索引细、查询粗」或反之的场景,必须显式分开设置,后文会专门展开。

2. 字符过滤器

一句话总结: 字符过滤器在切词前对原文做替换,常见的是 HTML 剥离与字符映射。

2.1 html_strip

curl -X PUT "localhost:9200/articles" -H 'Content-Type: application/json' -d '
{
  "settings": {
    "analysis": {
      "analyzer": {
        "html_analyzer": {
          "type": "custom",
          "char_filter": ["html_strip"],
          "tokenizer": "standard",
          "filter": ["lowercase", "stop"]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "content": { "type": "text", "analyzer": "html_analyzer" }
    }
  }
}'

html_strip 会把 <b>粗体</b> 处理成 粗体,同时把 &amp; 解码成 &。

2.2 mapping 字符过滤器

curl -X PUT "localhost:9200/products" -H 'Content-Type: application/json' -d '
{
  "settings": {
    "analysis": {
      "char_filter": {
        "normalize_symbols": {
          "type": "mapping",
          "mappings": ["٠ => 0", "١ => 1", "٢ => 2", "- => -", "+ => +"]
        }
      },
      "analyzer": {
        "product_analyzer": {
          "type": "custom",
          "char_filter": ["normalize_symbols"],
          "tokenizer": "standard",
          "filter": ["lowercase"]
        }
      }
    }
  }
}'

mapping 过滤器支持 => 单字符映射,也支持用逗号分隔的等价组(如 "a,b => c" 表示 a 与 b 都映射为 c)。

2.3 pattern_replace 正则替换

{
  "char_filter": {
    "remove_dashes": {
      "type": "pattern_replace",
      "pattern": "(\\d)-(\\d)",
      "replacement": "$1$2"
    }
  }
}

上面的规则把 138-0013-8000 归一化成 13800138000,适合电话号码检索。正则必须小心灾难性回溯,避免在大文本上卡死。

3. 分词器的选型

一句话总结: 中文必须用专门分词器,英文默认 standard 已够用,结构化字段用 keyword。

3.1 standard 与 whitespace 的取舍

curl -X POST "localhost:9200/_analyze?pretty" -H 'Content-Type: application/json' -d '
{ "tokenizer": "whitespace", "text": "hello,world foo-bar" }'

whitespace 输出 hello,world 与 foo-bar 两个 token,标点原样保留。这适合日志、代码片段这类「标点是语义一部分」的场景。

3.2 中文分词器

# 安装 IK 插件后定义自定义词典路径
curl -X PUT "localhost:9200/news" -H 'Content-Type: application/json' -d '
{
  "settings": {
    "analysis": {
      "analyzer": {
        "ik_smart_custom": { "type": "custom", "tokenizer": "ik_smart",
          "filter": ["lowercase"] },
        "ik_max_word_custom": { "type": "custom", "tokenizer": "ik_max_word",
          "filter": ["lowercase"] }
      }
    }
  },
  "mappings": {
    "properties": {
      "title": { "type": "text", "analyzer": "ik_max_word",
        "search_analyzer": "ik_smart" }
    }
  }
}'

IK 提供两种模式:ik_max_word 做最细粒度切分,召回高但索引大;ik_smart 做粗粒度切分,精度高但召回略低。经典组合是索引用 max_word、查询用 smart,兼顾召回与精度。

3.3 keyword 与 pattern 分词器

curl -X POST "localhost:9200/_analyze?pretty" -H 'Content-Type: application/json' -d '
{
  "tokenizer": { "type": "pattern", "pattern": "[/\\\\-]" },
  "text": "2026/10/01-release"
}'

输出 2026、10、01、release。pattern 分词器适合路径、版本号这类有固定分隔符的字段。

4. Token 过滤器链

一句话总结: token 过滤器是分析器里最灵活的一段,小写、停用词、词干、ngram 都在这里完成。

4.1 常用过滤器

curl -X PUT "localhost:9200/docs" -H 'Content-Type: application/json' -d '
{
  "settings": {
    "analysis": {
      "filter": {
        "english_stop": { "type": "stop",
          "stopwords": ["the", "a", "an", "of"] },
        "english_stem": { "type": "stemmer", "language": "english" }
      },
      "analyzer": {
        "english_custom": { "type": "custom", "tokenizer": "standard",
          "filter": ["lowercase", "asciifolding", "english_stop", "english_stem"] }
      }
    }
  }
}'

过滤器顺序影响结果:先小写再去停用词,才能匹配到小写形式的停用词表。asciifolding 把 café 转成 cafe,对多语言混排很有用。

4.2 ngram 与 edge_ngram

curl -X PUT "localhost:9200/autocomplete" -H 'Content-Type: application/json' -d '
{
  "settings": {
    "analysis": {
      "filter": {
        "edge_ngram_filter": { "type": "edge_ngram", "min_gram": 2, "max_gram": 15 }
      },
      "analyzer": {
        "autocomplete_index": { "type": "custom", "tokenizer": "standard",
          "filter": ["lowercase", "edge_ngram_filter"] },
        "autocomplete_search": { "type": "custom", "tokenizer": "standard",
          "filter": ["lowercase"] }
      }
    }
  },
  "mappings": {
    "properties": {
      "name": { "type": "text", "analyzer": "autocomplete_index",
        "search_analyzer": "autocomplete_search" }
    }
  }
}'

这是索引期与查询期分离的经典案例:索引期把 elasticsearch 切成 el、ela、elas……,查询期只用普通分词,于是输入 ela 就能命中。

4.3 过滤器链的性能代价

ngram 的 max_gram 每加 1,索引体积近似线性增长。同义词展开会显著放大 token 数量。生产上要在召回效果与索引成本之间实测取平衡。

5. 同义词与拼音

一句话总结: 同义词提升召回,拼音支持首字母与全拼检索,两者都通过 token 过滤器挂到分析器上。

5.1 同义词过滤器

curl -X PUT "localhost:9200/synonyms_demo" -H 'Content-Type: application/json' -d '
{
  "settings": {
    "analysis": {
      "filter": {
        "my_synonyms": { "type": "synonym",
          "synonyms": ["手机, 移动电话, 蜂窝电话", "笔记本 => 笔记本电脑",
            "es, elasticsearch"] }
      },
      "analyzer": {
        "synonym_analyzer": { "type": "custom", "tokenizer": "ik_smart",
          "filter": ["lowercase", "my_synonyms"] }
      }
    }
  }
}'

a, b, c 是等价语法,三者互相映射;a => b 是单向映射,只在查询时把 a 换成 b。生产上词典通常放在 synonyms.txt 里,用 synonyms_path 引用,并配合 _reload_search_analyzers 热更新:

curl -X POST "localhost:9200/synonyms_demo/_reload_search_analyzers?pretty"

5.2 拼音检索

curl -X PUT "localhost:9200/contacts" -H 'Content-Type: application/json' -d '
{
  "settings": {
    "analysis": {
      "filter": {
        "pinyin_filter": { "type": "pinyin", "keep_first_letter": true,
          "keep_full_pinyin": true, "keep_original": true,
          "remove_duplicated_term": true }
      },
      "analyzer": {
        "pinyin_analyzer": { "type": "custom", "tokenizer": "keyword",
          "filter": ["pinyin_filter"] }
      }
    }
  },
  "mappings": {
    "properties": {
      "name": { "type": "text", "analyzer": "ik_smart",
        "fields": { "pinyin": { "type": "text", "analyzer": "pinyin_analyzer" } } }
    }
  }
}'

这里用了 multi-field:name 走中文分词,name.pinyin 走拼音分析器。查询时用 multi_match 同时打两个字段,就能同时支持中文与拼音输入。

5.3 同义词与拼音的组合顺序

顺序错会导致同义词失效:同义词词典里写的是汉字,若先转拼音,汉字形态已经不存在了。因此 filter 数组里 synonym 必须在 pinyin 之前。

6. analyze API 调试

一句话总结: analyze API 是分析器调试的唯一权威工具,能看到每一步的 token 与偏移量。

6.1 用指定分析器观察

curl -X POST "localhost:9200/synonyms_demo/_analyze?pretty" -H 'Content-Type: application/json' -d '
{ "analyzer": "synonym_analyzer", "text": "手机" }'

返回的每个 token 都带 position、start_offset、end_offset,可以用来核对同义词是否真的展开了。

6.2 逐段验证

# 只验证分词器
curl -X POST "localhost:9200/_analyze?pretty" -H 'Content-Type: application/json' -d '
{ "tokenizer": "ik_smart", "text": "中华人民共和国" }'

# 验证过滤器链(不给分词器则默认 standard)
curl -X POST "localhost:9200/_analyze?pretty" -H 'Content-Type: application/json' -d '
{ "filter": ["lowercase", "my_synonyms"], "text": "ES 与 Elasticsearch" }'

6.3 用 explain 模式看中间结果

curl -X POST "localhost:9200/synonyms_demo/_analyze?pretty" -H 'Content-Type: application/json' -d '
{ "field": "name", "text": "手机", "explain": true }'

6.4 用 termvectors 看已索引的词

curl -X GET "localhost:9200/contacts/_termvectors/1?fields=name.pinyin&pretty"

当怀疑「analyze 结果对但搜不到」时,termvectors 往往能揭示真相:文档里实际索引的 token 与预期不符。

7. 索引期与查询期分析器分离

一句话总结: 索引期追求召回,查询期追求精度,两者分离是自定义分析器设计的核心原则。

7.1 为什么必须分离

以自动补全为例:写入 elasticsearch 要切成 el、ela 等前缀,查询 ela 却不应被切成前缀,否则会匹配到所有以 el 开头的词。若两者用同一个分析器,效果必然打折。

7.2 配置方式

{
  "properties": {
    "title": { "type": "text", "analyzer": "ik_max_word",
      "search_analyzer": "ik_smart" },
    "name": { "type": "text", "analyzer": "autocomplete_index",
      "search_analyzer": "autocomplete_search" }
  }
}

7.3 多字段策略

curl -X PUT "localhost:9200/multi_field_demo" -H 'Content-Type: application/json' -d '
{
  "mappings": {
    "properties": {
      "content": { "type": "text", "analyzer": "ik_smart",
        "fields": {
          "keyword": { "type": "keyword", "ignore_above": 256 },
          "pinyin": { "type": "text", "analyzer": "pinyin_analyzer" }
        } }
    }
  }
}'

一个字段同时支持中文分词检索、精确匹配与拼音检索,代价是索引体积变大。按实际查询模式裁剪子字段,是控制成本的关键。

8. 总结

环节要点
三段式结构字符过滤器整体替换、分词器切词、token 过滤器逐个改写
字符过滤器html_strip 剥离标签,mapping 做符号归一,pattern_replace 正则替换
分词器选型中文用 IK 或 jieba,结构化字段用 keyword,路径用 pattern
过滤器链小写、asciifolding、停用词、词干、ngram 依次生效,顺序敏感
同义词等价与映射两种语法,外置词典可热更新
拼音pinyin 插件转全拼与首字母,常与中文分析器做成 multi-field
调试手段analyze 看切词,termvectors 看实际索引,explain 看逐段变化
索引查询分离索引期保召回、查询期保精度,用 search_analyzer 与多字段实现

分析器是检索质量的源头,选错分词器或漏配同义词,后面再多的打分调优也补不回来。把三段流水线搭对、用 analyze API 验证、用多字段兼顾多种查询模式,文本检索就成功了一半。下一篇我们讲如何用搜索模板把查询参数化,让复杂 DSL 可以复用与集中管理。

延伸阅读

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「elasticsearch」更多文章

  1. 可搜索快照与冻结层:把冷数据放进对象存储还能查
  2. 分页与深度分页:from/size、search_after、PIT 与 scroll
  3. 嵌套与父子关联查询:nested、join 字段与性能取舍