23. MongoDB Atlas Search 全文与向量搜索实战

Atlas Search 基于 Lucene 的全文检索、compound/autocomplete/fuzzy 查询、分析与高亮、Atlas Vector Search kNN 向量检索、分面聚合与生产调优

传统 MongoDB 的 $text 与正则查询在"相关度排序"“模糊纠错"“同义词"等搜索场景力不从心,而独立 Elasticsearch 集群又带来双写与运维负担。MongoDB Atlas Search 把 Lucene 索引直接嵌入 Atlas 托管集群,让开发者在同一个 MongoDB 实例上获得近实时全文搜索、拼写容错与向量检索能力,无需维护第二套搜索基础设施。

重要澄清:Atlas Search 与 Atlas Vector Search 是 MongoDB Atlas 托管服务的功能,自建 MongoDB 社区版/企业版不包含该模块。本文命令均面向 Atlas 集群,命令行演示通过 mongosh 连接 Atlas 后执行。

本文将介绍 Atlas Search 的 Lucene 基础、索引定义、全文/复合/自动补全/模糊查询、分析与高亮、向量检索、分面统计,以及与聚合管道的配合和生产调优。

1. Lucene 基础与 Atlas Search 架构

Atlas Search 的检索能力构建在 Apache Lucene 之上。Lucene 通过倒排索引(inverted index)把文本字段拆分为词项(term),并记录词项到文档的映射关系与位置信息,从而支持全文匹配、相关度打分与短语查询。

// 文档 -> Lucene 倒排索引(概念示意)
// "The quick brown fox jumps" 分词后形成:
// term "quick" -> docId: 1, positions: [1]
// term "brown" -> docId: 1, positions: [2]
// term "fox"   -> docId: 1, positions: [3]

Atlas Search 的架构分三层:Lucene 索引文件存储于 Atlas 集群的存储层;搜索节点(search nodes)承担索引与查询的计算;客户端通过 $search 聚合阶段与 $searchMeta 阶段与索引交互。

// 在聚合管道中使用 $search 进行全文检索
db.articles.aggregate([
  {
    $search: {
      index: "default",
      text: { query: "mongodb sharding", path: { wildcard: "*" } }
    }
  },
  { $limit: 10 },
  { $project: { title: 1, score: { $meta: "searchScore" } } }
])

与 $text 的核心区别在于:$text 依赖普通索引的文本功能,受限较多(每集合一个文本索引、不支持相关度细分、不支持同义词与模糊);Atlas Search 索引是独立的 Lucene 索引,每个集合可以定义多个搜索索引,且字段级分析与打分粒度更细。

能力$textAtlas Search
索引数量每集合 1 个可多个 search index
相关度打分固定公式BM25 可调可解释
模糊/同义词不支持fuzzy / synonyms 支持
高亮不支持支持
分面统计不支持facet 支持
向量检索不支持Atlas Vector Search

2. 创建 Search Index

Search Index 可以通过 Atlas 控制台、Atlas Admin API 或 $searchIndex 相关命令创建。以控制台定义为例,索引本质是一份 JSON 配置,核心是 mappings 与 analyzer。

// Search Index 定义(Atlas UI 或 API 提交)
{
  "mappings": {
    "dynamic": true,
    "fields": {
      "title": { "type": "string", "analyzer": "lucene.standard", "searchAnalyzer": "lucene.standard" },
      "tags":  { "type": "string", "analyzer": "lucene.keyword" },
      "views": { "type": "number" },
      "publishedAt": { "type": "date" }
    }
  },
  "analyzer": "lucene.standard"
}

字段类型(fields 类型映射)决定了该字段参与哪些搜索操作:string 参与文本检索,number/date 参与范围过滤与排序,geo 参与地理位置检索。dynamic: true 时未显式声明的字段会被自动推断类型,但自动推断的 string 字段默认按 lucene.standard 分词,往往不符合中文场景需求。

中文检索的关键:中文没有天然空格分词,lucene.standard 会把中文按字切分。对中文标题/正文建议使用 lucene.zh 或 lucene.smartcn 分析器(依赖 ICU/SmartChinese 分词),或使用 lucene.keyword + ngram 组合。

// 查看集合已有搜索索引
db.articles.aggregate([{ $listSearchIndexes: {} }])

// 通过 API 创建索引(Atlas Admin API 示例,使用 curl 语义)
// POST /api/atlas/v2/groups/{groupId}/clusters/{clusterName}/fts/indexes
// { "collectionName": "articles", "database": "blog", "name": "default",
//   "mappings": { "dynamic": true } }

索引创建是异步的,构建期间不影响既有写入。大集合的索引构建耗时取决于数据量,构建完成后新数据会近实时(秒级延迟)进入搜索索引。

3. 全文查询:text / compound / autocomplete / fuzzy

Atlas Search 查询通过聚合管道的 $search 阶段表达。最常用的是 text 与 compound。

3.1 text 全文检索与相关度

db.articles.aggregate([
  { $search: {
      index: "default",
      text: { query: "分布式事务", path: ["title", "content"] }
  }},
  { $limit: 20 },
  { $project: {
      title: 1,
      score: { $meta: "searchScore" }   // BM25 得分
  }}
])

query 支持字符串与数组(多词条 OR 语义),path 可以是单个字段、字段数组或 { wildcard: "*" }。BM25 得分可通过 $meta: "searchScore" 读取,用于自定义排序阈值。

3.2 compound 复合查询

compound 通过 must、mustNot、should、filter 组合多个子查询,实现"语义 OR + 必须匹配 + 过滤"的复杂表达式。

db.articles.aggregate([
  { $search: {
      index: "default",
      compound: {
        must: [ { text: { query: "mongodb 分片", path: "content" } } ],
        mustNot: [ { text: { query: "广告", path: "content" } } ],
        should: [
          { text: { query: "实战", path: "title", score: { boost: { value: 3 } } } }
        ],
        filter: [
          { range: { path: "views", gte: 1000 } },
          { range: { path: "publishedAt", lte: new Date("2026-09-27") } }
        ],
        minimumShouldMatch: 1
      }
  }},
  { $limit: 20 }
])

should 子句的命中会提升相关度得分,filter 子句只做过滤不参与打分。合理使用 minimumShouldMatch 可以避免 OR 过宽导致的相关度稀释。

3.3 autocomplete 自动补全与 fuzzy 模糊纠错

// 自动补全:输入前缀即时提示
db.products.aggregate([
  { $search: {
      index: "default",
      autocomplete: { query: "mongodb at", path: "name" }
  }},
  { $limit: 8 },
  { $project: { name: 1 } }
])

// 模糊检索:容忍 1 个编辑距离(拼写错误容错)
db.products.aggregate([
  { $search: {
      index: "default",
      text: { query: "mongob", path: "name", fuzzy: { maxEdits: 1, maxExpansions: 50 } }
  }},
  { $limit: 8 }
])

autocomplete 需要字段在索引中使用 lucene.ngram 或 edge-ngram 分析器(或默认的 autocomplete 类型)才能生效;fuzzy 基于 Levenshtein 编辑距离,maxEdits 可取 1 或 2,编辑距离越大约消耗更多资源。

操作适用场景关键参数
text全文相关度检索path、fuzzy
compound多条件组合打分must/should/filter/boost
autocomplete搜索框即时提示ngram 分析器
fuzzy拼写错误容错maxEdits、maxExpansions
range数值/日期过滤gte/lte
geoShape地理范围查询circle/geometry

4. 分析与高亮

分析(Analysis)是把原始文本变为可检索词项的过程。Atlas Search 的分析器体系继承 Lucene:analyzer 用于索引期分词,searchAnalyzer 用于查询期分词,两者不一致时可实现"查询更宽松、索引更精确"的效果。

// 自定义分析器:中文 SmartCN + 停用词
// (在索引定义中声明)
{
  "analyzer": "lucene.smartcn",
  "searchAnalyzer": "lucene.smartcn",
  "mappings": {
    "fields": {
      "title": { "type": "string", "analyzer": "lucene.smartcn", "searchAnalyzer": "lucene.smartcn" }
    }
  }
}

常用的 Lucene 分析器包括:lucene.standard(通用英语)、lucene.keyword(不分词,适合标签/ID)、lucene.english(词干还原)、lucene.smartcn/lucene.zh(中文)、lucene.ngram(模糊/补全)。

高亮(highlight)用于在搜索结果中标记命中片段:

db.articles.aggregate([
  { $search: {
      index: "default",
      text: { query: "分片集群", path: "content" },
      highlight: { path: "content" }
  }},
  { $limit: 5 },
  { $project: {
      title: 1,
      highlights: { $meta: "searchHighlights" }   // 命中片段数组
  }}
])

searchHighlights 返回形如 { text: "分片集群是 MongoDB 水平扩展的核心……", score: 2.1 } 的片段,前端渲染时用 <em> 包裹命中词即可。高亮不支持 highlight 之外的字段裁剪,且对动态映射字段仅返回命中的片段而非整字段。

高亮性能提示:对超长正文做高亮会显著增加查询耗时,建议限制 path 字段长度或在文档中预存摘要字段,对摘要做高亮而非全文。

5. Atlas Vector Search 向量检索

Atlas Vector Search 让 MongoDB 直接支持 k-最近邻(kNN)向量检索,是 RAG(检索增强生成)与语义搜索的核心能力。它基于 HNSW(Hierarchical Navigable Small World)图索引,把高维向量组织为分层邻居图,实现近似最近邻(ANN)检索。

// 创建向量搜索索引(Atlas 控制台 / API)
{
  "type": "vectorSearch",
  "fields": [
    {
      "type": "vector",
      "path": "embedding",
      "numDimensions": 768,
      "similarity": "cosine",
      "quantization": "scalar"
    }
  ]
}

相似度度量支持 cosine、euclidean 与 dotProduct。numDimensions 必须与实际向量维度一致(如 OpenAI text-embedding-3-small 为 1536 维,BGE-M3 为 1024 维)。

// 执行向量检索:查询向量 = "如何排查 MongoDB 慢查询"
db.articles.aggregate([
  {
    $vectorSearch: {
      index: "vector_index",
      path: "embedding",
      queryVector: [0.01, -0.02, ... /* 768 维查询向量 */],
      numCandidates: 100,     // HNSW 图遍历候选数
      limit: 10,              // 返回 top-k
      filter: { category: "database" }   // 可选的预过滤
    }
  },
  { $project: { title: 1, score: { $meta: "vectorSearchScore" } } }
])
参数作用调优建议
numCandidates图搜索候选池大小越大召回越准,消耗越高;经验值 20~100
limit返回条数与业务分页配合
filter预过滤条件减小搜索空间,提升准确率
similarity度量方式语义检索常用 cosine

向量检索与全文检索可以混合使用:先 $vectorSearch 做语义召回,再通过 $search 的 compound 叠加关键词约束,或用 $match/$facet 做后过滤。Atlas Vector Search 还可以配合 embedding 生成管道自动为新文档计算向量。

6. 分面与聚合(Facets & $searchMeta)

分面导航(facets)是电商、内容站的高频需求:在搜索结果旁展示"分类/价格区间/标签"的计数,供用户逐级筛选。Atlas Search 提供 $searchMeta 阶段,在不返回文档的前提下仅返回分面统计,成本远低于 $group。

db.products.aggregate([
  { $searchMeta: {
      index: "default",
      facet: {
        operator: { text: { query: "laptop", path: "name" } },
        facets: {
          brandFacet:  { type: "string", path: "brand" },
          priceFacet:  { type: "number", path: "price", boundaries: [0, 3000, 6000, 10000], default: "other" },
          ratingFacet: { type: "number", path: "rating", boundaries: [0, 3, 4, 5] }
        }
      }
  }}
])

返回结果大致为:facet.brandFacet.buckets: [{ _id: "Apple", count: 42 }, ...]。string 分面按词项计数,number 分面按 boundaries 划分区间,date 分面按时间范围计数。

$search 阶段本身也能与聚合管道其他阶段协同:先 $search 召回,再 $match 后过滤,然后 $group/$sort/$facet 做业务聚合,最后 $skip/$limit 分页。

// $search + 聚合:按品牌聚合销量 Top 商品
db.orders.aggregate([
  { $search: { index: "default", text: { query: "无线耳机", path: "productName" } } },
  { $group: { _id: "$brand", totalQty: { $sum: "$qty" } } },
  { $sort: { totalQty: -1 } },
  { $limit: 5 }
])

注意:$search 必须是聚合管道的第一阶段(或紧跟在少数允许的前置阶段之后)。需要先做权限/租户过滤时,可将过滤条件放进 $search 的 compound.filter,而不是在 $search 之后用 $match——后者无法利用 Lucene 索引裁剪。

7. 与 $text / 聚合管道的配合

$text 与 Atlas Search 可以共存:$text 适合简单的关键词过滤与存量代码兼容,Atlas Search 承担相关度检索。两者的结果在存储层共享同一份数据,但 $text 走普通文本索引,Atlas Search 走独立 Lucene 索引,因此两套查询的代价模型不同。

// 场景:先用 $text 快速过滤,再用 $search 排序
// (生产上更推荐:把过滤条件全部下沉到 $search.filter)
db.articles.aggregate([
  { $search: {
      index: "default",
      compound: {
        filter: [
          { text: { query: "数据库", path: "category" } },
          { range: { path: "publishedAt", gte: new Date("2026-01-01") } }
        ],
        should: [ { text: { query: "分片", path: "content", score: { boost: { value: 2 } } } } ]
      }
  }},
  { $facet: {
      total: [{ $count: "count" }],
      results: [{ $sort: { searchScore: -1 } }, { $skip: 0 }, { $limit: 10 }]
  }}
])

结合 $facet 可以一次性返回总数与分页结果,减少往返。注意 searchScore 字段在 $project 中需要用 $meta 取回,不能在 $facet 内直接排序。

8. 生产调优

8.1 索引定义与资源

  • 字段映射使用显式 fields 而非 dynamic: true,避免意外索引大文本字段
  • 不需要搜索的字段(如内部代码)不要映射进索引
  • 中文搜索优先 SmartCN;英文标题用 standard 即可

8.2 查询性能

  • $limit 尽量前移,避免检索大量文档后丢弃
  • filter 子句(日期、状态、租户)务必下沉到 $search,减少后续 $match
  • 避免在 $search 后对未索引字段排序;排序字段需在索引映射中声明为 sortable: true
// 声明可排序字段
{
  "mappings": {
    "fields": {
      "views": { "type": "number", "sortable": true },
      "publishedAt": { "type": "date", "sortable": true }
    }
  }
}

8.3 容量与延迟

  • 向量索引占用显著内存:768 维、1M 文档约需要数百 MB~GB 级内存,选择合适 quantization(scalar/int8 压缩)降低内存
  • 高写入场景考虑在低峰期重建向量索引,避免 HNSW 图频繁更新
  • 观察 Atlas 的 Search 指标面板:查询延迟、索引构建进度、被拒绝的查询数
指标含义参考阈值
查询延迟 p95$search 阶段耗时< 50ms
索引延迟文档到可检索的时延秒级
内存使用率Lucene/向量索引驻留内存不触发分页
高亮命中率高亮片段产出质量视业务

9. 典型场景总结

Atlas Search 的适用面可以概括为三条:需要全文相关度与拼写容错的站内搜索;需要分面导航的电商/内容平台;需要语义检索与向量召回的新型 AI 应用。它无法替代专用搜索引擎在极端规模下的深度定制,但对绝大多数 MongoDB 用户而言,是在不引入第二套存储的前提下获得企业级检索能力的最短路径。

延伸阅读

继续阅读

探索更多技术文章

浏览归档,发现更多关于系统设计、工具链和工程实践的内容。

全部文章 返回首页

「mongodb」更多文章

  1. 27. MongoDB 多租户与隔离架构设计
  2. 26. MongoDB 监控与可观测性实践
  3. 25. MongoDB WiredTiger 存储引擎深入