传统 MongoDB 的 $text 与正则查询在"相关度排序"“模糊纠错"“同义词"等搜索场景力不从心,而独立 Elasticsearch 集群又带来双写与运维负担。MongoDB Atlas Search 把 Lucene 索引直接嵌入 Atlas 托管集群,让开发者在同一个 MongoDB 实例上获得近实时全文搜索、拼写容错与向量检索能力,无需维护第二套搜索基础设施。
重要澄清:Atlas Search 与 Atlas Vector Search 是 MongoDB Atlas 托管服务的功能,自建 MongoDB 社区版/企业版不包含该模块。本文命令均面向 Atlas 集群,命令行演示通过
mongosh连接 Atlas 后执行。
本文将介绍 Atlas Search 的 Lucene 基础、索引定义、全文/复合/自动补全/模糊查询、分析与高亮、向量检索、分面统计,以及与聚合管道的配合和生产调优。
1. Lucene 基础与 Atlas Search 架构
Atlas Search 的检索能力构建在 Apache Lucene 之上。Lucene 通过倒排索引(inverted index)把文本字段拆分为词项(term),并记录词项到文档的映射关系与位置信息,从而支持全文匹配、相关度打分与短语查询。
// 文档 -> Lucene 倒排索引(概念示意)
// "The quick brown fox jumps" 分词后形成:
// term "quick" -> docId: 1, positions: [1]
// term "brown" -> docId: 1, positions: [2]
// term "fox" -> docId: 1, positions: [3]
Atlas Search 的架构分三层:Lucene 索引文件存储于 Atlas 集群的存储层;搜索节点(search nodes)承担索引与查询的计算;客户端通过 $search 聚合阶段与 $searchMeta 阶段与索引交互。
// 在聚合管道中使用 $search 进行全文检索
db.articles.aggregate([
{
$search: {
index: "default",
text: { query: "mongodb sharding", path: { wildcard: "*" } }
}
},
{ $limit: 10 },
{ $project: { title: 1, score: { $meta: "searchScore" } } }
])
与 $text 的核心区别在于:$text 依赖普通索引的文本功能,受限较多(每集合一个文本索引、不支持相关度细分、不支持同义词与模糊);Atlas Search 索引是独立的 Lucene 索引,每个集合可以定义多个搜索索引,且字段级分析与打分粒度更细。
| 能力 | $text | Atlas Search |
|---|---|---|
| 索引数量 | 每集合 1 个 | 可多个 search index |
| 相关度打分 | 固定公式 | BM25 可调可解释 |
| 模糊/同义词 | 不支持 | fuzzy / synonyms 支持 |
| 高亮 | 不支持 | 支持 |
| 分面统计 | 不支持 | facet 支持 |
| 向量检索 | 不支持 | Atlas Vector Search |
2. 创建 Search Index
Search Index 可以通过 Atlas 控制台、Atlas Admin API 或 $searchIndex 相关命令创建。以控制台定义为例,索引本质是一份 JSON 配置,核心是 mappings 与 analyzer。
// Search Index 定义(Atlas UI 或 API 提交)
{
"mappings": {
"dynamic": true,
"fields": {
"title": { "type": "string", "analyzer": "lucene.standard", "searchAnalyzer": "lucene.standard" },
"tags": { "type": "string", "analyzer": "lucene.keyword" },
"views": { "type": "number" },
"publishedAt": { "type": "date" }
}
},
"analyzer": "lucene.standard"
}
字段类型(fields 类型映射)决定了该字段参与哪些搜索操作:string 参与文本检索,number/date 参与范围过滤与排序,geo 参与地理位置检索。dynamic: true 时未显式声明的字段会被自动推断类型,但自动推断的 string 字段默认按 lucene.standard 分词,往往不符合中文场景需求。
中文检索的关键:中文没有天然空格分词,
lucene.standard会把中文按字切分。对中文标题/正文建议使用lucene.zh或lucene.smartcn分析器(依赖 ICU/SmartChinese 分词),或使用lucene.keyword+ ngram 组合。
// 查看集合已有搜索索引
db.articles.aggregate([{ $listSearchIndexes: {} }])
// 通过 API 创建索引(Atlas Admin API 示例,使用 curl 语义)
// POST /api/atlas/v2/groups/{groupId}/clusters/{clusterName}/fts/indexes
// { "collectionName": "articles", "database": "blog", "name": "default",
// "mappings": { "dynamic": true } }
索引创建是异步的,构建期间不影响既有写入。大集合的索引构建耗时取决于数据量,构建完成后新数据会近实时(秒级延迟)进入搜索索引。
3. 全文查询:text / compound / autocomplete / fuzzy
Atlas Search 查询通过聚合管道的 $search 阶段表达。最常用的是 text 与 compound。
3.1 text 全文检索与相关度
db.articles.aggregate([
{ $search: {
index: "default",
text: { query: "分布式事务", path: ["title", "content"] }
}},
{ $limit: 20 },
{ $project: {
title: 1,
score: { $meta: "searchScore" } // BM25 得分
}}
])
query 支持字符串与数组(多词条 OR 语义),path 可以是单个字段、字段数组或 { wildcard: "*" }。BM25 得分可通过 $meta: "searchScore" 读取,用于自定义排序阈值。
3.2 compound 复合查询
compound 通过 must、mustNot、should、filter 组合多个子查询,实现"语义 OR + 必须匹配 + 过滤"的复杂表达式。
db.articles.aggregate([
{ $search: {
index: "default",
compound: {
must: [ { text: { query: "mongodb 分片", path: "content" } } ],
mustNot: [ { text: { query: "广告", path: "content" } } ],
should: [
{ text: { query: "实战", path: "title", score: { boost: { value: 3 } } } }
],
filter: [
{ range: { path: "views", gte: 1000 } },
{ range: { path: "publishedAt", lte: new Date("2026-09-27") } }
],
minimumShouldMatch: 1
}
}},
{ $limit: 20 }
])
should 子句的命中会提升相关度得分,filter 子句只做过滤不参与打分。合理使用 minimumShouldMatch 可以避免 OR 过宽导致的相关度稀释。
3.3 autocomplete 自动补全与 fuzzy 模糊纠错
// 自动补全:输入前缀即时提示
db.products.aggregate([
{ $search: {
index: "default",
autocomplete: { query: "mongodb at", path: "name" }
}},
{ $limit: 8 },
{ $project: { name: 1 } }
])
// 模糊检索:容忍 1 个编辑距离(拼写错误容错)
db.products.aggregate([
{ $search: {
index: "default",
text: { query: "mongob", path: "name", fuzzy: { maxEdits: 1, maxExpansions: 50 } }
}},
{ $limit: 8 }
])
autocomplete 需要字段在索引中使用 lucene.ngram 或 edge-ngram 分析器(或默认的 autocomplete 类型)才能生效;fuzzy 基于 Levenshtein 编辑距离,maxEdits 可取 1 或 2,编辑距离越大约消耗更多资源。
| 操作 | 适用场景 | 关键参数 |
|---|---|---|
| text | 全文相关度检索 | path、fuzzy |
| compound | 多条件组合打分 | must/should/filter/boost |
| autocomplete | 搜索框即时提示 | ngram 分析器 |
| fuzzy | 拼写错误容错 | maxEdits、maxExpansions |
| range | 数值/日期过滤 | gte/lte |
| geoShape | 地理范围查询 | circle/geometry |
4. 分析与高亮
分析(Analysis)是把原始文本变为可检索词项的过程。Atlas Search 的分析器体系继承 Lucene:analyzer 用于索引期分词,searchAnalyzer 用于查询期分词,两者不一致时可实现"查询更宽松、索引更精确"的效果。
// 自定义分析器:中文 SmartCN + 停用词
// (在索引定义中声明)
{
"analyzer": "lucene.smartcn",
"searchAnalyzer": "lucene.smartcn",
"mappings": {
"fields": {
"title": { "type": "string", "analyzer": "lucene.smartcn", "searchAnalyzer": "lucene.smartcn" }
}
}
}
常用的 Lucene 分析器包括:lucene.standard(通用英语)、lucene.keyword(不分词,适合标签/ID)、lucene.english(词干还原)、lucene.smartcn/lucene.zh(中文)、lucene.ngram(模糊/补全)。
高亮(highlight)用于在搜索结果中标记命中片段:
db.articles.aggregate([
{ $search: {
index: "default",
text: { query: "分片集群", path: "content" },
highlight: { path: "content" }
}},
{ $limit: 5 },
{ $project: {
title: 1,
highlights: { $meta: "searchHighlights" } // 命中片段数组
}}
])
searchHighlights 返回形如 { text: "分片集群是 MongoDB 水平扩展的核心……", score: 2.1 } 的片段,前端渲染时用 <em> 包裹命中词即可。高亮不支持 highlight 之外的字段裁剪,且对动态映射字段仅返回命中的片段而非整字段。
高亮性能提示:对超长正文做高亮会显著增加查询耗时,建议限制
path字段长度或在文档中预存摘要字段,对摘要做高亮而非全文。
5. Atlas Vector Search 向量检索
Atlas Vector Search 让 MongoDB 直接支持 k-最近邻(kNN)向量检索,是 RAG(检索增强生成)与语义搜索的核心能力。它基于 HNSW(Hierarchical Navigable Small World)图索引,把高维向量组织为分层邻居图,实现近似最近邻(ANN)检索。
// 创建向量搜索索引(Atlas 控制台 / API)
{
"type": "vectorSearch",
"fields": [
{
"type": "vector",
"path": "embedding",
"numDimensions": 768,
"similarity": "cosine",
"quantization": "scalar"
}
]
}
相似度度量支持 cosine、euclidean 与 dotProduct。numDimensions 必须与实际向量维度一致(如 OpenAI text-embedding-3-small 为 1536 维,BGE-M3 为 1024 维)。
// 执行向量检索:查询向量 = "如何排查 MongoDB 慢查询"
db.articles.aggregate([
{
$vectorSearch: {
index: "vector_index",
path: "embedding",
queryVector: [0.01, -0.02, ... /* 768 维查询向量 */],
numCandidates: 100, // HNSW 图遍历候选数
limit: 10, // 返回 top-k
filter: { category: "database" } // 可选的预过滤
}
},
{ $project: { title: 1, score: { $meta: "vectorSearchScore" } } }
])
| 参数 | 作用 | 调优建议 |
|---|---|---|
| numCandidates | 图搜索候选池大小 | 越大召回越准,消耗越高;经验值 20~100 |
| limit | 返回条数 | 与业务分页配合 |
| filter | 预过滤条件 | 减小搜索空间,提升准确率 |
| similarity | 度量方式 | 语义检索常用 cosine |
向量检索与全文检索可以混合使用:先 $vectorSearch 做语义召回,再通过 $search 的 compound 叠加关键词约束,或用 $match/$facet 做后过滤。Atlas Vector Search 还可以配合 embedding 生成管道自动为新文档计算向量。
6. 分面与聚合(Facets & $searchMeta)
分面导航(facets)是电商、内容站的高频需求:在搜索结果旁展示"分类/价格区间/标签"的计数,供用户逐级筛选。Atlas Search 提供 $searchMeta 阶段,在不返回文档的前提下仅返回分面统计,成本远低于 $group。
db.products.aggregate([
{ $searchMeta: {
index: "default",
facet: {
operator: { text: { query: "laptop", path: "name" } },
facets: {
brandFacet: { type: "string", path: "brand" },
priceFacet: { type: "number", path: "price", boundaries: [0, 3000, 6000, 10000], default: "other" },
ratingFacet: { type: "number", path: "rating", boundaries: [0, 3, 4, 5] }
}
}
}}
])
返回结果大致为:facet.brandFacet.buckets: [{ _id: "Apple", count: 42 }, ...]。string 分面按词项计数,number 分面按 boundaries 划分区间,date 分面按时间范围计数。
$search 阶段本身也能与聚合管道其他阶段协同:先 $search 召回,再 $match 后过滤,然后 $group/$sort/$facet 做业务聚合,最后 $skip/$limit 分页。
// $search + 聚合:按品牌聚合销量 Top 商品
db.orders.aggregate([
{ $search: { index: "default", text: { query: "无线耳机", path: "productName" } } },
{ $group: { _id: "$brand", totalQty: { $sum: "$qty" } } },
{ $sort: { totalQty: -1 } },
{ $limit: 5 }
])
注意:
$search必须是聚合管道的第一阶段(或紧跟在少数允许的前置阶段之后)。需要先做权限/租户过滤时,可将过滤条件放进$search的compound.filter,而不是在$search之后用$match——后者无法利用 Lucene 索引裁剪。
7. 与 $text / 聚合管道的配合
$text 与 Atlas Search 可以共存:$text 适合简单的关键词过滤与存量代码兼容,Atlas Search 承担相关度检索。两者的结果在存储层共享同一份数据,但 $text 走普通文本索引,Atlas Search 走独立 Lucene 索引,因此两套查询的代价模型不同。
// 场景:先用 $text 快速过滤,再用 $search 排序
// (生产上更推荐:把过滤条件全部下沉到 $search.filter)
db.articles.aggregate([
{ $search: {
index: "default",
compound: {
filter: [
{ text: { query: "数据库", path: "category" } },
{ range: { path: "publishedAt", gte: new Date("2026-01-01") } }
],
should: [ { text: { query: "分片", path: "content", score: { boost: { value: 2 } } } } ]
}
}},
{ $facet: {
total: [{ $count: "count" }],
results: [{ $sort: { searchScore: -1 } }, { $skip: 0 }, { $limit: 10 }]
}}
])
结合 $facet 可以一次性返回总数与分页结果,减少往返。注意 searchScore 字段在 $project 中需要用 $meta 取回,不能在 $facet 内直接排序。
8. 生产调优
8.1 索引定义与资源
- 字段映射使用显式
fields而非dynamic: true,避免意外索引大文本字段 - 不需要搜索的字段(如内部代码)不要映射进索引
- 中文搜索优先 SmartCN;英文标题用 standard 即可
8.2 查询性能
$limit尽量前移,避免检索大量文档后丢弃filter子句(日期、状态、租户)务必下沉到$search,减少后续$match- 避免在
$search后对未索引字段排序;排序字段需在索引映射中声明为sortable: true
// 声明可排序字段
{
"mappings": {
"fields": {
"views": { "type": "number", "sortable": true },
"publishedAt": { "type": "date", "sortable": true }
}
}
}
8.3 容量与延迟
- 向量索引占用显著内存:768 维、1M 文档约需要数百 MB~GB 级内存,选择合适
quantization(scalar/int8 压缩)降低内存 - 高写入场景考虑在低峰期重建向量索引,避免 HNSW 图频繁更新
- 观察 Atlas 的 Search 指标面板:查询延迟、索引构建进度、被拒绝的查询数
| 指标 | 含义 | 参考阈值 |
|---|---|---|
| 查询延迟 p95 | $search 阶段耗时 | < 50ms |
| 索引延迟 | 文档到可检索的时延 | 秒级 |
| 内存使用率 | Lucene/向量索引驻留内存 | 不触发分页 |
| 高亮命中率 | 高亮片段产出质量 | 视业务 |
9. 典型场景总结
Atlas Search 的适用面可以概括为三条:需要全文相关度与拼写容错的站内搜索;需要分面导航的电商/内容平台;需要语义检索与向量召回的新型 AI 应用。它无法替代专用搜索引擎在极端规模下的深度定制,但对绝大多数 MongoDB 用户而言,是在不引入第二套存储的前提下获得企业级检索能力的最短路径。
延伸阅读
继续阅读
探索更多技术文章
浏览归档,发现更多关于系统设计、工具链和工程实践的内容。