只有一个真正的方法来做到这一点。你必须索引你的数据关键字和搜索它与带状疱疹分析:
看到这个再现:
首先,我们将创建两个自定义分析:关键字和带状疱疹:
PUT test
{
"settings": {
"analysis": {
"analyzer": {
"my_analyzer_keyword": {
"type": "custom",
"tokenizer": "keyword",
"filter": [
"asciifolding",
"lowercase"
]
},
"my_analyzer_shingle": {
"type": "custom",
"tokenizer": "standard",
"filter": [
"asciifolding",
"lowercase",
"shingle"
]
}
}
}
},
"mappings": {
"your_type": {
"properties": {
"keyword": {
"type": "string",
"index_analyzer": "my_analyzer_keyword",
"search_analyzer": "my_analyzer_shingle"
}
}
}
}
}
现在,让我们创建一个使用你给我们一些样本数据:
POST /test/your_type/1
{
"id": 1,
"keyword": "thousand eyes"
}
POST /test/your_type/2
{
"id": 2,
"keyword": "facebook"
}
POST /test/your_type/3
{
"id": 3,
"keyword": "superdoc"
}
POST /test/your_type/4
{
"id": 4,
"keyword": "quora"
}
POST /test/your_type/5
{
"id": 5,
"keyword": "your story"
}
POST /test/your_type/6
{
"id": 6,
"keyword": "Surgery"
}
POST /test/your_type/7
{
"id": 7,
"keyword": "lending club"
}
POST /test/your_type/8
{
"id": 8,
"keyword": "ad roll"
}
POST /test/your_type/9
{
"id": 9,
"keyword": "the honest company"
}
POST /test/your_type/10
{
"id": 10,
"keyword": "Draft kings"
}
最后查询运行搜索:
POST /test/your_type/_search
{
"query": {
"match": {
"keyword": "I saw the news of lending club on facebook, your story and quora"
}
}
}
这是结果:
{
"took": 6,
"timed_out": false,
"_shards": {
"total": 5,
"successful": 5,
"failed": 0
},
"hits": {
"total": 4,
"max_score": 0.009332742,
"hits": [
{
"_index": "test",
"_type": "your_type",
"_id": "2",
"_score": 0.009332742,
"_source": {
"id": 2,
"keyword": "facebook"
}
},
{
"_index": "test",
"_type": "your_type",
"_id": "7",
"_score": 0.009332742,
"_source": {
"id": 7,
"keyword": "lending club"
}
},
{
"_index": "test",
"_type": "your_type",
"_id": "4",
"_score": 0.009207102,
"_source": {
"id": 4,
"keyword": "quora"
}
},
{
"_index": "test",
"_type": "your_type",
"_id": "5",
"_score": 0.0014755741,
"_source": {
"id": 5,
"keyword": "your story"
}
}
]
}
}
那么它在幕后?
- 它将您的文档索引为整个关键字(它将整个字符串作为单个标记发出)。我还添加了asciifolding过滤器,因此它将字母标准化,即
é
变为e
)和小写字母过滤器(不区分大小写的搜索)。因此,例如Draft kings
被索引为draft kings
- 现在搜索分析器使用相同的逻辑,除了它的标记器正在发出单词标记并且在其上创建了带状疱疹(标记的组合),这将与您的关键字匹配步。
是任何人能够在ElasticSearch的5.x版本运行它,似乎映射类型应该从字符串改为文字,index_analyzer只是分析,但我试图执行一个搜索 – mac
@mac让当too_many_clauses错误我试图让你为你工作! –
@mac我能够运行查询,但他们没有带回任何数据。我已经在GitHub上记录了这个问题:https://github.com/elastic/elasticsearch/issues/26989 –