Compare commits

..
Author SHA1 Message Date
Dominik Schmidt 54b62f42e5 feat(search): hierarchy tokenizer for bleve path fields
Path was a keyword, so the descendant lookup behind delete/move/restore/purge, the scoped search and the KQL path predicate expanded into one term searcher per descendant and OOM-killed the server on large folders (#1269, #3469).

Path is now analyzed into its ancestor prefixes, like path_hierarchy in OpenSearch: ./a/b.txt becomes ., ./a, ./a/b.txt. A folder's descendants are every document carrying the folder's path as a term, so all three call sites are a single term query. Schema 4 -> 5, v4 never shipped.

The same tokenizer with tag_depth is registered as the geohash analyzer, so #3272 can add its geohash field without another schema change.
2026-09-10 16:45:59 +00:00
Viktor Scharf e503c2c5eb add insecure to search reindex (#3505) 2026-09-10 16:29:37 +02:00
17 changed files with 305 additions and 113 deletions

No files matched your search

-61
View File
@@ -1,66 +1,5 @@
# Changelog
## [8.0.0](https://github.com/opencloud-eu/opencloud/releases/tag/v8.0.0) - 2026-09-09
### ❤️ Thanks to all contributors! ❤️
@butonic, @dschmidt, @fredrikblau, @fschade, @maki5, @pbleser-oc, @v-scharf, @zerox80
### 💥 Breaking changes
- feat(search): check the index schema on startup and refuse breaking changes [[#3197](https://github.com/opencloud-eu/opencloud/pull/3197)]
- refactor: reflection-based search mapping + location geopoint [[#3345](https://github.com/opencloud-eu/opencloud/pull/3345)]
- fix(search): make openSearch and bleve behave the same [[#3408](https://github.com/opencloud-eu/opencloud/pull/3408)]
### ✅ Tests
- api-tests: add search retry to all search tests [[#3500](https://github.com/opencloud-eu/opencloud/pull/3500)]
- test(search): fail the parity suite when the committed matrix is stale [[#3423](https://github.com/opencloud-eu/opencloud/pull/3423)]
- ci: run search API and e2e suites against OpenSearch [[#3379](https://github.com/opencloud-eu/opencloud/pull/3379)]
- test(search): engine parity suite for bleve and opensearch [[#3418](https://github.com/opencloud-eu/opencloud/pull/3418)]
### 🐛 Bug Fixes
- fix(search): extract facets from the main tika document only [[#3484](https://github.com/opencloud-eu/opencloud/pull/3484)]
- fix(thumbnails): bound declared image dimensions before decoding [[#3457](https://github.com/opencloud-eu/opencloud/pull/3457)]
- test(search): re-search until the expected files are in the result [[#3488](https://github.com/opencloud-eu/opencloud/pull/3488)]
- fix(config): correct pending version annotations [[#3487](https://github.com/opencloud-eu/opencloud/pull/3487)]
- Activitylog event handler split [[#3241](https://github.com/opencloud-eu/opencloud/pull/3241)]
- test(search): wait for expected properties and documents [[#3486](https://github.com/opencloud-eu/opencloud/pull/3486)]
- fix: log jwt expired on debug level instead of error [[#3463](https://github.com/opencloud-eu/opencloud/pull/3463)]
- fix(proxy): restrict JWT signed urls to the allowed HTTP methods [[#3481](https://github.com/opencloud-eu/opencloud/pull/3481)]
- fix: notification handling for share removal and space membership expiry [[#3257](https://github.com/opencloud-eu/opencloud/pull/3257)]
- fix: posix cli commands [[#3348](https://github.com/opencloud-eu/opencloud/pull/3348)]
### 📈 Enhancement
- feat(graph): expand thumbnails on driveItems [[#3471](https://github.com/opencloud-eu/opencloud/pull/3471)]
- graph: expose lockInfo on driveItems [[#3444](https://github.com/opencloud-eu/opencloud/pull/3444)]
- graph: expose @libre.graph.shareTypes on driveItems [[#3438](https://github.com/opencloud-eu/opencloud/pull/3438)]
- feat(search): live photo facet [[#3202](https://github.com/opencloud-eu/opencloud/pull/3202)]
- feat: support $expand=children on the driveItem endpoint [[#3445](https://github.com/opencloud-eu/opencloud/pull/3445)]
- feat(search): motion photo facet [[#3200](https://github.com/opencloud-eu/opencloud/pull/3200)]
- feat(search): video facet [[#3201](https://github.com/opencloud-eu/opencloud/pull/3201)]
- graph: expose pendingOperations on driveItems [[#3437](https://github.com/opencloud-eu/opencloud/pull/3437)]
- feat(search): extract more data from tika 4 (if available) [[#3198](https://github.com/opencloud-eu/opencloud/pull/3198)]
- graph: expose following state, tags and allowed actions on driveItems [[#3113](https://github.com/opencloud-eu/opencloud/pull/3113)]
- feat(search): scope searches to a drive via the driveId field [[#3424](https://github.com/opencloud-eu/opencloud/pull/3424)]
- chore(policies): disable gRPC or event handlers by configuration + add metrics [[#3287](https://github.com/opencloud-eu/opencloud/pull/3287)]
- moved ShareCreated event consumer from frontend to shared service [[#3389](https://github.com/opencloud-eu/opencloud/pull/3389)]
### 📦️ Dependencies
- build(deps): bump go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp from 0.70.0 to 0.71.0 [[#3441](https://github.com/opencloud-eu/opencloud/pull/3441)]
- build(deps): bump go.opentelemetry.io/contrib/zpages from 0.70.0 to 0.71.0 [[#3442](https://github.com/opencloud-eu/opencloud/pull/3442)]
- build(deps): bump go.opentelemetry.io/otel from 1.45.0 to 1.46.0 [[#3426](https://github.com/opencloud-eu/opencloud/pull/3426)]
- build(deps): bump google.golang.org/grpc from 1.83.1 to 1.83.2 [[#3429](https://github.com/opencloud-eu/opencloud/pull/3429)]
- build(deps): bump github.com/blevesearch/bleve/v2 from 2.6.0 to 2.6.1 [[#3417](https://github.com/opencloud-eu/opencloud/pull/3417)]
- build(deps): bump github.com/sirupsen/logrus from 1.10.0 to 1.10.1 [[#3409](https://github.com/opencloud-eu/opencloud/pull/3409)]
- build(deps): bump github.com/opensearch-project/opensearch-go/v4 from 4.6.0 to 4.7.3 [[#3216](https://github.com/opencloud-eu/opencloud/pull/3216)]
- build(deps): bump github.com/KimMachineGun/automemlimit from 0.7.5 to 1.0.0 [[#3404](https://github.com/opencloud-eu/opencloud/pull/3404)]
- build(deps): bump github.com/nats-io/nats.go from 1.52.0 to 1.53.1 [[#3403](https://github.com/opencloud-eu/opencloud/pull/3403)]
- build(deps): bump github.com/open-policy-agent/opa from 1.19.0 to 1.19.1 [[#3402](https://github.com/opencloud-eu/opencloud/pull/3402)]
## [7.5.0](https://github.com/opencloud-eu/opencloud/releases/tag/v7.5.0) - 2026-08-25
### ❤️ Thanks to all contributors! ❤️
+2 -2
View File
@@ -13,7 +13,7 @@ Fill the new index by indexing all spaces again:
```shell
# the service keeps running while it happens
opencloud search index --all-spaces
opencloud search index --all-spaces --insecure
```
Once the new index is filled, every index but the one with the highest
@@ -31,7 +31,7 @@ The new index is a directory next to the old `bleve` one, both in
bleve index cannot be copied, index all spaces again:
```shell
opencloud search index --all-spaces
opencloud search index --all-spaces --insecure
```
Once the new index is filled, every directory but the one with the highest
+2 -2
View File
@@ -124,14 +124,14 @@ opencloud search index --space $SPACE_ID
It can also be used to re-index all spaces:
```shell
opencloud search index --all-spaces
opencloud search index --all-spaces --insecure
```
Please note that a reindex only picks up new or changed files. Files that have already been indexed are not scanned again, even if the configuration or the whole extractor has been changed. To force a full rescan (re-running the extractor on every file) you need to use the `force-rescan` flag:
```shell
opencloud search index --all-spaces --force-rescan
opencloud search index --all-spaces --force-rescan --insecure
```
## Metrics
+3 -7
View File
@@ -75,14 +75,10 @@ func (b *Backend) Search(_ context.Context, sir *searchService.SearchIndexReques
},
)
// Scope below the space root: restrict at query level so totals and
// paging respect the path too. Path is a case-preserving keyword
// (paths act as references, /Foo and /foo are distinct), so the exact
// folder or the folder prefix matches all of, and only, the scope.
// paging respect the path too. The folder term matches the folder and
// its descendants (see PathAnalyzer).
if requestedPath := utils.MakeRelativePath(sir.Ref.Path); requestedPath != "." {
q.Conjuncts = append(q.Conjuncts, query.NewDisjunctionQuery([]query.Query{
&query.TermQuery{FieldVal: "Path", Term: requestedPath},
&query.PrefixQuery{FieldVal: "Path", Prefix: requestedPath + "/"},
}))
q.Conjuncts = append(q.Conjuncts, &query.TermQuery{FieldVal: "Path", Term: requestedPath})
}
}
-8
View File
@@ -1,8 +1,6 @@
package bleve
import (
"regexp"
bleveSearch "github.com/blevesearch/bleve/v2/search"
storageProvider "github.com/cs3org/go-cs3apis/cs3/storage/provider/v1beta1"
@@ -11,8 +9,6 @@ import (
"github.com/opencloud-eu/opencloud/services/search/pkg/search"
)
var queryEscape = regexp.MustCompile(`([` + regexp.QuoteMeta(`+=&|><!(){}[]^\"~*?:\/`) + `\-\s])`)
func getFieldValue[T any](m map[string]any, key string) (out T) {
val, ok := m[key]
if !ok {
@@ -84,7 +80,3 @@ func hitToFacet[T any](fields map[string]any, prefix string) *T {
func matchToResource(match *bleveSearch.DocumentMatch) *search.Resource {
return mapping.Deserialize[search.Resource](match.Fields)
}
func escapeQuery(s string) string {
return queryEscape.ReplaceAllString(s, "\\$1")
}
@@ -0,0 +1,51 @@
package bleve
import (
"github.com/blevesearch/bleve/v2"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"github.com/opencloud-eu/opencloud/pkg/log"
"github.com/opencloud-eu/opencloud/services/search/pkg/search"
)
var _ = Describe("searchResourcesByPath", func() {
var idx bleve.Index
BeforeEach(func() {
var err error
idx, _, err = NewIndex(GinkgoT().TempDir(), log.NopLogger())
Expect(err).ToNot(HaveOccurred())
DeferCleanup(func() { Expect(idx.Close()).To(Succeed()) })
})
ids := func(rootID, lookupPath string) []string {
res, err := searchResourcesByPath(rootID, lookupPath, idx)
Expect(err).ToNot(HaveOccurred())
out := make([]string, 0, len(res))
for _, r := range res {
out = append(out, r.ID)
}
return out
}
It("returns all of, and only, the descendants of the folder in its space", func() {
const rootA, rootB = "s$a!root", "s$b!root"
batch := idx.NewBatch()
for _, r := range []search.Resource{
{ID: "s$a!big", RootID: rootA, Path: "./big", Type: 2},
{ID: "s$a!f1", RootID: rootA, Path: "./big/f1.txt", Type: 1},
{ID: "s$a!f2", RootID: rootA, Path: "./big/sub/f2.txt", Type: 1},
{ID: "s$a!big2", RootID: rootA, Path: "./big2/x.txt", Type: 1},
{ID: "s$b!clone", RootID: rootB, Path: "./big/f1.txt", Type: 1},
{ID: "s$a!odd", RootID: rootA, Path: `./odd name*[1]/file:with spaces?.txt`, Type: 1},
} {
Expect(batch.Index(r.ID, r)).To(Succeed())
}
Expect(idx.Batch(batch)).To(Succeed())
Expect(ids(rootA, "./big")).To(ConsistOf("s$a!f1", "s$a!f2"))
Expect(ids(rootA, "./odd name*[1]")).To(ConsistOf("s$a!odd"))
Expect(ids(rootA, ".")).To(ConsistOf("s$a!big", "s$a!f1", "s$a!f2", "s$a!big2", "s$a!odd"))
})
})
@@ -0,0 +1,82 @@
package hierarchy
import (
"bytes"
"strconv"
"github.com/blevesearch/bleve/v2/analysis"
"github.com/blevesearch/bleve/v2/registry"
)
// emits every prefix up to a level: "./a/b" -> ".", "./a", "./a/b" with
// delimiter "/", one level per byte without. tag_depth prepends "<depth>/".
const Name = "hierarchy"
type Tokenizer struct {
delimiter []byte
tagDepth bool
}
func (t *Tokenizer) Tokenize(input []byte) analysis.TokenStream {
if len(input) == 0 {
return nil
}
var out analysis.TokenStream
emit := func(depth, end int) {
term := input[:end]
if t.tagDepth {
term = strconv.AppendInt(make([]byte, 0, end+4), int64(depth), 10)
term = append(term, '/')
term = append(term, input[:end]...)
}
out = append(out, &analysis.Token{
Term: term,
Position: depth,
Start: 0,
End: end,
Type: analysis.AlphaNumeric,
})
}
if len(t.delimiter) == 0 {
for i := range input {
emit(i+1, i+1)
}
return out
}
depth := 0
for start := 0; start <= len(input); {
i := bytes.Index(input[start:], t.delimiter)
if i < 0 {
if start < len(input) {
depth++
emit(depth, len(input))
}
break
}
if i > 0 {
depth++
emit(depth, start+i)
}
start += i + len(t.delimiter)
}
return out
}
func Constructor(config map[string]interface{}, _ *registry.Cache) (analysis.Tokenizer, error) {
t := &Tokenizer{}
if d, ok := config["delimiter"].(string); ok {
t.delimiter = []byte(d)
}
if v, ok := config["tag_depth"].(bool); ok {
t.tagDepth = v
}
return t, nil
}
func init() {
if err := registry.RegisterTokenizer(Name, Constructor); err != nil {
panic(err)
}
}
@@ -0,0 +1,13 @@
package hierarchy_test
import (
"testing"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
func TestHierarchy(t *testing.T) {
RegisterFailHandler(Fail)
RunSpecs(t, "hierarchy tokenizer")
}
@@ -0,0 +1,56 @@
package hierarchy_test
import (
"github.com/blevesearch/bleve/v2/analysis"
"github.com/blevesearch/bleve/v2/registry"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"github.com/opencloud-eu/opencloud/services/search/pkg/bleve/hierarchy"
)
func terms(ts analysis.TokenStream) []string {
out := make([]string, 0, len(ts))
for _, t := range ts {
out = append(out, string(t.Term))
}
return out
}
func tokenize(config map[string]any, input string) []string {
tok, err := hierarchy.Constructor(config, registry.NewCache())
Expect(err).ToNot(HaveOccurred())
return terms(tok.Tokenize([]byte(input)))
}
var _ = Describe("hierarchy tokenizer", func() {
path := map[string]any{"delimiter": "/"}
geohash := map[string]any{"tag_depth": true}
DescribeTable("emits every prefix up to a level boundary",
func(config map[string]any, input string, want []string) {
Expect(tokenize(config, input)).To(Equal(want))
},
Entry("relative path", path, "./a/b.txt", []string{".", "./a", "./a/b.txt"}),
Entry("space root", path, ".", []string{"."}),
Entry("trailing delimiter is not a level", path, "./a/", []string{".", "./a"}),
Entry("delimiter only", path, "/", []string{}),
Entry("leading delimiter", path, "/abs/x", []string{"/abs", "/abs/x"}),
Entry("double delimiter", path, "./a//b", []string{".", "./a", "./a//b"}),
Entry("spaces and special characters stay literal", path, "./odd name*[1]/f:x?.txt",
[]string{".", "./odd name*[1]", "./odd name*[1]/f:x?.txt"}),
Entry("empty input", path, "", []string{}),
Entry("geohash, one level per byte, depth tagged", geohash, "u4pru",
[]string{"1/u", "2/u4", "3/u4p", "4/u4pr", "5/u4pru"}),
)
It("keeps byte offsets on the source value", func() {
tok, err := hierarchy.Constructor(path, registry.NewCache())
Expect(err).ToNot(HaveOccurred())
ts := tok.Tokenize([]byte("./a/b"))
Expect(ts).To(HaveLen(3))
Expect(ts[2].Start).To(Equal(0))
Expect(ts[2].End).To(Equal(5))
Expect(ts[2].Position).To(Equal(3))
})
})
+48 -5
View File
@@ -19,6 +19,7 @@ import (
storageProvider "github.com/cs3org/go-cs3apis/cs3/storage/provider/v1beta1"
"github.com/opencloud-eu/opencloud/pkg/log"
"github.com/opencloud-eu/opencloud/services/search/pkg/bleve/hierarchy"
searchmapping "github.com/opencloud-eu/opencloud/services/search/pkg/mapping"
"github.com/opencloud-eu/opencloud/services/search/pkg/search"
)
@@ -208,6 +209,40 @@ func NewMapping() (mapping.IndexMapping, error) {
if err != nil {
return nil, err
}
// path: every ancestor prefix is a term, so one term query matches a folder
// and all of its descendants
err = indexMapping.AddCustomTokenizer("path_hierarchy", map[string]any{
"type": hierarchy.Name,
"delimiter": "/",
})
if err != nil {
return nil, err
}
err = indexMapping.AddCustomAnalyzer(searchmapping.PathAnalyzer, map[string]any{
"type": custom.Name,
"tokenizer": "path_hierarchy",
})
if err != nil {
return nil, err
}
// geohash: no field uses it yet. It is part of the v5 schema so that #3272
// can add its geohash field additively: new fields reconcile at startup,
// a changed analysis block does not (classifyStoredMapping), so the names
// and the config below must not change.
err = indexMapping.AddCustomTokenizer("geohash_hierarchy", map[string]any{
"type": hierarchy.Name,
"tag_depth": true,
})
if err != nil {
return nil, err
}
err = indexMapping.AddCustomAnalyzer(searchmapping.GeohashAnalyzer, map[string]any{
"type": custom.Name,
"tokenizer": "geohash_hierarchy",
})
if err != nil {
return nil, err
}
return indexMapping, nil
}
@@ -226,11 +261,15 @@ func searchResourceByID(id string, index bleve.Index) (*search.Resource, error)
return matchToResource(res.Hits[0]), nil
}
// searchResourcesByPath returns the descendants of the folder at lookupPath.
// The folder term matches the folder and everything below it in one term
// query (see PathAnalyzer); the folder itself is dropped from the result.
func searchResourcesByPath(rootID string, lookupPath string, index bleve.Index) ([]*search.Resource, error) {
q := bleve.NewConjunctionQuery(
bleve.NewQueryStringQuery("RootID:"+rootID),
bleve.NewQueryStringQuery("Path:"+escapeQuery(lookupPath+"/*")),
)
rootQuery := bleve.NewTermQuery(rootID)
rootQuery.SetField("RootID")
pathQuery := bleve.NewTermQuery(lookupPath)
pathQuery.SetField("Path")
q := bleve.NewConjunctionQuery(rootQuery, pathQuery)
bleveReq := bleve.NewSearchRequest(q)
bleveReq.Size = math.MaxInt
bleveReq.Fields = []string{"*"}
@@ -241,7 +280,11 @@ func searchResourcesByPath(rootID string, lookupPath string, index bleve.Index)
resources := make([]*search.Resource, 0, res.Hits.Len())
for _, match := range res.Hits {
resources = append(resources, matchToResource(match))
resource := matchToResource(match)
if resource.Path == lookupPath {
continue
}
resources = append(resources, resource)
}
return resources, nil
+19 -1
View File
@@ -160,7 +160,7 @@
"fields": [
{
"type": "text",
"analyzer": "keyword",
"analyzer": "path",
"store": true,
"index": true,
"include_term_vectors": true,
@@ -1259,7 +1259,25 @@
"type": "regexp"
}
},
"tokenizers": {
"geohash_hierarchy": {
"tag_depth": true,
"type": "hierarchy"
},
"path_hierarchy": {
"delimiter": "/",
"type": "hierarchy"
}
},
"analyzers": {
"geohash": {
"tokenizer": "geohash_hierarchy",
"type": "custom"
},
"path": {
"tokenizer": "path_hierarchy",
"type": "custom"
},
"words": {
"char_filters": [
"dot_to_space"
+6 -3
View File
@@ -60,7 +60,6 @@ func buildBleveDocMapping(t reflect.Type, overrides map[string]FieldOpts, prefix
}
if fieldType == TypeKeyword || fieldType == TypePath {
// bleve has no path tokenizer, so a path is a plain keyword here.
base := bleveKeywordMapping(fieldType, opts)
doc.AddFieldMappingsAt(fi.Name, base)
if opts.caseInsensitive() {
@@ -84,8 +83,9 @@ func buildBleveDocMapping(t reflect.Type, overrides map[string]FieldOpts, prefix
return doc, err
}
// bleveKeywordMapping is a case-preserving keyword field; path fields stay out
// of _all by default.
// bleveKeywordMapping is a case-preserving keyword field; path fields are
// analyzed into their ancestor prefixes (see PathAnalyzer) and stay out of
// _all by default.
func bleveKeywordMapping(fieldType string, opts FieldOpts) *bleveMapping.FieldMapping {
fm := bleve.NewKeywordFieldMapping()
switch {
@@ -94,6 +94,9 @@ func bleveKeywordMapping(fieldType string, opts FieldOpts) *bleveMapping.FieldMa
case fieldType == TypePath:
fm.IncludeInAll = false
}
if fieldType == TypePath {
fm.Analyzer = PathAnalyzer
}
return fm
}
+9
View File
@@ -26,6 +26,15 @@ const WordsSuffix = "_words"
// WordsAnalyzer names the analyzer both engines register for the words sibling.
const WordsAnalyzer = "words"
// PathAnalyzer names the bleve analyzer for TypePath fields: every ancestor
// prefix is a term, like path_hierarchy in OpenSearch.
const PathAnalyzer = "path"
// GeohashAnalyzer names the bleve analyzer for a geohash: every prefix is a
// depth-tagged term (1/u, 2/u4, ...), so a terms facet with TermPrefix
// "<precision>/" is a geohash grid at that precision.
const GeohashAnalyzer = "geohash"
// FieldOpts overrides the default type inference for a struct field. Keys in
// the override map are json-tag names (e.g. "Name", "location", "audio.artist"),
// not Go field names.
+2 -2
View File
@@ -27,7 +27,7 @@ func Reconcile(index string, r SchemaReconciler, logger log.Logger) (Classificat
case VerdictAdditive:
persisted, err := r.ApplyAdditive()
if persisted {
logger.Warn().Strs("fields", classification.NewFields).Str("index", index).Msg("extended the search index mapping with new fields; documents indexed before the upgrade do not contain them and queries on these fields will miss those documents until they are re-indexed; to re-index everything run: opencloud search index --all-spaces --force-rescan")
logger.Warn().Strs("fields", classification.NewFields).Str("index", index).Msg("extended the search index mapping with new fields; documents indexed before the upgrade do not contain them and queries on these fields will miss those documents until they are re-indexed; to re-index everything run: opencloud search index --all-spaces --force-rescan --insecure")
}
if err != nil {
return classification, err
@@ -40,5 +40,5 @@ func Reconcile(index string, r SchemaReconciler, logger log.Logger) (Classificat
// LogNewIndexCreated logs that a fresh, empty index was created and how to
// backfill it. The create path does not run through Reconcile.
func LogNewIndexCreated(logger log.Logger, index string) {
logger.Info().Str("index", index).Msg("created a new empty search index; if this OpenCloud instance already held files, they are not in it yet, index them by running: opencloud search index --all-spaces --force-rescan")
logger.Info().Str("index", index).Msg("created a new empty search index; if this OpenCloud instance already held files, they are not in it yet, index them by running: opencloud search index --all-spaces --force-rescan --insecure")
}
+7 -11
View File
@@ -123,6 +123,13 @@ func walk(offset int, nodes []ast.Node) (bleveQuery.Query, int, error) {
var q bleveQuery.Query = bleveQuery.NewQueryStringQuery(k + ":" + v)
switch {
case searchQuery.FieldIsPath(n.Key) && !isWildcard:
// the folder term matches the folder itself and its descendants
// (see PathAnalyzer); a query string would analyze the value into
// its prefixes and match everything under the root
tq := bleveQuery.NewTermQuery(val)
tq.SetField(k)
q = tq
case n.Exact && !isWildcard:
// = matches the whole value, on the lowercased sibling for
// case-insensitive fields
@@ -140,17 +147,6 @@ func walk(offset int, nodes []ast.Node) (bleveQuery.Query, int, error) {
bq.SetMinShould(1)
q = bq
}
if searchQuery.FieldIsPath(n.Key) {
// bleve has no path hierarchy analyzer, unlike OpenSearch: match the
// folder itself and its descendants (`\/*`). A BooleanQuery keeps
// this atomic; a DisjunctionQuery would be redistributed by an
// enclosing AND (mapBinary treats a left disjunction as an OR-chain).
bq := bleve.NewBooleanQuery()
bq.AddShould(q, bleveQuery.NewQueryStringQuery(k+":"+v+`\/*`))
bq.SetMinShould(1)
q = bq
}
if prev == nil {
prev = q
} else {
@@ -51,23 +51,17 @@ func Test_compile(t *testing.T) {
wantErr: false,
},
{
// path fields expand to match the folder itself and its descendants,
// since bleve has no path hierarchy analyzer.
// one term matches the folder itself and its descendants
name: `path:/Foo`,
args: &ast.Ast{
Nodes: []ast.Node{
&ast.StringNode{Key: "path", Value: "/Foo"},
},
},
// a BooleanQuery (should: exact OR descendants), not a DisjunctionQuery,
// so an enclosing AND does not redistribute the folder-itself clause.
want: func() query.Query {
bq := query.NewBooleanQuery(nil, []query.Query{
query.NewQueryStringQuery(`Path:\/Foo`),
query.NewQueryStringQuery(`Path:\/Foo\/*`),
}, nil)
bq.SetMinShould(1)
return query.NewConjunctionQuery([]query.Query{bq})
tq := query.NewTermQuery("/Foo")
tq.SetField("Path")
return query.NewConjunctionQuery([]query.Query{tq})
}(),
wantErr: false,
},
+1 -1
View File
@@ -29,7 +29,7 @@ import (
// on a breaking mapping change: each version gets its own index (OpenSearch name
// suffix, bleve path suffix), so the service builds a fresh index instead of
// colliding with the old one. No migration; reindex to populate.
const SchemaVersion = 4
const SchemaVersion = 5
var scopeRegex = regexp.MustCompile(`scope:\s*([^" "\n\r]*)`)