sum/min/max/avg over a numeric field: bleve folds the accumulators from doc values, OpenSearch reads them from a stats aggregation, the service layer reduces them to the value of the kind after the cross-space merge. The proto gains MetricDefinition, MetricKind and Metric, which travels as accumulators (sum, count, min, max) so the merge works whatever the kind; value is set after the merge and absent for a metric without a single value. A metric is only valid on a numeric field, and only one of bucketDefinition and metricDefinition at a time. AGG-06, 18 and 21 in the parity suite.
Numeric and date ranges on both engines: bleve folds them from doc values, OpenSearch uses range and date_range aggregations. The proto gains BucketRange and bucketDefinition.ranges, the aggregation package the range parser (bounds are numbers or RFC3339 dates, one kind per aggregation, anything else is rejected), the kind of an option and the validation of ranges against the field type. Every requested range is answered in request order, an empty one with a count of zero, with no space answering too; size does not cut ranges, keyAsNumber sorts them by their lower bound. Range keys read from..to. AGG-04, 05 and 07 to 09 in the parity suite.
Term buckets on both engines: bleve folds them from doc values (its facets read numbers as prefix-coded terms), OpenSearch uses terms aggregations. A request gets every bucket, up to 65535 per space (the OpenSearch default for search.max_buckets, held on bleve as well); beyond that it is refused. The proto gains AggregationOption, BucketDefinition (sort, minimum count), AggregationResult and Bucket. The new aggregation package merges the per-space results by position, so several aggregations on one field stay apart, and shapes them per BucketDefinition (minimum count, sort, size). It also validates aggregations against the index mapping (a field the index knows and a client may name, of a type whose values are buckets; the mapping overrides mark the internal fields): the search service relies on it, the graph endpoint asks it first to save the round trip, next to its own checks against the spec, and with the field names of the request so an error reads like the request. A request names a field like the driveItem property, exactly so. A number is keyed as it is, a bool as true or false on both engines, an empty value is no bucket. Pinned in the parity suite as AGG-01 to 03, 17, 34, 37 and 39; the paging rows AGG-41 to 43 pin the first commit through the same matrix machinery, which also shortens long answer lists in the README.
MS Graph style search endpoint: searchRequest with KQL queryString, from/size pagination (values outside the spec's bounds are rejected), entityTypes validated to driveItem, hits as driveItems with webUrl, parentReference, remoteItem for shares, the facets and @libre.graph.permissions.actions.allowedValues, which the search service now projects onto the Entity from the same permission set as the WebDAV report. Properties the endpoint does not evaluate yet (aggregations, aggregationFilters, sortProperties) answer 501 instead of being ignored, an unknown $expand answers 400, and a failing search service keeps its status. Hits reuse what the drive item listing has, which moves a little: the thumbnails helper is named for the drive item it fills (setDriveItemThumbnailsByID, it served shares only), the web URL of an id gets a helper (webURLForID), and the facets come out of the proto through one generic mapping.FromProto instead of a copy per facet.
The search service takes a `from` offset next to the page size and applies it after the cross-space merge; every space answers the full prefix up to from+size, since an offset cannot be distributed. Pages are stable: both engines and the merge break score ties by id, and OpenSearch counts every match (track_total_hits). No engine pages beyond the first 10000 matches, the OpenSearch result window, so a from and size reaching beyond it are rejected on both engines. `page_size` becomes optional on the wire: absent is the default of 200, 0 asks for no matches, -1 for all; WebDAV and the tags listing set it accordingly, and the WebDAV report's `<oc:offset>`, parsed since the report exists and never read, now pages the merged list. The gRPC wrapper passes the request through and keys its cache by the whole request.
Spaces that are disabled while reindexing are skipped. They are now
remembered in a NATS KV bucket (including whether a forced rescan was
requested) and reindexed when the SpaceEnabled event comes in. The
entry is removed when the space gets deleted.
The command no longer times out after 10 minutes when the processing
hasn't been finished yet. Instead it shows progress and logs successful
and failed spaces. The process can also be aborted using ctrl+c.
Example output:
```
$ bin/opencloud search index --all-spaces --insecure --force-rescan
[1/602] indexed space a9033d65-6c13-4556-b923-f321ab33a9ca$eb6a8a0e-2e62-4b40-bce9-360063236676!eb6a8a0e-2e62-4b40-bce9-360063236676 in 5.501190908s
[2/602] indexed space a9033d65-6c13-4556-b923-f321ab33a9ca$34b7c565-da3b-4f53-8559-825e87db4011!34b7c565-da3b-4f53-8559-825e87db4011 in 5.935075697s
[3/602] indexed space a9033d65-6c13-4556-b923-f321ab33a9ca$e08c6026-24f6-4a9e-a805-ed335b34b8da!e08c6026-24f6-4a9e-a805-ed335b34b8da in 6.016350958s
^Caborted, indexing has been stopped
```
Fixes#2592
That can be helpful when the search service configuration has changed,
e.g. by enabling TIKA. Previously files that had already been indexed
were not indexed again and thus were no part of the fulltext index.
Fixes#2285Fixes#2578
Maintaining the positioning of the files from v2 to reduce cognitive
load.
Indentation of yaml files now matches `.editorconfig`.
All mock files regenerated.
Added empty `{}` following convention from `mockery init` etc.
Removed directory specification where it would already match.
This re-adds the check for go being installed before including the
bingo variables make file to avoid repeating errors about missing a
missing go binary when running 'make node-generate' in the ci (the node
container doesn't have go installed)
Adds the remote item id to search `REPORT` responses for shared resources and resources that are part of such. This id represents the id of the original resource that has been shared (= the remote item) and is needed for clients to correctly resolve their locations.