Files
LocalAI/core/http/react-ui/e2e/router-template.spec.js
T
mudler-agentandEttore Di Giacinto 99043b442c feat(router): route with native decision models (#12449)
* fix(schema): preserve SystemOne image inputs

Assisted-by: OpenAI

* test(schema): follow Ginkgo conventions for decision inputs

Assisted-by: OpenAI

* feat(llama-cpp): dispatch native decisions through Score

Upgrade the stock dependency and reconcile Score/TTS patches. Reuse native decision parsing, tasks, formatting and response-reader cleanup; preserve ordinary scoring admission and guard older dependencies.

Assisted-by: OpenAI

* refactor(systemone): share request and model validation

Assisted-by: OpenAI:gpt-5

* fix(systemone): preserve HTTP wire-byte validation limit

Keep structural validation separate from the serialized internal request bound so HTML escaping cannot reject valid HTTP payloads.

Assisted-by: OpenAI:gpt-5

* feat(systemone): bound images and account native decisions

Preserve public wire limits independently from router serialization. Reject unsupported NER images, map native request/capability errors, and stamp explicit usage once. Advertise decisions for stock llama-cpp.

Assisted-by: OpenAI

* fix(systemone): record usage on registered native route

Exercise real registration and billing with a mock native backend. Reject empty native responses, malformed image URLs, trailing JSON, and wire overflow including whitespace.

Assisted-by: OpenAI

* feat(router): add lazy native decision transport

Bind named models through internal ModelSystemOne calls with shared validation and bounded abandoned operations. Remove request and echoed-error contents from decision traces.

Assisted-by: OpenAI:gpt-5

* feat(router): classify overlapping policies with native decisions

Ask independent noul questions, validate probabilities and preserve first-superset routing. Wire the central factory with config-sensitive invalidation and cancellation-safe resolution. Document native framing and bounded operation limits.

Assisted-by: OpenAI:gpt-5

* feat(gallery): add pinned Julia-1 native decision model

Add a separate text-only llama-cpp Q8 entry with pinned Apache-2.0 source provenance and checksum. Installed using the gallery installer and exercised choice, score and noul on CPU.

Assisted-by: OpenAI

* test(router): verify native decisions through central factory

Add an opt-in real-model Ginkgo integration covering the native Go loader and C++ transport, token usage, independent overlapping labels, and candidate selection. Document owned-server execution and the intentionally non-quality threshold.

Assisted-by: Codex:gpt-5

* fix(llama-cpp): align upstream pin and preserve decision signatures

Advance to bed0a856 without losing the automated upstream bump. Detect full-request fill_task support at compile time and forward every question for Nimble framing while retaining the earlier native signature. Preserve reconciled SCORE/TTS patches; add standalone compatibility coverage.

Assisted-by: Codex:gpt-5

* feat(gallery): add native decision family defaults

Pin Laya, Kev-4B, lev, OpenJev and Nimble artifacts. Verify Laya/Kev/lev gallery installs and CPU contracts on both native pins; clearly mark OpenJev/Nimble runtime validation pending and their noncommercial licenses.

Assisted-by: OpenAI

* docs(decisions): clarify integrated Nimble prerequisite

Record the exact combined backend pin while retaining pending OpenJev and Nimble installation/runtime validation status.

Assisted-by: Codex:gpt-5

* fix(gallery): indent native decision model sequences

Match repository yamllint indentation for Laya, Kev, lev and OpenJev list fields. Parsed gallery data is unchanged; reproduce CI gallery lint failure before the whitespace-only fix and pass the same command afterward.

Assisted-by: Codex:gpt-5

* docs(decisions): record OpenJev and Nimble CPU validation

Record gallery installation, checksum/metadata verification and multiquestion native smoke results on bed0a856. Retain noncommercial and text-only limitations without accuracy or deterministic-output claims.

Assisted-by: OpenAI

* fix(ui): expose native Decisions router classifiers

Select classifier models using metadata-driven capability routing, retain tuned thresholds, and validate native decision selections before saving. Cover both native backends and create/save/reopen in the real React editor.

Assisted-by: Codex:gpt-5

* fix(router): exclude aliases from native decision discovery

Check the originally named config before advertising native Decisions eligibility. Retain target capability inheritance for ordinary generation aliases. Exercise the actual capabilities endpoint with native models on both backends, aliases, and disabled models.

Assisted-by: Codex:gpt-5

* feat(systemone): share bounded multimodal input validation

Preserve text wire limits while admitting bounded PNG/JPEG decision input. Share collection and header validation across internal and public callers and keep the native runner response budget independent.

Assisted-by: OpenAI:API-assistant

* fix(systemone): bound admission lifetimes and validate complete images

Retain shared admission leases through actual work completion, including abandoned internal operations. Decode bounded image pixels, cap public native responses before usage stamping, and preserve oversized malformed text status precedence.

Assisted-by: OpenAI:API-assistant

* fix(router): classify images before media fetching

Preserve ordered structured probes for native decisions. Defer OpenAI
media preparation until routing selects the served model, so rejected
decision URLs cannot trigger downloads before shared validation.

Guard direct image collection with context-aware shared admission.
Keep text classifiers and embedding caches from discarding image input.
Retain fail-closed classifier configuration and runtime fallback policy.

Add middleware, typed-content, admission, cancellation and cache tests.

Assisted-by: OpenAI:API-assistant

* fix(router): bound extraction before serialization

Check probe budgets before copying text or marshaling message state.
Count JSON escaping so oversized internal inputs fail before allocation.

Preserve typed Anthropic blocks through selected-model conversion and
fallback. Keep retry coverage in Ginkgo without global test registration.

Assisted-by: OpenAI

* fix(router): bound supported probe serialization

Arbitrary structs can bypass the probe budget through pointer marshalers,
string tags, and promoted fields. Accept concrete chat schema types and
plain JSON values instead of emulating arbitrary struct serialization.

Budget escaped direct prompts before marshaling so raw length cannot hide
serialized expansion. Preserve runtime fallback and reject oversized
input before invoking the decision runner.

Add Ginkgo allocation, boundary, and marshaler invocation regressions.
Six-package tests, three-package race tests, and full-T2 delta lint pass.

Assisted-by: OpenAI:GPT-5 golangci-lint

* feat(decisions): enable bounded OpenJev images

Validate native decision images before permissive media parsing and pixel
allocation. Require both decision image support and a vision projector;
missing or audio-only projectors cannot silently become text decisions.

Pin the OpenJev Q8 projector and document its license and disk footprint.
Add native safety tests, canonical limit parity, gallery and load-option
checks, and a reproducible CPU direct-RPC contrasting-image smoke.

Assisted-by: OpenAI:GPT-5

* fix(decisions): reject incomplete image streams

stb accepts corrupt PNG Adler checksums and truncated JPEG scans.
Use bounded zlib validation and strict libjpeg decoding before parsing.
Keep dimension and aggregate pixel checks ahead of decoder allocations.

Wire decoder dependencies into native builds and runtime packaging.
Add regressions for appended EOI and embedded marker bypasses.

Assisted-by: OpenAI:GPT-5

* fix(ci): gate native decision image validation

Run the decoder security tests outside the stdlib-only native suite.
Fetch vendor headers at the backend pin and provision decoder dependencies.
Gate Go limit parity and production CMake wiring without model downloads.

Assisted-by: OpenAI:GPT-5

* test(decisions): cover multimodal public API paths

Exercise shared image contracts through the registered HTTP routes and
external mock backend. Add opt-in cached gallery installation and real
OpenJev image decisions through SystemOne and both routing APIs.

Assisted-by: Codex:gpt-5

* test(decisions): assert isolation and cache bypass

Observe external RPC calls and compare complete classifier history.
Winner-only and cache-miss checks could hide dropped history or cache use.

Give real inference its own application and model directory so shared
backend mappings and loaded processes cannot affect mixed suite order.

Assisted-by: OpenAI:ChatGPT

* test(decisions): isolate fixture globals

Disable optional global services in the isolated HTTP fixture and register
cleanup before setup assertions. Verify meter provider identity survives
fixture creation and destruction.

Snapshot observed usage before assertions so failures cannot retain the
mutex. Require a successful usage stamp before checking error responses.

Assisted-by: Codex:gpt-5 golangci-lint

* fix(application): honor optional telemetry controls

Skip failover gauge registration when metrics are disabled. Register
against the application meter rather than looking up the global provider.

Allow embedders to retain the bounded routing log without billing stats.
Keep the existing default when stats are disabled. The isolated HTTP
fixture uses this option without losing its native router assertions.

Assisted-by: Codex:gpt-5 golangci-lint

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-10-04 09:34:21 +02:00

278 lines
15 KiB
JavaScript

import { test, expect } from './coverage-fixtures'
import YAML from 'yaml'
// Router template + structured editor regression tests.
//
// The historical regression was: the "Create routing model" button
// loaded the model editor with an array-shaped `router.candidates`
// value, which crashed when a code-editor field received it instead
// of a string ("(intermediate value).split is not a function").
//
// The current schema is also covered:
// - classifier choices preserve the existing router creation flow
// - router.policies surfaces in its own structured editor (label +
// description rows with duplicate detection)
// - router.candidates is the structured {model, labels[]} editor;
// labels are chips populated from router.policies via FormContext
// - router.embedding_cache.* surface as labelled fields with the
// correct components (model-select / slider)
// - router.activation_threshold and the two embedding_cache slider
// fields render with slider min/max/step from the registry
const ROUTER_METADATA = {
sections: [
{ id: 'general', label: 'General', icon: 'settings', order: 0 },
{ id: 'other', label: 'Other', icon: 'more-horizontal', order: 100 },
],
fields: [
{ path: 'name', yaml_key: 'name', go_type: 'string', ui_type: 'string',
section: 'general', label: 'Model Name', component: 'input', order: 0 },
{
path: 'router.classifier', yaml_key: 'classifier', go_type: 'string', ui_type: 'string',
section: 'other', label: 'Classifier', component: 'select',
options: [{ value: 'score', label: 'Score (Arch-Router-style)' }, { value: 'decisions', label: 'Decisions (native probabilities)' }, { value: 'colbert', label: 'Colbert (reranker)' }, { value: 'knn', label: 'KNN (labelled corpus)' }],
description: 'Picks a candidate by scoring every policy label against the prompt. Only "score" is shipped today.',
order: 230,
},
{
path: 'router.classifier_model', yaml_key: 'classifier_model', go_type: 'string', ui_type: 'string',
section: 'other', label: 'Classifier Model', component: 'model-select', autocomplete_provider: 'models:score',
autocomplete_by: { field: 'router.classifier', providers: { decisions: 'models:decisions', colbert: 'models:rerank', knn: '' } },
description: 'Loaded LocalAI model the score classifier asks to rank each policy label.',
order: 231,
},
{
path: 'router.fallback', yaml_key: 'fallback', go_type: 'string', ui_type: 'string',
section: 'other', label: 'Fallback Model', component: 'model-select', autocomplete_provider: 'models:chat',
description: 'Model used when no candidate covers the active label set.',
order: 232,
},
{
path: 'router.activation_threshold', yaml_key: 'activation_threshold', go_type: 'float64', ui_type: 'float',
section: 'other', label: 'Activation Threshold', component: 'slider',
min: 0, max: 1, step: 0.05,
description: 'Softmax-probability floor a policy must clear to join the active label set.',
order: 233,
},
{
path: 'router.policies', yaml_key: 'policies', go_type: '[]RouterPolicy', ui_type: 'object',
section: 'other', label: 'Policies', component: 'router-policies',
description: 'Label vocabulary the classifier scores over.',
order: 235,
},
{
path: 'router.candidates', yaml_key: 'candidates', go_type: '[]RouterCandidate', ui_type: 'object',
section: 'other', label: 'Candidates', component: 'router-candidates',
description: 'Routing table: each entry binds a downstream model to a set of policy labels.',
order: 236,
},
{
path: 'router.embedding_cache.embedding_model', yaml_key: 'embedding_model', go_type: 'string', ui_type: 'string',
section: 'other', label: 'L2 Cache: Embedding Model', component: 'model-select', autocomplete_provider: 'models',
description: 'Embedding model used by the L2 decision cache.',
order: 237,
},
{
path: 'router.embedding_cache.similarity_threshold', yaml_key: 'similarity_threshold', go_type: 'float64', ui_type: 'float',
section: 'other', label: 'L2 Cache: Similarity Threshold', component: 'slider',
min: 0, max: 1, step: 0.01,
description: 'Cosine-similarity floor a cache candidate must clear to count as a hit.',
order: 238,
},
],
}
const MIDDLEWARE_STATUS = {
pii: { enabled_globally: false, patterns: [], models: [], recent_event_count: 0 },
router: { configured: false, models: [], recent_decision_count: 0, available_classifiers: ['score'] },
mitm: { running: false, listen_addr: '', configured_addr: '', host_owners: {}, host_conflicts: {}, models: [], ca_available: false, ca_cert_url: '' },
}
test.describe('Router template — create flow', () => {
test.beforeEach(async ({ page }) => {
await page.route('**/api/auth/status', (route) =>
route.fulfill({
contentType: 'application/json',
body: JSON.stringify({ authEnabled: false, staticApiKeyRequired: false, providers: [] }),
})
)
await page.route('**/api/middleware/status', (route) =>
route.fulfill({ contentType: 'application/json', body: JSON.stringify(MIDDLEWARE_STATUS) })
)
await page.route('**/api/router/decisions?**', (route) =>
route.fulfill({ contentType: 'application/json', body: JSON.stringify({ decisions: [] }) })
)
await page.route('**/api/pii/events?**', (route) =>
route.fulfill({ contentType: 'application/json', body: JSON.stringify({ events: [] }) })
)
await page.route('**/api/models/config-metadata*', (route) =>
route.fulfill({ contentType: 'application/json', body: JSON.stringify(ROUTER_METADATA) })
)
await page.route('**/api/models/config-metadata/autocomplete/**', (route) =>
route.fulfill({ contentType: 'application/json', body: JSON.stringify({ values: [] }) })
)
// Surface any uncaught render-time error so the assertion fails
// with a useful message rather than the test silently passing.
page.on('pageerror', (err) => {
throw new Error(`uncaught page error: ${err.message}`)
})
})
test('Routing tab links to the model editor with the router template loaded', async ({ page }) => {
await page.goto('/app/middleware')
await page.getByRole('button', { name: /Routing/i }).click()
// Empty-state button is the primary CTA.
await page.getByRole('button', { name: /Create routing model/i }).click()
// Editor loads on a /app/model-editor URL with template=router.
await expect(page).toHaveURL(/\/app\/model-editor.*template=router/)
})
test('Router template renders without crashing on structured candidates/policies', async ({ page }) => {
// Navigate straight to the create-with-template URL. This was the
// regression that crashed with "(intermediate value).split is not
// a function" when the template's array-shaped router.candidates
// fell into a code-editor wrapper.
await page.goto('/app/model-editor?template=router')
// The react-router error overlay must not appear.
await expect(page.getByText(/Unexpected Application Error/i)).toHaveCount(0)
// Editor surface visible. Template URL is "create mode", so the
// heading reads "Add Model" rather than "Model Editor".
await expect(page.locator('h1.page-title')).toBeVisible({ timeout: 10_000 })
// Top-level field labels seeded by the template are visible.
// embedding_cache.* fields are surfaced via "Add Field" search
// rather than active by default — separate spec covers them.
await expect(page.getByText('Classifier').first()).toBeVisible()
await expect(page.getByText('Policies').first()).toBeVisible()
await expect(page.getByText('Candidates').first()).toBeVisible()
await expect(page.getByText('Activation Threshold').first()).toBeVisible()
})
test('Classifier defaults to score', async ({ page }) => {
await page.goto('/app/model-editor?template=router')
// SearchableSelect renders the current option's *label* inside the
// trigger button. The initial option is
// "Score (Arch-Router-style)", pre-selected by the template.
await expect(page.getByText('Score (Arch-Router-style)').first()).toBeVisible({ timeout: 10_000 })
})
test('Policies editor renders structured rows with label + description fields', async ({ page }) => {
await page.goto('/app/model-editor?template=router')
// The template seeds three example policies. Their labels are
// pre-populated in input fields with monospace styling — the
// editor signature is "Add policy" button + label/description
// input pairs.
await expect(page.getByRole('button', { name: /Add policy/i }).first()).toBeVisible()
// Pre-seeded labels visible as input values. RouterPoliciesEditor
// renders each label in an input with a recognisable placeholder;
// assert on their values by position.
const labelInputs = page.locator('input[placeholder^="label ("]')
await expect(labelInputs.nth(0)).toHaveValue('code-generation')
await expect(labelInputs.nth(1)).toHaveValue('casual-chat')
await expect(labelInputs.nth(2)).toHaveValue('math-reasoning')
})
test('Candidates editor renders {model, labels} rows with policy-aware label chips', async ({ page }) => {
await page.goto('/app/model-editor?template=router')
// "Add candidate" is the signature of the new RouterCandidatesEditor.
await expect(page.getByRole('button', { name: /Add candidate/i }).first()).toBeVisible()
// Each candidate row should expose move-up/move-down controls,
// a model picker, and label chips. The chip for a known policy
// label appears as a button with the policy's label text.
// Pre-seeded template: candidate[0] has labels=['casual-chat'];
// candidate[1] has labels=['code-generation', 'casual-chat', 'math-reasoning'].
//
// The chips appear inside a flex row of buttons. Using getByRole
// with the exact name catches typos/regressions cleanly.
await expect(page.getByRole('button', { name: 'casual-chat' }).first()).toBeVisible()
await expect(page.getByRole('button', { name: 'code-generation' }).first()).toBeVisible()
await expect(page.getByRole('button', { name: 'math-reasoning' }).first()).toBeVisible()
})
test('Adding a duplicate policy label flags the duplicate row', async ({ page }) => {
await page.goto('/app/model-editor?template=router')
// Add a new empty policy row, then type a duplicate of the
// existing 'casual-chat'. The duplicate detection in
// RouterPoliciesEditor sets a warning border via inline style.
await page.getByRole('button', { name: /Add policy/i }).first().click()
// Find the newly-added empty label input (placeholder catches it).
const newLabel = page.locator('input[placeholder*="label (e.g. code-generation)"]').last()
await newLabel.fill('casual-chat')
// Both rows now hold the same label. The duplicate-detection
// logic flags the row visually; we assert on the title attribute
// RouterPoliciesEditor sets on the input when duplicate=true.
await expect(
page.locator('input[title="Duplicate label — candidates won\'t be able to distinguish them"]').first()
).toBeVisible()
})
test('Decisions picker creates, saves and reopens exact model and tuned threshold', async ({ page }) => {
const models = [
{ id: 'openjev-llama', capabilities: ['decisions'] },
{ id: 'LayaGLiNERDecide-vllm', capabilities: ['decisions'] },
{ id: 'TevKev-vllm', capabilities: ['decisions'] },
{ id: 'NimbleCLM-vllm', capabilities: ['decisions'] },
{ id: 'generic-ner', capabilities: ['token_classify'] },
{ id: 'chat-target', capabilities: ['chat'] },
]
await page.route('**/v1/models/capabilities', route => route.fulfill({ json: { data: models } }))
await page.route('**/api/models/capabilities', route => route.fulfill({ json: { data: [{ id: 'arch-score', capabilities: ['FLAG_SCORE'] }] } }))
let saved
await page.route('**/models/import', async route => {
saved = route.request().postDataJSON()
await route.fulfill({ json: { success: true } })
})
await page.route('**/api/models/edit/smart-router', route => route.fulfill({ json: { config: YAML.stringify(saved) } }))
await page.goto('/app/model-editor?template=router')
const modelRow = page.locator('.form-row').filter({ has: page.locator('.form-row__label-text', { hasText: /^Classifier Model$/ }) })
const threshold = page.locator('.form-row').filter({ hasText: 'Activation Threshold' }).locator('input[type="range"]')
await modelRow.locator('input').fill('arch-score')
await threshold.focus()
for (let i = 0; i < 5; i++) await threshold.press('ArrowRight')
await page.getByText('Score (Arch-Router-style)', { exact: true }).click()
await page.getByRole('option', { name: 'Decisions (native probabilities)' }).click()
await expect(modelRow.locator('input')).toHaveValue('')
await expect(threshold).toHaveValue('0.65')
await modelRow.locator('input').click()
for (const id of ['openjev-llama', 'LayaGLiNERDecide-vllm', 'TevKev-vllm', 'NimbleCLM-vllm']) {
await expect(page.getByRole('option', { name: id })).toBeVisible()
}
await expect(page.getByRole('option', { name: 'generic-ner' })).toHaveCount(0)
await expect(page.getByRole('option', { name: 'chat-target' })).toHaveCount(0)
await page.getByRole('option', { name: 'NimbleCLM-vllm' }).click()
await page.getByRole('button', { name: /Create Model$/ }).click()
await expect(page).toHaveURL(/model-editor\/smart-router/)
expect(saved.router).toMatchObject({ classifier: 'decisions', classifier_model: 'NimbleCLM-vllm', activation_threshold: 0.65 })
await page.reload()
await expect(page.getByText('Decisions (native probabilities)', { exact: true })).toBeVisible()
await expect(modelRow.locator('input')).toHaveValue('NimbleCLM-vllm')
await expect(threshold).toHaveValue('0.65')
await page.getByText('Decisions (native probabilities)', { exact: true }).click()
await page.getByRole('option', { name: 'Colbert (reranker)' }).click()
await expect(modelRow.locator('input')).toHaveValue('')
await expect(threshold).toHaveValue('0.65')
await page.getByText('Colbert (reranker)', { exact: true }).click()
await page.getByRole('option', { name: 'KNN (labelled corpus)' }).click()
await expect(modelRow).toHaveCount(0)
await expect(threshold).toHaveValue('0.65')
await page.getByText('KNN (labelled corpus)', { exact: true }).click()
await page.getByRole('option', { name: 'Decisions (native probabilities)' }).click()
await modelRow.locator('input').fill('generic-ner')
await page.getByRole('button', { name: /Save Changes$/ }).click()
await expect(page.getByText('Save failed: Select an eligible native Decisions classifier model')).toBeVisible()
})
})