feat(router): route with native decision models (#12449)

* fix(schema): preserve SystemOne image inputs

Assisted-by: OpenAI

* test(schema): follow Ginkgo conventions for decision inputs

Assisted-by: OpenAI

* feat(llama-cpp): dispatch native decisions through Score

Upgrade the stock dependency and reconcile Score/TTS patches. Reuse native decision parsing, tasks, formatting and response-reader cleanup; preserve ordinary scoring admission and guard older dependencies.

Assisted-by: OpenAI

* refactor(systemone): share request and model validation

Assisted-by: OpenAI:gpt-5

* fix(systemone): preserve HTTP wire-byte validation limit

Keep structural validation separate from the serialized internal request bound so HTML escaping cannot reject valid HTTP payloads.

Assisted-by: OpenAI:gpt-5

* feat(systemone): bound images and account native decisions

Preserve public wire limits independently from router serialization. Reject unsupported NER images, map native request/capability errors, and stamp explicit usage once. Advertise decisions for stock llama-cpp.

Assisted-by: OpenAI

* fix(systemone): record usage on registered native route

Exercise real registration and billing with a mock native backend. Reject empty native responses, malformed image URLs, trailing JSON, and wire overflow including whitespace.

Assisted-by: OpenAI

* feat(router): add lazy native decision transport

Bind named models through internal ModelSystemOne calls with shared validation and bounded abandoned operations. Remove request and echoed-error contents from decision traces.

Assisted-by: OpenAI:gpt-5

* feat(router): classify overlapping policies with native decisions

Ask independent noul questions, validate probabilities and preserve first-superset routing. Wire the central factory with config-sensitive invalidation and cancellation-safe resolution. Document native framing and bounded operation limits.

Assisted-by: OpenAI:gpt-5

* feat(gallery): add pinned Julia-1 native decision model

Add a separate text-only llama-cpp Q8 entry with pinned Apache-2.0 source provenance and checksum. Installed using the gallery installer and exercised choice, score and noul on CPU.

Assisted-by: OpenAI

* test(router): verify native decisions through central factory

Add an opt-in real-model Ginkgo integration covering the native Go loader and C++ transport, token usage, independent overlapping labels, and candidate selection. Document owned-server execution and the intentionally non-quality threshold.

Assisted-by: Codex:gpt-5

* fix(llama-cpp): align upstream pin and preserve decision signatures

Advance to bed0a856 without losing the automated upstream bump. Detect full-request fill_task support at compile time and forward every question for Nimble framing while retaining the earlier native signature. Preserve reconciled SCORE/TTS patches; add standalone compatibility coverage.

Assisted-by: Codex:gpt-5

* feat(gallery): add native decision family defaults

Pin Laya, Kev-4B, lev, OpenJev and Nimble artifacts. Verify Laya/Kev/lev gallery installs and CPU contracts on both native pins; clearly mark OpenJev/Nimble runtime validation pending and their noncommercial licenses.

Assisted-by: OpenAI

* docs(decisions): clarify integrated Nimble prerequisite

Record the exact combined backend pin while retaining pending OpenJev and Nimble installation/runtime validation status.

Assisted-by: Codex:gpt-5

* fix(gallery): indent native decision model sequences

Match repository yamllint indentation for Laya, Kev, lev and OpenJev list fields. Parsed gallery data is unchanged; reproduce CI gallery lint failure before the whitespace-only fix and pass the same command afterward.

Assisted-by: Codex:gpt-5

* docs(decisions): record OpenJev and Nimble CPU validation

Record gallery installation, checksum/metadata verification and multiquestion native smoke results on bed0a856. Retain noncommercial and text-only limitations without accuracy or deterministic-output claims.

Assisted-by: OpenAI

* fix(ui): expose native Decisions router classifiers

Select classifier models using metadata-driven capability routing, retain tuned thresholds, and validate native decision selections before saving. Cover both native backends and create/save/reopen in the real React editor.

Assisted-by: Codex:gpt-5

* fix(router): exclude aliases from native decision discovery

Check the originally named config before advertising native Decisions eligibility. Retain target capability inheritance for ordinary generation aliases. Exercise the actual capabilities endpoint with native models on both backends, aliases, and disabled models.

Assisted-by: Codex:gpt-5

* feat(systemone): share bounded multimodal input validation

Preserve text wire limits while admitting bounded PNG/JPEG decision input. Share collection and header validation across internal and public callers and keep the native runner response budget independent.

Assisted-by: OpenAI:API-assistant

* fix(systemone): bound admission lifetimes and validate complete images

Retain shared admission leases through actual work completion, including abandoned internal operations. Decode bounded image pixels, cap public native responses before usage stamping, and preserve oversized malformed text status precedence.

Assisted-by: OpenAI:API-assistant

* fix(router): classify images before media fetching

Preserve ordered structured probes for native decisions. Defer OpenAI
media preparation until routing selects the served model, so rejected
decision URLs cannot trigger downloads before shared validation.

Guard direct image collection with context-aware shared admission.
Keep text classifiers and embedding caches from discarding image input.
Retain fail-closed classifier configuration and runtime fallback policy.

Add middleware, typed-content, admission, cancellation and cache tests.

Assisted-by: OpenAI:API-assistant

* fix(router): bound extraction before serialization

Check probe budgets before copying text or marshaling message state.
Count JSON escaping so oversized internal inputs fail before allocation.

Preserve typed Anthropic blocks through selected-model conversion and
fallback. Keep retry coverage in Ginkgo without global test registration.

Assisted-by: OpenAI

* fix(router): bound supported probe serialization

Arbitrary structs can bypass the probe budget through pointer marshalers,
string tags, and promoted fields. Accept concrete chat schema types and
plain JSON values instead of emulating arbitrary struct serialization.

Budget escaped direct prompts before marshaling so raw length cannot hide
serialized expansion. Preserve runtime fallback and reject oversized
input before invoking the decision runner.

Add Ginkgo allocation, boundary, and marshaler invocation regressions.
Six-package tests, three-package race tests, and full-T2 delta lint pass.

Assisted-by: OpenAI:GPT-5 golangci-lint

* feat(decisions): enable bounded OpenJev images

Validate native decision images before permissive media parsing and pixel
allocation. Require both decision image support and a vision projector;
missing or audio-only projectors cannot silently become text decisions.

Pin the OpenJev Q8 projector and document its license and disk footprint.
Add native safety tests, canonical limit parity, gallery and load-option
checks, and a reproducible CPU direct-RPC contrasting-image smoke.

Assisted-by: OpenAI:GPT-5

* fix(decisions): reject incomplete image streams

stb accepts corrupt PNG Adler checksums and truncated JPEG scans.
Use bounded zlib validation and strict libjpeg decoding before parsing.
Keep dimension and aggregate pixel checks ahead of decoder allocations.

Wire decoder dependencies into native builds and runtime packaging.
Add regressions for appended EOI and embedded marker bypasses.

Assisted-by: OpenAI:GPT-5

* fix(ci): gate native decision image validation

Run the decoder security tests outside the stdlib-only native suite.
Fetch vendor headers at the backend pin and provision decoder dependencies.
Gate Go limit parity and production CMake wiring without model downloads.

Assisted-by: OpenAI:GPT-5

* test(decisions): cover multimodal public API paths

Exercise shared image contracts through the registered HTTP routes and
external mock backend. Add opt-in cached gallery installation and real
OpenJev image decisions through SystemOne and both routing APIs.

Assisted-by: Codex:gpt-5

* test(decisions): assert isolation and cache bypass

Observe external RPC calls and compare complete classifier history.
Winner-only and cache-miss checks could hide dropped history or cache use.

Give real inference its own application and model directory so shared
backend mappings and loaded processes cannot affect mixed suite order.

Assisted-by: OpenAI:ChatGPT

* test(decisions): isolate fixture globals

Disable optional global services in the isolated HTTP fixture and register
cleanup before setup assertions. Verify meter provider identity survives
fixture creation and destruction.

Snapshot observed usage before assertions so failures cannot retain the
mutex. Require a successful usage stamp before checking error responses.

Assisted-by: Codex:gpt-5 golangci-lint

* fix(application): honor optional telemetry controls

Skip failover gauge registration when metrics are disabled. Register
against the application meter rather than looking up the global provider.

Allow embedders to retain the bounded routing log without billing stats.
Keep the existing default when stats are disabled. The isolated HTTP
fixture uses this option without losing its native router assertions.

Assisted-by: Codex:gpt-5 golangci-lint

---------

Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
This commit is contained in:
mudler-agentandEttore Di Giacinto authored and GitHub committed 2026-10-04 09:34:21 +02:00
1 parent f035746db9
commit 99043b442c
108 files changed
+6263 -274

No files matched your search

+7
View File
@@ -83,6 +83,13 @@ target_link_libraries(${TARGET} PRIVATE ${_LLAMA_COMMON_TARGET} llama mtmd ${CMA
gRPC::${_REFLECTION}
gRPC::${_GRPC_GRPCPP}
protobuf::${_PROTOBUF_LIBPROTOBUF})
# Match grpc-server.cpp's native-decision guard: older forks do not need
# strict decision image decoders and must not acquire new dependencies.
if(EXISTS "${CMAKE_CURRENT_SOURCE_DIR}/../server/server-decision.cpp")
find_package(ZLIB REQUIRED)
find_package(JPEG REQUIRED)
target_link_libraries(${TARGET} PRIVATE ZLIB::ZLIB JPEG::JPEG)
endif()
target_compile_features(${TARGET} PRIVATE cxx_std_11)
if(TARGET BUILD_INFO)
add_dependencies(${TARGET} BUILD_INFO)
+1 -1
View File
@@ -1,5 +1,5 @@
LLAMA_VERSION?=a868c3e3c56657f7e8a6231190dbbe90e7dd86c0
LLAMA_VERSION?=bed0a856606ee4a24a164066f73d2379447033f5
LLAMA_REPO?=https://github.com/ggerganov/llama.cpp
CMAKE_ARGS?=
+18
View File
@@ -0,0 +1,18 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <type_traits>
#include <utility>
// Nimble framing needs every question. Detect the callable signature rather
// than a revision number so older decision-capable forks keep working too.
template <typename Decision, typename State, typename Questions, typename... Args>
void localai_fill_decision_task(const Decision & decision, const State & state,
const Questions & questions, Args &&... args) {
if constexpr (std::is_invocable_v<decltype(&Decision::fill_task),
const Decision &, const State &, const Questions &, Args...>) {
decision.fill_task(state, questions, std::forward<Args>(args)...);
} else {
decision.fill_task(state, std::forward<Args>(args)...);
}
}
@@ -0,0 +1,27 @@
// SPDX-License-Identifier: MIT
#include "decision_compat.h"
#include <cassert>
#include <vector>
struct legacy_decision {
void fill_task(const int & state, int question, int & result) const {
result = state + question;
}
};
struct full_request_decision {
const std::vector<int> * expected;
void fill_task(const int & state, const std::vector<int> & questions,
int question, int & result) const {
assert(&questions == expected); // no copy or singleton substitution
assert(questions.size() == 2);
result = state + question + questions[1];
}
};
int main() {
const std::vector<int> questions{3, 7};
int result = 0;
localai_fill_decision_task(legacy_decision{}, 2, questions, 3, result);
assert(result == 5);
localai_fill_decision_task(full_request_decision{&questions}, 2, questions, 3, result);
assert(result == 12);
}
+210
View File
@@ -0,0 +1,210 @@
// SPDX-License-Identifier: MIT
#pragma once
#include <nlohmann/json.hpp>
#include <stdexcept>
#include <string>
#include <vector>
#include <cstdint>
#include <cstring>
#include <memory>
#include <csetjmp>
#include <cstdio>
#include <jpeglib.h>
#include <zlib.h>
#include "stb/stb_image.h"
// Decision-only limits, mirrored from core/systemone/images.go. The verification
// script checks parity; these must not change ordinary chat or fork backends.
namespace localai_decision {
using json = nlohmann::ordered_json;
constexpr size_t max_images = 8;
constexpr size_t decoded_bytes = 8 << 20;
constexpr size_t encoded_bytes = 12 << 20;
constexpr size_t body_bytes = 16 << 20;
constexpr size_t text_bytes = 64 << 10;
constexpr size_t max_dimension = 4096;
constexpr size_t max_pixels = 16000000;
struct image_error : std::invalid_argument {
bool too_large;
image_error(const char * message, bool large=false) : std::invalid_argument(message), too_large(large) {}
};
inline bool supports_images(bool decision, bool vision) { return decision && vision; }
inline void require(bool ok, const char * message, bool large=false) {
if (!ok) throw image_error(message, large);
}
// libjpeg normally repairs premature EOF and incomplete entropy scans. Treat
// warnings as failures as well as fatal errors; an appended EOI cannot hide a
// short scan. Keep all mutable decoder state on the heap across longjmp.
struct jpeg_validator {
jpeg_decompress_struct decoder{};
jpeg_error_mgr errors{};
std::jmp_buf jump;
};
inline void jpeg_failure(j_common_ptr decoder) {
auto * state = static_cast<jpeg_validator *>(decoder->client_data);
std::longjmp(state->jump, 1);
}
inline void jpeg_message(j_common_ptr decoder, int level) {
if (level < 0) jpeg_failure(decoder);
}
inline void validate_jpeg(const std::vector<unsigned char> & raw, int width, int height) {
auto state = std::make_unique<jpeg_validator>();
auto * decoder = &state->decoder;
decoder->err = jpeg_std_error(&state->errors);
state->errors.error_exit = jpeg_failure;
state->errors.emit_message = jpeg_message;
decoder->client_data = state.get();
if (setjmp(state->jump)) {
jpeg_destroy_decompress(decoder);
throw image_error("invalid or incomplete JPEG");
}
jpeg_create_decompress(decoder);
jpeg_mem_src(decoder, raw.data(), raw.size());
jpeg_read_header(decoder, TRUE);
// The dimension/aggregate checks in validate_url precede all pixel or
// coefficient allocations. Verify both decoders saw the same dimensions.
if (decoder->image_width != unsigned(width) || decoder->image_height != unsigned(height)) {
jpeg_destroy_decompress(decoder);
throw image_error("inconsistent JPEG dimensions");
}
jpeg_start_decompress(decoder);
auto row = (*decoder->mem->alloc_sarray)(reinterpret_cast<j_common_ptr>(decoder),
JPOOL_IMAGE, decoder->output_width * decoder->output_components, 1);
while (decoder->output_scanline < decoder->output_height) {
jpeg_read_scanlines(decoder, row, 1);
}
jpeg_finish_decompress(decoder);
jpeg_destroy_decompress(decoder);
}
inline int digit(unsigned char c) {
if (c >= 'A' && c <= 'Z') return c-'A';
if (c >= 'a' && c <= 'z') return c-'a'+26;
if (c >= '0' && c <= '9') return c-'0'+52;
return c=='+' ? 62 : c=='/' ? 63 : -1;
}
inline uint32_t be32(const unsigned char * p) {
return uint32_t(p[0])<<24 | uint32_t(p[1])<<16 | uint32_t(p[2])<<8 | p[3];
}
inline uint32_t png_crc(const unsigned char * data, size_t size) {
uint32_t crc = 0xffffffffu;
for (size_t i = 0; i < size; ++i) {
crc ^= data[i];
for (int bit = 0; bit < 8; ++bit) crc = (crc >> 1) ^ (0xedb88320u & (0u - (crc & 1)));
}
return crc ^ 0xffffffffu;
}
inline void validate_url(const std::string & url, size_t & decoded, size_t & pixels) {
const auto comma = url.find(',');
const auto header = url.substr(0, comma);
bool png = header == "data:image/png;base64";
require(comma != std::string::npos && (png || header == "data:image/jpeg;base64"), "images must be PNG/JPEG base64 data URLs");
size_t n = url.size()-comma-1;
require(n > 0 && n%4 == 0, "invalid base64 length");
const char * data = url.data()+comma+1;
size_t pad = (data[n-1]=='=') + (data[n-2]=='=');
size_t size = n/4*3-pad;
require(size <= decoded_bytes-decoded, "decoded image aggregate exceeds limit", true);
std::vector<unsigned char> raw;
raw.reserve(size);
for (size_t i=0; i<n; i+=4) {
int a=digit(data[i]), b=digit(data[i+1]);
int c=data[i+2]=='=' ? 0 : digit(data[i+2]);
int d=data[i+3]=='=' ? 0 : digit(data[i+3]);
require(a>=0 && b>=0 && c>=0 && d>=0, "invalid base64 character");
require((data[i+2]!='=' && data[i+3]!='=') || i+4==n, "invalid base64 padding");
require(data[i+2]!='=' || (data[i+3]=='=' && (b&15)==0), "invalid base64 padding bits");
require(data[i+3]!='=' || data[i+2]=='=' || (c&3)==0, "invalid base64 padding bits");
raw.push_back((a<<2)|(b>>4));
if (data[i+2]!='=') raw.push_back((b<<4)|(c>>2));
if (data[i+3]!='=') raw.push_back((c<<6)|d);
}
decoded += raw.size();
require(png ? raw.size()>=24 && std::memcmp(raw.data(),"\x89PNG\r\n\x1a\n",8)==0
: raw.size()>=3 && raw[0]==255 && raw[1]==216 && raw[2]==255, "image MIME mismatch");
int w=0,h=0,c=0;
require(stbi_info_from_memory(raw.data(),raw.size(),&w,&h,&c)!=0 && w>0 && h>0, "invalid image header");
require(size_t(w)<=max_dimension && size_t(h)<=max_dimension, "image dimensions exceed limit", true);
size_t count=size_t(w)*size_t(h);
require(count<=max_pixels-pixels, "image pixel aggregate exceeds limit", true);
pixels += count;
if (png) {
// stb's PNG inflater grows independently of IHDR. Validate IDAT with a
// fixed output buffer first, preventing small-header decompression bombs.
// 16-bit RGBA plus Adam7 row filters fit this conservative pixel bound.
std::vector<unsigned char> idat;
size_t pos=8;
bool end=false;
while (pos+12<=raw.size()) {
size_t len=be32(raw.data()+pos);
require(len<=raw.size()-pos-12, "truncated PNG chunk");
require(png_crc(raw.data()+pos+4,len+4)==be32(raw.data()+pos+8+len), "invalid PNG checksum");
if (std::memcmp(raw.data()+pos+4,"IDAT",4)==0)
idat.insert(idat.end(),raw.begin()+pos+8,raw.begin()+pos+8+len);
if (std::memcmp(raw.data()+pos+4,"IEND",4)==0) { end=true; break; }
pos+=len+12;
}
require(end && !idat.empty(), "incomplete PNG");
std::vector<char> inflated(9*count+8*size_t(h)+1024);
z_stream stream{};
stream.next_in = idat.data();
stream.avail_in = static_cast<uInt>(idat.size());
stream.next_out = reinterpret_cast<Bytef *>(inflated.data());
stream.avail_out = static_cast<uInt>(inflated.size());
require(inflateInit(&stream) == Z_OK, "PNG inflater initialization failed");
int result = inflate(&stream, Z_FINISH);
bool complete = result == Z_STREAM_END && stream.avail_in == 0;
inflateEnd(&stream);
require(complete, "invalid or oversized PNG decompression");
} else {
validate_jpeg(raw, w, h);
}
auto * image=stbi_load_from_memory(raw.data(),raw.size(),&w,&h,&c,3);
require(image!=nullptr, "invalid image pixels");
stbi_image_free(image);
}
// Normalize only actual chat content parts, as in the canonical Go collector.
// Upstream parse_state understands image_url but not Anthropic source objects.
inline size_t validate(json & body, size_t wire_size) {
require(wire_size<=body_bytes, "decision request exceeds limit", true);
std::vector<const std::string *> urls;
auto add=[&](const json & value) {
require(value.is_string(), "image URL must be a string");
require(urls.size()<max_images, "too many decision images", true);
urls.push_back(&value.get_ref<const std::string &>());
};
if (body.contains("images") && !body["images"].is_null()) {
require(body["images"].is_array(), "images must be an array");
for (const auto & url : body["images"]) add(url);
}
auto state=body.find("state");
if (state!=body.end()) {
json * messages=&*state;
if (state->is_object() && state->contains("messages")) messages=&(*state)["messages"];
if (messages->is_array()) for (auto & msg : *messages) {
if (!msg.is_object() || !msg.contains("content") || !msg["content"].is_array()) continue;
for (auto & part : msg["content"]) {
if (!part.is_object() || !part.contains("type")) continue;
if (part["type"]=="image") {
require(part.contains("source") && part["source"].is_object(), "invalid image source");
auto & s=part["source"];
require(s.value("type",std::string())=="base64" && s.contains("media_type") && s["media_type"].is_string() && s.contains("data") && s["data"].is_string(), "invalid image source");
require(s["data"].get_ref<const std::string &>().size()<=encoded_bytes && s["media_type"].get_ref<const std::string &>().size()<=32, "image source exceeds limit",true);
std::string url="data:"+s["media_type"].get<std::string>()+";base64,"+s["data"].get<std::string>();
part=json{{"type","image_url"},{"image_url",{{"url",url}}}};
}
if (part["type"]=="image_url") {
require(part.contains("image_url"), "missing image URL");
auto & u=part["image_url"];
if (u.is_object()) { require(u.contains("url"), "missing image URL"); add(u["url"]); }
else add(u);
}
}
}
}
require(!urls.empty() || wire_size<=text_bytes, "text decision request exceeds limit",true);
size_t encoded=0,decoded=0,pixels=0;
for (const auto * u : urls) { require(u->size()<=encoded_bytes-encoded,"encoded image aggregate exceeds limit",true); encoded+=u->size(); }
for (const auto * u : urls) validate_url(*u,decoded,pixels);
return urls.size();
}
}
+115
View File
@@ -43,6 +43,12 @@
#if __has_include("server-stream.cpp")
#include "server-stream.cpp"
#endif
#if __has_include("server-decision.cpp")
#define LOCALAI_HAS_NATIVE_DECISIONS 1
#include "server-decision.cpp"
#include "decision_compat.h"
#include "decision_images.h"
#endif
#include "server-context.cpp"
// LocalAI
@@ -3301,6 +3307,107 @@ public:
// together in one batch, so a warm scoring call costs roughly one
// forward pass over the new prompt tokens plus one batched pass over
// the candidate tails.
#ifdef LOCALAI_HAS_NATIVE_DECISIONS
grpc::Status SystemOne(ServerContext* context, const backend::ScoreRequest* request,
backend::ScoreResponse* response) {
const auto & decision = ctx_server.impl->decision;
if (decision.type == COMMON_DECISION_TYPE_NONE) {
return grpc::Status(grpc::StatusCode::UNIMPLEMENTED, "This model is not a decision model");
}
try {
if (request->prompt().size() > localai_decision::body_bytes) {
return grpc::Status(grpc::StatusCode::RESOURCE_EXHAUSTED, "Decision request exceeds limit");
}
auto checked_body = localai_decision::json::parse(request->prompt());
const auto image_count = localai_decision::validate(checked_body, request->prompt().size());
const json body = json::parse(checked_body.dump());
if (image_count && !localai_decision::supports_images(decision.can_use_images(),
ctx_server.impl->mctx && mtmd_support_vision(ctx_server.impl->mctx))) {
return grpc::Status(grpc::StatusCode::UNIMPLEMENTED,
"This decision model requires an image-capable projector for image input");
}
const auto questions = decision.parse_questions(body);
std::vector<raw_buffer> files;
const json state = decision.parse_state(body, files);
if (context->IsCancelled()) {
return grpc::Status(grpc::StatusCode::CANCELLED, "Request cancelled by client");
}
auto rd = ctx_server.get_response_reader(); // destructor cancels outstanding tasks
std::vector<server_task> tasks;
size_t expected_results = 0;
for (const auto & question : questions) {
for (size_t variant = 0; variant < decision.n_variants(question); ++variant) {
server_task task(SERVER_TASK_TYPE_DECISION);
task.id = rd.get_new_id();
localai_fill_decision_task(decision, state, questions, question, variant, files,
ctx_server.impl->mctx, ctx_server.impl->init_opt, task);
tasks.push_back(std::move(task));
++expected_results;
}
}
if (decision.can_share_prompt()) {
tasks = server_decision_group_tasks(std::move(tasks), params_base.n_parallel);
}
rd.post_tasks(std::move(tasks));
auto results = rd.wait_for_all([&context]() { return context->IsCancelled(); });
if (results.is_terminated) {
return grpc::Status(grpc::StatusCode::CANCELLED, "Request cancelled by client");
}
if (results.error) {
auto code = grpc::StatusCode::INTERNAL;
const auto * error = dynamic_cast<server_task_result_error *>(results.error.get());
if (!error) {
return grpc::Status(grpc::StatusCode::INTERNAL, "Unexpected decision error type");
}
switch (error->err_type) {
case ERROR_TYPE_INVALID_REQUEST:
case ERROR_TYPE_EXCEED_CONTEXT_SIZE: code = grpc::StatusCode::INVALID_ARGUMENT; break;
case ERROR_TYPE_NOT_SUPPORTED: code = grpc::StatusCode::UNIMPLEMENTED; break;
default: break;
}
return grpc::Status(code, error->err_msg);
}
if (results.results.size() != expected_results) {
return grpc::Status(grpc::StatusCode::INTERNAL, "Unexpected decision result count");
}
json answers = json::object();
int64_t n_tokens = 0;
size_t index = 0;
for (const auto & question : questions) {
std::vector<std::vector<float>> scores;
for (size_t variant = 0; variant < decision.n_variants(question); ++variant) {
auto * result = dynamic_cast<server_task_result_decision *>(results.results[index++].get());
if (!result) {
return grpc::Status(grpc::StatusCode::INTERNAL, "Unexpected decision result type");
}
scores.push_back(result->scores);
n_tokens += result->n_tokens;
}
answers[question.id] = decision.format_answer(question, scores);
}
const auto output = json{
{"model", body.value("model", std::string())}, {"answers", answers},
{"usage", {{"input_tokens", n_tokens}, {"output_tokens", 0}}}
}.dump();
if (output.size() > localai_decision::text_bytes) {
return grpc::Status(grpc::StatusCode::RESOURCE_EXHAUSTED, "Decision response exceeds limit");
}
response->set_response_json(output);
return grpc::Status::OK;
} catch (const localai_decision::image_error & err) {
return grpc::Status(err.too_large ? grpc::StatusCode::RESOURCE_EXHAUSTED : grpc::StatusCode::INVALID_ARGUMENT, err.what());
} catch (const localai_decision::json::exception & err) {
return grpc::Status(grpc::StatusCode::INVALID_ARGUMENT, err.what());
} catch (const common_json_error & err) {
return grpc::Status(grpc::StatusCode::INVALID_ARGUMENT, err.what());
} catch (const std::invalid_argument & err) {
return grpc::Status(grpc::StatusCode::INVALID_ARGUMENT, err.what());
} catch (const std::exception & err) {
return grpc::Status(grpc::StatusCode::INTERNAL, err.what());
}
}
#endif
grpc::Status Score(ServerContext* context, const backend::ScoreRequest* request, backend::ScoreResponse* response) override {
auto auth = checkAuth(context);
if (!auth.ok()) return auth;
@@ -3309,6 +3416,14 @@ public:
if (params_base.model.path.empty()) {
return grpc::Status(grpc::StatusCode::FAILED_PRECONDITION, "Model not loaded");
}
if (request->question_type() == "systemone") {
#ifdef LOCALAI_HAS_NATIVE_DECISIONS
return SystemOne(context, request, response);
#else
return grpc::Status(grpc::StatusCode::UNIMPLEMENTED,
"Native decisions are unavailable in this llama.cpp fork backend");
#endif
}
#ifdef LOCALAI_LLAMA_CPP_NO_SCORE_TASK
(void) request;
(void) response;
+7
View File
@@ -39,6 +39,13 @@ GPU_LIB_SCRIPT="${REPO_ROOT}/scripts/build/package-gpu-libs.sh"
if [ -f "$GPU_LIB_SCRIPT" ]; then
echo "Packaging GPU libraries for BUILD_TYPE=${BUILD_TYPE:-cpu}..."
source "$GPU_LIB_SCRIPT" "$CURDIR/package/lib"
# Native decision validation links zlib/libjpeg. Collect actual ELF
# dependencies rather than assuming those libraries exist on the host.
for binary in "$CURDIR"/package/llama-cpp-*; do
if [ -f "$binary" ]; then
copy_elf_deps "$binary"
fi
done
package_gpu_libs
fi
@@ -1,20 +1,8 @@
From 75220a0d74892e3315f4042274b1efa6195868d8 Mon Sep 17 00:00:00 2001
From: Codex <codex@local>
Date: Mon, 10 Aug 2026 23:05:52 +0000
Subject: [PATCH 1/2] score-patch
---
common/common.cpp | 6 +-
common/common.h | 3 +
tools/server/server-context.cpp | 358 +++++++++++++++++++++++++++++++-
tools/server/server-task.h | 47 +++++
4 files changed, 405 insertions(+), 9 deletions(-)
diff --git a/common/common.cpp b/common/common.cpp
index 2e3f14c..0cec0dc 100644
index aca1949..428b120 100644
--- a/common/common.cpp
+++ b/common/common.cpp
@@ -1636,8 +1636,10 @@ struct llama_context_params common_context_params_to_llama(const common_params &
@@ -1662,8 +1662,10 @@ struct llama_context_params common_context_params_to_llama(const common_params &
auto cparams = llama_context_default_params();
cparams.n_ctx = params.n_ctx;
@@ -28,10 +16,10 @@ index 2e3f14c..0cec0dc 100644
cparams.n_outputs_max_per_seq = std::max(params.n_outputs_max_per_seq, 0);
cparams.n_batch = params.n_batch;
diff --git a/common/common.h b/common/common.h
index 878534d..4001df2 100644
index 04ffbcd..5daa023 100644
--- a/common/common.h
+++ b/common/common.h
@@ -445,6 +445,9 @@ struct common_params {
@@ -455,6 +455,9 @@ struct common_params {
int32_t n_keep = 0; // number of tokens to keep from initial prompt
int32_t n_chunks = -1; // max number of chunks to process (-1 = unlimited)
int32_t n_parallel = 1; // number of parallel sequences to decode
@@ -42,7 +30,7 @@ index 878534d..4001df2 100644
int32_t n_outputs_max = 0; // max outputs in a batch (0 = n_batch)
int32_t n_outputs_max_per_seq = 1; // max outputs per sequence
diff --git a/tools/server/server-context.cpp b/tools/server/server-context.cpp
index 3b5f6a1..d0e18e6 100644
index edb8e2d..9c88feb 100644
--- a/tools/server/server-context.cpp
+++ b/tools/server/server-context.cpp
@@ -48,6 +48,13 @@ static common_speculative_output_limits server_output_limits(const common_params
@@ -59,7 +47,7 @@ index 3b5f6a1..d0e18e6 100644
result.total = std::max<int32_t>(1, result.total);
result.per_seq = std::max<int32_t>(1, result.per_seq);
return result;
@@ -239,6 +246,26 @@ struct server_slot {
@@ -243,6 +250,26 @@ struct server_slot {
std::vector<completion_token_output> generated_token_probs;
@@ -86,7 +74,7 @@ index 3b5f6a1..d0e18e6 100644
bool has_next_token = true;
bool has_new_line = false;
bool truncated = false;
@@ -341,6 +368,10 @@ struct server_slot {
@@ -346,6 +373,10 @@ struct server_slot {
}
generated_tokens.clear();
generated_token_probs.clear();
@@ -97,7 +85,7 @@ index 3b5f6a1..d0e18e6 100644
json_schema = json();
task_prev = std::move(task);
@@ -2271,6 +2302,227 @@ private:
@@ -2308,6 +2339,227 @@ private:
queue_results.send(std::move(res));
}
@@ -325,15 +313,15 @@ index 3b5f6a1..d0e18e6 100644
//
// Functions to process the task
//
@@ -2407,6 +2661,7 @@ private:
case SERVER_TASK_TYPE_INFILL:
@@ -2468,6 +2720,7 @@ private:
case SERVER_TASK_TYPE_EMBEDDING:
case SERVER_TASK_TYPE_RERANK:
case SERVER_TASK_TYPE_DECISION:
+ case SERVER_TASK_TYPE_SCORE:
{
// special case: if input is provided via CLI, tokenize it first
// otherwise, no need to tokenize as it's already done inside the HTTP thread
@@ -2903,6 +3158,13 @@ private:
@@ -2986,6 +3239,13 @@ private:
break; // stop any further processing
}
}
@@ -347,7 +335,7 @@ index 3b5f6a1..d0e18e6 100644
}
void pre_decode() {
@@ -3222,6 +3484,16 @@ private:
@@ -3325,6 +3585,16 @@ private:
n_past = std::min(n_past, slot.alora_invocation_start - 1);
}
@@ -364,7 +352,7 @@ index 3b5f6a1..d0e18e6 100644
const auto n_cache_reuse = slot.task->params.n_cache_reuse;
const bool can_cache_reuse =
@@ -3455,8 +3727,12 @@ private:
@@ -3578,8 +3848,12 @@ private:
bool do_checkpoint = params_base.n_ctx_checkpoints > 0;
@@ -379,7 +367,7 @@ index 3b5f6a1..d0e18e6 100644
// make a checkpoint of the parts of the memory that cannot be rolled back.
// checkpoints are created only if:
@@ -3444,9 +3720,16 @@ private:
@@ -3670,13 +3944,46 @@ private:
// embedding requires all tokens in the batch to be output;
// MTP also wants logits at every prompt position so the
// streaming hook can mirror t_h_nextn into ctx_dft.
@@ -397,7 +385,7 @@ index 3b5f6a1..d0e18e6 100644
+ /* output = */ slot.need_embd() || need_score_logit,
/* is_prompt = */ true);
slot.prompt.tokens.push_back(cur_tok);
@@ -3454,2 +3737,28 @@ private:
+ // score tasks: break at the shared-prompt boundary so the checkpoint
+ // below lands exactly there — the other candidates of the same
+ // scoring call re-process only their own tokens. Also break at the
@@ -426,7 +414,8 @@ index 3b5f6a1..d0e18e6 100644
+
// break at the last user message, or at user messages at least min step past the last checkpoint
if (do_checkpoint && spans.is_user_start(slot.prompt.n_tokens())) {
@@ -3573,6 +3882,15 @@ private:
const auto pos = slot.prompt.n_tokens();
@@ -3719,6 +4026,15 @@ private:
const bool is_user_start = spans.is_user_start(n_tokens_start);
const bool is_last_user_message = n_tokens_start == last_user_pos;
@@ -442,7 +431,7 @@ index 3b5f6a1..d0e18e6 100644
// entire prompt has been processed
if (slot.prompt.n_tokens() == slot.task->n_tokens()) {
slot.state = SLOT_STATE_DONE_PROMPT;
@@ -3588,8 +3906,8 @@ private:
@@ -3734,8 +4050,8 @@ private:
slot.init_sampler();
} else {
// skip ordinary mid-prompt checkpoints, unless the batch starts a user
@@ -453,7 +442,7 @@ index 3b5f6a1..d0e18e6 100644
do_checkpoint = false;
}
}
@@ -3606,10 +3924,10 @@ private:
@@ -3752,10 +4068,10 @@ private:
// do not checkpoint after mtmd chunks
do_checkpoint = do_checkpoint && !has_mtmd;
@@ -466,7 +455,7 @@ index 3b5f6a1..d0e18e6 100644
n_tokens_start > slot.prompt.checkpoints.back().n_tokens + params_base.checkpoint_min_step);
SLT_DBG(slot, "main/do_checkpoint = %s, pos_min = %d, pos_max = %d\n", do_checkpoint ? "yes" : "no", pos_min, pos_max);
@@ -3772,6 +4090,13 @@ private:
@@ -3943,6 +4259,13 @@ private:
}
}
@@ -480,7 +469,7 @@ index 3b5f6a1..d0e18e6 100644
if (!is_inside_view(slot.i_batch)) {
// the required token not in this sub-batch, skip
return;
@@ -3793,6 +4118,25 @@ private:
@@ -3971,6 +4294,25 @@ private:
return;
}
@@ -507,12 +496,12 @@ index 3b5f6a1..d0e18e6 100644
// prompt evaluated for next-token prediction
diff --git a/tools/server/server-task.h b/tools/server/server-task.h
index 6275ec7..5bedf19 100644
index 8c5fa9a..4b38805 100644
--- a/tools/server/server-task.h
+++ b/tools/server/server-task.h
@@ -13,10 +13,25 @@
@@ -12,11 +12,26 @@
#include "server-common.h"
using json = nlohmann::ordered_json;
+// SERVER_TASK_TYPE_SCORE emits one logits output per candidate token (plus
+// the forced last-token output), and the context's output budget
@@ -532,11 +521,12 @@ index 6275ec7..5bedf19 100644
SERVER_TASK_TYPE_COMPLETION,
SERVER_TASK_TYPE_EMBEDDING,
SERVER_TASK_TYPE_RERANK,
SERVER_TASK_TYPE_DECISION,
+ SERVER_TASK_TYPE_SCORE,
SERVER_TASK_TYPE_INFILL,
SERVER_TASK_TYPE_CANCEL,
SERVER_TASK_TYPE_CONTROL,
@@ -153,6 +168,18 @@ struct server_task {
@@ -156,6 +171,18 @@ struct server_task {
task_params params;
server_tokens tokens;
@@ -555,7 +545,7 @@ index 6275ec7..5bedf19 100644
// only used by CLI, this allow tokenizing CLI inputs on server side
// we need this because mtmd_context and vocab are not accessible outside of server_context
bool cli = false;
@@ -197,6 +224,7 @@ struct server_task {
@@ -234,6 +261,7 @@ struct server_task {
switch (type) {
case SERVER_TASK_TYPE_COMPLETION:
case SERVER_TASK_TYPE_INFILL:
@@ -563,7 +553,7 @@ index 6275ec7..5bedf19 100644
return true;
default:
return false;
@@ -494,6 +522,25 @@ struct server_task_result_rerank : server_task_result {
@@ -509,6 +537,25 @@ struct server_task_result_decision : server_task_result {
virtual json to_json() override;
};
@@ -589,5 +579,3 @@ index 6275ec7..5bedf19 100644
struct server_task_result_error : server_task_result {
error_type err_type = ERROR_TYPE_SERVER;
std::string err_msg;
--
2.39.5
@@ -1,5 +1,5 @@
diff --git a/tools/mtmd/mtmd-helper-gen.cpp b/tools/mtmd/mtmd-helper-gen.cpp
index 1c58d3ae1..196cbd433 100644
index 5fb7ea9..b554bf2 100644
--- a/tools/mtmd/mtmd-helper-gen.cpp
+++ b/tools/mtmd/mtmd-helper-gen.cpp
@@ -50,29 +50,38 @@ static llama_token find_special_token(const llama_vocab * vocab, const std::stri
@@ -86,7 +86,7 @@ index 1c58d3ae1..196cbd433 100644
// the prompt above holds the whole text stream up to tts_eos, so every generated
// frame adds tts_pad on top of the codes embedding
@@ -302,31 +317,60 @@ public:
@@ -301,31 +316,60 @@ public:
}
int32_t get_output(int32_t * out_sample_rate, const char ** out_data, size_t * out_data_len, int64_t * out_n_samples) override {
@@ -156,7 +156,7 @@ index 1c58d3ae1..196cbd433 100644
private:
bool ensure_cache() {
if (specials_ok) {
@@ -370,7 +414,7 @@ private:
@@ -369,7 +413,7 @@ private:
LOG_ERR("mtmd_helper_gen_audio: mmproj has no speaker/audio encoder\n");
return false;
}
@@ -165,7 +165,7 @@ index 1c58d3ae1..196cbd433 100644
mtmd_input_text text{ marker.c_str(), marker.size(), false, true };
mtmd_input_chunks * chunks = mtmd_input_chunks_init();
const mtmd_bitmap * bptr = bitmap;
@@ -456,6 +500,9 @@ private:
@@ -455,6 +499,9 @@ private:
std::vector<float> h_state_buf;
mtmd_helper_gen_audio_outtype out_type = MTMD_HELPER_GEN_AUDIO_OUTTYPE_WAV;
std::vector<char> out_buf;
@@ -175,7 +175,7 @@ index 1c58d3ae1..196cbd433 100644
};
// settings that only live in the reference's per-pack yaml, not in the checkpoint
@@ -1024,6 +1071,14 @@ void mtmd_helper_gen_audio_reset(mtmd_helper_gen_audio * ctx) {
@@ -1022,6 +1069,14 @@ void mtmd_helper_gen_audio_reset(mtmd_helper_gen_audio * ctx) {
}
}
@@ -190,7 +190,7 @@ index 1c58d3ae1..196cbd433 100644
int32_t mtmd_helper_gen_audio_set_input(mtmd_helper_gen_audio * ctx, const mtmd_helper_gen_audio_inp * inp) {
if (!ctx->pipeline) {
LOG_ERR("mtmd_helper_gen_audio: unsupported or missing gen-audio pipeline\n");
@@ -1060,3 +1115,10 @@ int32_t mtmd_helper_gen_audio_get_output(mtmd_helper_gen_audio * ctx, int32_t *
@@ -1058,3 +1113,10 @@ int32_t mtmd_helper_gen_audio_get_output(mtmd_helper_gen_audio * ctx, int32_t *
}
return ctx->pipeline->get_output(out_sample_rate, out_data, out_data_len, out_n_samples);
}
@@ -202,10 +202,10 @@ index 1c58d3ae1..196cbd433 100644
+ return ctx->pipeline->flush();
+}
diff --git a/tools/mtmd/mtmd-helper.h b/tools/mtmd/mtmd-helper.h
index 832f7171a..3eaa01aab 100644
index 7436230..acbfefc 100644
--- a/tools/mtmd/mtmd-helper.h
+++ b/tools/mtmd/mtmd-helper.h
@@ -175,6 +175,7 @@ enum mtmd_helper_gen_audio_outtype {
@@ -204,6 +204,7 @@ enum mtmd_helper_gen_audio_outtype {
MTMD_HELPER_GEN_AUDIO_OUTTYPE_WAV, // WAV PCM 16-bit LE, mono
};
struct mtmd_helper_gen_audio_inp {
@@ -213,7 +213,7 @@ index 832f7171a..3eaa01aab 100644
llama_seq_id seq_id;
const char * prompt;
@@ -190,6 +191,8 @@ struct mtmd_helper_gen_audio_inp {
@@ -219,6 +220,8 @@ struct mtmd_helper_gen_audio_inp {
enum mtmd_helper_gen_audio_outtype out_type;
};
@@ -222,7 +222,7 @@ index 832f7171a..3eaa01aab 100644
MTMD_API mtmd_helper_gen_audio * mtmd_helper_gen_audio_init(
struct llama_context * lctx,
struct mtmd_context * mctx);
@@ -221,6 +224,8 @@ MTMD_API int32_t mtmd_helper_gen_audio_step_gen(
@@ -250,6 +253,8 @@ MTMD_API int32_t mtmd_helper_gen_audio_step_gen(
// out_data valid until next get_output() or reset() call
// out_n_samples (optional, can be NULL) receives the number of generated PCM samples
@@ -231,7 +231,7 @@ index 832f7171a..3eaa01aab 100644
MTMD_API int32_t mtmd_helper_gen_audio_get_output(
mtmd_helper_gen_audio * ctx,
int32_t * out_sample_rate,
@@ -228,6 +233,10 @@ MTMD_API int32_t mtmd_helper_gen_audio_get_output(
@@ -257,6 +262,10 @@ MTMD_API int32_t mtmd_helper_gen_audio_get_output(
size_t * out_data_len,
int64_t * out_n_samples);
@@ -242,7 +242,7 @@ index 832f7171a..3eaa01aab 100644
#ifdef __cplusplus
} // extern "C"
#endif
@@ -254,8 +263,41 @@ struct mtmd_helper_gen_audio_deleter {
@@ -283,8 +292,41 @@ struct mtmd_helper_gen_audio_deleter {
};
using gen_audio_ptr = std::unique_ptr<mtmd_helper_gen_audio, mtmd_helper_gen_audio_deleter>;
struct gen_audio {
@@ -285,7 +285,7 @@ index 832f7171a..3eaa01aab 100644
void reset() {
mtmd_helper_gen_audio_reset(ctx.get());
}
@@ -271,6 +313,9 @@ struct gen_audio {
@@ -300,6 +342,9 @@ struct gen_audio {
int32_t get_output(int32_t * out_sample_rate, const char ** out_data, size_t * out_data_len, int64_t * out_n_samples = nullptr) {
return mtmd_helper_gen_audio_get_output(ctx.get(), out_sample_rate, out_data, out_data_len, out_n_samples);
}
@@ -296,10 +296,10 @@ index 832f7171a..3eaa01aab 100644
} // namespace mtmd_helper
diff --git a/tools/server/server-context.cpp b/tools/server/server-context.cpp
index 9069463fe..b7fa1e534 100644
index 9c88feb..064bb51 100644
--- a/tools/server/server-context.cpp
+++ b/tools/server/server-context.cpp
@@ -16,6 +16,7 @@
@@ -17,6 +17,7 @@
#include "speculative.h"
#include "mtmd.h"
#include "mtmd-helper.h"
@@ -319,7 +319,7 @@ index 9069463fe..b7fa1e534 100644
}
auto result = common_speculative_get_output_limits(
@@ -212,6 +214,30 @@ struct server_slot {
@@ -215,6 +217,30 @@ struct server_slot {
mtmd_context * mctx = nullptr;
mtmd::batch_ptr mbatch = nullptr;
@@ -350,7 +350,7 @@ index 9069463fe..b7fa1e534 100644
// speculative decoding
common_speculative * spec;
@@ -391,6 +417,8 @@ struct server_slot {
@@ -395,6 +421,8 @@ struct server_slot {
// clear multimodal state
mbatch.reset();
@@ -359,9 +359,9 @@ index 9069463fe..b7fa1e534 100644
}
void init_sampler() const {
@@ -829,6 +857,14 @@ public:
mtmd_context * mctx = nullptr;
const llama_vocab * vocab = nullptr;
@@ -871,6 +899,14 @@ public:
server_decision_context decision;
+ bool has_cap_tts() const {
+ return mctx != nullptr && mtmd_gen_audio_get_info(mctx).type != MTMD_GEN_AUDIO_TYPE_NONE;
@@ -374,7 +374,7 @@ index 9069463fe..b7fa1e534 100644
server_queue queue_tasks;
server_response queue_results;
@@ -1288,6 +1324,10 @@ private:
@@ -1350,6 +1386,10 @@ private:
slot.mctx = mctx;
slot.prompt.tokens.has_mtmd = mctx != nullptr;
@@ -385,7 +385,7 @@ index 9069463fe..b7fa1e534 100644
SLT_TRC(slot, "new slot, n_ctx = %d\n", slot.n_ctx);
slot.callback_on_release = [this](int id_slot) {
@@ -1748,6 +1788,28 @@ private:
@@ -1830,6 +1870,28 @@ private:
SLT_DBG(slot, "launching slot : %s\n", safe_json_to_str(slot.to_json()).c_str());
@@ -414,7 +414,7 @@ index 9069463fe..b7fa1e534 100644
// initialize samplers
if (task.need_sampling()) {
try {
@@ -1765,6 +1827,9 @@ private:
@@ -1847,6 +1909,9 @@ private:
// TODO: getting pre sampling logits is not yet supported with backend sampling
use_backend_sampling &= !need_pre_sample_logits;
@@ -424,7 +424,7 @@ index 9069463fe..b7fa1e534 100644
// TODO: tmp until backend sampling is fully implemented
if (use_backend_sampling) {
llama_set_sampler(ctx_tgt, slot.id, common_sampler_get(slot.smpl.get()));
@@ -1783,9 +1848,13 @@ private:
@@ -1872,9 +1937,13 @@ private:
slot.task = std::make_unique<const server_task>(std::move(task));
@@ -441,7 +441,7 @@ index 9069463fe..b7fa1e534 100644
// reset server kill-switch counter
n_empty_consecutive = 0;
@@ -2050,6 +2119,18 @@ private:
@@ -2139,6 +2208,18 @@ private:
queue_results.send(std::move(res));
}
@@ -460,15 +460,18 @@ index 9069463fe..b7fa1e534 100644
void send_final_response(server_slot & slot) {
auto res = std::make_unique<server_task_result_cmpl_final>();
@@ -2556,6 +2637,7 @@ private:
case SERVER_TASK_TYPE_EMBEDDING:
@@ -2721,6 +2802,7 @@ private:
case SERVER_TASK_TYPE_RERANK:
case SERVER_TASK_TYPE_DECISION:
case SERVER_TASK_TYPE_SCORE:
+ case SERVER_TASK_TYPE_TTS:
{
// special case: if input is provided via CLI, tokenize it first
// otherwise, no need to tokenize as it's already done inside the HTTP thread
@@ -3007,1 +3089,9 @@ private:
@@ -3179,6 +3261,14 @@ private:
return;
}
+ // note: TTS slots bypass the shared batch entirely
+ try {
+ process_tts_slots();
@@ -478,7 +481,9 @@ index 9069463fe..b7fa1e534 100644
+ }
+
GGML_ASSERT(batch.slot_batched || batch.size() == 0);
@@ -3074,10 +3164,77 @@ private:
if (batch.slot_batched) {
@@ -3248,10 +3338,77 @@ private:
}
}
@@ -556,7 +561,7 @@ index 9069463fe..b7fa1e534 100644
if (slot.state == SLOT_STATE_GENERATING && slot.prompt.n_tokens() + 1 >= slot.n_ctx) {
if (!params_base.ctx_shift) {
// this check is redundant (for good)
@@ -3150,7 +3307,7 @@ private:
@@ -3324,7 +3481,7 @@ private:
// determine which slots are generating and drafting
iterate(slots, [&](server_slot & slot) {
@@ -565,7 +570,7 @@ index 9069463fe..b7fa1e534 100644
return;
}
@@ -3284,7 +3441,7 @@ private:
@@ -3458,7 +3615,7 @@ private:
return; // batch is full, skip remaining slots
}
@@ -574,16 +579,16 @@ index 9069463fe..b7fa1e534 100644
return;
}
@@ -4433,6 +4590,8 @@ server_context_meta server_context::get_meta() const {
@@ -4678,6 +4835,8 @@ server_context_meta server_context::get_meta() const {
/* has_inp_image */ impl->chat_params.allow_image,
/* has_inp_audio */ impl->chat_params.allow_audio,
/* has_inp_video */ impl->chat_params.allow_video,
+ /* has_cap_chat */ impl->has_cap_chat(),
+ /* has_cap_tts */ impl->has_cap_tts(),
/* json_ui_settings */ impl->json_ui_settings,
/* slot_n_ctx */ impl->get_slot_n_ctx(),
/* slot_n_ctx */ impl->n_ctx_slot(),
/* pooling_type */ llama_pooling_type(impl->ctx_tgt),
@@ -4512,6 +4671,11 @@ std::unique_ptr<server_res_generator> server_routes::handle_completions_impl(
@@ -4751,6 +4910,11 @@ std::unique_ptr<server_res_generator> server_routes::handle_completions_impl(
res->set_req(&req); // will also set spipe if needed
@@ -595,7 +600,7 @@ index 9069463fe..b7fa1e534 100644
int32_t sse_ping_interval = params.sse_ping_interval;
try {
@@ -5399,6 +5563,150 @@ void server_routes::init_routes() {
@@ -5776,6 +5940,150 @@ void server_routes::init_routes() {
return res;
};
@@ -747,10 +752,10 @@ index 9069463fe..b7fa1e534 100644
auto res = create_response();
diff --git a/tools/server/server-context.h b/tools/server/server-context.h
index f9ab1132b..610512678 100644
index c554bb9..f025792 100644
--- a/tools/server/server-context.h
+++ b/tools/server/server-context.h
@@ -22,6 +22,8 @@ struct server_context_meta {
@@ -23,6 +23,8 @@ struct server_context_meta {
bool has_inp_image;
bool has_inp_audio;
bool has_inp_video;
@@ -759,19 +764,19 @@ index f9ab1132b..610512678 100644
json json_ui_settings;
int slot_n_ctx;
enum llama_pooling_type pooling_type;
@@ -151,6 +153,7 @@ struct server_routes {
@@ -152,6 +154,7 @@ struct server_routes {
server_http_context::handler_t post_embeddings;
server_http_context::handler_t post_embeddings_oai;
server_http_context::handler_t post_rerank;
+ server_http_context::handler_t post_tts;
server_http_context::handler_t post_systemone;
server_http_context::handler_t get_lora_adapters;
server_http_context::handler_t post_lora_adapters;
diff --git a/tools/server/server-task.cpp b/tools/server/server-task.cpp
index 1ee677553..939630b8b 100644
index a5c33c0..7e59745 100644
--- a/tools/server/server-task.cpp
+++ b/tools/server/server-task.cpp
@@ -1497,6 +1497,17 @@ json server_task_result_rerank::to_json() {
@@ -1506,6 +1506,17 @@ json server_task_result_decision::to_json() {
};
}
@@ -790,7 +795,7 @@ index 1ee677553..939630b8b 100644
// server_task_result_error
//
diff --git a/tools/server/server-task.h b/tools/server/server-task.h
index 5bedf1987..e6ca67a65 100644
index 4b38805..d79ea4f 100644
--- a/tools/server/server-task.h
+++ b/tools/server/server-task.h
@@ -10,6 +10,7 @@
@@ -799,9 +804,9 @@ index 5bedf1987..e6ca67a65 100644
#include "server-common.h"
+#include "mtmd-helper.h"
using json = nlohmann::ordered_json;
@@ -42,6 +43,7 @@ enum server_task_type {
// SERVER_TASK_TYPE_SCORE emits one logits output per candidate token (plus
@@ -43,6 +44,7 @@ enum server_task_type {
SERVER_TASK_TYPE_SLOT_ERASE,
SERVER_TASK_TYPE_GET_LORA,
SERVER_TASK_TYPE_SET_LORA,
@@ -809,7 +814,7 @@ index 5bedf1987..e6ca67a65 100644
};
// TODO: change this to more generic "response_format" to replace the "format_response_*" in server-common
@@ -202,6 +204,9 @@ struct server_task {
@@ -225,6 +227,9 @@ struct server_task {
// used by SERVER_TASK_TYPE_SET_LORA
std::map<int, float> set_lora; // mapping adapter ID -> scale
@@ -819,15 +824,15 @@ index 5bedf1987..e6ca67a65 100644
server_task() = default;
server_task(server_task_type type) : type(type) {}
@@ -235,6 +240,7 @@ struct server_task {
@@ -249,6 +254,7 @@ struct server_task {
switch (type) {
case SERVER_TASK_TYPE_COMPLETION:
case SERVER_TASK_TYPE_INFILL:
+ case SERVER_TASK_TYPE_TTS:
return true;
default:
return false;
@@ -494,5 +500,15 @@ struct server_task_result_embd : server_task_result {
case SERVER_TASK_TYPE_DECISION:
return !decision.labels.empty();
@@ -521,5 +527,15 @@ struct server_task_result_embd : server_task_result {
json to_json_oaicompat();
};
+2
View File
@@ -45,6 +45,8 @@ done
cp -r CMakeLists.txt llama.cpp/tools/grpc-server/
cp -r grpc-server.cpp llama.cpp/tools/grpc-server/
cp -r decision_compat.h llama.cpp/tools/grpc-server/
cp -r decision_images.h llama.cpp/tools/grpc-server/
# Model-load diagnostics (included by grpc-server.cpp) and their standalone
# regression test.
cp -r model_load_error.h llama.cpp/tools/grpc-server/
@@ -0,0 +1,75 @@
# Native decision images
Run from the repository root after obtaining the pinned llama.cpp checkout:
```sh
bash backend/cpp/llama-cpp/tests/verify-decision-images.sh
```
Prerequisites: C++17, zlib and libjpeg development headers/libraries, Python 3
with Pillow (fixture generation only). Ubuntu: `libjpeg-dev zlib1g-dev`;
macOS: `brew install jpeg-turbo zlib`. CMake requires these libraries only when
the checkout has native decisions; older forks retain their dependency guard.
Both Docker builder paths install the packages in the shared compile stage,
including builds using cached base images. Darwin CI installs the Homebrew
packages. The llama.cpp packager collects the executable dependency closure, including
zlib/libjpeg. Darwin uses its existing dylib collection.
This compiles the production validation helper with upstream stb, zlib and
libjpeg. zlib requires a complete stream and valid Adler-32 within a fixed
output budget. libjpeg decodes all scans with warnings treated as errors, so
synthetic EOI recovery and short entropy scans cannot pass. Dimensions and
aggregate pixels are checked before decoder pixel/coefficient allocation. It checks
strict base64, MIME matching, count/byte/dimension/pixel limits, PNG decompression
bombs, invalid Adler-32 with valid chunk CRC, baseline/progressive JPEG,
missing EOI, truncated scans with appended EOI, embedded markers, truncation, chat-only collection, Anthropic normalization, empty-image
text limits, capability combinations, and parity with Go's canonical limits.
Fixtures are generated locally; no image or model downloads occur.
Build the native backend normally with `make backends/llama-cpp`. For a prepared
CPU checkout and extracted distro gRPC dependencies, the existing adapter is:
```sh
DEPS_ROOT=/path/to/deps CPU_BUILD=/path/to/llama.cpp/build-cpu \
bash backend/cpp/llama-cpp/tests/build-decision-bridge.sh
```
The adapter compiles the actual prepared grpc-server.cpp and links upstream CPU
libraries. It is not a replacement implementation or mocked backend.
Start the resulting `grpc-server --addr=127.0.0.1:50061`, then run:
```sh
PYTHONPATH=/path/to/build-decision-validation \
python3 backend/cpp/llama-cpp/tests/decision-image-smoke.py \
--model /models/OpenJev-Q4_K_M.gguf \
--projector /models/mmproj-OpenJev-Q8_0.gguf \
--fixtures backend/cpp/llama-cpp/llama.cpp/build-image-tests/fixtures.json
```
The smoke uses CPU only, four threads, one slot, 8192 context, batch 512. It first
loads without a projector and asserts explicit unsupported plus direct-RPC
safety errors, then loads with the projector and compares red/blue probabilities.
It requires existing weights and a checksum-verified projector; it downloads
nothing. These are direct RPC tests, **not** public HTTP/router E2E evidence.
Check the production CMake dependency block and the non-decision fork guard:
```sh
python3 backend/cpp/llama-cpp/tests/verify-image-build-wiring.py
```
This focused check does not replace a full backend or platform build.
## CI gate
`.github/workflows/decision-images.yml` runs both verification commands on
relevant pull requests and master pushes, or by manual dispatch. It installs
C++17, Python/Pillow, CMake, zlib and libjpeg development dependencies and fetches
only the two vendor headers at `LLAMA_VERSION` from the backend Makefile (not a
floating upstream branch). No model, projector, GPU or full backend build is
needed. The path filters include this workflow, the backend helper/tests/build
files and upstream pin, Go limits, Dockerfiles and Darwin dependency setup.
This gate is separate from `backend/cpp/run-unit-tests.sh`: that stdlib-only
runner discovers `*_test.cpp`, not the dependency-bearing `decision-images.cpp`.
+96
View File
@@ -0,0 +1,96 @@
# Native decision bridge validation
The stock dependency is pinned to `bed0a856606ee4a24a164066f73d2379447033f5`.
`Score(question_type="systemone")` uses upstream decision tasks internally, not
HTTP. Plain Score keeps its existing admission checks. Older dependencies without
`server-decision.cpp` return gRPC `UNIMPLEMENTED` for this request type.
## Decision signature compatibility
This pin includes upstream Nimble support, in addition to OpenJev, Lev, Kev,
and Laya. The native bridge forwards the complete parsed question collection
when upstream's `fill_task` accepts it, as required by Nimble's schema framing.
`decision_compat.h` detects the callable C++ signature at compile time; older
native-decision forks still use their original single-question signature.
Forks without native decision support retain the existing `UNIMPLEMENTED` guard.
The standalone `decision_compat_test.cpp` checks both signatures and that the
full collection is passed by reference, not replaced with a singleton.
It is automatically discovered by `backend/cpp/run-unit-tests.sh`.
Signature and compile validation do not establish Nimble model accuracy or
runtime support for every artifact. Nimble weights are not part of this test
fixture. The official `ggml-org/Bespoke-Nimble-9B-v3-GGUF` model card declares
CC-BY-NC-4.0; check its restrictions before deployment.
## CPU build
Use a fresh stock checkout at the pin in `backend/cpp/llama-cpp/llama.cpp`.
Do not reuse a customized developer checkout. Apply patches once:
```sh
cd backend/cpp/llama-cpp/llama.cpp
git apply --check ../patches/0001-add-server-task-type-score.patch
git apply ../patches/0001-add-server-task-type-score.patch
git apply --check ../patches/0002-add-server-task-type-tts.patch
git apply ../patches/0002-add-server-task-type-tts.patch
cmake -S . -B build-cpu -DGGML_NATIVE=OFF -DLLAMA_OPENSSL=OFF \
-DLLAMA_CURL=OFF -DBUILD_SHARED_LIBS=OFF
cmake --build build-cpu --target llama-server -j2
```
For the normal product build, start instead from a fresh unpatched checkout and
run `make -C backend/cpp/llama-cpp grpc-server JOBS=2`; preparation applies the
patches and stages the bridge. This requires CMake packages for gRPC, protobuf,
and Abseil, plus `protoc` and `grpc_cpp_plugin`.
`build-decision-bridge.sh` is an alternative link validation for distro packages
that lack `ProtobufConfig.cmake`. It uses the already patched CPU static libraries
and the bridge source staged by `prepare.sh`. Do not apply patches twice when
staging that source. Set `CPU_BUILD` to the absolute `build-cpu` directory,
`OUT_DIR` to a scratch output directory, and optionally `DEPS_ROOT` to the root
of **locally extracted** distro packages. It does not install or download anything.
It also generates Python bindings, requiring `grpc_python_plugin`.
## Direct RPC smoke
The upstream test fixture `ggml-org/tinylaya-for-testing-gguf` has file
`tinylaya-for-testing-Q8_0.gguf`, size **97,200,288 bytes**, SHA-256
`a8b2b8f7fe6b7e10a884c55bf72362b0a8701e40dc3f332831d58246e5fa0b70`.
This is a test model, not a production gallery recommendation.
Start the backend with cores disabled, using the matching library path when
validating extracted distro dependencies:
```sh
ulimit -c 0
export LD_LIBRARY_PATH="$DEPS_ROOT/usr/lib/x86_64-linux-gnu${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}"
"$OUT_DIR/grpc-server" --addr 127.0.0.1:50051
```
In another terminal (Python requires `grpcio` and `protobuf`):
```sh
PYTHONPATH="$OUT_DIR" python3 backend/cpp/llama-cpp/tests/decision_smoke.py \
--address 127.0.0.1:50051 --model "$MODEL_FILE"
```
The smoke covers multiquestion choice/score/noul, normalized probabilities,
positive input and explicit zero output usage, concurrent calls, invalid JSON,
client cancellation/recovery, ordinary Score disabled/enabled, and missing
metadata. It reloads models; do not share the backend with another test runner.
The missing-metadata test makes a temporary equal-length metadata-key rename of
the fixture and explicitly enables embeddings, as appropriate for this encoder.
Limits: immediate client cancellation does not prove interruption during active
evaluation. The tiny fixture finishes too quickly for a deterministic timing-only
assertion; a queue barrier or server-side instrumentation is needed for that gate.
TTS is compiled and linked, not runtime-tested by this text-only fixture. Older
pin compile validation does not establish every supported fork's full build.
For bounded image/projector support and its separate runtime checks, see
[README-decision-images.md](README-decision-images.md).
The metadata-stripped encoder with embeddings disabled and `-np 1` aborts
in warmup at `llama-context.cpp`'s output-budget assertion on **clean unpatched**
upstream at the pinned revision as well. This is a preexisting invalid-fixture
configuration hazard, not a decision-dispatch or Score/TTS patch regression.
Do not use that configuration as the missing-metadata test.
@@ -0,0 +1,50 @@
#!/usr/bin/env bash
# SPDX-License-Identifier: MIT
# CPU validation against an already prepared stock checkout. No downloads.
# Run prepare.sh on a clean checkout first. DEPS_ROOT is an optional extracted
# distro /usr tree parent, not a system installation. Output remains local.
set -euo pipefail
root=$(git rev-parse --show-toplevel)
backend="$root/backend/cpp/llama-cpp"
source="$backend/llama.cpp"
out=${OUT_DIR:-"$source/build-decision-validation"}
mkdir -p "$out"
deps=${DEPS_ROOT:-/}
export LD_LIBRARY_PATH="$deps/usr/lib/x86_64-linux-gnu${LD_LIBRARY_PATH:+:$LD_LIBRARY_PATH}"
export PKG_CONFIG_SYSROOT_DIR="$deps"
export PKG_CONFIG_PATH="$deps/usr/lib/x86_64-linux-gnu/pkgconfig${PKG_CONFIG_PATH:+:$PKG_CONFIG_PATH}"
# Build upstream libraries without the optional grpc CMake subdirectory: distro
# protobuf packages may expose FindProtobuf rather than ProtobufConfig.cmake.
# An existing CPU build can be supplied to avoid rebuilding them.
build=${CPU_BUILD:?set CPU_BUILD to the patched upstream CPU build directory}
protoc -I "$root/backend" --cpp_out="$out" --grpc_out="$out" \
--plugin=protoc-gen-grpc="$deps/usr/bin/grpc_cpp_plugin" "$root/backend/backend.proto"
protoc -I "$root/backend" --python_out="$out" --grpc_out="$out" \
--plugin=protoc-gen-grpc="$deps/usr/bin/grpc_python_plugin" "$root/backend/backend.proto"
includes=(-I"$out" -I"$deps/usr/include")
for dir in '' include common vendor ggml/include tools/mtmd tools/server; do
includes+=(-I"$source/$dir")
done
cxx=${CXX:-g++}
"$cxx" -O0 -std=c++17 -pthread "${includes[@]}" -c "$source/tools/grpc-server/grpc-server.cpp" -o "$out/grpc-server.o"
for file in backend.pb backend.grpc.pb; do
"$cxx" -O0 -std=c++17 -pthread "${includes[@]}" -c "$out/$file.cc" -o "$out/$file.o"
done
image_libs=()
if [[ -f "$source/tools/server/server-decision.cpp" ]]; then
image_libs=(-lz -ljpeg)
fi
libs=()
for lib in common/llama-common common/llama-common-base tools/mtmd/mtmd src/llama ggml/src/ggml ggml/src/ggml-cpu ggml/src/ggml-base vendor/hash/vendor-hash vendor/cpp-httplib/cpp-httplib; do
libs+=("$build/${lib%/*}/lib${lib##*/}.a")
done
# pkg-config emits a linker flag list, so intentional word splitting here.
# shellcheck disable=SC2046
"$cxx" -pthread "$out/grpc-server.o" "$out/backend.pb.o" "$out/backend.grpc.pb.o" "${libs[@]}" \
-L"$deps/usr/lib/x86_64-linux-gnu" -Wl,-rpath-link,"$deps/usr/lib/x86_64-linux-gnu" \
-lgrpc++_reflection $(pkg-config --libs grpc++) \
-labsl_flags_parse -labsl_flags_usage -labsl_flags_usage_internal \
-labsl_flags_commandlineflag -labsl_flags_commandlineflag_internal \
-labsl_flags_config -labsl_flags_internal -labsl_flags_reflection \
-labsl_flags_marshalling -lprotobuf "${image_libs[@]}" -ldl -lm -lgomp -o "$out/grpc-server"
printf 'Built %s\n' "$out/grpc-server"
@@ -0,0 +1,66 @@
# SPDX-License-Identifier: MIT
"""Direct RPC only; public API/router end-to-end tests are a separate gate."""
import argparse
import json
import math
import grpc
import backend_pb2 as pb
import backend_pb2_grpc as rpc
p = argparse.ArgumentParser()
p.add_argument('--address', default='127.0.0.1:50061')
p.add_argument('--model', required=True)
p.add_argument('--projector', required=True)
p.add_argument('--fixtures', required=True)
a = p.parse_args()
f = json.load(open(a.fixtures))
channel = grpc.insecure_channel(a.address, options=[('grpc.max_send_message_length', 20 << 20)])
grpc.channel_ready_future(channel).result(timeout=20)
s = rpc.BackendStub(channel)
def load(projector):
r = s.LoadModel(pb.ModelOptions(ModelFile=a.model, MMProj=projector,
ContextSize=8192, NBatch=512, Threads=4, NGPULayers=0,
Options=['parallel:1']), timeout=600)
assert r.success, r
print('LOAD', 'vision' if projector else 'no projector', 'PASS', flush=True)
def body(image):
return {'state': {}, 'images': [image], 'questions': {'color': {
'type': 'choice', 'instructions': 'What is the dominant color of the image?',
'criteria': {'red': None, 'blue': None}}}}
def score(b):
return s.Score(pb.ScoreRequest(question_type='systemone', prompt=json.dumps(b)), timeout=600)
def reject(b, code):
try:
score(b)
raise AssertionError('request unexpectedly accepted')
except grpc.RpcError as e:
assert e.code() == code, (e.code(), e.details())
print('REJECT', code.name, e.details(), flush=True)
load('')
reject(body(f['red']), grpc.StatusCode.UNIMPLEMENTED)
for k in ['dimension', 'pixels', 'jpeg_dimension', 'jpeg_pixels']:
reject(body(f[k]), grpc.StatusCode.RESOURCE_EXHAUSTED)
for k in ['bomb', 'truncated', 'bad_crc', 'bad_adler', 'jpeg_missing_eoi',
'jpeg_truncated_scan', 'jpeg_appended_eoi', 'jpeg_embedded_missing_eoi']:
reject(body(f[k]), grpc.StatusCode.INVALID_ARGUMENT)
reject(body('data:image/png;base64,AB=='), grpc.StatusCode.INVALID_ARGUMENT)
reject(body('https://example.invalid/a.png'), grpc.StatusCode.INVALID_ARGUMENT)
reject({'state': 'x' * (64 << 10)}, grpc.StatusCode.RESOURCE_EXHAUSTED)
reject({'state': 'x' * (16 << 20)}, grpc.StatusCode.RESOURCE_EXHAUSTED)
load(a.projector)
results = {}
for color in ['red', 'blue']:
r = json.loads(score(body(f[color])).response_json)
print('IMAGE', color, json.dumps(r), flush=True)
assert r['usage']['input_tokens'] > 0 and r['usage']['output_tokens'] == 0, r
probs = r['answers']['color']['probabilities']
assert all(math.isfinite(v) for v in probs.values()) and abs(sum(probs.values())-1)<1e-4, r
results[color] = probs
assert results['red']['red'] > results['blue']['red'], results
assert results['blue']['blue'] > results['red']['blue'], results
print('CONTRASTING IMAGE EXECUTION PASS', flush=True)
@@ -0,0 +1,49 @@
// SPDX-License-Identifier: MIT
#define STB_IMAGE_IMPLEMENTATION
#include "stb/stb_image.h"
#undef STB_IMAGE_IMPLEMENTATION
#include "decision_images.h"
#include <cassert>
#include <iostream>
#include <fstream>
using namespace localai_decision;
template<class F> void rejects(F f, bool large=false) {
try { f(); assert(false); } catch (const image_error & e) { assert(e.too_large == large); }
}
int main(int argc, char ** argv) {
assert(argc==2);
std::ifstream input(argv[1]);
json fixtures; input >> fixtures;
assert(!supports_images(true, false));
assert(!supports_images(false, true));
assert(supports_images(true, true));
auto check=[](json j) { return validate(j, j.dump().size()); };
check(json{{"state", "text"}});
for (auto images : {json(), json::array()}) {
rejects([&]{check(json{{"state",std::string(text_bytes, 'x')},{"images",images}});},true);
}
rejects([&]{check(json{{"state",std::string(body_bytes, 'x')}});},true);
rejects([&]{check(json{{"state",std::string(text_bytes, 'x')}});},true);
rejects([&]{check(json{{"images", {"https://invalid/image.png"}}});});
rejects([&]{check(json{{"images", {"data:image/png;base64,AAAA\n"}}});});
rejects([&]{check(json{{"images", {"data:image/png;base64,AB=="}}});});
rejects([&]{check(json{{"images", std::vector<std::string>(9,"x")}});},true);
rejects([&]{check(json{{"images", {std::string(encoded_bytes+1,'x')}}});},true);
rejects([&]{check(json{{"images", {"data:image/png;base64,"+std::string(12*1024*1024-24,'A')}}});},true);
for (auto key : {"dimension", "pixels", "jpeg_dimension", "jpeg_pixels"}) rejects([&]{check(json{{"images",{fixtures[key]}}});},true);
for (auto key : {"bomb", "truncated", "bad_crc", "bad_adler", "jpeg_missing_eoi", "jpeg_truncated_scan", "jpeg_appended_eoi", "jpeg_embedded_missing_eoi"}) rejects([&]{check(json{{"images",{fixtures[key]}}});});
// Aggregate pixels reject even when each individual image fits.
rejects([&]{check(json{{"images",{fixtures["aggregate"],fixtures["aggregate"]}}});},true);
for (auto key : {"red", "blue", "jpeg", "jpeg_progressive", "jpeg_embedded_marker"}) assert(check(json{{"images",{fixtures[key]}}})==1);
// Valid one-pixel PNG, and MIME mismatch.
std::string png="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNk+A8AAQUBAScY42YAAAAASUVORK5CYII=";
assert(check(json{{"images",{png}}})==1);
auto jpeg=png; jpeg.replace(5,9,"image/jpeg");
rejects([&]{check(json{{"images",{jpeg}}});});
json chat={{"state",{{"messages",json::array({{{"content",json::array({{{"type","image"},{"source",{{"type","base64"},{"media_type","image/png"},{"data",png.substr(22)}}}}})}}})}}}};
assert(validate(chat,chat.dump().size())==1);
assert(chat["state"]["messages"][0]["content"][0]["type"]=="image_url");
// Domain state is not chat content.
assert(check(json{{"state",{{"image_url","https://invalid"}}}})==0);
std::cout << "decision image safety PASS\n";
}
@@ -0,0 +1,68 @@
# SPDX-License-Identifier: MIT
# Requires generated backend_pb2{,_grpc}.py on PYTHONPATH and a running backend.
import argparse
import json, grpc, os, concurrent.futures
import tempfile
from pathlib import Path
parser = argparse.ArgumentParser()
parser.add_argument('--address', default='127.0.0.1:50051')
parser.add_argument('--model', required=True)
args = parser.parse_args()
import backend_pb2 as pb
import backend_pb2_grpc as rpc
channel=grpc.insecure_channel(args.address)
grpc.channel_ready_future(channel).result(timeout=10)
s=rpc.BackendStub(channel)
r=s.LoadModel(pb.ModelOptions(ModelFile=os.path.abspath(args.model),ContextSize=1024,NBatch=512,Threads=2,NGPULayers=0,Options=['parallel:2']),timeout=120)
assert r.success,r
print('LOAD PASS',flush=True)
body={'model':'tinylaya','state':'I was charged twice for my order last week and nobody has replied.','questions':{
'route':{'type':'choice','instructions':'Which team should handle this?','criteria':{'billing':'payments and refunds','shipping':None,'technical':None}},
'urgency':{'type':'score','instructions':'How urgent is this?','criteria':['can wait','this week','today','right now']},
'angry':{'type':'noul','instructions':'Is the customer angry?'}}}
def run():
r=json.loads(s.Score(pb.ScoreRequest(question_type='systemone',prompt=json.dumps(body)),timeout=60).response_json)
assert r['usage']['input_tokens']>0 and r['usage']['output_tokens']==0,r
a=r['answers']; assert set(a)==set(body['questions']),r
assert 0<=a['angry']['noul']<=1,r
for k in ['route','urgency']: assert abs(sum(a[k]['probabilities'].values())-1)<1e-4,r
assert a['route']['choice']==max(a['route']['probabilities'],key=a['route']['probabilities'].get),r
assert abs(a['urgency']['score']-sum(int(k)*v for k,v in a['urgency']['probabilities'].items()))<1e-4,r
return r
print('MULTIQUESTION',json.dumps(run()),flush=True)
with concurrent.futures.ThreadPoolExecutor(max_workers=2) as pool: list(pool.map(lambda _:run(),range(4)))
print('CONCURRENCY PASS',flush=True)
for payload in ['{',json.dumps({'state':'x','questions':{}})]:
try:s.Score(pb.ScoreRequest(question_type='systemone',prompt=payload),timeout=30);raise AssertionError('expected invalid')
except grpc.RpcError as e:assert e.code()==grpc.StatusCode.INVALID_ARGUMENT,e
print('INVALID PASS',flush=True)
f=s.Score.future(pb.ScoreRequest(question_type='systemone',prompt=json.dumps(body)),timeout=60);f.cancel()
try:f.result();raise AssertionError('expected cancelled')
except grpc.FutureCancelledError:pass
run();print('CANCEL AND RECOVERY PASS',flush=True)
try:s.Score(pb.ScoreRequest(prompt='hello',candidates=['world']),timeout=30);raise AssertionError('score unexpectedly enabled')
except grpc.RpcError as e:assert e.code()==grpc.StatusCode.FAILED_PRECONDITION,e
print('PLAIN SCORE DISABLED GUARD PASS',flush=True)
r=s.LoadModel(pb.ModelOptions(ModelFile=os.path.abspath(args.model),ContextSize=1024,NBatch=512,Threads=2,EnableScore=True,Options=['parallel:2']),timeout=120)
assert r.success,r
r=s.Score(pb.ScoreRequest(prompt='Hello',candidates=[' world',' there']),timeout=30)
assert len(r.candidates)==2 and all(c.num_tokens>0 for c in r.candidates),r
assert all(__import__('math').isfinite(c.log_prob) for c in r.candidates),r
print('PLAIN SCORE ENABLED PASS',flush=True)
# This fixture is an encoder: keep embeddings enabled when removing decision
# metadata, otherwise upstream warmup can exceed the one-slot output budget.
data = Path(args.model).read_bytes()
assert data.count(b'.decision.type') == 1, 'expected the tinylaya test fixture'
with tempfile.TemporaryDirectory() as directory:
model = Path(directory) / 'no-decision.gguf'
model.write_bytes(data.replace(b'.decision.type', b'.disabled.type'))
r=s.LoadModel(pb.ModelOptions(ModelFile=str(model),ContextSize=1024,NBatch=512,Threads=2,Embeddings=True),timeout=120)
assert r.success,r
try:
s.Score(pb.ScoreRequest(question_type='systemone',prompt=json.dumps(body)),timeout=30)
raise AssertionError('missing decision metadata accepted')
except grpc.RpcError as e:
assert e.code()==grpc.StatusCode.UNIMPLEMENTED,e
print('MISSING DECISION METADATA PASS',flush=True)
@@ -0,0 +1,66 @@
# SPDX-License-Identifier: MIT
"""Generate small compressed fixtures, including hostile IHDR/IDAT combinations."""
import base64
import json
import io
from PIL import Image
import struct
import sys
import zlib
def png(w, h, pixels):
def chunk(kind, data):
return struct.pack('>I', len(data)) + kind + data + struct.pack('>I', zlib.crc32(kind + data))
return b'\x89PNG\r\n\x1a\n' + chunk(b'IHDR', struct.pack('>IIBBBBB', w, h, 8, 2, 0, 0, 0)) + chunk(b'IDAT', zlib.compress(pixels)) + chunk(b'IEND', b'')
def url(raw):
return 'data:image/png;base64,' + base64.b64encode(raw).decode()
fixtures = {
'aggregate': url(png(3000, 3000, (b'\0' * 9001)*3000)),
'dimension': url(png(4097, 1, b'\0'*12292)),
'pixels': url(png(4096, 4096, b'\0')),
'bomb': url(png(1, 1, b'\0'*1000000)),
'truncated': url(png(1, 1, b'\0'*4)[:-15]),
'red': url(png(64, 64, (b'\0'+b'\xff\0\0'*64)*64)),
'blue': url(png(64, 64, (b'\0'+b'\0\0\xff'*64)*64)),
}
bad = bytearray(png(1, 1, b'\0'*4))
bad[29] ^= 1
fixtures['bad_crc'] = url(bad)
# CRC-valid IDAT with an invalid zlib Adler-32 checksum.
bad = bytearray(base64.b64decode(fixtures['red'].split(',')[1]))
pos = bad.index(b'IDAT')
n = struct.unpack('>I', bad[pos-4:pos])[0]
bad[pos+4+n-1] ^= 1
bad[pos+4+n:pos+8+n] = struct.pack('>I', zlib.crc32(bad[pos:pos+4+n]))
try:
zlib.decompress(bad[pos+4:pos+4+n])
raise AssertionError('invalid Adler-32 accepted')
except zlib.error:
pass
fixtures['bad_adler'] = url(bad)
def jpeg(w, h, progressive=False):
out = io.BytesIO()
Image.new('RGB', (w, h), 'red').save(out, format='JPEG', progressive=progressive)
return out.getvalue()
def jpg_url(raw):
return 'data:image/jpeg;base64,' + base64.b64encode(raw).decode()
jpg = jpeg(64, 64)
for key, raw in {
'jpeg': jpg,
'jpeg_progressive': jpeg(64, 64, True),
# EOI inside a comment is data, not an end marker.
'jpeg_embedded_marker': jpg[:2] + b'\xff\xfe\x00\x04\xff\xd9' + jpg[2:],
'jpeg_missing_eoi': jpg[:-2],
'jpeg_truncated_scan': jpg[:-30],
'jpeg_appended_eoi': jpg[:-30] + b'\xff\xd9',
'jpeg_embedded_missing_eoi': jpg[:2] + b'\xff\xfe\x00\x04\xff\xd9' + jpg[2:-2],
'jpeg_dimension': jpeg(4097, 1),
'jpeg_pixels': jpeg(4096, 4096),
}.items():
fixtures[key] = jpg_url(raw)
json.dump(fixtures, open(sys.argv[1], 'w'))
+22
View File
@@ -0,0 +1,22 @@
#!/usr/bin/env bash
# SPDX-License-Identifier: MIT
set -euo pipefail
root=$(git rev-parse --show-toplevel)
b="$root/backend/cpp/llama-cpp"
out="$b/llama.cpp/build-image-tests"
mkdir -p "$out"
${CXX:-g++} -std=c++17 -Wall -Wextra -I"$b" -I"$b/llama.cpp/vendor" "$b/tests/decision-images.cpp" -lz -ljpeg -o "$out/decision-images"
python3 "$b/tests/image-fixtures.py" "$out/fixtures.json"
"$out/decision-images" "$out/fixtures.json"
# Keep the native boundary in lockstep with canonical Go limits.
python3 - "$root" <<'PY'
import pathlib, re, sys
root=pathlib.Path(sys.argv[1])
go=(root/'core/systemone/images.go').read_text()
cpp=(root/'backend/cpp/llama-cpp/decision_images.h').read_text()
for g,c in [('MaxImages','max_images'),('MaxImageDecodedBytes','decoded_bytes'),('MaxImageEncodedBytes','encoded_bytes'),('MaxImageBodyBytes','body_bytes'),('MaxImageDimension','max_dimension'),('MaxImagePixels','max_pixels'),('MaxResponseBytes','text_bytes')]:
gv=re.search(r'\b'+g+r'\s*=\s*([^\n]+)',go)[1]
cv=re.search(r'\b'+c+r'\s*=\s*([^;]+)',cpp)[1]
assert eval(gv)==eval(cv),(g,c)
print('Go/native limit parity PASS')
PY
@@ -0,0 +1,38 @@
# SPDX-License-Identifier: MIT
"""Exercise the production CMake decoder dependency block, including old forks."""
import pathlib
import subprocess
import tempfile
backend = pathlib.Path(__file__).resolve().parents[1]
cmake = (backend / 'CMakeLists.txt').read_text()
start = cmake.index('if(EXISTS "${CMAKE_CURRENT_SOURCE_DIR}/../server/server-decision.cpp")')
block = cmake[start:cmake.index('endif()', start) + len('endif()')]
with tempfile.TemporaryDirectory() as tmp:
root = pathlib.Path(tmp)
source = root / 'grpc-server'
source.mkdir()
(root / 'server').mkdir()
(source / 'main.cpp').write_text('int main() {}\n')
(source / 'CMakeLists.txt').write_text('''cmake_minimum_required(VERSION 3.15)
project(decoder_wiring LANGUAGES CXX)
set(TARGET grpc-server)
add_executable(${TARGET} main.cpp)
''' + block + '''
get_target_property(libs ${TARGET} LINK_LIBRARIES)
if(EXPECT_DECODERS)
if(NOT "${libs}" STREQUAL "ZLIB::ZLIB;JPEG::JPEG")
message(FATAL_ERROR "Missing decoder links: ${libs}")
endif()
elseif(libs)
message(FATAL_ERROR "Old fork acquired decoder dependencies: ${libs}")
endif()
''')
subprocess.run(['cmake', '-S', str(source), '-B', str(root / 'fork'),
'-DCMAKE_DISABLE_FIND_PACKAGE_ZLIB=TRUE',
'-DCMAKE_DISABLE_FIND_PACKAGE_JPEG=TRUE'], check=True)
(root / 'server/server-decision.cpp').touch()
subprocess.run(['cmake', '-S', str(source), '-B', str(root / 'native'),
'-DEXPECT_DECODERS=ON'], check=True)
subprocess.run(['cmake', '--build', str(root / 'native')], check=True)
print('Production CMake decoder links and old-fork guard PASS')