mirror of
https://github.com/mudler/LocalAI.git
synced 2026-10-05 12:34:43 -04:00
* fix(schema): preserve SystemOne image inputs Assisted-by: OpenAI * test(schema): follow Ginkgo conventions for decision inputs Assisted-by: OpenAI * feat(llama-cpp): dispatch native decisions through Score Upgrade the stock dependency and reconcile Score/TTS patches. Reuse native decision parsing, tasks, formatting and response-reader cleanup; preserve ordinary scoring admission and guard older dependencies. Assisted-by: OpenAI * refactor(systemone): share request and model validation Assisted-by: OpenAI:gpt-5 * fix(systemone): preserve HTTP wire-byte validation limit Keep structural validation separate from the serialized internal request bound so HTML escaping cannot reject valid HTTP payloads. Assisted-by: OpenAI:gpt-5 * feat(systemone): bound images and account native decisions Preserve public wire limits independently from router serialization. Reject unsupported NER images, map native request/capability errors, and stamp explicit usage once. Advertise decisions for stock llama-cpp. Assisted-by: OpenAI * fix(systemone): record usage on registered native route Exercise real registration and billing with a mock native backend. Reject empty native responses, malformed image URLs, trailing JSON, and wire overflow including whitespace. Assisted-by: OpenAI * feat(router): add lazy native decision transport Bind named models through internal ModelSystemOne calls with shared validation and bounded abandoned operations. Remove request and echoed-error contents from decision traces. Assisted-by: OpenAI:gpt-5 * feat(router): classify overlapping policies with native decisions Ask independent noul questions, validate probabilities and preserve first-superset routing. Wire the central factory with config-sensitive invalidation and cancellation-safe resolution. Document native framing and bounded operation limits. Assisted-by: OpenAI:gpt-5 * feat(gallery): add pinned Julia-1 native decision model Add a separate text-only llama-cpp Q8 entry with pinned Apache-2.0 source provenance and checksum. Installed using the gallery installer and exercised choice, score and noul on CPU. Assisted-by: OpenAI * test(router): verify native decisions through central factory Add an opt-in real-model Ginkgo integration covering the native Go loader and C++ transport, token usage, independent overlapping labels, and candidate selection. Document owned-server execution and the intentionally non-quality threshold. Assisted-by: Codex:gpt-5 * fix(llama-cpp): align upstream pin and preserve decision signatures Advance to bed0a856 without losing the automated upstream bump. Detect full-request fill_task support at compile time and forward every question for Nimble framing while retaining the earlier native signature. Preserve reconciled SCORE/TTS patches; add standalone compatibility coverage. Assisted-by: Codex:gpt-5 * feat(gallery): add native decision family defaults Pin Laya, Kev-4B, lev, OpenJev and Nimble artifacts. Verify Laya/Kev/lev gallery installs and CPU contracts on both native pins; clearly mark OpenJev/Nimble runtime validation pending and their noncommercial licenses. Assisted-by: OpenAI * docs(decisions): clarify integrated Nimble prerequisite Record the exact combined backend pin while retaining pending OpenJev and Nimble installation/runtime validation status. Assisted-by: Codex:gpt-5 * fix(gallery): indent native decision model sequences Match repository yamllint indentation for Laya, Kev, lev and OpenJev list fields. Parsed gallery data is unchanged; reproduce CI gallery lint failure before the whitespace-only fix and pass the same command afterward. Assisted-by: Codex:gpt-5 * docs(decisions): record OpenJev and Nimble CPU validation Record gallery installation, checksum/metadata verification and multiquestion native smoke results on bed0a856. Retain noncommercial and text-only limitations without accuracy or deterministic-output claims. Assisted-by: OpenAI * fix(ui): expose native Decisions router classifiers Select classifier models using metadata-driven capability routing, retain tuned thresholds, and validate native decision selections before saving. Cover both native backends and create/save/reopen in the real React editor. Assisted-by: Codex:gpt-5 * fix(router): exclude aliases from native decision discovery Check the originally named config before advertising native Decisions eligibility. Retain target capability inheritance for ordinary generation aliases. Exercise the actual capabilities endpoint with native models on both backends, aliases, and disabled models. Assisted-by: Codex:gpt-5 * feat(systemone): share bounded multimodal input validation Preserve text wire limits while admitting bounded PNG/JPEG decision input. Share collection and header validation across internal and public callers and keep the native runner response budget independent. Assisted-by: OpenAI:API-assistant * fix(systemone): bound admission lifetimes and validate complete images Retain shared admission leases through actual work completion, including abandoned internal operations. Decode bounded image pixels, cap public native responses before usage stamping, and preserve oversized malformed text status precedence. Assisted-by: OpenAI:API-assistant * fix(router): classify images before media fetching Preserve ordered structured probes for native decisions. Defer OpenAI media preparation until routing selects the served model, so rejected decision URLs cannot trigger downloads before shared validation. Guard direct image collection with context-aware shared admission. Keep text classifiers and embedding caches from discarding image input. Retain fail-closed classifier configuration and runtime fallback policy. Add middleware, typed-content, admission, cancellation and cache tests. Assisted-by: OpenAI:API-assistant * fix(router): bound extraction before serialization Check probe budgets before copying text or marshaling message state. Count JSON escaping so oversized internal inputs fail before allocation. Preserve typed Anthropic blocks through selected-model conversion and fallback. Keep retry coverage in Ginkgo without global test registration. Assisted-by: OpenAI * fix(router): bound supported probe serialization Arbitrary structs can bypass the probe budget through pointer marshalers, string tags, and promoted fields. Accept concrete chat schema types and plain JSON values instead of emulating arbitrary struct serialization. Budget escaped direct prompts before marshaling so raw length cannot hide serialized expansion. Preserve runtime fallback and reject oversized input before invoking the decision runner. Add Ginkgo allocation, boundary, and marshaler invocation regressions. Six-package tests, three-package race tests, and full-T2 delta lint pass. Assisted-by: OpenAI:GPT-5 golangci-lint * feat(decisions): enable bounded OpenJev images Validate native decision images before permissive media parsing and pixel allocation. Require both decision image support and a vision projector; missing or audio-only projectors cannot silently become text decisions. Pin the OpenJev Q8 projector and document its license and disk footprint. Add native safety tests, canonical limit parity, gallery and load-option checks, and a reproducible CPU direct-RPC contrasting-image smoke. Assisted-by: OpenAI:GPT-5 * fix(decisions): reject incomplete image streams stb accepts corrupt PNG Adler checksums and truncated JPEG scans. Use bounded zlib validation and strict libjpeg decoding before parsing. Keep dimension and aggregate pixel checks ahead of decoder allocations. Wire decoder dependencies into native builds and runtime packaging. Add regressions for appended EOI and embedded marker bypasses. Assisted-by: OpenAI:GPT-5 * fix(ci): gate native decision image validation Run the decoder security tests outside the stdlib-only native suite. Fetch vendor headers at the backend pin and provision decoder dependencies. Gate Go limit parity and production CMake wiring without model downloads. Assisted-by: OpenAI:GPT-5 * test(decisions): cover multimodal public API paths Exercise shared image contracts through the registered HTTP routes and external mock backend. Add opt-in cached gallery installation and real OpenJev image decisions through SystemOne and both routing APIs. Assisted-by: Codex:gpt-5 * test(decisions): assert isolation and cache bypass Observe external RPC calls and compare complete classifier history. Winner-only and cache-miss checks could hide dropped history or cache use. Give real inference its own application and model directory so shared backend mappings and loaded processes cannot affect mixed suite order. Assisted-by: OpenAI:ChatGPT * test(decisions): isolate fixture globals Disable optional global services in the isolated HTTP fixture and register cleanup before setup assertions. Verify meter provider identity survives fixture creation and destruction. Snapshot observed usage before assertions so failures cannot retain the mutex. Require a successful usage stamp before checking error responses. Assisted-by: Codex:gpt-5 golangci-lint * fix(application): honor optional telemetry controls Skip failover gauge registration when metrics are disabled. Register against the application meter rather than looking up the global provider. Allow embedders to retain the bounded routing log without billing stats. Keep the existing default when stats are disabled. The isolated HTTP fixture uses this option without losing its native router assertions. Assisted-by: Codex:gpt-5 golangci-lint --------- Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
211 lines
10 KiB
C++
211 lines
10 KiB
C++
// SPDX-License-Identifier: MIT
|
|
#pragma once
|
|
#include <nlohmann/json.hpp>
|
|
#include <stdexcept>
|
|
#include <string>
|
|
#include <vector>
|
|
#include <cstdint>
|
|
#include <cstring>
|
|
#include <memory>
|
|
#include <csetjmp>
|
|
#include <cstdio>
|
|
#include <jpeglib.h>
|
|
#include <zlib.h>
|
|
#include "stb/stb_image.h"
|
|
|
|
// Decision-only limits, mirrored from core/systemone/images.go. The verification
|
|
// script checks parity; these must not change ordinary chat or fork backends.
|
|
namespace localai_decision {
|
|
using json = nlohmann::ordered_json;
|
|
constexpr size_t max_images = 8;
|
|
constexpr size_t decoded_bytes = 8 << 20;
|
|
constexpr size_t encoded_bytes = 12 << 20;
|
|
constexpr size_t body_bytes = 16 << 20;
|
|
constexpr size_t text_bytes = 64 << 10;
|
|
constexpr size_t max_dimension = 4096;
|
|
constexpr size_t max_pixels = 16000000;
|
|
struct image_error : std::invalid_argument {
|
|
bool too_large;
|
|
image_error(const char * message, bool large=false) : std::invalid_argument(message), too_large(large) {}
|
|
};
|
|
inline bool supports_images(bool decision, bool vision) { return decision && vision; }
|
|
inline void require(bool ok, const char * message, bool large=false) {
|
|
if (!ok) throw image_error(message, large);
|
|
}
|
|
// libjpeg normally repairs premature EOF and incomplete entropy scans. Treat
|
|
// warnings as failures as well as fatal errors; an appended EOI cannot hide a
|
|
// short scan. Keep all mutable decoder state on the heap across longjmp.
|
|
struct jpeg_validator {
|
|
jpeg_decompress_struct decoder{};
|
|
jpeg_error_mgr errors{};
|
|
std::jmp_buf jump;
|
|
};
|
|
inline void jpeg_failure(j_common_ptr decoder) {
|
|
auto * state = static_cast<jpeg_validator *>(decoder->client_data);
|
|
std::longjmp(state->jump, 1);
|
|
}
|
|
inline void jpeg_message(j_common_ptr decoder, int level) {
|
|
if (level < 0) jpeg_failure(decoder);
|
|
}
|
|
inline void validate_jpeg(const std::vector<unsigned char> & raw, int width, int height) {
|
|
auto state = std::make_unique<jpeg_validator>();
|
|
auto * decoder = &state->decoder;
|
|
decoder->err = jpeg_std_error(&state->errors);
|
|
state->errors.error_exit = jpeg_failure;
|
|
state->errors.emit_message = jpeg_message;
|
|
decoder->client_data = state.get();
|
|
if (setjmp(state->jump)) {
|
|
jpeg_destroy_decompress(decoder);
|
|
throw image_error("invalid or incomplete JPEG");
|
|
}
|
|
jpeg_create_decompress(decoder);
|
|
jpeg_mem_src(decoder, raw.data(), raw.size());
|
|
jpeg_read_header(decoder, TRUE);
|
|
// The dimension/aggregate checks in validate_url precede all pixel or
|
|
// coefficient allocations. Verify both decoders saw the same dimensions.
|
|
if (decoder->image_width != unsigned(width) || decoder->image_height != unsigned(height)) {
|
|
jpeg_destroy_decompress(decoder);
|
|
throw image_error("inconsistent JPEG dimensions");
|
|
}
|
|
jpeg_start_decompress(decoder);
|
|
auto row = (*decoder->mem->alloc_sarray)(reinterpret_cast<j_common_ptr>(decoder),
|
|
JPOOL_IMAGE, decoder->output_width * decoder->output_components, 1);
|
|
while (decoder->output_scanline < decoder->output_height) {
|
|
jpeg_read_scanlines(decoder, row, 1);
|
|
}
|
|
jpeg_finish_decompress(decoder);
|
|
jpeg_destroy_decompress(decoder);
|
|
}
|
|
inline int digit(unsigned char c) {
|
|
if (c >= 'A' && c <= 'Z') return c-'A';
|
|
if (c >= 'a' && c <= 'z') return c-'a'+26;
|
|
if (c >= '0' && c <= '9') return c-'0'+52;
|
|
return c=='+' ? 62 : c=='/' ? 63 : -1;
|
|
}
|
|
inline uint32_t be32(const unsigned char * p) {
|
|
return uint32_t(p[0])<<24 | uint32_t(p[1])<<16 | uint32_t(p[2])<<8 | p[3];
|
|
}
|
|
inline uint32_t png_crc(const unsigned char * data, size_t size) {
|
|
uint32_t crc = 0xffffffffu;
|
|
for (size_t i = 0; i < size; ++i) {
|
|
crc ^= data[i];
|
|
for (int bit = 0; bit < 8; ++bit) crc = (crc >> 1) ^ (0xedb88320u & (0u - (crc & 1)));
|
|
}
|
|
return crc ^ 0xffffffffu;
|
|
}
|
|
inline void validate_url(const std::string & url, size_t & decoded, size_t & pixels) {
|
|
const auto comma = url.find(',');
|
|
const auto header = url.substr(0, comma);
|
|
bool png = header == "data:image/png;base64";
|
|
require(comma != std::string::npos && (png || header == "data:image/jpeg;base64"), "images must be PNG/JPEG base64 data URLs");
|
|
size_t n = url.size()-comma-1;
|
|
require(n > 0 && n%4 == 0, "invalid base64 length");
|
|
const char * data = url.data()+comma+1;
|
|
size_t pad = (data[n-1]=='=') + (data[n-2]=='=');
|
|
size_t size = n/4*3-pad;
|
|
require(size <= decoded_bytes-decoded, "decoded image aggregate exceeds limit", true);
|
|
std::vector<unsigned char> raw;
|
|
raw.reserve(size);
|
|
for (size_t i=0; i<n; i+=4) {
|
|
int a=digit(data[i]), b=digit(data[i+1]);
|
|
int c=data[i+2]=='=' ? 0 : digit(data[i+2]);
|
|
int d=data[i+3]=='=' ? 0 : digit(data[i+3]);
|
|
require(a>=0 && b>=0 && c>=0 && d>=0, "invalid base64 character");
|
|
require((data[i+2]!='=' && data[i+3]!='=') || i+4==n, "invalid base64 padding");
|
|
require(data[i+2]!='=' || (data[i+3]=='=' && (b&15)==0), "invalid base64 padding bits");
|
|
require(data[i+3]!='=' || data[i+2]=='=' || (c&3)==0, "invalid base64 padding bits");
|
|
raw.push_back((a<<2)|(b>>4));
|
|
if (data[i+2]!='=') raw.push_back((b<<4)|(c>>2));
|
|
if (data[i+3]!='=') raw.push_back((c<<6)|d);
|
|
}
|
|
decoded += raw.size();
|
|
require(png ? raw.size()>=24 && std::memcmp(raw.data(),"\x89PNG\r\n\x1a\n",8)==0
|
|
: raw.size()>=3 && raw[0]==255 && raw[1]==216 && raw[2]==255, "image MIME mismatch");
|
|
int w=0,h=0,c=0;
|
|
require(stbi_info_from_memory(raw.data(),raw.size(),&w,&h,&c)!=0 && w>0 && h>0, "invalid image header");
|
|
require(size_t(w)<=max_dimension && size_t(h)<=max_dimension, "image dimensions exceed limit", true);
|
|
size_t count=size_t(w)*size_t(h);
|
|
require(count<=max_pixels-pixels, "image pixel aggregate exceeds limit", true);
|
|
pixels += count;
|
|
if (png) {
|
|
// stb's PNG inflater grows independently of IHDR. Validate IDAT with a
|
|
// fixed output buffer first, preventing small-header decompression bombs.
|
|
// 16-bit RGBA plus Adam7 row filters fit this conservative pixel bound.
|
|
std::vector<unsigned char> idat;
|
|
size_t pos=8;
|
|
bool end=false;
|
|
while (pos+12<=raw.size()) {
|
|
size_t len=be32(raw.data()+pos);
|
|
require(len<=raw.size()-pos-12, "truncated PNG chunk");
|
|
require(png_crc(raw.data()+pos+4,len+4)==be32(raw.data()+pos+8+len), "invalid PNG checksum");
|
|
if (std::memcmp(raw.data()+pos+4,"IDAT",4)==0)
|
|
idat.insert(idat.end(),raw.begin()+pos+8,raw.begin()+pos+8+len);
|
|
if (std::memcmp(raw.data()+pos+4,"IEND",4)==0) { end=true; break; }
|
|
pos+=len+12;
|
|
}
|
|
require(end && !idat.empty(), "incomplete PNG");
|
|
std::vector<char> inflated(9*count+8*size_t(h)+1024);
|
|
z_stream stream{};
|
|
stream.next_in = idat.data();
|
|
stream.avail_in = static_cast<uInt>(idat.size());
|
|
stream.next_out = reinterpret_cast<Bytef *>(inflated.data());
|
|
stream.avail_out = static_cast<uInt>(inflated.size());
|
|
require(inflateInit(&stream) == Z_OK, "PNG inflater initialization failed");
|
|
int result = inflate(&stream, Z_FINISH);
|
|
bool complete = result == Z_STREAM_END && stream.avail_in == 0;
|
|
inflateEnd(&stream);
|
|
require(complete, "invalid or oversized PNG decompression");
|
|
} else {
|
|
validate_jpeg(raw, w, h);
|
|
}
|
|
auto * image=stbi_load_from_memory(raw.data(),raw.size(),&w,&h,&c,3);
|
|
require(image!=nullptr, "invalid image pixels");
|
|
stbi_image_free(image);
|
|
}
|
|
// Normalize only actual chat content parts, as in the canonical Go collector.
|
|
// Upstream parse_state understands image_url but not Anthropic source objects.
|
|
inline size_t validate(json & body, size_t wire_size) {
|
|
require(wire_size<=body_bytes, "decision request exceeds limit", true);
|
|
std::vector<const std::string *> urls;
|
|
auto add=[&](const json & value) {
|
|
require(value.is_string(), "image URL must be a string");
|
|
require(urls.size()<max_images, "too many decision images", true);
|
|
urls.push_back(&value.get_ref<const std::string &>());
|
|
};
|
|
if (body.contains("images") && !body["images"].is_null()) {
|
|
require(body["images"].is_array(), "images must be an array");
|
|
for (const auto & url : body["images"]) add(url);
|
|
}
|
|
auto state=body.find("state");
|
|
if (state!=body.end()) {
|
|
json * messages=&*state;
|
|
if (state->is_object() && state->contains("messages")) messages=&(*state)["messages"];
|
|
if (messages->is_array()) for (auto & msg : *messages) {
|
|
if (!msg.is_object() || !msg.contains("content") || !msg["content"].is_array()) continue;
|
|
for (auto & part : msg["content"]) {
|
|
if (!part.is_object() || !part.contains("type")) continue;
|
|
if (part["type"]=="image") {
|
|
require(part.contains("source") && part["source"].is_object(), "invalid image source");
|
|
auto & s=part["source"];
|
|
require(s.value("type",std::string())=="base64" && s.contains("media_type") && s["media_type"].is_string() && s.contains("data") && s["data"].is_string(), "invalid image source");
|
|
require(s["data"].get_ref<const std::string &>().size()<=encoded_bytes && s["media_type"].get_ref<const std::string &>().size()<=32, "image source exceeds limit",true);
|
|
std::string url="data:"+s["media_type"].get<std::string>()+";base64,"+s["data"].get<std::string>();
|
|
part=json{{"type","image_url"},{"image_url",{{"url",url}}}};
|
|
}
|
|
if (part["type"]=="image_url") {
|
|
require(part.contains("image_url"), "missing image URL");
|
|
auto & u=part["image_url"];
|
|
if (u.is_object()) { require(u.contains("url"), "missing image URL"); add(u["url"]); }
|
|
else add(u);
|
|
}
|
|
}
|
|
}
|
|
}
|
|
require(!urls.empty() || wire_size<=text_bytes, "text decision request exceeds limit",true);
|
|
size_t encoded=0,decoded=0,pixels=0;
|
|
for (const auto * u : urls) { require(u->size()<=encoded_bytes-encoded,"encoded image aggregate exceeds limit",true); encoded+=u->size(); }
|
|
for (const auto * u : urls) validate_url(*u,decoded,pixels);
|
|
return urls.size();
|
|
}
|
|
}
|