Compare commits

...

60 Commits

Author SHA1 Message Date
Ettore Di Giacinto
2a34785810 fix(nemo-speech-cpp): restore std::binary_function for MeCab on libc++
NEMO_SPEECH_TTS_WITH_JA=ON compiles Open JTalk's bundled MeCab, and
mecab/src/dictionary.cpp derives a comparator from std::binary_function,
which C++17 removed. libstdc++ still ships it as deprecated-but-present
under -std=gnu++17, so Linux never notices. libc++ compiles it out and
the macOS arm64 build dies with "no template named 'binary_function' in
namespace 'std'".

This is ours, not an upstream regression: upstream defaults both
NEMO_SPEECH_TTS_WITH_JA and NEMO_SPEECH_TTS_WITH_ZH to OFF and the OSS
drop carries no CI at all, so that target is never built there. Upstream
does already carry the equivalent workaround for MSVC's STL
(_HAS_AUTO_PTR_ETC plus /FIfunctional) but has no libc++ branch.

libc++ gates the two templates on
_LIBCPP_ENABLE_CXX17_REMOVED_UNARY_BINARY_FUNCTION, and has since LLVM
16, older than any clang Xcode still ships. The name is the whole
problem: _LIBCPP_ENABLE_CXX17_REMOVED_BINDERS covers bind1st, bind2nd,
ptr_fun and mem_fun and not unary_function or binary_function, and the
umbrella _LIBCPP_ENABLE_CXX17_REMOVED_FEATURES no longer exists in
libcxx at all. A wrong name preprocesses fine and fixes nothing.

Applied through CMAKE_CXX_FLAGS rather than to the one target, because
the tokenizer CMakeLists is upstream's and sources/ is a pinned
checkout. Project-wide is also the safer scope: the macro decides
whether libc++'s internal __binary_function alias resolves to
std::binary_function or to __binary_function_keep_layout_base, a base
class of std::less and friends, so defining it for a subset of
translation units would give those class templates two spellings in one
binary. Both bases are empty and, at C++17, carry identical members, so
the define changes no layout and no ABI.

Darwin only. On Linux the branch is unreachable and the macro is not a
name libstdc++ knows, so it would be inert even if taken; a Linux
configure with the flag forced on puts it on all 23 C++ TUs of
nemo_speech_openjtalk_frontend including dictionary.cpp at -std=gnu++17,
and on none of the 16 C TUs.

Mandarin needs nothing: cppjieba v5.6.7 and limonp have no removed C++17
constructs left (limonp replaced std::not1 and std::bind2nd with
lambdas) and cppjieba's own CI builds macos-14 and macos-latest at C++11
through C++20.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-07 00:54:57 +00:00
Ettore Di Giacinto
00c60a50c2 fix(nemo-speech-cpp): skip the CUDA-only ggml patch series on darwin
The macOS backend build died in patch-ggml:

    scripts/apply-ggml-patches.sh: line 56: mapfile: command not found
    make[1]: *** [patch-ggml] Error 127

mapfile is a bash 4 builtin (and its -d flag needs 4.4). macOS ships bash
3.2.57 as /bin/bash and GitHub's runner images add no newer one, so the
bare `bash` the recipe resolves from PATH cannot run upstream's script.

Rather than hunt for a capable bash that the runner does not have, drop
the step where it does nothing. ggml-patches/ is a CUDA series: every
kernel it adds is under src/ggml-cuda/, and its whole footprint outside
that directory is an op enum plus prototype in include/ggml.h, the
constructor and a name-table entry in src/ggml.c, and two ggml-cpu lines
that make the CUDA-only op report unsupported and abort. Nothing it
touches is compiled into a Metal kernel or changes a CPU one.

The project's own references to patch-only ggml symbols sit behind
NEMO_SPEECH_FUSED_RELPOS_ATTN and NEMO_SPEECH_FASTCONFORMER_CUDA_FUSIONS,
which cmake already forces OFF without GGML_CUDA, or behind
NEMO_SPEECH_GGML_PATCHED itself, which guards a GGML_TENSOR_FLAG_Q8_PLANAR
write that a non-CUDA buffer throws before reaching. So passing
NEMO_SPEECH_GGML_PATCHED=OFF costs the Metal build nothing, and it is
required once the series is skipped: that flag is what stops the ASR
sources referencing a tensor flag stock ggml does not define.

This is upstream's own Metal configuration. Its metal-* and vulkan-*
CMake presets inherit the cpu-* ones, which set NEMO_SPEECH_GGML_PATCHED
to OFF; docker/Dockerfile and scripts/windows/build.ps1 do the same for
their non-CUDA targets. LocalAI's Makefile never passed the flag at all
and so inherited the CUDA default everywhere.

Linux is untouched and keeps applying the series, including its
idempotency and its hard failure on a patch that does not apply. The gate
is the same uname test the WITH_NORM block above already uses, and both
branches keep the order-only clone prerequisite, which on a WITH_NORM=OFF
tree is the only thing that pulls sources/ in.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-07 00:34:24 +00:00
Ettore Di Giacinto
eb4863e2f7 fix(nemo-speech-cpp): audit the gosec unsafe and file-inclusion sites
gosec flags 13 alerts on this backend: one G304 and twelve G103. Each was
checked individually rather than blanket-suppressed, and each annotation
states what makes that particular site safe.

The G304 at audio.go is a false positive. The opened path is
filepath.Join of a directory the function just created with os.MkdirTemp
and a constant basename; the request-controlled path is the input to
AudioToWav and never reaches the open.

The twelve G103 sites are the package's three established shapes, and
every one was verified against them: cstr and pinPtr take the address of
something pinned on the line above and return it one-way (nothing in the
package converts either result back, which is what keeps checkptr out of
it under -race), and each *Create hands C a stack-local POD config whose
uintptr members are cstr allocations or pinPtr addresses held by a pinner
the loader unpins only after the call. The two slice-building sites are
bounded by construction: DiarSegments is handed exactly len(buf) with the
buffer sized under maxDiarSegments and a reported count larger than it
rejected rather than sliced to, and the TTS callback copies out a slice
whose length is the length the runtime declared for that buffer.

Separately, sampleRateOf gets a real fix rather than an annotation.
go-audio reads the WAV header's sample rate from an unsigned 32-bit field
into an int, so a header claiming more than 2^31-1 passed the "> 0" test
and then narrowed to a NEGATIVE rate, which the runtime would take as a
resampling ratio. AudioToWav cannot produce one today, but that is a
property of another package and this function exists precisely because
the rate is read back rather than assumed, so the bound is enforced here
and pinned by a spec.

The four remaining integer narrowings are annotated with the bound that
makes each safe: the WAV payload length is already checked against
maxWAVDataBytes, the speaker count is bounded by maxDiarSegments, and the
two segment ids are the proto's own int32 wire type.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 23:55:15 +00:00
Ettore Di Giacinto
6f301d1cf4 fix(nemo-speech-cpp): map every C status, not just the NMT one
INVALID_ARGUMENT was translated to codes.InvalidArgument at exactly one of
sixteen C call sites. Everywhere else a non-zero status collapsed to
codes.Internal, so the same backend answered an unsupported language pair with
HTTP 400 and an unknown TTS voice, which is the same class of caller mistake
against the same process, with HTTP 500. Status 4 is CANCELLED on the ASR and
TTS surfaces and was reported as a backend failure rather than as the consumer
having stopped listening.

asr.h, tts.h and nmt.h each declare their own status enum and diar.h reuses the
ASR one; the values they share agree, and the single divergence is that NMT
declares no CANCELLED because nemo_speech_nmt_translate has no callback for a
consumer to stop with. That is an absence, not a disagreement, so one table
serves all three. status.go carries it, with the header line numbers and a note
that a pin bump has to recheck it: purego binds by name and the status crosses
as a bare int32, so nothing in the build or the linker can see a drift.

New specs cover the whole enum, unknown values, and one real INVALID_ARGUMENT
per family driven through the shared objects rather than through the Go mapping
asserting against itself.

Also add UsecaseChat to this backend's capability entry, which the docs already
told operators to set for translation models. chat is a gallery filter key and
completion is not, so GET /api/backends/usecases would have greyed the Chat
filter out and hidden a Riva-Translate gallery entry from the one filter that
fits it. The flag gates no endpoint; it makes the model eligible as the default
chat model and puts it in the web UI chat picker, both of which Predict and
PredictStream already serve.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 23:13:42 +00:00
Ettore Di Giacinto
b15c010db6 docs(nemo-speech-cpp): correct the translation limits, the macOS gap and the TTS conversion
Three factual errors found in review, all of them the kind a user would act on.

The translation limits were described backwards. Input longer than the 1024-token
context is rejected, not truncated: translator.cpp throws "nmt: prompt too long
(N tokens) for context 1024", which reaches the caller as a failed request. What
is silently cut is the output, by the max_new_tokens loop at 256. The bullet now
separates the two and says which one fails quietly.

The macOS gap covers TTS text normalization as well. Both directions sit behind
the single NEMO_SPEECH_WITH_NORM flag, which the Makefile forces off on Darwin,
so tn_dir is as inert there as itn_dir. Neither fails the load: both warn and
carry on. pnc_model really is unaffected, since punctuation is compiled in
unconditionally. The tn_dir row in the option reference gained the caveat the
itn_dir row already had.

The TTS conversion procedure produced a model that could not load. It converted
MagpieTTS and stopped, leaving no NanoCodec, which the same page lists as
required; following it gave "no NanoCodec GGUF found next to ...". Both halves
are now there, each with the download that feeds it, so the block runs top to
bottom on a clean machine.

Also: any negative gpu value pins TTS to the CPU, not only -1, and FLAG_CHAT
additionally surfaces the model in the web UI chat picker.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 22:44:19 +00:00
Ettore Di Giacinto
52b5a617d9 docs(nemo-speech-cpp): document the backend and list it in the importer
Adds docs/content/features/nemo-speech-cpp.md, alongside the audio.cpp page
that is its closest sibling, and cross-links it from the speech-to-text,
diarization, text-to-speech, backend-type and compatibility-table pages so the
backend is reachable from every surface that lists its modalities.

The page covers the architecture-to-family table, every option key with a model
YAML per family, the translation prefix directive, the acceleration matrix, and
the four limitations this backend ships with: Linux-only inverse text
normalization, suppressed interim streaming results, the library's default
translation context and generation limits, and the absence of gallery entries.

knownPrefOnlyBackends gains the backend so it appears in the /import-model
dropdown. It stays preference-only and AutoDetect=false: general.architecture
lives inside the GGUF where no remote-repo probe can read it, and a translation
model carries an ordinary LLM architecture with no NeMo-specific marker.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 22:27:04 +00:00
Ettore Di Giacinto
686669111d fix(nemo-speech-cpp): build the CUDA-13 Jetson image the l4t-cuda-13 key needs
The nvidia-l4t-cuda-13 capability pointed at nvidia-l4t-arm64-nemo-speech-cpp,
which is built on nvcr.io/nvidia/l4t-jetpack:r36.4.0 and therefore links ggml
against CUDA 12. A Jetson whose CUDA 13 runtime is present reports that
capability and would have pulled an image with no libcudart.so.12 to dlopen,
failing hard at load. That is worse than omitting the key: with no key
Capability() falls back to "default" and the host gets a working CPU build.

Fixed the way parakeet-cpp and moss-transcribe-cpp already do it, by shipping
the second L4T image rather than dropping the key. Nothing prevents building it
here: those peers use plain ubuntu:24.04 on ubuntu-24.04-arm with the same
Dockerfile.golang as this backend's other rows, and every package in the
nemo-speech-cpp apt gate exists on noble arm64.

Adds the -nvidia-l4t-cuda-13-arm64-nemo-speech-cpp matrix row and its two index
entries, repoints the key on both metas, and rewrites the capability-map comment,
which had the reasoning backwards.

Also adds the documentary inferBackendPath branch, matching all six sibling
*-cpp Go backends. Behaviour is unchanged; the generic golang fallthrough
already resolved this backend correctly.

The previous commit message said "all seven handlers" of the shared gRPC
wrapper. There are eight RPC entry points: seven are guarded by
checkModelIdentity and AudioTranscriptionLive is the unguarded eighth, which
that message already called out separately. Wording only.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 22:08:43 +00:00
Ettore Di Giacinto
568ae16ed5 feat(nemo-speech-cpp): register the backend and give its specs a CI job
Registers nemo-speech-cpp across every surface .agents/adding-backends.md
requires, and adds the CI job its unit suite never had.

backend/index.yaml gets the meta backend (capabilities map, no uri), a
development meta and 12 image entries. No amd and no intel capability keys:
upstream NeMo-Speech.cpp builds ggml with CUDA, Vulkan or Metal only, and
SystemState.Capability falls back to "default", so those hosts get the CPU
build rather than a tag that does not exist. The nvidia-cuda-* and
nvidia-l4t-cuda-* keys are present because getSystemCapabilities() refines an
NVIDIA host to them whenever the CUDA directory exists; without them every
modern CUDA host and Jetson would miss the map and quietly run on CPU.

.github/backend-matrix.yml gets 7 include rows and 1 includeDarwin row. No
hipblas and no sycl rows, for the same upstream reason. cpu and vulkan are
per-arch pairs sharing a tag-suffix so backend-merge-jobs builds a multi-arch
manifest: an ARM host with no NVIDIA GPU reports "default" and the Jetson image
does not cover it.

The CI job is the substantive part. make test-extra is dead on master, because
prepare-test-extra depends on a protogen-python target that does not exist and
no workflow invokes it anyway, so the entry added earlier in this series ran
nowhere. abi_test.go asserts the size and field offsets of every Go mirror
struct against the C ABI it is dlopened into, and those assertions are the only
defence against silent memory corruption after a purego symbol rename or an
upstream header change. tests-nemo-speech-cpp in test-extra.yml now executes
them on pull_request and on master, gated on the backend's own path filter.
The recipe sets NEMO_SPEECH_REQUIRE_LIBS=1, so a missing library fails rather
than skips. WITH_NORM=OFF skips the OpenFST leg and costs no coverage: nothing
in the four C ABI headers is conditional on it, so the layouts are identical.

Also registers the upstream pin with the bump bot, which the backend Makefile
already claimed but was never wired up, and adds the BackendCapabilities entry
so a hand-written model config gets a real usecase surface. PossibleUsecases is
the union of the four families and DefaultUsecases is transcript alone, the
audio-cpp pattern. No VoiceCloning key: MagpieTTS synthesizes from baked
speaker ids, not a reference clip.

No gallery entries: publishing converted GGUFs is a follow-up.

ModelIdentity needs no work in this backend. main.go serves through
grpc.StartServer, so every RPC lands on pkg/grpc's shared server wrapper first,
and checkModelIdentity is the first statement of all seven handlers this
backend implements. A second check inside NemoSpeech would be unreachable and
would risk diverging from the cross-language sentinel the router matches on.
AudioTranscriptionLive stays unguarded because TranscriptLiveRequest carries no
ModelIdentity field at all, which is a proto-level gap affecting every backend
and needs its own change.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 21:55:53 +00:00
Ettore Di Giacinto
fdce1dca6c test(nemo-speech-cpp): pin the three-segment pair tag in an NMT directive
The directive regex allowed an unbounded run of two-letter segments per side, but
nothing tested it: narrowing that run back to a single optional segment left every
spec green. resolve_tag accepts a ready pair tag in one field with the other empty
(src/nmt/langpairs.cc), and those tags run to three segments (en-zh-cn, pt-br-en),
so a shorter pattern does not mis-split the tag, it fails to match the directive at
all and the whole bracket is handed to the model as text to translate.

The justification on the regex was also wrong and is corrected: pt-br and zh-cn are
two segments and parse either way. It is the single-field form that needs the run.

Renames the NMT handle to n.nmt so it stops sharing a name with the translator
interface, following n.synth, which is shortened for the same reason.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 21:36:25 +00:00
Ettore Di Giacinto
14ed23ef82 feat(nemo-speech-cpp): surface NMT translation through Predict
nemo_speech_nmt_translate takes explicit source and target languages and has no
free-form generation or token-callback entry point, so there is no prompt in the
LLM sense. The pair comes from the source_language / target_language model
options, with an optional leading [src->tgt] directive as the only per-request
override, and PredictStream emits the whole translation as a single chunk
because the C API has nothing finer to give it.

Both RPCs wrap their body in withEngine so the family check and the C calls that
trust the handle share one acquisition of engineMu. PredictStream closes its
channel on every path, including the family rejection: this is the legacy
streaming contract, and pkg/grpc/server.go blocks on a drain goroutine that only
finishes when the channel closes, so leaving it open hangs the RPC rather than
failing it.

nmtTranslatorConfig is extracted so its four adjacent pointer fields can be
asserted against distinct sentinels. Transposing two of them changes neither the
struct size nor any field offset, so the layout assertions cannot see it.

Also removes goString, which had no production caller: every string-returning
symbol in abi.go is bound with a Go string return that purego converts itself.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 21:24:09 +00:00
Ettore Di Giacinto
579cbb0a42 feat(nemo-speech-cpp): implement TTS and streaming TTS
The PCM callback is compiled once per process behind a sync.Once, not once per
request and not once per load. purego.NewCallback writes into a fixed table of
2000 entries (purego/syscall_sysv.go) and never releases one, so a per-request
callback panics the backend process on the 2001st synthesis, and a per-load one
reaches the same ceiling on a server that swaps models. Synthesis is routed
through that single callback plus a user_data id: engineMu is per-model, one
process holds several models, so a single current-sink pointer would be
overwritten by two TTS models synthesizing at once.

Deviations from the brief, all verified against the real headers and proto:

  - TTS is TTS(*pb.TTSRequest) error and TTSStream is
    TTSStream(*pb.TTSRequest, chan []byte) error, per pkg/grpc/interface.go.
    The brief's context/pb.Result and server-stream forms do not implement the
    interface. The channel is closed on every path, including the family
    rejection, because pkg/grpc/server.go blocks on its drain goroutine and an
    unclosed channel hangs the RPC with the backend lock held.
  - The callback takes unsafe.Pointer, not uintptr. Converting a uintptr
    parameter back to a pointer is a checkptr violation that aborts under
    -race.
  - resolveSpeaker refuses to turn a negative number into a speaker index. -1
    is the C API's "use the default" sentinel, so the brief's rule would have
    made a request naming an invalid voice synthesize in the default voice
    instead of being rejected.

temperature and cfg_scale each write their override flag as well:
magpietts/runtime.cpp reads the float only when the flag is set, so a
temperature without it is silently discarded.

Also folds in Task 8's review finding on asr.go: the six bare -1 sentinels in
loadASR move to an asrDiarConfig builder reusing diarGeometryDefault, with
specs. src/asr/c_api.cpp applies left_context_frames at >= 0, so a dropped
sentinel pins the model geometry to 0 and no layout assertion can see it.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 21:01:45 +00:00
Ettore Di Giacinto
48c9e63c87 fix(nemo-speech-cpp): pin the diarizer geometry sentinels and cap the segment buffer
The six frame-geometry overrides were written as -1 with nothing
asserting it. c_api.cpp applies left_context_frames at >= 0 while the
other five need > 0, so a dropped sentinel there pins the model's left
context to zero, and the struct keeps exactly the same shape, which is
all the layout assertions can see. Extracting diarModelConfig makes the
values assertable: five specs now pin all six frame fields, the device
index, the declared size and the NULL preset, each frame field on its
own line so a missing sentinel names itself.

distinctSpeakers had a spec with three segments over three distinct
labels, which len(segs) satisfies just as well as the real thing. Four
segments over three labels makes it a spec that can fail.

collectSegments sized its buffer straight from a count the C side
reported, and make() panics rather than erroring on a length it cannot
satisfy, so an uninitialised size_t coming back across the ABI killed
the backend process instead of failing one request. A ceiling of 2^22
segments, upwards of 93 hours of audio at one 80 ms frame each, turns
that into a diagnosable error.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 20:32:35 +00:00
Ettore Di Giacinto
4e553d4743 feat(nemo-speech-cpp): implement standalone diarization
loadDiarizer creates the Sortformer diarizer and Diarize serves the RPC
over a diarization stream: decode, chunked push, finish, then the
count-then-fill segments protocol.

nemo_speech_diar_segment carries start_time and end_time in SECONDS
already, not frame indices, so no conversion happens on the way to
DiarizeSegment.start/end and the model's seconds-per-frame is not
involved at all. The speaker label is the runtime's 1-based tag as a
decimal string, matching what wordsToSegments emits on the ASR path, so
the same speaker reads the same way whether a caller diarized a file or
transcribed it.

The six frame-geometry overrides are written as -1 rather than left
zero. c_api.cpp applies left_context_frames when it is >= 0 while every
other override needs > 0, so a zeroed config would silently pin the left
context to zero and change the model's streaming geometry.

nemo_speech_diar_segments writes *count before it rejects a buffer that
is too small, so a rejected fill still reports the size to retry with.
collectSegments uses that rather than truncating, bounded at four
attempts because the RPC holds engineMu for its whole body and an
unbounded retry would block an unload behind it.

Two DiarizeRequest knobs map onto the segmentation config, and the
proto and header names cross over: min_duration_on is the C
min_duration_sec and min_duration_off is the C min_gap_sec. Six fields
have no equivalent in this pipeline and are logged rather than dropped
in silence: num_speakers, min_speakers and max_speakers (Sortformer's
capacity is fixed by the checkpoint), clustering_threshold (there is no
clustering stage), include_text (no ASR here) and threads.

The empty-PCM guard fires before the stream is opened, so a silent clip
never reaches a purego entry point that would dereference &pcm[0].

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 20:19:53 +00:00
Ettore Di Giacinto
55070db8f8 fix(nemo-speech-cpp): make live deltas concatenate and fill segment words
runLive wrote the inter-utterance separator into the accumulated
transcript but emitted the delta without it, so a two-utterance turn sent
"one." and "two." while the terminal result read "one. two.". The live
consumer is the one that really concatenates: the realtime semantic-VAD
path joins the accumulated deltas with the empty string and clears them
only at a turn reset, never at an endpoint, so the running caption read
"one.two.". The separator now goes into the delta, as it already did on
the file path, and the terminal text is the verbatim concatenation rather
than a trimmed rebuild.

TranscriptSegment.Words was never populated, so a request asking for
timestamp_granularities ["word"] came back with no words at all even
though the timings were decoded. wordsToSegments now attaches them,
gated on the granularity the same way parakeet-cpp gates it, so a
transcript that did not ask for word timestamps does not pay for them.

Also: the final that comes back from the tail flush no longer claims an
end-of-utterance. It is the end of the stream, not a user yielding the
turn, and eou is what the realtime turn detector acts on.

The comment explaining why interims are suppressed led with the runtime's
postprocessing. The wire contract is the stronger reason and now comes
first: consumers concatenate deltas, so forwarding a growing hypothesis
assembles to "hehellhelloHello.". The postprocessing only explains why no
diffing trick would rescue them. It is also ITN and strip_formatting
rather than punctuation, which is off by default here.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 19:59:14 +00:00
Ettore Di Giacinto
4d295044a5 feat(nemo-speech-cpp): implement streaming and live transcription
AudioTranscriptionStream drives a whole clip through the cache-aware
streaming API in 100 ms pushes, emitting each finalized utterance as a
delta and closing with the assembled result. AudioTranscriptionLive
serves the bidirectional RPC over the same session: config first, a ready
ack, deltas with word timings as utterances land, and a terminal result
when the caller closes its send side.

Both wrap their body in withEngine, so a stream holds engineMu for its
whole life and Free waits on it rather than destroying the recognizer
underneath a half-finished stream. That makes the way out load-bearing:
the file loop honours the request context between pushes, and the live
loop ends when the host closes the request channel, so a disconnected
client cannot pin the model against unload.

Only finals become deltas. The runtime applies punctuation and inverse
text normalization on finals only, so a final rewrites the utterance
rather than extending its interim, and delta on the wire is
newly-finalized text that consumers concatenate. Forwarding interims
would duplicate and mispunctuate every utterance.

The four streaming entry points sit behind an asrSession interface. No
NeMo GGUF is small enough to keep in the tree, so without that seam the
need-more-audio drain would have no test at all: nemo_speech_asr_stream_next
reports OK with a NULL handle when it wants more audio, which is a pause
rather than an end, and reading it either way round drops results or
spins forever.

Also folds in three items from the offline transcription review:

  - empty audio is now refused before anything crosses the ABI, not
    inside recognizeF32. The added integration spec caught the old
    ordering panicking on an unbound entry point instead of failing;
  - an undecodable sample rate is an error rather than 0, which this
    runtime reads as "already at the model rate" and would have made a
    wrong rate silently pitch-shift the audio;
  - AudioTranscription guards its result pointer instead of relying on
    an unstated invariant.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 19:43:46 +00:00
Ettore Di Giacinto
cf79798294 feat(nemo-speech-cpp): implement offline transcription
Create the ASR recognizer in loadASR and serve AudioTranscription.

Segment times are int64 nanoseconds, not seconds: the proto field is an
int64 that core/backend reads straight into a time.Duration, while the
runtime reports word offsets in milliseconds. Words are grouped into one
segment per consecutive speaker run, with the 1-based speaker tag carried
through and 0 (untagged) left unlabelled.

The whole RPC body runs inside withEngine so the family check and the C
calls happen under one acquisition of engineMu. Free runs without the
backend lock, so checking the family and then relocking would let a
teardown destroy the handle in the gap. The audio decode is inside the
closure too, which costs nothing: base.SingleThread already serialises
this backend's RPCs.

recognizeF32 guards zero-length PCM. &pcm[0] panics on an empty slice, so
Go never reaches the C side's own "empty audio" rejection, and a silent
clip or a truncated upload is ordinary input.

pkg/utils has no WAV decode helper, only the ffmpeg normalisation, so
audio.go pairs AudioToWav with go-audio the way parakeet-cpp does. It
returns the sample rate rather than a duration, since the C API resamples
off that number.

Also closes the write-side half of the race Task 5 fixed on the read
side: Load now holds engineMu across the family switch and the n.fam
commit, matching Free. The loaders still must not take it.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 19:16:52 +00:00
Ettore Di Giacinto
139c57a8c0 fix(nemo-speech-cpp): pin the load ordering, close the engineMu race
Three review items, plus a defect the race detector turned up.

The spec covering "no family selected after a failed load" wrote junk to
a .gguf, so Load returned at ggufArchitecture before a family was ever
chosen and the assertion was vacuous. Generalised the GGUF test helper
to take a string architecture, and added a spec that loads a magpietts
GGUF with no sibling codec, so familyFor succeeds and discoverTTSAssets
then fails. It self-guards on ggufArchitecture so it cannot degrade back
into the earlier path.

requireFamily read n.fam unlocked while Free wrote it under engineMu,
which the race detector confirms is a real race. pkg/grpc/server.go
calls Free without the backend lock every other RPC holds, so teardown
can land mid-request. withEngine now takes the lock, checks the family
and runs the body under one acquisition; two would leave a window for
Free to destroy the handle between check and use. The locking protocol
is stated in both directions for the RPCs still to be written.

Running -race also enables checkptr, which aborts on cstr's pointer
being read back by goString: converting a uintptr to a pointer is fatal
whenever the address lands in a Go allocation, so a pinned Go buffer can
never be dereferenced from Go. The pointer is for C alone. Both helpers
now document the one-way contract, and goString is tested against a real
C-owned string by rebinding the version symbol to return a raw char*.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 18:54:12 +00:00
Ettore Di Giacinto
14bdd6cd8d feat(nemo-speech-cpp): select the family at load and gate RPCs on it
Load sniffs the GGUF architecture, maps it to a family and dispatches to
the family's loader. requireFamily gates every other RPC, returning
Unimplemented naming both the loaded and the wanted family so a
misconfigured model YAML produces a message a user can act on.

The family is committed only once its loader has succeeded. A load that
fails part way through would otherwise leave the gate open on a handle
that was never created.

cstr uses runtime.Pinner rather than an ordinary Go allocation. The
address crosses the ABI as a uintptr, which the collector does not
trace, so incidental reachability through the release closure is not a
guarantee: a caller discarding that closure could have the bytes
collected before the create call reads them. Pinning is the sanctioned
mechanism, makes the release function do real work, and turns a dropped
release into a loud leaked-Pinner panic instead of silent corruption.

Free overrides the base no-op to destroy the handle and reset the
family. Every family owns C memory only its own destroy entry point can
release, so without this an unloaded model leaks an acoustic model.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 18:32:47 +00:00
Ettore Di Giacinto
272a219830 test(nemo-speech-cpp): run the ABI specs in CI and refuse to skip them
The layout assertions were inert. TEST_PATHS does not cover this backend and
the per-backend list in test-extra had no entry for it, so nothing invoked the
package's tests. Add it next to depth-anything-cpp, supertonic and vllm-cpp,
the group whose own test target carries its build prerequisites; stage-libs
already pulls the native build chain, so no prepare-test-extra entry is needed.

The skip guard was also loader-inconsistent: librariesPresent stats bare
filenames relative to the working directory while openLibraries resolves them
through the loader search path, so any invocation other than make test skipped
every library-backed spec and still reported green. NEMO_SPEECH_REQUIRE_LIBS=1
turns that into a failure naming the directory and the remedy, and the Makefile
test target sets it. Unset, the plain skip survives so a developer without a
build can still run the pure-Go layer specs.

Trim the default-value fingerprint from roughly forty assertions to eight. It
was pinning tunables such as threads and flush_partial_chunk, so a legitimate
pin bump would have failed with a message reading like a layout error. What
survives is only header-documented contract: the lone non-zero max_alternatives,
the run of -1 sentinels and the zero that witnesses where it stops. Verified the
narrowed spec still catches a mirror and offset table corrupted in lockstep,
which is the one class only this layer sees.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 18:14:15 +00:00
Ettore Di Giacinto
da51e952f7 feat(nemo-speech-cpp): bind the C ABI with layout assertions
purego binds by name at runtime and the config structs are passed by pointer,
so both a renamed symbol and a mismatched struct layout would otherwise survive
a green build. registerSymbols names the failing symbol, and the layout specs
compare each Go mirror against the size the library reports for itself, against
the offsets a C compiler produces for the installed headers, and against the
default values upstream writes into the structs it returns.

Two of the bindings differ from the plan because the headers do. The plan's
nemo_speech_diar_segments signature omits the segmentation-config pointer that
diar.h declares as the second parameter, which would have shifted the output
buffer, the capacity and the count pointer one position each. And
nemo_speech_diar_stream_push_f32 was missing from the symbol table although
standalone diarization cannot work without it.

Also close the two panic and equality gaps left in family.go: ValueString panics
on a mistyped general.architecture, and the self-codec guard compared a Cleaned
candidate path against an uncleaned one, so a doubled separator let the primary
GGUF be selected as its own codec.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 17:55:56 +00:00
Ettore Di Giacinto
2127a26523 feat(nemo-speech-cpp): detect model family and discover TTS assets
Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 17:36:29 +00:00
Ettore Di Giacinto
9d0ac115f5 feat(nemo-speech-cpp): guard the empty option value and warn on a bad gpu index
An empty value like "vad_model:" must stay empty, since callers read the
empty string as "unset". That branch of resolve() had no spec: dropping the
guard left every spec green while parseOptions started returning the models
directory itself. Add the spec that fails without the guard.

A known key with an unparseable value is a typo, not a config from a newer
backend, and "gpu:banna" failed expensively: the model loaded, produced
correct output, and ran on CPU with no signal anywhere. Log it. Unknown keys
stay silently ignored, which is what keeps configs forward compatible.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 17:29:19 +00:00
Ettore Di Giacinto
472d144c88 feat(nemo-speech-cpp): parse model options
Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 17:22:25 +00:00
Ettore Di Giacinto
75162ec270 fix(nemo-speech-cpp): move the backend apt gate below the expensive layers
Addresses the fourth review round.

The nemo-speech-cpp apt block sat immediately after the shared apt layer, above
the Vulkan SDK build, the CUDA and ROCm installs, the Go toolchain and the
protoc download. Docker keys each layer on its parent, so inserting a step there
re-keys everything below it: a byte-identical shared layer is not enough, and
merging as it stood would have forced all of those to re-execute once for every
Go backend image. Move it down beside the existing opus, crispasr and
sherpa-onnx gates, which sit after those layers for the same reason.

Checked the ordering both ways before moving. Nothing between the two positions
uses these packages: the Vulkan and opus blocks install their own ninja and
pkg-config, go install protoc-gen-go needs the Go toolchain rather than protoc,
and the protoc 27.1 step is a release-binary download that needs neither
protobuf-compiler nor libprotobuf-dev. Nothing in the block needs anything those
layers provide; it uses only apt, and the mirror rewrite from the first RUN
persists in the image. It also runs no update-alternatives, so the default
compiler stays untouched for later layers. The diff against master is now a
single additive hunk with no shared layer touched.

Also preflight ITN_PROTOC. configure gates a preset PROTOC on test -n alone, so
a path that does not exist is accepted and the error surfaces much later as a
bare "No such file or directory" from inside make -C src/proto. The pin
introduced that failure on a box whose only protoc is in /usr/local/bin, which
worked before. Check it alongside the gcc-12 check and name the ITN_PROTOC=
override in the message.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 17:15:32 +00:00
Ettore Di Giacinto
fd8df77c3b fix(nemo-speech-cpp): pin protoc for ITN and make the norm stack its own target
Addresses the third review round.

Dockerfile.golang installs protoc 27.1 into /usr/local/bin, ahead of /usr/bin,
while libprotobuf-dev is the distro's 3.21 on noble and 3.12 on jammy.
Sparrowhawk resolves protoc from PATH at make time (configure.ac uses
AC_CHECK_PROG, so PROTOC substitutes to the bare word, and src/proto/Makefile.am
invokes it) and commits no pregenerated stubs, so the rule always runs. Code
generated by 27.1 includes google/protobuf/runtime_version.h and a
PROTOBUF_VERSION guard the older headers lack, so the WITH_NORM build could not
complete. Pin PROTOC to the apt one for that step; configure documents that a
pre-set value wins. The apt protoc and libprotobuf-dev come from one source
package at one version, which is the property that makes this correct.

The text-normalization stack is now a target keyed on a file build_itn_deps.sh
actually produces, rather than a side effect of the runtime library rule. As a
side effect make could not see whether it existed, so once the library was up to
date the script could never run again: a tree built WITH_NORM=OFF could not move
to ON, and make test hard-failed with no escape but a full 345 MB clean. It is
now built on demand and reachable on its own as 'make itn'. Staging keys on the
prefix existing rather than on WITH_NORM, so it stages what the tree actually
built, and package.sh's closure guard remains the backstop.

An already-configured build tree also now wins over the platform default, so a
tree built WITH_NORM=OFF is not silently reconfigured to ON by a bare make test,
which is what demanded gcc-12 from developers who chose not to have it. An
explicit WITH_NORM= on the command line still overrides both, and the ITN rule
preflights for gcc-12 with an error that names the alternative.

Move ninja-build out of the shared apt layer into the existing BACKEND-gated
block. Dockerfile.golang serves 225 matrix entries and only this backend
configures with -G Ninja, so the common list is byte-identical to master again
and no other image loses its cache.

Drop libabsl-dev and correct the comment that justified it. No base image here
ships protobuf 25, so nothing needs the absl split, and the cmake glob looks in
/usr/lib rather than the multiarch directory Ubuntu actually uses, so the
package could never have contributed anything.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 17:01:01 +00:00
Ettore Di Giacinto
5e91f6b39f fix(nemo-speech-cpp): make the closure guard fail closed and give CI its toolchain
Addresses the second review round.

The dependency-closure guard failed open. Its glob expands once per pass, so
each pass advanced the closure by exactly one level, and the fixed count of five
passes then fell out of the loop without checking whether anything remained. An
eight-deep chain packaged six libraries, exited zero and reported success. That
is the case the guard was written for: asr to sparrowhawk to protobuf to absl
already runs several levels deep, so a WITH_NORM build could ship missing its
deepest libraries and fail at first dlopen. The loop now runs until the staged
set stops growing, and exhausting the bound is a hard error rather than a silent
exit.

For the same reason, a build image with neither readelf nor objdump no longer
warns and skips. It cannot show the package is complete, so it refuses to ship
it. The guard is entered only when there is something to check, so an empty
package cannot trip the new error.

Dockerfile.golang installed ninja-build only in the Vulkan branch while this
Makefile runs cmake -G Ninja unconditionally, so the CPU, cuBLAS and L4T images
could not configure at all. ninja-build moves to the common apt list; it does
not change CMake's default generator, so it is inert for the other backends.

gcc-12 was nowhere in the tree, yet WITH_NORM defaults ON and
build_itn_deps.sh needs it, so the committed default was unbuildable in CI.
Install it, with the protobuf, absl, re2 and autotools that Sparrowhawk and
OpenFST need, gated on BACKEND so the other Go images do not carry it. The list
follows upstream's own docker/Dockerfile, trimmed of the gRPC, portaudio and
python entries a BUILD_GRPC=OFF build does not use. Text normalization stays ON:
downgrading it silently would ship a backend advertising a feature it lacks.

Also: make test depend on stage-libs, so LD_LIBRARY_PATH is not an empty
directory on a clean tree, and add an engine target so Dockerfile.golang's
cacheable prebuild layer is not skipped and a CUDA build stops recompiling all
of upstream on every Go-side change.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 16:40:16 +00:00
Ettore Di Giacinto
1c82d44f88 fix(nemo-speech-cpp): make 'build' produce the package and bundle the ITN stack
Addresses the review of the scaffold commit.

backend/Dockerfile.golang runs 'make -C backend/go/$(BACKEND) build' and then
copies package/ into the final image, so 'build' has to end with a populated
package/. It only staged shared objects, which would have shipped an image with
no binary and no libraries at all. The old staging recipe is now stage-libs and
the chain is stage-libs, nemo-speech-cpp-grpc, package, build, matching every
sibling Go backend.

Text normalization was packaged incorrectly. nemo_speech_text_normalization is
STATIC but links sparrowhawk, fstfar and fst PUBLIC, so they land as DT_NEEDED
on libnemo_speech_asr.so, and they live in a project-local prefix that nothing
else provides. WITH_NORM stays ON by default on Linux, since normalization is a
wanted feature. Instead stage_libs now copies .deps/itn/lib when WITH_NORM=ON,
and package.sh bundles it.

Staging that prefix is still not enough on its own: Sparrowhawk drags in
protobuf, re2 and absl, which neither build_itn_deps.sh nor
package-system-libs.sh provides. Rather than hard-code another hand-maintained
list, package.sh now walks the DT_NEEDED entries of everything staged and copies
whatever is unresolved, skipping the core set and the GPU set that the shared
scripts already own. It fails at package time, not at first dlopen, when
something cannot be resolved. On a WITH_NORM=OFF build the closure is already
complete and it copies nothing.

Restore CGO_ENABLED=0 on the Go build to match whisper, parakeet-cpp and
omnivoice-cpp. Note that purego reaches dlopen through fakecgo, so the binary is
dynamically linked either way; what the flag changes is the NEEDED set, and
lib/ld.so routing in run.sh exists precisely because the binary is not static.

Replace the hand-rolled .patched sentinel with upstream's
scripts/apply-ggml-patches.sh. It applies the series in filename order, exits
non-zero when a patch does not apply, and detects "already applied" by comparing
the full-series tree hash rather than an mtime, so it is safe to run every time
and there is no sentinel left to go stale or to wedge the build when deleted. It
is wired as an order-only prerequisite so running it does not force a relink.

Also: correct the package.sh header, which claimed three shared objects when
there are five and none of the TTS ones carry a _c suffix; give 'make test' the
LD_LIBRARY_PATH the dlopen tests will need; document that a NEMO_SPEECH_VERSION
bump needs 'make purge'; and extend 'clean' to remove package/ and the ITN
libraries.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 16:22:37 +00:00
Ettore Di Giacinto
2c5f5bf1d2 feat(nemo-speech-cpp): scaffold the backend and upstream build
Adds the backend skeleton and the NeMo-Speech.cpp build, pinned at
2e12e2def8a98ed06666f7ee3ca94e7193e04be4. The Go side is deliberately a stub:
it dlopens the runtime and starts the gRPC server, later work fills in the
symbol table and the model logic.

Three details of the upstream layout differ from what the plan assumed, and the
build reflects the real tree:

* The TTS C ABI ships as libnemo_speech_tts, not libnemo_speech_tts_c. Upstream
  compiles c_api.cpp straight into the implementation library and only aliases
  the nemo_speech_tts_c CMake target, so no _c object exists on disk. ASR and
  NMT do build a real _c shim.
* Shared objects land in build/bin, since upstream points
  CMAKE_LIBRARY_OUTPUT_DIRECTORY at ${CMAKE_BINARY_DIR}/bin.
* The ASR and NMT _c shims carry a DT_NEEDED on libnemo_speech_asr and
  libnemo_speech_nmt, so those are staged and packaged alongside them.
  Otherwise dlopen fails at startup.

The ggml patch step uses an order-only prerequisite. cmake writes into the
checkout and bumps its mtime past the sentinel, which would otherwise re-run
git apply over an already-patched tree and break every incremental build.

Assisted-by: Claude Code:claude-opus-5[1m]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 15:58:07 +00:00
mudler's LocalAI [bot]
ea438cdeaf feat(vllm-cpp): wire the full engine config surface through engine_args (#11159)
The backend could configure four of the engine's knobs (block size, KV block
count, max sequence length, max concurrent sequences) out of a config surface
that is considerably larger. Speculative decoding, prefix caching, the
chunked-prefill token budget, the scheduling policy and the external KV
connector were reachable from vllm.cpp's own HTTP server and from nothing
LocalAI could write in a model config.

Config now goes through `engine_args:`, the same map the vLLM and SGLang
backends take, with keys spelled as vLLM's own CLI flags so a speculative_config
or kv_transfer_config block written for vLLM works verbatim. The legacy
`options:` list keeps working and reads every key too; engine_args wins where
both set one. Unknown keys are logged and ignored rather than fatal: the field
is shared with the other engines, so a config carrying their knobs must not take
the model down.

Two details worth knowing:

`enable_prefix_caching: false` maps to the ABI tri-state force-OFF (2), not 0.
0 means "let the model capability decide" and dense architectures default the
cache on, so collapsing the two would silently enable it against an explicit
false. enable_jump_forward (ABI v10) shares the encoding, deferring to
VT_ENABLE_JUMP_FORWARD instead of to the model.

The importer probes config.json on a vllm-cpp import and writes
speculative_config: {method: mtp} when the checkpoint declares an MTP head, the
safetensors analogue of the llama-cpp importer's GGUF probe. DFlash draft repos
are refused with a warning instead, since a drafter cannot serve alone and the
pairing is not derivable from either repo. The draft path is resolved against
LocalAI's model directory, because the engine only looks in a directory holding
config.json or in the HF cache and never downloads: the repo-id spelling the
vLLM docs teach used to die deep in the load with "draft checkpoint not found".

docs/content/features/text-generation.md gains a vllm.cpp section covering the
engine_args table, all three speculative methods, LMCache and the legacy list.
The backend had no documentation page before.

This replaces a branch that had gone stale behind master and carried its own
route to ABI v10, which #11386 has since landed in minimal form. Rebased onto
that as a single commit rather than replaying the intermediate steps, whose
ABI v9 mirrors no longer make sense against master's pin. The Darwin build
fixes for Apple Clang's gnu-folding-constant diagnostic on C++, Objective-C and
Objective-C++, originally authored by localai-org-maint-bot, are folded in here.

Verified: `make abi-check` agrees at v10; unit specs, core/config and
core/gallery/importers green; and the full e2e passes in 1330s against a CPU
libvllm.so reporting ABI v10 with Qwen_Qwen3.5-0.8B-Q4_K_M.gguf (load, blocking
completion, streaming, chat and tool calls).

Assisted-by: Claude:claude-fable-5 golangci-lint

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 12:10:56 +02:00
localai-org-maint-bot
32023f3cb9 gallery: add Qwen3.5 9B Defiant Fable variants (#11335)
Add the MTP and plain Q4_K_M GGUF builds with their shared vision projector so LocalAI users can select accelerated or fallback llama.cpp inference.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-06 09:07:43 +02:00
localai-org-maint-bot
1b69da3bd7 gallery: add Qwen3.5 9B HauhauCS variants (#11339)
Add Q4_K_M and Q8_0 builds of the popular refusal-removed Qwen3.5 9B fine-tune, including its multimodal projector.

Assisted-by: Codex:gpt-5 [web]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-06 09:07:06 +02:00
mudler's LocalAI [bot]
5c29a79246 chore(model-gallery): ⬆️ update checksum (#11382)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-06 09:06:41 +02:00
mudler's LocalAI [bot]
93bc537e99 chore: ⬆️ Update antirez/ds4 to b0309611041655f4e45671cfd9c9886aff161406 (#11381)
⬆️ Update antirez/ds4

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-06 09:06:28 +02:00
Nandana Dileep
147a5ee783 fix(react-ui): stop traces page crash when switching trace tabs (#11387)
Switching from Backend Traces back to API Traces crashed the page
with "can't access property status, e.response is undefined" (#11376).
The API table briefly renders the previous tab's backend rows while the
refetch effect is still pending, and those rows carry no `response`
envelope. The status column dereferenced it unguarded. Render a neutral
placeholder instead of throwing, and cover the tab-switch scenario with
a regression spec.

Assisted-by: opencode:big-pickle

Signed-off-by: Nandana Dileep <110280757+nandanadileep@users.noreply.github.com>
2026-08-06 09:06:08 +02:00
mudler's LocalAI [bot]
102d91414e fix(vllm-cpp): mirror the engine's ABI v10 so the backend loads again (#11386)
The Go bindings mirror vllm.h by hand and refuse a library whose
vllm_abi_version differs from what they were written against. Two
automated pin bumps (#11174, #11352) moved VLLM_CPP_VERSION onto engines
declaring ABI v10 while govllmcpp.go still mirrored v5, so every
vllm-cpp image built since then panics at startup on every platform:

  panic: vllm-cpp: ABI mismatch: library reports v10, backend built against v5

Grow both PODs to the v10 layout: vllm_model_params gains
speculative_config, enable_prefix_caching, max_num_batched_tokens,
scheduling_policy, kv_transfer_config and enable_jump_forward (88 bytes),
vllm_sampling_params gains the v8 logits-processor pair (136 bytes). The
offsets in the specs come from offsetof() against the pinned header. All
of the new fields are inert when zeroed, so the engine behaves exactly as
it did under v5; the backend sets none of them.

Nothing cross-checked the two files, which is why a blind pin bump could
ship a backend that cannot load. The library build now runs abi-check
first: it compares VLLM_ABI_VERSION in the fetched header against
abiVersion in govllmcpp.go and fails the build naming both, instead of
leaving the mismatch for a user's runtime.

Fixes #11379

Assisted-by: Claude:claude-fable-5 golangci-lint

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-06 09:05:41 +02:00
mudler's LocalAI [bot]
b8264b48ad chore: ⬆️ Update CrispStrobe/CrispASR to 21901d3f7c23554f072964828363e49ddbc2dc68 (#11383)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-06 09:03:25 +02:00
mudler's LocalAI [bot]
bfce3ccfb9 chore: ⬆️ Update leejet/stable-diffusion.cpp to c6beeef35526c6dc94b74a7fb69f9d2e6a2a7a12 (#11384)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-06 09:02:52 +02:00
mudler's LocalAI [bot]
c86f617f61 chore: ⬆️ Update ikawrakow/ik_llama.cpp to cf1aa57e1a0fabfd015831718fc99d1aec01ada5 (#11380)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-06 09:02:36 +02:00
mudler's LocalAI [bot]
8b059e7ad7 chore: ⬆️ Update 0xShug0/audio.cpp to 7efbb58def443722ea540d931dd3debee3e4d5e8 (#11378)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-06 09:02:22 +02:00
mudler's LocalAI [bot]
75839de46a docs: ⬆️ update docs version mudler/LocalAI (#11377)
⬆️ Update docs version mudler/LocalAI

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-06 09:02:09 +02:00
Richard Palethorpe
f8d3f31594 fix(vram): contain malformed GGUF metadata (#11374)
Recover parser panics at metadata boundaries, skip unneeded remote arrays, and use the parser's overflow-hardened release. Keep detached gallery workers and CrispASR probes from terminating their processes on malformed GGUF input. Disable startup warming in the provided Compose files as an operational fallback.

Assisted-by: Codex:gpt-5

Signed-off-by: Richard Palethorpe <io@richiejp.com>
2026-08-06 09:01:56 +02:00
mudler's LocalAI [bot]
1271b97a46 docs(blog): cover the terminal agent in the 4.8 post (#11372)
docs(blog): cover the terminal agent, and fix the counts in the intro

The 4.8 post never mentions that `local-ai chat` stopped being a REPL
and became an agent (#11291): the nib harness compiled into the binary,
with tool use behind an approval gate, sub-agents, MCP servers, plugins
and skills, auto-configured against the local instance. It also ships a
shell integration script for zsh, bash and fish that binds Ctrl+Space.

That is one of the larger user-facing changes in the release and it was
missing from both the post and the release-notes highlights. Added a
section after 3D generation, including the breaking changes for anyone
who had habits around the old REPL: `/clear` is gone in favour of
`/compact`, and a model switch now keeps the conversation.

While in the intro, corrected the counts. The post said 374 pull
requests in twenty-one days, which was accurate when it was drafted on
the 4th but not once v4.8.0 was tagged on the 5th. The published release
notes say 386 in twenty-two days, and the intro now matches them rather
than contradicting them.

For the record, neither figure is exactly right: `git log --format=%s
v4.7.1..v4.8.0 | grep -cE '\(#[0-9]+\)$'` counts 388 squash-merged pull
requests, and 389 from v4.7.0. The notes were cut before the last few
landed. Matching the published notes was the priority here, since that
is the artifact everyone else quotes, and 386 is the number already in
circulation.


Assisted-by: Claude:claude-opus-5 [Claude Code]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-05 16:46:27 +02:00
Ettore Di Giacinto
2c0e7c584d website: re-record the hero and gallery clips for the 4.8 UI
The two landing-page clips predated the v4.8.0 interface work (#11288,
#11305, #11307): the gallery clip showed the retired light-theme Install
Models table, and the hero clip toured the Nodes pages in a full browser
window while its caption promised a chat completion on CPU.

Both are re-recorded from a real local-ai built from v4.8.0, dark theme,
app chrome only:

- hero-ui.mp4: a chat completion on lfm2.5-1.2b-instruct streaming on
  CPU with the live tok/s meter, so the caption now matches the footage.
  The poster frame is regenerated from the new clip.
- gallery.mp4: the Discover rail and detail pane, the hardware
  recommendation lanes, the VRAM-by-context chart, and a real install
  with the live progress banner.

The hand-typed model count moves from 1,585 to 1,255 in the three places
it appears, matching the distinct-model count the recorded UI shows on
screen. The 3d-generation clip is untouched: the post-capture UI changes
do not show in its footage.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-fable-5
2026-08-05 12:46:44 +00:00
localai-org-maint-bot
fb444f917f gallery: add Agents-A1 4B variants (#11365)
Add the official Q4_K_M and Q8_0 GGUF builds with their matching vision projectors so the compact agentic model can be installed through LocalAI.

Assisted-by: Codex:gpt-5 [web]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-05 09:40:37 +02:00
localai-org-maint-bot
a05a790021 fix(ci): emit verifiable backend signature bundles (#11366)
Cosign v2.4.1 does not select the Sigstore bundle format by default, while LocalAI's verifier only consumes OCI bundle referrers. Request the format explicitly for both registries and guard the producer contract with a shell regression test.

Document strict backend integrity configuration and release-tag identities for operators.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-05 09:39:35 +02:00
localai-org-maint-bot
9f62401fca feat(traces): show in-flight API requests (#11368)
Register JSON API exchanges before their handlers run so the traces dashboard can surface active work. Replace the live entry with the completed persisted record under the same ID, and clean it up if a handler panics.

Assisted-by: Codex:gpt-5

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-05 09:37:05 +02:00
Ettore Di Giacinto
4a5c5e51b7 website: say plainly that engines are swappable behind the same API
The runtime section described the small core and on-demand backends but
never stated the simple fact readers look for: one model can run on
llama.cpp while the next loads on vLLM, SGLang or MLX, behind the same
endpoint, and switching is one line in the model's config.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-fable-5
2026-08-05 07:33:08 +00:00
mudler's LocalAI [bot]
c61b6f2286 docs(blog): new DeepSeek and Laguna numbers, visuals, humanizer pass (#11369)
* docs(blog): new DeepSeek and Laguna numbers, visuals, humanizer pass

vllm.cpp master moved 26 commits past what the post was written against,
and two results changed enough to matter. Both came from the same lever:
staging weights device-resident at load instead of reading them from the
GGUF mmap over unified memory, which the GB10 reads about 20% slower per
GEMV than device memory.

- DeepSeek-V4-Flash against DwarfStar: 0.997x parity becomes 1.144x
  ahead, 18.69 vs 16.33 tok/s decode, same generated tokens.
- Laguna-XS-2.1 against vLLM: 87% becomes 1.03x, 44.46 vs 43.10 tok/s.
  New row in the scoreboard.

Adds three visuals. A chart of throughput against every reference engine,
which is worth having now that the spread is 0.976 to 1.144 rather than a
flat line at parity. The Activity page with four installs running, and the
model detail pane with all four pocket-35b variants. Both screenshots were
recaptured on 2026-08-04 because #11288, #11305, #11307 and #11222 had all
changed those pages since the earlier set.

llama.cpp is deliberately absent from the chart: its 1.18x is a prefill
ratio, and putting it on the same axis as throughput ratios would be
comparing two different measurements.

Also carries the media the release notes embed, since a GitHub release
body needs URLs that survive publishing and drag-and-drop has no CLI.
Supersedes #11364.

Humanizer pass on the prose. The post had collected five exactness idioms
in one section (token-for-token, byte-exact twice, byte-identical,
token-identical). One is precision, five is a tic, so the 27B row keeps
its "token-for-token identical" where identical output is the actual
claim and the rest say what they mean. That also fixed a hyphen in
predicate position ("is token-identical").

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

* docs(blog): redraw the benchmark chart as a branded card

The Flint bar chart was generic: default palette, no brand, and drawn
from zero, which made five ratios between 0.976 and 1.144 look like five
bars of roughly equal length.

Redrawn in the style of recorder-for-agents' render-card.sh cards, the
same shape as the vllm.cpp README GIF. Palette taken from the two logos
rather than invented (LocalAI navy #0E2632 and teal #469AAF, vllm.cpp
teal #3AB4CA), SVG generated by a small JS loop so the geometry is exact
at any scale, headless Chrome to PNG at 2x.

The substantive change is that bars now run from the 1.00 parity line
instead of from zero. Deviation is what the data is about, so DeepSeek's
+14.4% and MLX-LM's -2.4% are both legible, and the one row that is
behind is the one row in amber. Each bar carries its ratio and the raw
measurement under it.

Keeps the .html source next to the .png so the chart is editable later:
change a number, re-run render-card.sh.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-05 09:14:51 +02:00
localai-org-maint-bot
0332e9729f gallery: add LFM2.5 2.6B variants (#11351)
Add LiquidAI official Q4_K_M and Q8_0 GGUF builds with linked variant selection and documented generation defaults.

Assisted-by: Codex:gpt-5 [Hugging Face]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-05 01:37:33 +02:00
mudler's LocalAI [bot]
a8d310573e chore: ⬆️ Update mudler/vllm.cpp to 0757cac231ecd571a83c4fd2f50805c9251fc225 (#11352)
⬆️ Update mudler/vllm.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-05 01:37:09 +02:00
mudler's LocalAI [bot]
144baaa809 chore: ⬆️ Update ggml-org/whisper.cpp to 306c88f4d1286aec1bf96e544632897886af5501 (#11353)
⬆️ Update ggml-org/whisper.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-05 01:36:56 +02:00
mudler's LocalAI [bot]
86c2e9a273 chore: ⬆️ Update leejet/stable-diffusion.cpp to ea7f0c87cfe4c673263b4c201c596c7f1cbe2528 (#11354)
⬆️ Update leejet/stable-diffusion.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-05 01:36:41 +02:00
mudler's LocalAI [bot]
89995d7535 chore: ⬆️ Update 0xShug0/audio.cpp to 238ab6a9e321c17de8e120559f57efeedaeb1345 (#11355)
⬆️ Update 0xShug0/audio.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-05 01:36:26 +02:00
mudler's LocalAI [bot]
1f4ec3bdf8 chore: ⬆️ Update CrispStrobe/CrispASR to ec730908a418b6032f9e69ded6186d3f042a7747 (#11356)
⬆️ Update CrispStrobe/CrispASR

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-05 01:36:13 +02:00
mudler's LocalAI [bot]
1466aaa9f7 chore: ⬆️ Update antirez/ds4 to 6747e7718dd08f00b680d0c16231f2d59ec3747e (#11357)
⬆️ Update antirez/ds4

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-05 01:36:01 +02:00
mudler's LocalAI [bot]
e6712844ee chore: ⬆️ Update ikawrakow/ik_llama.cpp to 6b55d2c7504f482e7c8ec6cbf22a19f3778c522b (#11358)
⬆️ Update ikawrakow/ik_llama.cpp

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-05 01:35:49 +02:00
mudler's LocalAI [bot]
b1d964ef7b chore(model-gallery): ⬆️ update checksum (#11359)
⬆️ Checksum updates in gallery/index.yaml

Signed-off-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: mudler <2420543+mudler@users.noreply.github.com>
2026-08-05 01:35:37 +02:00
mudler's LocalAI [bot]
0d342c61d8 docs(backends): correct the vllm-cpp description in the gallery (#11363)
This is the text users read in the backends list and the gallery, and it
was the last place still describing vllm.cpp as "a from-scratch C++20
port of vLLM created and maintained by the LocalAI team" with no
indication of maturity.

Three corrections, matching the v4.8 release notes and blog post:

- It leads with ALPHA. These are alpha development builds and llama-cpp
  stays the recommendation for production, which is the single most
  useful thing to know before clicking install.
- It is maintained by the LocalAI team but developed in its own
  repository and usable without LocalAI. vLLM is named for what it
  actually is, the reference implementation that output is checked
  against and benchmarked against, rather than just the thing that was
  ported.
- It records the featureset that has grown past vLLM: GGUF loading,
  speculative decoding and KV offload, alongside the architecture and
  hardware coverage that were already listed.

Also notes that the project is expected to be renamed, with the new name
still to be decided, so anyone who installs it now is not surprised
later.

vllm-cpp-development inherits all of this through the YAML anchor, so
both entries are covered by the one edit. Verified the file still parses
and that both entries carry the new text.


Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-05 01:35:21 +02:00
mudler's LocalAI [bot]
4fec33966a docs(blog): final figures for the 4.8 post, and the MLX provider (#11362)
* docs(blog): final figures for the 4.8 post, and the MLX provider

The cycle closed at 374 PRs over twenty-one days, not the 321 over
eighteen the post was written against. Corrects the summary, the opening
line, the contributor count and the gallery total, and moves the date to
the day the release is cut.

Adds the MLX GEMM provider (#11137), which merged after the post was
written and is the one number an Apple Silicon reader wants: 1.54x to
2.19x on an M4 with time to first token roughly halving, both arms
toggled on one binary. The +/-10% caveat travels with the table rather
than being left in the PR.

Two lines edited against the no-ai-slop skill while I was in the file,
the same pass #11324 ran over the engines post:

- The opener balanced two clauses across a colon and closed on "without
  lying to you", which is the built-to-be-quoted shape readers picked
  out of the HN thread. It is a flat statement now.
- "This is a new modality rather than a new backend under an existing
  one" is a binary contrast that says nothing the next clause does not.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

* docs(blog): call vllm.cpp alpha, and finish the no-ai-slop pass

vllm.cpp is not a released backend and the post read like it was. The
old wording buried the caveat in a block quote at the end of the section
and still said "first release of a young engine". It now says plainly,
before the caveat can be skipped, that these are alpha development
builds, that shipping them in 4.8 is about letting people try the thing
rather than recommending it, and that llama-cpp stays the default.

Also completes the no-ai-slop pass I had only half run. Counting the
lines built to be quoted, headings and section endings included, the post
is in reasonable shape: long flat informational stretches, tables
followed by a plain finding, headings that are labels rather than
epigram-verdicts. Three patterns survived, each one an item in eval.md:

- "and inverts that:" set the usual shape against ours across a colon.
  The sentence works without the frame.
- "Two things were conflated there: a signal, which needs one line, and
  the detail, which needs somewhere to put it" is a role-assignment pair.
  Says what happens instead.
- "The maturity statement from the release notes is worth repeating in
  full" is throat-clearing in front of a quote, and the quote is gone.

Left the rest alone. Minimum effective edit, not a rewrite.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

* docs(blog): present vllm.cpp as a community project, with its own numbers

The post described vllm.cpp as "a from-scratch port of vLLM, written and
maintained by the LocalAI team". Two things wrong with that. It is a
community project, and it has stopped being only a port: it loads GGUF,
runs on CPU, Metal and Vulkan, ships speculative decoding and KV offload,
and its benchmark page measures against llama.cpp, MLX-LM and DwarfStar
as well as vLLM, because those are the engines it competes with on that
hardware.

vLLM's role is now stated for what it is, the reference implementation.
Correctness is checked against it and the scoreboard is kept against it.
Also flags that the name will probably change, since it is drifting far
enough that vllm.cpp will eventually mislead.

Adds real numbers from the project's own docs/BENCHMARKS.md rather than
adjectives: 1.045x vLLM at concurrency 1 on Qwen3.6-27B NVFP4 with
token-for-token identical output, 1.010x and 1.013x at c16 and c32 on the
35B MoE and behind below that, prefill 1.18x over llama.cpp on CPU
aarch64, 97.6% of MLX-LM warm total on an M4. Upstream's own caution
travels with them: it treats c2 through c32 as ties because its noise
band is 0.5% and those margins are 0.7% to 1.7%.

Every figure was checked against ~/_git/vllm.cpp/docs/BENCHMARKS.md
rather than restated from memory. The heading is marked alpha to match
the section body.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

* docs(blog): say who maintains vllm.cpp, and add the DeepSeek Flash result

Two corrections to the previous commit.

"A community project" says nothing and was not quite true either. The
LocalAI team maintains vllm.cpp. Community-first is the intent, not a
description, so it now says that and says what backs it: its own
repository, its own docs, benchmark record and issue tracker, and it runs
without LocalAI anywhere in the picture.

Adds the DeepSeek-V4-Flash result, which makes the divergence point
better than any of the prose around it. That model does not run on vLLM
on a single GB10: every vLLM-loadable checkpoint is 156 GB or more
against a 119 GiB unified pool, and the only quant that fits is an
extreme-low-bit GGUF that vLLM cannot load. vllm.cpp reads GGUF and runs
it at 16.28 tok/s against ds4's 16.33, a parity result. Also notes MTP
speculative decoding, token-identical to vLLM's and about 4% faster at
concurrency 1.

Both figures checked against ~/_git/vllm.cpp/docs/BENCHMARKS.md.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

* docs(blog): lead the DeepSeek result with what we run, not with what vLLM cannot

The previous version opened on "that model does not run on vLLM on a
single GB10 at all". Wrong emphasis twice over: it makes a strong
negative claim about another project the headline, and it buries the
actual result, which is that vllm.cpp runs DeepSeek-V4-Flash at roughly
2-bit (IQ2_XXS mixed, about 80 GB) on a single DGX Spark and decodes at
16.28 tok/s against DwarfStar's 16.33.

The size constraint is still there, stated as the reason the quant is
what it is rather than as a point about vLLM: at 300B+ total parameters
even a 4-bit checkpoint is 156 GB or more, so a 2-bit GGUF is what fits
the Spark's 119 GiB unified pool.

The table row now names the quant and the box (IQ2_XXS, one DGX Spark)
instead of just "GGUF, GB10", since that is the part a reader with a
Spark wants.

Figures unchanged and still from ~/_git/vllm.cpp/docs/BENCHMARKS.md.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

* docs(blog): say the new name is undecided

"The name will probably change at some point" invited the obvious
question. It now says the rename is expected and the name is still to be
decided, which is the actual state.

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Assisted-by: Claude Code:claude-opus-5 [Read] [Edit] [Bash]

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
2026-08-05 01:23:49 +02:00
localai-org-maint-bot
8f52437c81 fix(gallery): describe Genesis Hermes model accurately (#11342)
Replace copied HauhauCS base-model text with metadata for the actual Genesis Hermes V6 artifact and link its upstream base model.

Assisted-by: Codex:gpt-5 [Hugging Face]

Co-authored-by: localai-org-maint-bot <306269227+localai-org-maint-bot@users.noreply.github.com>
2026-08-04 17:47:32 +02:00
94 changed files with 11005 additions and 149 deletions

View File

@@ -16,8 +16,7 @@ side (`pkg/oci/cosignverify` plus the gallery YAML).
per-arch manifest before checking signatures.
- **Storage:** Signatures are written as OCI 1.1 referrers
(`--registry-referrers-mode=oci-1-1`) in the new Sigstore bundle format
(current cosign releases do this by default; no `--new-bundle-format`
flag). No `:sha256-<hex>.sig` tag clutter.
(`--new-bundle-format`). No `:sha256-<hex>.sig` tag clutter.
- **Consumer:** `pkg/oci/cosignverify` discovers the bundle via the
referrers API, hands it to `sigstore-go`, and verifies it against the
policy declared in the gallery YAML (`Gallery.Verification`).
@@ -34,14 +33,15 @@ to sign. The job needs:
- `permissions: { id-token: write, contents: read }` at the job level so
the runner can exchange its GitHub OIDC token for a Fulcio cert.
- `sigstore/cosign-installer@v3` step (current cosign releases already
default to the new bundle format).
- `sigstore/cosign-installer@v3` step (the pinned cosign v2 release needs
`--new-bundle-format` explicitly).
- After each `docker buildx imagetools create`, resolve the resulting
list digest with `docker buildx imagetools inspect <tag> --format
'{{.Manifest.Digest}}'` and sign:
```sh
cosign sign --yes --recursive \
--new-bundle-format \
--registry-referrers-mode=oci-1-1 \
"${REGISTRY_REPO}@${DIGEST}"
```
@@ -70,7 +70,7 @@ entry (`backend/index.yaml`):
url: github:mudler/LocalAI/backend/index.yaml@master
verification:
issuer: "https://token.actions.githubusercontent.com"
identity_regex: "^https://github\\.com/mudler/LocalAI/\\.github/workflows/backend_merge\\.yml@refs/heads/master$"
identity_regex: "^https://github\\.com/mudler/LocalAI/\\.github/workflows/backend_merge\\.yml@refs/(heads/master|tags/.+)$"
# Optional revocation cutoff; advance during incident response.
# not_before: "2026-06-01T00:00:00Z"
```

View File

@@ -860,6 +860,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-12-nemo-speech-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "nemo-speech-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "8"
@@ -1911,6 +1924,19 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-nvidia-cuda-13-nemo-speech-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "nemo-speech-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
@@ -1963,6 +1989,24 @@ include:
backend: "parakeet-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
# The CUDA-13 counterpart to the JetPack r36.4.0 row in the nemo-speech-cpp
# block below. A Jetson whose CUDA 13 runtime is present reports the
# nvidia-l4t-cuda-13 capability, and pointing that key at the JetPack image
# would hand it a ggml linked against CUDA 12 whose libcudart.so.12 is not
# there to dlopen. Same base and runner as the parakeet-cpp row above.
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
platforms: 'linux/arm64'
skip-drivers: 'false'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-cuda-13-arm64-nemo-speech-cpp'
base-image: "ubuntu:24.04"
ubuntu-version: '2404'
runs-on: 'ubuntu-24.04-arm'
backend: "nemo-speech-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
- build-type: 'cublas'
cuda-major-version: "13"
cuda-minor-version: "0"
@@ -4183,6 +4227,86 @@ include:
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
# nemo-speech-cpp
#
# No hipblas and no sycl rows, unlike the parakeet-cpp block above: upstream
# NeMo-Speech.cpp builds ggml with CUDA, Vulkan or Metal only, so a ROCm or
# SYCL image would be a CPU build wearing a GPU tag.
#
# cpu and vulkan are per-arch pairs sharing one tag-suffix, so
# backend-merge-jobs assembles a multi-arch manifest from the two digests.
# The arm64 legs are not redundant with the Jetson image below: an ARM server
# with no NVIDIA GPU reports the "default" capability and would otherwise pull
# an amd64-only manifest.
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-cpu-nemo-speech-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "nemo-speech-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: ''
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-cpu-nemo-speech-cpp'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "nemo-speech-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/amd64'
platform-tag: 'amd64'
tag-latest: 'auto'
tag-suffix: '-gpu-vulkan-nemo-speech-cpp'
runs-on: 'ubuntu-latest'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "nemo-speech-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'vulkan'
cuda-major-version: ""
cuda-minor-version: ""
platforms: 'linux/arm64'
platform-tag: 'arm64'
tag-latest: 'auto'
tag-suffix: '-gpu-vulkan-nemo-speech-cpp'
runs-on: 'ubuntu-24.04-arm'
base-image: "ubuntu:24.04"
skip-drivers: 'false'
backend: "nemo-speech-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2404'
- build-type: 'cublas'
cuda-major-version: "12"
cuda-minor-version: "0"
platforms: 'linux/arm64'
skip-drivers: 'false'
tag-latest: 'auto'
tag-suffix: '-nvidia-l4t-arm64-nemo-speech-cpp'
base-image: "nvcr.io/nvidia/l4t-jetpack:r36.4.0"
runs-on: 'ubuntu-24.04-arm'
backend: "nemo-speech-cpp"
dockerfile: "./backend/Dockerfile.golang"
context: "./"
ubuntu-version: '2204'
# moss-transcribe-cpp
- build-type: ''
cuda-major-version: ""
@@ -6226,6 +6350,10 @@ includeDarwin:
tag-suffix: "-metal-darwin-arm64-moss-transcribe-cpp"
build-type: "metal"
lang: "go"
- backend: "nemo-speech-cpp"
tag-suffix: "-metal-darwin-arm64-nemo-speech-cpp"
build-type: "metal"
lang: "go"
- backend: "ced"
tag-suffix: "-metal-darwin-arm64-ced"
build-type: "metal"

View File

@@ -71,8 +71,8 @@ jobs:
# cosign signs each pushed manifest list with --recursive so the
# index and every per-arch entry get an attached Sigstore bundle.
# Recent cosign releases always emit the new bundle format, so
# there's no extra CLI flag to opt into it.
# The pinned cosign v2 release needs --new-bundle-format explicitly;
# the verifier only consumes OCI 1.1 Sigstore bundle referrers.
- name: Install cosign
if: github.event_name != 'pull_request'
uses: sigstore/cosign-installer@v3
@@ -159,6 +159,7 @@ jobs:
# manifest before checking signatures need the per-arch
# signatures, not just the list-level one.
cosign sign --yes --recursive \
--new-bundle-format \
--registry-referrers-mode=oci-1-1 \
"quay.io/go-skynet/local-ai-backends@${digest}"
@@ -185,6 +186,7 @@ jobs:
' <<< "$DOCKER_METADATA_OUTPUT_JSON")
digest=$(docker buildx imagetools inspect "$first_tag" --format '{{.Manifest.Digest}}')
cosign sign --yes --recursive \
--new-bundle-format \
--registry-referrers-mode=oci-1-1 \
"localai/localai-backends@${digest}"

View File

@@ -62,6 +62,10 @@ jobs:
variable: "MOSS_VERSION"
branch: "master"
file: "backend/go/moss-transcribe-cpp/Makefile"
- repository: "NVIDIA/NeMo-Speech.cpp"
variable: "NEMO_SPEECH_VERSION"
branch: "main"
file: "backend/go/nemo-speech-cpp/Makefile"
- repository: "localai-org/ced.cpp"
variable: "CED_VERSION"
branch: "main"

View File

@@ -50,6 +50,7 @@ jobs:
sherpa-onnx: ${{ steps.detect.outputs.sherpa-onnx }}
whisper: ${{ steps.detect.outputs.whisper }}
parakeet-cpp: ${{ steps.detect.outputs.parakeet-cpp }}
nemo-speech-cpp: ${{ steps.detect.outputs.nemo-speech-cpp }}
steps:
- name: Checkout repository
uses: actions/checkout@v7
@@ -900,6 +901,57 @@ jobs:
- name: Test magpie-tts-cpp
run: |
make --jobs=5 --output-sync=target -C backend/go/magpie-tts-cpp test
# Per-backend unit suite for nemo-speech-cpp. This job exists for one reason
# above all: abi_test.go asserts the size and field offsets of every Go mirror
# struct against the C ABI it is dlopened into. Those assertions are the only
# thing standing between a purego symbol rename or an upstream header change
# and silent memory corruption at run time, and they are worthless unless
# something executes them. `make -C backend/go/nemo-speech-cpp test` sets
# NEMO_SPEECH_REQUIRE_LIBS=1, which turns "library missing" from a skip into a
# failure, so this job cannot report green having checked nothing.
#
# The backend Makefile's `test` target depends on `stage-libs`, so it clones
# upstream at the pinned SHA and builds the native runtime itself. There is no
# separate build step for that reason, and no model download: the specs are
# ABI and pure-Go only.
#
# WITH_NORM=OFF skips the Sparrowhawk/OpenFST inverse-text-normalization
# stack, which is the single most expensive leg of the build and needs a gcc-12
# pin because OpenFST's templates ICE on gcc-13/14. It costs no coverage here:
# nothing in include/nemo_speech/{asr,tts,diar,nmt}.h is conditional on it (the
# only preprocessor conditionals in those headers are include guards,
# __cplusplus and the _WIN32 export macros), so every struct layout this suite
# checks is identical either way. The shipped images still build WITH_NORM=ON;
# that path is covered by the backend image build in backend_pr.yml.
tests-nemo-speech-cpp:
needs: detect-changes
if: needs.detect-changes.outputs.nemo-speech-cpp == 'true' || needs.detect-changes.outputs.run-all == 'true'
runs-on: ubuntu-latest
timeout-minutes: 90
steps:
- name: Clone
uses: actions/checkout@v7
with:
submodules: true
- name: Dependencies
run: |
sudo apt-get update
sudo apt-get install -y build-essential cmake ninja-build curl libopenblas-dev ffmpeg
- name: Setup Go
uses: actions/setup-go@v5
- name: Display Go version
run: go version
- name: Proto Dependencies
run: |
curl -L -s https://github.com/protocolbuffers/protobuf/releases/download/v26.1/protoc-26.1-linux-x86_64.zip -o protoc.zip && \
unzip -j -d /usr/local/bin protoc.zip bin/protoc && \
rm protoc.zip
go install google.golang.org/protobuf/cmd/protoc-gen-go@v1.34.2
go install google.golang.org/grpc/cmd/protoc-gen-go-grpc@1958fcbe2ca8bd93af633f11e97d44e567e945af
PATH="$PATH:$HOME/go/bin" make protogen-go
- name: Test nemo-speech-cpp
run: |
make --jobs=5 --output-sync=target -C backend/go/nemo-speech-cpp WITH_NORM=OFF test
# Per-backend smoke for rfdetr-cpp: builds the .so + Go binary and runs
# `make -C backend/go/rfdetr-cpp test`. test.sh fetches the small (~20 MB)
# rfdetr-nano-q8_0 GGUF from the published mudler/rfdetr-cpp-nano HF repo

View File

@@ -1,5 +1,5 @@
# Disable parallel execution for backend builds
.NOTPARALLEL: backends/diffusers backends/llama-cpp backends/turboquant backends/bonsai backends/outetts backends/piper backends/stablediffusion-ggml backends/trellis2cpp backends/trellis2cpp-darwin backends/whisper backends/crispasr backends/parakeet-cpp backends/moss-transcribe-cpp backends/faster-whisper backends/silero-vad backends/local-store backends/valkey-store backends/cloud-proxy backends/huggingface backends/rfdetr backends/rfdetr-cpp backends/insightface backends/speaker-recognition backends/kitten-tts backends/kokoro backends/chatterbox backends/llama-cpp-darwin backends/neutts build-darwin-python-backend build-darwin-go-backend backends/mlx backends/diffuser-darwin backends/mlx-vlm backends/mlx-audio backends/mlx-distributed backends/stablediffusion-ggml-darwin backends/vllm backends/vllm-omni backends/longcat-video backends/sglang backends/moonshine backends/pocket-tts backends/qwen-tts backends/faster-qwen3-tts backends/qwen-asr backends/nemo backends/voxcpm backends/whisperx backends/ace-step backends/acestep-cpp backends/fish-speech backends/voxtral backends/opus backends/trl backends/llama-cpp-quantization backends/kokoros backends/sam3-cpp backends/qwen3-tts-cpp backends/moss-tts-cpp backends/magpie-tts-cpp backends/vllm-cpp backends/omnivoice-cpp backends/vibevoice-cpp backends/localvqe backends/tinygrad backends/sherpa-onnx backends/ds4 backends/ds4-darwin backends/liquid-audio backends/supertonic backends/depth-anything-cpp backends/privacy-filter backends/privacy-filter-darwin backends/audio-cpp backends/audio-cpp-darwin
.NOTPARALLEL: backends/diffusers backends/llama-cpp backends/turboquant backends/bonsai backends/outetts backends/piper backends/stablediffusion-ggml backends/trellis2cpp backends/trellis2cpp-darwin backends/whisper backends/crispasr backends/parakeet-cpp backends/moss-transcribe-cpp backends/nemo-speech-cpp backends/faster-whisper backends/silero-vad backends/local-store backends/valkey-store backends/cloud-proxy backends/huggingface backends/rfdetr backends/rfdetr-cpp backends/insightface backends/speaker-recognition backends/kitten-tts backends/kokoro backends/chatterbox backends/llama-cpp-darwin backends/neutts build-darwin-python-backend build-darwin-go-backend backends/mlx backends/diffuser-darwin backends/mlx-vlm backends/mlx-audio backends/mlx-distributed backends/stablediffusion-ggml-darwin backends/vllm backends/vllm-omni backends/longcat-video backends/sglang backends/moonshine backends/pocket-tts backends/qwen-tts backends/faster-qwen3-tts backends/qwen-asr backends/nemo backends/voxcpm backends/whisperx backends/ace-step backends/acestep-cpp backends/fish-speech backends/voxtral backends/opus backends/trl backends/llama-cpp-quantization backends/kokoros backends/sam3-cpp backends/qwen3-tts-cpp backends/moss-tts-cpp backends/magpie-tts-cpp backends/vllm-cpp backends/omnivoice-cpp backends/vibevoice-cpp backends/localvqe backends/tinygrad backends/sherpa-onnx backends/ds4 backends/ds4-darwin backends/liquid-audio backends/supertonic backends/depth-anything-cpp backends/privacy-filter backends/privacy-filter-darwin backends/audio-cpp backends/audio-cpp-darwin
GOCMD=go
GOTEST=$(GOCMD) test
@@ -654,6 +654,7 @@ test-extra: prepare-test-extra
$(MAKE) -C backend/go/depth-anything-cpp test
$(MAKE) -C backend/go/supertonic test
$(MAKE) -C backend/go/vllm-cpp test
$(MAKE) -C backend/go/nemo-speech-cpp test
$(MAKE) -C backend/go/trellis2cpp test
$(MAKE) -C backend/go/valkey-store test
@@ -1298,6 +1299,7 @@ BACKEND_WHISPER = whisper|golang|.|false|true
BACKEND_CRISPASR = crispasr|golang|.|false|true
BACKEND_PARAKEET_CPP = parakeet-cpp|golang|.|false|true
BACKEND_MOSS_TRANSCRIBE_CPP = moss-transcribe-cpp|golang|.|false|true
BACKEND_NEMO_SPEECH_CPP = nemo-speech-cpp|golang|.|false|true
BACKEND_DEPTH_ANYTHING_CPP = depth-anything-cpp|golang|.|false|true
BACKEND_VOXTRAL = voxtral|golang|.|false|true
BACKEND_ACESTEP_CPP = acestep-cpp|golang|.|false|true
@@ -1400,6 +1402,7 @@ $(eval $(call generate-docker-build-target,$(BACKEND_WHISPER)))
$(eval $(call generate-docker-build-target,$(BACKEND_CRISPASR)))
$(eval $(call generate-docker-build-target,$(BACKEND_PARAKEET_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_MOSS_TRANSCRIBE_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_NEMO_SPEECH_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_DEPTH_ANYTHING_CPP)))
$(eval $(call generate-docker-build-target,$(BACKEND_VOXTRAL)))
$(eval $(call generate-docker-build-target,$(BACKEND_OPUS)))
@@ -1456,7 +1459,7 @@ $(eval $(call generate-docker-build-target,$(BACKEND_SUPERTONIC)))
docker-save-%: backend-images
docker save local-ai-backend:$* -o backend-images/$*.tar
docker-build-backends: docker-build-llama-cpp docker-build-ik-llama-cpp docker-build-turboquant docker-build-bonsai docker-build-ds4 docker-build-rerankers docker-build-vllm docker-build-vllm-omni docker-build-longcat-video docker-build-sglang docker-build-transformers docker-build-outetts docker-build-diffusers docker-build-kokoro docker-build-faster-whisper docker-build-crispasr docker-build-coqui docker-build-chatterbox docker-build-vibevoice docker-build-liquid-audio docker-build-moonshine docker-build-pocket-tts docker-build-qwen-tts docker-build-fish-speech docker-build-faster-qwen3-tts docker-build-qwen-asr docker-build-nemo docker-build-voxcpm docker-build-whisperx docker-build-ace-step docker-build-acestep-cpp docker-build-voxtral docker-build-mlx-distributed docker-build-trl docker-build-llama-cpp-quantization docker-build-tinygrad docker-build-kokoros docker-build-sam3-cpp docker-build-rfdetr-cpp docker-build-qwen3-tts-cpp docker-build-moss-tts-cpp docker-build-magpie-tts-cpp docker-build-vllm-cpp docker-build-omnivoice-cpp docker-build-vibevoice-cpp docker-build-localvqe docker-build-insightface docker-build-speaker-recognition docker-build-sherpa-onnx docker-build-cloud-proxy docker-build-supertonic docker-build-depth-anything-cpp docker-build-moss-transcribe-cpp docker-build-privacy-filter docker-build-trellis2cpp docker-build-valkey-store docker-build-audio-cpp
docker-build-backends: docker-build-llama-cpp docker-build-ik-llama-cpp docker-build-turboquant docker-build-bonsai docker-build-ds4 docker-build-rerankers docker-build-vllm docker-build-vllm-omni docker-build-longcat-video docker-build-sglang docker-build-transformers docker-build-outetts docker-build-diffusers docker-build-kokoro docker-build-faster-whisper docker-build-crispasr docker-build-coqui docker-build-chatterbox docker-build-vibevoice docker-build-liquid-audio docker-build-moonshine docker-build-pocket-tts docker-build-qwen-tts docker-build-fish-speech docker-build-faster-qwen3-tts docker-build-qwen-asr docker-build-nemo docker-build-voxcpm docker-build-whisperx docker-build-ace-step docker-build-acestep-cpp docker-build-voxtral docker-build-mlx-distributed docker-build-trl docker-build-llama-cpp-quantization docker-build-tinygrad docker-build-kokoros docker-build-sam3-cpp docker-build-rfdetr-cpp docker-build-qwen3-tts-cpp docker-build-moss-tts-cpp docker-build-magpie-tts-cpp docker-build-vllm-cpp docker-build-omnivoice-cpp docker-build-vibevoice-cpp docker-build-localvqe docker-build-insightface docker-build-speaker-recognition docker-build-sherpa-onnx docker-build-cloud-proxy docker-build-supertonic docker-build-depth-anything-cpp docker-build-moss-transcribe-cpp docker-build-nemo-speech-cpp docker-build-privacy-filter docker-build-trellis2cpp docker-build-valkey-store docker-build-audio-cpp
########################################################
### Mock Backend for E2E Tests

View File

@@ -248,6 +248,52 @@ RUN <<EOT bash
fi
EOT
# nemo-speech-cpp builds NVIDIA NeMo-Speech.cpp with text normalization enabled,
# which compiles the Sparrowhawk/OpenFST WFST stack from source via
# scripts/build_itn_deps.sh. That step needs gcc-12 specifically: OpenFST's
# template-heavy translation units ICE on gcc-13 and gcc-14 at -O2, so upstream
# pins gcc-12 for it while the runtime itself builds with the image default.
# No update-alternatives here, so the default compiler is untouched; the backend
# Makefile reaches gcc-12 by name for that one step.
#
# The rest is what build_itn_deps.sh and the WITH_NORM cmake block expect:
# protobuf (headers plus protoc, which must come from the same apt set so the
# generated stubs match the headers they compile against) and re2 for
# Sparrowhawk, and autotools because OpenFST and Sparrowhawk ship autoconf
# builds. ninja is not in the common apt list because this is the only Go
# backend that configures with -G Ninja, and that list is a layer shared by
# every backend image in the matrix.
#
# No libabsl-dev, despite upstream's Dockerfile installing it: upstream builds
# against protobuf 25, which splits its runtime across libabsl_*, whereas every
# base image in this matrix carries protobuf 3.21 (noble) or 3.12 (jammy), which
# has no absl dependency. The cmake block's file(GLOB ... /usr/lib/libabsl_*.so)
# would not match on Ubuntu anyway, since multiarch puts those under
# /usr/lib/<triplet>/.
#
# Placed down here with the other per-backend gates rather than next to the
# shared apt layer: Docker re-keys every layer below an inserted one, so adding
# a step above the Vulkan SDK, CUDA, Go and protoc layers would force all of
# them to re-execute once for every Go backend image, not just this one.
# Nothing between there and here needs any of these packages (the Vulkan and
# opus blocks install their own ninja and pkg-config, and the protoc download is
# a release binary that needs neither libprotobuf-dev nor protoc from apt), and
# nothing here needs anything those layers provide.
RUN <<EOT bash
if [ "${BACKEND}" = "nemo-speech-cpp" ]; then
set -e
apt-get update
apt-get install -y --no-install-recommends \
gcc-12 g++-12 \
ninja-build \
libprotobuf-dev protobuf-compiler \
libre2-dev \
autoconf automake libtool pkg-config
apt-get clean
rm -rf /var/lib/apt/lists/*
fi
EOT
RUN git config --global --add safe.directory /LocalAI
# Prebuild the native engine from a layer that depends on this backend's own

View File

@@ -9,7 +9,7 @@
# recipe is a make target (not a prepare.sh) so 'make purge && make' is a clean
# rebuild and so the bump bot can see the pin.
AUDIO_CPP_VERSION?=4e3aea2fd99aeaa5924e71c51eb2793846045332
AUDIO_CPP_VERSION?=7efbb58def443722ea540d931dd3debee3e4d5e8
AUDIO_CPP_REPO?=https://github.com/0xShug0/audio.cpp
CURRENT_MAKEFILE_DIR := $(dir $(abspath $(lastword $(MAKEFILE_LIST))))

View File

@@ -1,10 +1,10 @@
# ds4 backend Makefile.
#
# Upstream pin lives below as DS4_VERSION?=b7e9f0091139999b6c070a57590c447c5741da5c
# Upstream pin lives below as DS4_VERSION?=b0309611041655f4e45671cfd9c9886aff161406
# (.github/bump_deps.sh) can find and update it - matches the
# llama-cpp / ik-llama-cpp / turboquant convention.
DS4_VERSION?=b7e9f0091139999b6c070a57590c447c5741da5c
DS4_VERSION?=b0309611041655f4e45671cfd9c9886aff161406
DS4_REPO?=https://github.com/antirez/ds4
CURRENT_MAKEFILE_DIR := $(dir $(abspath $(lastword $(MAKEFILE_LIST))))

View File

@@ -1,5 +1,5 @@
IK_LLAMA_VERSION?=60389410a1ff01f9d37dcc6261db33b3183bdea2
IK_LLAMA_VERSION?=cf1aa57e1a0fabfd015831718fc99d1aec01ada5
LLAMA_REPO?=https://github.com/ikawrakow/ik_llama.cpp
CMAKE_ARGS?=

View File

@@ -8,7 +8,7 @@ JOBS?=$(shell nproc --ignore=1)
# CrispASR version (release tag)
CRISPASR_REPO?=https://github.com/CrispStrobe/CrispASR
CRISPASR_VERSION?=fe3caf8e363b27572dbdd1a9d37083f25e6decda
CRISPASR_VERSION?=21901d3f7c23554f072964828363e49ddbc2dc68
SO_TARGET?=libgocrispasr.so
CMAKE_ARGS+=-DBUILD_SHARED_LIBS=OFF

View File

@@ -67,7 +67,16 @@ const defaultTTSSampleRate = 24000
// resampling, so the WAV header must match it. Returns ok=false for non-piper
// models (key absent) or an unreadable file, letting the caller fall back to
// defaultTTSSampleRate.
func piperSampleRate(modelPath string) (int, bool) {
func piperSampleRate(modelPath string) (rate int, ok bool) {
// A malformed metadata length can make gguf-parser-go panic before it can
// return an error. Keep a bad voice file from crash-looping the backend.
defer func() {
if recover() != nil {
rate = 0
ok = false
}
}()
// Only scalar architecture keys are read, so skip the large array metadata
// (phoneme map) and mmap the header - same rationale as pkg/vram's reader.
f, err := gguf.ParseGGUFFile(modelPath, gguf.UseMMap(), gguf.SkipLargeMetadata())
@@ -78,7 +87,7 @@ func piperSampleRate(modelPath string) (int, bool) {
if !ok || kv.ValueType != gguf.GGUFMetadataValueTypeUint32 {
return 0, false
}
rate := int(kv.ValueUint32())
rate = int(kv.ValueUint32())
if rate <= 0 {
return 0, false
}

View File

@@ -3,6 +3,7 @@ package main
import (
"bytes"
"encoding/binary"
"math"
"os"
"path/filepath"
@@ -102,6 +103,24 @@ var _ = Describe("piper sample rate", func() {
_, ok := piperSampleRate(p)
Expect(ok).To(BeFalse())
})
It("returns ok=false instead of panicking on a malformed string length", func() {
p := filepath.Join(GinkgoT().TempDir(), "malformed.gguf")
var b bytes.Buffer
b.WriteString("GGUF")
Expect(binary.Write(&b, binary.LittleEndian, uint32(3))).To(Succeed())
Expect(binary.Write(&b, binary.LittleEndian, uint64(0))).To(Succeed())
Expect(binary.Write(&b, binary.LittleEndian, uint64(1))).To(Succeed())
key := "general.name"
Expect(binary.Write(&b, binary.LittleEndian, uint64(len(key)))).To(Succeed())
b.WriteString(key)
Expect(binary.Write(&b, binary.LittleEndian, ggufTypeString)).To(Succeed())
Expect(binary.Write(&b, binary.LittleEndian, uint64(math.MaxInt64))).To(Succeed())
Expect(os.WriteFile(p, b.Bytes(), 0o644)).To(Succeed())
_, ok := piperSampleRate(p)
Expect(ok).To(BeFalse())
})
})
// End-to-end through the built .so. Gated on CRISPASR_PIPER_MODEL_PATH (a

22
backend/go/nemo-speech-cpp/.gitignore vendored Normal file
View File

@@ -0,0 +1,22 @@
# Fetched upstream sources
sources/
# CMake build directories
build*/
# Packaging output
package/
# Compiled backend binary. The second name is what a bare `go build ./...` from
# this directory produces (it names the binary after the directory), as opposed
# to the -o name the Makefile asks for.
nemo-speech-cpp-grpc
/nemo-speech-cpp
# Shared libraries staged in-tree by the Makefile (cp from sources/). The
# SOVERSION suffix means the payload is libnemo_speech_*.so.1, hence both globs.
*.so
*.so.*
*.dylib
compile_commands.json

View File

@@ -0,0 +1,312 @@
# nemo-speech-cpp backend Makefile.
#
# Upstream pin lives below as NEMO_SPEECH_VERSION so .github/bump_deps.sh can
# find and update it, matching the parakeet-cpp / vibevoice-cpp convention.
#
# Bumping NEMO_SPEECH_VERSION is a no-op on an existing checkout: sources/ is a
# directory target, so make only clones when it is missing and never re-checks
# out an already-cloned tree. After a bump run 'make purge && make', the same
# rule the parakeet-cpp Makefile documents.
#
# 'build' is the entry point the backend image calls (backend/Dockerfile.golang
# runs 'make -C backend/go/$(BACKEND) build' and then copies package/), so it
# has to produce the binary and the package, not just the shared libraries.
NEMO_SPEECH_VERSION?=2e12e2def8a98ed06666f7ee3ca94e7193e04be4
NEMO_SPEECH_REPO?=https://github.com/NVIDIA/NeMo-Speech.cpp
GOCMD?=go
GO_TAGS?=
JOBS?=$(shell nproc 2>/dev/null || sysctl -n hw.ncpu 2>/dev/null || echo 4)
BUILD_TYPE?=
NATIVE?=false
# NEMO_SPEECH_CUBLAS_SHIM defaults ON upstream and builds a drop-in
# libcublas.so.13. LocalAI's CUDA images ship the real cuBLAS, so the shim would
# shadow it with a slower native GEMM. Always OFF here.
CMAKE_ARGS?=-DCMAKE_BUILD_TYPE=Release \
-DBUILD_SHARED_LIBS=OFF \
-DCMAKE_POSITION_INDEPENDENT_CODE=ON \
-DNEMO_SPEECH_CUBLAS_SHIM=OFF \
-DNEMO_SPEECH_BUILD_ASR=ON \
-DNEMO_SPEECH_BUILD_DIAR=ON \
-DNEMO_SPEECH_BUILD_TTS=ON \
-DNEMO_SPEECH_BUILD_NMT=ON \
-DNEMO_SPEECH_BUILD_CLI=OFF \
-DNEMO_SPEECH_BUILD_HTTP=OFF \
-DNEMO_SPEECH_BUILD_GRPC=OFF \
-DNEMO_SPEECH_WITH_FLASHLIGHT=OFF \
-DNEMO_SPEECH_TTS_WITH_ZH=ON \
-DNEMO_SPEECH_TTS_WITH_JA=ON
ifeq ($(NATIVE),false)
CMAKE_ARGS+=-DGGML_NATIVE=OFF
endif
# NEMO_SPEECH_TTS_WITH_JA=ON compiles Open JTalk's bundled MeCab, and
# mecab/src/dictionary.cpp derives a comparator from std::binary_function, which
# C++17 removed. libstdc++ still ships it as deprecated-but-present under
# -std=gnu++17, so Linux never notices; libc++ compiles it out and the build dies
# with "no template named 'binary_function' in namespace 'std'". Upstream's own
# CMakeLists already carries the equivalent workaround for MSVC's STL
# (_HAS_AUTO_PTR_ETC plus /FIfunctional) but has no libc++ branch, because
# NEMO_SPEECH_TTS_WITH_JA defaults OFF upstream and only LocalAI turns it on.
#
# libc++ gates the two templates on _LIBCPP_ENABLE_CXX17_REMOVED_UNARY_BINARY_FUNCTION,
# and has done since LLVM 16, which is older than any clang Xcode still ships.
# The name matters: the older _LIBCPP_ENABLE_CXX17_REMOVED_BINDERS covers
# bind1st/bind2nd/ptr_fun/mem_fun and NOT unary_function/binary_function, and the
# umbrella _LIBCPP_ENABLE_CXX17_REMOVED_FEATURES no longer exists at all. A wrong
# name is silently accepted by the preprocessor and fixes nothing.
#
# Applied through CMAKE_CXX_FLAGS rather than to the one target because the
# tokenizer CMakeLists is upstream's and this tree is a pinned checkout, not a
# patched one. Project-wide is also the safer scope: the macro decides whether
# libc++'s internal __binary_function alias resolves to std::binary_function or
# to __binary_function_keep_layout_base, which is a base class of std::less and
# friends, so defining it for a subset of translation units would give those
# class templates two spellings in one binary. Both bases are empty and, at
# C++17, carry identical members, so the project-wide define changes no layout
# and no ABI. On Linux the macro is not a name libstdc++ knows, so the branch is
# unreachable there and would be inert even if it were taken.
ifeq ($(shell uname -s),Darwin)
CXX_COMPAT_FLAGS?=-D_LIBCPP_ENABLE_CXX17_REMOVED_UNARY_BINARY_FUNCTION
else
CXX_COMPAT_FLAGS?=
endif
ifneq ($(strip $(CXX_COMPAT_FLAGS)),)
CMAKE_ARGS+=-DCMAKE_CXX_FLAGS=$(CXX_COMPAT_FLAGS)
endif
# scripts/build_itn_deps.sh installs the Sparrowhawk/OpenFST runtime here.
# NEMO_SPEECH_DEPENDENCY_PREFIX defaults to <src>/.deps upstream, and the ITN
# stack goes under its itn/ subdirectory. ITN_MARKER is a real output of that
# script (it prints exactly this file on success), so it can drive a make rule.
ITN_LIB_DIR=sources/NeMo-Speech.cpp/.deps/itn/lib
ITN_MARKER=$(ITN_LIB_DIR)/libsparrowhawk.so
ITN_CC?=gcc-12
ITN_CXX?=g++-12
# Pin protoc to the apt one. backend/Dockerfile.golang drops protoc 27.1 into
# /usr/local/bin, which precedes /usr/bin on PATH, while libprotobuf-dev is the
# distro's (3.21 on noble, 3.12 on jammy). Sparrowhawk resolves protoc from PATH
# at make time (configure.ac uses AC_CHECK_PROG, so PROTOC substitutes to the
# bare word, and src/proto/Makefile.am invokes $(PROTOC)), and it commits no
# pregenerated stubs, so this always runs. Code generated by 27.1 includes
# google/protobuf/runtime_version.h and a PROTOBUF_VERSION #error guard that the
# older headers do not have, so the mismatch breaks the build. configure honours
# a pre-set PROTOC ("Let the user override the test"), which is what this is.
ITN_PROTOC?=/usr/bin/protoc
# Text normalization is Linux-only: Sparrowhawk/OpenFST assume a GNU toolchain
# and the gcc-12 pin has no macOS analogue. Documented gap, see the spec.
#
# An already-configured build tree wins over the platform default. Without that,
# a tree configured WITH_NORM=OFF would silently try to reconfigure itself to ON
# on the next bare `make test`, which means demanding gcc-12 from a developer who
# deliberately built without it. An explicit WITH_NORM= on the command line still
# overrides both, since command-line variables beat ?= assignments.
CMAKE_CACHE=sources/NeMo-Speech.cpp/build/CMakeCache.txt
CACHED_WITH_NORM=$(shell sed -n 's/^NEMO_SPEECH_WITH_NORM:BOOL=//p' $(CMAKE_CACHE) 2>/dev/null)
ifeq ($(shell uname -s),Darwin)
WITH_NORM?=OFF
else ifneq ($(CACHED_WITH_NORM),)
WITH_NORM?=$(CACHED_WITH_NORM)
else
WITH_NORM?=ON
endif
CMAKE_ARGS+=-DNEMO_SPEECH_WITH_NORM=$(WITH_NORM)
ifeq ($(BUILD_TYPE),cublas)
CMAKE_ARGS+=-DGGML_CUDA=ON
else ifeq ($(BUILD_TYPE),vulkan)
CMAKE_ARGS+=-DGGML_VULKAN=ON
else ifeq ($(BUILD_TYPE),metal)
CMAKE_ARGS+=-DGGML_METAL=ON
endif
# ggml-patches/ is a CUDA series. Every kernel it adds lives under
# src/ggml-cuda/; the only files it touches outside that directory are enum and
# name-table entries in include/ggml.h and src/ggml.c plus, in ggml-cpu, a
# supports_op returning false and an abort case for the CUDA-only op. Upstream
# agrees: its metal-* and vulkan-* CMake presets inherit the cpu-* ones, which
# set NEMO_SPEECH_GGML_PATCHED=OFF, and every use of a patch-only symbol in the
# ASR sources sits behind NEMO_SPEECH_FUSED_RELPOS_ATTN /
# NEMO_SPEECH_FASTCONFORMER_CUDA_FUSIONS (both force-OFF without GGML_CUDA) or
# behind NEMO_SPEECH_GGML_PATCHED itself, which guards a Q8_PLANAR flag write
# that a non-CUDA buffer already throws before reaching.
#
# So on macOS the series buys nothing, and it cannot be applied there anyway:
# upstream's scripts/apply-ggml-patches.sh uses mapfile, a bash 4 builtin, and
# macOS ships bash 3.2 as the only bash on the runner's PATH. Skip the patch
# step and tell cmake the linked ggml is stock, which is exactly upstream's own
# Metal configuration. Linux keeps applying the series unchanged.
ifeq ($(shell uname -s),Darwin)
GGML_PATCHED?=OFF
else
GGML_PATCHED?=ON
endif
CMAKE_ARGS+=-DNEMO_SPEECH_GGML_PATCHED=$(GGML_PATCHED)
.PHONY: nemo-speech-cpp-grpc package build clean purge test all stage-libs patch-ggml engine itn
all: nemo-speech-cpp-grpc package
sources/NeMo-Speech.cpp:
mkdir -p sources
cd sources && git clone $(NEMO_SPEECH_REPO) NeMo-Speech.cpp
cd sources/NeMo-Speech.cpp && git checkout $(NEMO_SPEECH_VERSION)
# NMT links llama.cpp; ja needs open_jtalk; zh needs cppjieba. flashlight and
# kenlm are deliberately not initialized, they are out of scope.
cd sources/NeMo-Speech.cpp && git submodule update --init --recursive \
ggml llama.cpp third_party/open_jtalk third_party/cppjieba third_party/cpp-httplib
# NEMO_SPEECH_GGML_PATCHED defaults ON and silently assumes the ggml-patches
# series is applied. An unpatched checkout builds fine and produces wrong CUDA
# encoder output, so a failure here must stop the build rather than warn.
#
# Upstream's own script is the right tool: it applies the series in filename
# order, exits non-zero when a patch does not apply, and decides "already
# applied" by comparing the full-series tree hash rather than a timestamp. That
# makes it safe to run unconditionally, so there is no sentinel file to go stale
# or to wedge the build when deleted.
#
# Both branches keep the order-only clone prerequisite: it is the only thing
# that pulls sources/ in on a WITH_NORM=OFF tree, where the library rule has no
# other prerequisite left.
ifeq ($(GGML_PATCHED),ON)
patch-ggml: | sources/NeMo-Speech.cpp
cd sources/NeMo-Speech.cpp && bash scripts/apply-ggml-patches.sh
else
patch-ggml: | sources/NeMo-Speech.cpp
@echo "[ggml-patch] skipped: NEMO_SPEECH_GGML_PATCHED=$(GGML_PATCHED), the series is CUDA-only"
endif
# The Sparrowhawk/OpenFST text-normalization stack, as a target in its own right
# keyed on a file the build script actually produces.
#
# It used to be a side effect of the runtime library rule, which meant make had
# no idea whether it existed: once the library was up to date the script could
# never run again, so a tree built WITH_NORM=OFF could not be moved to ON, and
# anything that needed the ITN prefix was stuck demanding a full clean. As its
# own rule it is built on demand, rebuilt independently, and reachable directly
# with 'make itn'.
#
# OpenFST's templates ICE on gcc-13/14 at -O2, hence the gcc-12 pin for this one
# step; the runtime itself builds with the image default compiler.
$(ITN_MARKER): | sources/NeMo-Speech.cpp
@command -v $(ITN_CC) >/dev/null 2>&1 && command -v $(ITN_CXX) >/dev/null 2>&1 || { \
echo "ERROR: $(ITN_CC)/$(ITN_CXX) not found, and text normalization needs them:" >&2; \
echo " OpenFST's templates ICE on gcc-13 and gcc-14 at -O2." >&2; \
echo " Install them, or build this backend with WITH_NORM=OFF." >&2; \
exit 1; }
# configure's only gate on a preset PROTOC is test -n, so a path that does not
# exist is accepted here and surfaces much later as a bare "No such file or
# directory" from inside make -C src/proto. Check it up front instead.
@command -v $(ITN_PROTOC) >/dev/null 2>&1 || { \
echo "ERROR: protoc not found at $(ITN_PROTOC)." >&2; \
echo " Install the protobuf-compiler package, whose protoc matches" >&2; \
echo " the libprotobuf-dev headers Sparrowhawk compiles against, or" >&2; \
echo " point this at a matching one with ITN_PROTOC=/path/to/protoc." >&2; \
exit 1; }
cd sources/NeMo-Speech.cpp && CC=$(ITN_CC) CXX=$(ITN_CXX) PROTOC=$(ITN_PROTOC) \
JOBS=$(JOBS) scripts/build_itn_deps.sh
itn: $(ITN_MARKER)
# Only a WITH_NORM=ON build needs the ITN stack, and it must exist before cmake
# configures, since the WITH_NORM cmake block find_library()s into the prefix
# with REQUIRED.
ifeq ($(WITH_NORM),ON)
NEMO_RUNTIME_PREREQS=$(ITN_MARKER)
endif
# Upstream sets CMAKE_LIBRARY_OUTPUT_DIRECTORY to ${CMAKE_BINARY_DIR}/bin, so the
# shared objects land in build/bin rather than at the top of the build tree.
#
# patch-ggml is order-only: it is phony and therefore always runs, but an
# order-only prerequisite does not mark this target out of date, so an
# already-built tree is not relinked on every invocation.
sources/NeMo-Speech.cpp/build/bin/libnemo_speech_asr_c.so: $(NEMO_RUNTIME_PREREQS) | patch-ggml
cd sources/NeMo-Speech.cpp && cmake -B build -G Ninja $(CMAKE_ARGS)
cd sources/NeMo-Speech.cpp && cmake --build build -j$(JOBS)
# Stage the runtime next to the Go sources so purego.Dlopen finds it during
# local development and so package.sh has a single directory to bundle from.
#
# ASR and NMT build a dedicated _c shared object that links the C++ implementation
# in privately. TTS does not: upstream compiles its c_api.cpp straight into
# libnemo_speech_tts and only aliases the nemo_speech_tts_c CMake target, so the
# TTS C ABI ships without the _c suffix.
stage-libs: sources/NeMo-Speech.cpp/build/bin/libnemo_speech_asr_c.so
# -a keeps the SOVERSION symlink a symlink instead of duplicating the payload.
cp -af sources/NeMo-Speech.cpp/build/bin/libnemo_speech_asr_c.* .
cp -af sources/NeMo-Speech.cpp/build/bin/libnemo_speech_tts.* .
cp -af sources/NeMo-Speech.cpp/build/bin/libnemo_speech_nmt_c.* .
# The _c libraries are thin ABI shims with a DT_NEEDED on the C++
# implementation DSO, so dlopen fails without these next to them. TTS needs
# no counterpart, its implementation and ABI live in the same object.
cp -af sources/NeMo-Speech.cpp/build/bin/libnemo_speech_asr.* .
cp -af sources/NeMo-Speech.cpp/build/bin/libnemo_speech_nmt.* .
# nemo_speech_text_normalization is STATIC but links sparrowhawk, fstfar and
# fst PUBLIC, so those become DT_NEEDED on libnemo_speech_asr.so. They live in
# a project-local prefix that nothing else on the system provides, so without
# staging them here the packaged backend cannot dlopen at all.
#
# Keyed on the prefix existing rather than on WITH_NORM, so this stages what
# the tree actually built. A WITH_NORM=ON build cannot reach here without the
# prefix (the library rule takes ITN_MARKER as a prerequisite), and if a
# library that needs Sparrowhawk somehow arrives unstaged, package.sh's
# closure guard fails the build rather than shipping it.
@if [ -d "$(ITN_LIB_DIR)" ]; then \
echo "cp -af $(ITN_LIB_DIR)/*.so* ."; \
cp -af $(ITN_LIB_DIR)/*.so* .; \
fi
## Builds the native runtime and stops short of the Go binary. Everything it
## touches lives under sources/, a clone pinned by NEMO_SPEECH_VERSION, so
## nothing here can observe a change elsewhere in the LocalAI tree.
## Dockerfile.golang calls this from a layer that copies in this directory and
## nothing else, which keeps the multi-minute ggml/llama.cpp compile in the
## registry layer cache across builds whose only change is on the Go side.
## Without it that prebuild is skipped and a CUDA build recompiles all of
## upstream on every Go-side edit. See .agents/ci-caching.md.
engine: stage-libs
nemo-speech-cpp-grpc: stage-libs
# CGO_ENABLED=0 matches whisper / parakeet-cpp / omnivoice-cpp: the runtime is
# reached through purego.Dlopen, not cgo, and a static binary is what lets
# run.sh route execution through the packaged lib/ld.so.
CGO_ENABLED=0 $(GOCMD) build -tags "$(GO_TAGS)" -o nemo-speech-cpp-grpc .
# The dlopen tests need the staged shared objects on the loader path, the same
# way parakeet-cpp sets it up. Depends on stage-libs so that path is not an
# empty directory on a clean tree, which would fail the tests confusingly.
#
# NEMO_SPEECH_REQUIRE_LIBS turns a missing library from a skip into a failure.
# The ABI specs are the only thing standing between this backend and silent
# memory corruption, so a run that reaches them and quietly skips them is worse
# than one that fails: it reports green having checked nothing.
test: stage-libs
NEMO_SPEECH_REQUIRE_LIBS=1 LD_LIBRARY_PATH=$(CURDIR):$$LD_LIBRARY_PATH $(GOCMD) test ./... -count=1
package: nemo-speech-cpp-grpc
bash package.sh
# What backend/Dockerfile.golang invokes. It must leave both the binary and a
# populated package/ behind, because the final image stage copies package/.
build: package
clean:
# Every .so here is staged output (nemo runtime plus, on a WITH_NORM build,
# the ITN stack), and the SOVERSION suffix means the payload is *.so.1, so
# the globs have to reach past the .so.
rm -f nemo-speech-cpp-grpc
rm -f *.so *.so.* *.dylib
rm -rf package
rm -rf sources/NeMo-Speech.cpp/build
purge: clean
rm -rf sources

View File

@@ -0,0 +1,423 @@
package main
// purego binds by name at runtime and the config structs cross the ABI by
// pointer, so neither a renamed symbol nor a mis-laid-out mirror struct is
// visible to the compiler or the linker. Everything here is transcribed from
// sources/NeMo-Speech.cpp/include/nemo_speech/{asr,diar,tts,nmt}.h, and
// abi_test.go asserts it against the real shared objects.
import (
"fmt"
"unsafe"
"github.com/ebitengine/purego"
)
var (
asrLib uintptr
ttsLib uintptr
nmtLib uintptr
)
// ---- ASR ----
var (
ASRCreate func(cfg unsafe.Pointer, out *uintptr) int32
ASRDestroy func(recognizer uintptr)
ASRRecognizeF32 func(recognizer uintptr, options unsafe.Pointer, samples *float32, nSamples uint64, sampleRate int32, out *uintptr) int32
ASRStreamingRecognize func(recognizer uintptr, options unsafe.Pointer, out *uintptr) int32
ASRStreamPushF32 func(stream uintptr, samples *float32, nSamples uint64, sampleRate int32) int32
ASRStreamForceEndpoint func(stream uintptr) int32
ASRStreamFinish func(stream uintptr) int32
ASRStreamNext func(stream uintptr, out *uintptr) int32
ASRStreamClose func(stream uintptr)
ASRRecognitionOptionsDef func() cASRRecognitionOptions
ASRResultIsFinal func(result uintptr) bool
ASRResultAudioProcessed func(result uintptr) float32
ASRResultAlternativeCount func(result uintptr) uint64
ASRResultTranscript func(result uintptr, alt uint64) string
ASRResultConfidence func(result uintptr, alt uint64) float32
ASRResultWordCount func(result uintptr, alt uint64) uint64
ASRResultWordText func(result uintptr, alt, i uint64) string
ASRResultWordStartTime func(result uintptr, alt, i uint64) int32
ASRResultWordEndTime func(result uintptr, alt, i uint64) int32
ASRResultWordConfidence func(result uintptr, alt, i uint64) float32
ASRResultWordSpeakerTag func(result uintptr, alt, i uint64) int32
ASRResultLanguageCount func(result uintptr, alt uint64) uint64
ASRResultLanguageCode func(result uintptr, alt, i uint64) string
ASRResultDestroy func(result uintptr)
ASRLastError func() string
ASRVersion func() string
)
// ---- Diarization (exported from the ASR library) ----
var (
DiarCreate func(cfg unsafe.Pointer, out *uintptr) int32
DiarDestroy func(model uintptr)
DiarNumSpeakers func(model uintptr) int32
DiarSecondsPerFrame func(model uintptr) float64
DiarStreamOpen func(model uintptr, out *uintptr) int32
DiarStreamPushF32 func(stream uintptr, samples *float32, nSamples uint64, sampleRate int32) int32
DiarStreamFinish func(stream uintptr) int32
DiarStreamClose func(stream uintptr)
// cfg is the optional nemo_speech_diar_segmentation_config (NULL = library
// defaults). The two-call count-then-fill pattern is documented on the C
// declaration in diar.h.
DiarSegments func(stream uintptr, cfg unsafe.Pointer, out unsafe.Pointer, capacity uint64, count *uint64) int32
)
// ---- TTS ----
var (
TTSCreate func(cfg unsafe.Pointer, out *uintptr) int32
TTSDestroy func(synthesizer uintptr)
TTSSampleRate func(synthesizer uintptr) int32
TTSSpeakerCount func(synthesizer uintptr) int32
TTSSpeakerName func(synthesizer uintptr, i uint64) string
TTSSynthesizeText func(synthesizer uintptr, options unsafe.Pointer, text string, callback uintptr, userData uintptr, statsOut unsafe.Pointer) int32
TTSRuntimeConfigDefault func() cTTSRuntimeConfig
TTSSynthesisOptionsDefault func() cTTSSynthesisOptions
TTSLastError func() string
TTSVersion func() string
)
// ---- NMT ----
var (
NMTCreate func(cfg unsafe.Pointer, out *uintptr) int32
NMTDestroy func(translator uintptr)
NMTTranslate func(translator uintptr, texts *uintptr, nTexts uint64, source, target string, out *uintptr) int32
NMTResultCount func(result uintptr) uint64
NMTResultText func(result uintptr, i uint64) string
NMTResultLanguage func(result uintptr, i uint64) string
NMTResultDestroy func(result uintptr)
NMTLastError func() string
NMTVersion func() string
)
// ---- C struct mirrors ----
//
// Each mirrors a struct in include/nemo_speech/*.h field for field. The leading
// Size field is the C `size_t size` the runtime validates against its own
// sizeof, which is what makes a layout mismatch detectable at runtime instead
// of silently corrupting memory. Blank fields are System V AMD64 / AAPCS64
// padding: C inserts it implicitly, Go does not, so it has to be written out.
// See abi_test.go, which pins both every total size and every field offset.
type cASRBackendConfig struct {
Size uintptr
GPU int32
_ [4]byte // trailing pad to the struct's 8-byte alignment
}
type cASRModelConfig struct {
Size uintptr
Path uintptr
Name uintptr
}
type cASRVADConfig struct {
Size uintptr
ModelPath uintptr
EnableMasking bool
_ [3]byte
Onset float32
Offset float32
_ [4]byte
}
type cASRPostprocConfig struct {
Size uintptr
ProfanityListPath uintptr
ITNModelDir uintptr
PNCModelPath uintptr
}
type cASRDiarConfig struct {
Size uintptr
ModelPath uintptr
ChunkFrames int32
RightContextFrames int32
LeftContextFrames int32
FIFOFrames int32
SpkcacheFrames int32
UpdatePeriodFrames int32
}
type cASRRecognizerConfig struct {
Size uintptr
Backend uintptr
Model uintptr
Streaming uintptr
Decoder uintptr
VAD uintptr
Endpointing uintptr
Postproc uintptr
Diar uintptr
Batching uintptr
}
type cASRRecognitionOptions struct {
Size uintptr
RequestID uintptr
LanguageCode uintptr
InterimResults bool
EnableWordTimeOffsets bool
EnableAutomaticPunctuation bool
VerbatimTranscripts bool
ProfanityFilter bool
_ [3]byte
StopHistoryEouMs int32
_ [4]byte
SpeechContexts uintptr
SpeechContextCount uintptr
MaxAlternatives int32
EnableSpeakerDiarization bool
_ [3]byte
MaxSpeakerCount int32
_ [4]byte
}
// cDiarModelConfig mirrors nemo_speech_diar_model_config (diar.h). This is the
// standalone Sortformer pipeline's own config and is NOT cASRDiarConfig, which
// is the diarizer attached to a recognizer: this one carries gpu and preset,
// that one does not.
//
// The six frame counts are sentinel-sensitive. src/asr/c_api.cpp applies each
// one only when it is > 0, EXCEPT left_context_frames, which it applies when it
// is >= 0. A zero-valued struct would therefore pin the left context to 0
// rather than leave the preset's value alone, so loadDiarizer writes -1 into
// all six.
type cDiarModelConfig struct {
Size uintptr
ModelPath uintptr
GPU int32
_ [4]byte // pad to the alignment of the pointer that follows
Preset uintptr
// Encoder-frame geometry overrides, applied on top of the preset.
ChunkFrames int32
RightContextFrames int32
LeftContextFrames int32
FIFOFrames int32
SpkcacheFrames int32
UpdatePeriodFrames int32
}
// cDiarSegmentationConfig mirrors nemo_speech_diar_segmentation_config
// (diar.h): the NeMo ts_vad postprocessing applied when turning per-frame
// speaker probabilities into segments.
//
// onset and offset are float, the four durations are double. That mixture is
// the whole reason this mirror needs its offsets pinned: writing all six as
// float32 or all six as float64 both produce a struct C would read shifted.
type cDiarSegmentationConfig struct {
Size uintptr
Onset float32
Offset float32
PadOnsetSec float64
PadOffsetSec float64
MinGapSec float64
MinDurationSec float64
}
// cDiarSegment mirrors nemo_speech_diar_segment (diar.h), the element type
// nemo_speech_diar_segments fills.
//
// It has no leading size field: unlike the config structs it travels from C to
// Go, so there is no caller-declared size for the runtime to validate against.
// The times are already SECONDS (double), not frame indices, so nothing here
// needs the model's seconds-per-frame to be interpreted. Speaker is 1-based,
// matching WordInfo.speaker_tag on the ASR surface.
type cDiarSegment struct {
StartTime float64
EndTime float64
Speaker int32
_ [4]byte // trailing pad to the struct's 8-byte alignment
}
type cTTSModelConfig struct {
Size uintptr
MagpieModel uintptr
CodecModel uintptr
TokenizerModelDir uintptr
TextNormalizerModelDir uintptr
}
// cTTSRuntimeConfig mirrors nemo_speech_tts_runtime_config. The four backend /
// mode fields are C enums, which this toolchain lays out as int32.
type cTTSRuntimeConfig struct {
Size uintptr
Speaker int32
Threads int32
CodecThreads int32
Seed int32
Steps int32
TopK int32
ChunkFrames int32
CodecQueueDepth int32
CodecHistoryFrames int32
CodecFutureFrames int32
WindowMs int32
Temperature float32
OverrideTemperature bool
_ [3]byte
CFGScale float32
OverrideCFGScale bool
UseCFG bool
UseLocalTransformer bool
UseKVCache bool
UseStatefulCodec bool
CodecCPU bool
FlushPartialChunk bool
Verbose bool
LTBackend int32
SamplingBackend int32
UMAMode int32
LongformMode int32
LTFP32 bool
_ [7]byte
}
type cTTSSynthesizerConfig struct {
Size uintptr
Model uintptr
Runtime uintptr
DefaultLanguageCode uintptr
DefaultVoiceName uintptr
}
type cTTSSynthesisOptions struct {
Size uintptr
RequestID uintptr
LanguageCode uintptr
Speaker int32
Seed int32
Steps int32
TopK int32
Temperature float32
OverrideTemperature bool
_ [3]byte
CFGScale float32
OverrideCFGScale bool
_ [3]byte
VoiceName uintptr
OutputSampleRate int32
_ [4]byte
}
type cNMTBackendConfig struct {
Size uintptr
GPU int32
_ [4]byte
}
type cNMTModelConfig struct {
Size uintptr
Path uintptr
NCtx int32
_ [4]byte
}
type cNMTTranslatorConfig struct {
Size uintptr
Backend uintptr
Model uintptr
Generation uintptr
Pool uintptr
}
// symbol pairs a Go function pointer with its exported C name. Keeping the
// name next to the var means `nm -D libnemo_speech_asr_c.so.1 | grep nemo_speech`
// is enough to spot drift after a pin bump.
type symbol struct {
fn any
name string
lib *uintptr
}
func symbols() []symbol {
return []symbol{
{&ASRCreate, "nemo_speech_asr_create", &asrLib},
{&ASRDestroy, "nemo_speech_asr_destroy", &asrLib},
{&ASRRecognizeF32, "nemo_speech_asr_recognize_f32", &asrLib},
{&ASRStreamingRecognize, "nemo_speech_asr_streaming_recognize", &asrLib},
{&ASRStreamPushF32, "nemo_speech_asr_stream_push_f32", &asrLib},
{&ASRStreamForceEndpoint, "nemo_speech_asr_stream_force_endpoint", &asrLib},
{&ASRStreamFinish, "nemo_speech_asr_stream_finish", &asrLib},
{&ASRStreamNext, "nemo_speech_asr_stream_next", &asrLib},
{&ASRStreamClose, "nemo_speech_asr_stream_close", &asrLib},
{&ASRRecognitionOptionsDef, "nemo_speech_asr_recognition_options_default", &asrLib},
{&ASRResultIsFinal, "nemo_speech_asr_result_is_final", &asrLib},
{&ASRResultAudioProcessed, "nemo_speech_asr_result_audio_processed", &asrLib},
{&ASRResultAlternativeCount, "nemo_speech_asr_result_alternative_count", &asrLib},
{&ASRResultTranscript, "nemo_speech_asr_result_transcript", &asrLib},
{&ASRResultConfidence, "nemo_speech_asr_result_confidence", &asrLib},
{&ASRResultWordCount, "nemo_speech_asr_result_word_count", &asrLib},
{&ASRResultWordText, "nemo_speech_asr_result_word_text", &asrLib},
{&ASRResultWordStartTime, "nemo_speech_asr_result_word_start_time", &asrLib},
{&ASRResultWordEndTime, "nemo_speech_asr_result_word_end_time", &asrLib},
{&ASRResultWordConfidence, "nemo_speech_asr_result_word_confidence", &asrLib},
{&ASRResultWordSpeakerTag, "nemo_speech_asr_result_word_speaker_tag", &asrLib},
{&ASRResultLanguageCount, "nemo_speech_asr_result_language_count", &asrLib},
{&ASRResultLanguageCode, "nemo_speech_asr_result_language_code", &asrLib},
{&ASRResultDestroy, "nemo_speech_asr_result_destroy", &asrLib},
{&ASRLastError, "nemo_speech_asr_last_error", &asrLib},
{&ASRVersion, "nemo_speech_asr_version", &asrLib},
{&DiarCreate, "nemo_speech_diar_create", &asrLib},
{&DiarDestroy, "nemo_speech_diar_destroy", &asrLib},
{&DiarNumSpeakers, "nemo_speech_diar_num_speakers", &asrLib},
{&DiarSecondsPerFrame, "nemo_speech_diar_seconds_per_frame", &asrLib},
{&DiarStreamOpen, "nemo_speech_diar_stream_open", &asrLib},
{&DiarStreamPushF32, "nemo_speech_diar_stream_push_f32", &asrLib},
{&DiarStreamFinish, "nemo_speech_diar_stream_finish", &asrLib},
{&DiarStreamClose, "nemo_speech_diar_stream_close", &asrLib},
{&DiarSegments, "nemo_speech_diar_segments", &asrLib},
{&TTSCreate, "nemo_speech_tts_create", &ttsLib},
{&TTSDestroy, "nemo_speech_tts_destroy", &ttsLib},
{&TTSSampleRate, "nemo_speech_tts_sample_rate", &ttsLib},
{&TTSSpeakerCount, "nemo_speech_tts_speaker_count", &ttsLib},
{&TTSSpeakerName, "nemo_speech_tts_speaker_name", &ttsLib},
{&TTSSynthesizeText, "nemo_speech_tts_synthesize_text", &ttsLib},
{&TTSRuntimeConfigDefault, "nemo_speech_tts_runtime_config_default", &ttsLib},
{&TTSSynthesisOptionsDefault, "nemo_speech_tts_synthesis_options_default", &ttsLib},
{&TTSLastError, "nemo_speech_tts_last_error", &ttsLib},
{&TTSVersion, "nemo_speech_tts_version", &ttsLib},
{&NMTCreate, "nemo_speech_nmt_create", &nmtLib},
{&NMTDestroy, "nemo_speech_nmt_destroy", &nmtLib},
{&NMTTranslate, "nemo_speech_nmt_translate", &nmtLib},
{&NMTResultCount, "nemo_speech_nmt_result_count", &nmtLib},
{&NMTResultText, "nemo_speech_nmt_result_text", &nmtLib},
{&NMTResultLanguage, "nemo_speech_nmt_result_language", &nmtLib},
{&NMTResultDestroy, "nemo_speech_nmt_result_destroy", &nmtLib},
{&NMTLastError, "nemo_speech_nmt_last_error", &nmtLib},
{&NMTVersion, "nemo_speech_nmt_version", &nmtLib},
}
}
// registerSymbols binds every entry point. purego panics on a missing symbol,
// so this recovers and returns the offending name: after an upstream pin bump a
// rename must fail loudly at startup, not at first inference.
func registerSymbols() error {
for _, s := range symbols() {
if err := registerOne(s); err != nil {
return err
}
}
return nil
}
func registerOne(s symbol) (err error) {
defer func() {
if r := recover(); r != nil {
err = fmt.Errorf("nemo-speech-cpp: binding %q: %v", s.name, r)
}
}()
purego.RegisterLibFunc(s.fn, *s.lib, s.name)
return nil
}

View File

@@ -0,0 +1,317 @@
package main
import (
"os"
"unsafe"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
// requireLibs reports whether a missing shared library must fail the specs
// instead of skipping them.
//
// librariesPresent stats bare filenames relative to the working directory,
// while openLibraries resolves them through the loader search path, so the two
// can legitimately disagree. Under `make test` that is harmless because the
// stage-libs prerequisite puts the .so files in the working directory, but any
// other invocation would skip every library-backed spec and still report a
// green run. The Makefile sets NEMO_SPEECH_REQUIRE_LIBS=1 so no CI path can
// pass on a silent skip; leaving it unset keeps the pure-Go specs runnable on a
// checkout with no build.
func requireLibs() bool {
return os.Getenv("NEMO_SPEECH_REQUIRE_LIBS") == "1"
}
// librariesPresent reports whether a local build is available to bind against.
func librariesPresent() bool {
for _, n := range []string{
libraryName("NEMO_SPEECH_ASR_LIBRARY", "libnemo_speech_asr_c"),
libraryName("NEMO_SPEECH_TTS_LIBRARY", "libnemo_speech_tts"),
libraryName("NEMO_SPEECH_NMT_LIBRARY", "libnemo_speech_nmt_c"),
} {
if _, err := os.Stat(n); err != nil {
return false
}
}
return true
}
// layout is one expected number transcribed from the C headers.
type layout struct {
what string
got uintptr
want uintptr
}
// The `want` column is what a C compiler reports for the structs in
// include/nemo_speech/{asr,diar,tts,nmt}.h under the System V AMD64 / AAPCS64 rules
// both supported targets follow. Regenerate after an upstream pin bump with a
// throwaway program over the installed headers:
//
// printf('SIZE %%zu\n', sizeof(nemo_speech_asr_recognition_options));
// printf('OFF %%zu\n', offsetof(nemo_speech_asr_recognition_options, max_speaker_count));
//
// Sizes alone are not enough: two padding mistakes can cancel out and leave the
// total unchanged while every field between them reads from the wrong offset,
// so each mirror pins its field offsets too.
func structSizes() []layout {
return []layout{
{"cASRBackendConfig", unsafe.Sizeof(cASRBackendConfig{}), 16},
{"cASRModelConfig", unsafe.Sizeof(cASRModelConfig{}), 24},
{"cASRVADConfig", unsafe.Sizeof(cASRVADConfig{}), 32},
{"cASRPostprocConfig", unsafe.Sizeof(cASRPostprocConfig{}), 32},
{"cASRDiarConfig", unsafe.Sizeof(cASRDiarConfig{}), 40},
{"cASRRecognizerConfig", unsafe.Sizeof(cASRRecognizerConfig{}), 80},
{"cASRRecognitionOptions", unsafe.Sizeof(cASRRecognitionOptions{}), 72},
{"cDiarModelConfig", unsafe.Sizeof(cDiarModelConfig{}), 56},
{"cDiarSegmentationConfig", unsafe.Sizeof(cDiarSegmentationConfig{}), 48},
{"cDiarSegment", unsafe.Sizeof(cDiarSegment{}), 24},
{"cTTSModelConfig", unsafe.Sizeof(cTTSModelConfig{}), 40},
{"cTTSRuntimeConfig", unsafe.Sizeof(cTTSRuntimeConfig{}), 96},
{"cTTSSynthesizerConfig", unsafe.Sizeof(cTTSSynthesizerConfig{}), 40},
{"cTTSSynthesisOptions", unsafe.Sizeof(cTTSSynthesisOptions{}), 72},
{"cNMTBackendConfig", unsafe.Sizeof(cNMTBackendConfig{}), 16},
{"cNMTModelConfig", unsafe.Sizeof(cNMTModelConfig{}), 24},
{"cNMTTranslatorConfig", unsafe.Sizeof(cNMTTranslatorConfig{}), 40},
}
}
func structOffsets() []layout {
return []layout{
{"cASRBackendConfig.GPU", unsafe.Offsetof(cASRBackendConfig{}.GPU), 8},
{"cASRModelConfig.Path", unsafe.Offsetof(cASRModelConfig{}.Path), 8},
{"cASRModelConfig.Name", unsafe.Offsetof(cASRModelConfig{}.Name), 16},
{"cASRVADConfig.ModelPath", unsafe.Offsetof(cASRVADConfig{}.ModelPath), 8},
{"cASRVADConfig.EnableMasking", unsafe.Offsetof(cASRVADConfig{}.EnableMasking), 16},
{"cASRVADConfig.Onset", unsafe.Offsetof(cASRVADConfig{}.Onset), 20},
{"cASRVADConfig.Offset", unsafe.Offsetof(cASRVADConfig{}.Offset), 24},
{"cASRPostprocConfig.ProfanityListPath", unsafe.Offsetof(cASRPostprocConfig{}.ProfanityListPath), 8},
{"cASRPostprocConfig.ITNModelDir", unsafe.Offsetof(cASRPostprocConfig{}.ITNModelDir), 16},
{"cASRPostprocConfig.PNCModelPath", unsafe.Offsetof(cASRPostprocConfig{}.PNCModelPath), 24},
{"cASRDiarConfig.ModelPath", unsafe.Offsetof(cASRDiarConfig{}.ModelPath), 8},
{"cASRDiarConfig.ChunkFrames", unsafe.Offsetof(cASRDiarConfig{}.ChunkFrames), 16},
{"cASRDiarConfig.RightContextFrames", unsafe.Offsetof(cASRDiarConfig{}.RightContextFrames), 20},
{"cASRDiarConfig.LeftContextFrames", unsafe.Offsetof(cASRDiarConfig{}.LeftContextFrames), 24},
{"cASRDiarConfig.FIFOFrames", unsafe.Offsetof(cASRDiarConfig{}.FIFOFrames), 28},
{"cASRDiarConfig.SpkcacheFrames", unsafe.Offsetof(cASRDiarConfig{}.SpkcacheFrames), 32},
{"cASRDiarConfig.UpdatePeriodFrames", unsafe.Offsetof(cASRDiarConfig{}.UpdatePeriodFrames), 36},
{"cASRRecognizerConfig.Backend", unsafe.Offsetof(cASRRecognizerConfig{}.Backend), 8},
{"cASRRecognizerConfig.Model", unsafe.Offsetof(cASRRecognizerConfig{}.Model), 16},
{"cASRRecognizerConfig.Streaming", unsafe.Offsetof(cASRRecognizerConfig{}.Streaming), 24},
{"cASRRecognizerConfig.Decoder", unsafe.Offsetof(cASRRecognizerConfig{}.Decoder), 32},
{"cASRRecognizerConfig.VAD", unsafe.Offsetof(cASRRecognizerConfig{}.VAD), 40},
{"cASRRecognizerConfig.Endpointing", unsafe.Offsetof(cASRRecognizerConfig{}.Endpointing), 48},
{"cASRRecognizerConfig.Postproc", unsafe.Offsetof(cASRRecognizerConfig{}.Postproc), 56},
{"cASRRecognizerConfig.Diar", unsafe.Offsetof(cASRRecognizerConfig{}.Diar), 64},
{"cASRRecognizerConfig.Batching", unsafe.Offsetof(cASRRecognizerConfig{}.Batching), 72},
{"cASRRecognitionOptions.RequestID", unsafe.Offsetof(cASRRecognitionOptions{}.RequestID), 8},
{"cASRRecognitionOptions.LanguageCode", unsafe.Offsetof(cASRRecognitionOptions{}.LanguageCode), 16},
{"cASRRecognitionOptions.InterimResults", unsafe.Offsetof(cASRRecognitionOptions{}.InterimResults), 24},
{"cASRRecognitionOptions.EnableWordTimeOffsets", unsafe.Offsetof(cASRRecognitionOptions{}.EnableWordTimeOffsets), 25},
{"cASRRecognitionOptions.EnableAutomaticPunctuation", unsafe.Offsetof(cASRRecognitionOptions{}.EnableAutomaticPunctuation), 26},
{"cASRRecognitionOptions.VerbatimTranscripts", unsafe.Offsetof(cASRRecognitionOptions{}.VerbatimTranscripts), 27},
{"cASRRecognitionOptions.ProfanityFilter", unsafe.Offsetof(cASRRecognitionOptions{}.ProfanityFilter), 28},
{"cASRRecognitionOptions.StopHistoryEouMs", unsafe.Offsetof(cASRRecognitionOptions{}.StopHistoryEouMs), 32},
{"cASRRecognitionOptions.SpeechContexts", unsafe.Offsetof(cASRRecognitionOptions{}.SpeechContexts), 40},
{"cASRRecognitionOptions.SpeechContextCount", unsafe.Offsetof(cASRRecognitionOptions{}.SpeechContextCount), 48},
{"cASRRecognitionOptions.MaxAlternatives", unsafe.Offsetof(cASRRecognitionOptions{}.MaxAlternatives), 56},
{"cASRRecognitionOptions.EnableSpeakerDiarization", unsafe.Offsetof(cASRRecognitionOptions{}.EnableSpeakerDiarization), 60},
{"cASRRecognitionOptions.MaxSpeakerCount", unsafe.Offsetof(cASRRecognitionOptions{}.MaxSpeakerCount), 64},
{"cDiarModelConfig.ModelPath", unsafe.Offsetof(cDiarModelConfig{}.ModelPath), 8},
{"cDiarModelConfig.GPU", unsafe.Offsetof(cDiarModelConfig{}.GPU), 16},
{"cDiarModelConfig.Preset", unsafe.Offsetof(cDiarModelConfig{}.Preset), 24},
{"cDiarModelConfig.ChunkFrames", unsafe.Offsetof(cDiarModelConfig{}.ChunkFrames), 32},
{"cDiarModelConfig.RightContextFrames", unsafe.Offsetof(cDiarModelConfig{}.RightContextFrames), 36},
{"cDiarModelConfig.LeftContextFrames", unsafe.Offsetof(cDiarModelConfig{}.LeftContextFrames), 40},
{"cDiarModelConfig.FIFOFrames", unsafe.Offsetof(cDiarModelConfig{}.FIFOFrames), 44},
{"cDiarModelConfig.SpkcacheFrames", unsafe.Offsetof(cDiarModelConfig{}.SpkcacheFrames), 48},
{"cDiarModelConfig.UpdatePeriodFrames", unsafe.Offsetof(cDiarModelConfig{}.UpdatePeriodFrames), 52},
{"cDiarSegmentationConfig.Onset", unsafe.Offsetof(cDiarSegmentationConfig{}.Onset), 8},
{"cDiarSegmentationConfig.Offset", unsafe.Offsetof(cDiarSegmentationConfig{}.Offset), 12},
{"cDiarSegmentationConfig.PadOnsetSec", unsafe.Offsetof(cDiarSegmentationConfig{}.PadOnsetSec), 16},
{"cDiarSegmentationConfig.PadOffsetSec", unsafe.Offsetof(cDiarSegmentationConfig{}.PadOffsetSec), 24},
{"cDiarSegmentationConfig.MinGapSec", unsafe.Offsetof(cDiarSegmentationConfig{}.MinGapSec), 32},
{"cDiarSegmentationConfig.MinDurationSec", unsafe.Offsetof(cDiarSegmentationConfig{}.MinDurationSec), 40},
{"cDiarSegment.StartTime", unsafe.Offsetof(cDiarSegment{}.StartTime), 0},
{"cDiarSegment.EndTime", unsafe.Offsetof(cDiarSegment{}.EndTime), 8},
{"cDiarSegment.Speaker", unsafe.Offsetof(cDiarSegment{}.Speaker), 16},
{"cTTSModelConfig.MagpieModel", unsafe.Offsetof(cTTSModelConfig{}.MagpieModel), 8},
{"cTTSModelConfig.CodecModel", unsafe.Offsetof(cTTSModelConfig{}.CodecModel), 16},
{"cTTSModelConfig.TokenizerModelDir", unsafe.Offsetof(cTTSModelConfig{}.TokenizerModelDir), 24},
{"cTTSModelConfig.TextNormalizerModelDir", unsafe.Offsetof(cTTSModelConfig{}.TextNormalizerModelDir), 32},
{"cTTSRuntimeConfig.Speaker", unsafe.Offsetof(cTTSRuntimeConfig{}.Speaker), 8},
{"cTTSRuntimeConfig.Threads", unsafe.Offsetof(cTTSRuntimeConfig{}.Threads), 12},
{"cTTSRuntimeConfig.CodecThreads", unsafe.Offsetof(cTTSRuntimeConfig{}.CodecThreads), 16},
{"cTTSRuntimeConfig.Seed", unsafe.Offsetof(cTTSRuntimeConfig{}.Seed), 20},
{"cTTSRuntimeConfig.Steps", unsafe.Offsetof(cTTSRuntimeConfig{}.Steps), 24},
{"cTTSRuntimeConfig.TopK", unsafe.Offsetof(cTTSRuntimeConfig{}.TopK), 28},
{"cTTSRuntimeConfig.ChunkFrames", unsafe.Offsetof(cTTSRuntimeConfig{}.ChunkFrames), 32},
{"cTTSRuntimeConfig.CodecQueueDepth", unsafe.Offsetof(cTTSRuntimeConfig{}.CodecQueueDepth), 36},
{"cTTSRuntimeConfig.CodecHistoryFrames", unsafe.Offsetof(cTTSRuntimeConfig{}.CodecHistoryFrames), 40},
{"cTTSRuntimeConfig.CodecFutureFrames", unsafe.Offsetof(cTTSRuntimeConfig{}.CodecFutureFrames), 44},
{"cTTSRuntimeConfig.WindowMs", unsafe.Offsetof(cTTSRuntimeConfig{}.WindowMs), 48},
{"cTTSRuntimeConfig.Temperature", unsafe.Offsetof(cTTSRuntimeConfig{}.Temperature), 52},
{"cTTSRuntimeConfig.OverrideTemperature", unsafe.Offsetof(cTTSRuntimeConfig{}.OverrideTemperature), 56},
{"cTTSRuntimeConfig.CFGScale", unsafe.Offsetof(cTTSRuntimeConfig{}.CFGScale), 60},
{"cTTSRuntimeConfig.OverrideCFGScale", unsafe.Offsetof(cTTSRuntimeConfig{}.OverrideCFGScale), 64},
{"cTTSRuntimeConfig.UseCFG", unsafe.Offsetof(cTTSRuntimeConfig{}.UseCFG), 65},
{"cTTSRuntimeConfig.UseLocalTransformer", unsafe.Offsetof(cTTSRuntimeConfig{}.UseLocalTransformer), 66},
{"cTTSRuntimeConfig.UseKVCache", unsafe.Offsetof(cTTSRuntimeConfig{}.UseKVCache), 67},
{"cTTSRuntimeConfig.UseStatefulCodec", unsafe.Offsetof(cTTSRuntimeConfig{}.UseStatefulCodec), 68},
{"cTTSRuntimeConfig.CodecCPU", unsafe.Offsetof(cTTSRuntimeConfig{}.CodecCPU), 69},
{"cTTSRuntimeConfig.FlushPartialChunk", unsafe.Offsetof(cTTSRuntimeConfig{}.FlushPartialChunk), 70},
{"cTTSRuntimeConfig.Verbose", unsafe.Offsetof(cTTSRuntimeConfig{}.Verbose), 71},
{"cTTSRuntimeConfig.LTBackend", unsafe.Offsetof(cTTSRuntimeConfig{}.LTBackend), 72},
{"cTTSRuntimeConfig.SamplingBackend", unsafe.Offsetof(cTTSRuntimeConfig{}.SamplingBackend), 76},
{"cTTSRuntimeConfig.UMAMode", unsafe.Offsetof(cTTSRuntimeConfig{}.UMAMode), 80},
{"cTTSRuntimeConfig.LongformMode", unsafe.Offsetof(cTTSRuntimeConfig{}.LongformMode), 84},
{"cTTSRuntimeConfig.LTFP32", unsafe.Offsetof(cTTSRuntimeConfig{}.LTFP32), 88},
{"cTTSSynthesizerConfig.Model", unsafe.Offsetof(cTTSSynthesizerConfig{}.Model), 8},
{"cTTSSynthesizerConfig.Runtime", unsafe.Offsetof(cTTSSynthesizerConfig{}.Runtime), 16},
{"cTTSSynthesizerConfig.DefaultLanguageCode", unsafe.Offsetof(cTTSSynthesizerConfig{}.DefaultLanguageCode), 24},
{"cTTSSynthesizerConfig.DefaultVoiceName", unsafe.Offsetof(cTTSSynthesizerConfig{}.DefaultVoiceName), 32},
{"cTTSSynthesisOptions.RequestID", unsafe.Offsetof(cTTSSynthesisOptions{}.RequestID), 8},
{"cTTSSynthesisOptions.LanguageCode", unsafe.Offsetof(cTTSSynthesisOptions{}.LanguageCode), 16},
{"cTTSSynthesisOptions.Speaker", unsafe.Offsetof(cTTSSynthesisOptions{}.Speaker), 24},
{"cTTSSynthesisOptions.Seed", unsafe.Offsetof(cTTSSynthesisOptions{}.Seed), 28},
{"cTTSSynthesisOptions.Steps", unsafe.Offsetof(cTTSSynthesisOptions{}.Steps), 32},
{"cTTSSynthesisOptions.TopK", unsafe.Offsetof(cTTSSynthesisOptions{}.TopK), 36},
{"cTTSSynthesisOptions.Temperature", unsafe.Offsetof(cTTSSynthesisOptions{}.Temperature), 40},
{"cTTSSynthesisOptions.OverrideTemperature", unsafe.Offsetof(cTTSSynthesisOptions{}.OverrideTemperature), 44},
{"cTTSSynthesisOptions.CFGScale", unsafe.Offsetof(cTTSSynthesisOptions{}.CFGScale), 48},
{"cTTSSynthesisOptions.OverrideCFGScale", unsafe.Offsetof(cTTSSynthesisOptions{}.OverrideCFGScale), 52},
{"cTTSSynthesisOptions.VoiceName", unsafe.Offsetof(cTTSSynthesisOptions{}.VoiceName), 56},
{"cTTSSynthesisOptions.OutputSampleRate", unsafe.Offsetof(cTTSSynthesisOptions{}.OutputSampleRate), 64},
{"cNMTBackendConfig.GPU", unsafe.Offsetof(cNMTBackendConfig{}.GPU), 8},
{"cNMTModelConfig.Path", unsafe.Offsetof(cNMTModelConfig{}.Path), 8},
{"cNMTModelConfig.NCtx", unsafe.Offsetof(cNMTModelConfig{}.NCtx), 16},
{"cNMTTranslatorConfig.Backend", unsafe.Offsetof(cNMTTranslatorConfig{}.Backend), 8},
{"cNMTTranslatorConfig.Model", unsafe.Offsetof(cNMTTranslatorConfig{}.Model), 16},
{"cNMTTranslatorConfig.Generation", unsafe.Offsetof(cNMTTranslatorConfig{}.Generation), 24},
{"cNMTTranslatorConfig.Pool", unsafe.Offsetof(cNMTTranslatorConfig{}.Pool), 32},
}
}
var _ = Describe("C struct mirrors", func() {
// These need no shared object, so they run on any checkout and catch a
// transcription slip the moment it is introduced.
It("matches the C sizeof of every mirrored struct", func() {
for _, l := range structSizes() {
Expect(l.got).To(Equal(l.want), "%s: Go mirror is %d bytes, C says %d", l.what, l.got, l.want)
}
})
It("matches the C offset of every mirrored field", func() {
for _, l := range structOffsets() {
Expect(l.got).To(Equal(l.want), "%s: Go offset %d, C offset %d", l.what, l.got, l.want)
}
})
})
var _ = Describe("C ABI binding", func() {
BeforeEach(func() {
if !librariesPresent() {
if requireLibs() {
cwd, _ := os.Getwd()
Fail("NEMO_SPEECH_REQUIRE_LIBS=1 but the shared libraries are not in " + cwd +
": these specs are the ABI defence and must not be skipped." +
" Run make -C backend/go/nemo-speech-cpp stage-libs")
}
Skip("shared libraries not built, run make in backend/go/nemo-speech-cpp")
}
Expect(openLibraries()).To(Succeed())
})
It("resolves every bound symbol", func() {
Expect(symbols()).ToNot(BeEmpty())
for _, s := range symbols() {
Expect(registerOne(s)).To(Succeed())
}
})
// The library reports its own sizeof through the size field of each
// defaults struct. A Go mirror that disagrees means every field after the
// first divergence is read from the wrong offset, which no compiler or
// linker check would catch. The three structs below are the only ones with
// a defaults entry point, so they are the only ones the runtime can be
// asked about directly.
It("mirrors the C recognition-options struct layout", func() {
def := ASRRecognitionOptionsDef()
Expect(def.Size).To(Equal(unsafe.Sizeof(cASRRecognitionOptions{})),
"cASRRecognitionOptions does not match the C layout")
})
It("mirrors the C TTS runtime-config struct layout", func() {
def := TTSRuntimeConfigDefault()
Expect(def.Size).To(Equal(unsafe.Sizeof(cTTSRuntimeConfig{})),
"cTTSRuntimeConfig does not match the C layout")
})
It("mirrors the C TTS synthesis-options struct layout", func() {
def := TTSSynthesisOptionsDefault()
Expect(def.Size).To(Equal(unsafe.Sizeof(cTTSSynthesisOptions{})),
"cTTSSynthesisOptions does not match the C layout")
})
// A size match alone cannot see a field read from the wrong offset when two
// padding mistakes cancel out, and structOffsets checks the mirrors against
// numbers transcribed by the same hand that wrote them. This spec is the
// only layer independent of that transcription: it reads values back out of
// the running library, so a systematically wrong table cannot hide here.
//
// Deliberately narrow. An earlier version pinned roughly forty default
// values, which would make a legitimate pin bump (threads 4 to 8, or a
// flipped flush_partial_chunk) fail with a message that reads like a layout
// error. What survives is only the values that are contract, not tuning:
//
// - max_alternatives is the single non-zero in an otherwise memset-zero
// struct, and asr.h documents "<= 1 = 1-best (default)". It pins offset
// 56, deep in the tail past the bool run.
// - The synthesis-options run of four -1 sentinels, each documented in
// tts.h as "< 0 = synthesizer default", pins offsets 24 through 36, and
// temperature witnesses that the run stops exactly at offset 40. A
// mirror whose tail is shifted by one field spills a -1 into that zero.
// - Two -1 sentinels at the ends of the runtime config's long int32 run
// pin offset 20 and offset 40 without depending on any tunable.
//
// Sources: src/asr/c_api.cpp nemo_speech_asr_recognition_options_default,
// src/tts/magpietts/runtime.h MagpieRuntimeConfig, src/tts/c_api.cpp
// nemo_speech_tts_synthesis_options_default.
It("reads the documented default values back through the mirrors", func() {
asr := ASRRecognitionOptionsDef()
Expect(asr.MaxAlternatives).To(Equal(int32(1)))
rt := TTSRuntimeConfigDefault()
Expect(rt.Seed).To(Equal(int32(-1)))
Expect(rt.CodecHistoryFrames).To(Equal(int32(-1)))
opt := TTSSynthesisOptionsDefault()
Expect(opt.Speaker).To(Equal(int32(-1)))
Expect(opt.Seed).To(Equal(int32(-1)))
Expect(opt.Steps).To(Equal(int32(-1)))
Expect(opt.TopK).To(Equal(int32(-1)))
Expect(opt.Temperature).To(Equal(float32(0)))
})
It("reports a non-empty version from each library", func() {
Expect(ASRVersion()).ToNot(BeEmpty())
Expect(TTSVersion()).ToNot(BeEmpty())
Expect(NMTVersion()).ToNot(BeEmpty())
})
})

View File

@@ -0,0 +1,374 @@
package main
import (
"context"
"runtime"
"strconv"
"strings"
"time"
"unsafe"
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
"github.com/mudler/xlog"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
)
// asrWord is one decoded word with its millisecond offsets and 1-based speaker
// tag (0 when diarization was not requested).
type asrWord struct {
Text string
Start int32
End int32
Speaker int32
}
// pinPtr pins v for the lifetime of p and returns its address in the uintptr
// form the config structs carry.
//
// The config structs mirror C, so their pointer members are uintptr, which the
// collector does not trace. Everything reachable only through one of them is
// therefore invisible to the GC while C is reading it, exactly as described on
// cstr, and needs the same pin. runtime.KeepAlive would cover collection but
// says nothing about relocation, and the guarantee wanted here is that the
// address C holds stays the address of the object.
func pinPtr[T any](p *runtime.Pinner, v *T) uintptr {
p.Pin(v)
// #nosec G103 -- v is pinned into p on the previous line, so its address is
// stable and traced for as long as p lives; every caller defers p.Unpin only
// after the create call that reads it. One-way, like cstr: nothing converts
// this uintptr back to a pointer.
return uintptr(unsafe.Pointer(v))
}
// asrDiarConfig builds the config for the diarizer attached to a recognizer.
//
// Extracted from loadASR for the same reason diarModelConfig was extracted from
// loadDiarizer: the six frame counts are sentinel-sensitive and invisible to
// every other check in the tree. src/asr/c_api.cpp:151-165 applies five of them
// when they are > 0 but applies left_context_frames when it is >= 0, so a
// dropped -1 does not fall back to the model's own streaming geometry, it pins
// the left context to zero. The struct is the right shape either way, so the
// layout assertions in abi_test.go cannot see it and only a spec on this builder
// can.
//
// diarGeometryDefault is shared with the standalone diarizer rather than
// restated: it is the same sentinel, from the same rule, in the same runtime.
//
// modelPath is a C pointer from cstr, not a Go string, and the caller owns its
// release.
func asrDiarConfig(modelPath uintptr) cASRDiarConfig {
return cASRDiarConfig{
Size: unsafe.Sizeof(cASRDiarConfig{}),
ModelPath: modelPath,
ChunkFrames: diarGeometryDefault,
RightContextFrames: diarGeometryDefault,
LeftContextFrames: diarGeometryDefault,
FIFOFrames: diarGeometryDefault,
SpkcacheFrames: diarGeometryDefault,
UpdatePeriodFrames: diarGeometryDefault,
}
}
// loadASR creates the recognizer, attaching VAD, PnC, ITN and diarization when
// the corresponding options were set.
//
// Every field below is assigned by name against include/nemo_speech/asr.h. The
// sub-configs are optional pointers: a nil one means "library defaults", which
// is why each is populated only when its option was given rather than always
// being attached with empty strings.
//
// Each struct's Size is load-bearing, not decoration. The runtime decides a
// field is present with HAS_FIELD (src/asr/c_api.cpp), which tests the caller's
// size against offsetof(field) + sizeof(field), so a config sent with Size 0
// has every field ignored and the model silently loads with defaults.
//
// This must not take engineMu: Load is its only caller and already holds it.
func (n *NemoSpeech) loadASR(modelFile string) error {
// nemo_speech_asr_create deep-copies every const char* into a std::string
// (src/asr/c_api.cpp to_config, via str_or_empty) and retains no pointer
// afterwards, so pinning for the duration of the create call is both
// necessary and sufficient.
var pinner runtime.Pinner
defer pinner.Unpin()
pathP, freePath := cstr(modelFile)
defer freePath()
model := cASRModelConfig{Size: unsafe.Sizeof(cASRModelConfig{}), Path: pathP}
backend := cASRBackendConfig{Size: unsafe.Sizeof(cASRBackendConfig{}), GPU: n.opts.gpu}
cfg := cASRRecognizerConfig{
Size: unsafe.Sizeof(cASRRecognizerConfig{}),
Backend: pinPtr(&pinner, &backend),
Model: pinPtr(&pinner, &model),
}
var vad cASRVADConfig
if n.opts.vadModel != "" {
p, free := cstr(n.opts.vadModel)
defer free()
vad = cASRVADConfig{Size: unsafe.Sizeof(cASRVADConfig{}), ModelPath: p}
cfg.VAD = pinPtr(&pinner, &vad)
}
var postproc cASRPostprocConfig
if n.opts.itnDir != "" || n.opts.pncModel != "" {
itnP, freeITN := cstr(n.opts.itnDir)
defer freeITN()
pncP, freePNC := cstr(n.opts.pncModel)
defer freePNC()
postproc = cASRPostprocConfig{
Size: unsafe.Sizeof(cASRPostprocConfig{}),
ITNModelDir: itnP,
PNCModelPath: pncP,
}
cfg.Postproc = pinPtr(&pinner, &postproc)
}
var diar cASRDiarConfig
if n.opts.diarModel != "" {
p, free := cstr(n.opts.diarModel)
defer free()
diar = asrDiarConfig(p)
cfg.Diar = pinPtr(&pinner, &diar)
}
xlog.Info("nemo-speech-cpp: creating recognizer",
"gpu", n.opts.gpu,
"vad", n.opts.vadModel != "",
"pnc", n.opts.pncModel != "",
"itn", n.opts.itnDir != "",
"diarization", n.opts.diarModel != "")
// #nosec G103 -- cfg is a local POD struct passed as a pointer for the
// duration of this call only; every uintptr member it carries is either a
// cstr allocation or a pinPtr address, all pinned above and released by the
// defers, and nemo_speech_asr_create deep-copies and retains nothing.
if st := ASRCreate(unsafe.Pointer(&cfg), &n.recognizer); st != 0 {
return statusErrorf(st, "nemo-speech-cpp: asr create: %s", ASRLastError())
}
return nil
}
// recognizeF32 runs one offline decode and returns the result handle, which the
// caller must destroy.
//
// The empty-input guard is here rather than at the call site because &pcm[0]
// panics on a zero-length slice: Go never reaches the C side's own "empty
// audio" rejection. A silent clip or a truncated upload decodes to zero
// samples, which is ordinary input, not an exotic one.
//
// The caller must hold engineMu.
func recognizeF32(recognizer uintptr, opts *cASRRecognitionOptions, pcm []float32, sampleRate int32) (uintptr, error) {
if len(pcm) == 0 {
return 0, status.Error(codes.InvalidArgument, "nemo-speech-cpp: empty audio")
}
var result uintptr
// #nosec G103 -- opts is the caller's live struct, borrowed for this call
// only; its LanguageCode is a cstr allocation the caller keeps pinned across
// it. &pcm[0] is guarded by the empty check above and the length handed over
// is exactly len(pcm), so the runtime cannot read past the slice.
if st := ASRRecognizeF32(recognizer, unsafe.Pointer(opts),
&pcm[0], uint64(len(pcm)), sampleRate, &result); st != 0 {
return 0, statusErrorf(st, "nemo-speech-cpp: recognize: %s", ASRLastError())
}
return result, nil
}
// msToNanos converts a runtime word offset to the wire unit. The runtime
// reports milliseconds (src/asr/types.h); TranscriptSegment.start/end and
// TranscriptWord.start/end are int64 nanoseconds, which core/backend reads
// straight into a time.Duration.
func msToNanos(ms int32) int64 {
return int64(ms) * int64(time.Millisecond)
}
// extractWords pulls the top alternative's words out of a result handle.
func extractWords(result uintptr) []asrWord {
if ASRResultAlternativeCount(result) == 0 {
return nil
}
count := ASRResultWordCount(result, 0)
words := make([]asrWord, 0, count)
for i := uint64(0); i < count; i++ {
words = append(words, asrWord{
Text: ASRResultWordText(result, 0, i),
Start: ASRResultWordStartTime(result, 0, i),
End: ASRResultWordEndTime(result, 0, i),
Speaker: ASRResultWordSpeakerTag(result, 0, i),
})
}
return words
}
// wordsRequested reports whether the caller asked for word-level timestamps.
// The OpenAI transcription API gates word timings behind
// timestamp_granularities[] containing "word" and defaults to segment level
// otherwise; every backend here follows that contract (see
// backend/go/parakeet-cpp).
func wordsRequested(granularities []string) bool {
for _, g := range granularities {
if strings.EqualFold(strings.TrimSpace(g), "word") {
return true
}
}
return false
}
// wordsToSegments groups words into one segment per consecutive speaker run.
// Without diarization every word carries speaker 0, so this collapses to a
// single segment.
//
// The boundary is a CHANGE of speaker, not the first appearance of one: a
// conversation that returns to an earlier speaker has to start a new turn
// rather than reopen the old one.
//
// withWords additionally attaches the per-word timings that
// core/backend/transcript.go turns into the response's word list. It is off by
// default because the OpenAI contract asks for word timestamps explicitly, and
// a long transcript pays for every word twice otherwise.
func wordsToSegments(words []asrWord, withWords bool) []*pb.TranscriptSegment {
if len(words) == 0 {
return nil
}
var segs []*pb.TranscriptSegment
start := 0
flush := func(end int) {
run := words[start:end]
texts := make([]string, 0, len(run))
for _, w := range run {
texts = append(texts, w.Text)
}
seg := &pb.TranscriptSegment{
// #nosec G115 -- TranscriptSegment.Id is int32 on the wire, and segs
// holds one entry per speaker run over the words of a single decode
// result, which exhausts memory long before it reaches 2^31.
Id: int32(len(segs)),
Text: strings.Join(texts, " "),
Start: msToNanos(run[0].Start),
End: msToNanos(run[len(run)-1].End),
}
// The speaker tag is 1-based with 0 meaning untagged, so an undiarized
// run must stay unlabelled rather than be attributed to a speaker "0".
if run[0].Speaker > 0 {
seg.Speaker = strconv.Itoa(int(run[0].Speaker))
}
if withWords {
seg.Words = wordsToProto(run)
}
segs = append(segs, seg)
}
for i := 1; i < len(words); i++ {
if words[i].Speaker != words[start].Speaker {
flush(i)
start = i
}
}
flush(len(words))
return segs
}
// AudioTranscription decodes the audio at req.Dst and returns one offline
// transcription.
//
// The whole body runs inside withEngine, so the family check and the C calls
// that trust the handle happen under a single acquisition of engineMu. Decoding
// the audio is in there too: pkg/grpc/server.go already serialises RPCs on this
// backend through base.SingleThread, so the lock costs no concurrency, and the
// alternative (check, unlock, decode, relock) is the exact gap Free can land in.
func (n *NemoSpeech) AudioTranscription(ctx context.Context, req *pb.TranscriptRequest) (pb.TranscriptResult, error) {
var out *pb.TranscriptResult
if err := n.withEngine(familyASR, func() error {
r, err := n.transcribe(req)
out = r
return err
}); err != nil {
return pb.TranscriptResult{}, err
}
// transcribe returns a non-nil result whenever it returns a nil error, so
// this cannot fire today. It is a guard rather than a comment because the
// alternative to stating the invariant is a nil dereference in an RPC
// handler if a later edit ever adds a success path that forgets to set it.
if out == nil {
return pb.TranscriptResult{}, status.Error(codes.Internal,
"nemo-speech-cpp: transcription produced no result")
}
// Assembled field by field rather than dereferenced: the RPC signature
// returns the proto message by value, but the message embeds a mutex, so
// copying the struct is a copylocks violation. Every backend in this tree
// gets around it the same way, by only ever returning a composite literal.
return pb.TranscriptResult{
Text: out.Text,
Segments: out.Segments,
Language: out.Language,
Duration: out.Duration,
}, nil
}
// transcribe is AudioTranscription's body. The caller must hold engineMu.
func (n *NemoSpeech) transcribe(req *pb.TranscriptRequest) (*pb.TranscriptResult, error) {
if req.GetDst() == "" {
return nil, status.Error(codes.InvalidArgument,
"nemo-speech-cpp: TranscriptRequest.dst (audio path) is required")
}
pcm, sampleRate, err := decodeAudioMono16k(req.GetDst())
if err != nil {
return nil, status.Errorf(codes.InvalidArgument,
"nemo-speech-cpp: read audio: %v", err)
}
// Rejected here, before anything crosses the ABI, and not only inside
// recognizeF32: a silent or truncated upload decodes to zero samples, and
// there is no point building options and pinning strings for a request
// that cannot produce a transcript. recognizeF32 keeps its own guard as a
// precondition on the function.
if len(pcm) == 0 {
return nil, status.Error(codes.InvalidArgument, "nemo-speech-cpp: empty audio")
}
// A per-request language wins over the model-level default; both may be
// empty, which the runtime reads as auto/model default.
language := req.GetLanguage()
if language == "" {
language = n.opts.languageCode
}
langP, freeLang := cstr(language)
defer freeLang()
opts := ASRRecognitionOptionsDef()
opts.LanguageCode = langP
// Segments are built out of word offsets, so they are always asked for.
opts.EnableWordTimeOffsets = true
// Keyed on the recognizer owning a diar model, not on req.Diarize: asr.h
// documents that a request asking for diarization from a recognizer created
// without one fails with INVALID_ARGUMENT, and setting diar_model is already
// the operator's opt-in.
opts.EnableSpeakerDiarization = n.opts.diarModel != ""
result, err := recognizeF32(n.recognizer, &opts, pcm, sampleRate)
if err != nil {
return nil, err
}
defer ASRResultDestroy(result)
out := &pb.TranscriptResult{
Text: ASRResultTranscript(result, 0),
Segments: wordsToSegments(extractWords(result),
wordsRequested(req.GetTimestampGranularities())),
}
// Multilingual models report what they decided the audio was; monolingual
// ones report nothing, and an empty language is better than echoing back
// whatever the caller guessed.
if ASRResultLanguageCount(result, 0) > 0 {
out.Language = ASRResultLanguageCode(result, 0, 0)
}
if sampleRate > 0 {
out.Duration = float32(len(pcm)) / float32(sampleRate)
}
return out, nil
}

View File

@@ -0,0 +1,526 @@
package main
import (
"context"
"strings"
"unsafe"
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
"github.com/mudler/xlog"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
)
// streamChunkSamples is one push into a streaming session. At 16 kHz mono 1600
// samples is 100 ms, short enough that the decoder is polled often enough to
// see an endpoint promptly and short enough that a cancelled request stops
// within one push.
const streamChunkSamples = 1600
// The rates nemo_speech_asr_stream_push_f32 will resample from (asr.h). Outside
// this range the runtime has nothing to do with the audio, and 0 is NOT
// "unknown": it means "these samples are already at the model rate".
const (
minStreamSampleRate = 8000
maxStreamSampleRate = 96000
// TranscriptLiveConfig.sample_rate documents 0 as 16 kHz, which is a
// different meaning from the C API's 0, so it is resolved before the push.
defaultLiveSampleRate = 16000
)
// streamResult is one result lifted out of C memory. Everything is copied
// before nemo_speech_asr_result_destroy runs, so a streamResult outlives the
// handle it came from.
type streamResult struct {
Text string
Final bool
Words []asrWord
}
// asrSession is the streaming half of the ASR C API, narrowed to the four
// entry points the two streaming RPCs use.
//
// It is an interface because there is no NeMo GGUF small enough to keep in the
// tree, so the loops on top of it (chunking, the need-more-audio drain, the
// live config/reset protocol) would otherwise have no test at all. The seam is
// at the ABI, not at the model: a fake session scripts what the C API returns,
// it does not pretend to transcribe anything.
type asrSession interface {
// push buffers audio. It does not decode; next drives that.
push(pcm []float32, sampleRate int32) error
// finish flushes the decoder tail. The end-of-stream final then comes back
// from next.
finish() error
// next pulls one result. ok=false means the decoder needs more audio,
// which is a pause in the stream and not an error or an end.
next() (result streamResult, ok bool, err error)
close()
}
// sessionOpener creates a session for one language. n.openSession is the
// C-backed implementation.
type sessionOpener func(language string) (asrSession, error)
// cSession is the real asrSession, over one nemo_speech_asr_stream.
type cSession struct {
handle uintptr
}
func (s *cSession) push(pcm []float32, sampleRate int32) error {
// &pcm[0] panics on an empty slice, and an empty frame is ordinary input
// from a live caller: it is a keepalive, not audio.
if len(pcm) == 0 {
return nil
}
if st := ASRStreamPushF32(s.handle, &pcm[0], uint64(len(pcm)), sampleRate); st != 0 {
return statusErrorf(st, "nemo-speech-cpp: stream push: %s", ASRLastError())
}
return nil
}
func (s *cSession) finish() error {
if st := ASRStreamFinish(s.handle); st != 0 {
return statusErrorf(st, "nemo-speech-cpp: stream finish: %s", ASRLastError())
}
return nil
}
func (s *cSession) next() (streamResult, bool, error) {
var handle uintptr
if st := ASRStreamNext(s.handle, &handle); st != 0 {
return streamResult{}, false, statusErrorf(st, "nemo-speech-cpp: stream next: %s", ASRLastError())
}
// OK with a NULL handle is the documented "need more audio". Reading it as
// an error aborts every stream at the first gap; reading it as "keep
// pulling" spins forever.
if handle == 0 {
return streamResult{}, false, nil
}
// Destroyed here rather than by the caller: everything below is copied out
// of C memory into Go values, so nothing survives that would need it, and
// a caller that returned early would otherwise leak the result.
defer ASRResultDestroy(handle)
return streamResult{
Text: ASRResultTranscript(handle, 0),
Final: ASRResultIsFinal(handle),
Words: extractWords(handle),
}, true, nil
}
func (s *cSession) close() { ASRStreamClose(s.handle) }
// openSession starts a streaming recognition on the loaded recognizer.
//
// The caller must hold engineMu.
//
// nemo_speech_asr_streaming_recognize copies the options (src/asr/c_api.cpp
// to_options) and keeps no pointer into them, so the language buffer only has
// to stay pinned across this call, exactly as in loadASR.
func (n *NemoSpeech) openSession(language string) (asrSession, error) {
// A per-request language wins over the model-level default; both may be
// empty, which the runtime reads as auto/model default.
if language == "" {
language = n.opts.languageCode
}
langP, freeLang := cstr(language)
defer freeLang()
opts := ASRRecognitionOptionsDef()
opts.LanguageCode = langP
// Segments and the live word list are built out of word offsets, so they
// are always asked for.
opts.EnableWordTimeOffsets = true
// Keyed on the recognizer owning a diar model rather than on the request:
// asr.h documents that asking a recognizer created without one for
// diarization fails with INVALID_ARGUMENT.
opts.EnableSpeakerDiarization = n.opts.diarModel != ""
// interim_results is left off deliberately. The runtime emits interims from
// next() regardless of it, and they are filtered here rather than
// forwarded: see streamPCM's emit for why the wire contract cannot carry
// them.
var handle uintptr
// #nosec G103 -- opts is a local POD struct borrowed for this call only, and
// its one uintptr member (LanguageCode) is the cstr allocation pinned by the
// deferred freeLang above. to_options copies the struct, so nothing here
// outlives the call.
if st := ASRStreamingRecognize(n.recognizer, unsafe.Pointer(&opts), &handle); st != 0 {
return nil, statusErrorf(st, "nemo-speech-cpp: streaming recognize: %s", ASRLastError())
}
xlog.Debug("nemo-speech-cpp: streaming session open", "language", language)
return &cSession{handle: handle}, nil
}
// chunkPCM slices pcm into fixed-size chunks, leaving the final chunk short
// rather than padding it: silence padding would push audio the caller never
// sent through the encoder and shift the tail word timings.
func chunkPCM(pcm []float32, size int) [][]float32 {
if len(pcm) == 0 {
return nil
}
out := make([][]float32, 0, (len(pcm)+size-1)/size)
for off := 0; off < len(pcm); off += size {
out = append(out, pcm[off:min(off+size, len(pcm))])
}
return out
}
// drain pulls every result the session currently has, handing each to emit.
// It returns when the session reports it needs more audio, which is the loop's
// only terminating condition.
func drain(sess asrSession, emit func(streamResult) error) error {
for {
r, ok, err := sess.next()
if err != nil {
return err
}
if !ok {
return nil
}
if err := emit(r); err != nil {
return err
}
}
}
// streamPCM drives one whole clip through an open session, emitting each
// finalized utterance as a delta and closing with the assembled result.
//
// Only finals become deltas, and the reason is the wire contract:
// TranscriptStreamResponse.delta is newly-FINALIZED text that consumers
// CONCATENATE (core/http/endpoints/openai/transcription.go, and the realtime
// semantic-VAD path). An interim is the decoder's running hypothesis for the
// utterance in flight, so forwarding "he", "hell", "hello", "Hello." would
// assemble to "hehellhelloHello." rather than to the transcript. That the
// runtime also postprocesses finals only (build_result_ in
// src/asr/recognizer.cpp runs ITN and strip_formatting on the final, so it
// rewrites rather than extends the interim) means there is no diffing trick
// that would rescue them either.
//
// The cost is that the first delta of an utterance arrives at its endpoint
// rather than mid-word.
func streamPCM(ctx context.Context, sess asrSession, pcm []float32, sampleRate int32, wantWords bool, results chan<- *pb.TranscriptStreamResponse) error {
if len(pcm) == 0 {
return status.Error(codes.InvalidArgument, "nemo-speech-cpp: empty audio")
}
var (
full strings.Builder
segments []*pb.TranscriptSegment
// sawEndpoint records a final that arrived before the tail flush, i.e.
// a real endpoint rather than the end of the file.
sawEndpoint bool
flushing bool
tailText string
)
emit := func(r streamResult) error {
if !r.Final {
return nil
}
if flushing {
tailText += r.Text
} else {
sawEndpoint = true
}
if r.Text == "" {
return nil
}
// The separator is part of the delta, not added when assembling the
// final text, so concatenating the deltas reproduces FinalResult.Text
// exactly. Utterance transcripts carry no leading or trailing space of
// their own (the runner clears its buffer at each endpoint).
delta := r.Text
if full.Len() > 0 {
delta = " " + delta
}
full.WriteString(delta)
// One segment run per utterance, renumbered into the running sequence.
// wordsToSegments splits a run further on a speaker change, so a
// diarized utterance contributes one segment per turn.
segs := wordsToSegments(r.Words, wantWords)
if len(segs) == 0 {
// Word offsets were requested but a decoder head may still return
// none; a segment carrying just the text beats dropping it.
segs = []*pb.TranscriptSegment{{Text: r.Text}}
}
for _, s := range segs {
// #nosec G115 -- TranscriptSegment.Id is int32 on the wire, and
// segments holds one entry per speaker run per finalized utterance of
// a single request, which exhausts memory long before it reaches 2^31.
s.Id = int32(len(segments))
segments = append(segments, s)
}
results <- &pb.TranscriptStreamResponse{Delta: delta}
return nil
}
for _, chunk := range chunkPCM(pcm, streamChunkSamples) {
// The RPC body holds engineMu for the whole stream, so Free waits on
// it. Without this check a client that disconnected mid-file would pin
// the model against unload until the whole clip had been pushed.
if err := ctx.Err(); err != nil {
return status.Error(codes.Canceled, "nemo-speech-cpp: transcription cancelled")
}
if err := sess.push(chunk, sampleRate); err != nil {
return err
}
if err := drain(sess, emit); err != nil {
return err
}
}
flushing = true
if err := sess.finish(); err != nil {
return err
}
if err := drain(sess, emit); err != nil {
return err
}
final := &pb.TranscriptResult{
Text: full.String(),
Segments: segments,
// The tail flush returns whatever the decoder was still holding.
// Nothing held back after at least one endpoint means the last
// endpoint consumed the audio, which is what "the clip ended on an
// utterance boundary" means here. Text coming back means it ended
// mid-utterance.
Eou: sawEndpoint && tailText == "",
}
if sampleRate > 0 {
final.Duration = float32(len(pcm)) / float32(sampleRate)
}
results <- &pb.TranscriptStreamResponse{FinalResult: final}
return nil
}
// runLive drives one bidirectional live session. The protocol is the one
// documented on the RPC in backend.proto: a Config first, a ready ack once the
// session is open, deltas as utterances finalize, and a terminal result when
// the caller closes its send side.
//
// There is no context here on purpose. The gRPC host closes `in` when the
// stream context is cancelled (pkg/grpc/server.go's recv pump), so ranging
// over it is what stops this loop, and that is also what releases engineMu for
// a waiting Free.
func runLive(open sessionOpener, in <-chan *pb.TranscriptLiveRequest, out chan<- *pb.TranscriptLiveResponse) error {
first, ok := <-in
if !ok {
// The caller closed without sending anything. Nothing was opened, so
// there is nothing to report.
return nil
}
cfg := first.GetConfig()
if cfg == nil {
return status.Error(codes.InvalidArgument,
"nemo-speech-cpp: the first live message must carry a config")
}
rate, err := liveSampleRate(cfg)
if err != nil {
return err
}
sess, err := open(cfg.GetLanguage())
if err != nil {
return err
}
// A mid-stream config replaces sess, so this closes whichever session is
// current when the RPC unwinds.
defer func() { sess.close() }()
// Callers block on the first Recv waiting for this and degrade to
// non-live transcription when it does not arrive, so it goes out before
// any audio is read.
out <- &pb.TranscriptLiveResponse{Ready: true}
var (
full strings.Builder
flushing bool
)
emit := func(r streamResult) error {
// Finals only, for the same reason as streamPCM: an interim is a
// hypothesis the final rewrites, and delta is newly-finalized text.
if !r.Final || (r.Text == "" && len(r.Words) == 0) {
return nil
}
// The separator goes INTO the delta, exactly as in streamPCM, because
// the live consumer is the one that actually concatenates: the realtime
// semantic-VAD path joins the accumulated deltas with the empty string
// and only clears them at a turn reset, never at an endpoint. Adding
// the space when assembling the terminal text instead would make the
// running caption read "one.two." while the committed transcript read
// "one. two.".
delta := r.Text
if delta != "" && full.Len() > 0 {
delta = " " + delta
}
full.WriteString(delta)
out <- &pb.TranscriptLiveResponse{
Delta: delta,
// A final that arrives while audio is still coming IS the model's
// endpoint: the decoder resets its utterance there and the next one
// starts fresh, which is the turn boundary the realtime detector
// waits on. The final that comes back from the tail flush is the
// end of the STREAM, not a user yielding a turn, so it carries no
// eou even though the send side has already closed.
Eou: !flushing,
Words: wordsToProto(r.Words),
}
return nil
}
for req := range in {
switch payload := req.GetPayload().(type) {
case *pb.TranscriptLiveRequest_Config:
// A rate cannot change inside a stream (asr.h) and the decoder
// keeps utterance state, so a reconfigure has to be a fresh
// session rather than a reconfigured one.
newRate, err := liveSampleRate(payload.Config)
if err != nil {
return err
}
// Opened before the old one is closed so a failure here leaves a
// live session for the deferred close, not a dangling handle.
next, err := open(payload.Config.GetLanguage())
if err != nil {
return err
}
sess.close()
sess, rate = next, newRate
full.Reset()
case *pb.TranscriptLiveRequest_Audio:
pcm := payload.Audio.GetPcm()
if len(pcm) == 0 {
continue
}
if err := sess.push(pcm, rate); err != nil {
return err
}
if err := drain(sess, emit); err != nil {
return err
}
}
}
// Send side closed: flush the tail and emit the terminal result. Like the
// other backends' live path this carries Text only; per-utterance segments
// and the duration are the file path's concern.
flushing = true
if err := sess.finish(); err != nil {
return err
}
if err := drain(sess, emit); err != nil {
return err
}
// Not trimmed: the terminal text is the verbatim concatenation of the
// deltas, which is the invariant the concatenating consumers rely on. The
// first delta never carries the separator, so there is no leading space to
// trim off in the first place.
out <- &pb.TranscriptLiveResponse{
FinalResult: &pb.TranscriptResult{Text: full.String()},
}
return nil
}
// liveSampleRate resolves TranscriptLiveConfig.sample_rate to the rate the C
// API is given. The proto's 0 means 16 kHz; the C API's 0 means "already at the
// model rate", so the two cannot be forwarded to each other.
func liveSampleRate(cfg *pb.TranscriptLiveConfig) (int32, error) {
rate := cfg.GetSampleRate()
if rate == 0 {
return defaultLiveSampleRate, nil
}
if rate < minStreamSampleRate || rate > maxStreamSampleRate {
return 0, status.Errorf(codes.InvalidArgument,
"nemo-speech-cpp: unsupported live sample_rate %d (accepted: 0 or %d-%d Hz)",
rate, minStreamSampleRate, maxStreamSampleRate)
}
return rate, nil
}
// wordsToProto converts decoded words to the wire form. TranscriptWord.start
// and .end are int64 nanoseconds; the runtime reports milliseconds.
func wordsToProto(words []asrWord) []*pb.TranscriptWord {
if len(words) == 0 {
return nil
}
out := make([]*pb.TranscriptWord, len(words))
for i, w := range words {
out[i] = &pb.TranscriptWord{
Text: w.Text,
Start: msToNanos(w.Start),
End: msToNanos(w.End),
}
}
return out
}
// AudioTranscriptionStream decodes the audio at req.Dst through the streaming
// recognizer, emitting each finalized utterance as it lands.
//
// The body runs inside withEngine for the reason documented on withEngine, and
// that holds engineMu for the whole stream: Free waits rather than destroying
// the recognizer under a half-finished stream. streamPCM honours ctx so the
// wait is bounded by the client's disconnect rather than by its silence.
func (n *NemoSpeech) AudioTranscriptionStream(ctx context.Context, req *pb.TranscriptRequest, results chan *pb.TranscriptStreamResponse) error {
// The host ranges over this channel and only returns once it closes, so
// every path out of here, rejection included, has to close it.
defer close(results)
return n.withEngine(familyASR, func() error {
return n.transcribeStream(ctx, req, results)
})
}
// transcribeStream is AudioTranscriptionStream's body. The caller must hold
// engineMu.
func (n *NemoSpeech) transcribeStream(ctx context.Context, req *pb.TranscriptRequest, results chan<- *pb.TranscriptStreamResponse) error {
if req.GetDst() == "" {
return status.Error(codes.InvalidArgument,
"nemo-speech-cpp: TranscriptRequest.dst (audio path) is required")
}
// Checked before the decode so a client that has already gone away does
// not pay for an ffmpeg run, and so a cancellation is never reported as a
// broken file.
if err := ctx.Err(); err != nil {
return status.Error(codes.Canceled, "nemo-speech-cpp: transcription cancelled")
}
pcm, sampleRate, err := decodeAudioMono16k(req.GetDst())
if err != nil {
return status.Errorf(codes.InvalidArgument, "nemo-speech-cpp: read audio: %v", err)
}
// Before the session is opened, for the same reason as the offline path:
// there is no transcript to be had from zero samples, and opening a stream
// only to close it again asks the runtime to allocate decoder state for
// nothing.
if len(pcm) == 0 {
return status.Error(codes.InvalidArgument, "nemo-speech-cpp: empty audio")
}
sess, err := n.openSession(req.GetLanguage())
if err != nil {
return err
}
defer sess.close()
return streamPCM(ctx, sess, pcm, sampleRate,
wordsRequested(req.GetTimestampGranularities()), results)
}
// AudioTranscriptionLive serves the bidirectional live RPC over one streaming
// session. See runLive for the protocol and withEngine for the locking.
func (n *NemoSpeech) AudioTranscriptionLive(in <-chan *pb.TranscriptLiveRequest, out chan<- *pb.TranscriptLiveResponse) error {
defer close(out)
return n.withEngine(familyASR, func() error {
return runLive(n.openSession, in, out)
})
}

View File

@@ -0,0 +1,699 @@
package main
import (
"context"
"errors"
"path/filepath"
"time"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
)
// fakeSession is a scripted asrSession. It stands in for the streaming C API,
// not for a model: no NeMo GGUF is small enough to keep in the tree, and the
// need-more-audio drain is the easiest thing in this file to get subtly wrong
// (a mishandled NULL either spins forever or drops every result).
//
// script is one batch of results per drain. next() hands back the current
// batch one result at a time and then reports "need more audio" exactly once,
// which advances to the next batch. That is precisely the C contract:
// nemo_speech_asr_stream_next returns OK with a NULL handle when the decoder
// has consumed the buffered audio, and the loop must resume after the next
// push rather than treat it as the end of the stream.
type fakeSession struct {
script [][]streamResult
batch int
pos int
pushed [][]float32
rates []int32
finished int
closed int
pushErr error
finishErr error
nextErr error
}
func (f *fakeSession) push(pcm []float32, sampleRate int32) error {
if f.pushErr != nil {
return f.pushErr
}
f.pushed = append(f.pushed, pcm)
f.rates = append(f.rates, sampleRate)
return nil
}
func (f *fakeSession) finish() error {
if f.finishErr != nil {
return f.finishErr
}
f.finished++
return nil
}
func (f *fakeSession) next() (streamResult, bool, error) {
if f.nextErr != nil {
return streamResult{}, false, f.nextErr
}
if f.batch >= len(f.script) {
return streamResult{}, false, nil
}
if f.pos >= len(f.script[f.batch]) {
f.batch++
f.pos = 0
return streamResult{}, false, nil
}
r := f.script[f.batch][f.pos]
f.pos++
return r, true, nil
}
func (f *fakeSession) close() { f.closed++ }
// samples returns the flat concatenation of everything pushed, so a spec can
// assert the whole clip reached the engine without caring how it was sliced.
func (f *fakeSession) samples() []float32 {
var out []float32
for _, c := range f.pushed {
out = append(out, c...)
}
return out
}
// collect drains a response channel into a slice. The channels are unbuffered
// in the specs on purpose: a producer that stops honouring cancellation would
// otherwise fill a buffer and look healthy.
func collect[T any](ch chan T) chan []T {
done := make(chan []T, 1)
go func() {
var got []T
for v := range ch {
got = append(got, v)
}
done <- got
}()
return done
}
var _ = Describe("chunkPCM", func() {
It("splits into equal chunks when evenly divisible", func() {
chunks := chunkPCM(make([]float32, 400), 100)
Expect(chunks).To(HaveLen(4))
for _, c := range chunks {
Expect(c).To(HaveLen(100))
}
})
// Padding the tail with silence would push phantom audio through the
// encoder and shift the tail word timings, so the final chunk stays short.
It("makes the final chunk short rather than padding it", func() {
chunks := chunkPCM(make([]float32, 250), 100)
Expect(chunks).To(HaveLen(3))
Expect(chunks[2]).To(HaveLen(50))
})
It("returns one chunk when the input is shorter than the chunk size", func() {
chunks := chunkPCM(make([]float32, 10), 100)
Expect(chunks).To(HaveLen(1))
Expect(chunks[0]).To(HaveLen(10))
})
It("returns nothing for empty input", func() {
Expect(chunkPCM(nil, 100)).To(BeEmpty())
Expect(chunkPCM([]float32{}, 100)).To(BeEmpty())
})
// Every spec above works on all-zero audio, so none of them can tell a
// correct slicing from one that reorders or repeats windows. Audio fed out
// of order still decodes, it just decodes to nonsense.
It("preserves sample order across the chunk boundaries", func() {
pcm := []float32{1, 2, 3, 4, 5}
chunks := chunkPCM(pcm, 2)
Expect(chunks).To(HaveLen(3))
Expect(chunks[0]).To(Equal([]float32{1, 2}))
Expect(chunks[1]).To(Equal([]float32{3, 4}))
Expect(chunks[2]).To(Equal([]float32{5}))
})
})
var _ = Describe("drain", func() {
It("emits every result in a batch and stops on need-more-audio", func() {
sess := &fakeSession{script: [][]streamResult{
{{Text: "a"}, {Text: "b", Final: true}},
{{Text: "c"}},
}}
var got []string
Expect(drain(sess, func(r streamResult) error {
got = append(got, r.Text)
return nil
})).To(Succeed())
Expect(got).To(Equal([]string{"a", "b"}))
})
// The NULL handle is a pause, not an end: the next drain, after more audio
// has been pushed, must pick the stream back up.
It("resumes on the next drain after a need-more-audio pause", func() {
sess := &fakeSession{script: [][]streamResult{{{Text: "a"}}, {{Text: "b"}}}}
var got []string
emit := func(r streamResult) error { got = append(got, r.Text); return nil }
Expect(drain(sess, emit)).To(Succeed())
Expect(drain(sess, emit)).To(Succeed())
Expect(got).To(Equal([]string{"a", "b"}))
})
It("returns nothing and no error for a stream with no results ready", func() {
var got []string
Expect(drain(&fakeSession{}, func(r streamResult) error {
got = append(got, r.Text)
return nil
})).To(Succeed())
Expect(got).To(BeEmpty())
})
It("propagates a failure from the runtime", func() {
sess := &fakeSession{nextErr: errors.New("boom")}
Expect(drain(sess, func(streamResult) error { return nil })).To(MatchError(ContainSubstring("boom")))
})
It("stops pulling once emit fails", func() {
sess := &fakeSession{script: [][]streamResult{{{Text: "a"}, {Text: "b"}}}}
Expect(drain(sess, func(streamResult) error {
return errors.New("send failed")
})).To(MatchError(ContainSubstring("send failed")))
Expect(sess.pos).To(Equal(1))
})
})
var _ = Describe("streamPCM", func() {
streamWords := func(ctx context.Context, sess asrSession, pcm []float32, rate int32, wantWords bool) ([]*pb.TranscriptStreamResponse, error) {
GinkgoHelper()
results := make(chan *pb.TranscriptStreamResponse)
done := collect(results)
err := streamPCM(ctx, sess, pcm, rate, wantWords, results)
close(results)
return <-done, err
}
stream := func(ctx context.Context, sess asrSession, pcm []float32, rate int32) ([]*pb.TranscriptStreamResponse, error) {
GinkgoHelper()
return streamWords(ctx, sess, pcm, rate, false)
}
It("pushes the whole clip in chunks at the clip's own sample rate", func() {
sess := &fakeSession{}
pcm := make([]float32, streamChunkSamples*2+7)
_, err := stream(context.Background(), sess, pcm, 16000)
Expect(err).ToNot(HaveOccurred())
Expect(sess.pushed).To(HaveLen(3))
Expect(sess.samples()).To(HaveLen(len(pcm)))
for _, r := range sess.rates {
Expect(r).To(Equal(int32(16000)))
}
})
It("finishes the stream once, after the last chunk", func() {
sess := &fakeSession{}
_, err := stream(context.Background(), sess, make([]float32, 10), 16000)
Expect(err).ToNot(HaveOccurred())
Expect(sess.finished).To(Equal(1))
})
// Interims are the decoder's running hypothesis for the utterance in
// flight. The wire contract is that delta is newly FINALIZED text and that
// concatenating the deltas reproduces the transcript, so forwarding an
// interim would duplicate every word it later re-sends inside the final.
It("emits a delta per final and nothing for interims", func() {
sess := &fakeSession{script: [][]streamResult{{
{Text: "hel"},
{Text: "hello"},
{Text: "Hello.", Final: true},
}}}
got, err := stream(context.Background(), sess, make([]float32, 10), 16000)
Expect(err).ToNot(HaveOccurred())
var deltas []string
for _, r := range got {
if r.GetDelta() != "" {
deltas = append(deltas, r.GetDelta())
}
}
Expect(deltas).To(Equal([]string{"Hello."}))
})
It("reproduces the final transcript by concatenating the deltas", func() {
sess := &fakeSession{script: [][]streamResult{{
{Text: "One.", Final: true},
{Text: "Two.", Final: true},
}}}
got, err := stream(context.Background(), sess, make([]float32, 10), 16000)
Expect(err).ToNot(HaveOccurred())
var joined string
var final *pb.TranscriptResult
for _, r := range got {
joined += r.GetDelta()
if r.GetFinalResult() != nil {
final = r.GetFinalResult()
}
}
Expect(final).ToNot(BeNil())
Expect(final.GetText()).To(Equal("One. Two."))
Expect(joined).To(Equal(final.GetText()))
})
It("sends the terminal final result last and only once", func() {
sess := &fakeSession{script: [][]streamResult{{{Text: "hi", Final: true}}}}
got, err := stream(context.Background(), sess, make([]float32, 10), 16000)
Expect(err).ToNot(HaveOccurred())
Expect(got).ToNot(BeEmpty())
var finals int
for _, r := range got {
if r.GetFinalResult() != nil {
finals++
}
}
Expect(finals).To(Equal(1))
Expect(got[len(got)-1].GetFinalResult()).ToNot(BeNil())
})
It("reports the clip duration in seconds", func() {
sess := &fakeSession{}
got, err := stream(context.Background(), sess, make([]float32, 8000), 16000)
Expect(err).ToNot(HaveOccurred())
Expect(got[len(got)-1].GetFinalResult().GetDuration()).To(BeNumerically("~", 0.5, 1e-6))
})
It("builds per-utterance segments with nanosecond timestamps", func() {
sess := &fakeSession{script: [][]streamResult{{
{Text: "one", Final: true, Words: []asrWord{{Text: "one", Start: 0, End: 500}}},
{Text: "two", Final: true, Words: []asrWord{{Text: "two", Start: 900, End: 1400}}},
}}}
got, err := stream(context.Background(), sess, make([]float32, 10), 16000)
Expect(err).ToNot(HaveOccurred())
segs := got[len(got)-1].GetFinalResult().GetSegments()
Expect(segs).To(HaveLen(2))
Expect(segs[0].GetId()).To(Equal(int32(0)))
Expect(segs[1].GetId()).To(Equal(int32(1)))
Expect(time.Duration(segs[1].GetStart())).To(Equal(900 * time.Millisecond))
Expect(time.Duration(segs[1].GetEnd())).To(Equal(1400 * time.Millisecond))
})
// core/backend/transcript.go builds the response's word list out of
// TranscriptSegment.Words, so leaving it unset makes
// timestamp_granularities: ["word"] come back empty.
It("attaches the word timings only when they were asked for", func() {
script := func() [][]streamResult {
return [][]streamResult{{{Text: "one", Final: true,
Words: []asrWord{{Text: "one", Start: 100, End: 500}}}}}
}
got, err := streamWords(context.Background(), &fakeSession{script: script()}, make([]float32, 10), 16000, true)
Expect(err).ToNot(HaveOccurred())
segs := got[len(got)-1].GetFinalResult().GetSegments()
Expect(segs[0].GetWords()).To(HaveLen(1))
Expect(segs[0].GetWords()[0].GetText()).To(Equal("one"))
Expect(time.Duration(segs[0].GetWords()[0].GetStart())).To(Equal(100 * time.Millisecond))
got, err = streamWords(context.Background(), &fakeSession{script: script()}, make([]float32, 10), 16000, false)
Expect(err).ToNot(HaveOccurred())
segs = got[len(got)-1].GetFinalResult().GetSegments()
Expect(segs[0].GetText()).To(Equal("one"))
Expect(segs[0].GetWords()).To(BeEmpty())
})
// The flush that nemo_speech_asr_stream_finish triggers returns whatever the
// decoder was still holding. Nothing held back means the last endpoint
// consumed the audio, which is exactly "the clip ended on an utterance
// boundary"; text coming back means it ended mid-utterance.
It("marks eou when the tail flush had nothing left to emit", func() {
sess := &fakeSession{script: [][]streamResult{
{{Text: "done.", Final: true}},
{{Text: "", Final: true}},
}}
got, err := stream(context.Background(), sess, make([]float32, 10), 16000)
Expect(err).ToNot(HaveOccurred())
Expect(got[len(got)-1].GetFinalResult().GetEou()).To(BeTrue())
})
It("does not mark eou when the tail flush produced text", func() {
sess := &fakeSession{script: [][]streamResult{
{{Text: "done.", Final: true}},
{{Text: "and more", Final: true}},
}}
got, err := stream(context.Background(), sess, make([]float32, 10), 16000)
Expect(err).ToNot(HaveOccurred())
Expect(got[len(got)-1].GetFinalResult().GetEou()).To(BeFalse())
})
// The RPC body runs inside withEngine, so it holds the engine mutex for the
// whole stream and Free waits on it. A loop that ignored cancellation would
// pin the model against unload for as long as a disconnected client's audio
// takes to push.
It("stops promptly when the request context is cancelled", func() {
ctx, cancel := context.WithCancel(context.Background())
cancel()
sess := &fakeSession{}
_, err := stream(ctx, sess, make([]float32, streamChunkSamples*4), 16000)
Expect(status.Code(err)).To(Equal(codes.Canceled))
Expect(sess.pushed).To(BeEmpty())
})
It("reports a push failure", func() {
sess := &fakeSession{pushErr: errors.New("push blew up")}
_, err := stream(context.Background(), sess, make([]float32, 10), 16000)
Expect(err).To(MatchError(ContainSubstring("push blew up")))
})
It("reports a finish failure", func() {
sess := &fakeSession{finishErr: errors.New("finish blew up")}
_, err := stream(context.Background(), sess, make([]float32, 10), 16000)
Expect(err).To(MatchError(ContainSubstring("finish blew up")))
})
})
var _ = Describe("runLive", func() {
// live drives runLive against a fake opener and returns everything the RPC
// wrote plus the sessions it opened.
live := func(reqs []*pb.TranscriptLiveRequest, script ...[][]streamResult) ([]*pb.TranscriptLiveResponse, []*fakeSession, error) {
GinkgoHelper()
var opened []*fakeSession
open := func(language string) (asrSession, error) {
s := &fakeSession{}
if len(opened) < len(script) {
s.script = script[len(opened)]
}
opened = append(opened, s)
return s, nil
}
in := make(chan *pb.TranscriptLiveRequest)
out := make(chan *pb.TranscriptLiveResponse)
done := collect(out)
go func() {
defer close(in)
for _, r := range reqs {
in <- r
}
}()
err := runLive(open, in, out)
close(out)
return <-done, opened, err
}
cfg := func(rate int32) *pb.TranscriptLiveRequest {
return &pb.TranscriptLiveRequest{Payload: &pb.TranscriptLiveRequest_Config{
Config: &pb.TranscriptLiveConfig{SampleRate: rate},
}}
}
audio := func(pcm ...float32) *pb.TranscriptLiveRequest {
return &pb.TranscriptLiveRequest{Payload: &pb.TranscriptLiveRequest_Audio{
Audio: &pb.TranscriptLiveAudio{Pcm: pcm},
}}
}
It("requires the first message to carry a config", func() {
_, opened, err := live([]*pb.TranscriptLiveRequest{audio(1, 2, 3)})
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(opened).To(BeEmpty())
})
It("returns without error when the caller closes without sending anything", func() {
got, opened, err := live(nil)
Expect(err).ToNot(HaveOccurred())
Expect(got).To(BeEmpty())
Expect(opened).To(BeEmpty())
})
// Callers block on the first Recv waiting for this ack, and degrade to
// non-live transcription when it does not arrive.
It("acknowledges a successful open before any transcript", func() {
got, _, err := live([]*pb.TranscriptLiveRequest{cfg(0)})
Expect(err).ToNot(HaveOccurred())
Expect(got).ToNot(BeEmpty())
Expect(got[0].GetReady()).To(BeTrue())
})
// The proto documents 0 as "16 kHz". The C API reads 0 as "these samples
// are already at the model rate" and skips resampling, so forwarding the
// zero through would silently mean something else.
It("resolves the default sample rate to 16 kHz before pushing", func() {
_, opened, err := live([]*pb.TranscriptLiveRequest{cfg(0), audio(1, 2, 3)})
Expect(err).ToNot(HaveOccurred())
Expect(opened).To(HaveLen(1))
Expect(opened[0].rates).To(Equal([]int32{16000}))
})
It("pushes at the configured sample rate", func() {
_, opened, err := live([]*pb.TranscriptLiveRequest{cfg(8000), audio(1, 2, 3)})
Expect(err).ToNot(HaveOccurred())
Expect(opened[0].rates).To(Equal([]int32{8000}))
Expect(opened[0].samples()).To(Equal([]float32{1, 2, 3}))
})
It("rejects a sample rate the runtime cannot resample", func() {
_, opened, err := live([]*pb.TranscriptLiveRequest{cfg(4000)})
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(opened).To(BeEmpty())
})
It("ignores an empty audio frame instead of pushing it", func() {
_, opened, err := live([]*pb.TranscriptLiveRequest{cfg(0), audio()})
Expect(err).ToNot(HaveOccurred())
Expect(opened[0].pushed).To(BeEmpty())
})
It("streams a delta with its words and marks the utterance boundary", func() {
got, _, err := live(
[]*pb.TranscriptLiveRequest{cfg(0), audio(1)},
[][]streamResult{{
{Text: "partial"},
{Text: "Hello there.", Final: true, Words: []asrWord{
{Text: "Hello", Start: 100, End: 400},
{Text: "there", Start: 400, End: 900},
}},
}},
)
Expect(err).ToNot(HaveOccurred())
var deltas []*pb.TranscriptLiveResponse
for _, r := range got {
if r.GetDelta() != "" {
deltas = append(deltas, r)
}
}
Expect(deltas).To(HaveLen(1))
Expect(deltas[0].GetDelta()).To(Equal("Hello there."))
Expect(deltas[0].GetEou()).To(BeTrue())
Expect(deltas[0].GetWords()).To(HaveLen(2))
Expect(time.Duration(deltas[0].GetWords()[1].GetStart())).To(Equal(400 * time.Millisecond))
Expect(time.Duration(deltas[0].GetWords()[1].GetEnd())).To(Equal(900 * time.Millisecond))
})
It("finishes and closes the session when the caller closes the send side", func() {
got, opened, err := live(
[]*pb.TranscriptLiveRequest{cfg(0), audio(1)},
[][]streamResult{{{Text: "one.", Final: true}}, {{Text: "two.", Final: true}}},
)
Expect(err).ToNot(HaveOccurred())
Expect(opened[0].finished).To(Equal(1))
Expect(opened[0].closed).To(Equal(1))
Expect(got[len(got)-1].GetFinalResult()).ToNot(BeNil())
Expect(got[len(got)-1].GetFinalResult().GetText()).To(Equal("one. two."))
})
// The live path is the one with a consumer that really concatenates: the
// realtime semantic-VAD path joins the accumulated deltas with the empty
// string and clears them only at a turn reset, never at an utterance
// boundary. A separator added when assembling the terminal text instead of
// inside the delta makes the running caption read "one.two." while the
// committed transcript reads "one. two.".
It("reproduces the final transcript by concatenating the deltas", func() {
got, _, err := live(
[]*pb.TranscriptLiveRequest{cfg(0), audio(1)},
[][]streamResult{{{Text: "one.", Final: true}}, {{Text: "two.", Final: true}}},
)
Expect(err).ToNot(HaveOccurred())
var joined string
var final *pb.TranscriptResult
for _, r := range got {
joined += r.GetDelta()
if r.GetFinalResult() != nil {
final = r.GetFinalResult()
}
}
Expect(final).ToNot(BeNil())
Expect(final.GetText()).To(Equal("one. two."))
Expect(joined).To(Equal(final.GetText()))
})
// Eou is the model's endpoint, which is a user yielding the turn. The final
// that comes back from the tail flush is the end of the stream: the send
// side has already closed, so reporting a turn boundary there tells the
// turn detector something that did not happen.
It("marks the endpoint finals but not the tail flush", func() {
got, _, err := live(
[]*pb.TranscriptLiveRequest{cfg(0), audio(1)},
[][]streamResult{{{Text: "one.", Final: true}}, {{Text: "two.", Final: true}}},
)
Expect(err).ToNot(HaveOccurred())
var eous []bool
for _, r := range got {
if r.GetDelta() != "" {
eous = append(eous, r.GetEou())
}
}
Expect(eous).To(Equal([]bool{true, false}))
})
// A rate cannot change inside a stream and the decoder keeps no state
// across a reset, so a second config has to be a fresh session, not a
// reconfigured one.
It("opens a fresh session on a mid-stream config and drops the old transcript", func() {
got, opened, err := live(
[]*pb.TranscriptLiveRequest{cfg(0), audio(1), cfg(0), audio(2)},
[][]streamResult{{{Text: "dropped.", Final: true}}},
[][]streamResult{{{Text: "kept.", Final: true}}},
)
Expect(err).ToNot(HaveOccurred())
Expect(opened).To(HaveLen(2))
Expect(opened[0].closed).To(Equal(1))
Expect(got[len(got)-1].GetFinalResult().GetText()).To(Equal("kept."))
})
It("reports a push failure and still closes the session", func() {
var opened []*fakeSession
open := func(string) (asrSession, error) {
s := &fakeSession{pushErr: errors.New("push blew up")}
opened = append(opened, s)
return s, nil
}
in := make(chan *pb.TranscriptLiveRequest, 2)
in <- cfg(0)
in <- audio(1, 2)
close(in)
out := make(chan *pb.TranscriptLiveResponse, 8)
err := runLive(open, in, out)
Expect(err).To(MatchError(ContainSubstring("push blew up")))
Expect(opened[0].closed).To(Equal(1))
})
It("propagates a failure to open the session", func() {
open := func(string) (asrSession, error) { return nil, errors.New("no streaming here") }
in := make(chan *pb.TranscriptLiveRequest, 1)
in <- cfg(0)
close(in)
out := make(chan *pb.TranscriptLiveResponse, 8)
Expect(runLive(open, in, out)).To(MatchError(ContainSubstring("no streaming here")))
})
})
var _ = Describe("AudioTranscriptionStream", func() {
run := func(ctx context.Context, n *NemoSpeech, req *pb.TranscriptRequest) ([]*pb.TranscriptStreamResponse, error) {
GinkgoHelper()
results := make(chan *pb.TranscriptStreamResponse)
done := collect(results)
err := n.AudioTranscriptionStream(ctx, req, results)
return <-done, err
}
// The RPC owns the channel: the gRPC host ranges over it and only returns
// once it closes, so a rejection path that forgets to close hangs the call
// instead of failing it.
It("closes the results channel on every rejection path", func() {
for _, n := range []*NemoSpeech{{fam: familyTTS}, {fam: familyASR}, {}} {
_, err := run(context.Background(), n, &pb.TranscriptRequest{})
Expect(err).To(HaveOccurred())
Expect(n.engineMu.TryLock()).To(BeTrue())
n.engineMu.Unlock()
}
})
It("refuses a model loaded as another family", func() {
n := &NemoSpeech{fam: familyTTS}
_, err := run(context.Background(), n, &pb.TranscriptRequest{Dst: "x.wav"})
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
Expect(err.Error()).To(ContainSubstring("tts"))
})
It("requires a destination path", func() {
n := &NemoSpeech{fam: familyASR}
_, err := run(context.Background(), n, &pb.TranscriptRequest{})
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
})
// Cancellation is checked before the decode so a client that has already
// gone away does not pay for an ffmpeg run, and so the check cannot be
// mistaken for the decode failing.
It("returns cancelled without touching the audio", func() {
ctx, cancel := context.WithCancel(context.Background())
cancel()
n := &NemoSpeech{fam: familyASR}
_, err := run(ctx, n, &pb.TranscriptRequest{
Dst: filepath.Join(GinkgoT().TempDir(), "absent.wav"),
})
Expect(status.Code(err)).To(Equal(codes.Canceled))
})
It("reports an audio file it cannot read", func() {
n := &NemoSpeech{fam: familyASR}
_, err := run(context.Background(), n, &pb.TranscriptRequest{
Dst: filepath.Join(GinkgoT().TempDir(), "absent.wav"),
})
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
})
// Same ordering constraint as the offline path: a clip that decodes to no
// samples has to be refused before a session is opened, which is also
// before any bound entry point is called. Nothing is loaded here, so a
// guard placed after the open would panic instead of failing.
It("refuses a decodable clip that carries no samples, before opening a session", func() {
path := filepath.Join(GinkgoT().TempDir(), "silence.wav")
writeMono16kWAV(path, 0)
n := &NemoSpeech{fam: familyASR}
var err error
Expect(func() {
_, err = run(context.Background(), n, &pb.TranscriptRequest{Dst: path})
}).ToNot(Panic())
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(err.Error()).To(ContainSubstring("empty audio"))
})
})
var _ = Describe("AudioTranscriptionLive", func() {
It("refuses a model loaded as another family and closes the output", func() {
n := &NemoSpeech{fam: familyNMT}
in := make(chan *pb.TranscriptLiveRequest)
close(in)
out := make(chan *pb.TranscriptLiveResponse)
done := collect(out)
err := n.AudioTranscriptionLive(in, out)
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
Expect(<-done).To(BeEmpty())
Expect(n.engineMu.TryLock()).To(BeTrue())
n.engineMu.Unlock()
})
It("refuses an unloaded model", func() {
n := &NemoSpeech{}
in := make(chan *pb.TranscriptLiveRequest)
close(in)
out := make(chan *pb.TranscriptLiveResponse)
done := collect(out)
err := n.AudioTranscriptionLive(in, out)
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
Expect(<-done).To(BeEmpty())
})
})

View File

@@ -0,0 +1,367 @@
package main
import (
"context"
"math"
"os"
"path/filepath"
"time"
"unsafe"
"github.com/go-audio/audio"
"github.com/go-audio/wav"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
)
// writeMono16kWAV writes `frames` samples of 16 kHz mono 16-bit silence.
// That is already AudioToWav's target format, so the decode path copies the
// file through instead of shelling out to ffmpeg, which the test host may not
// have.
func writeMono16kWAV(path string, frames int) {
GinkgoHelper()
f, err := os.Create(path)
Expect(err).ToNot(HaveOccurred())
enc := wav.NewEncoder(f, 16000, 16, 1, 1)
Expect(enc.Write(&audio.IntBuffer{
Format: &audio.Format{NumChannels: 1, SampleRate: 16000},
SourceBitDepth: 16,
Data: make([]int, frames),
})).To(Succeed())
Expect(enc.Close()).To(Succeed())
Expect(f.Close()).To(Succeed())
}
var _ = Describe("wordsToSegments", func() {
It("groups words into one segment per speaker run", func() {
words := []asrWord{
{Text: "hello", Start: 0, End: 400, Speaker: 1},
{Text: "there", Start: 400, End: 800, Speaker: 1},
{Text: "hi", Start: 900, End: 1200, Speaker: 2},
}
segs := wordsToSegments(words, false)
Expect(segs).To(HaveLen(2))
Expect(segs[0].Text).To(Equal("hello there"))
Expect(segs[1].Text).To(Equal("hi"))
})
// A run is bounded by a CHANGE of speaker, not by the speaker id being new.
// Grouping that keyed on the id itself (a map, or a comparison against the
// first word) would merge the two A turns into one segment spanning B, and
// the three-word spec above cannot see that because it never returns to an
// earlier speaker.
It("starts a new segment when an earlier speaker takes another turn", func() {
words := []asrWord{
{Text: "one", Start: 0, End: 100, Speaker: 1},
{Text: "two", Start: 100, End: 200, Speaker: 2},
{Text: "three", Start: 200, End: 300, Speaker: 1},
}
segs := wordsToSegments(words, false)
Expect(segs).To(HaveLen(3))
Expect(segs[0].Text).To(Equal("one"))
Expect(segs[1].Text).To(Equal("two"))
Expect(segs[2].Text).To(Equal("three"))
})
// TranscriptSegment.start/end are int64 nanoseconds, not seconds:
// core/backend/transcript.go reads them straight into a time.Duration. The
// runtime reports word offsets in milliseconds (src/asr/types.h:46).
It("converts millisecond word times to nanoseconds", func() {
words := []asrWord{{Text: "a", Start: 1500, End: 2250, Speaker: 0}}
segs := wordsToSegments(words, false)
Expect(segs).To(HaveLen(1))
Expect(time.Duration(segs[0].Start)).To(Equal(1500 * time.Millisecond))
Expect(time.Duration(segs[0].End)).To(Equal(2250 * time.Millisecond))
})
It("spans a segment from its first word's start to its last word's end", func() {
words := []asrWord{
{Text: "a", Start: 100, End: 200, Speaker: 0},
{Text: "b", Start: 500, End: 900, Speaker: 0},
}
segs := wordsToSegments(words, false)
Expect(segs).To(HaveLen(1))
Expect(time.Duration(segs[0].Start)).To(Equal(100 * time.Millisecond))
Expect(time.Duration(segs[0].End)).To(Equal(900 * time.Millisecond))
})
It("produces a single segment when no speaker tags are present", func() {
words := []asrWord{
{Text: "a", Start: 0, End: 100, Speaker: 0},
{Text: "b", Start: 100, End: 200, Speaker: 0},
}
segs := wordsToSegments(words, false)
Expect(segs).To(HaveLen(1))
Expect(segs[0].Text).To(Equal("a b"))
})
It("returns no segments for no words", func() {
Expect(wordsToSegments(nil, false)).To(BeEmpty())
Expect(wordsToSegments([]asrWord{}, false)).To(BeEmpty())
})
It("numbers the segments from zero in order", func() {
words := []asrWord{
{Text: "a", Speaker: 1},
{Text: "b", Speaker: 2},
{Text: "c", Speaker: 3},
}
segs := wordsToSegments(words, false)
Expect(segs).To(HaveLen(3))
for i, s := range segs {
Expect(s.Id).To(Equal(int32(i)))
}
})
// TranscriptSegment.Words is what core/backend/transcript.go turns into the
// response's word list, so an unset one makes timestamp_granularities:
// ["word"] come back empty however good the timings were.
It("attaches the per-word timings only when they were asked for", func() {
words := []asrWord{
{Text: "a", Start: 0, End: 100},
{Text: "b", Start: 100, End: 250},
}
with := wordsToSegments(words, true)
Expect(with[0].Words).To(HaveLen(2))
Expect(with[0].Words[1].Text).To(Equal("b"))
Expect(time.Duration(with[0].Words[1].Start)).To(Equal(100 * time.Millisecond))
Expect(time.Duration(with[0].Words[1].End)).To(Equal(250 * time.Millisecond))
Expect(wordsToSegments(words, false)[0].Words).To(BeEmpty())
})
// A speaker change splits the run, and each segment must carry only its own
// words rather than the whole utterance's.
It("gives each speaker run only its own words", func() {
segs := wordsToSegments([]asrWord{
{Text: "a", Speaker: 1},
{Text: "b", Speaker: 2},
}, true)
Expect(segs).To(HaveLen(2))
Expect(segs[0].Words).To(HaveLen(1))
Expect(segs[0].Words[0].Text).To(Equal("a"))
Expect(segs[1].Words[0].Text).To(Equal("b"))
})
// The C ABI documents the speaker tag as 1-based with 0 meaning "untagged",
// so a run of untagged words must not come back attributed to a speaker
// literally named "0".
It("labels a diarized run and leaves an untagged one unlabelled", func() {
Expect(wordsToSegments([]asrWord{{Text: "a", Speaker: 2}}, false)[0].Speaker).To(Equal("2"))
Expect(wordsToSegments([]asrWord{{Text: "a", Speaker: 0}}, false)[0].Speaker).To(BeEmpty())
})
})
var _ = Describe("wordsRequested", func() {
It("recognises the OpenAI word granularity in any casing or padding", func() {
Expect(wordsRequested([]string{"word"})).To(BeTrue())
Expect(wordsRequested([]string{"segment", " Word "})).To(BeTrue())
})
It("defaults to segment level", func() {
Expect(wordsRequested(nil)).To(BeFalse())
Expect(wordsRequested([]string{"segment"})).To(BeFalse())
})
})
var _ = Describe("recognizeF32", func() {
// &pcm[0] panics on a zero-length slice, and a silent or empty upload is
// ordinary input rather than an exotic one. The C side rejects empty audio
// too, but Go never gets that far.
It("refuses empty audio instead of indexing an empty slice", func() {
// A zero options struct is enough: the guard has to fire before the
// options are ever handed across the ABI, and building real ones would
// need the library bound, which this spec deliberately does not.
opts := cASRRecognitionOptions{}
for _, pcm := range [][]float32{nil, {}} {
var (
handle uintptr
err error
)
Expect(func() { handle, err = recognizeF32(0, &opts, pcm, 16000) }).ToNot(Panic())
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(err.Error()).To(ContainSubstring("empty audio"))
Expect(handle).To(BeZero())
}
})
})
var _ = Describe("AudioTranscription", func() {
// The gate has to fire before anything expensive: a model loaded as TTS
// cannot transcribe whatever the request says, and reading the audio first
// would report a file problem for a configuration one.
It("refuses a model loaded as another family, before it reads the audio", func() {
n := &NemoSpeech{fam: familyTTS}
_, err := n.AudioTranscription(context.Background(), &pb.TranscriptRequest{
Dst: filepath.Join(GinkgoT().TempDir(), "does-not-exist.wav"),
})
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
Expect(err.Error()).To(ContainSubstring("tts"))
})
It("refuses an unloaded model", func() {
n := &NemoSpeech{}
_, err := n.AudioTranscription(context.Background(), &pb.TranscriptRequest{Dst: "ignored.wav"})
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
})
// A lock leaked on a rejection path deadlocks the next request rather than
// failing it, which is far harder to diagnose than the failure itself.
It("releases the engine lock on every rejection path", func() {
n := &NemoSpeech{fam: familyTTS}
_, err := n.AudioTranscription(context.Background(), &pb.TranscriptRequest{Dst: "x.wav"})
Expect(err).To(HaveOccurred())
Expect(n.engineMu.TryLock()).To(BeTrue())
n.engineMu.Unlock()
})
It("reports an audio file it cannot read", func() {
n := &NemoSpeech{fam: familyASR}
_, err := n.AudioTranscription(context.Background(), &pb.TranscriptRequest{
Dst: filepath.Join(GinkgoT().TempDir(), "absent.wav"),
})
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(n.engineMu.TryLock()).To(BeTrue())
n.engineMu.Unlock()
})
It("requires a destination path", func() {
n := &NemoSpeech{fam: familyASR}
_, err := n.AudioTranscription(context.Background(), &pb.TranscriptRequest{})
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
})
// The whole rejection path end to end, on the input that actually reaches
// it: a silent or truncated upload decodes to zero samples, and the guard
// has to fire between the decode and the ABI. Nothing is loaded here (no
// recognizer, and the specs that bind the library may not have run), so
// this also pins the ORDER: a guard placed after the options are built
// calls a nil-bound entry point and panics rather than failing.
It("refuses a decodable clip that carries no samples", func() {
path := filepath.Join(GinkgoT().TempDir(), "silence.wav")
writeMono16kWAV(path, 0)
n := &NemoSpeech{fam: familyASR, recognizer: 0}
var err error
Expect(func() {
_, err = n.AudioTranscription(context.Background(), &pb.TranscriptRequest{Dst: path})
}).ToNot(Panic())
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(err.Error()).To(ContainSubstring("empty audio"))
Expect(n.engineMu.TryLock()).To(BeTrue())
n.engineMu.Unlock()
})
})
var _ = Describe("sampleRateOf", func() {
// 0 is not "unknown" to this runtime: nemo_speech_asr_recognize_f32 and
// nemo_speech_asr_stream_push_f32 both read a 0 rate as "these samples are
// already at the model rate" and skip resampling. Falling back to it for an
// undecodable header would silently pitch-shift the audio instead of
// failing, so an unknown rate has to be an error.
It("rejects a buffer whose format the decoder did not fill in", func() {
_, err := sampleRateOf(&audio.IntBuffer{})
Expect(err).To(HaveOccurred())
})
It("rejects a non-positive sample rate", func() {
_, err := sampleRateOf(&audio.IntBuffer{Format: &audio.Format{SampleRate: 0, NumChannels: 1}})
Expect(err).To(HaveOccurred())
})
// The WAV header carries the sample rate as an unsigned 32-bit field, which
// go-audio widens to int. Anything above the int32 range therefore passes a
// "> 0" test and then narrows to a NEGATIVE rate, which the runtime would take
// as a resampling ratio rather than reject. The failure is silent, so the
// bound is asserted rather than left to the caller.
//
// Written as a conversion plus one rather than as the constant MaxInt32+1:
// the untyped form does not fit an int on a 32-bit build and would not
// compile there, while this wraps to a negative rate the same guard rejects.
It("rejects a rate that would not survive the narrowing to int32", func() {
_, err := sampleRateOf(&audio.IntBuffer{
Format: &audio.Format{SampleRate: int(math.MaxInt32) + 1, NumChannels: 1},
})
Expect(err).To(HaveOccurred())
})
It("returns the decoded rate", func() {
rate, err := sampleRateOf(&audio.IntBuffer{Format: &audio.Format{SampleRate: 22050, NumChannels: 1}})
Expect(err).ToNot(HaveOccurred())
Expect(rate).To(Equal(int32(22050)))
})
})
var _ = Describe("decodeAudioMono16k", func() {
It("decodes a 16 kHz mono WAV to float32 samples at its own rate", func() {
path := filepath.Join(GinkgoT().TempDir(), "silence.wav")
writeMono16kWAV(path, 800)
pcm, rate, err := decodeAudioMono16k(path)
Expect(err).ToNot(HaveOccurred())
Expect(rate).To(Equal(int32(16000)))
Expect(pcm).To(HaveLen(800))
})
// A zero-frame WAV is what a truncated upload decodes to, and it is the
// input recognizeF32's guard exists for.
It("decodes a WAV with no frames to an empty slice", func() {
path := filepath.Join(GinkgoT().TempDir(), "empty.wav")
writeMono16kWAV(path, 0)
pcm, _, err := decodeAudioMono16k(path)
Expect(err).ToNot(HaveOccurred())
Expect(pcm).To(BeEmpty())
})
It("reports a file that does not exist", func() {
_, _, err := decodeAudioMono16k(filepath.Join(GinkgoT().TempDir(), "nope.wav"))
Expect(err).To(HaveOccurred())
})
})
// The six frame counts on the recognizer-attached diarizer are
// sentinel-sensitive and invisible to every other check in the tree.
// src/asr/c_api.cpp:151-165 applies five of them when they are > 0 but applies
// left_context_frames when it is >= 0, so a dropped -1 does not fall back to
// the model's own streaming geometry, it pins the left context to zero. The
// struct is the right shape either way, so abi_test.go's layout assertions
// cannot see it.
var _ = Describe("asrDiarConfig", func() {
It("keeps the model path it was given", func() {
Expect(asrDiarConfig(42).ModelPath).To(Equal(uintptr(42)))
})
// A config sent with the wrong size has every field past it ignored by
// HAS_FIELD, and the diarizer attaches with defaults instead of failing.
It("declares the size the runtime validates against", func() {
Expect(asrDiarConfig(42).Size).To(Equal(unsafe.Sizeof(cASRDiarConfig{})))
})
It("leaves every frame count at the sentinel that means default", func() {
cfg := asrDiarConfig(42)
Expect(cfg.ChunkFrames).To(Equal(diarGeometryDefault))
Expect(cfg.RightContextFrames).To(Equal(diarGeometryDefault))
Expect(cfg.LeftContextFrames).To(Equal(diarGeometryDefault))
Expect(cfg.FIFOFrames).To(Equal(diarGeometryDefault))
Expect(cfg.SpkcacheFrames).To(Equal(diarGeometryDefault))
Expect(cfg.UpdatePeriodFrames).To(Equal(diarGeometryDefault))
})
// Stated separately from the field-by-field assertions above: the whole
// group is only "unset" to the runtime while the sentinel stays negative,
// and zero is a value it would apply to the left context.
It("uses a negative sentinel, not zero", func() {
Expect(diarGeometryDefault).To(BeNumerically("<", 0))
})
})

View File

@@ -0,0 +1,86 @@
package main
import (
"errors"
"math"
"os"
"path/filepath"
"github.com/go-audio/audio"
"github.com/go-audio/wav"
"github.com/mudler/LocalAI/pkg/utils"
)
// decodeAudioMono16k converts an arbitrary audio file to 16 kHz mono PCM and
// returns the float32 samples together with the rate they are actually at.
//
// pkg/utils exposes the ffmpeg normalisation (AudioToWav) but no decode, so
// every Go ASR backend pairs it with go-audio itself. This mirrors
// backend/go/parakeet-cpp rather than adding a shared helper: the backends
// differ in what they need back (parakeet wants a duration, this one wants the
// sample rate to hand to the runtime), so a shared signature would be a
// lowest-common-denominator of both.
func decodeAudioMono16k(path string) ([]float32, int32, error) {
dir, err := os.MkdirTemp("", "nemo-speech")
if err != nil {
return nil, 0, err
}
defer func() { _ = os.RemoveAll(dir) }()
// A WAV already at 16 kHz mono 16-bit is hardlinked or copied through
// without spawning ffmpeg, so the common case costs nothing.
converted := filepath.Join(dir, "converted.wav")
if err := utils.AudioToWav(path, converted); err != nil {
return nil, 0, err
}
// #nosec G304 -- converted is filepath.Join of a directory this function just
// created with os.MkdirTemp and a constant basename. The request-controlled
// path is the INPUT to AudioToWav and never reaches this open.
fh, err := os.Open(converted)
if err != nil {
return nil, 0, err
}
defer func() { _ = fh.Close() }()
buf, err := wav.NewDecoder(fh).FullPCMBuffer()
if err != nil {
return nil, 0, err
}
// The rate is read back from the decoded file rather than assumed to be
// 16000. AudioToWav always lands there today, but the runtime resamples
// anything from 8 to 96 kHz off this number, so a wrong one would not fail,
// it would silently pitch-shift the audio and quietly degrade the transcript.
rate, err := sampleRateOf(buf)
if err != nil {
return nil, 0, err
}
return buf.AsFloat32Buffer().Data, rate, nil
}
// sampleRateOf reads the decoded rate back off the buffer.
//
// It is an error rather than a zero fallback because 0 is not "unknown" to this
// runtime: nemo_speech_asr_recognize_f32 and nemo_speech_asr_stream_push_f32
// both read a 0 rate as "these samples are already at the model rate" and skip
// resampling (include/nemo_speech/asr.h). Handing 0 over for a header the
// decoder could not read would not fail, it would silently pitch-shift the
// audio and quietly degrade the transcript, which is the same failure the
// caller comment warns about for a wrong rate.
//
// The upper bound is what makes the narrowing to int32 safe rather than merely
// unlikely. go-audio reads the WAV header's sample rate as an unsigned 32-bit
// field into an int, so on a 64-bit build a header claiming more than 2^31-1
// survives the "> 0" test and then narrows to a NEGATIVE rate, which the runtime
// would take as a resampling ratio. Nothing this backend decodes can reach that
// today (AudioToWav either passes through a WAV it has confirmed is exactly
// 16 kHz or runs ffmpeg with -ar 16000), but that is a property of a helper in
// another package, and this function exists precisely because the rate is read
// back rather than assumed.
func sampleRateOf(buf *audio.IntBuffer) (int32, error) {
if buf.Format == nil || buf.Format.SampleRate <= 0 || buf.Format.SampleRate > math.MaxInt32 {
return 0, errors.New("nemo-speech-cpp: decoded audio has no usable sample rate")
}
return int32(buf.Format.SampleRate), nil
}

View File

@@ -0,0 +1,502 @@
package main
import (
"strconv"
"unsafe"
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
"github.com/mudler/xlog"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
)
// diarSegmentsMaxAttempts bounds the count-then-fill retry.
//
// On a finished stream the count is stable and one attempt is always enough.
// The bound exists because the RPC holds engineMu for its whole body, so a
// runtime whose count kept growing would not merely spin, it would block the
// unload behind it.
const diarSegmentsMaxAttempts = 4
// maxDiarSegments caps the buffer collectSegments will allocate from a count
// the C side reported.
//
// make() panics rather than erroring on a length it cannot satisfy, and a
// panic in an RPC handler takes the backend process down, so an uninitialised
// or corrupted size_t coming back across the ABI would kill the model rather
// than fail the request. The ceiling turns that into a diagnosable error.
//
// It is set far above anything real: a segment spans at least one 80 ms frame,
// so 2^22 segments is upwards of 93 hours of audio, and the buffer itself
// would already be 100 MB at 24 bytes each.
const maxDiarSegments = 1 << 22
// diarSegmenter is the result half of the diarization C API: the two-call
// protocol nemo_speech_diar_segments documents.
//
// The two calls are the same C function with a different `out`, but they are
// separate methods here because their contracts differ. countSegments passes
// out=NULL, which the runtime answers by writing *count and returning OK
// without touching a buffer. fillSegments passes a real buffer and gets
// INVALID_ARGUMENT if it is too short, having written *count first, which is
// what makes a growth retry possible at all.
type diarSegmenter interface {
// countSegments is the size query. It never fails for lack of a buffer.
countSegments() (uint64, error)
// fillSegments fills buf and returns the count the runtime reported. That
// count is meaningful even alongside an error: on a short buffer the
// runtime writes it before rejecting the call.
fillSegments(buf []cDiarSegment) (uint64, error)
}
// diarStream is one diarization job over the C API, narrowed to what the RPC
// uses.
//
// It is an interface for the same reason asrSession is: no Sortformer GGUF is
// small enough to keep in the tree, so without a seam at the ABI the loop on
// top of it (the empty guard, chunking, finish-before-query, the growth retry)
// would have no test at all. A fake here scripts what C returns; it does not
// pretend to diarize anything.
type diarStream interface {
diarSegmenter
push(pcm []float32, sampleRate int32) error
finish() error
close()
}
// diarStreamOpener creates a job. n.openDiarStream is the C-backed one.
//
// The segmentation config is handed over at open time rather than per query
// because it belongs to the whole job: every segments call on one stream must
// use the same postprocessing or the segment ids would not be comparable
// between calls.
type diarStreamOpener func(cfg *cDiarSegmentationConfig) (diarStream, error)
// cDiarStream is the real diarStream, over one nemo_speech_diar_stream.
type cDiarStream struct {
handle uintptr
cfg *cDiarSegmentationConfig
}
// cfgPtr hands the segmentation config to C, or NULL when the request asked
// for no postprocessing. NULL is not the same as a zeroed struct in spirit
// even though src/asr/c_api.cpp treats them alike today: diar.h documents NULL
// as "library defaults", so it is the one form that cannot be invalidated by a
// future field whose sentinel is not zero.
func (s *cDiarStream) cfgPtr() unsafe.Pointer {
if s.cfg == nil {
return nil
}
// #nosec G103 -- a plain *T to unsafe.Pointer conversion of a non-nil,
// GC-traced field. cDiarSegmentationConfig is pure scalars (no uintptr
// members to pin) and the stream owns it for its whole life, so the only
// requirement is that it outlive the DiarSegments call, which it does.
return unsafe.Pointer(s.cfg)
}
func (s *cDiarStream) push(pcm []float32, sampleRate int32) error {
// &pcm[0] panics on an empty slice before the C side ever sees the call.
if len(pcm) == 0 {
return nil
}
if st := DiarStreamPushF32(s.handle, &pcm[0], uint64(len(pcm)), sampleRate); st != 0 {
return statusErrorf(st, "nemo-speech-cpp: diarization push: %s", ASRLastError())
}
return nil
}
func (s *cDiarStream) finish() error {
if st := DiarStreamFinish(s.handle); st != 0 {
return statusErrorf(st, "nemo-speech-cpp: diarization finish: %s", ASRLastError())
}
return nil
}
func (s *cDiarStream) close() { DiarStreamClose(s.handle) }
func (s *cDiarStream) countSegments() (uint64, error) {
var count uint64
// out=NULL and capacity=0: the size query. The runtime reads capacity only
// once it has a buffer to check it against.
if st := DiarSegments(s.handle, s.cfgPtr(), nil, 0, &count); st != 0 {
return 0, statusErrorf(st,
"nemo-speech-cpp: diarization segment count: %s", ASRLastError())
}
return count, nil
}
func (s *cDiarStream) fillSegments(buf []cDiarSegment) (uint64, error) {
if len(buf) == 0 {
// A NULL out would silently turn this into a second size query, and the
// caller would read it as "filled nothing" rather than "asked nothing".
return 0, status.Error(codes.Internal,
"nemo-speech-cpp: diarization segment fill needs a buffer")
}
var count uint64
// #nosec G103 -- &buf[0] is guarded by the empty check above, and the
// capacity handed over is exactly len(buf), so the runtime cannot write past
// the caller's allocation. collectSegments sizes buf under maxDiarSegments
// and rejects a reported count larger than it rather than slicing to it.
st := DiarSegments(s.handle, s.cfgPtr(), unsafe.Pointer(&buf[0]), uint64(len(buf)), &count)
if st != 0 {
// count is returned alongside the error on purpose: a too-small buffer
// is rejected only after the runtime has written the size it wanted.
return count, statusErrorf(st,
"nemo-speech-cpp: diarization segments: %s", ASRLastError())
}
return count, nil
}
// diarGeometryDefault is the sentinel that means "keep the preset's value" for
// every one of nemo_speech_diar_model_config's six frame counts.
//
// It has to be negative, not zero, and that is not a style choice.
// src/asr/c_api.cpp:497-512 applies five of the six overrides when they are
// > 0 but applies left_context_frames when it is >= 0, so a zero-valued config
// reads as "unset" for five fields and as an explicit left context of zero for
// the sixth. That silently changes the model's streaming geometry, and no
// layout assertion can see it because the struct is the right shape either way.
const diarGeometryDefault int32 = -1
// diarModelConfig builds the create-time config for the standalone diarizer.
//
// Extracted from loadDiarizer purely so the sentinels above can be asserted:
// they are invisible to every other check in the tree, including the layout
// assertions, so a spec pinning them is the only thing standing between a
// dropped -1 and a quietly mis-configured model.
//
// modelPath is a C pointer from cstr, not a Go string, and the caller owns its
// release. preset is deliberately left NULL, which diar.h reads as "streaming".
// The "offline" preset is a different accuracy/latency tradeoff for long files
// and is worth exposing, but not on an unverified guess: no Sortformer GGUF
// exists here to measure the difference on.
func diarModelConfig(modelPath uintptr, gpu int32) cDiarModelConfig {
return cDiarModelConfig{
Size: unsafe.Sizeof(cDiarModelConfig{}),
ModelPath: modelPath,
GPU: gpu,
ChunkFrames: diarGeometryDefault,
RightContextFrames: diarGeometryDefault,
LeftContextFrames: diarGeometryDefault,
FIFOFrames: diarGeometryDefault,
SpkcacheFrames: diarGeometryDefault,
UpdatePeriodFrames: diarGeometryDefault,
}
}
// loadDiarizer creates the standalone Sortformer diarizer.
//
// This must not take engineMu: Load is its only caller and already holds it.
func (n *NemoSpeech) loadDiarizer(modelFile string) error {
pathP, freePath := cstr(modelFile)
defer freePath()
cfg := diarModelConfig(pathP, n.opts.gpu)
xlog.Info("nemo-speech-cpp: creating diarizer", "gpu", n.opts.gpu)
// #nosec G103 -- cfg is a local POD struct borrowed for this call only. Its
// only uintptr member is ModelPath, the cstr allocation pinned by the
// deferred freePath above (Preset is deliberately NULL), and
// nemo_speech_diar_create deep-copies the path and retains nothing.
if st := DiarCreate(unsafe.Pointer(&cfg), &n.diarizer); st != 0 {
return statusErrorf(st, "nemo-speech-cpp: diarizer create: %s", ASRLastError())
}
return nil
}
// openDiarStream starts a diarization job on the loaded model.
//
// The caller must hold engineMu.
func (n *NemoSpeech) openDiarStream(cfg *cDiarSegmentationConfig) (diarStream, error) {
var handle uintptr
if st := DiarStreamOpen(n.diarizer, &handle); st != 0 {
return nil, statusErrorf(st,
"nemo-speech-cpp: diarization stream open: %s", ASRLastError())
}
return &cDiarStream{handle: handle, cfg: cfg}, nil
}
// sizeofDiarSegmentationConfig is the size the runtime validates the config
// against. It is a function so the specs can assert the value the config
// actually carries rather than restate the number.
func sizeofDiarSegmentationConfig() uintptr {
return unsafe.Sizeof(cDiarSegmentationConfig{})
}
// segmentationConfig maps the request's postprocessing knobs onto
// nemo_speech_diar_segmentation_config, or returns nil when none were set.
//
// Only two of DiarizeRequest's tuning fields have a real equivalent here, and
// both are exact rather than approximate: NeMo's ts_vad postprocessing is the
// same algorithm the proto's wording describes.
//
// - min_duration_on ("discard segments shorter than this") is min_duration_sec
// ("drop segments shorter than this"), which c_api.cpp assigns to
// DiarSegmentationCfg.min_duration_on.
// - min_duration_off ("merge gaps shorter than this") is min_gap_sec ("fill
// silence gaps shorter than this"), assigned to min_duration_off.
//
// The names cross over between the proto and the C header, which is exactly the
// kind of transposition a layout assertion cannot see, so each mapping is
// pinned by its own spec.
//
// Nothing is written for a non-positive value: the runtime tests every field
// with > 0 and keeps its default otherwise, so a zero here means "unset" on
// both sides.
func segmentationConfig(req *pb.DiarizeRequest) *cDiarSegmentationConfig {
cfg := cDiarSegmentationConfig{Size: sizeofDiarSegmentationConfig()}
var set bool
if v := req.GetMinDurationOn(); v > 0 {
cfg.MinDurationSec = float64(v)
set = true
}
if v := req.GetMinDurationOff(); v > 0 {
cfg.MinGapSec = float64(v)
set = true
}
if !set {
return nil
}
return &cfg
}
// unsupportedRequestFields names the DiarizeRequest fields this backend cannot
// honour, so they are logged rather than silently dropped.
//
// Each is a deliberate omission, not a gap waiting to be filled:
//
// - num_speakers, min_speakers, max_speakers: Sortformer is end-to-end and
// its speaker capacity is fixed by the checkpoint (v2: 4).
// nemo_speech_diar_num_speakers reports that capacity, it does not set it,
// and there is no config field for a target count.
// - clustering_threshold: there is no clustering stage. The nearest knob is
// the onset/offset probability hysteresis, which is a different quantity on
// a different scale, so mapping one onto the other would invent an
// equivalence the header does not have.
// - include_text: this pipeline carries no ASR at all (diar.h: "no ASR
// involved"). Word-level speaker tags on a transcript are the ASR surface's
// job, through diar_model plus enable_speaker_diarization.
// - threads: neither nemo_speech_diar_model_config nor the segmentation
// config has a thread count.
func unsupportedRequestFields(req *pb.DiarizeRequest) []string {
var out []string
if req.GetNumSpeakers() != 0 {
out = append(out, "num_speakers")
}
if req.GetMinSpeakers() != 0 {
out = append(out, "min_speakers")
}
if req.GetMaxSpeakers() != 0 {
out = append(out, "max_speakers")
}
if req.GetClusteringThreshold() != 0 {
out = append(out, "clustering_threshold")
}
if req.GetIncludeText() {
out = append(out, "include_text")
}
if req.GetThreads() != 0 {
out = append(out, "threads")
}
return out
}
// collectSegments runs the count-then-fill protocol and returns the segments.
//
// The growth retry is not defensive padding. nemo_speech_diar_segments writes
// *count and only then rejects a buffer that is too small, so the size a
// rejected call reports is the size to retry with; without the retry a stream
// that gained a segment between the two calls would fail the whole request.
// Truncating to the first count instead would be worse still, dropping turns
// with nothing to show for it.
func collectSegments(s diarSegmenter) ([]cDiarSegment, error) {
want, err := s.countSegments()
if err != nil {
return nil, err
}
for range diarSegmentsMaxAttempts {
if want == 0 {
// No segments means no fill: the fill call needs a non-empty buffer
// to be distinguishable from a second size query.
return nil, nil
}
if want > maxDiarSegments {
return nil, status.Errorf(codes.Internal,
"nemo-speech-cpp: diarization reported %d segments, above the %d ceiling", want, maxDiarSegments)
}
buf := make([]cDiarSegment, want)
got, fillErr := s.fillSegments(buf)
if fillErr == nil {
if got > want {
// The runtime cannot report this on success (it rejects a short
// buffer instead), so it means the ABI is not what this code
// thinks it is. Slicing to it would read past the allocation.
return nil, status.Errorf(codes.Internal,
"nemo-speech-cpp: diarization returned %d segments for a %d-segment buffer", got, want)
}
return buf[:got], nil
}
// A count that did not grow means the call failed for some other
// reason, and retrying the same size would just fail the same way.
if got <= want {
return nil, fillErr
}
want = got
}
return nil, status.Error(codes.Internal,
"nemo-speech-cpp: diarization segment count kept growing, giving up")
}
// toDiarizeSegments converts the runtime's segments to the wire form.
//
// No unit conversion happens here, and that is the point: nemo_speech_diar_segment
// carries start_time and end_time in SECONDS already (diar.h), and
// DiarizeSegment.start/end are seconds too. The frame indices the model works
// in never reach this layer, so nemo_speech_diar_seconds_per_frame is not
// involved. The narrowing to float32 is the proto's choice of type; at 80 ms
// resolution it is lossless for any clip short enough to hold in memory.
//
// The speaker label is the runtime's 1-based tag rendered as a decimal string,
// which is what wordsToSegments emits for the ASR path. The same speaker has to
// read the same way whether the caller diarized a file or transcribed it.
func toDiarizeSegments(in []cDiarSegment) []*pb.DiarizeSegment {
if len(in) == 0 {
return nil
}
out := make([]*pb.DiarizeSegment, 0, len(in))
for i, s := range in {
out = append(out, &pb.DiarizeSegment{
Id: int32(i),
Start: float32(s.StartTime),
End: float32(s.EndTime),
Speaker: strconv.Itoa(int(s.Speaker)),
})
}
return out
}
// distinctSpeakers counts the speaker labels present in the segments.
//
// This is what DiarizeResponse.num_speakers is documented to hold, and it is
// NOT nemo_speech_diar_num_speakers: that reports the checkpoint's capacity
// (four for Sortformer v2), so a two-person interview would come back claiming
// four speakers.
func distinctSpeakers(segs []*pb.DiarizeSegment) int32 {
seen := make(map[string]struct{}, len(segs))
for _, s := range segs {
seen[s.GetSpeaker()] = struct{}{}
}
// #nosec G115 -- seen holds at most one entry per segment, and collectSegments
// refuses any count above maxDiarSegments (2^22), so this is orders of
// magnitude below the int32 the proto field is.
return int32(len(seen))
}
// diarizePCM drives one whole clip through a diarization job.
//
// The caller must hold engineMu.
func diarizePCM(open diarStreamOpener, pcm []float32, sampleRate int32, cfg *cDiarSegmentationConfig) (*pb.DiarizeResponse, error) {
// Before the stream is opened, not inside the push: a silent or truncated
// upload decodes to zero samples, &pcm[0] panics on that, and there is no
// diarization to be had from it anyway.
if len(pcm) == 0 {
return nil, status.Error(codes.InvalidArgument, "nemo-speech-cpp: empty audio")
}
stream, err := open(cfg)
if err != nil {
return nil, err
}
defer stream.close()
// Chunked rather than pushed whole so the runtime advances as it goes
// instead of buffering the entire clip before the first chunk boundary.
for _, chunk := range chunkPCM(pcm, streamChunkSamples) {
if err := stream.push(chunk, sampleRate); err != nil {
return nil, err
}
}
// Before the query, always: finish is what labels the audio tail, so
// segmenting first drops the last turn of every clip.
if err := stream.finish(); err != nil {
return nil, err
}
raw, err := collectSegments(stream)
if err != nil {
return nil, err
}
segs := toDiarizeSegments(raw)
out := &pb.DiarizeResponse{
Segments: segs,
NumSpeakers: distinctSpeakers(segs),
}
// 0 is the proto's "unknown" and the C API's "already at the model rate",
// so a rate that means the latter must not be divided by.
if sampleRate > 0 {
out.Duration = float32(len(pcm)) / float32(sampleRate)
}
// Language and the per-segment text stay empty: there is no ASR in this
// pipeline to fill them, and the proto documents both as optional.
return out, nil
}
// Diarize labels who spoke when in the audio at req.Dst.
//
// The whole body runs inside withEngine, so the family check and the C calls
// that trust the handle happen under a single acquisition of engineMu. The
// audio decode is in there too, for the reason documented on
// AudioTranscription: the backend already serialises RPCs, so the wider hold
// costs nothing, and the narrower one is the gap Free can land in.
func (n *NemoSpeech) Diarize(req *pb.DiarizeRequest) (pb.DiarizeResponse, error) {
var out *pb.DiarizeResponse
if err := n.withEngine(familyDiarization, func() error {
r, err := n.diarize(req)
out = r
return err
}); err != nil {
return pb.DiarizeResponse{}, err
}
if out == nil {
return pb.DiarizeResponse{}, status.Error(codes.Internal,
"nemo-speech-cpp: diarization produced no result")
}
// Assembled field by field rather than dereferenced: the RPC returns the
// message by value and the message embeds a mutex, so copying the struct is
// a copylocks violation.
return pb.DiarizeResponse{
Segments: out.Segments,
NumSpeakers: out.NumSpeakers,
Duration: out.Duration,
Language: out.Language,
}, nil
}
// diarize is Diarize's body. The caller must hold engineMu.
func (n *NemoSpeech) diarize(req *pb.DiarizeRequest) (*pb.DiarizeResponse, error) {
if req.GetDst() == "" {
return nil, status.Error(codes.InvalidArgument,
"nemo-speech-cpp: DiarizeRequest.dst (audio path) is required")
}
// Logged rather than rejected: a client that asks for a speaker count still
// wants the diarization it can have, and a request that names a field this
// backend drops should say so somewhere the operator can find it.
if dropped := unsupportedRequestFields(req); len(dropped) > 0 {
xlog.Warn("nemo-speech-cpp: ignoring diarization request fields this model has no equivalent for",
"fields", dropped)
}
pcm, sampleRate, err := decodeAudioMono16k(req.GetDst())
if err != nil {
return nil, status.Errorf(codes.InvalidArgument, "nemo-speech-cpp: read audio: %v", err)
}
return diarizePCM(n.openDiarStream, pcm, sampleRate, segmentationConfig(req))
}

View File

@@ -0,0 +1,540 @@
package main
import (
"errors"
"path/filepath"
"unsafe"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
)
// fakeDiarStream scripts what the C API returns for one diarization job.
//
// There is no Sortformer GGUF in the tree, so this is the only way the loop on
// top of the ABI (the empty guard, chunking, the count-then-fill protocol, the
// buffer growth retry) gets tested at all. It fakes the C contract, not the
// model: segs is whatever nemo_speech_diar_segments would have produced.
type fakeDiarStream struct {
segs []cDiarSegment
// countErr and fillErrs script failures. fillErrs is consumed one entry per
// fillSegments call so a growth retry can be scripted.
countErr error
fillErrs []error
// queryCount, when non-zero, is what the size query reports instead of
// len(segs), so a runtime that under-reported can be scripted.
queryCount uint64
// growTo, when non-zero, is the count reported by the FIRST fillSegments
// call, standing in for a runtime whose segment list outgrew the size query.
growTo uint64
pushed [][]float32
rates []int32
finished int
closed int
counts int
fills int
// opened records the segmentation config the opener was handed.
cfg *cDiarSegmentationConfig
}
func (f *fakeDiarStream) push(pcm []float32, rate int32) error {
f.pushed = append(f.pushed, pcm)
f.rates = append(f.rates, rate)
return nil
}
func (f *fakeDiarStream) finish() error {
f.finished++
return nil
}
func (f *fakeDiarStream) close() { f.closed++ }
func (f *fakeDiarStream) countSegments() (uint64, error) {
f.counts++
if f.countErr != nil {
return 0, f.countErr
}
if f.queryCount > 0 {
return f.queryCount, nil
}
return uint64(len(f.segs)), nil
}
func (f *fakeDiarStream) fillSegments(buf []cDiarSegment) (uint64, error) {
f.fills++
var err error
if len(f.fillErrs) > 0 {
err, f.fillErrs = f.fillErrs[0], f.fillErrs[1:]
}
if f.fills == 1 && f.growTo > 0 {
// The runtime writes *count before it rejects a short buffer, so a
// growth failure still reports the count the caller needs.
return f.growTo, err
}
n := copy(buf, f.segs)
return uint64(n), err
}
// allPushed flattens what the fake received, so a spec can assert the audio
// arrived intact regardless of how it was chunked.
func (f *fakeDiarStream) allPushed() []float32 {
var out []float32
for _, c := range f.pushed {
out = append(out, c...)
}
return out
}
func (f *fakeDiarStream) opener() diarStreamOpener {
return func(cfg *cDiarSegmentationConfig) (diarStream, error) {
f.cfg = cfg
return f, nil
}
}
var _ = Describe("Diarize", func() {
It("refuses when the loaded model is not a diarization model", func() {
n := &NemoSpeech{fam: familyASR}
_, err := n.Diarize(&pb.DiarizeRequest{Dst: "/tmp/whatever.wav"})
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
})
It("refuses on a model that was never loaded", func() {
n := &NemoSpeech{}
_, err := n.Diarize(&pb.DiarizeRequest{Dst: "/tmp/whatever.wav"})
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
})
// The family gate has to run before anything reads the request, or a
// misrouted request would be reported as a bad path rather than as a model
// that cannot diarize.
It("reports a missing audio path on a diarization model", func() {
n := &NemoSpeech{fam: familyDiarization}
_, err := n.Diarize(&pb.DiarizeRequest{})
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(err.Error()).To(ContainSubstring("dst"))
})
It("reports audio it cannot read", func() {
n := &NemoSpeech{fam: familyDiarization}
missing := filepath.Join(GinkgoT().TempDir(), "absent.wav")
_, err := n.Diarize(&pb.DiarizeRequest{Dst: missing})
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(err.Error()).To(ContainSubstring("read audio"))
})
// A rejection must leave the mutex free, or the next request deadlocks
// rather than fails.
It("releases the engine lock on every rejection path", func() {
n := &NemoSpeech{fam: familyDiarization}
_, err := n.Diarize(&pb.DiarizeRequest{})
Expect(err).To(HaveOccurred())
Expect(n.engineMu.TryLock()).To(BeTrue())
n.engineMu.Unlock()
})
})
var _ = Describe("diarizePCM", func() {
// Task 7 found that a purego-bound entry point reached with zero samples
// panics on &pcm[0], so a silent clip must be rejected before the stream is
// ever opened, not inside the push.
It("rejects empty audio without opening a stream", func() {
opened := false
open := func(*cDiarSegmentationConfig) (diarStream, error) {
opened = true
return &fakeDiarStream{}, nil
}
_, err := diarizePCM(open, nil, 16000, nil)
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(err.Error()).To(ContainSubstring("empty audio"))
Expect(opened).To(BeFalse())
})
It("pushes the whole clip, finishes, and closes the stream", func() {
pcm := make([]float32, streamChunkSamples*2+7)
for i := range pcm {
pcm[i] = float32(i)
}
f := &fakeDiarStream{}
_, err := diarizePCM(f.opener(), pcm, 16000, nil)
Expect(err).ToNot(HaveOccurred())
Expect(f.allPushed()).To(Equal(pcm))
Expect(f.pushed).To(HaveLen(3), "the clip must be chunked, not pushed whole")
Expect(f.rates).To(HaveEach(int32(16000)))
Expect(f.finished).To(Equal(1))
Expect(f.closed).To(Equal(1))
})
// Segments must come from a finished stream: the tail of the audio is only
// labelled by finish, so asking first silently drops the last turn.
It("finishes the stream before it asks for segments", func() {
f := &fakeDiarStream{segs: []cDiarSegment{{StartTime: 0, EndTime: 1, Speaker: 1}}}
f.fillErrs = nil
var finishedAtCount int
wrapped := func(cfg *cDiarSegmentationConfig) (diarStream, error) {
f.cfg = cfg
return &countObserver{fakeDiarStream: f, seen: &finishedAtCount}, nil
}
_, err := diarizePCM(wrapped, []float32{1, 2, 3}, 16000, nil)
Expect(err).ToNot(HaveOccurred())
Expect(finishedAtCount).To(Equal(1), "the size query ran before finish")
})
It("converts the runtime's seconds straight through and numbers the segments", func() {
f := &fakeDiarStream{segs: []cDiarSegment{
{StartTime: 0, EndTime: 0.8, Speaker: 1},
{StartTime: 0.8, EndTime: 2.0, Speaker: 2},
}}
res, err := diarizePCM(f.opener(), []float32{1, 2, 3}, 16000, nil)
Expect(err).ToNot(HaveOccurred())
Expect(res.Segments).To(HaveLen(2))
Expect(res.Segments[0].GetId()).To(Equal(int32(0)))
Expect(res.Segments[0].GetStart()).To(BeNumerically("~", 0.0, 1e-6))
Expect(res.Segments[0].GetEnd()).To(BeNumerically("~", 0.8, 1e-6))
Expect(res.Segments[0].GetSpeaker()).To(Equal("1"))
Expect(res.Segments[1].GetId()).To(Equal(int32(1)))
Expect(res.Segments[1].GetStart()).To(BeNumerically("~", 0.8, 1e-6))
Expect(res.Segments[1].GetEnd()).To(BeNumerically("~", 2.0, 1e-6))
Expect(res.Segments[1].GetSpeaker()).To(Equal("2"))
})
It("reports the clip duration in seconds", func() {
f := &fakeDiarStream{}
res, err := diarizePCM(f.opener(), make([]float32, 32000), 16000, nil)
Expect(err).ToNot(HaveOccurred())
Expect(res.GetDuration()).To(BeNumerically("~", 2.0, 1e-6))
})
// 0 is the proto's documented "unknown", and it is also the C API's "these
// samples are already at the model rate", so a rate that cannot be trusted
// must not be turned into a duration.
It("reports no duration when the sample rate is unknown", func() {
f := &fakeDiarStream{}
res, err := diarizePCM(f.opener(), make([]float32, 32000), 0, nil)
Expect(err).ToNot(HaveOccurred())
Expect(res.GetDuration()).To(BeZero())
})
It("hands the segmentation config to the opener", func() {
f := &fakeDiarStream{}
cfg := &cDiarSegmentationConfig{MinDurationSec: 0.5}
_, err := diarizePCM(f.opener(), []float32{1}, 16000, cfg)
Expect(err).ToNot(HaveOccurred())
Expect(f.cfg).To(BeIdenticalTo(cfg))
})
It("closes the stream when the segment query fails", func() {
f := &fakeDiarStream{countErr: errors.New("boom")}
_, err := diarizePCM(f.opener(), []float32{1}, 16000, nil)
Expect(err).To(MatchError(ContainSubstring("boom")))
Expect(f.closed).To(Equal(1))
})
// The pipeline carries no ASR, so text and language stay empty whatever the
// caller asked for.
It("leaves the transcript fields empty", func() {
f := &fakeDiarStream{segs: []cDiarSegment{{StartTime: 0, EndTime: 1, Speaker: 1}}}
res, err := diarizePCM(f.opener(), []float32{1}, 16000, nil)
Expect(err).ToNot(HaveOccurred())
Expect(res.GetLanguage()).To(BeEmpty())
Expect(res.Segments[0].GetText()).To(BeEmpty())
})
})
// countObserver records how many size queries had run by the time finish was
// called, so the ordering can be asserted without reaching into diarizePCM.
type countObserver struct {
*fakeDiarStream
seen *int
}
func (c *countObserver) finish() error {
*c.seen = c.counts + 1 // finish must run before the first query
return c.fakeDiarStream.finish()
}
var _ = Describe("distinctSpeakers", func() {
It("counts labels, not segments", func() {
segs := []*pb.DiarizeSegment{
{Speaker: "1"}, {Speaker: "2"}, {Speaker: "1"}, {Speaker: "2"}, {Speaker: "1"},
}
Expect(distinctSpeakers(segs)).To(Equal(int32(2)))
})
It("counts a single-speaker recording as one", func() {
segs := []*pb.DiarizeSegment{{Speaker: "1"}, {Speaker: "1"}, {Speaker: "1"}}
Expect(distinctSpeakers(segs)).To(Equal(int32(1)))
})
// Four segments over three labels, not three over three: with the segment
// count and the label count equal, a `return len(segs)` would satisfy this
// spec and it would assert nothing.
It("counts every distinct label once", func() {
segs := []*pb.DiarizeSegment{{Speaker: "1"}, {Speaker: "2"}, {Speaker: "3"}, {Speaker: "2"}}
Expect(distinctSpeakers(segs)).To(Equal(int32(3)))
})
It("is zero with no segments", func() {
Expect(distinctSpeakers(nil)).To(BeZero())
})
// The response field is documented as the count of distinct labels in
// `segments`, which is not the model's capacity: Sortformer v2 can label
// four speakers whatever the clip actually contains.
It("reports what the segments contain, not the model capacity", func() {
f := &fakeDiarStream{segs: []cDiarSegment{
{StartTime: 0, EndTime: 1, Speaker: 1},
{StartTime: 1, EndTime: 2, Speaker: 2},
{StartTime: 2, EndTime: 3, Speaker: 1},
}}
res, err := diarizePCM(f.opener(), []float32{1}, 16000, nil)
Expect(err).ToNot(HaveOccurred())
Expect(res.GetNumSpeakers()).To(Equal(int32(2)))
})
It("reports no speakers when the runtime found no segments", func() {
f := &fakeDiarStream{}
res, err := diarizePCM(f.opener(), []float32{1}, 16000, nil)
Expect(err).ToNot(HaveOccurred())
Expect(res.Segments).To(BeEmpty())
Expect(res.GetNumSpeakers()).To(BeZero())
})
})
var _ = Describe("collectSegments", func() {
It("skips the fill entirely when there is nothing to collect", func() {
f := &fakeDiarStream{}
segs, err := collectSegments(f)
Expect(err).ToNot(HaveOccurred())
Expect(segs).To(BeEmpty())
Expect(f.counts).To(Equal(1))
Expect(f.fills).To(BeZero(), "a zero count must not be followed by a fill")
})
It("sizes the buffer from the query and fills it", func() {
f := &fakeDiarStream{segs: []cDiarSegment{
{StartTime: 0, EndTime: 1, Speaker: 1},
{StartTime: 1, EndTime: 2, Speaker: 2},
}}
segs, err := collectSegments(f)
Expect(err).ToNot(HaveOccurred())
Expect(segs).To(HaveLen(2))
Expect(segs[1].Speaker).To(Equal(int32(2)))
Expect(f.counts).To(Equal(1))
Expect(f.fills).To(Equal(1))
})
// nemo_speech_diar_segments writes *count and only then rejects a buffer
// that is too small, so the rejected call still reports the size to retry
// with. Truncating instead would silently drop turns.
It("grows the buffer and retries when the count outran the query", func() {
f := &fakeDiarStream{
segs: []cDiarSegment{
{StartTime: 0, EndTime: 1, Speaker: 1},
{StartTime: 1, EndTime: 2, Speaker: 2},
{StartTime: 2, EndTime: 3, Speaker: 1},
},
// The query saw two, the fill found three and rejected the buffer.
queryCount: 2,
growTo: 3,
fillErrs: []error{errors.New("capacity too small (need 3)")},
}
segs, err := collectSegments(f)
Expect(err).ToNot(HaveOccurred())
Expect(segs).To(HaveLen(3))
Expect(f.fills).To(Equal(2))
})
It("propagates a failure that is not about capacity", func() {
f := &fakeDiarStream{
segs: []cDiarSegment{{StartTime: 0, EndTime: 1, Speaker: 1}},
fillErrs: []error{errors.New("boom")},
}
_, err := collectSegments(f)
Expect(err).To(MatchError(ContainSubstring("boom")))
Expect(f.fills).To(Equal(1), "a non-capacity failure must not be retried")
})
// make() panics on a length it cannot satisfy, and a panic in an RPC
// handler kills the backend process. A count that could only come from an
// uninitialised or corrupted size_t must fail the request instead.
It("refuses an implausible count rather than trying to allocate it", func() {
f := &fakeDiarStream{queryCount: maxDiarSegments + 1}
_, err := collectSegments(f)
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.Internal))
Expect(err.Error()).To(ContainSubstring("ceiling"))
Expect(f.fills).To(BeZero(), "nothing must be allocated or filled for a bad count")
})
It("still accepts a count right at the ceiling", func() {
// Only the guard is under test, so the fill is scripted to report zero
// rather than actually materialising a hundred megabytes of segments.
f := &fakeDiarStream{queryCount: maxDiarSegments}
segs, err := collectSegments(f)
Expect(err).ToNot(HaveOccurred())
Expect(segs).To(BeEmpty())
Expect(f.fills).To(Equal(1))
})
It("propagates a failed size query", func() {
f := &fakeDiarStream{countErr: errors.New("no stream")}
_, err := collectSegments(f)
Expect(err).To(MatchError(ContainSubstring("no stream")))
Expect(f.fills).To(BeZero())
})
// A runtime whose count grew on every attempt would otherwise loop forever
// holding engineMu, which blocks the unload too.
It("gives up rather than retrying forever", func() {
f := &fakeDiarStream{segs: []cDiarSegment{{StartTime: 0, EndTime: 1, Speaker: 1}}}
g := &alwaysGrowing{fakeDiarStream: f}
_, err := collectSegments(g)
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.Internal))
Expect(f.fills).To(Equal(diarSegmentsMaxAttempts))
})
})
// alwaysGrowing reports a bigger count on every fill, which is the pathological
// case the attempt bound exists for.
type alwaysGrowing struct {
*fakeDiarStream
n uint64
}
func (a *alwaysGrowing) fillSegments([]cDiarSegment) (uint64, error) {
a.n += 10
a.fills++
return a.n, errors.New("capacity too small")
}
// The six frame counts are the one part of the create config that no other
// check in the tree can see. The layout assertions pin the struct's shape, and
// a wrong VALUE keeps that shape exactly, so without these specs deleting a
// sentinel is invisible: c_api.cpp applies left_context_frames at >= 0, so a
// dropped -1 there silently pins the model's left context to zero.
var _ = Describe("diarModelConfig", func() {
It("declares its own size so the runtime accepts the fields", func() {
Expect(diarModelConfig(0, -1).Size).To(Equal(unsafe.Sizeof(cDiarModelConfig{})))
})
It("carries the model path and the configured device", func() {
cfg := diarModelConfig(0xDEADBEEF, 2)
Expect(cfg.ModelPath).To(Equal(uintptr(0xDEADBEEF)))
Expect(cfg.GPU).To(Equal(int32(2)))
})
It("passes the CPU sentinel through untouched", func() {
Expect(diarModelConfig(0, -1).GPU).To(Equal(int32(-1)))
})
// Asserted field by field rather than as a whole struct so a failure names
// the sentinel that went missing.
It("leaves every frame-geometry override at the negative sentinel", func() {
cfg := diarModelConfig(0, -1)
Expect(cfg.ChunkFrames).To(Equal(int32(-1)), "chunk_frames")
Expect(cfg.RightContextFrames).To(Equal(int32(-1)), "right_context_frames")
Expect(cfg.FIFOFrames).To(Equal(int32(-1)), "fifo_frames")
Expect(cfg.SpkcacheFrames).To(Equal(int32(-1)), "spkcache_frames")
Expect(cfg.UpdatePeriodFrames).To(Equal(int32(-1)), "update_period_frames")
// Called out on its own because it is the only one of the six the
// runtime applies at >= 0: zero here is a valid explicit left context,
// not "unset", so this is the field a dropped sentinel actually breaks.
Expect(cfg.LeftContextFrames).To(Equal(int32(-1)), "left_context_frames")
Expect(cfg.LeftContextFrames).To(BeNumerically("<", 0),
"left_context_frames is applied at >= 0, so a non-negative value pins the geometry")
})
// The preset selects the streaming geometry wholesale, so it has to stay
// NULL until there is a model to verify a different one against.
It("leaves the preset unset", func() {
Expect(diarModelConfig(0, -1).Preset).To(BeZero())
})
})
var _ = Describe("segmentationConfig", func() {
// A request that set nothing must stay NULL on the C side: diar.h documents
// NULL as "library defaults", and those defaults are NeMo's callhome-tuned
// values for this checkpoint rather than zeros.
It("is absent when the request asked for no postprocessing", func() {
Expect(segmentationConfig(&pb.DiarizeRequest{})).To(BeNil())
})
It("maps min_duration_on onto the minimum segment duration", func() {
cfg := segmentationConfig(&pb.DiarizeRequest{MinDurationOn: 0.4})
Expect(cfg).ToNot(BeNil())
Expect(cfg.MinDurationSec).To(BeNumerically("~", 0.4, 1e-6))
Expect(cfg.MinGapSec).To(BeZero())
})
It("maps min_duration_off onto the gap fill", func() {
cfg := segmentationConfig(&pb.DiarizeRequest{MinDurationOff: 0.25})
Expect(cfg).ToNot(BeNil())
Expect(cfg.MinGapSec).To(BeNumerically("~", 0.25, 1e-6))
Expect(cfg.MinDurationSec).To(BeZero())
})
It("declares its own size so the runtime accepts the fields", func() {
cfg := segmentationConfig(&pb.DiarizeRequest{MinDurationOn: 0.4})
Expect(cfg.Size).To(Equal(sizeofDiarSegmentationConfig()))
})
// The onset/offset hysteresis is not a clustering threshold and Sortformer
// has no clustering stage at all, so mapping one onto the other would be an
// invented equivalence. It has to stay unset.
It("ignores fields this pipeline has no equivalent for", func() {
Expect(segmentationConfig(&pb.DiarizeRequest{
NumSpeakers: 2,
MinSpeakers: 1,
MaxSpeakers: 4,
ClusteringThreshold: 0.7,
IncludeText: true,
Threads: 8,
})).To(BeNil())
})
It("ignores non-positive values, which the runtime reads as unset", func() {
Expect(segmentationConfig(&pb.DiarizeRequest{MinDurationOn: -1, MinDurationOff: 0})).To(BeNil())
})
})
var _ = Describe("unsupportedRequestFields", func() {
It("is empty for a request this backend can honour in full", func() {
Expect(unsupportedRequestFields(&pb.DiarizeRequest{
Dst: "/tmp/a.wav",
MinDurationOn: 0.4,
MinDurationOff: 0.2,
})).To(BeEmpty())
})
It("names every field it had to drop", func() {
Expect(unsupportedRequestFields(&pb.DiarizeRequest{
NumSpeakers: 2,
MinSpeakers: 1,
MaxSpeakers: 4,
ClusteringThreshold: 0.7,
IncludeText: true,
Threads: 8,
})).To(ConsistOf(
"num_speakers", "min_speakers", "max_speakers",
"clustering_threshold", "include_text", "threads",
))
})
})

View File

@@ -0,0 +1,116 @@
package main
import (
"fmt"
"os"
"path/filepath"
"strings"
gguf "github.com/gpustack/gguf-parser-go"
)
// auxOnlyArchitectures are converted NeMo components that attach to a primary
// model but are never loadable on their own. Pointing a model config at one is
// a configuration mistake worth naming explicitly.
var auxOnlyArchitectures = map[string]string{
"nemo-nano-codec": "a TTS codec, set it with the codec_model option on a magpietts model",
"vad": "a VAD model, set it with the vad_model option on an asr model",
"pnc": "a punctuation model, set it with the pnc_model option on an asr model",
}
// familyFor maps a GGUF general.architecture value to a model family.
//
// Unknown architectures resolve to NMT rather than an error: NMT GGUFs come
// from llama.cpp's converter and carry an ordinary LLM architecture, so there
// is no NeMo-specific string to match. The user selected this backend
// explicitly, which is the signal that the model is meant for it.
func familyFor(arch string) (family, error) {
if reason, ok := auxOnlyArchitectures[arch]; ok {
return familyUnknown, fmt.Errorf(
"nemo-speech-cpp: %q is %s, not a model that can be loaded directly", arch, reason)
}
switch arch {
case "asr":
return familyASR, nil
case "sortformer":
return familyDiarization, nil
case "magpietts":
return familyTTS, nil
}
return familyNMT, nil
}
// ggufArchitecture reads general.architecture from a GGUF file.
func ggufArchitecture(path string) (string, error) {
f, err := gguf.ParseGGUFFile(path, gguf.UseMMap(), gguf.SkipLargeMetadata())
if err != nil {
return "", fmt.Errorf("nemo-speech-cpp: parse gguf %q: %w", path, err)
}
kv, found := f.Header.MetadataKV.Index([]string{"general.architecture"})
if found == 0 {
return "", fmt.Errorf("nemo-speech-cpp: %q has no general.architecture key", path)
}
arch := kv["general.architecture"]
// ValueString panics on a mistyped key, and a hand-written or half-converted
// GGUF is exactly where that happens. This function is the load-time guard;
// it reports, it does not take the process down.
if arch.ValueType != gguf.GGUFMetadataValueTypeString {
return "", fmt.Errorf(
"nemo-speech-cpp: %q has a non-string general.architecture (type %v)", path, arch.ValueType)
}
return arch.ValueString(), nil
}
// discoverTTSAssets fills in codecModel and tokenizerDir when they were not set
// explicitly, by scanning the primary GGUF's own directory.
//
// A missing asset is a hard error rather than a warning: the runtime would
// otherwise load and emit garbage audio, which surfaces far from the cause.
func discoverTTSAssets(primaryGGUF string, o *loadOptions) error {
dir := filepath.Dir(primaryGGUF)
if o.codecModel == "" {
entries, err := os.ReadDir(dir)
if err != nil {
return fmt.Errorf("nemo-speech-cpp: scan %q for a codec model: %w", dir, err)
}
for _, e := range entries {
if e.IsDir() {
continue
}
name := e.Name()
candidate := filepath.Join(dir, name)
// Skip the primary model itself: a file called nanocodec-magpie.gguf
// would otherwise be selected as its own codec. Compare basenames,
// because candidate is Cleaned by filepath.Join while primaryGGUF
// arrives as the caller wrote it, so "/models//magpie.gguf" would
// slip past a whole-path equality.
if name == filepath.Base(primaryGGUF) {
continue
}
if strings.Contains(strings.ToLower(name), "nanocodec") ||
strings.Contains(strings.ToLower(name), "nano-codec") {
o.codecModel = candidate
break
}
}
}
if o.codecModel == "" {
return fmt.Errorf(
"nemo-speech-cpp: no NanoCodec GGUF found next to %q, set the codec_model option",
primaryGGUF)
}
if o.tokenizerDir == "" {
candidate := filepath.Join(dir, "extracted")
if st, err := os.Stat(candidate); err == nil && st.IsDir() {
o.tokenizerDir = candidate
}
}
if o.tokenizerDir == "" {
return fmt.Errorf(
"nemo-speech-cpp: no tokenizer directory found next to %q, set the tokenizer_dir option",
primaryGGUF)
}
return nil
}

View File

@@ -0,0 +1,168 @@
package main
import (
"encoding/binary"
"os"
"path/filepath"
gguf "github.com/gpustack/gguf-parser-go"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
var _ = Describe("familyFor", func() {
It("maps the NeMo architectures to their families", func() {
for arch, want := range map[string]family{
"asr": familyASR,
"sortformer": familyDiarization,
"magpietts": familyTTS,
} {
got, err := familyFor(arch)
Expect(err).ToNot(HaveOccurred(), "arch %q", arch)
Expect(got).To(Equal(want), "arch %q", arch)
}
})
It("treats an unknown architecture as NMT", func() {
// NMT GGUFs are produced by llama.cpp's converter, so they carry an LLM
// architecture such as qwen3 rather than a NeMo-specific string.
got, err := familyFor("qwen3")
Expect(err).ToNot(HaveOccurred())
Expect(got).To(Equal(familyNMT))
})
It("rejects an auxiliary-only architecture as a primary model", func() {
for _, arch := range []string{"nemo-nano-codec", "vad", "pnc"} {
_, err := familyFor(arch)
Expect(err).To(HaveOccurred(), "arch %q", arch)
Expect(err.Error()).To(ContainSubstring(arch))
}
})
})
// ggufWithArchValue builds a minimal GGUF v3 carrying general.architecture as
// its single metadata entry, with the caller's value type and encoded value.
func ggufWithArchValue(valueType gguf.GGUFMetadataValueType, value []byte) []byte {
const key = "general.architecture"
var b []byte
b = append(b, 'G', 'G', 'U', 'F')
b = binary.LittleEndian.AppendUint32(b, 3) // version
b = binary.LittleEndian.AppendUint64(b, 0) // tensor count
b = binary.LittleEndian.AppendUint64(b, 1) // metadata kv count
b = binary.LittleEndian.AppendUint64(b, uint64(len(key)))
b = append(b, key...)
b = binary.LittleEndian.AppendUint32(b, uint32(valueType))
return append(b, value...)
}
// writeGGUFWithUint32Arch writes a minimal GGUF v3 whose single metadata entry
// is general.architecture typed UINT32 rather than STRING. Handwritten and
// half-converted files really do carry mistyped keys, and the parser hands them
// back rather than rejecting them.
func writeGGUFWithUint32Arch(path string) {
b := ggufWithArchValue(gguf.GGUFMetadataValueTypeUint32, binary.LittleEndian.AppendUint32(nil, 7))
ExpectWithOffset(1, os.WriteFile(path, b, 0o600)).To(Succeed())
}
// writeGGUFWithArch writes a minimal GGUF v3 that parses cleanly and reports
// arch as its general.architecture. It is the only way to reach the code past
// ggufArchitecture in a test, since there are no real NeMo GGUFs to point at.
func writeGGUFWithArch(path, arch string) {
v := binary.LittleEndian.AppendUint64(nil, uint64(len(arch)))
v = append(v, arch...)
ExpectWithOffset(1, os.WriteFile(path, ggufWithArchValue(gguf.GGUFMetadataValueTypeString, v), 0o600)).To(Succeed())
}
var _ = Describe("ggufArchitecture", func() {
It("returns an error rather than panicking on a file that is not a GGUF", func() {
p := filepath.Join(GinkgoT().TempDir(), "not-a-model.gguf")
Expect(os.WriteFile(p, []byte("definitely not a gguf header"), 0o600)).To(Succeed())
_, err := ggufArchitecture(p)
Expect(err).To(HaveOccurred())
Expect(err.Error()).To(ContainSubstring(p))
})
It("returns an error rather than panicking when general.architecture is not a string", func() {
p := filepath.Join(GinkgoT().TempDir(), "mistyped-arch.gguf")
writeGGUFWithUint32Arch(p)
arch, err := ggufArchitecture(p)
Expect(err).To(HaveOccurred())
Expect(arch).To(BeEmpty())
Expect(err.Error()).To(ContainSubstring("general.architecture"))
})
})
var _ = Describe("discoverTTSAssets", func() {
var dir string
BeforeEach(func() {
dir = GinkgoT().TempDir()
})
write := func(name string) string {
p := filepath.Join(dir, name)
Expect(os.WriteFile(p, []byte("x"), 0o600)).To(Succeed())
return p
}
It("finds a sibling nanocodec gguf and extracted dir", func() {
primary := write("magpie.f16.gguf")
codec := write("nemo-nano-codec-22khz.f16.gguf")
Expect(os.Mkdir(filepath.Join(dir, "extracted"), 0o755)).To(Succeed())
o := loadOptions{}
Expect(discoverTTSAssets(primary, &o)).To(Succeed())
Expect(o.codecModel).To(Equal(codec))
Expect(o.tokenizerDir).To(Equal(filepath.Join(dir, "extracted")))
})
It("never selects the primary gguf as its own codec", func() {
// A file named so it would match a naive *.gguf scan.
primary := write("nanocodec-magpie.gguf")
Expect(os.Mkdir(filepath.Join(dir, "extracted"), 0o755)).To(Succeed())
o := loadOptions{}
err := discoverTTSAssets(primary, &o)
Expect(err).To(HaveOccurred())
Expect(err.Error()).To(ContainSubstring("codec_model"))
})
It("never selects the primary gguf as its own codec through an uncleaned path", func() {
// LocalAI joins the model directory and the model name itself, so a
// trailing separator on ModelPath produces a doubled slash here. The
// self-codec guard has to survive that.
write("nanocodec-magpie.gguf")
primary := dir + "//nanocodec-magpie.gguf"
Expect(os.Mkdir(filepath.Join(dir, "extracted"), 0o755)).To(Succeed())
o := loadOptions{}
err := discoverTTSAssets(primary, &o)
Expect(err).To(HaveOccurred())
Expect(o.codecModel).To(BeEmpty())
Expect(err.Error()).To(ContainSubstring("codec_model"))
})
It("does not overwrite explicitly configured paths", func() {
primary := write("magpie.f16.gguf")
write("nemo-nano-codec.gguf")
Expect(os.Mkdir(filepath.Join(dir, "extracted"), 0o755)).To(Succeed())
o := loadOptions{codecModel: "/explicit/codec.gguf", tokenizerDir: "/explicit/tok"}
Expect(discoverTTSAssets(primary, &o)).To(Succeed())
Expect(o.codecModel).To(Equal("/explicit/codec.gguf"))
Expect(o.tokenizerDir).To(Equal("/explicit/tok"))
})
It("names the missing option key when the tokenizer dir cannot be found", func() {
primary := write("magpie.f16.gguf")
write("nemo-nano-codec.gguf")
o := loadOptions{}
err := discoverTTSAssets(primary, &o)
Expect(err).To(HaveOccurred())
Expect(err.Error()).To(ContainSubstring("tokenizer_dir"))
})
})

View File

@@ -0,0 +1,76 @@
package main
// Started internally by LocalAI, one gRPC server per loaded model.
//
// Binds NVIDIA NeMo-Speech.cpp through purego. The runtime splits its C ABI
// across three shared objects: asr (which also exports the diarization
// symbols), tts, and nmt. Library names can be overridden with
// NEMO_SPEECH_ASR_LIBRARY / _TTS_LIBRARY / _NMT_LIBRARY, mirroring the
// PARAKEET_LIBRARY convention in the sibling backends.
//
// The naming is asymmetric on purpose: upstream links a dedicated
// libnemo_speech_asr_c / libnemo_speech_nmt_c around a private C++ core, but
// compiles the TTS c_api straight into libnemo_speech_tts and only aliases the
// nemo_speech_tts_c CMake target, so there is no libnemo_speech_tts_c on disk.
import (
"flag"
"fmt"
"os"
"runtime"
"github.com/ebitengine/purego"
grpc "github.com/mudler/LocalAI/pkg/grpc"
)
var addr = flag.String("addr", "localhost:50051", "the address to connect to")
// libSuffix is the platform's shared-object extension.
func libSuffix() string {
if runtime.GOOS == "darwin" {
return ".dylib"
}
return ".so"
}
// libraryName resolves an override env var, falling back to the platform name.
func libraryName(envVar, base string) string {
if v := os.Getenv(envVar); v != "" {
return v
}
return base + libSuffix()
}
func main() {
flag.Parse()
if err := openLibraries(); err != nil {
panic(err)
}
if err := grpc.StartServer(*addr, &NemoSpeech{}); err != nil {
panic(err)
}
}
// openLibraries dlopens the three C ABI shared objects. All three are opened
// eagerly so a packaging mistake fails at startup with a clear message rather
// than at first inference of one particular family.
func openLibraries() error {
for _, l := range []struct {
env string
base string
dst *uintptr
}{
{"NEMO_SPEECH_ASR_LIBRARY", "libnemo_speech_asr_c", &asrLib},
{"NEMO_SPEECH_TTS_LIBRARY", "libnemo_speech_tts", &ttsLib},
{"NEMO_SPEECH_NMT_LIBRARY", "libnemo_speech_nmt_c", &nmtLib},
} {
name := libraryName(l.env, l.base)
h, err := purego.Dlopen(name, purego.RTLD_NOW|purego.RTLD_GLOBAL)
if err != nil {
return fmt.Errorf("nemo-speech-cpp: dlopen %q: %w", name, err)
}
*l.dst = h
}
return registerSymbols()
}

View File

@@ -0,0 +1,273 @@
package main
import (
"errors"
"fmt"
"runtime"
"sync"
"unsafe"
"github.com/mudler/LocalAI/pkg/grpc/base"
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
"github.com/mudler/xlog"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
)
// family is the model family selected at load time from the GGUF architecture.
type family int
const (
familyUnknown family = iota
familyASR
familyDiarization
familyTTS
familyNMT
)
func (f family) String() string {
switch f {
case familyASR:
return "asr"
case familyDiarization:
return "diarization"
case familyTTS:
return "tts"
case familyNMT:
return "nmt"
}
return "unknown"
}
// NemoSpeech is one loaded model. Exactly one of the handles is non-zero,
// matching fam.
type NemoSpeech struct {
base.SingleThread
fam family
opts loadOptions
// engineMu guards fam and the handles, and serializes calls into the C
// runtime for this model. Its participants today are withEngine and Free;
// the per-family RPCs in Tasks 6 to 9 join it by routing through withEngine.
engineMu sync.Mutex
// synth and nmt are shortened rather than spelled out: synthesizer and
// translator are the names of the two RPC-side interfaces those handles are
// wrapped in (tts.go, nmt.go), and a field sharing a name with an interface in
// the same package makes every construction site read as a conversion.
recognizer uintptr
diarizer uintptr
synth uintptr
nmt uintptr
}
// cstr allocates a NUL-terminated C string and returns its pointer plus a
// release function. The empty string maps to a null pointer because the C API
// treats NULL and "" as equivalent for every optional field.
//
// The address leaves the Go type system as a uintptr, which the collector does
// not trace, so the bytes are pinned for as long as C may read them. Pinning is
// the only mechanism with a documented guarantee here: the config structs hold
// raw addresses, and an unpinned Go allocation is free to be collected (and, in
// principle, moved) the moment its last traced reference dies.
//
// The returned pointer is for C only, and the direction is one-way. Converting
// it back to an unsafe.Pointer to read the bytes from Go is checked by checkptr
// (which -race turns on) and kills the process with
//
// fatal error: checkptr: pointer arithmetic result points to invalid allocation
//
// as soon as the address lands inside a Go allocation, which is exactly what
// this produces. C reading it is fine because C is not instrumented; Go reading
// it back is not.
//
// The caller MUST defer the release function immediately, in the same statement
// that takes the pointer. Dropping it leaks the pin, which the runtime reports
// at the next collection as:
//
// runtime.Pinner: found leaking pinned pointer; forgot to call Unpin()?
//
// That is loud and wrong-looking on purpose: the alternative failure mode is C
// reading freed memory, which shows up as rare corruption with no trace back
// to here.
func cstr(s string) (uintptr, func()) {
if s == "" {
return 0, func() {}
}
b := append([]byte(s), 0)
pin := new(runtime.Pinner)
pin.Pin(&b[0])
// #nosec G103 -- b is non-empty (s != "" above) and &b[0] is pinned on the
// previous line, so the address C receives cannot be collected or moved
// until the returned release runs. One-way by construction: the doc comment
// above forbids converting this uintptr back, which is what keeps checkptr
// (and therefore -race) out of it.
return uintptr(unsafe.Pointer(&b[0])), func() {
if pin == nil {
return
}
pin.Unpin()
pin = nil
}
}
// There is deliberately no inverse of cstr in this package. Every C entry point
// that returns a string is bound in abi.go with a Go `string` return, which
// purego converts from the char* itself, so a hand-rolled reader would have no
// production caller and would exist only as an unsafe helper waiting to be
// pointed at the wrong kind of address. Reach for purego's conversion instead;
// if a future symbol genuinely needs the raw char* (to tell NULL from ""), bind
// it as uintptr at that call site, where the ownership can be reasoned about.
// requireFamily gates an RPC on the family selected at load time. Returning
// Unimplemented rather than a nil dereference means a misconfigured model YAML
// produces a message a user can act on.
//
// Callers must already hold engineMu: Free writes n.fam under it, so an
// unlocked read here is a data race. Use withEngine rather than calling this
// directly.
func (n *NemoSpeech) requireFamily(want family) error {
if n.fam != want {
return status.Errorf(codes.Unimplemented,
"nemo-speech-cpp: this model was loaded as %s, not %s", n.fam, want)
}
return nil
}
// withEngine runs fn holding engineMu, having first checked the family.
//
// Every RPC must go through this rather than calling requireFamily on its own.
// pkg/grpc/server.go takes the backend lock around each RPC but calls Free
// without it, so a teardown can land mid-request. Checking the family and then
// making the C calls that trust it under two separate acquisitions leaves a
// window in which Free destroys the handle, and the request goes on to use a
// zeroed one.
func (n *NemoSpeech) withEngine(want family, fn func() error) error {
n.engineMu.Lock()
defer n.engineMu.Unlock()
if err := n.requireFamily(want); err != nil {
return err
}
return fn()
}
func (n *NemoSpeech) Load(opts *pb.ModelOptions) error {
modelFile := opts.GetModelFile()
if modelFile == "" {
return errors.New("nemo-speech-cpp: ModelFile is required")
}
// Free writes fam and the handles under engineMu and runs without the
// backend lock that serialises the RPCs (pkg/grpc/server.go), so the
// load-side writes to those same fields need the same protection: without
// it this is the write-side half of the race withEngine closed on the read
// side. n.opts is in here too, since the loaders read it.
//
// The loaders called below must NOT take engineMu themselves; sync.Mutex is
// not reentrant and this is why.
n.engineMu.Lock()
defer n.engineMu.Unlock()
n.opts = parseOptions(opts.GetOptions(), opts.GetModelPath())
arch, err := ggufArchitecture(modelFile)
if err != nil {
return err
}
fam, err := familyFor(arch)
if err != nil {
return err
}
xlog.Info("nemo-speech-cpp: loading model", "arch", arch, "family", fam.String())
// fam is committed only once the family-specific loader has succeeded.
// requireFamily is the gate every RPC goes through, so a half-loaded model
// that kept its family would route requests at a handle that was never
// created.
switch fam {
case familyASR:
err = n.loadASR(modelFile)
case familyDiarization:
err = n.loadDiarizer(modelFile)
case familyTTS:
if err = discoverTTSAssets(modelFile, &n.opts); err == nil {
err = n.loadTTS(modelFile)
}
case familyNMT:
err = n.loadNMT(modelFile)
default:
err = fmt.Errorf("nemo-speech-cpp: unhandled family for architecture %q", arch)
}
if err != nil {
return err
}
n.fam = fam
return nil
}
// Free destroys the runtime handle created at load time.
//
// base.SingleThread.Free is a no-op that derived backends are expected to
// override, and every family here owns C memory that only its own destroy
// entry point can release, so without this an unloaded model leaks a whole
// acoustic model. Clearing fam as well means an RPC that races the unload is
// refused by the gate rather than handed a dangling handle, but that only holds
// for callers that took engineMu, which today means callers that went through
// withEngine.
func (n *NemoSpeech) Free() error {
n.engineMu.Lock()
defer n.engineMu.Unlock()
// Guarded on the handle, not on fam: a load that failed part way through
// leaves fam unset, and the destroy functions are nil pointers until
// openLibraries has bound them.
// Each is tested independently rather than switched on: the one-handle
// invariant is an invariant, and if it ever broke, a switch would silently
// leak the others.
if n.recognizer != 0 {
ASRDestroy(n.recognizer)
n.recognizer = 0
}
if n.diarizer != 0 {
DiarDestroy(n.diarizer)
n.diarizer = 0
}
if n.synth != 0 {
TTSDestroy(n.synth)
n.synth = 0
}
if n.nmt != 0 {
NMTDestroy(n.nmt)
n.nmt = 0
}
n.fam = familyUnknown
return nil
}
// The loaders are one per family: loadASR in asr.go, loadDiarizer in diar.go,
// loadTTS in tts.go and loadNMT in nmt.go. Each populates its config structs
// from n.opts and stores the handle in the matching field.
//
// Locking protocol, in both directions:
//
// - Every RPC must hold engineMu across its family check AND its C calls,
// which means wrapping its body in withEngine. Free runs without the
// backend lock (pkg/grpc/server.go:1019), so anything that checks the
// family and then releases the lock before calling C can have the handle
// destroyed underneath it. asr.go's AudioTranscription is the worked
// example: even the audio decode sits inside the closure, because the
// backend already serialises RPCs through base.SingleThread and so the
// wider hold costs nothing.
// - A loader must NOT take engineMu. Load holds it across the whole switch,
// and sync.Mutex is not reentrant, so locking in a loader deadlocks.
//
// One consequence the streaming RPCs have to plan around: a stream whose body
// is wrapped in withEngine holds engineMu for the WHOLE stream, so Free blocks
// until the stream ends rather than tearing the handle out from under it. That
// is the behaviour we want (a half-closed stream over a destroyed recognizer
// has no good outcome), but it means an unload waits on a client that has
// stopped sending, so a streaming loop must have its own way out: honour the
// request context and stop on it, rather than blocking forever on the next
// chunk.

View File

@@ -0,0 +1,13 @@
package main
import (
"testing"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
func TestNemoSpeech(t *testing.T) {
RegisterFailHandler(Fail)
RunSpecs(t, "nemo-speech-cpp Backend Suite")
}

View File

@@ -0,0 +1,260 @@
package main
import (
"errors"
"os"
"path/filepath"
"runtime"
"sync"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
)
var _ = Describe("requireFamily", func() {
It("accepts the loaded family", func() {
n := &NemoSpeech{fam: familyASR}
Expect(n.requireFamily(familyASR)).To(Succeed())
})
It("rejects a mismatched family with Unimplemented and names both", func() {
n := &NemoSpeech{fam: familyTTS}
err := n.requireFamily(familyASR)
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
Expect(err.Error()).To(ContainSubstring("tts"))
Expect(err.Error()).To(ContainSubstring("asr"))
})
It("rejects an unloaded model", func() {
n := &NemoSpeech{}
err := n.requireFamily(familyASR)
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
})
It("rejects every family when the model is unloaded", func() {
n := &NemoSpeech{}
for _, f := range []family{familyASR, familyDiarization, familyTTS, familyNMT} {
Expect(n.requireFamily(f)).To(HaveOccurred(), "family %s must be gated on an unloaded model", f)
}
})
})
// The brief's round-trip spec (cstr then a reader) cannot exist: cstr pins a Go
// allocation, and converting a uintptr back into a pointer to Go memory is a
// checkptr violation that aborts the process under -race. So cstr is asserted
// on what is observable without dereferencing its result.
var _ = Describe("cstr", func() {
It("returns a non-null pointer for a non-empty string", func() {
p, free := cstr("hello")
defer free()
Expect(p).ToNot(BeZero())
})
It("returns a null pointer for the empty string", func() {
// The C API documents NULL and "" as equivalent for optional fields, and
// passing NULL avoids allocating for every unset option.
p, free := cstr("")
defer free()
Expect(p).To(BeZero())
})
// The pin has to hold for the whole create call, which spans at least one
// safepoint. A collection must therefore neither move nor invalidate the
// address that C was handed.
It("keeps the pointer stable across a garbage collection", func() {
p, free := cstr("/models/nemo/parakeet.gguf")
defer free()
before := p
runtime.GC()
runtime.GC()
Expect(p).To(Equal(before))
})
It("survives releasing more than once", func() {
_, free := cstr("twice")
free()
Expect(free).ToNot(Panic())
})
// A dropped release leaks the pin, and the runtime turns that into a process
// abort at some later collection. Nothing can catch it, so this only pins the
// contract in prose: release in the same statement that takes the pointer.
It("releases without panicking when used as documented", func() {
Expect(func() {
p, free := cstr("released")
defer free()
_ = p
}).ToNot(Panic())
})
})
// pkg/grpc/server.go:1019 calls Free without taking the backend lock every
// other RPC holds, so a teardown really can land while a request is in flight.
// The family check and the C calls that trust it therefore have to happen under
// engineMu together, or Free can destroy the handle in the gap between them.
var _ = Describe("engine locking", func() {
It("serialises a teardown against an in-flight request", func() {
n := &NemoSpeech{fam: familyASR}
var wg sync.WaitGroup
wg.Add(2)
go func() {
defer GinkgoRecover()
defer wg.Done()
for i := 0; i < 2000; i++ {
// Errors are expected once the teardown wins the race; what must
// not happen is an unsynchronised read of the family.
_ = n.withEngine(familyASR, func() error { return nil })
}
}()
go func() {
defer GinkgoRecover()
defer wg.Done()
for i := 0; i < 2000; i++ {
Expect(n.Free()).To(Succeed())
}
}()
wg.Wait()
})
It("refuses the body when the family does not match, and still unlocks", func() {
n := &NemoSpeech{fam: familyTTS}
called := false
err := n.withEngine(familyASR, func() error {
called = true
return nil
})
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
Expect(called).To(BeFalse())
// A lock leaked on the rejection path would deadlock the next request
// rather than fail it, so prove the mutex is free afterwards.
Expect(n.engineMu.TryLock()).To(BeTrue())
n.engineMu.Unlock()
})
It("propagates the body's error and still unlocks", func() {
n := &NemoSpeech{fam: familyASR}
boom := errors.New("boom")
Expect(n.withEngine(familyASR, func() error { return boom })).To(MatchError(boom))
Expect(n.engineMu.TryLock()).To(BeTrue())
n.engineMu.Unlock()
})
})
var _ = Describe("Free", func() {
// The destroy entry points are nil function values until openLibraries has
// bound them, so an unloaded model must not reach them. LocalAI frees every
// backend it shuts down, including one whose Load failed.
It("is a no-op on a model that was never loaded", func() {
n := &NemoSpeech{}
Expect(n.Free()).To(Succeed())
})
It("is idempotent", func() {
n := &NemoSpeech{}
Expect(n.Free()).To(Succeed())
Expect(n.Free()).To(Succeed())
})
It("does not reach the runtime for a load that failed part way through", func() {
n := &NemoSpeech{}
path := filepath.Join(GinkgoT().TempDir(), "broken.gguf")
Expect(os.WriteFile(path, []byte("broken"), 0o600)).To(Succeed())
Expect(n.Load(&pb.ModelOptions{ModelFile: path})).ToNot(Succeed())
Expect(n.Free()).To(Succeed())
})
})
var _ = Describe("Load", func() {
It("rejects an empty model file", func() {
n := &NemoSpeech{}
err := n.Load(&pb.ModelOptions{})
Expect(err).To(HaveOccurred())
Expect(err.Error()).To(ContainSubstring("ModelFile"))
})
It("reports a model file that does not exist", func() {
n := &NemoSpeech{}
missing := filepath.Join(GinkgoT().TempDir(), "absent.gguf")
err := n.Load(&pb.ModelOptions{ModelFile: missing})
Expect(err).To(HaveOccurred())
Expect(err.Error()).To(ContainSubstring("absent.gguf"))
})
It("reports a file that is not a GGUF", func() {
n := &NemoSpeech{}
path := filepath.Join(GinkgoT().TempDir(), "notagguf.gguf")
Expect(os.WriteFile(path, []byte("this is not a gguf file at all"), 0o600)).To(Succeed())
err := n.Load(&pb.ModelOptions{ModelFile: path})
Expect(err).To(HaveOccurred())
Expect(err.Error()).To(ContainSubstring("nemo-speech-cpp"))
})
// A failed load must not leave a family selected, or the RPC gate would wave
// requests through to a nil handle.
It("leaves no family selected when the load fails", func() {
n := &NemoSpeech{}
path := filepath.Join(GinkgoT().TempDir(), "broken.gguf")
Expect(os.WriteFile(path, []byte("broken"), 0o600)).To(Succeed())
Expect(n.Load(&pb.ModelOptions{ModelFile: path})).ToNot(Succeed())
Expect(n.fam).To(Equal(familyUnknown))
Expect(n.requireFamily(familyASR)).To(HaveOccurred())
})
// The load path picks a family and only then runs that family's loader, so
// there is a window where the family is known and the load still fails.
// Committing n.fam before the loader runs would leave the RPC gate open on a
// handle that was never created, and pkg/grpc/server.go keeps serving the
// instance after a failed LoadModel, so the next request really would reach
// it. TTS is the only family whose loader can fail before touching C.
It("does not select the family until that family's loader has succeeded", func() {
dir := GinkgoT().TempDir()
path := filepath.Join(dir, "magpie.f16.gguf")
writeGGUFWithArch(path, "magpietts")
// Self-guard: if the handwritten GGUF ever stops parsing, Load would fail
// at ggufArchitecture instead, before a family is ever chosen, and the
// assertions below would pass without exercising the ordering at all.
Expect(ggufArchitecture(path)).To(Equal("magpietts"))
// No sibling codec in the directory, so discoverTTSAssets fails after
// familyFor has already resolved familyTTS.
n := &NemoSpeech{}
err := n.Load(&pb.ModelOptions{ModelFile: path})
Expect(err).To(HaveOccurred())
Expect(err.Error()).To(ContainSubstring("codec_model"))
Expect(n.fam).To(Equal(familyUnknown))
Expect(n.requireFamily(familyTTS)).To(HaveOccurred())
})
It("closes the family gate again after a free", func() {
n := &NemoSpeech{fam: familyASR}
Expect(n.Free()).To(Succeed())
Expect(n.fam).To(Equal(familyUnknown))
Expect(n.requireFamily(familyASR)).To(HaveOccurred())
})
It("parses the model options before it touches the model file", func() {
// The options are what tell a TTS load where its codec lives, so they have
// to be in place before any family-specific loader runs.
n := &NemoSpeech{}
path := filepath.Join(GinkgoT().TempDir(), "broken.gguf")
Expect(os.WriteFile(path, []byte("broken"), 0o600)).To(Succeed())
Expect(n.Load(&pb.ModelOptions{
ModelFile: path,
ModelPath: "/models",
Options: []string{"gpu:2", "codec_model:codec.gguf"},
})).ToNot(Succeed())
Expect(n.opts.gpu).To(Equal(int32(2)))
Expect(n.opts.codecModel).To(Equal("/models/codec.gguf"))
})
})

View File

@@ -0,0 +1,368 @@
package main
import (
"regexp"
"runtime"
"strings"
"unsafe"
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
"github.com/mudler/xlog"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
)
// pairDirective matches a leading "[src->tgt] " override.
//
// Each side is an unbounded run of two-letter segments, not one or two of them.
// Either side may also be omitted, which keeps the model-level default for it:
// resolve_tag accepts a READY pair tag in one field with the other empty
// (src/nmt/langpairs.cc:167-172), so "[->en-de]" names a pair for one request.
//
// The two rules together are what force the unbounded run. A regional code on
// its own is only two segments (pt-br, zh-cn, es-us) and would parse under a
// stricter pattern; it is the SINGLE-FIELD form of a regional pair that runs to
// three (en-zh-cn, en-zh-tw, en-es-us, en-pt-br, pt-br-en, zh-tw-en). And the
// failure is not a mis-split: a pattern too short to cover the tag does not
// match the directive at all, so the whole bracket survives into the text and
// is handed to the model as something to translate.
//
// The codes are not normalised or validated here. normalize_language_code
// lowercases and folds BCP-47 down to a supported base, and is_supported has the
// authoritative table; duplicating either would be a second source of truth that
// drifts on the next pin bump.
var pairDirective = regexp.MustCompile(`^\[\s*([a-zA-Z]{2}(?:-[a-zA-Z]{2})*)?\s*->\s*([a-zA-Z]{2}(?:-[a-zA-Z]{2})*)?\s*\]\s*`)
// translator is the NMT half of the C API, narrowed to what Predict uses.
//
// It is an interface for the same reason synthesizer and diarStream are: no
// Riva-Translate GGUF is small enough to keep in the tree, so the layer above
// the ABI (pair resolution, validation, the text array, the single-chunk stream)
// would otherwise have no test at all. A fake here scripts what the C API
// returns; it does not pretend to translate anything.
type translator interface {
// translate returns one translation per input text, in order.
translate(texts []string, source, target string) ([]string, error)
}
// cTranslator is the real translator, over one nemo_speech_nmt_translator.
type cTranslator struct {
handle uintptr
}
// nmtTexts builds the `const char* const* texts` argument and returns it with
// the release the caller MUST defer.
//
// Two levels need pinning, not one. cstr pins each string's bytes, but the array
// carrying their addresses is a separate Go allocation holding uintptrs: the
// collector neither traces through it nor is obliged to leave it where it is,
// and C dereferences it for the whole call. Pinning only the strings would leave
// the array itself free to move out from under the runtime.
//
// An empty element is refused rather than passed on. cstr maps "" to NULL and
// src/nmt/c_api.cpp maps a NULL element back to "" (str_or_empty), so a blank
// text would come back as a confident translation of nothing rather than an
// error.
func nmtTexts(texts []string) ([]uintptr, func(), error) {
pin := new(runtime.Pinner)
// The pin is released first so that the array stops being pinned before the
// strings it points at do.
frees := []func(){pin.Unpin}
release := func() {
for _, f := range frees {
f()
}
}
if len(texts) == 0 {
return nil, release, status.Error(codes.InvalidArgument,
"nemo-speech-cpp: nothing to translate")
}
ptrs := make([]uintptr, len(texts))
for i, t := range texts {
if t == "" {
return nil, release, status.Error(codes.InvalidArgument,
"nemo-speech-cpp: nothing to translate")
}
p, free := cstr(t)
frees = append(frees, free)
ptrs[i] = p
}
pin.Pin(&ptrs[0])
return ptrs, release, nil
}
func (t *cTranslator) translate(texts []string, source, target string) ([]string, error) {
ptrs, release, err := nmtTexts(texts)
if err != nil {
release()
return nil, err
}
defer release()
// source and target cross as Go strings: purego NUL-terminates and copies
// them itself for the duration of the call, and c_api.cpp deep-copies both
// into std::string before doing anything with them.
var result uintptr
if st := NMTTranslate(t.handle, &ptrs[0], uint64(len(ptrs)), source, target, &result); st != 0 {
// An unsupported language pair arrives here as INVALID_ARGUMENT
// (src/nmt/translator.cpp throws std::invalid_argument, which
// src/nmt/c_api.cpp's guard maps to it), which statusErrorf turns into
// the caller-facing code rather than Internal.
return nil, statusErrorf(st, "nemo-speech-cpp: translate: %s", NMTLastError())
}
defer NMTResultDestroy(result)
count := NMTResultCount(result)
out := make([]string, 0, count)
for i := uint64(0); i < count; i++ {
out = append(out, NMTResultText(result, i))
}
return out, nil
}
// nmtTranslatorConfig builds the create-time config.
//
// Extracted from loadNMT so its four adjacent pointer fields can be asserted
// against distinct sentinels. Backend, Model, Generation and Pool are all
// uintptr and all sit next to each other, so transposing two of them changes
// neither the struct's size nor any field's offset: the layout assertions in
// abi_test.go are blind to it, and what it produces at runtime is the backend
// config being read as the model config.
//
// Generation and Pool stay NULL, which nmt.h documents as "library defaults":
// max_new_tokens (256) and contexts (1) are create-time settings this backend
// has no option to fill them from, and PredictOptions carries no per-request
// equivalent that a create-time config could honour anyway.
//
// backend and model are pinned addresses, not Go pointers, and the caller owns
// the pins.
func nmtTranslatorConfig(backend, model uintptr) cNMTTranslatorConfig {
return cNMTTranslatorConfig{
Size: unsafe.Sizeof(cNMTTranslatorConfig{}),
Backend: backend,
Model: model,
}
}
// loadNMT creates the Riva-Translate translator.
//
// This must not take engineMu: Load is its only caller and already holds it.
func (n *NemoSpeech) loadNMT(modelFile string) error {
// nemo_speech_nmt_create deep-copies the path into a std::string
// (src/nmt/c_api.cpp to_config, via str_or_empty) and retains no pointer
// afterwards, so pinning for the duration of the create call is both
// necessary and sufficient.
var pinner runtime.Pinner
defer pinner.Unpin()
pathP, freePath := cstr(modelFile)
defer freePath()
// NCtx is left at 0, which to_config reads as "keep the default" (it applies
// the field only when > 0) and which the runtime resolves to 1024 tokens.
// That is sized for the sentence-length input Riva-Translate is built for,
// and raising it costs one n_ctx-sized KV cache per pooled context, so it
// wants a deliberate option rather than a guess made here.
model := cNMTModelConfig{Size: unsafe.Sizeof(cNMTModelConfig{}), Path: pathP}
// BackendConfig.gpu defaults to 0 in C++ (device 0), not to CPU, and
// to_config assigns it unconditionally, so the option's own -1 default is
// what keeps an unconfigured model on the CPU.
backend := cNMTBackendConfig{Size: unsafe.Sizeof(cNMTBackendConfig{}), GPU: n.opts.gpu}
cfg := nmtTranslatorConfig(pinPtr(&pinner, &backend), pinPtr(&pinner, &model))
xlog.Info("nemo-speech-cpp: creating translator",
"gpu", n.opts.gpu,
"source_language", n.opts.sourceLanguage,
"target_language", n.opts.targetLanguage)
// #nosec G103 -- cfg is a local POD struct borrowed for this call only. Its
// Backend and Model members are pinPtr addresses held by the pinner unpinned
// on return, Model.Path is the cstr allocation freed by the defer above, and
// nemo_speech_nmt_create deep-copies everything it reads.
if st := NMTCreate(unsafe.Pointer(&cfg), &n.nmt); st != 0 {
return statusErrorf(st, "nemo-speech-cpp: nmt create: %s", NMTLastError())
}
return nil
}
// languagePair resolves the languages for one request and returns the text to
// translate.
//
// nemo_speech_nmt_translate takes explicit source and target languages and has
// no free-form generation entry point at all, so there is no prompt in the LLM
// sense to carry an instruction. The pair therefore comes from the model
// options, and a leading "[src->tgt]" directive is the only per-request control
// Predict can offer.
func (n *NemoSpeech) languagePair(prompt string) (source, target, text string) {
source, target = n.opts.sourceLanguage, n.opts.targetLanguage
m := pairDirective.FindStringSubmatch(prompt)
if m == nil {
return source, target, strings.TrimSpace(prompt)
}
// An omitted side keeps the model-level default rather than blanking it.
if m[1] != "" {
source = m[1]
}
if m[2] != "" {
target = m[2]
}
// The directive must not survive into the text: the runtime wraps it in a
// chat template (src/nmt/langpairs.cc build_prompt), so anything left here is
// translated along with the sentence.
return source, target, strings.TrimSpace(prompt[len(m[0]):])
}
// unsupportedPredictFields names the PredictOptions fields a caller may have set
// that this C API has no way to honour, so they are logged rather than silently
// dropped.
//
// The list is deliberately narrow. Everything nemo_speech_nmt_translate accepts
// is in its five arguments: a translator, the texts, and two language codes.
// Everything else in PredictOptions is therefore unsupported, and naming all of
// it would log on every single request, because LocalAI fills the sampling
// defaults in from the model config whether or not the user asked for them.
//
// So the sampling and decoding knobs (temperature, top_p, top_k, min_p, seed,
// tokens, repeat/frequency/presence penalties, mirostat, tfz, typical_p,
// stop_prompts, prompt caching, rope scaling, n_draft, logit_bias) are ignored
// silently: there is no field for any of them on either side of the ABI.
// max_new_tokens and n_ctx exist but are CREATE-time settings on the translator,
// not per-request ones, so PredictOptions.Tokens has nowhere to go either.
//
// What is named here is the structural asks: requests that only make sense
// against a general language model, where honouring them partially would be
// worse than saying nothing at all.
func unsupportedPredictFields(opts *pb.PredictOptions) []string {
var out []string
if opts.GetGrammar() != "" {
out = append(out, "grammar")
}
if opts.GetTools() != "" {
out = append(out, "tools")
}
if len(opts.GetImages()) > 0 {
out = append(out, "images")
}
if len(opts.GetVideos()) > 0 {
out = append(out, "videos")
}
if len(opts.GetAudios()) > 0 {
out = append(out, "audios")
}
if opts.GetNegativePrompt() != "" {
out = append(out, "negative_prompt")
}
if opts.GetLogprobs() > 0 {
out = append(out, "logprobs")
}
return out
}
// translateText runs one translation and returns it.
//
// The two rejections happen before anything crosses the ABI. An empty text would
// otherwise reach the runtime as a NULL element (see nmtTexts), and a missing
// target would come back as "unsupported language pair: -> ", which names
// neither the option the operator has to set nor the request that failed.
func translateText(t translator, source, target, text string) (string, error) {
if text == "" {
return "", status.Error(codes.InvalidArgument,
"nemo-speech-cpp: PredictOptions.prompt is required, it is the text to translate")
}
if target == "" {
return "", status.Error(codes.InvalidArgument,
"nemo-speech-cpp: no target language: set the target_language model option, "+
"or prefix the prompt with a [src->tgt] directive")
}
out, err := t.translate([]string{text}, source, target)
if err != nil {
return "", err
}
// One text in, one translation out. A call that returned OK with none is a
// runtime bug, and the empty string it would hand back reaches the user as a
// successful but blank completion with nothing anywhere to say why.
if len(out) == 0 {
return "", status.Error(codes.Internal, "nemo-speech-cpp: translation produced no result")
}
return out[0], nil
}
// streamTranslation runs one translation and puts the whole of it on out as a
// single chunk.
//
// That is a limit of the C API and not a shortcut taken here.
// nemo_speech_nmt_translate has no token callback and no incremental result: it
// returns once the decode has finished, with the completed text. There is
// nothing finer to stream, and splitting the finished string into fake chunks
// would imitate progress that never happened.
//
// out is not closed here. PredictStream owns it, and closing it in one of two
// places depending on how far the request got is how a stream ends up
// half-closed.
func streamTranslation(t translator, source, target, text string, out chan<- string) error {
translated, err := translateText(t, source, target, text)
if err != nil {
return err
}
out <- translated
return nil
}
// resolveRequest is the shared front half of both RPCs: it names what it is
// dropping and works out the pair and the text.
func (n *NemoSpeech) resolveRequest(opts *pb.PredictOptions) (source, target, text string) {
// Logged rather than rejected, for the reason the diarization path logs its
// own dropped fields: a caller that asked for something extra still wants the
// translation it can have, and a request naming a field this backend drops
// should say so where an operator can find it.
if dropped := unsupportedPredictFields(opts); len(dropped) > 0 {
xlog.Warn("nemo-speech-cpp: ignoring request fields this model has no equivalent for",
"fields", dropped)
}
return n.languagePair(opts.GetPrompt())
}
// Predict translates PredictOptions.Prompt.
//
// The whole body runs inside withEngine, so the family check and the C calls
// that trust the handle happen under a single acquisition of engineMu. See the
// handoff notes at the bottom of nemospeech.go: Free runs without the backend
// lock, so anything that checks the family and then releases the lock before
// calling C can have the handle destroyed underneath it.
func (n *NemoSpeech) Predict(opts *pb.PredictOptions) (string, error) {
var out string
if err := n.withEngine(familyNMT, func() error {
source, target, text := n.resolveRequest(opts)
s, err := translateText(&cTranslator{handle: n.nmt}, source, target, text)
out = s
return err
}); err != nil {
return "", err
}
return out, nil
}
// PredictStream translates PredictOptions.Prompt and emits the result on
// results.
//
// results is closed on EVERY path, including the family rejection and a
// validation failure, and the close is deferred outside withEngine so that a
// rejected family still closes it. This is the LEGACY streaming contract, which
// is the opposite of PredictStreamRich's: pkg/grpc/server.go:529 calls this and
// then blocks on a drain goroutine that only finishes when the channel closes,
// so a channel left open does not fail the request, it hangs the RPC and, with
// the backend lock still held, every request queued behind it. The rich variant
// is the one whose channel the host closes; this one is not.
func (n *NemoSpeech) PredictStream(opts *pb.PredictOptions, results chan string) error {
defer close(results)
return n.withEngine(familyNMT, func() error {
source, target, text := n.resolveRequest(opts)
return streamTranslation(&cTranslator{handle: n.nmt}, source, target, text, results)
})
}

View File

@@ -0,0 +1,416 @@
package main
import (
"errors"
"unsafe"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
)
// fakeTranslator scripts the C API's answer and records what it was asked, so
// the layer above the ABI (pair resolution, validation, the single-element text
// array) has a test at all. No Riva-Translate GGUF is small enough to keep in
// the tree, and this pretends to translate nothing.
type fakeTranslator struct {
texts []string
source, target string
calls int
out []string
err error
}
func (f *fakeTranslator) translate(texts []string, source, target string) ([]string, error) {
f.calls++
f.texts = texts
f.source = source
f.target = target
return f.out, f.err
}
// collectStrings drains ch until it closes and hands back everything it saw.
// The host does the same, so a channel this backend forgets to close hangs the
// RPC rather than failing it.
func collectStrings(ch chan string) chan []string {
done := make(chan []string, 1)
go func() {
var got []string
for s := range ch {
got = append(got, s)
}
done <- got
}()
return done
}
var _ = Describe("languagePair", func() {
It("uses the configured pair and returns the prompt unchanged", func() {
n := &NemoSpeech{opts: loadOptions{sourceLanguage: "en", targetLanguage: "de"}}
src, tgt, text := n.languagePair("hello world")
Expect(src).To(Equal("en"))
Expect(tgt).To(Equal("de"))
Expect(text).To(Equal("hello world"))
})
// nemo_speech_nmt_translate takes explicit languages and has no prompt path,
// so an inline directive is the only way a caller can pick a pair per request.
It("honours an inline pair directive and strips it from the text", func() {
n := &NemoSpeech{opts: loadOptions{sourceLanguage: "en", targetLanguage: "de"}}
src, tgt, text := n.languagePair("[en->fr] hello world")
Expect(src).To(Equal("en"))
Expect(tgt).To(Equal("fr"))
Expect(text).To(Equal("hello world"))
})
// The directive has to be gone from what reaches the model: the runtime
// wraps the text in a chat template (src/nmt/langpairs.cc build_prompt), so
// a leftover "[en->fr]" would be translated along with the sentence.
It("leaves no trace of the directive in the translated text", func() {
n := &NemoSpeech{opts: loadOptions{targetLanguage: "de"}}
_, _, text := n.languagePair("[en->fr] hello world")
Expect(text).ToNot(ContainSubstring("["))
Expect(text).ToNot(ContainSubstring("->"))
Expect(text).ToNot(ContainSubstring("fr"))
Expect(text).To(Equal("hello world"))
})
It("leaves an unparseable directive in the text", func() {
n := &NemoSpeech{opts: loadOptions{sourceLanguage: "en", targetLanguage: "de"}}
src, tgt, text := n.languagePair("[not a directive] hi")
Expect(src).To(Equal("en"))
Expect(tgt).To(Equal("de"))
Expect(text).To(Equal("[not a directive] hi"))
})
It("trims surrounding whitespace from the text", func() {
n := &NemoSpeech{opts: loadOptions{sourceLanguage: "en", targetLanguage: "de"}}
_, _, text := n.languagePair(" hello ")
Expect(text).To(Equal("hello"))
})
// The model's own tags carry region subtags (src/nmt/langpairs.cc: en-zh-cn,
// pt-br, es-us), so a directive that only accepted bare two-letter codes
// could not name half the pairs the runtime supports.
It("accepts a regional code on either side", func() {
n := &NemoSpeech{}
src, tgt, text := n.languagePair("[pt-br->en] ola")
Expect(src).To(Equal("pt-br"))
Expect(tgt).To(Equal("en"))
Expect(text).To(Equal("ola"))
src, tgt, _ = n.languagePair("[en->zh-cn] hi")
Expect(src).To(Equal("en"))
Expect(tgt).To(Equal("zh-cn"))
})
// resolve_tag accepts a ready pair tag in one field with the other empty, so
// a directive that names only one side must keep the configured value for the
// other rather than blanking it.
It("keeps the configured code for a side the directive omits", func() {
n := &NemoSpeech{opts: loadOptions{sourceLanguage: "en", targetLanguage: "de"}}
src, tgt, text := n.languagePair("[->fr] hello")
Expect(src).To(Equal("en"))
Expect(tgt).To(Equal("fr"))
Expect(text).To(Equal("hello"))
src, tgt, _ = n.languagePair("[fr->] hello")
Expect(src).To(Equal("fr"))
Expect(tgt).To(Equal("de"))
})
// resolve_tag (src/nmt/langpairs.cc:167-172) accepts a READY pair tag in one
// field with the other empty, and the model's own tags run to three segments
// (en-zh-cn, en-zh-tw, en-es-us, en-pt-br). That single-field three-segment
// form is the case a two-segment pattern cannot express: it does not merely
// mis-split the tag, it fails to match the directive at all, so the whole
// bracket survives into the text and is handed to the model as something to
// translate.
//
// Two-segment codes like pt-br and zh-cn are NOT this case; they parse either
// way.
It("accepts a three-segment pair tag given in one side of the directive", func() {
n := &NemoSpeech{opts: loadOptions{targetLanguage: "de"}}
src, tgt, text := n.languagePair("[->en-zh-cn] hi")
Expect(src).To(BeEmpty())
Expect(tgt).To(Equal("en-zh-cn"))
Expect(text).To(Equal("hi"))
src, tgt, text = n.languagePair("[pt-br-en->] hola")
Expect(src).To(Equal("pt-br-en"))
Expect(tgt).To(Equal("de"))
Expect(text).To(Equal("hola"))
})
It("does not treat a bracketed sentence as a directive", func() {
n := &NemoSpeech{opts: loadOptions{targetLanguage: "de"}}
_, _, text := n.languagePair("[see figure 1] the cat sat")
Expect(text).To(Equal("[see figure 1] the cat sat"))
})
})
var _ = Describe("nmtTranslatorConfig", func() {
// Backend, Model, Generation and Pool are four adjacent same-typed pointers.
// Transposing two of them changes neither the struct's size nor any field's
// offset, so the layout assertions in abi_test.go cannot see it, and the
// failure it produces is the runtime reading the backend config as the model
// config. Distinct sentinels are the only thing that catches it.
It("wires each pointer into its own field", func() {
cfg := nmtTranslatorConfig(0xB, 0xD)
Expect(cfg.Backend).To(Equal(uintptr(0xB)))
Expect(cfg.Model).To(Equal(uintptr(0xD)))
})
// NULL is what nmt.h documents as "library defaults" for a subsystem config,
// and this backend has no option to fill either of them from.
It("leaves the generation and pool configs null", func() {
cfg := nmtTranslatorConfig(0xB, 0xD)
Expect(cfg.Generation).To(BeZero())
Expect(cfg.Pool).To(BeZero())
})
// The runtime decides a field is present with HAS_FIELD, which tests the
// caller's size against offsetof + sizeof (src/nmt/c_api.cpp), so a config
// sent with Size 0 has every field ignored and the model loads from a path
// it was never given.
It("declares its own size", func() {
Expect(nmtTranslatorConfig(0xB, 0xD).Size).To(Equal(unsafe.Sizeof(cNMTTranslatorConfig{})))
})
})
var _ = Describe("nmtTexts", func() {
It("produces one non-null pointer per text", func() {
ptrs, release, err := nmtTexts([]string{"one", "two", "three"})
Expect(err).ToNot(HaveOccurred())
defer release()
Expect(ptrs).To(HaveLen(3))
for i, p := range ptrs {
Expect(p).ToNot(BeZero(), "texts[%d] must not be NULL", i)
}
// Distinct addresses: one buffer reused for every element would make the
// runtime translate the last text three times.
Expect(ptrs[0]).ToNot(Equal(ptrs[1]))
Expect(ptrs[1]).ToNot(Equal(ptrs[2]))
})
// cstr maps "" to NULL and src/nmt/c_api.cpp maps a NULL element back to "",
// so a blank text would be answered with a translation of nothing instead of
// an error.
It("refuses an empty element", func() {
_, release, err := nmtTexts([]string{"one", ""})
Expect(release).ToNot(BeNil())
release()
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
})
It("refuses an empty batch", func() {
_, release, err := nmtTexts(nil)
Expect(release).ToNot(BeNil())
release()
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
})
It("survives releasing more than once", func() {
_, release, err := nmtTexts([]string{"once"})
Expect(err).ToNot(HaveOccurred())
release()
Expect(release).ToNot(Panic())
})
})
var _ = Describe("translateText", func() {
It("passes the resolved pair and the text through to the runtime", func() {
f := &fakeTranslator{out: []string{"hallo welt"}}
got, err := translateText(f, "en", "de", "hello world")
Expect(err).ToNot(HaveOccurred())
Expect(got).To(Equal("hallo welt"))
Expect(f.texts).To(Equal([]string{"hello world"}))
Expect(f.source).To(Equal("en"))
Expect(f.target).To(Equal("de"))
})
// A single-pair model is configured with target_language alone, and
// resolve_tag accepts a ready tag in one field with the other empty.
It("allows an empty source language", func() {
f := &fakeTranslator{out: []string{"ciao"}}
_, err := translateText(f, "", "en-it", "hi")
Expect(err).ToNot(HaveOccurred())
Expect(f.source).To(BeEmpty())
})
It("rejects a missing target language and names the option to set", func() {
f := &fakeTranslator{}
_, err := translateText(f, "en", "", "hello")
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(err.Error()).To(ContainSubstring("target_language"))
Expect(f.calls).To(BeZero())
})
It("rejects an empty text without calling the runtime", func() {
f := &fakeTranslator{}
_, err := translateText(f, "en", "de", "")
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(f.calls).To(BeZero())
})
It("propagates a runtime failure", func() {
boom := errors.New("boom")
_, err := translateText(&fakeTranslator{err: boom}, "en", "de", "hello")
Expect(err).To(MatchError(boom))
})
// A call that returned OK with no translations is a runtime bug, and the
// empty string it would hand back reaches the user as a successful but blank
// completion with nothing anywhere to say why.
It("refuses a result that carries no translation", func() {
_, err := translateText(&fakeTranslator{}, "en", "de", "hello")
Expect(status.Code(err)).To(Equal(codes.Internal))
})
It("takes the first translation when the runtime returns several", func() {
f := &fakeTranslator{out: []string{"first", "second"}}
got, err := translateText(f, "en", "de", "hello")
Expect(err).ToNot(HaveOccurred())
Expect(got).To(Equal("first"))
})
})
var _ = Describe("Predict", func() {
It("refuses a model loaded as another family", func() {
n := &NemoSpeech{fam: familyASR}
out, err := n.Predict(&pb.PredictOptions{Prompt: "hello"})
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
Expect(out).To(BeEmpty())
// A lock leaked on the rejection path deadlocks the next request rather
// than failing it.
Expect(n.engineMu.TryLock()).To(BeTrue())
n.engineMu.Unlock()
})
It("refuses an unloaded model", func() {
n := &NemoSpeech{}
_, err := n.Predict(&pb.PredictOptions{Prompt: "hello"})
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
})
// The validation has to happen before anything crosses the ABI: nothing is
// loaded here, so a guard placed after the C call would panic on a nil
// function value instead of failing the request.
It("rejects an empty prompt before it reaches the runtime", func() {
n := &NemoSpeech{fam: familyNMT, opts: loadOptions{targetLanguage: "de"}}
var err error
Expect(func() {
_, err = n.Predict(&pb.PredictOptions{})
}).ToNot(Panic())
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
})
It("rejects a request with no target language, before it reaches the runtime", func() {
n := &NemoSpeech{fam: familyNMT}
var err error
Expect(func() {
_, err = n.Predict(&pb.PredictOptions{Prompt: "hello"})
}).ToNot(Panic())
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(err.Error()).To(ContainSubstring("target_language"))
})
})
var _ = Describe("PredictStream", func() {
// pkg/grpc/server.go drains this channel from a goroutine and then blocks on
// that goroutine finishing, so a channel left open does not fail the request,
// it hangs the RPC and every request queued behind the backend lock.
It("closes the channel when the family does not match", func() {
n := &NemoSpeech{fam: familyTTS}
ch := make(chan string)
done := collectStrings(ch)
err := n.PredictStream(&pb.PredictOptions{Prompt: "hello"}, ch)
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
Expect(<-done).To(BeEmpty())
Expect(n.engineMu.TryLock()).To(BeTrue())
n.engineMu.Unlock()
})
It("closes the channel when the request is rejected", func() {
n := &NemoSpeech{fam: familyNMT}
ch := make(chan string)
done := collectStrings(ch)
err := n.PredictStream(&pb.PredictOptions{Prompt: "hello"}, ch)
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(<-done).To(BeEmpty())
})
It("closes the channel on an unloaded model", func() {
n := &NemoSpeech{}
ch := make(chan string)
done := collectStrings(ch)
Expect(n.PredictStream(&pb.PredictOptions{Prompt: "hi"}, ch)).ToNot(Succeed())
Expect(<-done).To(BeEmpty())
})
// The C API has no token callback, so the whole translation is one chunk.
// The seam is the only place that can be asserted without a model.
It("emits the whole translation as a single chunk", func() {
f := &fakeTranslator{out: []string{"hallo welt"}}
ch := make(chan string)
done := collectStrings(ch)
Expect(streamTranslation(f, "en", "de", "hello world", ch)).To(Succeed())
close(ch)
Expect(<-done).To(Equal([]string{"hallo welt"}))
})
It("emits nothing when the translation fails", func() {
f := &fakeTranslator{err: errors.New("boom")}
ch := make(chan string)
done := collectStrings(ch)
Expect(streamTranslation(f, "en", "de", "hello", ch)).ToNot(Succeed())
close(ch)
Expect(<-done).To(BeEmpty())
})
})
var _ = Describe("unsupportedPredictFields", func() {
It("names nothing for a plain translation request", func() {
Expect(unsupportedPredictFields(&pb.PredictOptions{Prompt: "hello"})).To(BeEmpty())
})
// The sampling knobs are deliberately absent from this list: LocalAI fills
// them in from the model config on every request, so warning about them
// would log on every translation and say nothing.
It("stays quiet about sampling parameters the runtime has no field for", func() {
Expect(unsupportedPredictFields(&pb.PredictOptions{
Prompt: "hello",
Temperature: 0.7,
TopP: 0.9,
TopK: 40,
Seed: 42,
Tokens: 256,
})).To(BeEmpty())
})
It("names the asks the C API cannot serve at all", func() {
got := unsupportedPredictFields(&pb.PredictOptions{
Prompt: "hello",
Grammar: "root ::= x",
Tools: `[{"type":"function"}]`,
Images: []string{"a.png"},
Videos: []string{"a.mp4"},
Audios: []string{"a.wav"},
NegativePrompt: "no",
Logprobs: 3,
})
Expect(got).To(ConsistOf("grammar", "tools", "images", "videos", "audios",
"negative_prompt", "logprobs"))
})
})

View File

@@ -0,0 +1,100 @@
package main
import (
"path/filepath"
"strconv"
"strings"
"github.com/mudler/xlog"
)
// loadOptions holds the parsed model-level options. Path fields are resolved
// against ModelOptions.ModelPath at parse time so every consumer sees an
// absolute path.
type loadOptions struct {
// ASR
vadModel string
pncModel string
diarModel string
itnDir string
languageCode string
// TTS
codecModel string
tokenizerDir string
tnDir string
// NMT
sourceLanguage string
targetLanguage string
// gpu is the device index passed to the runtime's backend config.
// -1 selects CPU, matching the C API's own sentinel.
gpu int32
}
// splitOption splits on the FIRST colon so values may themselves contain one.
func splitOption(o string) (key, value string, ok bool) {
i := strings.Index(o, ":")
if i < 0 {
return "", "", false
}
return strings.TrimSpace(o[:i]), strings.TrimSpace(o[i+1:]), true
}
// resolve makes a relative asset path absolute against the models directory.
// Empty stays empty so callers can test for "unset".
func resolve(base, p string) string {
if p == "" || filepath.IsAbs(p) {
return p
}
return filepath.Join(base, p)
}
// parseOptions reads the backend "key:value" option slice. Unknown keys are
// ignored rather than rejected, so a config written for a newer backend still
// loads on an older one.
func parseOptions(opts []string, modelPath string) loadOptions {
o := loadOptions{gpu: -1}
for _, oo := range opts {
key, value, ok := splitOption(oo)
if !ok {
continue
}
switch key {
case "vad_model":
o.vadModel = resolve(modelPath, value)
case "pnc_model":
o.pncModel = resolve(modelPath, value)
case "diar_model":
o.diarModel = resolve(modelPath, value)
case "itn_dir":
o.itnDir = resolve(modelPath, value)
case "language_code":
o.languageCode = value
case "codec_model":
o.codecModel = resolve(modelPath, value)
case "tokenizer_dir":
o.tokenizerDir = resolve(modelPath, value)
case "tn_dir":
o.tnDir = resolve(modelPath, value)
case "source_language":
o.sourceLanguage = value
case "target_language":
o.targetLanguage = value
case "gpu":
// An unknown key is ignored for forward compatibility, but a known key
// with an unparseable value is a typo, and this one fails expensively:
// the model still loads and still produces correct output, just on CPU
// and far slower, with nothing anywhere to say why.
n, err := strconv.ParseInt(value, 10, 32)
if err != nil {
xlog.Warn("nemo-speech-cpp: ignoring unparseable option value, falling back to CPU",
"key", key, "value", value)
continue
}
o.gpu = int32(n)
}
}
return o
}

View File

@@ -0,0 +1,70 @@
package main
import (
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
var _ = Describe("parseOptions", func() {
It("parses every known key", func() {
o := parseOptions([]string{
"vad_model:silero.gguf",
"pnc_model:pnc.gguf",
"diar_model:sortformer.gguf",
"itn_dir:tn_configs",
"language_code:es-ES",
"codec_model:nanocodec.gguf",
"tokenizer_dir:extracted",
"tn_dir:tn",
"source_language:en",
"target_language:de",
}, "/models")
Expect(o.vadModel).To(Equal("/models/silero.gguf"))
Expect(o.pncModel).To(Equal("/models/pnc.gguf"))
Expect(o.diarModel).To(Equal("/models/sortformer.gguf"))
Expect(o.itnDir).To(Equal("/models/tn_configs"))
Expect(o.languageCode).To(Equal("es-ES"))
Expect(o.codecModel).To(Equal("/models/nanocodec.gguf"))
Expect(o.tokenizerDir).To(Equal("/models/extracted"))
Expect(o.tnDir).To(Equal("/models/tn"))
Expect(o.sourceLanguage).To(Equal("en"))
Expect(o.targetLanguage).To(Equal("de"))
})
It("leaves absolute paths untouched", func() {
o := parseOptions([]string{"vad_model:/abs/silero.gguf"}, "/models")
Expect(o.vadModel).To(Equal("/abs/silero.gguf"))
})
It("ignores unknown keys and entries without a separator", func() {
o := parseOptions([]string{"nonsense", "unknown_key:value"}, "/models")
Expect(o).To(Equal(loadOptions{gpu: -1}))
})
It("trims whitespace around keys and values", func() {
o := parseOptions([]string{" language_code : en-US "}, "/models")
Expect(o.languageCode).To(Equal("en-US"))
})
It("keeps a value containing a colon intact", func() {
// URIs must survive the split on the FIRST colon.
o := parseOptions([]string{"tokenizer_dir:/a/b:c"}, "/models")
Expect(o.tokenizerDir).To(Equal("/a/b:c"))
})
It("leaves an empty value empty so callers can detect unset", func() {
o := parseOptions([]string{"vad_model:"}, "/models")
Expect(o.vadModel).To(BeEmpty())
})
It("defaults gpu to -1 meaning CPU", func() {
o := parseOptions(nil, "/models")
Expect(o.gpu).To(Equal(int32(-1)))
})
It("parses an explicit gpu index", func() {
o := parseOptions([]string{"gpu:0"}, "/models")
Expect(o.gpu).To(Equal(int32(0)))
})
})

View File

@@ -0,0 +1,161 @@
#!/bin/bash
#
# Bundle the nemo-speech-cpp-grpc binary, the five nemo_speech shared objects,
# the text-normalization stack on a WITH_NORM build, the core runtime libs
# (libc/libstdc++/libgomp + ld.so) and the GPU runtime for the active BUILD_TYPE
# so the package is self-contained. Mirrors backend/go/whisper/package.sh;
# run.sh routes the (CGO_ENABLED=0) binary through lib/ld.so so the packaged
# libc is used instead of the host's.
#
# Five, not three: ASR and NMT each ship a thin _c ABI shim plus the
# implementation DSO it depends on, while TTS ships one object with no _c
# suffix at all.
set -e
CURDIR=$(dirname "$(realpath "$0")")
REPO_ROOT="${CURDIR}/../../.."
mkdir -p "$CURDIR/package/lib"
cp -avf "$CURDIR/nemo-speech-cpp-grpc" "$CURDIR/package/"
cp -avf "$CURDIR/run.sh" "$CURDIR/package/"
# The runtime ships three C ABI shared objects, not one. All three are
# required: main.go dlopens them eagerly, so a package missing any of them
# fails at startup. ASR and NMT expose the ABI through a dedicated _c library;
# TTS compiles its c_api into libnemo_speech_tts itself and has no _c variant,
# hence the asymmetric list. purego.Dlopen resolves them via the
# NEMO_SPEECH_*_LIBRARY paths that run.sh points at lib/.
#
# libnemo_speech_asr and libnemo_speech_nmt are in the list because the matching
# _c shims carry a DT_NEEDED on them: dlopen of the shim fails without the
# implementation DSO alongside it.
for lib in libnemo_speech_asr_c libnemo_speech_asr libnemo_speech_tts libnemo_speech_nmt_c libnemo_speech_nmt; do
cp -avf "$CURDIR"/${lib}.so* "$CURDIR/package/lib/" 2>/dev/null || true
cp -avf "$CURDIR"/${lib}*.dylib "$CURDIR/package/lib/" 2>/dev/null || true
if ! ls "$CURDIR"/package/lib/${lib}.* >/dev/null 2>&1; then
echo "ERROR: ${lib} shared library not found in $CURDIR, run 'make' first" >&2
exit 1
fi
done
# Text normalization (WITH_NORM=ON, Linux only) links Sparrowhawk and OpenFST
# into libnemo_speech_asr.so. Those live in a project-local prefix that the
# Makefile stages here, so anything staged that is not a nemo_speech object is
# part of that stack. Absent on a WITH_NORM=OFF build, which is why this is a
# glob that tolerates no matches rather than a required list.
shopt -s nullglob
for so in "$CURDIR"/*.so "$CURDIR"/*.so.* "$CURDIR"/*.dylib; do
case "$(basename "$so")" in
libnemo_speech_*) continue ;;
esac
cp -avf "$so" "$CURDIR/package/lib/"
done
shopt -u nullglob
# Detect architecture and copy the core runtime libs the shared objects link
# against, plus the matching dynamic loader as lib/ld.so.
source "$CURDIR/../../../scripts/build/package-system-libs.sh" "$CURDIR/package/lib" ""
# Dependency-closure guard.
#
# The lists above are maintained by hand, and the WITH_NORM build in particular
# pulls in transitive dependencies nobody enumerated: Sparrowhawk drags in
# protobuf, re2 and absl, none of which package-system-libs.sh provides. Rather
# than hard-code that set, walk the DT_NEEDED entries of everything staged and
# copy whatever is still unresolved. On a WITH_NORM=OFF build the closure is
# already complete, so this copies nothing.
#
# Skipped deliberately: the core runtime set that package-system-libs.sh owns,
# and the GPU stack that package-gpu-libs.sh owns.
shopt -s nullglob
staged_libs=("$CURDIR"/package/lib/*.so*)
shopt -u nullglob
if [ "$(uname)" != "Darwin" ] && [ "${#staged_libs[@]}" -gt 0 ]; then
# No silent skip. If the closure cannot be checked, the package cannot be
# shown to be complete, and shipping an unverified one is the failure this
# guard exists to prevent.
if command -v readelf >/dev/null 2>&1; then
read_needed() { readelf -d "$1" 2>/dev/null | sed -n 's/.*(NEEDED).*\[\(.*\)\]/\1/p'; }
elif command -v objdump >/dev/null 2>&1; then
read_needed() { objdump -p "$1" 2>/dev/null | awk '$1 == "NEEDED" { print $2 }'; }
else
echo "ERROR: neither readelf nor objdump is available, so the dependency" >&2
echo " closure of ${#staged_libs[@]} staged libraries cannot be verified." >&2
echo " Install binutils in the build image; refusing to ship an" >&2
echo " unverified package." >&2
exit 1
fi
is_provided() {
case "$1" in
ld-linux*|libc.so.6|libstdc++.so.6|libgcc_s.so.1|libm.so.6|libgomp.so.1) return 0 ;;
libdl.so.2|librt.so.1|libpthread.so.0) return 0 ;;
libcuda*|libcudart*|libcublas*|libcublasLt*|libnvrtc*|libnvidia*) return 0 ;;
libamdhip*|libhsa*|librocm*|libze_*|libOpenCL*|libvulkan*) return 0 ;;
esac
[ -e "$CURDIR/package/lib/$1" ]
}
# Walk until the staged set stops growing. The glob below expands once per
# pass, so each pass advances the closure by exactly one dependency level;
# a copied library can itself pull in new dependencies.
#
# CLOSURE_MAX_PASSES is a runaway guard, not a depth limit. Exhausting it
# means the walk never converged and the package is therefore incomplete,
# which has to fail the build: a fixed pass count that just falls out of the
# loop would silently ship a package missing its deepest libraries, and
# libnemo_speech_asr -> sparrowhawk -> protobuf -> absl already runs several
# levels deep.
CLOSURE_MAX_PASSES="${CLOSURE_MAX_PASSES:-64}"
converged=0
for (( pass=1; pass<=CLOSURE_MAX_PASSES; pass++ )); do
missing=0
for so in "$CURDIR"/package/lib/*.so*; do
[ -f "$so" ] || continue
for need in $(read_needed "$so"); do
# Written as an if rather than "is_provided && continue" so a
# false return cannot trip set -e via the AND-list exit status.
if is_provided "$need"; then
continue
fi
# Resolve against the staging dir first, then the system loader.
src="$(LD_LIBRARY_PATH="$CURDIR:$CURDIR/package/lib:${LD_LIBRARY_PATH:-}" \
ldd "$so" 2>/dev/null | awk -v n="$need" '$1 == n { print $3 }' | head -1)"
if [ -z "$src" ] || [ ! -e "$src" ]; then
echo "ERROR: $(basename "$so") needs $need and it could not be resolved." >&2
echo " The packaged backend would fail to dlopen at runtime." >&2
exit 1
fi
cp -aLvf "$src" "$CURDIR/package/lib/$need"
missing=1
done
done
if [ "$missing" -eq 0 ]; then
converged=1
break
fi
done
if [ "$converged" -ne 1 ]; then
echo "ERROR: the dependency closure was still growing after" >&2
echo " $CLOSURE_MAX_PASSES passes, so the package is incomplete and" >&2
echo " would fail to dlopen at runtime. Refusing to ship it." >&2
exit 1
fi
fi
# Package GPU libraries (CUDA/ROCm/Intel/Vulkan loader + ICDs + drivers)
# based on BUILD_TYPE so the backend can reach the GPU without the runtime
# base image shipping those drivers.
GPU_LIB_SCRIPT="${REPO_ROOT}/scripts/build/package-gpu-libs.sh"
if [ -f "$GPU_LIB_SCRIPT" ]; then
echo "Packaging GPU libraries for BUILD_TYPE=${BUILD_TYPE:-cpu}..."
source "$GPU_LIB_SCRIPT" "$CURDIR/package/lib"
package_gpu_libs
fi
echo "Packaging completed successfully"
ls -liah "$CURDIR/package/" "$CURDIR/package/lib/"

View File

@@ -0,0 +1,28 @@
#!/bin/bash
set -e
CURDIR=$(dirname "$(realpath "$0")")
# The runtime splits its C ABI across three shared objects, so each gets its
# own override variable. main.go reads exactly these names.
if [ "$(uname)" = "Darwin" ]; then
export DYLD_LIBRARY_PATH="$CURDIR/lib:"$CURDIR":${DYLD_LIBRARY_PATH:-}"
export NEMO_SPEECH_ASR_LIBRARY="$CURDIR/lib/libnemo_speech_asr_c.dylib"
export NEMO_SPEECH_TTS_LIBRARY="$CURDIR/lib/libnemo_speech_tts.dylib"
export NEMO_SPEECH_NMT_LIBRARY="$CURDIR/lib/libnemo_speech_nmt_c.dylib"
else
export LD_LIBRARY_PATH="$CURDIR/lib:"$CURDIR":${LD_LIBRARY_PATH:-}"
export NEMO_SPEECH_ASR_LIBRARY="$CURDIR/lib/libnemo_speech_asr_c.so"
export NEMO_SPEECH_TTS_LIBRARY="$CURDIR/lib/libnemo_speech_tts.so"
export NEMO_SPEECH_NMT_LIBRARY="$CURDIR/lib/libnemo_speech_nmt_c.so"
fi
# If a self-contained ld.so was packaged, route through it so the
# packaged libc / libstdc++ are used instead of the host's (matches the
# whisper backend's runtime layout). Linux only.
if [ -f "$CURDIR/lib/ld.so" ]; then
echo "Using lib/ld.so"
exec "$CURDIR/lib/ld.so" "$CURDIR/nemo-speech-cpp-grpc" "$@"
fi
exec "$CURDIR/nemo-speech-cpp-grpc" "$@"

View File

@@ -0,0 +1,102 @@
package main
import (
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
)
// The status values every C entry point in this backend returns.
//
// There is no single C enum to mirror. asr.h:43-49, tts.h:38-44 and nmt.h:40-45
// each declare their own, and diar.h has none of its own at all: it includes
// asr.h and types every diarization function as nemo_speech_asr_status
// (diar.h:18). The names the three surfaces share carry the same numbers:
//
// value asr.h tts.h nmt.h
// 0 NEMO_SPEECH_ASR_OK NEMO_SPEECH_TTS_OK NEMO_SPEECH_NMT_OK
// 1 NEMO_SPEECH_ASR_ERROR_INVALID_... NEMO_SPEECH_TTS_ERROR_INVALID_... NEMO_SPEECH_NMT_ERROR_INVALID_...
// 2 NEMO_SPEECH_ASR_ERROR_OUT_OF_MEM. NEMO_SPEECH_TTS_ERROR_OUT_OF_MEM. NEMO_SPEECH_NMT_ERROR_OUT_OF_MEM.
// 3 NEMO_SPEECH_ASR_ERROR_RUNTIME NEMO_SPEECH_TTS_ERROR_RUNTIME NEMO_SPEECH_NMT_ERROR_RUNTIME
// 4 NEMO_SPEECH_ASR_ERROR_CANCELLED NEMO_SPEECH_TTS_ERROR_CANCELLED (not declared)
//
// The one divergence is 4, and it is an absence rather than a disagreement. ASR
// and TTS both drive a consumer callback that can ask for the work to stop, and
// cancellation is what they report when it does; nemo_speech_nmt_translate takes
// no callback and returns only when the decode has finished, so the NMT surface
// has no cancellation to name. That is why one table can serve all three: 4 is
// not some other NMT status that would be mislabelled, it is a value the NMT
// surface never produces.
//
// Recheck this table after an upstream pin bump. A status added to one header
// and not the others is exactly the shape of change that would break the single
// mapping, and nothing in the build or the linker can see it: purego binds by
// name, and the return value is a bare int32 on the Go side.
const (
statusOK int32 = 0
statusInvalidArgument int32 = 1
statusOutOfMemory int32 = 2
statusRuntime int32 = 3
statusCancelled int32 = 4
)
// statusCode maps a C status onto the gRPC code the caller should be told.
//
// What the mapping is really carrying is whose mistake the failure was.
// INVALID_ARGUMENT is what every guard in src/{asr,tts,nmt}/c_api.cpp returns
// for a std::invalid_argument from the runtime, and the things that throw it are
// requests: an unknown voice_name (src/tts/synthesizer.cpp), an unsupported
// language pair (src/nmt/translator.cpp), an out-of-range sample rate. Reporting
// those as Internal turns a 400 into a 500 and sends the user hunting for a
// broken model or a broken backend instead of fixing the request.
//
// OUT_OF_MEMORY is a resource limit rather than a defect, which is what
// ResourceExhausted means, and it is the one failure a client can sensibly
// retry later or retry smaller. CANCELLED is the consumer having stopped
// listening, which is not a failure of this backend at all: the streaming sinks
// return false once their client is gone (see ttsDeliverPCM), and the runtime
// turns that into status 4.
//
// RUNTIME, and anything a future pin adds that this table has not been taught,
// stay Internal. An unrecognised status is precisely the case where the backend
// does not know whose fault it was, and Internal is the honest answer.
func statusCode(st int32) codes.Code {
switch st {
case statusOK:
return codes.OK
case statusInvalidArgument:
return codes.InvalidArgument
case statusOutOfMemory:
return codes.ResourceExhausted
case statusCancelled:
return codes.Canceled
case statusRuntime:
// Named rather than folded into the default so this switch reads as the
// whole enum. A status the table has never heard of is a different thing
// from a runtime error even though both answer Internal, and a reader
// checking the mapping against the headers should not have to work out
// which arm RUNTIME lands in.
return codes.Internal
default:
return codes.Internal
}
}
// statusErrorf builds the gRPC error for a failed C call.
//
// Every C call site in this backend goes through this rather than through
// status.Errorf directly, and that is the whole point of it existing: the
// mapping used to be written out at exactly one of sixteen call sites, so the
// same backend answered an unsupported language pair with InvalidArgument and an
// unknown TTS voice, which is the same class of caller mistake against the same
// process, with Internal.
//
// The OK guard is not defensive noise. status.Errorf(codes.OK, ...) returns a
// nil error, so a call site that built its error without first checking the
// status would report a hard C failure as a successful request with no
// diagnostic anywhere. Returning Internal instead keeps that mistake loud.
func statusErrorf(st int32, format string, args ...any) error {
if st == statusOK {
return status.Errorf(codes.Internal, format, args...)
}
return status.Errorf(statusCode(st), format, args...)
}

View File

@@ -0,0 +1,109 @@
package main
import (
"os"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
)
var _ = Describe("C status mapping", func() {
// The whole enum, so a status that quietly moves to a different code is
// visible here rather than in a bug report about an HTTP 500. The names on
// the left are transcribed from asr.h:43-49, tts.h:38-44 and nmt.h:40-45;
// see the table in status.go for how the three surfaces line up.
DescribeTable("maps each declared C status onto a gRPC code",
func(st int32, want codes.Code) {
Expect(statusCode(st)).To(Equal(want))
},
Entry("OK", statusOK, codes.OK),
Entry("INVALID_ARGUMENT", statusInvalidArgument, codes.InvalidArgument),
Entry("OUT_OF_MEMORY", statusOutOfMemory, codes.ResourceExhausted),
Entry("RUNTIME", statusRuntime, codes.Internal),
Entry("CANCELLED", statusCancelled, codes.Canceled),
)
// A pin bump that adds a status this table has never been taught must not
// guess. Internal is the honest answer when the backend does not know whose
// mistake the failure was.
DescribeTable("reports an unknown status as Internal",
func(st int32) {
Expect(statusCode(st)).To(Equal(codes.Internal))
},
Entry("one past the last declared value", int32(5)),
Entry("far past it", int32(99)),
Entry("negative", int32(-1)),
)
It("carries the mapped code and the formatted message into the error", func() {
err := statusErrorf(statusInvalidArgument, "nemo-speech-cpp: %s: %d", "synthesize", 7)
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(err.Error()).To(ContainSubstring("nemo-speech-cpp: synthesize: 7"))
})
// status.Errorf(codes.OK, ...) returns nil, so a call site that built its
// error without checking the status first would turn a hard C failure into a
// silent success with no diagnostic anywhere.
It("never returns nil, not even for OK", func() {
err := statusErrorf(statusOK, "nemo-speech-cpp: should not happen")
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.Internal))
})
})
// These drive real C statuses out of the real shared objects, one per family,
// rather than asserting the Go mapping against itself.
//
// A NULL handle is the one failure every surface can be provoked into without a
// model: nemo_speech_asr_recognize_f32 and nemo_speech_nmt_translate check the
// handle up front, nemo_speech_tts_synthesize_text does the same, and the
// diarization stream entry points throw std::invalid_argument for a dead stream,
// which src/asr/c_api.cpp's guard maps to the same status. All four are the
// caller's mistake, and the point of the exercise is that all four now come back
// as InvalidArgument instead of Internal.
var _ = Describe("C status mapping at the call sites", func() {
BeforeEach(func() {
if !librariesPresent() {
if requireLibs() {
cwd, _ := os.Getwd()
Fail("NEMO_SPEECH_REQUIRE_LIBS=1 but the shared libraries are not in " + cwd +
": these specs are the ABI defence and must not be skipped." +
" Run make -C backend/go/nemo-speech-cpp stage-libs")
}
Skip("shared libraries not built, run make in backend/go/nemo-speech-cpp")
}
Expect(openLibraries()).To(Succeed())
})
It("reports an ASR INVALID_ARGUMENT as InvalidArgument", func() {
opts := ASRRecognitionOptionsDef()
// Non-empty PCM on purpose: recognizeF32 rejects an empty slice itself,
// which would prove nothing about what the C side returned.
_, err := recognizeF32(0, &opts, []float32{0, 0, 0}, 16000)
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
})
It("reports a diarization INVALID_ARGUMENT as InvalidArgument", func() {
err := (&cDiarStream{}).finish()
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
})
It("reports a TTS INVALID_ARGUMENT as InvalidArgument", func() {
s := &cSynthesizer{}
err := s.synthesize(&pb.TTSRequest{Text: "hello"}, "en", func([]byte) bool { return true })
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
})
It("reports an NMT INVALID_ARGUMENT as InvalidArgument", func() {
_, err := (&cTranslator{}).translate([]string{"hello"}, "en", "de")
Expect(err).To(HaveOccurred())
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
})
})

View File

@@ -0,0 +1,600 @@
package main
import (
"bytes"
"math"
"os"
"runtime"
"strconv"
"sync"
"unsafe"
"github.com/ebitengine/purego"
laudio "github.com/mudler/LocalAI/pkg/audio"
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
"github.com/mudler/xlog"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
)
// The backend-preference enum from include/nemo_speech/tts.h. The C type is an
// enum, which this toolchain lays out as int32, so the values are written here
// rather than inferred.
const (
ttsBackendAuto int32 = 0
ttsBackendCPU int32 = 1
)
// maxWAVDataBytes is the largest PCM payload a RIFF WAV can describe.
//
// Both size fields in the header are uint32, so a longer payload would not
// merely be unusual, it would wrap and produce a file whose header disagrees
// with its contents. At 22.05 kHz mono 16-bit that ceiling is about 27 hours of
// speech, so nothing real is being refused.
//
// Typed int64 rather than left untyped so the comparison below is the same one
// on every architecture: an untyped constant this large does not fit in a
// 32-bit int and would not compile there at all.
const maxWAVDataBytes int64 = math.MaxUint32 - laudio.WAVHeaderSize
// wavStreamingSize is the placeholder both size fields carry while the total
// length is still unknown. Players read it as "stream until the socket closes".
const wavStreamingSize = 0xFFFFFFFF
// ttsSink receives one PCM chunk, already copied into Go memory.
//
// It returns false to cancel the synthesis in progress: that is the C
// callback's only way to stop work early, and the runtime turns it into
// NEMO_SPEECH_TTS_ERROR_CANCELLED.
type ttsSink func(pcm []byte) bool
// ttsSinkTable maps the user_data value handed to C back to the Go sink the
// chunk belongs to.
//
// A single "current sink" pointer would be enough for one model, since every
// RPC holds that model's engineMu for its whole body. It is not enough for the
// process: engineMu is per-NemoSpeech, one backend process can hold several
// loaded models, and the callback below is shared by all of them, so two TTS
// models synthesizing at once would overwrite each other's sink. The id is what
// keeps them apart.
//
// The id is an integer and never a Go pointer. user_data crosses into C as a
// void*, which the collector does not trace, so a Go pointer parked there would
// have exactly the lifetime problem cstr documents.
type ttsSinkTable struct {
mu sync.Mutex
next uintptr
sinks map[uintptr]ttsSink
}
var ttsSinks = &ttsSinkTable{sinks: map[uintptr]ttsSink{}}
// register adds sink and returns its id together with the release the caller
// MUST defer. Ids start at 1 so a zeroed or stale user_data cannot resolve to
// somebody else's sink.
func (t *ttsSinkTable) register(sink ttsSink) (uintptr, func()) {
t.mu.Lock()
defer t.mu.Unlock()
t.next++
id := t.next
t.sinks[id] = sink
return id, func() {
t.mu.Lock()
defer t.mu.Unlock()
delete(t.sinks, id)
}
}
// lookup returns the sink for id, or nil once it has been released.
func (t *ttsSinkTable) lookup(id uintptr) ttsSink {
t.mu.Lock()
defer t.mu.Unlock()
return t.sinks[id]
}
var (
ttsCallbackOnce sync.Once
ttsCallbackFn uintptr
)
// ttsPCMCallback returns the C function pointer the runtime drives PCM through,
// compiling it on first use.
//
// Exactly one is ever created per process, and that is a hard requirement
// rather than a tidiness argument. purego.NewCallback writes into a fixed table
// of maxCB = 2000 entries (purego/syscall_sysv.go) and never releases an entry,
// so a callback compiled per request panics the whole backend process with
// "purego: the maximum number of callbacks has been reached" on the 2001st
// synthesis. Per model load is not safe either: a server that swaps models
// reaches the same ceiling, just later and even less predictably. Routing every
// synthesis through one callback plus a user_data id is what keeps the count at
// one for the life of the process.
func ttsPCMCallback() uintptr {
ttsCallbackOnce.Do(func() { ttsCallbackFn = purego.NewCallback(ttsDeliverPCM) })
return ttsCallbackFn
}
// ttsDeliverPCM is the body of that callback: nemo_speech_tts_pcm_callback,
// which the runtime invokes on its own thread for each chunk it produces.
//
// The bytes are copied rather than aliased. The pointer addresses a std::string
// the runtime owns and reuses for the next chunk (src/tts/c_api.cpp
// make_callback), so a slice over it would be rewritten under the consumer as
// soon as this returns.
func ttsDeliverPCM(pcm unsafe.Pointer, nBytes uint64, userData uintptr) bool {
// The table's lock is released before the sink runs, which matters because
// TTSStream's sink blocks on a channel send until its client drains it.
// Holding the lock across that would stall every other model's callback
// behind one slow consumer.
sink := ttsSinks.lookup(userData)
if sink == nil {
// The request that registered this sink has already returned, so there
// is nowhere to put the audio. false cancels rather than letting the
// runtime synthesize to completion into a consumer that stopped
// listening.
return false
}
// c_api.cpp filters empty chunks before calling us, so this is belt and
// braces: unsafe.Slice on a null pointer is what it protects against.
if pcm == nil || nBytes == 0 {
return true
}
buf := make([]byte, nBytes)
// #nosec G103 -- pcm and nBytes are the C-owned buffer and its length from
// one callback invocation, both null/zero-checked above. The slice is read
// only, its length is the length the runtime declared for that buffer, and it
// is copied into Go memory here and never retained past this return.
copy(buf, unsafe.Slice((*byte)(pcm), nBytes))
return sink(buf)
}
// synthesizer is the TTS half of the C API, narrowed to what the two RPCs use.
//
// It is an interface for the same reason asrSession and diarStream are: no
// MagpieTTS GGUF is small enough to keep in the tree, so the logic layered on
// top of the ABI (validation, the WAV framing, chunk ordering) would otherwise
// have no test at all. The seam is at the ABI, not at the model: a fake here
// scripts what the C API emits, it does not pretend to synthesize anything.
type synthesizer interface {
// sampleRate is the rate the PCM chunks arrive at.
sampleRate() int32
// synthesize maps req onto the runtime's per-request options and runs one
// synthesis, handing each chunk to sink as it is produced.
synthesize(req *pb.TTSRequest, defaultLanguage string, sink ttsSink) error
}
// cSynthesizer is the real synthesizer, over one nemo_speech_tts_synthesizer.
type cSynthesizer struct {
handle uintptr
}
func (s *cSynthesizer) sampleRate() int32 { return TTSSampleRate(s.handle) }
func (s *cSynthesizer) synthesize(req *pb.TTSRequest, defaultLanguage string, sink ttsSink) error {
// Started from the runtime's own defaults, not from a zero struct: every
// numeric field here is sentinel-sensitive (speaker/seed < 0, steps/top_k
// <= 0 all mean "use the synthesizer's value"), and a zeroed struct would
// read as speaker 0, seed 0 and zero decoding steps.
opts := TTSSynthesisOptionsDefault()
// A per-request language wins over the model-level default; both may be
// empty, which the runtime resolves to the synthesizer's own default.
language := req.GetLanguage()
if language == "" {
language = defaultLanguage
}
langP, freeLang := cstr(language)
defer freeLang()
opts.LanguageCode = langP
speaker, voiceName := resolveSpeaker(req.GetVoice())
opts.Speaker = speaker
voiceP, freeVoice := cstr(voiceName)
defer freeVoice()
opts.VoiceName = voiceP
applySynthesisParams(&opts, req.GetParams())
id, release := ttsSinks.register(sink)
defer release()
// stats_out is NULL: nemo_speech_tts_synthesis_stats is 300-odd bytes of
// timing detail with nowhere to go on either RPC, and the C API documents
// NULL as the way to decline it.
// #nosec G103 -- opts is a local POD struct borrowed for this call only. Its
// two uintptr members (LanguageCode, VoiceName) are cstr allocations pinned
// by the defers above, and this entry point is synchronous, so it returns
// before those pins are released even though the callbacks run off-thread.
st := TTSSynthesizeText(s.handle, unsafe.Pointer(&opts), req.GetText(), ttsPCMCallback(), id, nil)
if st != 0 {
// An unknown voice_name arrives here as INVALID_ARGUMENT
// (src/tts/synthesizer.cpp throws std::invalid_argument, which
// src/tts/c_api.cpp's guard maps to it), and a consumer that stopped
// reading arrives as CANCELLED. Neither is this backend's failure, so
// neither goes out as Internal.
return statusErrorf(st, "nemo-speech-cpp: synthesize: %s", TTSLastError())
}
return nil
}
// resolveSpeaker splits a request's voice into the two fields the C API has for
// it: a speaker index and a voice name.
//
// nemo_speech_tts_synthesis_options.voice_name is documented as ignored
// whenever speaker >= 0, and src/tts/synthesizer.cpp only calls resolve_speaker
// when options.speaker is negative, so the two are alternatives and never a
// pair. A named voice must therefore leave the index at -1 or the name is
// silently dropped.
//
// The numeric split cannot change what the runtime picks: resolve_speaker parses
// a numeric voice_name itself, so anything this function passes through as a
// name and that happens to be a number lands on the same speaker anyway. What it
// must not do is let a NEGATIVE number through as an index. "-1" is not a
// speaker, it is the sentinel for "use the default", and treating it as an index
// would turn a request naming an invalid voice into one that quietly synthesizes
// in the default voice instead of being rejected.
func resolveSpeaker(voice string) (int32, string) {
if voice == "" {
return -1, ""
}
if idx, err := strconv.ParseInt(voice, 10, 32); err == nil && idx >= 0 {
return int32(idx), ""
}
return -1, voice
}
// applySynthesisParams maps TTSRequest.params onto the runtime's per-request
// options.
//
// Only the five knobs nemo_speech_tts_synthesis_options actually has are read.
// An unset or unparseable value leaves the field alone rather than resetting it:
// the struct arrives carrying the runtime's defaults, and params is documented
// as "unset leaves the backend's configured defaults".
//
// The sentinels are the reason each write is guarded rather than unconditional.
// src/tts/magpietts/runtime.cpp takes the request's seed only when it is >= 0
// and its steps and top_k only when they are > 0, so writing a parsed 0 or a
// negative would not merely be ignored, it would erase the option's meaning for
// a caller who passed "0" expecting something.
//
// temperature and cfg_scale each need their override flag set as well. The
// runtime reads the float only when the flag is true and otherwise falls back to
// the synthesizer's config, so a temperature written without its flag is
// silently discarded.
func applySynthesisParams(o *cTTSSynthesisOptions, params map[string]string) {
if len(params) == 0 {
return
}
if v, ok := parseInt32Param(params["seed"]); ok && v >= 0 {
o.Seed = v
}
if v, ok := parseInt32Param(params["steps"]); ok && v > 0 {
o.Steps = v
}
if v, ok := parseInt32Param(params["top_k"]); ok && v > 0 {
o.TopK = v
}
if v, ok := parseFloat32Param(params["temperature"]); ok {
o.Temperature = v
o.OverrideTemperature = true
}
if v, ok := parseFloat32Param(params["cfg_scale"]); ok {
o.CFGScale = v
o.OverrideCFGScale = true
}
}
// parseInt32Param reads one params entry. ok is false for an absent or
// unparseable value, which the caller reads as "leave the default".
func parseInt32Param(v string) (int32, bool) {
if v == "" {
return 0, false
}
n, err := strconv.ParseInt(v, 10, 32)
if err != nil {
xlog.Warn("nemo-speech-cpp: ignoring unparseable TTS parameter", "value", v)
return 0, false
}
return int32(n), true
}
func parseFloat32Param(v string) (float32, bool) {
if v == "" {
return 0, false
}
f, err := strconv.ParseFloat(v, 32)
if err != nil {
xlog.Warn("nemo-speech-cpp: ignoring unparseable TTS parameter", "value", v)
return 0, false
}
return float32(f), true
}
// ttsModelConfig builds the create-time model config.
//
// Extracted from loadTTS and asserted field by field because three adjacent
// members of nemo_speech_tts_model_config are same-typed paths. Swapping two of
// them changes neither the struct's size nor any field's offset, so the layout
// assertions in abi_test.go cannot see it, and the failure it produces is the
// runtime loading the codec as the acoustic model.
//
// Every argument is a C pointer from cstr, not a Go string, and the caller owns
// the releases. tnDir may be null: text_normalizer_model_dir is optional and an
// empty one leaves the text unchanged.
func ttsModelConfig(magpieModel, codecModel, tokenizerDir, tnDir uintptr) cTTSModelConfig {
return cTTSModelConfig{
Size: unsafe.Sizeof(cTTSModelConfig{}),
MagpieModel: magpieModel,
CodecModel: codecModel,
TokenizerModelDir: tokenizerDir,
TextNormalizerModelDir: tnDir,
}
}
// ttsRuntimeBackend maps the backend's gpu option onto the TTS runtime's
// backend preference.
//
// nemo_speech_tts_runtime_config has no device index at all, only a three-way
// AUTO/CPU/CUDA preference, so a gpu option naming a particular device cannot be
// honoured and AUTO is the honest answer for it. A negative gpu is different: it
// is the option's documented "CPU" across this whole backend (asr.h: "-1 = CPU")
// and it is also the default, so it has to pin the preference rather than leave
// the runtime free to pick CUDA.
func ttsRuntimeBackend(gpu int32) int32 {
if gpu < 0 {
return ttsBackendCPU
}
return ttsBackendAuto
}
// loadTTS creates the MagpieTTS synthesizer.
//
// It runs after discoverTTSAssets, so codecModel and tokenizerDir are already
// resolved and non-empty; tnDir stays optional.
//
// This must not take engineMu: Load is its only caller and already holds it.
func (n *NemoSpeech) loadTTS(modelFile string) error {
// nemo_speech_tts_create deep-copies every const char* into a std::string
// (src/tts/c_api.cpp, via str_or_empty) and keeps no pointer afterwards, so
// pinning across the create call is both necessary and sufficient.
var pinner runtime.Pinner
defer pinner.Unpin()
magpieP, freeMagpie := cstr(modelFile)
defer freeMagpie()
codecP, freeCodec := cstr(n.opts.codecModel)
defer freeCodec()
tokenizerP, freeTokenizer := cstr(n.opts.tokenizerDir)
defer freeTokenizer()
tnP, freeTN := cstr(n.opts.tnDir)
defer freeTN()
model := ttsModelConfig(magpieP, codecP, tokenizerP, tnP)
rt := TTSRuntimeConfigDefault()
backend := ttsRuntimeBackend(n.opts.gpu)
rt.LTBackend = backend
rt.SamplingBackend = backend
// The codec is a separate graph with its own placement, so a CPU-only
// request has to say so here too or it would still try to run on the GPU.
rt.CodecCPU = backend == ttsBackendCPU
langP, freeLang := cstr(n.opts.languageCode)
defer freeLang()
cfg := cTTSSynthesizerConfig{
Size: unsafe.Sizeof(cTTSSynthesizerConfig{}),
Model: pinPtr(&pinner, &model),
Runtime: pinPtr(&pinner, &rt),
DefaultLanguageCode: langP,
}
xlog.Info("nemo-speech-cpp: creating synthesizer",
"gpu", n.opts.gpu,
"codec", n.opts.codecModel,
"tokenizer", n.opts.tokenizerDir,
"text_normalizer", n.opts.tnDir != "")
// Compiled before the handle exists so that a full callback table fails the
// load, where the operator can see it, rather than the first synthesis.
ttsPCMCallback()
// #nosec G103 -- cfg is a local POD struct borrowed for this call only. Model
// and Runtime are pinPtr addresses held by the pinner unpinned on return, the
// paths they carry are cstr allocations freed by the defers above, and
// nemo_speech_tts_create deep-copies every string it reads.
if st := TTSCreate(unsafe.Pointer(&cfg), &n.synth); st != 0 {
return statusErrorf(st, "nemo-speech-cpp: tts create: %s", TTSLastError())
}
return nil
}
// validateTTSRequest rejects what the runtime would reject, before anything
// crosses the ABI, and names the fields this backend drops.
//
// Empty text is checked here rather than left to the C side for the error code:
// src/tts/synthesizer.cpp throws "text is required", which arrives as a status
// this layer would otherwise report as Internal, and an empty prompt is a client
// mistake, not a backend failure.
//
// instructions is logged rather than rejected, for the reason the diarization
// path logs its own dropped fields: a caller that asked for an expressive style
// still wants the audio it can have, and a request naming something this backend
// silently ignores should say so where an operator can find it. There is nothing
// to map it onto, because MagpieTTS conditions on a speaker, not on a prose
// style description: nemo_speech_tts_synthesis_options has speaker and
// voice_name and no free-text field at all.
func validateTTSRequest(req *pb.TTSRequest) error {
if req.GetText() == "" {
return status.Error(codes.InvalidArgument, "nemo-speech-cpp: TTSRequest.text is required")
}
if req.GetInstructions() != "" {
xlog.Warn("nemo-speech-cpp: ignoring TTSRequest.instructions, this model has no equivalent")
}
return nil
}
// outputSampleRate reads the rate the synthesizer emits at.
//
// A non-positive rate is refused rather than passed on. nemo_speech_tts_sample_rate
// answers 0 for a null handle, and a WAV header carrying 0 is not a slightly
// wrong file, it is one no player can decode and one whose duration is
// undefined.
func outputSampleRate(s synthesizer) (uint32, error) {
rate := s.sampleRate()
if rate <= 0 {
return 0, status.Error(codes.Internal,
"nemo-speech-cpp: the synthesizer reported no sample rate")
}
return uint32(rate), nil
}
// wavFile frames PCM as a complete WAV: a header with real sizes, then the
// samples.
//
// pcm is little-endian signed 16-bit mono, which is what the runtime's callback
// delivers, and is exactly what pkg/audio's header describes, so nothing is
// converted on the way through.
func wavFile(pcm []byte, sampleRate uint32) ([]byte, error) {
if int64(len(pcm)) > maxWAVDataBytes {
return nil, status.Errorf(codes.Internal,
"nemo-speech-cpp: synthesis produced %d bytes, more than a WAV header can describe", len(pcm))
}
var buf bytes.Buffer
// #nosec G115 -- len(pcm) is checked against maxWAVDataBytes (MaxUint32 minus
// the header) immediately above, so the narrowing to uint32 cannot wrap.
h := laudio.NewWAVHeaderWithRate(uint32(len(pcm)), sampleRate)
if err := h.Write(&buf); err != nil {
return nil, status.Errorf(codes.Internal, "nemo-speech-cpp: write WAV header: %v", err)
}
buf.Write(pcm)
return buf.Bytes(), nil
}
// streamingWAVHeader is the first chunk of a streamed synthesis: the same header
// with both sizes left unknown, since the total length is not known until the
// synthesis ends.
//
// NewWAVHeaderWithRate derives ChunkSize from the payload length, so the RIFF
// size has to be overwritten as well: 36 + 0xFFFFFFFF wraps to 35, which is a
// smaller number than the header itself.
func streamingWAVHeader(sampleRate uint32) []byte {
h := laudio.NewWAVHeaderWithRate(wavStreamingSize, sampleRate)
h.ChunkSize = wavStreamingSize
var buf bytes.Buffer
// Write only fails on the writer, and bytes.Buffer does not fail.
_ = h.Write(&buf)
return buf.Bytes()
}
// synthesizeWAV runs one synthesis and writes the whole result to dst.
func synthesizeWAV(s synthesizer, req *pb.TTSRequest, defaultLanguage string) error {
if err := validateTTSRequest(req); err != nil {
return err
}
if req.GetDst() == "" {
return status.Error(codes.InvalidArgument,
"nemo-speech-cpp: TTSRequest.dst (output path) is required")
}
// Read before the synthesis rather than after: it is what the header is
// built from, and failing on a bad handle here costs nothing, where failing
// after costs the whole synthesis.
rate, err := outputSampleRate(s)
if err != nil {
return err
}
var pcm []byte
err = s.synthesize(req, defaultLanguage, func(chunk []byte) bool {
pcm = append(pcm, chunk...)
return true
})
if err != nil {
return err
}
// A synthesis that returned OK having emitted nothing is a runtime bug, but
// the file it would produce is a valid empty WAV, which reaches the user as
// silence with no error anywhere.
if len(pcm) == 0 {
return status.Error(codes.Internal, "nemo-speech-cpp: synthesis produced no audio")
}
out, err := wavFile(pcm, rate)
if err != nil {
return err
}
if err := os.WriteFile(req.GetDst(), out, 0o600); err != nil {
return status.Errorf(codes.Internal, "nemo-speech-cpp: write %q: %v", req.GetDst(), err)
}
return nil
}
// streamWAV runs one synthesis and emits a WAV header followed by each PCM
// chunk as the runtime produces it.
//
// The header is the backend's job, not the caller's: pkg/grpc/server.go only
// ever sets Reply.Audio on this path and never Reply.Message, and
// core/backend/tts.go's own header branch is keyed on Message, so a backend that
// emitted bare PCM would stream something no client could decode. sherpa-onnx
// and magpie-tts-cpp both do the same.
//
// out is not closed here. TTSStream owns it, and closing it in one of two places
// depending on how far the request got is how a stream ends up half-closed.
func streamWAV(s synthesizer, req *pb.TTSRequest, defaultLanguage string, out chan<- []byte) error {
if err := validateTTSRequest(req); err != nil {
return err
}
rate, err := outputSampleRate(s)
if err != nil {
return err
}
out <- streamingWAVHeader(rate)
return s.synthesize(req, defaultLanguage, func(chunk []byte) bool {
out <- chunk
return true
})
}
// TTS synthesizes req.Text and writes a WAV to req.Dst.
//
// The whole body runs inside withEngine, so the family check and the C calls
// that trust the handle happen under a single acquisition of engineMu. See the
// handoff notes at the bottom of nemospeech.go: Free runs without the backend
// lock, so anything that checks the family and then releases the lock before
// calling C can have the handle destroyed underneath it.
func (n *NemoSpeech) TTS(req *pb.TTSRequest) error {
return n.withEngine(familyTTS, func() error {
return synthesizeWAV(&cSynthesizer{handle: n.synth}, req, n.opts.languageCode)
})
}
// TTSStream synthesizes req.Text and emits the audio on results as it is
// produced.
//
// results is closed on EVERY path, including the family rejection and a
// validation failure, and the close is deferred outside withEngine so that a
// rejected family still closes it. pkg/grpc/server.go drains this channel from a
// goroutine and then blocks on that goroutine finishing, so a channel left open
// does not fail the request, it hangs the RPC and, because the backend lock is
// still held, every request behind it.
//
// Holding engineMu for the whole stream is deliberate and is the consequence
// documented on the locking protocol: an unload waits for the stream to end
// rather than destroying the synthesizer underneath it. There is no unbounded
// wait here, because unlike the ASR streams this one is driven by the runtime
// and ends when the text does, not when a client decides to stop sending.
func (n *NemoSpeech) TTSStream(req *pb.TTSRequest, results chan []byte) error {
defer close(results)
return n.withEngine(familyTTS, func() error {
return streamWAV(&cSynthesizer{handle: n.synth}, req, n.opts.languageCode, results)
})
}

View File

@@ -0,0 +1,727 @@
package main
import (
"encoding/binary"
"errors"
"go/ast"
"go/parser"
"go/token"
"os"
"path/filepath"
"sync"
"unsafe"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"google.golang.org/grpc/codes"
"google.golang.org/grpc/status"
laudio "github.com/mudler/LocalAI/pkg/audio"
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
)
// puregoCallbackTableSize is the hard ceiling purego compiles callbacks into:
// maxCB in purego/syscall_sysv.go, which panics rather than growing once it is
// full and never releases an entry. Read off the module source for v0.10.0
// rather than assumed, because the whole point of the specs below is that
// exceeding it kills the process.
const puregoCallbackTableSize = 2000
// fakeSynthesizer scripts what the TTS C API emits for one synthesis.
//
// There is no MagpieTTS GGUF in the tree, so this is the only way the logic on
// top of the ABI (validation, WAV framing, chunk ordering, channel closure)
// gets tested at all. It fakes the C contract, not the model: chunks is
// whatever nemo_speech_tts_synthesize_text would have handed the callback.
type fakeSynthesizer struct {
rate int32
chunks [][]byte
err error
calls int
gotReq *pb.TTSRequest
gotLang string
cancelled bool
}
func (f *fakeSynthesizer) sampleRate() int32 { return f.rate }
func (f *fakeSynthesizer) synthesize(req *pb.TTSRequest, defaultLanguage string, sink ttsSink) error {
f.calls++
f.gotReq = req
f.gotLang = defaultLanguage
for _, c := range f.chunks {
if !sink(c) {
f.cancelled = true
break
}
}
return f.err
}
var _ = Describe("resolveSpeaker", func() {
It("passes a numeric voice through as a speaker index", func() {
idx, name := resolveSpeaker("3")
Expect(idx).To(Equal(int32(3)))
Expect(name).To(BeEmpty())
})
It("passes a named voice through as a name with no index", func() {
// voice_name is ignored whenever speaker >= 0 (tts.h, and
// synthesizer.cpp only calls resolve_speaker for a negative speaker), so
// a named voice must leave the index negative or the name is dropped.
idx, name := resolveSpeaker("Aria")
Expect(idx).To(Equal(int32(-1)))
Expect(name).To(Equal("Aria"))
})
It("leaves both unset for an empty voice so the synthesizer default wins", func() {
idx, name := resolveSpeaker("")
Expect(idx).To(Equal(int32(-1)))
Expect(name).To(BeEmpty())
})
// A negative number is the C API's sentinel for "use the default", not a
// speaker. Passing it through as an index would turn a request naming an
// invalid voice into one that quietly synthesizes in the default voice.
// Handed on as a name instead, resolve_speaker rejects it.
It("does not let a negative number become a speaker index", func() {
idx, name := resolveSpeaker("-1")
Expect(idx).To(Equal(int32(-1)))
Expect(name).To(Equal("-1"))
})
It("treats a non-numeric voice that merely starts with digits as a name", func() {
idx, name := resolveSpeaker("3-alpha")
Expect(idx).To(Equal(int32(-1)))
Expect(name).To(Equal("3-alpha"))
})
It("keeps speaker 0 addressable", func() {
// 0 is a real speaker index, and the only sentinel here is < 0.
idx, name := resolveSpeaker("0")
Expect(idx).To(BeZero())
Expect(name).To(BeEmpty())
})
})
var _ = Describe("applySynthesisParams", func() {
// The struct the runtime hands out: speaker/seed/steps/top_k all -1, the
// overrides off. Written literally rather than taken from
// TTSSynthesisOptionsDefault so the specs run without the shared libraries.
defaults := func() cTTSSynthesisOptions {
return cTTSSynthesisOptions{
Size: unsafe.Sizeof(cTTSSynthesisOptions{}),
Speaker: -1,
Seed: -1,
Steps: -1,
TopK: -1,
}
}
It("leaves every default alone for an absent params map", func() {
o := defaults()
applySynthesisParams(&o, nil)
Expect(o).To(Equal(defaults()))
})
It("leaves every default alone for an empty params map", func() {
o := defaults()
applySynthesisParams(&o, map[string]string{})
Expect(o).To(Equal(defaults()))
})
It("maps the five knobs the C options struct actually has", func() {
o := defaults()
applySynthesisParams(&o, map[string]string{
"seed": "42",
"steps": "12",
"top_k": "80",
"temperature": "0.7",
"cfg_scale": "1.5",
})
Expect(o.Seed).To(Equal(int32(42)))
Expect(o.Steps).To(Equal(int32(12)))
Expect(o.TopK).To(Equal(int32(80)))
Expect(o.Temperature).To(BeNumerically("~", 0.7, 1e-6))
Expect(o.CFGScale).To(BeNumerically("~", 1.5, 1e-6))
})
// magpietts/runtime.cpp reads options.temperature only when
// override_temperature is true and otherwise falls back to the
// synthesizer's config, so a temperature written without its flag is
// silently discarded and the request looks like it was honoured.
It("sets the override flag with the temperature", func() {
o := defaults()
applySynthesisParams(&o, map[string]string{"temperature": "0.4"})
Expect(o.OverrideTemperature).To(BeTrue())
Expect(o.OverrideCFGScale).To(BeFalse(), "cfg_scale was not asked for")
})
It("sets the override flag with the cfg scale", func() {
o := defaults()
applySynthesisParams(&o, map[string]string{"cfg_scale": "2"})
Expect(o.OverrideCFGScale).To(BeTrue())
Expect(o.OverrideTemperature).To(BeFalse(), "temperature was not asked for")
})
It("keeps the defaults when a value cannot be parsed", func() {
o := defaults()
applySynthesisParams(&o, map[string]string{
"seed": "many",
"steps": "",
"top_k": "8.5",
"temperature": "warm",
"cfg_scale": "-",
})
Expect(o).To(Equal(defaults()))
})
// The runtime takes a request's seed only when it is >= 0 and its steps and
// top_k only when they are > 0. Writing a parsed 0 or a negative would not
// be ignored downstream, it would erase the sentinel that means "use the
// synthesizer's value".
It("refuses values that would erase a sentinel", func() {
o := defaults()
applySynthesisParams(&o, map[string]string{
"seed": "-5",
"steps": "0",
"top_k": "0",
})
Expect(o.Seed).To(Equal(int32(-1)))
Expect(o.Steps).To(Equal(int32(-1)))
Expect(o.TopK).To(Equal(int32(-1)))
})
It("keeps seed 0, which is a real seed", func() {
o := defaults()
applySynthesisParams(&o, map[string]string{"seed": "0"})
Expect(o.Seed).To(BeZero())
})
// TTSRequest carries fields with no equivalent in
// nemo_speech_tts_synthesis_options. They must not be smuggled in through a
// param name that happens to match.
It("ignores params the C options struct has no field for", func() {
o := defaults()
applySynthesisParams(&o, map[string]string{
"top_p": "0.9",
"repetition_penalty": "1.1",
"speed": "1.2",
"instructions": "cheerful",
})
Expect(o).To(Equal(defaults()))
})
})
// The PCM callback is the one resource in this backend with a hard, silent,
// process-wide ceiling: purego compiles each into a fixed table of 2000 entries
// and never releases one, so a callback built per request takes the whole
// backend process down with a panic after 2000 syntheses. Nothing about a
// handful of manual calls shows that.
var _ = Describe("ttsPCMCallback", func() {
It("compiles a usable callback", func() {
Expect(ttsPCMCallback()).ToNot(BeZero())
})
It("compiles exactly one callback however many times it is asked", func() {
first := ttsPCMCallback()
// One more than the table holds: a callback compiled per call panics
// with "purego: the maximum number of callbacks has been reached"
// before this loop ends, which is precisely the production failure.
for i := 0; i <= puregoCallbackTableSize; i++ {
Expect(ttsPCMCallback()).To(Equal(first),
"call %d returned a different callback, so a new one was compiled", i)
}
})
// A source-level assertion, deliberately, because the failure it guards
// against is invisible from inside the process: the way a per-request
// callback gets reintroduced is by someone calling purego.NewCallback at the
// synthesis site instead of going through ttsPCMCallback, and no in-process
// spec can reach that call without a MagpieTTS GGUF to synthesize with.
// Funnelling every compile through one accessor is what the whole design
// rests on, so the single call site is the invariant worth pinning.
It("compiles callbacks from exactly one place in the TTS path", func() {
fset := token.NewFileSet()
file, err := parser.ParseFile(fset, "tts.go", nil, 0)
Expect(err).ToNot(HaveOccurred())
// Counted over the syntax tree rather than by grepping the text: the
// doc comment on ttsPCMCallback names purego.NewCallback too, and a
// spec that cannot tell an explanation from a call would be pinning the
// prose.
var sites []string
ast.Inspect(file, func(n ast.Node) bool {
call, ok := n.(*ast.CallExpr)
if !ok {
return true
}
sel, ok := call.Fun.(*ast.SelectorExpr)
if !ok || sel.Sel.Name != "NewCallback" {
return true
}
if pkg, ok := sel.X.(*ast.Ident); ok && pkg.Name == "purego" {
sites = append(sites, fset.Position(call.Pos()).String())
}
return true
})
Expect(sites).To(HaveLen(1),
"every callback must be compiled through ttsPCMCallback, which memoises it")
})
})
var _ = Describe("the PCM sink table", func() {
It("routes a chunk to the sink registered for that id", func() {
var got []byte
id, release := ttsSinks.register(func(pcm []byte) bool {
got = pcm
return true
})
defer release()
src := []byte{1, 2, 3, 4}
Expect(ttsDeliverPCM(unsafe.Pointer(&src[0]), uint64(len(src)), id)).To(BeTrue())
Expect(got).To(Equal([]byte{1, 2, 3, 4}))
})
// The pointer addresses a std::string the runtime reuses for the next
// chunk, so a slice over it would be rewritten under the consumer.
It("copies the chunk out of the runtime's buffer", func() {
var got []byte
id, release := ttsSinks.register(func(pcm []byte) bool {
got = pcm
return true
})
defer release()
src := []byte{9, 8, 7}
Expect(ttsDeliverPCM(unsafe.Pointer(&src[0]), uint64(len(src)), id)).To(BeTrue())
src[0], src[1], src[2] = 0, 0, 0
Expect(got).To(Equal([]byte{9, 8, 7}))
})
It("gives each registration its own id", func() {
idA, releaseA := ttsSinks.register(func([]byte) bool { return true })
defer releaseA()
idB, releaseB := ttsSinks.register(func([]byte) bool { return true })
defer releaseB()
Expect(idA).ToNot(Equal(idB))
Expect(idA).ToNot(BeZero(), "id 0 is what a zeroed user_data would carry")
Expect(idB).ToNot(BeZero())
})
// Two models synthesizing at once share one callback, and engineMu is
// per-model, so nothing serialises them against each other.
It("keeps concurrent sinks apart", func() {
var mu sync.Mutex
got := map[uintptr][]byte{}
var wg sync.WaitGroup
for i := range 16 {
wg.Add(1)
go func() {
defer GinkgoRecover()
defer wg.Done()
src := []byte{byte(i)}
var mine []byte
id, release := ttsSinks.register(func(pcm []byte) bool {
mine = pcm
return true
})
defer release()
Expect(ttsDeliverPCM(unsafe.Pointer(&src[0]), 1, id)).To(BeTrue())
mu.Lock()
defer mu.Unlock()
got[id] = mine
}()
}
wg.Wait()
Expect(got).To(HaveLen(16))
for id, pcm := range got {
Expect(pcm).To(HaveLen(1), "sink %d received the wrong chunk", id)
}
})
// A released id means the request has returned. Answering true would leave
// the runtime synthesizing into nothing while the RPC that owns the lock
// waits for it.
It("cancels the synthesis when the sink is gone", func() {
id, release := ttsSinks.register(func([]byte) bool { return true })
release()
src := []byte{1}
Expect(ttsDeliverPCM(unsafe.Pointer(&src[0]), 1, id)).To(BeFalse())
})
It("cancels for a user_data that was never registered", func() {
src := []byte{1}
Expect(ttsDeliverPCM(unsafe.Pointer(&src[0]), 1, 0)).To(BeFalse())
})
It("accepts an empty chunk without touching the pointer", func() {
id, release := ttsSinks.register(func([]byte) bool {
Fail("an empty chunk must not reach the sink")
return true
})
defer release()
Expect(ttsDeliverPCM(nil, 0, id)).To(BeTrue())
})
It("passes the sink's cancellation back to the runtime", func() {
id, release := ttsSinks.register(func([]byte) bool { return false })
defer release()
src := []byte{1}
Expect(ttsDeliverPCM(unsafe.Pointer(&src[0]), 1, id)).To(BeFalse())
})
})
var _ = Describe("ttsModelConfig", func() {
// Three adjacent same-typed path fields: swapping two changes neither the
// struct size nor any offset, so abi_test.go's layout assertions cannot see
// it and the runtime would load the codec as the acoustic model.
It("assigns each path to its own field", func() {
cfg := ttsModelConfig(1, 2, 3, 4)
Expect(cfg.MagpieModel).To(Equal(uintptr(1)))
Expect(cfg.CodecModel).To(Equal(uintptr(2)))
Expect(cfg.TokenizerModelDir).To(Equal(uintptr(3)))
Expect(cfg.TextNormalizerModelDir).To(Equal(uintptr(4)))
})
// A config sent with the wrong size has every field past it ignored by
// HAS_FIELD, and the model loads with defaults instead of failing.
It("declares the size the runtime validates against", func() {
Expect(ttsModelConfig(1, 2, 3, 4).Size).To(Equal(unsafe.Sizeof(cTTSModelConfig{})))
})
It("leaves an unset text normalizer null", func() {
Expect(ttsModelConfig(1, 2, 3, 0).TextNormalizerModelDir).To(BeZero())
})
})
var _ = Describe("ttsRuntimeBackend", func() {
// -1 is this backend's documented "CPU" everywhere (asr.h: "-1 = CPU") and
// it is also the default, so it has to pin the preference rather than leave
// the runtime free to pick CUDA.
It("pins CPU for a negative gpu option", func() {
Expect(ttsRuntimeBackend(-1)).To(Equal(ttsBackendCPU))
})
// nemo_speech_tts_runtime_config has no device index at all, so a request
// for a particular device cannot be honoured and AUTO is the honest answer.
It("leaves the choice to the runtime when a device was named", func() {
Expect(ttsRuntimeBackend(0)).To(Equal(ttsBackendAuto))
Expect(ttsRuntimeBackend(3)).To(Equal(ttsBackendAuto))
})
})
var _ = Describe("WAV framing", func() {
// 16-bit mono little-endian, the format the runtime's callback delivers.
pcm := []byte{0x01, 0x00, 0xff, 0x7f, 0x00, 0x80}
It("writes a header the audio helpers can read back", func() {
out, err := wavFile(pcm, 22050)
Expect(err).ToNot(HaveOccurred())
body, rate := laudio.ParseWAV(out)
Expect(rate).To(Equal(22050))
Expect(body).To(Equal(pcm))
})
It("describes the payload it actually carries", func() {
out, err := wavFile(pcm, 22050)
Expect(err).ToNot(HaveOccurred())
Expect(out).To(HaveLen(laudio.WAVHeaderSize + len(pcm)))
Expect(string(out[0:4])).To(Equal("RIFF"))
Expect(string(out[8:12])).To(Equal("WAVE"))
Expect(binary.LittleEndian.Uint32(out[4:8])).To(Equal(uint32(36 + len(pcm))))
Expect(binary.LittleEndian.Uint32(out[40:44])).To(Equal(uint32(len(pcm))))
Expect(binary.LittleEndian.Uint16(out[22:24])).To(Equal(uint16(1)), "mono")
Expect(binary.LittleEndian.Uint16(out[34:36])).To(Equal(uint16(16)), "16-bit")
Expect(binary.LittleEndian.Uint32(out[24:28])).To(Equal(uint32(22050)))
// byte rate = sample rate * block align, and a wrong one plays back at
// the wrong speed in players that trust it.
Expect(binary.LittleEndian.Uint32(out[28:32])).To(Equal(uint32(22050 * 2)))
})
It("carries whatever rate the synthesizer reported", func() {
out, err := wavFile(pcm, 44100)
Expect(err).ToNot(HaveOccurred())
_, rate := laudio.ParseWAV(out)
Expect(rate).To(Equal(44100))
})
Describe("the streaming header", func() {
It("is a complete header on its own", func() {
h := streamingWAVHeader(22050)
Expect(h).To(HaveLen(laudio.WAVHeaderSize))
Expect(string(h[0:4])).To(Equal("RIFF"))
Expect(string(h[8:12])).To(Equal("WAVE"))
Expect(binary.LittleEndian.Uint32(h[24:28])).To(Equal(uint32(22050)))
})
// NewWAVHeaderWithRate derives ChunkSize from the payload length, so
// leaving it alone would write 36 + 0xFFFFFFFF, which wraps to 35: a
// RIFF size smaller than the header itself.
It("leaves both sizes unknown rather than wrapping", func() {
h := streamingWAVHeader(22050)
Expect(binary.LittleEndian.Uint32(h[4:8])).To(Equal(uint32(0xFFFFFFFF)))
Expect(binary.LittleEndian.Uint32(h[40:44])).To(Equal(uint32(0xFFFFFFFF)))
})
})
})
var _ = Describe("synthesizeWAV", func() {
var dst string
BeforeEach(func() {
dst = filepath.Join(GinkgoT().TempDir(), "out.wav")
})
It("writes one WAV holding every chunk in order", func() {
s := &fakeSynthesizer{rate: 22050, chunks: [][]byte{{1, 0}, {2, 0}, {3, 0}}}
Expect(synthesizeWAV(s, &pb.TTSRequest{Text: "hello", Dst: dst}, "")).To(Succeed())
out, err := os.ReadFile(dst)
Expect(err).ToNot(HaveOccurred())
body, rate := laudio.ParseWAV(out)
Expect(rate).To(Equal(22050))
Expect(body).To(Equal([]byte{1, 0, 2, 0, 3, 0}))
})
It("hands the request and the model default language to the runtime", func() {
s := &fakeSynthesizer{rate: 22050, chunks: [][]byte{{1, 0}}}
req := &pb.TTSRequest{Text: "hello", Dst: dst, Voice: "Aria"}
Expect(synthesizeWAV(s, req, "it-IT")).To(Succeed())
Expect(s.gotReq).To(Equal(req))
Expect(s.gotLang).To(Equal("it-IT"))
})
It("rejects an empty text before it reaches the runtime", func() {
s := &fakeSynthesizer{rate: 22050}
err := synthesizeWAV(s, &pb.TTSRequest{Dst: dst}, "")
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(s.calls).To(BeZero())
})
// instructions has no equivalent in nemo_speech_tts_synthesis_options, which
// conditions on a speaker rather than a prose style. Dropping it must not
// fail the request: the caller still wants the audio it can have.
It("synthesizes anyway for a request carrying instructions it cannot honour", func() {
s := &fakeSynthesizer{rate: 22050, chunks: [][]byte{{1, 0}}}
instructions := "speak cheerfully"
Expect(synthesizeWAV(s, &pb.TTSRequest{
Text: "hello",
Dst: dst,
Instructions: &instructions,
}, "")).To(Succeed())
Expect(dst).To(BeAnExistingFile())
})
It("rejects a request with no destination", func() {
s := &fakeSynthesizer{rate: 22050}
err := synthesizeWAV(s, &pb.TTSRequest{Text: "hello"}, "")
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(s.calls).To(BeZero())
})
// A zero rate is what a null handle reports. The file it would produce is
// undecodable, and the synthesis that produced it would be wasted.
It("refuses to write a file at an unusable sample rate", func() {
s := &fakeSynthesizer{rate: 0, chunks: [][]byte{{1, 0}}}
err := synthesizeWAV(s, &pb.TTSRequest{Text: "hello", Dst: dst}, "")
Expect(status.Code(err)).To(Equal(codes.Internal))
Expect(s.calls).To(BeZero())
Expect(dst).ToNot(BeAnExistingFile())
})
It("propagates a synthesis failure and writes nothing", func() {
boom := errors.New("boom")
s := &fakeSynthesizer{rate: 22050, chunks: [][]byte{{1, 0}}, err: boom}
Expect(synthesizeWAV(s, &pb.TTSRequest{Text: "hello", Dst: dst}, "")).To(MatchError(boom))
Expect(dst).ToNot(BeAnExistingFile())
})
// An empty WAV is a valid file, so this would otherwise reach the user as
// silence with no error anywhere.
It("fails rather than write a silent file when nothing was produced", func() {
s := &fakeSynthesizer{rate: 22050}
err := synthesizeWAV(s, &pb.TTSRequest{Text: "hello", Dst: dst}, "")
Expect(status.Code(err)).To(Equal(codes.Internal))
Expect(dst).ToNot(BeAnExistingFile())
})
It("reports a destination it cannot write", func() {
s := &fakeSynthesizer{rate: 22050, chunks: [][]byte{{1, 0}}}
bad := filepath.Join(GinkgoT().TempDir(), "no-such-dir", "out.wav")
err := synthesizeWAV(s, &pb.TTSRequest{Text: "hello", Dst: bad}, "")
Expect(status.Code(err)).To(Equal(codes.Internal))
Expect(err.Error()).To(ContainSubstring("out.wav"))
})
})
var _ = Describe("streamWAV", func() {
// drain collects everything streamWAV emits. The channel is buffered
// because streamWAV sends inline, so an unbuffered one would deadlock the
// spec rather than fail it.
drain := func(s synthesizer, req *pb.TTSRequest) ([][]byte, error) {
out := make(chan []byte, 16)
err := streamWAV(s, req, "", out)
close(out)
var got [][]byte
for c := range out {
got = append(got, c)
}
return got, err
}
It("emits the header first, then each chunk as it arrives", func() {
s := &fakeSynthesizer{rate: 22050, chunks: [][]byte{{1, 0}, {2, 0}}}
got, err := drain(s, &pb.TTSRequest{Text: "hello"})
Expect(err).ToNot(HaveOccurred())
Expect(got).To(HaveLen(3))
Expect(got[0]).To(Equal(streamingWAVHeader(22050)))
Expect(got[1]).To(Equal([]byte{1, 0}))
Expect(got[2]).To(Equal([]byte{2, 0}))
})
// pkg/grpc/server.go only ever sets Reply.Audio, so core/backend's own
// header branch (keyed on Reply.Message) never runs and a backend that
// emitted bare PCM would stream something no client could decode.
It("owns the header rather than leaving it to the caller", func() {
s := &fakeSynthesizer{rate: 44100, chunks: [][]byte{{1, 0}}}
got, err := drain(s, &pb.TTSRequest{Text: "hello"})
Expect(err).ToNot(HaveOccurred())
Expect(string(got[0][0:4])).To(Equal("RIFF"))
Expect(binary.LittleEndian.Uint32(got[0][24:28])).To(Equal(uint32(44100)))
})
It("rejects an empty text before emitting anything", func() {
s := &fakeSynthesizer{rate: 22050, chunks: [][]byte{{1, 0}}}
got, err := drain(s, &pb.TTSRequest{})
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(got).To(BeEmpty())
Expect(s.calls).To(BeZero())
})
It("emits no header at an unusable sample rate", func() {
s := &fakeSynthesizer{rate: 0, chunks: [][]byte{{1, 0}}}
got, err := drain(s, &pb.TTSRequest{Text: "hello"})
Expect(status.Code(err)).To(Equal(codes.Internal))
Expect(got).To(BeEmpty())
})
It("propagates a synthesis failure after the chunks it did emit", func() {
boom := errors.New("boom")
s := &fakeSynthesizer{rate: 22050, chunks: [][]byte{{1, 0}}, err: boom}
got, err := drain(s, &pb.TTSRequest{Text: "hello"})
Expect(err).To(MatchError(boom))
Expect(got).To(HaveLen(2))
})
// streamWAV must not close the channel: TTSStream owns it, and closing in
// one of two places depending on how far the request got is how a stream
// ends up double-closed.
It("leaves the channel open for its caller to close", func() {
s := &fakeSynthesizer{rate: 22050, chunks: [][]byte{{1, 0}}}
out := make(chan []byte, 4)
Expect(streamWAV(s, &pb.TTSRequest{Text: "hello"}, "", out)).To(Succeed())
Expect(func() { close(out) }).ToNot(Panic())
})
})
var _ = Describe("the TTS RPCs", func() {
It("refuses TTS on a model loaded as another family", func() {
n := &NemoSpeech{fam: familyASR}
err := n.TTS(&pb.TTSRequest{Text: "hello", Dst: "/tmp/out.wav"})
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
})
It("refuses TTS on an unloaded model", func() {
n := &NemoSpeech{}
Expect(status.Code(n.TTS(&pb.TTSRequest{Text: "hello", Dst: "/tmp/out.wav"}))).
To(Equal(codes.Unimplemented))
})
It("releases the engine lock after a refusal", func() {
n := &NemoSpeech{fam: familyASR}
Expect(n.TTS(&pb.TTSRequest{Text: "hello", Dst: "/tmp/out.wav"})).ToNot(Succeed())
Expect(n.engineMu.TryLock()).To(BeTrue())
n.engineMu.Unlock()
})
// pkg/grpc/server.go drains this channel from a goroutine and then blocks
// on that goroutine finishing, so a channel left open does not fail the
// request, it hangs the RPC with the backend lock still held. Every exit
// path has to close it.
Describe("TTSStream channel closure", func() {
// streamed runs TTSStream the way the server does and returns once the
// channel has been closed, so a spec that hangs is a real hang.
streamed := func(n *NemoSpeech, req *pb.TTSRequest) ([][]byte, error) {
ch := make(chan []byte, 16)
done := make(chan [][]byte, 1)
go func() {
defer GinkgoRecover()
var got [][]byte
for c := range ch {
got = append(got, c)
}
done <- got
}()
err := n.TTSStream(req, ch)
var got [][]byte
Eventually(done).Should(Receive(&got))
return got, err
}
It("closes the channel when the family does not match", func() {
n := &NemoSpeech{fam: familyASR}
got, err := streamed(n, &pb.TTSRequest{Text: "hello"})
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
Expect(got).To(BeEmpty())
})
It("closes the channel when the model was never loaded", func() {
n := &NemoSpeech{}
_, err := streamed(n, &pb.TTSRequest{Text: "hello"})
Expect(status.Code(err)).To(Equal(codes.Unimplemented))
})
// familyTTS with a zero handle: validation has to reject this before
// anything reaches the C entry points, which are nil function values
// until openLibraries has bound them.
It("closes the channel when the request is rejected", func() {
n := &NemoSpeech{fam: familyTTS}
got, err := streamed(n, &pb.TTSRequest{})
Expect(status.Code(err)).To(Equal(codes.InvalidArgument))
Expect(got).To(BeEmpty())
})
It("releases the engine lock afterwards", func() {
n := &NemoSpeech{fam: familyTTS}
_, err := streamed(n, &pb.TTSRequest{})
Expect(err).To(HaveOccurred())
Expect(n.engineMu.TryLock()).To(BeTrue())
n.engineMu.Unlock()
})
})
// The same guard on the offline path: a rejected request must not reach a
// nil C function through a zero handle.
It("rejects an invalid TTS request without touching the runtime", func() {
n := &NemoSpeech{fam: familyTTS}
Expect(status.Code(n.TTS(&pb.TTSRequest{Dst: "/tmp/out.wav"}))).To(Equal(codes.InvalidArgument))
Expect(status.Code(n.TTS(&pb.TTSRequest{Text: "hello"}))).To(Equal(codes.InvalidArgument))
})
})

View File

@@ -8,7 +8,7 @@ JOBS?=$(shell nproc --ignore=1)
# stablediffusion.cpp (ggml)
STABLEDIFFUSION_GGML_REPO?=https://github.com/leejet/stable-diffusion.cpp
STABLEDIFFUSION_GGML_VERSION?=db99efdd6d2a43c7937fd55b3359206c680a75b0
STABLEDIFFUSION_GGML_VERSION?=c6beeef35526c6dc94b74a7fb69f9d2e6a2a7a12
CMAKE_ARGS+=-DGGML_MAX_NAME=128

View File

@@ -11,7 +11,7 @@ JOBS?=$(shell nproc --ignore=1 2>/dev/null || sysctl -n hw.ncpu 2>/dev/null || e
# vllm.cpp version
VLLM_CPP_REPO?=https://github.com/mudler/vllm.cpp
VLLM_CPP_VERSION?=9d1fad3cde0acb95eb0bb0a1025f40a0eb614147
VLLM_CPP_VERSION?=0757cac231ecd571a83c4fd2f50805c9251fc225
# MLX GEMM provider (darwin/metal only; see the metal branch below for why).
# Consumed as the prebuilt pip wheel: building MLX from source needs `xcrun
@@ -96,6 +96,12 @@ endif
UNAME_S := $(shell uname -s)
ifeq ($(UNAME_S),Darwin)
LIB=libvllm.dylib
# Apple Clang diagnoses a pair of constant-folded array bounds in the Metal
# build as a GNU extension. Disable that diagnostic for both Objective-C and
# C++ because vllm.cpp appends target-local -Werror after these global flags.
CMAKE_ARGS+=-DCMAKE_CXX_FLAGS=-Wno-gnu-folding-constant
CMAKE_ARGS+=-DCMAKE_OBJC_FLAGS=-Wno-gnu-folding-constant
CMAKE_ARGS+=-DCMAKE_OBJCXX_FLAGS=-Wno-gnu-folding-constant
else
LIB=libvllm.so
endif
@@ -133,7 +139,26 @@ MLX_STAMP=
MLX_CMAKE_ARGS=
endif
# govllmcpp.go mirrors vllm.h by hand, and the only guard against the two
# drifting apart is the vllm_abi_version check inside registerLib - which fires
# at runtime, on the user's machine, taking down every model load (issue
# #11379). Compare the two here instead, so moving VLLM_CPP_VERSION past the
# mirrors turns the build red while the header is still around to diff.
abi-check: sources/vllm.cpp
@engine=$$(sed -n 's/^#define VLLM_ABI_VERSION \([0-9][0-9]*\).*/\1/p' sources/vllm.cpp/include/vllm.h); \
backend=$$(sed -n 's/^const abiVersion = \([0-9][0-9]*\).*/\1/p' govllmcpp.go); \
if [ -z "$$engine" ] || [ -z "$$backend" ]; then \
echo "vllm-cpp: cannot read the ABI version (engine='$$engine' backend='$$backend')" >&2; exit 1; \
fi; \
if [ "$$engine" != "$$backend" ]; then \
echo "vllm-cpp: ABI mismatch: vllm.cpp $(VLLM_CPP_VERSION) is v$$engine, govllmcpp.go mirrors v$$backend." >&2; \
echo " Update the struct mirrors and abiVersion in govllmcpp.go (and the offsets in vllmcpp_test.go) to v$$engine." >&2; \
exit 1; \
fi; \
echo "vllm-cpp: ABI v$$engine matches the pinned engine"
$(LIB): sources/vllm.cpp $(MLX_STAMP)
$(MAKE) abi-check
mkdir -p build && \
cd build && \
cmake ../sources/vllm.cpp $(CMAKE_ARGS) $(MLX_CMAKE_ARGS) && \
@@ -154,6 +179,8 @@ clean: purge
purge:
rm -rf build
.PHONY: abi-check
.NOTPARALLEL:
# The unit specs are pure Go (struct mirrors, option mapping, load

View File

@@ -6,7 +6,7 @@ safetensors + GGUF loading, CUDA / CPU / Metal / Vulkan) with no Python at
inference time.
The backend dlopens the engine's stable C ABI (`libvllm`, `include/vllm.h`,
ABI v2) through purego:
ABI v10) through purego:
- `Load` -> `vllm_engine_load`: accepts a `.gguf` file or a HF-style model
directory (`config.json` + safetensors). `context_size` maps to
@@ -29,6 +29,12 @@ ABI v2) through purego:
LocalAI's Go-side grammar-constrained tool calling; JSON-schema / regex /
choice constraints are also exposed by the ABI.
The struct mirrors in `govllmcpp.go` are hand-written against one ABI version,
and the engine refuses to load against any other. Moving `VLLM_CPP_VERSION` in
the Makefile therefore means updating `abiVersion` plus the mirrors (and their
offsets in `vllmcpp_test.go`) in the same change; `make abi-check` compares the
pinned header against the bindings and the library build runs it first.
Model config example:
```yaml

View File

@@ -109,6 +109,16 @@ func (v *VllmCpp) Load(opts *pb.ModelOptions) error {
v.opts = parseOptions(opts)
// A DFlash draft is a second checkpoint the engine opens by path, and the
// engine never downloads one. Resolve it against LocalAI's models directory
// now so a repo-id spelling works, and so a missing draft fails here with an
// actionable message rather than as an HF-cache miss inside the load.
resolvedSpec, err := resolveDraftModelPath(v.opts.speculativeConfig, opts.ModelPath)
if err != nil {
return err
}
v.opts.speculativeConfig = resolvedSpec
mp := defaultModelParams()
if v.opts.blockSize > 0 {
mp.BlockSize = v.opts.blockSize
@@ -116,34 +126,62 @@ func (v *VllmCpp) Load(opts *pb.ModelOptions) error {
if v.opts.numBlocks > 0 {
mp.NumBlocks = v.opts.numBlocks
}
// Sequence-length precedence, narrowest source last: context_size is the
// generic LocalAI knob every backend honours, max_model_len is the
// vLLM-specific one, and engine_args.max_model_len is the explicit
// vllm-cpp override.
if opts.ContextSize > 0 {
mp.MaxModelLen = opts.ContextSize
}
if opts.MaxModelLen > 0 {
mp.MaxModelLen = opts.MaxModelLen
}
if v.opts.maxModelLen > 0 {
mp.MaxModelLen = v.opts.maxModelLen
}
if v.opts.maxNumSeqs > 0 {
mp.MaxNumSeqs = v.opts.maxNumSeqs
}
if v.opts.maxNumBatchedTokens > 0 {
mp.MaxNumBatchedTokens = v.opts.maxNumBatchedTokens
}
mp.EnablePrefixCaching = v.opts.enablePrefixCaching
mp.EnableJumpForward = v.opts.enableJumpForward
// Every string below is borrowed by C for the duration of the load call
// only (the library copies what it keeps), so the backing slices just have
// to outlive vllmEngineLoad - hence the single KeepAlive after it.
modelC := cString(model)
mp.ModelPath = uintptr(unsafe.Pointer(&modelC[0])) // #nosec G103 -- borrowed by C for the load call only
var toolParserC, reasoningParserC []byte
if v.opts.toolParser != "" {
toolParserC = cString(v.opts.toolParser)
mp.ToolParser = uintptr(unsafe.Pointer(&toolParserC[0])) // #nosec G103 -- borrowed by C for the load call only
}
if v.opts.reasoningParser != "" {
reasoningParserC = cString(v.opts.reasoningParser)
mp.ReasoningParser = uintptr(unsafe.Pointer(&reasoningParserC[0])) // #nosec G103 -- borrowed by C for the load call only
keep := [][]byte{modelC}
setStr := func(dst *uintptr, s string) {
if s == "" {
return
}
b := cString(s)
keep = append(keep, b)
*dst = uintptr(unsafe.Pointer(&b[0])) // #nosec G103 -- borrowed by C for the load call only
}
setStr(&mp.ToolParser, v.opts.toolParser)
setStr(&mp.ReasoningParser, v.opts.reasoningParser)
setStr(&mp.SpeculativeConfig, v.opts.speculativeConfig)
setStr(&mp.KVTransferConfig, v.opts.kvTransferConfig)
setStr(&mp.SchedulingPolicy, v.opts.schedulingPolicy)
setStr(&mp.TokenizerConfigPath, v.opts.tokenizerConfigPath)
xlog.Info("[vllm-cpp] Load", "model", model, "engine", vllmVersion(),
"blockSize", mp.BlockSize, "numBlocks", mp.NumBlocks,
"maxModelLen", mp.MaxModelLen, "maxNumSeqs", mp.MaxNumSeqs)
"maxModelLen", mp.MaxModelLen, "maxNumSeqs", mp.MaxNumSeqs,
"maxNumBatchedTokens", mp.MaxNumBatchedTokens,
"prefixCaching", triStateName(mp.EnablePrefixCaching),
"jumpForward", triStateName(mp.EnableJumpForward),
"schedulingPolicy", v.opts.schedulingPolicy,
"speculativeConfig", v.opts.speculativeConfig,
"kvTransferConfig", v.opts.kvTransferConfig)
var engine uintptr
rc := vllmEngineLoad(unsafe.Pointer(&mp), unsafe.Pointer(&engine)) // #nosec G103 -- POD out-params
runtime.KeepAlive(modelC)
runtime.KeepAlive(toolParserC)
runtime.KeepAlive(reasoningParserC)
runtime.KeepAlive(keep)
if rc != vllmOK {
return fmt.Errorf("vllm-cpp: engine load failed: %s", vllmLastError())
}

View File

@@ -1,6 +1,6 @@
package main
// purego bindings for the vllm.cpp stable C ABI (include/vllm.h, ABI v2).
// purego bindings for the vllm.cpp stable C ABI (include/vllm.h, ABI v10).
//
// The structs below are hand-mirrored PODs of the C declarations, with
// explicit padding so the Go layout matches the C layout on linux/darwin
@@ -17,29 +17,65 @@ import (
"github.com/ebitengine/purego"
)
// abiVersion is the VLLM_ABI_VERSION this file mirrors (vllm.h).
const abiVersion = 5
// abiVersion is the VLLM_ABI_VERSION this file mirrors (vllm.h). It must track
// the header of the VLLM_CPP_VERSION pinned in the Makefile: the build checks
// the two against each other, because a mismatch is only caught at runtime by
// registerLib, where it takes the backend down on every load (issue #11379).
const abiVersion = 10
// The ABI's tri-state toggles (enable_prefix_caching ABI v7,
// enable_jump_forward ABI v10) share one encoding: 0 is NOT "off", it is
// "defer" - to the model capability for prefix caching, to the environment for
// jump forward. Only 2 is an explicit off.
const (
triStateDefer int32 = 0
triStateOn int32 = 1
triStateOff int32 = 2
)
// triStateName renders a tri-state for the load log line, where "0" would
// otherwise read as "off" rather than "whatever the default resolves to".
func triStateName(state int32) string {
switch state {
case triStateOn:
return "on"
case triStateOff:
return "off"
default:
return "model-default"
}
}
// vllm_status (vllm.h).
const (
vllmOK = 0
)
// cModelParams mirrors vllm_model_params.
// cModelParams mirrors vllm_model_params. The int32 fields sit in pairs so the
// interior needs no padding on LP64, but the struct is 8-aligned (it holds
// pointers) and ends on a lone int32, so the trailing pad is explicit. Offsets
// and total size are asserted in vllmcpp_test.go.
type cModelParams struct {
ModelPath uintptr // const char*
TokenizerConfigPath uintptr // const char*
TokenizerConfigPath uintptr // const char*; NULL = <model_dir>/... (ABI v9)
BlockSize int32
NumBlocks int32
MaxModelLen int32
MaxNumSeqs int32
ToolParser uintptr // const char*; NULL = auto-detect (ABI v4)
ReasoningParser uintptr // const char*; NULL = auto-detect (ABI v5)
SpeculativeConfig uintptr // const char* JSON; NULL = no speculation (ABI v6)
EnablePrefixCaching int32 // tri-state 0/1/2 (ABI v7)
MaxNumBatchedTokens int32 // <= 0 = per-arch default (ABI v9)
SchedulingPolicy uintptr // const char*; NULL = "fcfs" (ABI v9)
KVTransferConfig uintptr // const char* JSON; NULL = no connector (ABI v9)
EnableJumpForward int32 // tri-state 0/1/2 (ABI v10)
_ [4]byte // trailing pad to the struct's 8-byte alignment
}
// cSamplingParams mirrors vllm_sampling_params (ABI v2, structured fields
// included). Padding matches the C compiler's: the uint64 seed is 8-aligned,
// and each pointer following an int32 is 8-aligned.
// cSamplingParams mirrors vllm_sampling_params (structured fields included).
// Padding matches the C compiler's: the uint64 seed is 8-aligned, and each
// pointer following an int32 is 8-aligned.
type cSamplingParams struct {
Temperature float32
TopP float32
@@ -65,6 +101,12 @@ type cSamplingParams struct {
StructuredGrammar uintptr // const char*
StructuredJSONObject int32
_ [4]byte
// ABI v8 tail. LocalAI installs no custom logits processor, but the fields
// MUST be mirrored: the C side reads them off the pointer we hand it, so a
// Go struct that stopped at StructuredJSONObject would have the engine read
// 16 bytes past our allocation and call whatever garbage sat there.
LogitsProcessor uintptr // vllm_logits_processor; NULL = none
LogitsProcessorUserData uintptr // void*
}
// cCompletion mirrors vllm_completion.

View File

@@ -1,30 +1,80 @@
package main
// Engine-sizing knobs carried through the model config's free-form
// `options:` list ("key:value" entries), mirroring how the other in-house
// backends pass engine-specific settings that have no proto field.
// Load-time engine configuration, from two config surfaces:
//
// - `engine_args:` (ModelOptions.EngineArgs, a JSON object) is the canonical
// one. Keys are spelled exactly as vLLM's own CLI flags, so a config written
// against vLLM works verbatim here - `speculative_config` and
// `kv_transfer_config` in particular take the same JSON documents vLLM's
// --speculative-config / --kv-transfer-config accept, and are handed to the
// engine unparsed.
// - `options:` (the free-form "key:value" list) is the older surface this
// backend shipped with. It is still honoured so existing configs keep
// working; engine_args wins on any key set in both.
//
// Anything unrecognised is ignored rather than fatal: the engine validates the
// documents it is given and reports a precise error at load, and a config that
// also carries knobs for a different backend must not fail the load here.
import (
"encoding/json"
"fmt"
"os"
"path"
"path/filepath"
"strconv"
"strings"
pb "github.com/mudler/LocalAI/pkg/grpc/proto"
"github.com/mudler/xlog"
)
type loadOptions struct {
blockSize int32 // KV block size (tokens/block); engine default 32.
numBlocks int32 // KV blocks to allocate; engine default 256.
maxNumSeqs int32 // max concurrent sequences; engine default 8.
// Max sequence length. Also settable through the model config's
// context_size / max_model_len; see Load for the precedence.
maxModelLen int32
// Per-step chunked-prefill token budget (ABI v9). 0 = the engine's
// bounded per-arch default.
maxNumBatchedTokens int32
// Automatic prefix caching tri-state (ABI v7): 0 = the model-capability
// default, 1 = force on, 2 = force off.
enablePrefixCaching int32
// Jump-forward decoding tri-state (ABI v10), SGLang's grammar-speed subset:
// 0 = defer to the environment (VT_ENABLE_JUMP_FORWARD, default off),
// 1 = force on, 2 = force off.
enableJumpForward int32
// Scheduler admission policy (ABI v9): "" = fcfs, else fcfs|priority|lpm.
schedulingPolicy string
// Engine-side parser selection (ABI v4/v5). Empty = the engine
// auto-detects from the chat template; "none" disables the reasoning
// split; unknown names fail the first chat call.
toolParser string
reasoningParser string
// Speculative decoding (ABI v6), as vLLM's --speculative-config JSON:
// {"method":"mtp"|"dflash"|"ngram", ...}. Empty = no speculation.
speculativeConfig string
// External KV connector / LMCache (ABI v9), as vLLM's --kv-transfer-config
// JSON. Empty = no connector.
kvTransferConfig string
// Override for the tokenizer_config.json the chat template is read from
// (ABI v9). Empty = <model_dir>/tokenizer_config.json.
tokenizerConfigPath string
}
func parseOptions(opts *pb.ModelOptions) loadOptions {
lo := loadOptions{}
for _, o := range opts.GetOptions() {
applyOptionsList(&lo, opts.GetOptions())
applyEngineArgs(&lo, opts.GetEngineArgs())
return lo
}
// applyOptionsList reads the legacy free-form "key:value" list. strings.Cut
// splits on the FIRST colon only, so a JSON object value survives intact.
func applyOptionsList(lo *loadOptions, options []string) {
for _, o := range options {
k, v, found := strings.Cut(o, ":")
if !found {
continue
@@ -36,13 +86,211 @@ func parseOptions(opts *pb.ModelOptions) loadOptions {
lo.numBlocks = parseInt32(v, lo.numBlocks)
case "max_num_seqs":
lo.maxNumSeqs = parseInt32(v, lo.maxNumSeqs)
case "tool_parser":
case "max_num_batched_tokens":
lo.maxNumBatchedTokens = parseInt32(v, lo.maxNumBatchedTokens)
case "max_model_len":
lo.maxModelLen = parseInt32(v, lo.maxModelLen)
case "scheduling_policy", "schedule_policy":
lo.schedulingPolicy = strings.TrimSpace(v)
case "tool_parser", "tool_call_parser":
lo.toolParser = strings.TrimSpace(v)
case "reasoning_parser":
lo.reasoningParser = strings.TrimSpace(v)
case "speculative_config":
lo.speculativeConfig = strings.TrimSpace(v)
case "kv_transfer_config":
lo.kvTransferConfig = strings.TrimSpace(v)
case "tokenizer_config", "tokenizer_config_path":
lo.tokenizerConfigPath = strings.TrimSpace(v)
case "enable_prefix_caching", "enable_radix_attention":
if b, err := strconv.ParseBool(strings.TrimSpace(v)); err == nil {
lo.enablePrefixCaching = boolTriState(b)
}
case "enable_jump_forward":
if b, err := strconv.ParseBool(strings.TrimSpace(v)); err == nil {
lo.enableJumpForward = boolTriState(b)
}
}
}
return lo
}
// applyEngineArgs overlays the `engine_args:` JSON object. A document that does
// not parse is logged and skipped: engine_args is shared with the other engines
// (the vLLM and SGLang backends read the same field), so a stray key must not
// take the model down.
func applyEngineArgs(lo *loadOptions, engineArgs string) {
if strings.TrimSpace(engineArgs) == "" {
return
}
var args map[string]any
if err := json.Unmarshal([]byte(engineArgs), &args); err != nil {
xlog.Warn("[vllm-cpp] ignoring unparseable engine_args", "error", err)
return
}
for k, v := range args {
switch k {
case "block_size":
lo.blockSize = jsonInt32(v, lo.blockSize)
case "num_blocks":
lo.numBlocks = jsonInt32(v, lo.numBlocks)
case "max_num_seqs":
lo.maxNumSeqs = jsonInt32(v, lo.maxNumSeqs)
case "max_num_batched_tokens":
lo.maxNumBatchedTokens = jsonInt32(v, lo.maxNumBatchedTokens)
case "max_model_len":
lo.maxModelLen = jsonInt32(v, lo.maxModelLen)
case "scheduling_policy", "schedule_policy":
lo.schedulingPolicy = jsonString(v, lo.schedulingPolicy)
case "tool_parser", "tool_call_parser":
lo.toolParser = jsonString(v, lo.toolParser)
case "reasoning_parser":
lo.reasoningParser = jsonString(v, lo.reasoningParser)
case "tokenizer_config", "tokenizer_config_path":
lo.tokenizerConfigPath = jsonString(v, lo.tokenizerConfigPath)
case "speculative_config":
lo.speculativeConfig = jsonDocument(v, lo.speculativeConfig, k)
case "kv_transfer_config":
lo.kvTransferConfig = jsonDocument(v, lo.kvTransferConfig, k)
case "enable_prefix_caching", "enable_radix_attention":
if b, ok := v.(bool); ok {
lo.enablePrefixCaching = boolTriState(b)
}
case "enable_jump_forward":
if b, ok := v.(bool); ok {
lo.enableJumpForward = boolTriState(b)
}
default:
xlog.Debug("[vllm-cpp] ignoring unknown engine_args key", "key", k)
}
}
}
// boolTriState maps a YAML/JSON boolean onto the ABI's tri-state encoding. An
// explicit `false` must reach the engine as force-OFF (2), NOT as the 0 that
// means "defer". The difference is real in both directions: prefix caching
// defaults ON for dense archs and OFF for hybrid ones, and jump forward defers
// to VT_ENABLE_JUMP_FORWARD.
func boolTriState(on bool) int32 {
if on {
return triStateOn
}
return triStateOff
}
// jsonDocument normalises an object-valued engine_args entry to a JSON string
// for the C ABI. YAML nesting arrives as a map (the natural spelling); a
// pre-encoded JSON string is accepted too, since a config round-tripped through
// a flat store may carry it that way.
func jsonDocument(v any, fallback string, key string) string {
switch t := v.(type) {
case string:
if strings.TrimSpace(t) == "" {
return fallback
}
return t
default:
buf, err := json.Marshal(t)
if err != nil {
xlog.Warn("[vllm-cpp] ignoring unencodable engine_args value", "key", key, "error", err)
return fallback
}
return string(buf)
}
}
func jsonString(v any, fallback string) string {
s, ok := v.(string)
if !ok {
return fallback
}
return strings.TrimSpace(s)
}
// jsonInt32 accepts the float64 a JSON number decodes to, plus the string
// spelling a YAML config may produce. Non-positive values keep the fallback:
// every knob this covers uses "<= 0 means the engine default".
func jsonInt32(v any, fallback int32) int32 {
switch t := v.(type) {
case float64:
if t <= 0 || t > 1<<31-1 {
return fallback
}
return int32(t)
case string:
return parseInt32(t, fallback)
default:
return fallback
}
}
// resolveDraftModelPath rewrites a DFlash draft reference into an absolute path
// the engine can actually open.
//
// The engine resolves `speculative_config.model` against a directory containing
// config.json, or against ~/.cache/huggingface/hub/models--<org>--<repo>/
// snapshots/* - and it NEVER downloads. LocalAI keeps models in its own
// directory, so a bare HF repo id (the spelling the vLLM docs teach) misses the
// HF cache and dies deep in the load with "draft checkpoint not found", which
// reads like a broken checkpoint rather than a missing download.
//
// So: try the reference as given, then the last path segment under the models
// dir (`z-lab/Qwen3.6-27B-DFlash` -> `<models>/Qwen3.6-27B-DFlash`, which is
// what LocalAI's own downloader produces), then the whole reference under the
// models dir. If none exist, fail HERE with a message naming both what was
// asked for and where we looked.
//
// mtp and ngram carry no separate draft checkpoint, so they pass through. A
// document that does not parse also passes through: the engine owns config
// validation and produces the better error.
func resolveDraftModelPath(speculativeConfig, modelsDir string) (string, error) {
if strings.TrimSpace(speculativeConfig) == "" {
return speculativeConfig, nil
}
var spec map[string]any
if err := json.Unmarshal([]byte(speculativeConfig), &spec); err != nil {
return speculativeConfig, nil
}
if method, _ := spec["method"].(string); !strings.EqualFold(method, "dflash") {
return speculativeConfig, nil
}
ref, _ := spec["model"].(string)
ref = strings.TrimSpace(ref)
if ref == "" {
return "", fmt.Errorf(
"vllm-cpp: speculative_config method %q requires a \"model\" key naming the draft checkpoint", "dflash")
}
candidates := []string{ref}
if modelsDir != "" {
if base := path.Base(filepath.ToSlash(ref)); base != "" && base != "." && base != "/" {
candidates = append(candidates, filepath.Join(modelsDir, base))
}
candidates = append(candidates, filepath.Join(modelsDir, filepath.FromSlash(ref)))
}
for _, c := range candidates {
if _, err := os.Stat(filepath.Join(c, "config.json")); err != nil {
continue
}
abs, err := filepath.Abs(c)
if err != nil {
abs = c
}
spec["model"] = abs
out, err := json.Marshal(spec)
if err != nil {
return "", fmt.Errorf("vllm-cpp: re-encoding speculative_config: %w", err)
}
xlog.Info("[vllm-cpp] resolved DFlash draft checkpoint", "reference", ref, "path", abs)
return string(out), nil
}
return "", fmt.Errorf(
"vllm-cpp: DFlash draft checkpoint %q not found (looked in: %s). "+
"The engine does not download drafts - install the draft model into LocalAI first, "+
"or set speculative_config.model to an absolute path to a directory containing config.json",
ref, strings.Join(candidates, ", "))
}
func parseInt32(s string, fallback int32) int32 {

View File

@@ -16,10 +16,17 @@ func TestVllmCpp(t *testing.T) {
RunSpecs(t, "vllm-cpp suite")
}
// The Go POD mirrors must match the C struct layout of vllm.h (ABI v2)
// The Go POD mirrors must match the C struct layout of vllm.h (ABI v10)
// byte-for-byte: these offsets are the C offsets on LP64 (linux/darwin
// amd64+arm64). A failure here means govllmcpp.go drifted from vllm.h.
var _ = Describe("C ABI struct mirrors", func() {
It("declares the ABI version the pinned engine reports", func() {
// VLLM_ABI_VERSION in the vllm.h of VLLM_CPP_VERSION (Makefile).
// Moving the pin past this without growing the mirrors below ships a
// backend that refuses every load at startup (issue #11379).
Expect(abiVersion).To(Equal(10))
})
It("cModelParams matches vllm_model_params", func() {
var p cModelParams
Expect(unsafe.Offsetof(p.ModelPath)).To(Equal(uintptr(0)))
@@ -30,10 +37,18 @@ var _ = Describe("C ABI struct mirrors", func() {
Expect(unsafe.Offsetof(p.MaxNumSeqs)).To(Equal(uintptr(28)))
Expect(unsafe.Offsetof(p.ToolParser)).To(Equal(uintptr(32)))
Expect(unsafe.Offsetof(p.ReasoningParser)).To(Equal(uintptr(40)))
Expect(unsafe.Sizeof(p)).To(Equal(uintptr(48)))
Expect(unsafe.Offsetof(p.SpeculativeConfig)).To(Equal(uintptr(48)))
Expect(unsafe.Offsetof(p.EnablePrefixCaching)).To(Equal(uintptr(56)))
Expect(unsafe.Offsetof(p.MaxNumBatchedTokens)).To(Equal(uintptr(60)))
Expect(unsafe.Offsetof(p.SchedulingPolicy)).To(Equal(uintptr(64)))
Expect(unsafe.Offsetof(p.KVTransferConfig)).To(Equal(uintptr(72)))
Expect(unsafe.Offsetof(p.EnableJumpForward)).To(Equal(uintptr(80)))
// 88, not 84: the struct is 8-aligned (it holds pointers), so the
// trailing int32 is padded out. Go pads identically.
Expect(unsafe.Sizeof(p)).To(Equal(uintptr(88)))
})
It("cSamplingParams matches vllm_sampling_params (ABI v2)", func() {
It("cSamplingParams matches vllm_sampling_params (ABI v8)", func() {
var p cSamplingParams
Expect(unsafe.Offsetof(p.Temperature)).To(Equal(uintptr(0)))
Expect(unsafe.Offsetof(p.TopP)).To(Equal(uintptr(4)))
@@ -55,7 +70,9 @@ var _ = Describe("C ABI struct mirrors", func() {
Expect(unsafe.Offsetof(p.NStructuredChoice)).To(Equal(uintptr(96)))
Expect(unsafe.Offsetof(p.StructuredGrammar)).To(Equal(uintptr(104)))
Expect(unsafe.Offsetof(p.StructuredJSONObject)).To(Equal(uintptr(112)))
Expect(unsafe.Sizeof(p)).To(Equal(uintptr(120)))
Expect(unsafe.Offsetof(p.LogitsProcessor)).To(Equal(uintptr(120)))
Expect(unsafe.Offsetof(p.LogitsProcessorUserData)).To(Equal(uintptr(128)))
Expect(unsafe.Sizeof(p)).To(Equal(uintptr(136)))
})
It("cCompletion matches vllm_completion", func() {
@@ -68,6 +85,23 @@ var _ = Describe("C ABI struct mirrors", func() {
})
})
// Pin/mirror skew is the failure mode this backend is most exposed to: the Go
// PODs above are hand-written against one VLLM_ABI_VERSION, and the Makefile
// pins the vllm.cpp commit that produces it. This spec catches drift without
// needing model weights - set VLLM_CPP_LIBRARY to a built libvllm and it binds
// every symbol and compares the library's reported ABI against the mirrors'.
var _ = Describe("real library ABI handshake", func() {
It("binds every symbol and reports the ABI the mirrors were written against", func() {
lib := os.Getenv("VLLM_CPP_LIBRARY")
if lib == "" {
Skip("VLLM_CPP_LIBRARY not set; skipping the real-library handshake")
}
Expect(registerLib(lib)).To(Succeed())
Expect(vllmABIVersion()).To(Equal(int32(abiVersion)))
Expect(vllmVersion()).NotTo(BeEmpty())
})
})
var _ = Describe("parseOptions", func() {
It("extracts the engine sizing knobs", func() {
lo := parseOptions(&pb.ModelOptions{Options: []string{
@@ -83,6 +117,129 @@ var _ = Describe("parseOptions", func() {
}})
Expect(lo).To(Equal(loadOptions{}))
})
It("carries a speculative_config JSON value through the legacy options list", func() {
// strings.Cut splits on the FIRST colon only, so a JSON object value
// survives the "key:value" spelling intact.
lo := parseOptions(&pb.ModelOptions{Options: []string{
`speculative_config:{"method":"mtp","num_speculative_tokens":1}`,
}})
Expect(lo.speculativeConfig).To(Equal(`{"method":"mtp","num_speculative_tokens":1}`))
})
})
var _ = Describe("engine_args", func() {
It("maps every load knob onto the C model params", func() {
lo := parseOptions(&pb.ModelOptions{EngineArgs: `{
"block_size": 64,
"num_blocks": 1024,
"max_model_len": 16384,
"max_num_seqs": 32,
"max_num_batched_tokens": 8192,
"enable_prefix_caching": true,
"scheduling_policy": "lpm",
"tool_parser": "qwen3",
"reasoning_parser": "deepseek_r1",
"tokenizer_config": "/models/tok/tokenizer_config.json"
}`})
Expect(lo.blockSize).To(Equal(int32(64)))
Expect(lo.numBlocks).To(Equal(int32(1024)))
Expect(lo.maxModelLen).To(Equal(int32(16384)))
Expect(lo.maxNumSeqs).To(Equal(int32(32)))
Expect(lo.maxNumBatchedTokens).To(Equal(int32(8192)))
Expect(lo.enablePrefixCaching).To(Equal(int32(1)))
Expect(lo.schedulingPolicy).To(Equal("lpm"))
Expect(lo.toolParser).To(Equal("qwen3"))
Expect(lo.reasoningParser).To(Equal("deepseek_r1"))
Expect(lo.tokenizerConfigPath).To(Equal("/models/tok/tokenizer_config.json"))
})
It("re-marshals a nested speculative_config object to JSON for the engine", func() {
lo := parseOptions(&pb.ModelOptions{EngineArgs: `{
"speculative_config": {"method": "mtp", "num_speculative_tokens": 1}
}`})
Expect(lo.speculativeConfig).To(MatchJSON(`{"method":"mtp","num_speculative_tokens":1}`))
})
It("re-marshals a nested kv_transfer_config object (LMCache) to JSON", func() {
lo := parseOptions(&pb.ModelOptions{EngineArgs: `{
"kv_transfer_config": {
"kv_connector": "LMCacheConnector",
"kv_role": "kv_both",
"kv_connector_extra_config": {"host": "127.0.0.1", "port": 65432}
}
}`})
Expect(lo.kvTransferConfig).To(MatchJSON(`{
"kv_connector":"LMCacheConnector",
"kv_role":"kv_both",
"kv_connector_extra_config":{"host":"127.0.0.1","port":65432}
}`))
})
It("accepts a pre-encoded JSON string for the object-valued knobs", func() {
// A config written by hand (or round-tripped through a flat store) may
// carry the object as a string; both spellings reach the engine the same.
lo := parseOptions(&pb.ModelOptions{EngineArgs: `{
"speculative_config": "{\"method\":\"ngram\",\"num_speculative_tokens\":4}"
}`})
Expect(lo.speculativeConfig).To(MatchJSON(`{"method":"ngram","num_speculative_tokens":4}`))
})
It("maps enable_prefix_caching false onto the force-OFF tri-state", func() {
// The C ABI tri-state is 0=model default, 1=on, 2=off, so an explicit
// `false` must NOT collapse to the 0 that means "let the model decide".
lo := parseOptions(&pb.ModelOptions{EngineArgs: `{"enable_prefix_caching": false}`})
Expect(lo.enablePrefixCaching).To(Equal(int32(2)))
})
It("leaves the prefix-caching tri-state at the model default when unset", func() {
lo := parseOptions(&pb.ModelOptions{EngineArgs: `{"max_num_seqs": 4}`})
Expect(lo.enablePrefixCaching).To(Equal(int32(0)))
})
It("accepts the radix-attention alias upstream documents for prefix caching", func() {
lo := parseOptions(&pb.ModelOptions{EngineArgs: `{"enable_radix_attention": true}`})
Expect(lo.enablePrefixCaching).To(Equal(int32(1)))
})
It("maps enable_jump_forward onto its own tri-state", func() {
// ABI v10. Same tri-state shape as prefix caching, and the same trap:
// an explicit false must be force-OFF (2), not the 0 that defers to the
// environment.
on := parseOptions(&pb.ModelOptions{EngineArgs: `{"enable_jump_forward": true}`})
Expect(on.enableJumpForward).To(Equal(int32(1)))
off := parseOptions(&pb.ModelOptions{EngineArgs: `{"enable_jump_forward": false}`})
Expect(off.enableJumpForward).To(Equal(int32(2)))
unset := parseOptions(&pb.ModelOptions{EngineArgs: `{"max_num_seqs": 4}`})
Expect(unset.enableJumpForward).To(Equal(int32(0)))
})
It("reads enable_jump_forward from the legacy options list too", func() {
lo := parseOptions(&pb.ModelOptions{Options: []string{"enable_jump_forward:true"}})
Expect(lo.enableJumpForward).To(Equal(int32(1)))
})
It("lets engine_args override the legacy options list", func() {
lo := parseOptions(&pb.ModelOptions{
Options: []string{"max_num_seqs:8", "block_size:16"},
EngineArgs: `{"max_num_seqs": 64}`,
})
Expect(lo.maxNumSeqs).To(Equal(int32(64))) // engine_args wins
Expect(lo.blockSize).To(Equal(int32(16))) // untouched keys survive
})
It("ignores malformed engine_args rather than failing the load", func() {
lo := parseOptions(&pb.ModelOptions{
Options: []string{"max_num_seqs:8"},
EngineArgs: `{not json`,
})
Expect(lo.maxNumSeqs).To(Equal(int32(8)))
})
It("ignores unknown keys", func() {
lo := parseOptions(&pb.ModelOptions{EngineArgs: `{"gpu_memory_utilization": 0.9}`})
Expect(lo).To(Equal(loadOptions{}))
})
})
var _ = Describe("samplingFromPredict", func() {
@@ -135,6 +292,91 @@ var _ = Describe("samplingFromPredict", func() {
})
})
// The engine resolves speculative_config.model against a local directory or
// ~/.cache/huggingface/hub ONLY - it never downloads. LocalAI keeps models in
// its own directory, so a bare repo id would miss the HF cache and fail deep in
// the load with a confusing "draft checkpoint not found". Resolve it here.
var _ = Describe("resolveDraftModelPath", func() {
var modelsDir string
BeforeEach(func() {
modelsDir = GinkgoT().TempDir()
})
// draftDir creates a plausible draft checkpoint under models/.
draftDir := func(name string) string {
d := filepath.Join(modelsDir, name)
Expect(os.MkdirAll(d, 0o750)).To(Succeed())
Expect(os.WriteFile(filepath.Join(d, "config.json"), []byte("{}"), 0o600)).To(Succeed())
return d
}
It("rewrites a repo id to the matching directory in the models dir", func() {
want := draftDir("Qwen3.6-27B-DFlash")
spec := `{"method":"dflash","model":"z-lab/Qwen3.6-27B-DFlash"}`
out, err := resolveDraftModelPath(spec, modelsDir)
Expect(err).ToNot(HaveOccurred())
Expect(out).To(MatchJSON(`{"method":"dflash","model":"` + want + `"}`))
})
It("rewrites a models-dir-relative path", func() {
want := draftDir("drafts__dflash")
spec := `{"method":"dflash","model":"drafts__dflash"}`
out, err := resolveDraftModelPath(spec, modelsDir)
Expect(err).ToNot(HaveOccurred())
Expect(out).To(ContainSubstring(want))
})
It("leaves an absolute path that already resolves alone", func() {
abs := draftDir("elsewhere")
spec := `{"method":"dflash","model":"` + abs + `"}`
out, err := resolveDraftModelPath(spec, modelsDir)
Expect(err).ToNot(HaveOccurred())
Expect(out).To(MatchJSON(spec))
})
It("fails with an actionable error when the draft is nowhere on disk", func() {
// Silently passing the repo id through would surface as an HF-cache
// miss inside the engine, which reads as "your model is broken".
spec := `{"method":"dflash","model":"z-lab/Not-Downloaded"}`
_, err := resolveDraftModelPath(spec, modelsDir)
Expect(err).To(HaveOccurred())
Expect(err.Error()).To(ContainSubstring("z-lab/Not-Downloaded"))
Expect(err.Error()).To(ContainSubstring(modelsDir))
})
It("requires a model key for dflash", func() {
_, err := resolveDraftModelPath(`{"method":"dflash"}`, modelsDir)
Expect(err).To(HaveOccurred())
Expect(err.Error()).To(ContainSubstring("model"))
})
It("leaves mtp and ngram configs untouched", func() {
// Neither has a separate draft checkpoint to resolve.
for _, spec := range []string{
`{"method":"mtp"}`,
`{"method":"ngram","num_speculative_tokens":4}`,
} {
out, err := resolveDraftModelPath(spec, modelsDir)
Expect(err).ToNot(HaveOccurred())
Expect(out).To(MatchJSON(spec))
}
})
It("passes a malformed document through for the engine to reject", func() {
// The engine owns config validation and produces the better message.
out, err := resolveDraftModelPath(`{not json`, modelsDir)
Expect(err).ToNot(HaveOccurred())
Expect(out).To(Equal(`{not json`))
})
It("is a no-op on an empty config", func() {
out, err := resolveDraftModelPath("", modelsDir)
Expect(err).ToNot(HaveOccurred())
Expect(out).To(BeEmpty())
})
})
var _ = Describe("validModelPath", func() {
It("accepts a .gguf file", func() {
dir := GinkgoT().TempDir()

View File

@@ -8,7 +8,7 @@ JOBS?=$(shell nproc --ignore=1)
# whisper.cpp version
WHISPER_REPO?=https://github.com/ggml-org/whisper.cpp
WHISPER_CPP_VERSION?=64d57d3df5c8dacee098577257edcaa154bf5ef3
WHISPER_CPP_VERSION?=306c88f4d1286aec1bf96e544632897886af5501
SO_TARGET?=libgowhisper.so
CMAKE_ARGS+=-DBUILD_SHARED_LIBS=OFF

View File

@@ -193,12 +193,22 @@
alias: "vllm-cpp"
license: apache-2.0
description: |
vllm.cpp is a from-scratch C++20 port of vLLM created and maintained by the LocalAI team.
It mirrors vLLM's V1 architecture (paged KV cache, continuous batching, prefix caching,
scheduler, sampler) on a portable tensor runtime with no Python, PyTorch or ggml at
inference time. It loads Hugging Face safetensors and GGUF checkpoints, supports
structured output (JSON schema / regex / choice / GBNF grammar) enforced in-engine,
and runs on CPU, NVIDIA CUDA (Blackwell-family), Apple Metal and Vulkan.
ALPHA development builds. Try it, but llama-cpp stays the recommendation for
production use.
vllm.cpp is an Apache-2.0 C++20 inference engine maintained by the LocalAI team,
developed in its own repository and usable without LocalAI. It began as a port of
vLLM and keeps vLLM as its reference implementation, checking output against it and
benchmarking against it, while growing a featureset of its own. It implements vLLM's
V1 architecture (paged KV cache, continuous batching, prefix caching, scheduler,
sampler) on a portable tensor runtime with no Python, PyTorch or ggml at inference
time. It loads GGUF as well as Hugging Face safetensors, supports structured output
(JSON schema / regex / choice / GBNF grammar) enforced in-engine, ships speculative
decoding and KV offload, and runs on CPU, NVIDIA CUDA (Blackwell-family), Apple
Metal and Vulkan.
The project is expected to be renamed as it diverges further from vLLM; the new
name is still to be decided.
urls:
- https://github.com/mudler/vllm.cpp
tags:
@@ -282,6 +292,55 @@
nvidia-cuda-12: "cuda12-parakeet-cpp"
nvidia-l4t-cuda-12: "nvidia-l4t-arm64-parakeet-cpp"
nvidia-l4t-cuda-13: "cuda13-nvidia-l4t-arm64-parakeet-cpp"
- &nemospeechcpp
name: "nemo-speech-cpp"
alias: "nemo-speech-cpp"
license: apache-2.0
icon: https://avatars.githubusercontent.com/u/1728152?s=200&v=4
description: |
NVIDIA NeMo-Speech.cpp, a C++/ggml runtime for NVIDIA Nemotron Speech models.
One backend serves four model families, selected automatically from the GGUF
general.architecture key: automatic speech recognition (offline, cache-aware
streaming and live transcription, with optional Silero VAD, punctuation,
inverse text normalization and Sortformer speaker diarization attached),
standalone Sortformer diarization, MagpieTTS text-to-speech over NanoCodec,
and Riva-Translate text translation. Runs on CPU, NVIDIA CUDA, Vulkan,
NVIDIA Jetson (L4T) and Apple Metal.
urls:
- https://github.com/NVIDIA/NeMo-Speech.cpp
tags:
- audio-transcription
- text-to-speech
- diarization
- text-to-text
- CPU
- GPU
- CUDA
- Metal
# No amd and no intel key on purpose: upstream NeMo-Speech.cpp has no ROCm/HIP
# and no SYCL backend, so there is nothing to point those at. A host reporting
# either capability falls through to "default" (SystemState.Capability) and
# gets the CPU build, which is the honest answer rather than a broken tag.
#
# Listing only nvidia-l4t would be a silent downgrade: a Jetson that reports a
# CUDA-refined capability would miss the map and fall back to the CPU build.
#
# The two nvidia-l4t-cuda-* keys point at DIFFERENT images on purpose. The
# JetPack r36.4.0 base links ggml against CUDA 12, so serving it to a host that
# reports nvidia-l4t-cuda-13 would fail at dlopen on a missing libcudart.so.12.
# That is worse than no key at all, since a missing key falls back to a working
# CPU build. Hence the separate cuda13 L4T image, as parakeet-cpp and
# moss-transcribe-cpp both do.
capabilities:
default: "cpu-nemo-speech-cpp"
nvidia: "cuda12-nemo-speech-cpp"
metal: "metal-nemo-speech-cpp"
vulkan: "vulkan-nemo-speech-cpp"
nvidia-l4t: "nvidia-l4t-arm64-nemo-speech-cpp"
nvidia-cuda-13: "cuda13-nemo-speech-cpp"
nvidia-cuda-12: "cuda12-nemo-speech-cpp"
nvidia-l4t-cuda-12: "nvidia-l4t-arm64-nemo-speech-cpp"
nvidia-l4t-cuda-13: "cuda13-nvidia-l4t-arm64-nemo-speech-cpp"
- &mosstranscribecpp
name: "moss-transcribe-cpp"
alias: "moss-transcribe-cpp"
@@ -3264,6 +3323,89 @@
uri: "quay.io/go-skynet/local-ai-backends:master-gpu-nvidia-cuda-13-parakeet-cpp"
mirrors:
- localai/localai-backends:master-gpu-nvidia-cuda-13-parakeet-cpp
## nemo-speech-cpp
- !!merge <<: *nemospeechcpp
name: "nemo-speech-cpp-development"
capabilities:
default: "cpu-nemo-speech-cpp-development"
nvidia: "cuda12-nemo-speech-cpp-development"
metal: "metal-nemo-speech-cpp-development"
vulkan: "vulkan-nemo-speech-cpp-development"
nvidia-l4t: "nvidia-l4t-arm64-nemo-speech-cpp-development"
nvidia-cuda-13: "cuda13-nemo-speech-cpp-development"
nvidia-cuda-12: "cuda12-nemo-speech-cpp-development"
nvidia-l4t-cuda-12: "nvidia-l4t-arm64-nemo-speech-cpp-development"
nvidia-l4t-cuda-13: "cuda13-nvidia-l4t-arm64-nemo-speech-cpp-development"
- !!merge <<: *nemospeechcpp
name: "cpu-nemo-speech-cpp"
uri: "quay.io/go-skynet/local-ai-backends:latest-cpu-nemo-speech-cpp"
mirrors:
- localai/localai-backends:latest-cpu-nemo-speech-cpp
- !!merge <<: *nemospeechcpp
name: "cpu-nemo-speech-cpp-development"
uri: "quay.io/go-skynet/local-ai-backends:master-cpu-nemo-speech-cpp"
mirrors:
- localai/localai-backends:master-cpu-nemo-speech-cpp
- !!merge <<: *nemospeechcpp
name: "cuda12-nemo-speech-cpp"
uri: "quay.io/go-skynet/local-ai-backends:latest-gpu-nvidia-cuda-12-nemo-speech-cpp"
mirrors:
- localai/localai-backends:latest-gpu-nvidia-cuda-12-nemo-speech-cpp
- !!merge <<: *nemospeechcpp
name: "cuda12-nemo-speech-cpp-development"
uri: "quay.io/go-skynet/local-ai-backends:master-gpu-nvidia-cuda-12-nemo-speech-cpp"
mirrors:
- localai/localai-backends:master-gpu-nvidia-cuda-12-nemo-speech-cpp
- !!merge <<: *nemospeechcpp
name: "cuda13-nemo-speech-cpp"
uri: "quay.io/go-skynet/local-ai-backends:latest-gpu-nvidia-cuda-13-nemo-speech-cpp"
mirrors:
- localai/localai-backends:latest-gpu-nvidia-cuda-13-nemo-speech-cpp
- !!merge <<: *nemospeechcpp
name: "cuda13-nemo-speech-cpp-development"
uri: "quay.io/go-skynet/local-ai-backends:master-gpu-nvidia-cuda-13-nemo-speech-cpp"
mirrors:
- localai/localai-backends:master-gpu-nvidia-cuda-13-nemo-speech-cpp
- !!merge <<: *nemospeechcpp
name: "vulkan-nemo-speech-cpp"
uri: "quay.io/go-skynet/local-ai-backends:latest-gpu-vulkan-nemo-speech-cpp"
mirrors:
- localai/localai-backends:latest-gpu-vulkan-nemo-speech-cpp
- !!merge <<: *nemospeechcpp
name: "vulkan-nemo-speech-cpp-development"
uri: "quay.io/go-skynet/local-ai-backends:master-gpu-vulkan-nemo-speech-cpp"
mirrors:
- localai/localai-backends:master-gpu-vulkan-nemo-speech-cpp
- !!merge <<: *nemospeechcpp
name: "nvidia-l4t-arm64-nemo-speech-cpp"
uri: "quay.io/go-skynet/local-ai-backends:latest-nvidia-l4t-arm64-nemo-speech-cpp"
mirrors:
- localai/localai-backends:latest-nvidia-l4t-arm64-nemo-speech-cpp
- !!merge <<: *nemospeechcpp
name: "nvidia-l4t-arm64-nemo-speech-cpp-development"
uri: "quay.io/go-skynet/local-ai-backends:master-nvidia-l4t-arm64-nemo-speech-cpp"
mirrors:
- localai/localai-backends:master-nvidia-l4t-arm64-nemo-speech-cpp
- !!merge <<: *nemospeechcpp
name: "cuda13-nvidia-l4t-arm64-nemo-speech-cpp"
uri: "quay.io/go-skynet/local-ai-backends:latest-nvidia-l4t-cuda-13-arm64-nemo-speech-cpp"
mirrors:
- localai/localai-backends:latest-nvidia-l4t-cuda-13-arm64-nemo-speech-cpp
- !!merge <<: *nemospeechcpp
name: "cuda13-nvidia-l4t-arm64-nemo-speech-cpp-development"
uri: "quay.io/go-skynet/local-ai-backends:master-nvidia-l4t-cuda-13-arm64-nemo-speech-cpp"
mirrors:
- localai/localai-backends:master-nvidia-l4t-cuda-13-arm64-nemo-speech-cpp
- !!merge <<: *nemospeechcpp
name: "metal-nemo-speech-cpp"
uri: "quay.io/go-skynet/local-ai-backends:latest-metal-darwin-arm64-nemo-speech-cpp"
mirrors:
- localai/localai-backends:latest-metal-darwin-arm64-nemo-speech-cpp
- !!merge <<: *nemospeechcpp
name: "metal-nemo-speech-cpp-development"
uri: "quay.io/go-skynet/local-ai-backends:master-metal-darwin-arm64-nemo-speech-cpp"
mirrors:
- localai/localai-backends:master-metal-darwin-arm64-nemo-speech-cpp
## moss-transcribe-cpp
- !!merge <<: *mosstranscribecpp
name: "moss-transcribe-cpp-development"

View File

@@ -420,6 +420,47 @@ var BackendCapabilities = map[string]BackendCapability{
DefaultUsecases: []string{UsecaseTranscript},
Description: "NVIDIA NeMo Parakeet ASR (parakeet.cpp)",
},
// nemo-speech-cpp is one gRPC server in front of four NeMo-Speech.cpp model
// families, picked at load time from the GGUF general.architecture key, so
// PossibleUsecases is their UNION and no single model serves all of it: an
// asr model transcribes (and diarizes, when a Sortformer model is attached
// through options), a sortformer model only diarizes, a magpietts model only
// synthesizes, and a Riva-Translate model only answers Predict.
//
// UsecaseChat sits alongside UsecaseCompletion for the translation family
// because Predict and PredictStream are exactly the RPCs /v1/chat/completions
// drives, and chat is what a translation model is useful through: each turn
// goes in as the prompt and comes back translated. The flag is not a gate on
// any endpoint (a request naming the model explicitly is served either way);
// what it buys is being eligible as the default chat model when a request
// names none (core/http/routes/openai.go) and appearing in the React UI's
// chat model picker (CAP_CHAT in react-ui/src/utils/capabilities.js).
//
// Leaving it out is not neutral: chat is a gallery filter key and completion
// is not (usecaseFilters in core/http/routes/ui_api.go), so
// GET /api/backends/usecases would grey the Chat filter out and hide a
// Riva-Translate gallery entry from the one filter that fits it.
//
// DefaultUsecases is transcript alone because that is the only family whose
// weights a bare `backend: nemo-speech-cpp` config is likely to name; a model
// of any other family should pin its own known_usecases.
//
// No VoiceCloning key: MagpieTTS synthesizes from baked speaker ids, not from
// a reference clip, so advertising cloning would accept a `voice:
// "profile:<id>"` request the backend cannot serve.
"nemo-speech-cpp": {
GRPCMethods: []GRPCMethod{
MethodAudioTranscription, MethodDiarize,
MethodTTS, MethodTTSStream,
MethodPredict, MethodPredictStream,
},
PossibleUsecases: []string{
UsecaseTranscript, UsecaseDiarization, UsecaseTTS,
UsecaseCompletion, UsecaseChat,
},
DefaultUsecases: []string{UsecaseTranscript},
Description: "NVIDIA NeMo-Speech.cpp: one server for Nemotron ASR (offline, streaming and live), Sortformer diarization, MagpieTTS synthesis and Riva-Translate translation; the model's GGUF architecture decides which",
},
"qwen-asr": {
GRPCMethods: []GRPCMethod{MethodAudioTranscription},
PossibleUsecases: []string{UsecaseTranscript},

View File

@@ -124,6 +124,39 @@ var _ = Describe("GetBackendCapability", func() {
})
})
// nemo-speech-cpp fronts four model families from one server, and its
// PossibleUsecases is their union. The entry has to stay in step with what
// docs/content/features/nemo-speech-cpp.md tells operators to put in
// known_usecases: nothing validates known_usecases against PossibleUsecases, so
// a flag the docs recommend and the map omits fails silently, and the place it
// surfaces is the gallery. GET /api/backends/usecases is derived from this list
// and greys out the filters missing from it, so a recommended-but-unlisted flag
// hides the very models it was recommended for.
var _ = Describe("nemo-speech-cpp capabilities", func() {
It("advertises every usecase its four families serve", func() {
capability := GetBackendCapability("nemo-speech-cpp")
Expect(capability).NotTo(BeNil())
Expect(capability.PossibleUsecases).To(ContainElements(
UsecaseTranscript, UsecaseDiarization, UsecaseTTS,
UsecaseCompletion, UsecaseChat))
})
// Chat is the translation family's usecase, and it needs both Predict RPCs:
// /v1/chat/completions streams through PredictStream and answers
// non-streaming requests through Predict.
It("backs the chat usecase with the RPCs chat actually drives", func() {
capability := GetBackendCapability("nemo-speech-cpp")
Expect(capability.GRPCMethods).To(ContainElements(MethodPredict, MethodPredictStream))
})
// Defaults stay conservative: a bare `backend: nemo-speech-cpp` with no
// known_usecases is overwhelmingly an ASR model, and every other family is
// expected to pin its own flags.
It("still defaults to transcript alone", func() {
Expect(DefaultUsecasesForBackendCap("nemo-speech-cpp")).To(Equal([]string{UsecaseTranscript}))
})
})
// audio-cpp advertises voice cloning from the backend itself and ships
// audio-cpp-chatterbox, whose family serves cloning and NOT plain TTS, so a
// reference clip is the only way to use it. Without a capability entry

117
core/config/vllm_spec.go Normal file
View File

@@ -0,0 +1,117 @@
package config
// Speculative-decoding auto-defaults for the vllm-cpp backend, the safetensors
// counterpart of the GGUF/llama.cpp hook in mtp.go.
//
// The two engines detect and spell the same feature differently. llama.cpp
// reads `<arch>.nextn_predict_layers` out of the GGUF header and takes
// `spec_type:draft-mtp` in `options:`; vllm.cpp reads `mtp_num_hidden_layers`
// out of the checkpoint's config.json and takes vLLM's own
// `--speculative-config` JSON, which LocalAI carries in `engine_args`. The
// engine resolves the draft depth and the default k itself, so the config only
// has to name the method.
import (
"encoding/json"
"github.com/mudler/xlog"
)
// hfSpecConfig is the subset of a HuggingFace config.json that decides whether
// speculative decoding can be auto-enabled.
type hfSpecConfig struct {
ModelType string `json:"model_type"`
// MtpNumHiddenLayers is the MTP head depth (upstream speculative.py reads
// it as n_predict for the qwen3_5 / qwen3_5_moe families).
MtpNumHiddenLayers uint32 `json:"mtp_num_hidden_layers"`
// DFlashConfig marks a z-lab DFlash DRAFT checkpoint (mask_token_id +
// target_layer_ids). Its presence means this repo is a draft, not a
// servable target.
DFlashConfig json.RawMessage `json:"dflash_config"`
// TextConfig is where multimodal checkpoints nest the language-model
// config, and therefore the MTP depth.
TextConfig *hfSpecConfig `json:"text_config"`
}
// parseHFSpecConfig decodes the speculative-relevant subset of a config.json.
// A document that does not parse yields nothing rather than an error: detection
// is best-effort and must never break an import.
func parseHFSpecConfig(configJSON []byte) (hfSpecConfig, bool) {
if len(configJSON) == 0 {
return hfSpecConfig{}, false
}
var c hfSpecConfig
if err := json.Unmarshal(configJSON, &c); err != nil {
xlog.Debug("[vllm-spec] config.json did not parse; skipping detection", "error", err)
return hfSpecConfig{}, false
}
return c, true
}
// IsDFlashDraftConfig reports whether a HuggingFace config.json describes a
// DFlash DRAFT checkpoint. Unlike MTP - whose head ships inside the target
// checkpoint's `mtp.*` tensors - a DFlash draft is its own repo that can only
// run paired with a target it verifies against, so it must never be configured
// as a standalone model.
func IsDFlashDraftConfig(configJSON []byte) bool {
c, ok := parseHFSpecConfig(configJSON)
if !ok {
return false
}
return len(c.DFlashConfig) > 0 ||
(c.TextConfig != nil && len(c.TextConfig.DFlashConfig) > 0)
}
// HasSafetensorsMTPHead reports whether a HuggingFace config.json declares a
// self-speculating Multi-Token Prediction head, returning its depth. The depth
// is informational: vllm.cpp resolves n_predict and the default
// num_speculative_tokens from the checkpoint itself.
//
// DFlash drafts are excluded for the same reason `gemma4-assistant` GGUFs are
// excluded from the llama.cpp hook: they carry head metadata but cannot
// self-speculate.
//
// NOTE this is a safetensors-only signal. vllm.cpp rejects an MTP config over a
// GGUF source, because the `mtp.*` draft tensors only exist in the safetensors
// checkpoint - so the GGUF import path must not use this.
func HasSafetensorsMTPHead(configJSON []byte) (uint32, bool) {
c, ok := parseHFSpecConfig(configJSON)
if !ok {
return 0, false
}
if IsDFlashDraftConfig(configJSON) {
return 0, false
}
n := c.MtpNumHiddenLayers
if n == 0 && c.TextConfig != nil {
n = c.TextConfig.MtpNumHiddenLayers
}
return n, n > 0
}
// ApplyVLLMSpeculativeDefaults enables MTP speculative decoding in cfg's
// engine_args when nothing is configured there yet. It is a no-op when the user
// already set a speculative_config, so an explicit choice (a different method,
// an explicit k, a DFlash draft) is never clobbered.
//
// `layers` is the detected head depth and is only used for the diagnostic log
// line - the engine derives the real k from the checkpoint.
func ApplyVLLMSpeculativeDefaults(cfg *ModelConfig, layers uint32) {
if cfg == nil {
return
}
if _, set := cfg.EngineArgs["speculative_config"]; set {
xlog.Debug("[vllm-spec] MTP head detected but speculative_config already configured; leaving user choice intact",
"name", cfg.Name, "mtp_num_hidden_layers", layers)
return
}
if cfg.EngineArgs == nil {
cfg.EngineArgs = map[string]any{}
}
// Only the method: vllm.cpp defaults num_speculative_tokens to the
// checkpoint's own n_predict (speculative.py:865-875), which is the right
// value far more reliably than anything guessable here.
cfg.EngineArgs["speculative_config"] = map[string]any{"method": "mtp"}
xlog.Info("[vllm-spec] MTP head detected; enabling mtp speculative decoding",
"name", cfg.Name, "mtp_num_hidden_layers", layers)
}

View File

@@ -0,0 +1,117 @@
package config_test
import (
. "github.com/mudler/LocalAI/core/config"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
var _ = Describe("vllm-cpp speculative-decoding auto-defaults", func() {
Context("HasSafetensorsMTPHead", func() {
It("detects a top-level mtp_num_hidden_layers", func() {
n, ok := HasSafetensorsMTPHead([]byte(`{
"model_type": "qwen3_5_moe",
"mtp_num_hidden_layers": 1
}`))
Expect(ok).To(BeTrue())
Expect(n).To(Equal(uint32(1)))
})
It("detects the head nested under text_config", func() {
// Multimodal checkpoints nest the language-model config, which is
// where the MTP depth lives (mirrors the engine's own resolution
// off config.raw text_config).
n, ok := HasSafetensorsMTPHead([]byte(`{
"model_type": "qwen3_5_moe",
"text_config": {"mtp_num_hidden_layers": 2}
}`))
Expect(ok).To(BeTrue())
Expect(n).To(Equal(uint32(2)))
})
It("reports no head when the key is absent", func() {
n, ok := HasSafetensorsMTPHead([]byte(`{"model_type": "llama"}`))
Expect(ok).To(BeFalse())
Expect(n).To(BeZero())
})
It("reports no head for a zero depth", func() {
_, ok := HasSafetensorsMTPHead([]byte(`{"mtp_num_hidden_layers": 0}`))
Expect(ok).To(BeFalse())
})
It("ignores a DFlash draft checkpoint", func() {
// A DFlash draft is a SEPARATE checkpoint that cannot serve alone:
// it needs a target to verify against. Same exclusion the GGUF path
// makes for gemma4-assistant drafts.
_, ok := HasSafetensorsMTPHead([]byte(`{
"model_type": "qwen3_dflash",
"mtp_num_hidden_layers": 1,
"dflash_config": {"mask_token_id": 151666, "target_layer_ids": [0, 1]}
}`))
Expect(ok).To(BeFalse())
})
It("reports no head on unparseable JSON", func() {
_, ok := HasSafetensorsMTPHead([]byte(`{not json`))
Expect(ok).To(BeFalse())
})
It("reports no head on empty input", func() {
_, ok := HasSafetensorsMTPHead(nil)
Expect(ok).To(BeFalse())
})
})
Context("IsDFlashDraftConfig", func() {
It("recognises a draft by its dflash_config block", func() {
Expect(IsDFlashDraftConfig([]byte(`{
"dflash_config": {"mask_token_id": 151666, "target_layer_ids": [0]}
}`))).To(BeTrue())
})
It("does not flag an ordinary checkpoint", func() {
Expect(IsDFlashDraftConfig([]byte(`{"model_type": "qwen3_5_moe"}`))).To(BeFalse())
})
})
Context("ApplyVLLMSpeculativeDefaults", func() {
It("writes the mtp method into engine_args", func() {
cfg := &ModelConfig{Name: "qwen"}
ApplyVLLMSpeculativeDefaults(cfg, 1)
Expect(cfg.EngineArgs).To(HaveKey("speculative_config"))
spec, ok := cfg.EngineArgs["speculative_config"].(map[string]any)
Expect(ok).To(BeTrue())
Expect(spec["method"]).To(Equal("mtp"))
})
It("leaves an existing speculative_config alone", func() {
cfg := &ModelConfig{
Name: "qwen",
LLMConfig: LLMConfig{
EngineArgs: map[string]any{
"speculative_config": map[string]any{"method": "ngram", "num_speculative_tokens": 4},
},
},
}
ApplyVLLMSpeculativeDefaults(cfg, 1)
spec := cfg.EngineArgs["speculative_config"].(map[string]any)
Expect(spec["method"]).To(Equal("ngram"))
})
It("preserves unrelated engine_args keys", func() {
cfg := &ModelConfig{
Name: "qwen",
LLMConfig: LLMConfig{EngineArgs: map[string]any{"max_num_seqs": 32}},
}
ApplyVLLMSpeculativeDefaults(cfg, 1)
Expect(cfg.EngineArgs).To(HaveKeyWithValue("max_num_seqs", 32))
Expect(cfg.EngineArgs).To(HaveKey("speculative_config"))
})
It("tolerates a nil config", func() {
Expect(func() { ApplyVLLMSpeculativeDefaults(nil, 1) }).ToNot(Panic())
})
})
})

View File

@@ -9,6 +9,7 @@ import (
"time"
"github.com/mudler/LocalAI/core/config"
"github.com/mudler/LocalAI/pkg/concurrency"
"github.com/mudler/LocalAI/pkg/system"
"github.com/mudler/LocalAI/pkg/vram"
"github.com/mudler/xlog"
@@ -101,7 +102,7 @@ func WarmEstimateCache(ctx context.Context, galleries []config.Gallery, systemSt
return
}
go func() {
concurrency.SafeGo(func() {
started := time.Now()
models, err := AvailableGalleryModelsCached(galleries, systemState)
@@ -131,7 +132,7 @@ func WarmEstimateCache(ctx context.Context, galleries []config.Gallery, systemSt
for i := 0; i < cfg.Concurrency; i++ {
wg.Add(1)
go func() {
concurrency.SafeGo(func() {
defer wg.Done()
for m := range cursor {
// Per entry, not for the run: one unreachable weight file
@@ -164,7 +165,7 @@ func WarmEstimateCache(ctx context.Context, galleries []config.Gallery, systemSt
cancel()
}
}()
})
}
feed:
@@ -183,7 +184,7 @@ func WarmEstimateCache(ctx context.Context, galleries []config.Gallery, systemSt
return
}
xlog.Info("gallery caches warmed", "estimates", warmed, "variants", warmedVariants, "of", len(models), "took", time.Since(started).Round(time.Second))
}()
})
}
// EstimateWarmConfigFromEnv reads the warm-up bounds from the environment,

View File

@@ -1,11 +1,20 @@
package gallery_test
import (
"bytes"
"context"
"encoding/binary"
"math"
"net/http"
"net/http/httptest"
"os"
"path/filepath"
"time"
gguf "github.com/gpustack/gguf-parser-go"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
"gopkg.in/yaml.v3"
"github.com/mudler/LocalAI/core/config"
"github.com/mudler/LocalAI/core/gallery"
@@ -57,6 +66,46 @@ var _ = Describe("VRAM estimate warm-up", func() {
Consistently(func() bool { return true }, "100ms").Should(BeTrue())
})
It("does not crash the server when remote GGUF metadata is malformed", func() {
payload := warmMalformedGGUF()
requested := make(chan struct{})
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
select {
case <-requested:
default:
close(requested)
}
http.ServeContent(w, r, "model.gguf", time.Time{}, bytes.NewReader(payload))
}))
DeferCleanup(server.Close)
galleryPath := filepath.Join(state.Model.ModelsPath, "malformed-gallery.yaml")
index, err := yaml.Marshal([]gallery.GalleryModel{{Metadata: gallery.Metadata{
Name: "malformed-gguf",
AdditionalFiles: []gallery.File{{
Filename: "model.gguf",
URI: server.URL + "/model.gguf",
}},
}}})
Expect(err).NotTo(HaveOccurred())
Expect(os.WriteFile(galleryPath, index, 0600)).To(Succeed())
cfg := gallery.DefaultEstimateWarmConfig
cfg.Limit = 1
cfg.Concurrency = 1
cfg.Contexts = []uint32{8192}
gallery.WarmEstimateCache(context.Background(), []config.Gallery{{
Name: "malformed",
URL: "file://" + galleryPath,
}}, state, cfg)
Eventually(requested, "2s").Should(BeClosed())
// The warm-up is detached. Give its parser time to consume the response;
// before the recovery boundary, that goroutine panicked and killed the
// entire test process (and the LocalAI server in production).
Consistently(func() bool { return true }, "300ms").Should(BeTrue())
})
Describe("configuration from the environment", func() {
AfterEach(func() {
os.Unsetenv("LOCALAI_VRAM_WARM_LIMIT")
@@ -113,3 +162,19 @@ var _ = Describe("VRAM estimate warm-up", func() {
})
})
func warmMalformedGGUF() []byte {
payload := make([]byte, 0, 128)
payload = binary.LittleEndian.AppendUint32(payload, uint32(gguf.GGUFMagicGGUFLe))
payload = binary.LittleEndian.AppendUint32(payload, uint32(gguf.GGUFVersionV3))
payload = binary.LittleEndian.AppendUint64(payload, 0)
payload = binary.LittleEndian.AppendUint64(payload, 1)
key := "tokenizer.ggml.tokens"
payload = binary.LittleEndian.AppendUint64(payload, uint64(len(key)))
payload = append(payload, key...)
payload = binary.LittleEndian.AppendUint32(payload, uint32(gguf.GGUFMetadataValueTypeArray))
payload = binary.LittleEndian.AppendUint32(payload, uint32(gguf.GGUFMetadataValueTypeString))
payload = binary.LittleEndian.AppendUint64(payload, 1)
payload = binary.LittleEndian.AppendUint64(payload, math.MaxUint64)
return payload
}

View File

@@ -298,7 +298,15 @@ func (i *LlamaCPPImporter) Import(details Details) (gallery.ModelConfig, error)
// imported configs already carry spec_type:draft-mtp before the model is
// ever loaded - users see it in the YAML preview rather than discovering
// it after the first start.
maybeApplyMTPDefaults(&modelConfig, details, &cfg)
//
// vllm-cpp is excluded on both counts: `spec_type:*` are llama.cpp option
// keys it does not read, and vllm.cpp rejects an MTP config over a GGUF
// source outright (the `mtp.*` draft tensors exist only in the safetensors
// checkpoint). Its MTP auto-config runs in the vllm importer instead, over
// the safetensors config.json.
if backend != "vllm-cpp" {
maybeApplyMTPDefaults(&modelConfig, details, &cfg)
}
data, err := yaml.Marshal(modelConfig)
if err != nil {
@@ -401,7 +409,10 @@ func maybeApplyMTPDefaults(modelConfig *config.ModelConfig, details Details, cfg
}
}()
f, err := gguf.ParseGGUFFileRemote(ctx, probeURL)
// MTP markers are architecture scalars. Avoid allocating tokenizer and
// other large arrays from an untrusted remote header; panic recovery cannot
// contain a fatal out-of-memory condition.
f, err := gguf.ParseGGUFFileRemote(ctx, probeURL, gguf.SkipLargeMetadata())
if err != nil {
xlog.Debug("[mtp-importer] failed to read remote GGUF header for MTP detection", "uri", probeURL, "error", err)
return

View File

@@ -1,13 +1,21 @@
package importers
import (
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"path/filepath"
"strings"
"time"
"github.com/mudler/LocalAI/core/config"
"github.com/mudler/LocalAI/core/gallery"
"github.com/mudler/LocalAI/core/schema"
"github.com/mudler/LocalAI/pkg/downloader"
"github.com/mudler/LocalAI/pkg/httpclient"
"github.com/mudler/xlog"
"go.yaml.in/yaml/v2"
)
@@ -107,6 +115,12 @@ func (i *VLLMImporter) Import(details Details) (gallery.ModelConfig, error) {
// vllm python backend, so use_tokenizer_template carries over), but
// tool/reasoning parsing is the engine's own autoparser pipeline -
// the vllm-python tool_parser/reasoning_parser options don't apply.
//
// Auto-detect a Multi-Token Prediction head, the safetensors analogue
// of the llama-cpp importer's GGUF hook, so a freshly imported
// Qwen3.5 / Qwen3.6 config already carries speculative decoding in its
// engine_args instead of leaving the throughput on the table.
maybeApplyVLLMSpeculativeDefaults(&modelConfig, details)
} else {
// Auto-detect tool_parser and reasoning_parser for known model families.
// Surfacing them in the generated YAML lets users see and edit the choices.
@@ -132,3 +146,89 @@ func (i *VLLMImporter) Import(details Details) (gallery.ModelConfig, error) {
ConfigFile: string(data),
}, nil
}
// maxSpecConfigProbeBytes caps the config.json body we read. Real ones are a
// few KB; the cap keeps a hostile or mislabelled URL from streaming into the
// importer.
const maxSpecConfigProbeBytes = 1 << 20 // 1 MiB
// specConfigProbeTimeout bounds the config.json fetch. Detection is an
// optimisation, so it must never hold an import open for long.
const specConfigProbeTimeout = 30 * time.Second
// specConfigFetcher is the seam the config.json probe goes through, so tests can
// drive the whole import path without a network round trip.
var specConfigFetcher = fetchProbeBody
// maybeApplyVLLMSpeculativeDefaults fetches the repository's config.json and,
// when it declares a Multi-Token Prediction head, enables MTP speculative
// decoding in the emitted engine_args. This is the safetensors counterpart of
// the llama-cpp importer's GGUF header probe.
//
// Every failure is non-fatal and logged at debug: a network blip, a private
// repo, or a config.json this doesn't understand must leave the import working
// exactly as it did before, just without the speculative default.
func maybeApplyVLLMSpeculativeDefaults(modelConfig *config.ModelConfig, details Details) {
probeURL := vllmSpecProbeURL(details)
if probeURL == "" {
return
}
body, err := specConfigFetcher(probeURL)
if err != nil {
xlog.Debug("[vllm-spec-importer] could not read config.json for MTP detection", "uri", probeURL, "error", err)
return
}
applySpecFromConfigJSON(modelConfig, body, details.URI)
}
// applySpecFromConfigJSON is the decision half of the probe, split out so it can
// be exercised without a network round trip.
func applySpecFromConfigJSON(modelConfig *config.ModelConfig, body []byte, uri string) {
if config.IsDFlashDraftConfig(body) {
// A DFlash draft cannot serve on its own - it only proposes tokens for
// a target model to verify. Say so rather than emitting a config that
// would fail at load.
xlog.Warn("[vllm-spec-importer] this repository is a DFlash DRAFT checkpoint, not a servable model; "+
"import the TARGET model and point engine_args.speculative_config at this repo "+
`({"method":"dflash","model":"<this repo>"})`, "uri", uri)
return
}
n, ok := config.HasSafetensorsMTPHead(body)
if !ok {
return
}
config.ApplyVLLMSpeculativeDefaults(modelConfig, n)
}
// vllmSpecProbeURL returns the HTTP(S) URL of the repository's config.json, or
// "" when the import isn't backed by a HuggingFace repo we can fetch from (a
// local directory import, an OCI artifact, ...).
func vllmSpecProbeURL(details Details) string {
if details.HuggingFace == nil || details.HuggingFace.ModelID == "" {
return ""
}
return resolveHTTPProbe(downloader.HuggingFacePrefix + details.HuggingFace.ModelID + "/config.json")
}
// fetchProbeBody GETs a small remote JSON document under a short timeout.
func fetchProbeBody(url string) ([]byte, error) {
ctx, cancel := context.WithTimeout(context.Background(), specConfigProbeTimeout)
defer cancel()
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return nil, err
}
resp, err := httpclient.NewWithTimeout(specConfigProbeTimeout).Do(req)
if err != nil {
return nil, err
}
defer func() { _ = resp.Body.Close() }()
if resp.StatusCode != http.StatusOK {
return nil, fmt.Errorf("unexpected status %d", resp.StatusCode)
}
return io.ReadAll(io.LimitReader(resp.Body, maxSpecConfigProbeBytes))
}

View File

@@ -0,0 +1,118 @@
package importers
import (
"encoding/json"
"errors"
"github.com/mudler/LocalAI/core/config"
hfapi "github.com/mudler/LocalAI/pkg/huggingface-api"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
var _ = Describe("vllm-cpp speculative auto-config (importer)", func() {
Context("applySpecFromConfigJSON", func() {
It("enables mtp when the checkpoint declares an MTP head", func() {
cfg := &config.ModelConfig{Name: "qwen3.5"}
applySpecFromConfigJSON(cfg, []byte(`{
"model_type": "qwen3_5_moe",
"mtp_num_hidden_layers": 1
}`), "huggingface://Qwen/Qwen3.5-A3B")
Expect(cfg.EngineArgs).To(HaveKeyWithValue("speculative_config",
map[string]any{"method": "mtp"}))
})
It("leaves a plain checkpoint untouched", func() {
cfg := &config.ModelConfig{Name: "llama"}
applySpecFromConfigJSON(cfg, []byte(`{"model_type": "llama"}`), "huggingface://meta/llama")
Expect(cfg.EngineArgs).To(BeEmpty())
})
It("refuses to configure a DFlash draft as a servable model", func() {
// The draft only proposes tokens; configuring it standalone would
// produce a model that cannot load.
cfg := &config.ModelConfig{Name: "dflash-draft"}
applySpecFromConfigJSON(cfg, []byte(`{
"model_type": "qwen3_dflash",
"dflash_config": {"mask_token_id": 151666, "target_layer_ids": [0, 1]}
}`), "huggingface://z-lab/Qwen3.6-27B-DFlash")
Expect(cfg.EngineArgs).To(BeEmpty())
})
It("survives a config.json it cannot parse", func() {
cfg := &config.ModelConfig{Name: "weird"}
Expect(func() {
applySpecFromConfigJSON(cfg, []byte(`<html>404</html>`), "huggingface://a/b")
}).ToNot(Panic())
Expect(cfg.EngineArgs).To(BeEmpty())
})
})
Context("Import over a repository with an MTP head", func() {
var restore func()
BeforeEach(func() {
original := specConfigFetcher
restore = func() { specConfigFetcher = original }
})
AfterEach(func() { restore() })
importWith := func(backend, configJSON string) string {
specConfigFetcher = func(string) ([]byte, error) {
return []byte(configJSON), nil
}
importer := &VLLMImporter{}
out, err := importer.Import(Details{
URI: "huggingface://Qwen/Qwen3.5-A3B",
Preferences: json.RawMessage(`{"backend": "` + backend + `"}`),
HuggingFace: &hfapi.ModelDetails{ModelID: "Qwen/Qwen3.5-A3B"},
})
Expect(err).ToNot(HaveOccurred())
return out.ConfigFile
}
It("emits engine_args.speculative_config for vllm-cpp", func() {
yaml := importWith("vllm-cpp", `{"model_type":"qwen3_5_moe","mtp_num_hidden_layers":1}`)
Expect(yaml).To(ContainSubstring("engine_args:"))
Expect(yaml).To(ContainSubstring("speculative_config:"))
Expect(yaml).To(ContainSubstring("method: mtp"))
})
It("emits nothing speculative for the python vllm backend", func() {
// The python backend has its own speculative surface and its own
// version-dependent MTP support; this hook is vllm-cpp only.
yaml := importWith("vllm", `{"model_type":"qwen3_5_moe","mtp_num_hidden_layers":1}`)
Expect(yaml).NotTo(ContainSubstring("speculative_config"))
})
It("emits nothing speculative when the probe fails", func() {
specConfigFetcher = func(string) ([]byte, error) {
return nil, errors.New("network down")
}
importer := &VLLMImporter{}
out, err := importer.Import(Details{
URI: "huggingface://Qwen/Qwen3.5-A3B",
Preferences: json.RawMessage(`{"backend": "vllm-cpp"}`),
HuggingFace: &hfapi.ModelDetails{ModelID: "Qwen/Qwen3.5-A3B"},
})
Expect(err).ToNot(HaveOccurred())
Expect(out.ConfigFile).NotTo(ContainSubstring("speculative_config"))
})
})
Context("vllmSpecProbeURL", func() {
It("resolves the repository's config.json to an HTTPS URL", func() {
url := vllmSpecProbeURL(Details{
URI: "huggingface://Qwen/Qwen3.5-A3B",
HuggingFace: &hfapi.ModelDetails{ModelID: "Qwen/Qwen3.5-A3B"},
})
Expect(url).To(ContainSubstring("Qwen/Qwen3.5-A3B"))
Expect(url).To(HaveSuffix("config.json"))
Expect(url).To(HavePrefix("https://"))
})
It("skips the probe when there is no HuggingFace repo behind the import", func() {
Expect(vllmSpecProbeURL(Details{URI: "/models/local-dir"})).To(BeEmpty())
})
})
})

View File

@@ -37,6 +37,17 @@ var knownPrefOnlyBackends = []schema.KnownBackend{
// ASR
{Name: "whisperx", Modality: "asr", AutoDetect: false, Description: "WhisperX transcription (preference-only)"},
{Name: "crispasr", Modality: "asr", AutoDetect: false, Description: "CrispASR multi-architecture transcription (preference-only)"},
// nemo-speech-cpp serves four families from one backend, picked at load time
// from the GGUF general.architecture key. Modality is a single string and the
// import form chips on a fixed key set
// (core/http/react-ui/src/components/ModalityChips.jsx), so "asr" is the one
// it carries and the other three are named in the description rather than in
// a key the UI would bucket as "other".
// No importer: general.architecture lives inside the GGUF and cannot be read
// from a remote HuggingFace repo, and the NMT family carries an ordinary LLM
// architecture (qwen3), so even a local probe has no NeMo-specific string to
// match on.
{Name: "nemo-speech-cpp", Modality: "asr", AutoDetect: false, Description: "NVIDIA NeMo-Speech.cpp: Nemotron ASR (offline, streaming and live), Sortformer diarization, MagpieTTS speech synthesis and Riva-Translate translation, chosen from the model's GGUF architecture (preference-only)"},
// TTS
{Name: "mlx-audio", Modality: "tts", AutoDetect: false, Description: "MLX-Audio text-to-speech models (auto-detected; pref-only fallback)"},
{Name: "kokoros", Modality: "tts", AutoDetect: false, Description: "Kokoros TTS (preference-only)"},

View File

@@ -297,6 +297,36 @@ var _ = Describe("Backend Endpoints", func() {
"the description must name the modalities the single Modality field cannot")
})
It("advertises nemo-speech-cpp as a preference-only backend naming its other families", func() {
req := httptest.NewRequest(http.MethodGet, "/backends/known", nil)
rec := httptest.NewRecorder()
app.ServeHTTP(rec, req)
var payload []schema.KnownBackend
Expect(json.Unmarshal(rec.Body.Bytes(), &payload)).To(Succeed())
byName := map[string]schema.KnownBackend{}
for _, b := range payload {
byName[b.Name] = b
}
entry, ok := byName["nemo-speech-cpp"]
Expect(ok).To(BeTrue(), "nemo-speech-cpp must appear in the import form dropdown")
// AutoDetect=false is the honest answer, not an omission: the family
// comes from general.architecture inside the GGUF, which no remote-repo
// probe can read, and the translation family carries an ordinary LLM
// architecture that nothing NeMo-specific distinguishes.
Expect(entry.AutoDetect).To(BeFalse(),
"nemo-speech-cpp has no remote-detectable signal: the family lives inside the GGUF")
Expect(entry.Modality).To(Equal("asr"))
// The single Modality field cannot carry the other three families, so
// the description has to name them or the dropdown claims this backend
// only transcribes.
Expect(entry.Description).To(ContainSubstring("diarization"))
Expect(entry.Description).To(ContainSubstring("synthesis"))
Expect(entry.Description).To(ContainSubstring("translation"))
})
It("is sorted by Modality then Name", func() {
req := httptest.NewRequest(http.MethodGet, "/backends/known", nil)
rec := httptest.NewRecorder()

View File

@@ -60,6 +60,7 @@ type APIExchange struct {
}
var traceBuffer *circularbuffer.Queue[APIExchange]
var inFlightTraces = make(map[string]APIExchange)
var mu sync.Mutex
var logChan = make(chan traceCommand, 100)
var traceIDSeq atomic.Uint64
@@ -126,16 +127,17 @@ func initializeTracing(dataPath string, maxItems int) {
continue
}
exchange := *command.exchange
mu.Lock()
delete(inFlightTraces, exchange.ID)
if traceBuffer != nil {
traceBuffer.Enqueue(exchange)
}
mu.Unlock()
if command.store != nil {
if err := command.store.Append(exchange.ID, exchange); err != nil {
xlog.Warn("Failed to persist API trace", "error", err)
}
}
mu.Lock()
if traceBuffer != nil {
traceBuffer.Enqueue(exchange)
}
mu.Unlock()
}
}()
})
@@ -261,6 +263,38 @@ func TraceMiddleware(app *application.Application) echo.MiddlewareFunc {
// tens of MB, which then locks the admin Traces UI fetching the
// JSON dump faster than the 5s auto-refresh.
maxBodyBytes := app.ApplicationConfig().TracingMaxBodyBytes
requestHeaders := redactSensitiveHeaders(c.Request().Header)
requestBody, requestTruncated := truncateForTrace(body, maxBodyBytes)
exchange := APIExchange{
ID: nextTraceID(),
Timestamp: startTime,
ClientIP: c.RealIP(),
UserAgent: c.Request().UserAgent(),
Request: APIExchangeRequest{
Method: c.Request().Method,
Path: c.Path(),
Headers: &requestHeaders,
Body: &requestBody,
BodyTruncated: requestTruncated,
BodyBytes: len(body),
},
}
if user := auth.GetUser(c); user != nil {
exchange.UserID = user.ID
exchange.UserName = user.Name
}
mu.Lock()
inFlightTraces[exchange.ID] = exchange
mu.Unlock()
queued := false
defer func() {
if queued {
return
}
mu.Lock()
delete(inFlightTraces, exchange.ID)
mu.Unlock()
}()
// Wrap response writer to capture body
resBody := new(bytes.Buffer)
@@ -287,47 +321,27 @@ func TraceMiddleware(app *application.Application) echo.MiddlewareFunc {
// the trace endpoint is admin-only but the buffer is also reachable
// via any heap-dump-style introspection, and tokens shouldn't
// outlive the request that carried them.
requestHeaders := redactSensitiveHeaders(c.Request().Header)
requestBody, requestTruncated := truncateForTrace(body, maxBodyBytes)
responseHeaders := redactSensitiveHeaders(c.Response().Header())
responseBody := make([]byte, resBody.Len())
copy(responseBody, resBody.Bytes())
exchange := APIExchange{
ID: nextTraceID(),
Timestamp: startTime,
Duration: time.Since(startTime),
ClientIP: c.RealIP(),
UserAgent: c.Request().UserAgent(),
Request: APIExchangeRequest{
Method: c.Request().Method,
Path: c.Path(),
Headers: &requestHeaders,
Body: &requestBody,
BodyTruncated: requestTruncated,
BodyBytes: len(body),
},
Response: APIExchangeResponse{
Status: status,
Headers: &responseHeaders,
Body: &responseBody,
BodyTruncated: mw.truncated,
BodyBytes: mw.totalBytes,
},
exchange.Duration = time.Since(startTime)
exchange.Response = APIExchangeResponse{
Status: status,
Headers: &responseHeaders,
Body: &responseBody,
BodyTruncated: mw.truncated,
BodyBytes: mw.totalBytes,
}
if handlerErr != nil {
exchange.Error = handlerErr.Error()
}
if user := auth.GetUser(c); user != nil {
exchange.UserID = user.ID
exchange.UserName = user.Name
}
mu.Lock()
store := traceStore
mu.Unlock()
select {
case logChan <- traceCommand{exchange: &exchange, store: store}:
queued = true
default:
xlog.Warn("Trace channel full, dropping trace")
}
@@ -345,6 +359,10 @@ func GetTraces() []APIExchange {
return []APIExchange{}
}
traces := traceBuffer.Values()
for _, exchange := range inFlightTraces {
exchange.Duration = time.Since(exchange.Timestamp)
traces = append(traces, exchange)
}
mu.Unlock()
slices.SortFunc(traces, func(a, b APIExchange) int {

View File

@@ -0,0 +1,108 @@
// SPDX-License-Identifier: MIT
package middleware
import (
"net/http"
"net/http/httptest"
"time"
"github.com/labstack/echo/v4"
"github.com/mudler/LocalAI/core/application"
"github.com/mudler/LocalAI/core/config"
"github.com/mudler/LocalAI/pkg/system"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
var _ = Describe("live API traces", func() {
newApp := func(root string) *application.Application {
app, err := application.New(
config.EnableTracing,
config.WithDataPath(root),
config.WithDisableLocalAIAssistant(true),
config.WithDisableStats(true),
config.WithSystemState(&system.SystemState{
Model: system.Model{ModelsPath: root},
Backend: system.Backend{BackendsPath: root},
}),
)
Expect(err).NotTo(HaveOccurred())
DeferCleanup(func() { Expect(app.Shutdown()).To(Succeed()) })
ClearTraces()
return app
}
It("lists a request while its handler is still running", func() {
root := GinkgoT().TempDir()
app := newApp(root)
started := make(chan struct{})
release := make(chan struct{})
DeferCleanup(func() {
select {
case <-release:
default:
close(release)
}
})
handler := TraceMiddleware(app)(func(c echo.Context) error {
close(started)
<-release
return c.NoContent(http.StatusNoContent)
})
e := echo.New()
req := httptest.NewRequest(http.MethodPost, "/slow", http.NoBody)
req.Header.Set(echo.HeaderContentType, echo.MIMEApplicationJSON)
rec := httptest.NewRecorder()
ctx := e.NewContext(req, rec)
ctx.SetPath("/slow")
done := make(chan error, 1)
go func() {
done <- handler(ctx)
}()
<-started
var running APIExchange
Eventually(func() bool {
traces := GetTraces()
if len(traces) != 1 {
return false
}
running = traces[0]
return running.Request.Path == "/slow"
}).Should(BeTrue())
Expect(running.Response.Status).To(Equal(0))
Expect(running.Duration).To(BeNumerically(">", 0))
close(release)
Expect(<-done).To(Succeed())
Eventually(func() []APIExchange { return GetTraces() }).Should(ConsistOf(
And(
HaveField("ID", running.ID),
HaveField("Response.Status", http.StatusNoContent),
HaveField("Duration", BeNumerically(">", time.Duration(0))),
),
))
})
It("removes an in-flight trace when the handler panics", func() {
app := newApp(GinkgoT().TempDir())
handler := TraceMiddleware(app)(func(echo.Context) error {
panic("handler panic")
})
e := echo.New()
req := httptest.NewRequest(http.MethodPost, "/panic", http.NoBody)
req.Header.Set(echo.HeaderContentType, echo.MIMEApplicationJSON)
ctx := e.NewContext(req, httptest.NewRecorder())
ctx.SetPath("/panic")
func() {
defer func() { _ = recover() }()
_ = handler(ctx)
}()
Expect(GetTraces()).To(BeEmpty())
})
})

View File

@@ -0,0 +1,65 @@
import { test, expect } from './coverage-fixtures.js'
test('marks an API trace with no response status as in progress', async ({ page }) => {
await page.route('**/api/traces?*', route => route.fulfill({
json: [{
id: 'running-1',
timestamp: '2026-08-05T02:00:00Z',
duration: 2_000_000_000,
request: { method: 'POST', path: '/v1/chat/completions' },
response: { status: 0 },
}],
headers: { 'X-Total-Count': '1' },
}))
await page.route('**/api/backend-traces?*', route => route.fulfill({ json: [] }))
await page.goto('/app/traces')
const row = page.locator('tbody tr').filter({ hasText: '/v1/chat/completions' })
await expect(row.getByText('Running', { exact: true })).toBeVisible()
await expect(row.locator('[title="In progress"]')).toBeVisible()
await expect(row.locator('.fa-check-circle')).toHaveCount(0)
})
// Regression for #11376: switching from Backend Traces back to API Traces
// used to crash the page. `traces` holds whichever list was fetched last, so
// right after `setActiveTab('api')` — before the refetch effect lands — the
// API table renders the previous tab's backend rows, which carry no
// `response` envelope. The status column must tolerate that instead of
// dereferencing `trace.response.status` and tearing down the React tree.
test('switching from backend to API traces with a response-less row does not crash', async ({ page }) => {
const pageErrors = []
page.on('pageerror', (e) => pageErrors.push(e.message))
await page.route('**/api/traces?*', route => route.fulfill({
json: [{
id: 'api-1',
timestamp: '2026-08-05T02:00:00Z',
request: { method: 'POST', path: '/v1/chat/completions' },
response: { status: 200 },
}],
headers: { 'X-Total-Count': '1' },
}))
await page.route('**/api/backend-traces?*', route => route.fulfill({
json: [{
id: 'backend-1',
type: 'llm',
timestamp: '2026-08-05T02:00:00Z',
model_name: 'mock-model',
summary: 'generated a reply',
}],
headers: { 'X-Total-Count': '1' },
}))
await page.goto('/app/traces')
await expect(page.locator('tbody tr').filter({ hasText: '/v1/chat/completions' })).toBeVisible()
await page.getByRole('button', { name: /Backend Traces/ }).click()
await expect(page.locator('tbody tr').filter({ hasText: 'generated a reply' })).toBeVisible()
await page.getByRole('button', { name: /API Traces/ }).click()
// The stale backend row renders in the API table for one frame; the status
// column falls back to a neutral placeholder rather than throwing.
await expect(page.locator('tbody tr').filter({ hasText: '/v1/chat/completions' })).toBeVisible()
expect(pageErrors).toEqual([])
})

View File

@@ -664,10 +664,18 @@ export default function Traces() {
<td><span className="badge badge-info">{trace.request?.method || '-'}</span></td>
<td className="text-mono text-sm">{trace.request?.path || '-'}</td>
<td className="text-sub cell-clip" title={trace.user_name || trace.user_id || ''}>{trace.user_name || trace.user_id || '-'}</td>
<td><span className={`badge ${(trace.response?.status || 0) < 400 ? 'badge-success' : 'badge-error'}`}>{trace.response?.status || '-'}</span></td>
<td>
{trace.response?.status === 0
? <span className="badge badge-info">Running</span>
: trace.response?.status == null
? <span className="badge badge--soft">-</span>
: <span className={`badge ${trace.response.status < 400 ? 'badge-success' : 'badge-error'}`}>{trace.response.status}</span>}
</td>
<td><LatencyCell ns={trace.duration} max={slowestTrace} /></td>
<td className="text-center">
{trace.error
{trace.response?.status === 0
? <i className="fas fa-spinner fa-spin text-primary" title="In progress" />
: trace.error
? <i className="fas fa-times-circle text-error" title={trace.error} />
: <i className="fas fa-check-circle text-success" />}
</td>

View File

@@ -74,6 +74,9 @@ services:
GODEBUG: "netdns=go"
# Paths
MODELS_PATH: /models
# Avoid probing remote gallery GGUF metadata during container startup.
# Remove this line or set a positive limit to opt back into cache warming.
LOCALAI_VRAM_WARM_LIMIT: "0"
volumes:
- frontend_models:/models
- frontend_data:/data

View File

@@ -18,6 +18,9 @@ services:
- .env
environment:
- MODELS_PATH=/models
# Avoid probing remote gallery GGUF metadata during container startup.
# Remove this line or set a positive limit to opt back into cache warming.
- LOCALAI_VRAM_WARM_LIMIT=0
# - DEBUG=true
## Agents (LocalAGI) - https://localai.io/features/agents/
# - LOCALAI_DISABLE_AGENTS=false

View File

@@ -477,6 +477,11 @@ then on.
| `LOCALAI_VRAM_WARM_LIMIT` | `300` | How many gallery entries to warm at startup, estimates and variants alike. Set to `0` to disable the warm-up entirely. |
| `LOCALAI_VRAM_WARM_CONCURRENCY` | `4` | How many estimates to run at once. |
The provided Docker Compose configurations set `LOCALAI_VRAM_WARM_LIMIT=0`
as a defensive default, so container startup does not probe remote GGUF files.
Remove that override or set it to a positive number to opt into background
warming.
```bash
# Air-gapped, or you would rather not make the requests at all
LOCALAI_VRAM_WARM_LIMIT=0 local-ai run

View File

@@ -9,10 +9,11 @@ url = "/features/audio-diarization/"
Speaker diarization answers the question **"who spoke when?"** - given an audio clip with multiple speakers, it returns time-stamped segments labelled with a stable speaker ID (`SPEAKER_00`, `SPEAKER_01`, …).
LocalAI exposes this through the `/v1/audio/diarization` endpoint, modelled after `/v1/audio/transcriptions`. Three backends are supported today:
LocalAI exposes this through the `/v1/audio/diarization` endpoint, modelled after `/v1/audio/transcriptions`. Four backends are supported today:
- **[sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx)** - pyannote-3.0 segmentation + a speaker-embedding extractor (3D-Speaker, NeMo, WeSpeaker) + fast clustering. Pure diarization - no transcription cost. Recommended when you only need speaker turns.
- **[vibevoice.cpp](https://github.com/microsoft/VibeVoice)** - produces speaker-labelled segments as a by-product of its long-form ASR pass, so you can optionally get a transcript per segment for free.
- **[NeMo-Speech.cpp](https://github.com/NVIDIA/NeMo-Speech.cpp)** - NVIDIA Sortformer, served standalone by the [NeMo-Speech.cpp backend]({{%relref "features/nemo-speech-cpp" %}}). It is end to end, so the speaker capacity is fixed by the checkpoint and the count hints are ignored. The same backend can instead put speaker tags on a transcript, by attaching a Sortformer model to an ASR one.
- **[audio.cpp](https://github.com/0xShug0/audio.cpp)** - the `sortformer_diar` family, served by the multi-modality [audio.cpp backend]({{%relref "features/audio-cpp" %}}).
Because diarization is exposed as a regular OpenAI-compatible endpoint, any HTTP client works. There is no Python dependency on pyannote or NeMo on the consumer side.

View File

@@ -14,6 +14,7 @@ The transcription endpoint allows to convert audio files to text. The endpoint s
- **[parakeet-cpp](https://github.com/mudler/parakeet.cpp)**: A C++/ggml port of NVIDIA NeMo Parakeet (FastConformer TDT/CTC/RNNT/hybrid). Runs quantized GGUFs on CPU or GPU, emits word-level timestamps, and supports cache-aware streaming (the `realtime_eou` model surfaces end-of-utterance events).
- **llama-cpp**: Route transcription to any multimodal-audio GGUF model served by the `llama-cpp` backend (e.g. [Qwen3-ASR](https://huggingface.co/ggml-org/Qwen3-ASR-0.6B-GGUF), Voxtral, Qwen2-Audio). Under the hood the request is converted into a chat completion with the audio attached via the model's audio encoder - the same path the upstream llama.cpp server uses. Set `backend: llama-cpp` in the model YAML and point `mmproj` at the matching audio encoder.
- **voxtral**: Voxtral-family models served by a dedicated backend
- **[NeMo-Speech.cpp](https://github.com/NVIDIA/NeMo-Speech.cpp)**: NVIDIA's C++/ggml runtime for the Nemotron Speech models. Serves offline, streaming and live transcription, with VAD, punctuation, inverse text normalization and Sortformer speaker tags attached through model options, and covers diarization, speech synthesis and translation from the same backend. See the [NeMo-Speech.cpp backend]({{%relref "features/nemo-speech-cpp" %}}) page for the model options.
- **[audio.cpp](https://github.com/0xShug0/audio.cpp)**: Multi-family GGML audio engine. Serves transcription and forced alignment from families such as `nemotron_asr`, `qwen3_asr`, `citrinet_asr`, `higgs_audio_stt` and `voxtral_realtime`, and covers diarization, VAD, TTS and source separation from the same backend. See the [audio.cpp backend]({{%relref "features/audio-cpp" %}}) page for the model options.
The endpoint input supports all the audio formats supported by `ffmpeg`.

View File

@@ -72,6 +72,44 @@ tags:
- "text-generation"
```
### Verifying OCI Backends
Backend galleries can require keyless Sigstore signatures for every OCI image
they provide. Add a `verification` policy to the gallery configuration, then
enable strict integrity mode:
```bash
export LOCALAI_BACKEND_GALLERIES='[{"name":"localai","url":"github:mudler/LocalAI/backend/index.yaml@master","verification":{"issuer":"https://token.actions.githubusercontent.com","identity_regex":"^https://github\\.com/mudler/LocalAI/\\.github/workflows/backend_merge\\.yml@refs/(heads/master|tags/.+)$"}}]'
export LOCALAI_REQUIRE_BACKEND_INTEGRITY=1
local-ai run
```
The policy pins the Fulcio issuer and the GitHub Actions workflow identity that
signed the image. The identity expression covers development images produced
from `master` and release images produced from tags. Use a narrower expression
if your deployment only accepts one release channel.
Without strict mode, an OCI gallery without a verification policy installs
with a warning. With strict mode, LocalAI refuses galleries without a policy,
images without a compatible Sigstore bundle, and signatures that do not match
the configured identity. Existing images published before bundle signing was
enabled must be rebuilt or re-signed before strict deployments can install
them.
An optional `not_before` RFC3339 value revokes signatures logged before that
time. Advance it after a signing-workflow compromise, then rebuild or re-sign
the trusted images:
```json
{
"verification": {
"issuer": "https://token.actions.githubusercontent.com",
"identity_regex": "^https://github\\.com/mudler/LocalAI/\\.github/workflows/backend_merge\\.yml@refs/(heads/master|tags/.+)$",
"not_before": "2026-08-05T00:00:00Z"
}
}
```
## Pre-installing Backends
You can pre-install backends when starting LocalAI using the `LOCALAI_EXTERNAL_BACKENDS` environment variable:
@@ -130,8 +168,8 @@ For getting started, see the available backends in LocalAI here: https://github.
LocalAI supports various types of backends:
- **LLM Backends**: For running language models (e.g., llama.cpp, vLLM, vllm.cpp, SGLang, transformers, MLX)
- **Speech-to-Text Backends**: For transcription, forced alignment and speaker diarization (e.g., whisper.cpp, parakeet.cpp, moss-transcribe.cpp, faster-whisper, NeMo, [audio.cpp]({{%relref "features/audio-cpp" %}}))
- **Text-to-Speech Backends**: For speech synthesis (e.g., piper, Kokoro, VibeVoice, Qwen3-TTS, [audio.cpp]({{%relref "features/audio-cpp" %}}))
- **Speech-to-Text Backends**: For transcription, forced alignment and speaker diarization (e.g., whisper.cpp, parakeet.cpp, moss-transcribe.cpp, [NeMo-Speech.cpp]({{%relref "features/nemo-speech-cpp" %}}), faster-whisper, NeMo, [audio.cpp]({{%relref "features/audio-cpp" %}}))
- **Text-to-Speech Backends**: For speech synthesis (e.g., piper, Kokoro, VibeVoice, Qwen3-TTS, [NeMo-Speech.cpp]({{%relref "features/nemo-speech-cpp" %}}), [audio.cpp]({{%relref "features/audio-cpp" %}}))
- **Sound Generation Backends**: For music and audio generation (e.g., ACE-Step, [audio.cpp]({{%relref "features/audio-cpp" %}}))
- **Sound Classification Backends**: For sound-event classification / audio tagging - identifying everyday sounds like baby cry, glass breaking, alarms (e.g., ced.cpp)
- **Image & Video Generation Backends**: For diffusion and audio-conditioned avatar models (e.g., stable-diffusion.cpp, diffusers, vLLM-Omni, [LongCat-Video]({{%relref "features/video-generation" %}}))

View File

@@ -0,0 +1,335 @@
+++
disableToc = false
title = "NeMo-Speech.cpp backend"
weight = 39
url = "/features/nemo-speech-cpp/"
+++
[NeMo-Speech.cpp](https://github.com/NVIDIA/NeMo-Speech.cpp) is NVIDIA's Apache-2.0
C++/ggml runtime for the Nemotron Speech models. LocalAI exposes it through the native
`nemo-speech-cpp` backend, which serves four model families from one installed backend:
transcription, speaker diarization, speech synthesis and text translation.
NeMo-Speech.cpp is developed by NVIDIA, not by the LocalAI project.
## Installing
```bash
local-ai backends install nemo-speech-cpp
```
Or install it from the **Backends** page in the web UI. `nemo-speech-cpp` is a
preference-only backend: LocalAI never picks it automatically during model import,
because the family is decided by the GGUF's `general.architecture` key, which cannot be
read from a remote repository, and because a translation model carries an ordinary LLM
architecture with no NeMo-specific marker at all. Set `backend: nemo-speech-cpp` in the
model YAML, or select it explicitly in the import form.
## How the family is chosen
The backend reads `general.architecture` from the GGUF at load time and picks the family
from it. Nothing else in the config selects it.
| `general.architecture` | Family | Serves |
|---|---|---|
| `asr` | Transcription | `/v1/audio/transcriptions`, the same endpoint with `stream=true`, and the realtime live-transcription path |
| `sortformer` | Diarization | `/v1/audio/diarization` |
| `magpietts` | Text to speech | `/v1/audio/speech`, `/tts`, including streamed audio |
| `nemo-nano-codec`, `vad`, `pnc` | *(none)* | Auxiliary assets. Loading one directly is refused. |
| anything else | Translation | `/v1/chat/completions` and `/v1/completions` |
Two rows need explaining.
**The auxiliary architectures** are converted NeMo components that attach to a primary
model and cannot run on their own. Pointing `parameters.model` at one fails the load with
a message naming the option it belongs on: the codec belongs on `codec_model` of a
`magpietts` model, the VAD on `vad_model` of an `asr` model, the punctuation model on
`pnc_model` of an `asr` model.
**Everything else is translation**, and that is deliberate rather than a fallback that
happens to catch it. Riva-Translate GGUFs are produced by llama.cpp's converter and carry
an ordinary LLM architecture such as `qwen3`, so there is no NeMo-specific string to match
on. Selecting this backend explicitly is the signal that the model is meant for it.
A request for the wrong family is refused with `UNIMPLEMENTED` naming the family the model
was loaded as, rather than failing somewhere inside the runtime.
## Model YAML
Options are `key:value` entries in the `options:` list, split on the **first** colon so a
value may contain more. Every path option is resolved relative to the models directory
when it is not absolute. An unknown key is ignored rather than rejected, so a config
written for a newer backend still loads on an older one.
### Transcription
```yaml
name: nemotron-asr
backend: nemo-speech-cpp
parameters:
model: asr.gguf
known_usecases:
- FLAG_TRANSCRIPT
options:
# All optional. Each attaches a converted NeMo component to the recognizer.
- vad_model:vad.gguf
- pnc_model:pnc-bert-base-en.q8_0.gguf
- itn_dir:sparrowhawk_grammars
- diar_model:diarization.gguf
- language_code:en-US
# GPU device index. Omit it, or set -1, to run on the CPU.
- gpu:0
```
Attaching `diar_model` is what turns on per-word speaker tags: the backend asks for
diarization only when the recognizer was created with a diarization model, because the
runtime refuses a request for it otherwise. Transcript segments are then cut at each
change of speaker. Without it there is a single unlabelled segment. Word-level timings are
returned only when the request asks for them with
`timestamp_granularities[]=word`, following the OpenAI contract.
`language_code` is the model-level default. A per-request `language` wins over it, and
both may be left empty, which the runtime reads as the model's own default.
### Diarization
A `sortformer` model is standalone: this pipeline has no ASR in it, so segments carry
speaker labels and timings but no text.
```yaml
name: sortformer-diarization
backend: nemo-speech-cpp
parameters:
model: diarization.gguf
known_usecases:
- FLAG_DIARIZATION
options:
- gpu:0
```
Sortformer is end to end and its speaker capacity is fixed by the checkpoint, so
`num_speakers`, `min_speakers`, `max_speakers` and `clustering_threshold` have nothing to
map onto and are logged and dropped. `include_text` is dropped for the same reason: there
is no ASR in this pipeline. `min_duration_on` and `min_duration_off` are honoured. To get
speaker labels *on a transcript*, use an `asr` model with `diar_model` instead.
### Text to speech
```yaml
name: magpie-tts
backend: nemo-speech-cpp
parameters:
model: magpie-tts/magpietts.gguf
known_usecases:
- FLAG_TTS
options:
# Both are required, and both are auto-discovered when unset (see below).
- codec_model:magpie-tts/nanocodec.gguf
- tokenizer_dir:magpie-tts/extracted
# Optional Sparrowhawk text-normalization grammars applied to the input text.
- tn_dir:tts_grammars
- language_code:en-US
- gpu:0
```
`codec_model` and `tokenizer_dir` are both required, and both are discovered from the
model's own directory when they are not set: a sibling file whose name contains
`nanocodec` or `nano-codec` becomes the codec, and a sibling directory named `extracted`
becomes the tokenizer directory. If discovery finds nothing, the load fails naming the
option to set. That is a hard error rather than a warning because the alternative is a
synthesizer that loads and emits noise.
Voices are selected by the request's `voice` field. A non-negative integer is used as a
speaker index; anything else is passed through as a voice name. MagpieTTS conditions on a
speaker, not on a prose style, so `instructions` has nothing to map onto and is logged and
ignored. Per-request `params` are read for `seed`, `steps`, `top_k`, `temperature` and
`cfg_scale`; anything else is left at the synthesizer's configured value.
The TTS runtime has a three-way CPU/CUDA/auto preference rather than a device index, so
`gpu` with a non-negative value means "let the runtime choose" here, and any negative
value (`-1` is the default) pins it to the CPU.
### Translation
```yaml
name: riva-translate
backend: nemo-speech-cpp
parameters:
model: translate.q8_0.gguf
known_usecases:
- FLAG_COMPLETION
- FLAG_CHAT
options:
- source_language:en
- target_language:de
- gpu:0
```
Both flags run through the same pair of RPCs, `Predict` and `PredictStream`, which is why
both are listed for this backend in `core/config/backend_capabilities.go`. `FLAG_COMPLETION`
is the plain shape: prompt in, translation out. `FLAG_CHAT` adds three things, and each
turn of the conversation is translated on its own: the model becomes eligible as the
default chat model when a request names none, it appears in the web UI's chat model
picker, and it stays visible under the gallery's Chat filter (`completion` is not a
gallery filter, `chat` is). Neither flag gates a request that names the model explicitly.
### Option reference
| Option | Family | Default | Meaning |
|---|---|---|---|
| `vad_model:<path>` | ASR | unset | Converted Silero VAD GGUF, attached to the recognizer. |
| `pnc_model:<path>` | ASR | unset | Punctuation and capitalization BERT GGUF. |
| `diar_model:<path>` | ASR | unset | Sortformer GGUF. Setting it is the opt-in that enables per-word speaker tags. |
| `itn_dir:<path>` | ASR | unset | Sparrowhawk grammar directory for inverse text normalization, or a parent directory whose children are named per language (`en`, `es`, …). Linux only, see the limitations below. |
| `language_code:<code>` | ASR, TTS | unset | Model-level default language. A per-request `language` overrides it. |
| `codec_model:<path>` | TTS | auto-discovered | NanoCodec GGUF. Required. |
| `tokenizer_dir:<path>` | TTS | auto-discovered | Directory holding the extracted MagpieTTS tokenizer assets. Required. |
| `tn_dir:<path>` | TTS | unset | Sparrowhawk text-normalization grammars applied to the input text. Linux only, see the limitations below. |
| `source_language:<code>` | Translation | unset | Default source language. May be overridden per request. |
| `target_language:<code>` | Translation | unset | Default target language. A request with neither this nor a per-request override is refused. |
| `gpu:<n>` | all | `-1` (CPU) | Device index. `-1` selects the CPU. A value that is not an integer is ignored with a warning and the model runs on the CPU. |
## Translating
Translation is reached through the text endpoints: the prompt is the text to translate,
and the languages come from `source_language` and `target_language`.
```bash
curl http://localhost:8080/v1/chat/completions -H "Content-Type: application/json" -d '{
"model": "riva-translate",
"messages": [{"role": "user", "content": "The quick brown fox jumps over the lazy dog."}]
}'
```
A leading `[src->tgt]` prefix on the prompt overrides the configured pair for that one
request:
```json
{"role": "user", "content": "[en->zh-cn] The quick brown fox jumps over the lazy dog."}
```
Either side may be left out to keep the model-level default for it, so `[->de]` changes
only the target. Three-segment pair tags such as `en-zh-cn`, `en-pt-br` and `zh-tw-en`
parse correctly. The prefix is stripped before the text reaches the model.
The C API takes a source and a target language and has no free-form generation entry
point, so there is no prompt in the LLM sense. Sampling parameters, tools, grammars and
attached media have no equivalent and are ignored; the structural ones (`grammar`,
`tools`, `images`, `videos`, `audios`, `negative_prompt`, `logprobs`) are named in the log
when a request sets them. Streaming works, but the runtime returns the finished
translation in one piece, so the whole result arrives as a single chunk rather than token
by token.
## Streaming transcription
Streaming transcription and the realtime live path both emit a delta per **finalized**
utterance. Interim hypotheses are produced by the runtime and deliberately dropped: the
wire contract defines `delta` as newly finalized text that consumers concatenate, and the
runtime rewrites a final rather than extending its interims (it runs ITN and formatting
stripping on the final only). Forwarding interims would assemble "he", "hell", "hello",
"Hello." into `hehellhelloHello.` rather than into the transcript, and no diffing trick
recovers it.
The cost is latency: the first delta of an utterance arrives at its endpoint rather than
mid-word. That is the trade this backend takes.
## Acceleration
| Variant | Platform |
|---|---|
| CPU | linux/amd64, linux/arm64 |
| CUDA 12 | linux/amd64 |
| CUDA 13 | linux/amd64 |
| Vulkan | linux/amd64, linux/arm64 |
| NVIDIA Jetson (L4T), CUDA 12 | linux/arm64 |
| NVIDIA Jetson (L4T), CUDA 13 | linux/arm64 |
| Metal | darwin/arm64 |
There is **no AMD ROCm and no Intel SYCL** support, because upstream NeMo-Speech.cpp has
no HIP and no SYCL backend, so there is nothing to build against. A host reporting either
capability gets the CPU build, which is the honest answer rather than a broken image. If
you want NeMo ASR on an AMD or Intel GPU, use
[parakeet-cpp](https://github.com/mudler/parakeet.cpp) instead, which is described on the
[Audio to text]({{%relref "features/audio-to-text" %}}) page.
## Limitations
- **Text normalization is Linux only.** The macOS build ships without it: the
Sparrowhawk/OpenFST stack assumes a GNU toolchain, and the gcc-12 pin it needs (OpenFST's
templates fail to compile on gcc-13 and gcc-14 at `-O2`) has no macOS equivalent. One
build flag covers both directions, so on macOS **both** `itn_dir` (inverse text
normalization on ASR output) and `tn_dir` (text normalization on TTS input) do nothing.
A `tn_dir` set on a macOS build logs a warning and carries on with normalization
disabled. `pnc_model` is unaffected: punctuation is always compiled in.
- **Interim streaming results are suppressed**, as described above. This costs latency.
- **Translation runs at the library's default limits**: 1024 tokens of context and 256
new tokens per call. Neither is configurable from the model YAML, because both are
create-time settings on the translator and raising the context costs one context-sized
KV cache per pooled context. The two limits fail differently. Input longer than the
context is **rejected**, with `nmt: prompt too long (N tokens) for context 1024`
surfacing as a failed request, so you will know. Output longer than 256 tokens is
**silently cut**: generation simply stops at the limit and the truncated translation is
returned as if it were complete. Translate a sentence or a paragraph at a time rather
than a whole document.
- **There are no gallery entries yet.** Models have to be converted with upstream's
converter and configured by hand, as below. This is a follow-up, not an oversight.
## Converting models
Upstream has a single conversion entry point for every family. `SOURCE` may be a `.nemo`
archive, an extracted NeMo checkpoint, a local Hugging Face directory, or a Hugging Face
repository ID.
```bash
git clone https://github.com/NVIDIA/NeMo-Speech.cpp
cd NeMo-Speech.cpp
pip install -r requirements.txt
python3 convert_model.py nvidia/nemotron-speech-streaming-en-0.6b \
--outfile /models/asr.gguf
```
Copy the resulting GGUF into your models directory and write the YAML above against it.
Text to speech is the one family that needs more than a single conversion. It loads **two**
GGUFs, the MagpieTTS token generator and the NanoCodec decoder, and it needs the tokenizer
assets that live inside the MagpieTTS `.nemo` archive rather than in the GGUF. Extract the
archive and keep the extracted directory next to the converted model, so that
`tokenizer_dir` (or the `extracted/` auto-discovery) has something to point at, and put the
codec GGUF in the same directory so `codec_model` (or its auto-discovery) finds it.
```bash
# 1. MagpieTTS: download the .nemo, extract it for the tokenizer, convert it
hf download nvidia/magpie_tts_multilingual_357m --revision v2602 \
--local-dir /models/magpie-tts
mkdir -p /models/magpie-tts/extracted
tar -xf /models/magpie-tts/magpie_tts_multilingual_357m.nemo \
-C /models/magpie-tts/extracted
python3 convert_model.py /models/magpie-tts/extracted \
--outfile /models/magpie-tts/magpietts.gguf
# 2. NanoCodec: no tokenizer, it is a codec decoder. The filename carries
# "nanocodec" so the auto-discovery in the YAML above picks it up.
hf download nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps \
--local-dir /models/magpie-tts/nano-codec
python3 convert_model.py \
/models/magpie-tts/nano-codec/nemo-nano-codec-22khz-1.89kbps-21.5fps.nemo \
--outfile /models/magpie-tts/nanocodec.gguf
```
Skipping the codec leaves the model unloadable: the backend fails with
`no NanoCodec GGUF found next to ...` naming the `codec_model` option.
Translation additionally uses the pinned llama.cpp converter, which has to be initialized
first:
```bash
git submodule update --init llama.cpp
pip install -r llama.cpp/requirements/requirements-convert_hf_to_gguf.txt
python3 convert_model.py nvidia/Riva-Translate-4B-Instruct-v2 \
--outfile /models/translate.q8_0.gguf --outtype q8_0
```
See upstream's [model conversion
guide](https://github.com/NVIDIA/NeMo-Speech.cpp/blob/main/docs/model-conversion.md) for
the per-architecture defaults and options.

View File

@@ -918,6 +918,200 @@ options:
The full list of registered parsers lives in `sglang.srt.function_call`
and `sglang.srt.parser.reasoning_parser`.
### vllm.cpp
[vllm.cpp](https://github.com/mudler/vllm.cpp) is the LocalAI team's C++ port of
vLLM: the same continuous-batching scheduler, paged KV cache and prefix caching,
with no Python at inference time. It consumes either a HuggingFace safetensors
model directory or a `.gguf` file, and applies the model's chat template,
tool-call parsing and reasoning split engine-side.
#### Setup
```yaml
name: vllm-cpp
backend: vllm-cpp
parameters:
model: "Qwen/Qwen3-4B"
context_size: 8192
template:
use_tokenizer_template: true
```
#### Configuring the engine with `engine_args`
The same `engine_args:` map the vLLM and SGLang backends accept is honoured
here, with keys spelled exactly as vLLM's own CLI flags - so a `speculative_config`
or `kv_transfer_config` block written for vLLM works verbatim. Unknown keys are
ignored rather than fatal; the engine validates the documents it is handed and
reports a precise error at load.
```yaml
name: qwen35-a3b
backend: vllm-cpp
parameters:
model: "Qwen/Qwen3.5-A3B"
context_size: 16384
template:
use_tokenizer_template: true
engine_args:
# KV cache sizing: num_blocks * block_size tokens of cache.
block_size: 32
num_blocks: 1024
# Concurrency and the per-step chunked-prefill token budget.
max_num_seqs: 32
max_num_batched_tokens: 8192
# Automatic prefix caching. Omit to keep the model's own default
# (on for dense models, off for hybrid / attention-free ones).
enable_prefix_caching: true
# Scheduler admission order: fcfs (default), priority, or lpm
# (cache-aware longest-prefix-match; needs prefix caching to have any effect).
scheduling_policy: lpm
```
| Key | Meaning | Default |
|-----|---------|---------|
| `block_size` | KV-cache block size, in tokens per block | 32 |
| `num_blocks` | KV-cache blocks to allocate | 256 |
| `max_model_len` | Max sequence length; also settable as `context_size` / `max_model_len` | model config |
| `max_num_seqs` | Max concurrent sequences the scheduler admits | 8 |
| `max_num_batched_tokens` | Per-step chunked-prefill token budget | per-arch (2048 dense, 4096/8192 MoE) |
| `enable_prefix_caching` | Automatic prefix caching; `enable_radix_attention` is an accepted alias | model default |
| `enable_jump_forward` | Jump-forward decoding, which emits grammar-forced tokens without a model step. Only affects constrained requests (`grammar`, JSON schema) | off |
| `scheduling_policy` | `fcfs`, `priority`, or `lpm` | `fcfs` |
| `tool_parser` / `reasoning_parser` | Force a parser instead of chat-template auto-detection | auto |
| `tokenizer_config` | Override the `tokenizer_config.json` the chat template is read from | `<model_dir>/tokenizer_config.json` |
| `speculative_config` | Speculative decoding (see below) | disabled |
| `kv_transfer_config` | External KV connector / LMCache (see below) | none |
Raising `max_num_batched_tokens` lets more prefill land in a single step, at the
cost of decode latency for requests queued behind it. The default deliberately
does not scale with `max_num_seqs`, which is what keeps a large concurrent
prefill from blowing up the per-step activation on the hybrid architectures.
`enable_prefix_caching` and `enable_jump_forward` are tri-state at the engine
boundary: omitting the key defers to a default (the model's own capability for
prefix caching, an environment variable for jump forward), while an explicit
`false` forces the feature off. Those are genuinely different - prefix caching
defaults *on* for dense models - so write the key only when you mean to override.
#### Speculative decoding
`speculative_config:` takes the same JSON object as vLLM's
`--speculative-config`. Three methods are supported.
> **Architecture limit.** At the current engine pin, `mtp` and `dflash` are
> **Qwen3.5 / Qwen3.6 only**. The engine builds a widened speculative KV cache
> directly for those families rather than through the model registry, so a
> speculative config on any other architecture (Llama, GLM, Gemma, Mistral, ...)
> will not work regardless of checkpoint format. `ngram` needs no draft weights
> and is not subject to this limit.
> **Format support.** `mtp` and `dflash` now work from a `.gguf` target as well
> as safetensors. An MTP head is read from the GGUF's `nextn.*` tensors when the
> file declares `<arch>.nextn_predict_layers`; a GGUF exported WITHOUT the head
> (converted with `--no-mtp`, or predating llama.cpp's Qwen3.5 MTP support) is
> refused at load naming that as the reason. A DFlash draft may itself be a
> `dflash`-arch GGUF, and the target may be a GGUF too. `ngram` needs no draft
> weights and works on any format.
**MTP** (Multi-Token Prediction) uses a draft head shipped inside the target
checkpoint's own `mtp.*` tensors, so there is no second model to download. It
requires a **safetensors** checkpoint - the `mtp.*` tensors do not survive GGUF
conversion, and an MTP config over a `.gguf` model is rejected at load.
```yaml
engine_args:
speculative_config:
method: mtp
# Optional; defaults to the checkpoint's own head depth, which is
# usually the right value. Must be a multiple of that depth.
num_speculative_tokens: 1
```
**DFlash** uses a separate block-diffusion drafter that proposes a whole block
of tokens in one non-autoregressive forward pass. Unlike MTP, the draft is its
own checkpoint, so `model:` is **required**:
```yaml
engine_args:
speculative_config:
method: dflash
model: z-lab/Qwen3.6-27B-DFlash
num_speculative_tokens: 4
```
The draft shares the *target's* `embed_tokens` and `lm_head`, so both must come
from the same model family and the target must be safetensors.
**The engine does not download the draft.** `model:` is resolved, in order,
as a path as given, then as the last path segment under LocalAI's models
directory (`z-lab/Qwen3.6-27B-DFlash``<models>/Qwen3.6-27B-DFlash`, which is
what LocalAI's own downloader produces), then as the whole reference under the
models directory. Install the draft into LocalAI first, or give an absolute path
to a directory containing `config.json`. If none of those resolve, the load
fails immediately naming every location that was tried, rather than reporting a
missing checkpoint from inside the engine.
**N-gram** needs no draft model at all - it proposes from the prompt's own
suffix history. `num_speculative_tokens` is required:
```yaml
engine_args:
speculative_config:
method: ngram
num_speculative_tokens: 4
prompt_lookup_min: 5
prompt_lookup_max: 5
```
> **Auto-configuration on import.** When you import a safetensors repository
> with `backend: vllm-cpp`, LocalAI reads the checkpoint's `config.json` and, if
> it declares an MTP head (`mtp_num_hidden_layers`), writes
> `speculative_config: {method: mtp}` into the generated `engine_args` for you.
> An explicit `speculative_config` in your own config is never overwritten.
> Importing a DFlash *draft* repository is refused with a warning: a drafter
> cannot serve on its own, so import the target model and point
> `speculative_config.model` at the draft.
#### External KV cache with LMCache
`kv_transfer_config:` takes vLLM's `--kv-transfer-config` JSON and selects an
external KV-cache connector. The `lm://` LMCache client lets prefill KV be
stored to and reloaded from a shared `lmcache.v1.server`, so a prefix computed
by one replica does not have to be recomputed by the next:
```yaml
engine_args:
kv_transfer_config:
kv_connector: LMCacheConnector
kv_role: kv_both # required whenever kv_connector is set
kv_connector_extra_config:
host: 127.0.0.1
port: 65432
```
`kv_role` is one of `kv_producer` (store only), `kv_consumer` (load only), or
`kv_both`. An unregistered connector name, a missing role, or a malformed
document fails the load with an explicit error rather than silently running
without the cache.
#### Legacy `options:` list
Earlier versions configured this backend through the flat `options:` list, and
those configs keep working. Every key in the table above is still read from
there in `key:value` form, and `engine_args` wins on any key set in both:
```yaml
options:
- max_num_seqs:32
- enable_prefix_caching:true
```
New configs should prefer `engine_args:`, which is the only place the nested
`speculative_config` / `kv_transfer_config` documents can be written naturally
rather than as a single-line JSON string.
### Transformers
[Transformers](https://huggingface.co/docs/transformers/index) is a State-of-the-art Machine Learning library for PyTorch, TensorFlow, and JAX.

View File

@@ -774,6 +774,33 @@ clip. See the
[audio.cpp backend]({{%relref "features/audio-cpp" %}}) page for the full option list,
including the `load.` and `session.` namespaces and the supertonic packaging caveat.
### NeMo-Speech.cpp (MagpieTTS)
[NeMo-Speech.cpp](https://github.com/NVIDIA/NeMo-Speech.cpp) is NVIDIA's C++/ggml runtime
for the Nemotron Speech models. Its `magpietts` family synthesizes speech through a
NanoCodec decoder, and the same installed backend also covers transcription, diarization
and translation.
```yaml
name: magpie-tts
backend: nemo-speech-cpp
parameters:
model: magpie-tts/magpietts.gguf
known_usecases:
- FLAG_TTS
options:
- codec_model:magpie-tts/nanocodec.gguf
- tokenizer_dir:magpie-tts/extracted
- gpu:0
```
The codec and the tokenizer directory are both required. Both are discovered from the
model's own directory when left unset: a sibling file whose name contains `nanocodec` and
a sibling directory named `extracted`. MagpieTTS conditions on a speaker rather than on a
prose style, so `voice` selects a speaker index or a voice name and `instructions` has no
equivalent. See the [NeMo-Speech.cpp backend]({{%relref "features/nemo-speech-cpp" %}})
page for the full option list and the conversion steps.
## Response format
To provide some compatibility with OpenAI API regarding `response_format`, ffmpeg must be installed (or a docker image including ffmpeg used) to leverage converting the generated wav file before the api provide its response.

View File

@@ -9,6 +9,11 @@ LocalAI can retain recent API exchanges and backend operations for inspection
on the **Traces** page in the management interface. Enable tracing in runtime
settings or with the existing tracing configuration.
API requests appear while they are still running. Their elapsed duration
updates when the page refreshes, and the result column marks them as in
progress until the response completes. In-flight requests live only in memory;
the completed exchange is what LocalAI adds to the bounded, persistent history.
API and backend trace histories are persisted in separate directories below
the configured data path. They are restored after a clean service restart,
whether or not authentication is enabled.

View File

@@ -49,6 +49,7 @@ All backends listed here can be installed on demand from the [Backend Gallery]({
| [NeMo](https://github.com/NVIDIA/NeMo) | NVIDIA NeMo ASR toolkit | CPU, CUDA 12/13, ROCm, Intel SYCL, Metal |
| [sherpa-onnx](https://k2-fsa.github.io/sherpa/onnx/) | Sherpa-ONNX ASR (Whisper, Paraformer, SenseVoice) and TTS | CPU, CUDA 12, Metal |
| [audio.cpp](https://github.com/0xShug0/audio.cpp) | Multi-family GGML audio engine: transcription, forced alignment and speaker diarization (`nemotron_asr`, `qwen3_asr`, `voxtral_realtime`, `sortformer_diar` and more). See [audio.cpp backend](/features/audio-cpp/) | CPU, CUDA 12/13, Metal |
| [NeMo-Speech.cpp](https://github.com/NVIDIA/NeMo-Speech.cpp) | NVIDIA C++/GGML runtime for Nemotron Speech: offline, streaming and live transcription, Sortformer diarization, MagpieTTS synthesis and Riva-Translate translation from one backend. See [NeMo-Speech.cpp backend](/features/nemo-speech-cpp/) | CPU, CUDA 12/13, Vulkan, Metal, Jetson L4T |
## Text-to-Speech
@@ -77,6 +78,7 @@ All backends listed here can be installed on demand from the [Backend Gallery]({
| [MLX-Audio](https://github.com/Blaizzy/mlx-audio) | Audio models on Apple Silicon | CPU, CUDA 12/13, Metal, Jetson L4T |
| [liquid-audio](https://github.com/Liquid4All/liquid-audio) | LFM2 end-to-end speech-to-speech, ASR, and TTS | CPU, CUDA 12/13, ROCm, Intel SYCL, Jetson L4T |
| [audio.cpp](https://github.com/0xShug0/audio.cpp) | Multi-family GGML audio engine: TTS, voice cloning and voice design (`supertonic`, `vibevoice`, `qwen3_tts`, `chatterbox` and more). See [audio.cpp backend](/features/audio-cpp/) | CPU, CUDA 12/13, Metal |
| [NeMo-Speech.cpp](https://github.com/NVIDIA/NeMo-Speech.cpp) | NVIDIA MagpieTTS over NanoCodec, served by the same backend as Nemotron ASR. See [NeMo-Speech.cpp backend](/features/nemo-speech-cpp/) | CPU, CUDA 12/13, Vulkan, Metal, Jetson L4T |
## Music & Sound Generation

View File

@@ -1,3 +1,3 @@
{
"version": "v4.7.1"
"version": "v4.8.0"
}

View File

@@ -1,4 +1,101 @@
---
- &qwen3-5-9b-defiant-fable
name: "qwen3.5-9b-defiant-fable-mtp"
variants:
- model: qwen3.5-9b-defiant-fable
url: "github:mudler/LocalAI/gallery/virtual.yaml@master"
urls:
- https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF
description: |
Qwen3.5 9B Defiant Fable is an Apache-2.0 multimodal fine-tune for
reasoning, coding, creative writing, and roleplay. It retains the 256K
context window and vision support of Qwen3.5 while reducing refusals.
This default entry uses the NEO-imatrix Q4_K_M build with multi-token
prediction enabled for faster generation.
license: apache-2.0
icon: https://huggingface.co/DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF/resolve/main/defiant-fable-9b.png
tags:
- llm
- gguf
- cpu
- gpu
- qwen3.5
- reasoning
- coding
- creative-writing
- uncensored
- vision
- multimodal
- mtp
last_checked: "2026-08-04"
overrides:
backend: llama-cpp
function:
automatic_tool_parsing_fallback: true
grammar:
disable: true
known_usecases:
- chat
- vision
mmproj: llama-cpp/mmproj/qwen3.5-9b-defiant-fable/mmproj-BF16.gguf
options:
- use_jinja:true
- spec_type:draft-mtp
- spec_n_max:6
- spec_p_min:0.75
parameters:
model: llama-cpp/models/qwen3.5-9b-defiant-fable/Qwen3.5-9B-The-Defiant-Fable-Uncnr-Heretic-NEO-MAX-MTP-Q4_K_M.gguf
template:
use_tokenizer_template: true
files:
- filename: llama-cpp/models/qwen3.5-9b-defiant-fable/Qwen3.5-9B-The-Defiant-Fable-Uncnr-Heretic-NEO-MAX-MTP-Q4_K_M.gguf
uri: huggingface://DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF/Qwen3.5-9B-The-Defiant-Fable-Uncnr-Heretic-NEO-MAX-MTP-Q4_K_M.gguf
sha256: d7eb4fac9389d53fa576f64a6ff53e914a00bc7705dc354d1065887565147320
- filename: llama-cpp/mmproj/qwen3.5-9b-defiant-fable/mmproj-BF16.gguf
uri: huggingface://DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF/mmproj-BF16.gguf
sha256: 853698ce7aa6c7ba732478bad280240969ddf7b0fcbf93900046f63903a83383
- !!merge <<: *qwen3-5-9b-defiant-fable
name: "qwen3.5-9b-defiant-fable"
variants: []
description: |
Qwen3.5 9B Defiant Fable in the plain NEO-imatrix Q4_K_M GGUF format.
This fallback offers the same multimodal reasoning, coding, and creative
capabilities without enabling multi-token prediction.
tags:
- llm
- gguf
- cpu
- gpu
- qwen3.5
- reasoning
- coding
- creative-writing
- uncensored
- vision
- multimodal
overrides:
backend: llama-cpp
function:
automatic_tool_parsing_fallback: true
grammar:
disable: true
known_usecases:
- chat
- vision
mmproj: llama-cpp/mmproj/qwen3.5-9b-defiant-fable/mmproj-BF16.gguf
options:
- use_jinja:true
parameters:
model: llama-cpp/models/qwen3.5-9b-defiant-fable/Qwen3.5-9B-The-Defiant-Fable-Uncnr-Heretic-NEO-MAX-Q4_K_M.gguf
template:
use_tokenizer_template: true
files:
- filename: llama-cpp/models/qwen3.5-9b-defiant-fable/Qwen3.5-9B-The-Defiant-Fable-Uncnr-Heretic-NEO-MAX-Q4_K_M.gguf
uri: huggingface://DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF/Qwen3.5-9B-The-Defiant-Fable-Uncnr-Heretic-NEO-MAX-Q4_K_M.gguf
sha256: d33db5e583b9c9251402e876443791bc979f12af934bfb0630eadfb456279f84
- filename: llama-cpp/mmproj/qwen3.5-9b-defiant-fable/mmproj-BF16.gguf
uri: huggingface://DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF/mmproj-BF16.gguf
sha256: 853698ce7aa6c7ba732478bad280240969ddf7b0fcbf93900046f63903a83383
- &nemotron-3-embed-1b
name: "nemotron-3-embed-1b-q4"
url: "github:mudler/LocalAI/gallery/virtual.yaml@master"
@@ -785,35 +882,18 @@
- name: "qwen3.6-35b-a3b-uncensored-genesis-hermes-v6"
url: "github:mudler/LocalAI/gallery/virtual.yaml@master"
urls:
- https://huggingface.co/HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
- https://huggingface.co/LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V6-GGUF
description: |
# Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
Qwen3.6-35B-A3B Uncensored Genesis Hermes V6 is LuffyTheFox's multimodal,
agentic derivative of HauhauCS's uncensored Qwen3.6-35B-A3B model. It
combines Genesis tensor calibration with Hermes function-calling data while
retaining the 35B mixture-of-experts architecture, roughly 3B active
parameters per token, and the native 262K-token context window.
> **Join the Discord** for updates, roadmaps, projects, or just to chat.
Qwen3.6-35B-A3B uncensored by HauhauCS. **0/465 Refusals.**
> **HuggingFace's "Hardware Compatibility" widget doesn't recognize K_P quants** — it may show fewer files than actually exist. Click **"View +X variants"** or go to **Files and versions** to see all available downloads.
## About
No changes to datasets or capabilities. Fully functional, 100% of what the original authors intended - just without the refusals.
These are meant to be the best lossless uncensored models out there.
## Aggressive Variant
Stronger uncensoring — model is fully unlocked and won't refuse prompts. May occasionally append short disclaimers (baked into base model training, not refusals) but full content is always generated.
For a more conservative uncensor that keeps some safety guardrails, check the Balanced variant when it's available.
## Downloads
All quants generated with importance matrix (imatrix) for optimal quality preservation on abliterated weights.
## What are K_P quants?
...
This entry installs the Q8_0 GGUF together with its F16 multimodal projector
for llama.cpp. The model card recommends Jinja chat templates and at least a
128K context for its thinking behavior. License: Apache-2.0.
license: "apache-2.0"
tags:
- llm
@@ -2009,7 +2089,7 @@
files:
- filename: ds4flash.gguf
uri: https://huggingface.co/unsloth/DeepSeek-V4-Flash-GGUF
sha256: 1bfdafd1c288eb1b2bcb629ee9e1b7567dcf0abbe4d20995905a3c3465e9bd1e
sha256: ea3dc48cb9797ea1bfaa8a74d8a819756b06b16e8fbaa30728ad2cd0a643c605
- name: "qwopus3.6-35b-a3b-coder-mtp"
url: "github:mudler/LocalAI/gallery/virtual.yaml@master"
urls:
@@ -2108,6 +2188,83 @@
- filename: llama-cpp/models/Qwen-AgentWorld-35B-A3B-GGUF/Qwen-AgentWorld-35B-A3B-UD-Q4_K_M.gguf
sha256: e7a8eafdd8013443b6bcc4b6fb47b2d2025f772d359650b9ceb7d75971e22cad
uri: https://huggingface.co/unsloth/Qwen-AgentWorld-35B-A3B-GGUF/resolve/main/Qwen-AgentWorld-35B-A3B-UD-Q4_K_M.gguf
- &agents-a1-4b
name: "agents-a1-4b"
variants:
- model: agents-a1-4b-q8
url: "github:mudler/LocalAI/gallery/virtual.yaml@master"
urls:
- https://huggingface.co/InternScience/Agents-A1-4B
- https://huggingface.co/InternScience/Agents-A1-4B-Q4_K_M-GGUF
description: |
Agents-A1-4B is InternScience's Apache-2.0 dense 4B agentic model, based on
Qwen3.5. It is trained for long-horizon search, engineering and scientific
research, instruction following, tool use, and multimodal tasks. This entry
uses the official Q4_K_M GGUF quantization and vision projector.
license: "apache-2.0"
tags:
- llm
- gguf
- vision
- multimodal
- gpu
- cpu
icon: https://huggingface.co/InternScience/Agents-A1-4B/resolve/main/figures/logo_nobg.png
overrides:
backend: llama-cpp
function:
automatic_tool_parsing_fallback: true
grammar:
disable: true
known_usecases:
- chat
mmproj: llama-cpp/mmproj/Agents-A1-4B-Q4_K_M/Agents-A1-4B-mmproj.gguf
options:
- use_jinja:true
parameters:
model: llama-cpp/models/Agents-A1-4B-Q4_K_M/Agents-A1-4B-Q4_K_M.gguf
template:
use_tokenizer_template: true
files:
- filename: llama-cpp/models/Agents-A1-4B-Q4_K_M/Agents-A1-4B-Q4_K_M.gguf
sha256: d93c393a9bd5139a4b5cfe24d31ef553c5a497bfb8afec178a354ecbf508f062
uri: huggingface://InternScience/Agents-A1-4B-Q4_K_M-GGUF/Agents-A1-4B-Q4_K_M.gguf
- filename: llama-cpp/mmproj/Agents-A1-4B-Q4_K_M/Agents-A1-4B-mmproj.gguf
sha256: 254145e7e03e9e8d3120813fac8033ffa04e411eb6d70a198833504935681084
uri: huggingface://InternScience/Agents-A1-4B-Q4_K_M-GGUF/Agents-A1-4B-mmproj.gguf
- !!merge <<: *agents-a1-4b
name: "agents-a1-4b-q8"
variants: []
urls:
- https://huggingface.co/InternScience/Agents-A1-4B
- https://huggingface.co/InternScience/Agents-A1-4B-Q8_0-GGUF
description: |
Agents-A1-4B is InternScience's Apache-2.0 dense 4B agentic model, based on
Qwen3.5. It is trained for long-horizon search, engineering and scientific
research, instruction following, tool use, and multimodal tasks. This entry
uses the official Q8_0 GGUF quantization and vision projector.
overrides:
backend: llama-cpp
function:
automatic_tool_parsing_fallback: true
grammar:
disable: true
known_usecases:
- chat
mmproj: llama-cpp/mmproj/Agents-A1-4B-Q8_0/Agents-A1-4B-mmproj.gguf
options:
- use_jinja:true
parameters:
model: llama-cpp/models/Agents-A1-4B-Q8_0/Agents-A1-4B-Q8_0.gguf
template:
use_tokenizer_template: true
files:
- filename: llama-cpp/models/Agents-A1-4B-Q8_0/Agents-A1-4B-Q8_0.gguf
sha256: c327f66e820dae550bd230394595071c79f48c88d411b452d013ee4b5999fcea
uri: huggingface://InternScience/Agents-A1-4B-Q8_0-GGUF/Agents-A1-4B-Q8_0.gguf
- filename: llama-cpp/mmproj/Agents-A1-4B-Q8_0/Agents-A1-4B-mmproj.gguf
sha256: 254145e7e03e9e8d3120813fac8033ffa04e411eb6d70a198833504935681084
uri: huggingface://InternScience/Agents-A1-4B-Q8_0-GGUF/Agents-A1-4B-mmproj.gguf
- name: "ornith-1.0-9b"
variants:
- model: ornith-1.0-9b-mtp
@@ -2631,6 +2788,83 @@
- filename: llama-cpp/models/LFM2.5-1.2B-Instruct-GGUF/LFM2.5-1.2B-Instruct-Q4_K_M.gguf
sha256: b1b3de114215d9507409a662a501a631095a479a419584e8a2ded6304b19b4f5
uri: https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF/resolve/main/LFM2.5-1.2B-Instruct-Q4_K_M.gguf
- &lfm2-5-2-6b
name: "lfm2.5-2.6b"
url: "github:mudler/LocalAI/gallery/virtual.yaml@master"
urls:
- https://huggingface.co/LiquidAI/LFM2.5-2.6B
- https://huggingface.co/LiquidAI/LFM2.5-2.6B-GGUF
description: |
LFM2.5-2.6B is LiquidAI's compact, text-only reasoning model for on-device
agentic workloads. It has 2.69B parameters, a 128K-token context window,
multilingual support, and post-training for tool use, instruction following,
data extraction, RAG, and multi-step agents. This entry uses the recommended
Q4_K_M GGUF quantization from LiquidAI's official repository.
license: "other"
tags:
- llm
- gguf
- reasoning
- cpu
- gpu
icon: https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png
variants:
- model: lfm2.5-2.6b-q8
overrides:
backend: llama-cpp
context_size: 131072
function:
automatic_tool_parsing_fallback: true
grammar:
disable: true
known_usecases:
- chat
- completion
options:
- use_jinja:true
parameters:
model: llama-cpp/models/LFM2.5-2.6B-GGUF/LFM2.5-2.6B-Q4_K_M.gguf
repeat_penalty: 1.1
temperature: 0.1
top_k: 50
template:
use_tokenizer_template: true
files:
- filename: llama-cpp/models/LFM2.5-2.6B-GGUF/LFM2.5-2.6B-Q4_K_M.gguf
sha256: 79fdf00351b46cf26f020aead28d01889886be87c55fa0eb907e6f9b00bfee14
uri: https://huggingface.co/LiquidAI/LFM2.5-2.6B-GGUF/resolve/main/LFM2.5-2.6B-Q4_K_M.gguf
- !!merge <<: *lfm2-5-2-6b
name: "lfm2.5-2.6b-q8"
description: |
LFM2.5-2.6B is LiquidAI's compact, text-only reasoning model for on-device
agentic workloads. It has 2.69B parameters, a 128K-token context window,
multilingual support, and post-training for tool use, instruction following,
data extraction, RAG, and multi-step agents. This entry uses the higher-quality
Q8_0 GGUF quantization from LiquidAI's official repository.
variants: null
overrides:
backend: llama-cpp
context_size: 131072
function:
automatic_tool_parsing_fallback: true
grammar:
disable: true
known_usecases:
- chat
- completion
options:
- use_jinja:true
parameters:
model: llama-cpp/models/LFM2.5-2.6B-GGUF/LFM2.5-2.6B-Q8_0.gguf
repeat_penalty: 1.1
temperature: 0.1
top_k: 50
template:
use_tokenizer_template: true
files:
- filename: llama-cpp/models/LFM2.5-2.6B-GGUF/LFM2.5-2.6B-Q8_0.gguf
sha256: 36587fdf27bdfc69caf2637273679a0870ec155162161bde6fd16e8c70bdb757
uri: https://huggingface.co/LiquidAI/LFM2.5-2.6B-GGUF/resolve/main/LFM2.5-2.6B-Q8_0.gguf
- name: "qwopus3.6-27b-coder-compat-mtp"
url: "github:mudler/LocalAI/gallery/virtual.yaml@master"
urls:
@@ -2670,6 +2904,86 @@
- filename: llama-cpp/mmproj/Qwopus3.6-27B-Coder-Compat-MTP-GGUF/mmproj-F32.gguf
sha256: 32f7ea0600c07272547da401d460f8abbd980f3a57b69d6df87be0e2505e0b9c
uri: https://huggingface.co/Jackrong/Qwopus3.6-27B-Coder-Compat-MTP-GGUF/resolve/main/mmproj-F32.gguf
- &qwen3-5-9b-hauhaucs-aggressive
name: "qwen3.5-9b-hauhaucs-aggressive"
variants:
- model: qwen3.5-9b-hauhaucs-aggressive-q8
url: "github:mudler/LocalAI/gallery/virtual.yaml@master"
urls:
- https://huggingface.co/Qwen/Qwen3.5-9B
- https://huggingface.co/HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive
description: |
Qwen3.5 9B Aggressive is HauhauCS's refusal-removed fine-tune of the
multimodal Qwen3.5 9B model. It retains the base model's reasoning, tool
use, image and video understanding, and 262K-token native context window.
This entry uses the balanced Q4_K_M GGUF quantization and includes the
matching BF16 multimodal projector. The Q8_0 variant offers higher fidelity.
license: "apache-2.0"
tags:
- llm
- gguf
- cpu
- gpu
- qwen
- multimodal
- uncensored
icon: https://qianwen-res.oss-cn-beijing.aliyuncs.com/logo_qwen.jpg
last_checked: "2026-08-04"
overrides:
backend: llama-cpp
function:
automatic_tool_parsing_fallback: true
grammar:
disable: true
known_usecases:
- chat
mmproj: llama-cpp/mmproj/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q4_K_M/mmproj-Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-BF16.gguf
options:
- use_jinja:true
parameters:
model: llama-cpp/models/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q4_K_M/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf
template:
use_tokenizer_template: true
files:
- filename: llama-cpp/models/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q4_K_M/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf
sha256: 2ca636d9e81d3d23ca9b60c234fe185d30ec082eeba69ce770fdb0c76559a4f5
uri: huggingface://HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q4_K_M.gguf
- filename: llama-cpp/mmproj/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q4_K_M/mmproj-Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-BF16.gguf
sha256: 05f662501f8bd45607b079723a3e238a4e888fd085a10a53f4057a0e250f6934
uri: huggingface://HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive/mmproj-Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-BF16.gguf
- !!merge <<: *qwen3-5-9b-hauhaucs-aggressive
name: "qwen3.5-9b-hauhaucs-aggressive-q8"
variants: []
description: |
Qwen3.5 9B Aggressive is HauhauCS's refusal-removed fine-tune of the
multimodal Qwen3.5 9B model. It retains the base model's reasoning, tool
use, image and video understanding, and 262K-token native context window.
This entry uses the higher-fidelity Q8_0 GGUF quantization and includes the
matching BF16 multimodal projector.
overrides:
backend: llama-cpp
function:
automatic_tool_parsing_fallback: true
grammar:
disable: true
known_usecases:
- chat
mmproj: llama-cpp/mmproj/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q8_0/mmproj-Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-BF16.gguf
options:
- use_jinja:true
parameters:
model: llama-cpp/models/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q8_0/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q8_0.gguf
template:
use_tokenizer_template: true
files:
- filename: llama-cpp/models/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q8_0/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q8_0.gguf
sha256: 99e7f2201c0046b05d2825e4d8be6a2efad2b87b071cd55d37bdd9fbe201a58b
uri: huggingface://HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q8_0.gguf
- filename: llama-cpp/mmproj/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-Q8_0/mmproj-Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-BF16.gguf
sha256: 05f662501f8bd45607b079723a3e238a4e888fd085a10a53f4057a0e250f6934
uri: huggingface://HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive/mmproj-Qwen3.5-9B-Uncensored-HauhauCS-Aggressive-BF16.gguf
# DFlash speculative-decoding pairs (upstream llama.cpp `draft-dflash`).
# Each entry ships a full target model plus a small block-diffusion drafter
# (z-lab DFlash, converted with upstream convert_hf_to_gguf.py, GGUF arch

2
go.mod
View File

@@ -24,7 +24,7 @@ require (
github.com/gofrs/flock v0.13.0
github.com/google/go-containerregistry v0.21.6
github.com/google/uuid v1.6.0
github.com/gpustack/gguf-parser-go v0.24.0
github.com/gpustack/gguf-parser-go v0.25.0
github.com/hpcloud/tail v1.0.0
github.com/ipfs/go-log v1.0.5
github.com/jaypipes/ghw v0.24.0

4
go.sum
View File

@@ -666,8 +666,8 @@ github.com/gorilla/css v1.0.1/go.mod h1:BvnYkspnSzMmwRK+b8/xgNPLiIuNZr6vbZBTPQ2A
github.com/gorilla/websocket v1.4.2/go.mod h1:YR8l580nyteQvAITg2hZ9XVh4b55+EU/adAjf1fMHhE=
github.com/gorilla/websocket v1.5.4-0.20250319132907-e064f32e3674 h1:JeSE6pjso5THxAzdVpqr6/geYxZytqFMBCOtn/ujyeo=
github.com/gorilla/websocket v1.5.4-0.20250319132907-e064f32e3674/go.mod h1:r4w70xmWCQKmi1ONH4KIaBptdivuRPyosB9RmPlGEwA=
github.com/gpustack/gguf-parser-go v0.24.0 h1:tdJceXYp9e5RhE9RwVYIuUpir72Jz2D68NEtDXkKCKc=
github.com/gpustack/gguf-parser-go v0.24.0/go.mod h1:y4TwTtDqFWTK+xvprOjRUh+dowgU2TKCX37vRKvGiZ0=
github.com/gpustack/gguf-parser-go v0.25.0 h1:1AMBhMKtI24nTtn588Bq53FqNiOvEw1x9Nb4HbRrThs=
github.com/gpustack/gguf-parser-go v0.25.0/go.mod h1:y4TwTtDqFWTK+xvprOjRUh+dowgU2TKCX37vRKvGiZ0=
github.com/grpc-ecosystem/go-grpc-middleware v1.4.0 h1:UH//fgunKIs4JdUbpDl1VZCDaL56wXCB/5+wF6uHfaI=
github.com/grpc-ecosystem/go-grpc-middleware v1.4.0/go.mod h1:g5qyo/la0ALbONm6Vbp88Yd8NsDy6rZz+RcrMPxvld8=
github.com/grpc-ecosystem/grpc-gateway v1.16.0/go.mod h1:BDjrQk3hbvj6Nolgz8mAMFbcEtjT1g+wF4CSlocrBnw=

View File

@@ -2,6 +2,7 @@ package vram
import (
"context"
"fmt"
"strings"
gguf "github.com/gpustack/gguf-parser-go"
@@ -10,7 +11,18 @@ import (
type defaultGGUFReader struct{}
func (defaultGGUFReader) ReadMetadata(ctx context.Context, uri string) (*GGUFMeta, error) {
func (defaultGGUFReader) ReadMetadata(ctx context.Context, uri string) (meta *GGUFMeta, err error) {
// gguf-parser-go parses lengths supplied by the file and has historically
// panicked on values that cannot fit in a Go slice. Metadata can come from
// an untrusted remote host, and this reader is also used by a background
// gallery worker, where an escaped panic would terminate the whole server.
defer func() {
if recovered := recover(); recovered != nil {
meta = nil
err = fmt.Errorf("read GGUF metadata: parser panic: %v", recovered)
}
}()
u := downloader.URI(uri)
urlStr := u.ResolveURL()
@@ -28,7 +40,10 @@ func (defaultGGUFReader) ReadMetadata(ctx context.Context, uri string) (*GGUFMet
if !u.LooksLikeHTTPURL() {
return nil, nil
}
f, err := gguf.ParseGGUFFileRemote(ctx, urlStr)
// The estimator only consumes architecture scalars. Tokenizer arrays can
// be very large and are unnecessary here, so avoid downloading or
// allocating them for remote files just as the local path does above.
f, err := gguf.ParseGGUFFileRemote(ctx, urlStr, gguf.SkipLargeMetadata())
if err != nil {
return nil, err
}

View File

@@ -0,0 +1,115 @@
package vram_test
import (
"bytes"
"context"
"encoding/binary"
"math"
"net/http"
"net/http/httptest"
"time"
gguf "github.com/gpustack/gguf-parser-go"
"github.com/mudler/LocalAI/pkg/vram"
. "github.com/onsi/ginkgo/v2"
. "github.com/onsi/gomega"
)
var _ = Describe("DefaultGGUFReader", func() {
It("reads architecture scalars from a valid remote GGUF", func() {
server := serveGGUF(validRemoteGGUF())
meta, err := vram.DefaultGGUFReader().ReadMetadata(context.Background(), server.URL+"/model.gguf")
Expect(err).NotTo(HaveOccurred())
Expect(meta).To(Equal(&vram.GGUFMeta{
BlockCount: 32,
EmbeddingLength: 4096,
HeadCount: 32,
HeadCountKV: 8,
MaximumContextLength: 8192,
}))
})
It("rejects an overflowing tokenizer array without allocating it", func() {
server := serveGGUF(malformedGGUFArray(math.MaxUint64))
_, err := vram.DefaultGGUFReader().ReadMetadata(context.Background(), server.URL+"/model.gguf")
Expect(err).To(HaveOccurred())
Expect(err.Error()).NotTo(ContainSubstring("parser panic"),
"large tokenizer metadata should be skipped with a bounds error")
})
It("converts a parser panic from malformed string metadata to an error", func() {
server := serveGGUF(malformedGGUFString(uint64(math.MaxInt64)))
_, err := vram.DefaultGGUFReader().ReadMetadata(context.Background(), server.URL+"/model.gguf")
Expect(err).To(HaveOccurred())
Expect(err.Error()).To(ContainSubstring("parser panic"))
})
})
func serveGGUF(payload []byte) *httptest.Server {
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
http.ServeContent(w, r, "model.gguf", time.Time{}, bytes.NewReader(payload))
}))
DeferCleanup(server.Close)
return server
}
func malformedGGUFString(length uint64) []byte {
payload := ggufHeader(1)
payload = appendGGUFString(payload, "general.name")
payload = binary.LittleEndian.AppendUint32(payload, uint32(gguf.GGUFMetadataValueTypeString))
payload = binary.LittleEndian.AppendUint64(payload, length)
return payload
}
func validRemoteGGUF() []byte {
payload := ggufHeader(6)
payload = appendGGUFStringValue(payload, "general.architecture", "llama")
payload = appendGGUFUint32(payload, "llama.block_count", 32)
payload = appendGGUFUint32(payload, "llama.embedding_length", 4096)
payload = appendGGUFUint32(payload, "llama.attention.head_count", 32)
payload = appendGGUFUint32(payload, "llama.attention.head_count_kv", 8)
payload = appendGGUFUint32(payload, "llama.context_length", 8192)
return payload
}
func malformedGGUFArray(itemLength uint64) []byte {
payload := ggufHeader(1)
payload = appendGGUFString(payload, "tokenizer.ggml.tokens")
payload = binary.LittleEndian.AppendUint32(payload, uint32(gguf.GGUFMetadataValueTypeArray))
payload = binary.LittleEndian.AppendUint32(payload, uint32(gguf.GGUFMetadataValueTypeString))
payload = binary.LittleEndian.AppendUint64(payload, 1)
payload = binary.LittleEndian.AppendUint64(payload, itemLength)
return payload
}
func ggufHeader(metadataCount uint64) []byte {
payload := make([]byte, 0, 128)
payload = binary.LittleEndian.AppendUint32(payload, uint32(gguf.GGUFMagicGGUFLe))
payload = binary.LittleEndian.AppendUint32(payload, uint32(gguf.GGUFVersionV3))
payload = binary.LittleEndian.AppendUint64(payload, 0)
payload = binary.LittleEndian.AppendUint64(payload, metadataCount)
return payload
}
func appendGGUFString(payload []byte, value string) []byte {
payload = binary.LittleEndian.AppendUint64(payload, uint64(len(value)))
return append(payload, value...)
}
func appendGGUFStringValue(payload []byte, key, value string) []byte {
payload = appendGGUFString(payload, key)
payload = binary.LittleEndian.AppendUint32(payload, uint32(gguf.GGUFMetadataValueTypeString))
return appendGGUFString(payload, value)
}
func appendGGUFUint32(payload []byte, key string, value uint32) []byte {
payload = appendGGUFString(payload, key)
payload = binary.LittleEndian.AppendUint32(payload, uint32(gguf.GGUFMetadataValueTypeUint32))
return binary.LittleEndian.AppendUint32(payload, value)
}

View File

@@ -0,0 +1,13 @@
#!/usr/bin/env bash
set -euo pipefail
WORKFLOW="$(dirname "$(realpath "$0")")/../../.github/workflows/backend_merge.yml"
sign_commands=$(grep -Ec -- '^[[:space:]]+cosign sign([[:space:]]|$)' "$WORKFLOW" || true)
bundle_flags=$(grep -Ec -- '^[[:space:]]+--new-bundle-format([[:space:]]|$)' "$WORKFLOW" || true)
if [ "$sign_commands" -ne 2 ] || [ "$bundle_flags" -ne "$sign_commands" ]; then
echo "FAIL: every backend signing command must request the new bundle format (commands=$sign_commands flags=$bundle_flags)"
exit 1
fi
echo "PASS: backend signing emits Sigstore bundles for both registries"

View File

@@ -50,6 +50,14 @@ export function inferBackendPath(item) {
if (item.backend === "magpie-tts-cpp") {
return `backend/go/magpie-tts-cpp/`;
}
// nemo-speech-cpp is a Go backend (Dockerfile.golang) wrapping NVIDIA's
// NeMo-Speech.cpp ggml runtime via purego, living in
// backend/go/nemo-speech-cpp/. Same explicit-branch rationale as its siblings
// above: the generic golang fallthrough would also resolve it, but this
// documents the mapping and guards a future dockerfile-suffix change.
if (item.backend === "nemo-speech-cpp") {
return `backend/go/nemo-speech-cpp/`;
}
// trellis2cpp is a Go backend (Dockerfile.golang) wrapping the trellis2.cpp
// ggml port via purego, living in backend/go/trellis2cpp/. Keep the mapping
// explicit so a future dockerfile-suffix change cannot break path filtering.

View File

@@ -1,14 +1,14 @@
---
title: "What landed in LocalAI 4.8"
date: 2026-08-01
date: 2026-08-04
author: "Ettore Di Giacinto"
category: "Release"
tags: ["release", "vllm.cpp", "audio.cpp", "3d", "gallery", "distributed", "performance"]
summary: "A new inference engine, 3D generation, one backend that serves six audio endpoints, and a web interface 3.48x lighter. 321 pull requests in eighteen days."
tags: ["release", "vllm.cpp", "audio.cpp", "3d", "agent", "gallery", "distributed", "performance"]
summary: "A new inference engine, a terminal agent in the CLI, 3D generation, and a web interface 3.48x lighter. 386 pull requests in twenty-two days."
extracss: ["blog.css"]
---
LocalAI 4.8.0 is out. It took eighteen days and 321 merged pull requests, and it pulls in two directions at once: three new things LocalAI can do that it could not do before, and a long list of places where it now does the old things without lying to you.
LocalAI 4.8.0 is out, after twenty-two days and 386 merged pull requests. There are four new things LocalAI can do, and a lot of repair work on things it already did.
The full notes list everything. This post covers the parts that change what you do day to day, with the pull request numbers so you can read the diffs.
@@ -36,6 +36,11 @@ The third one was `/api/traces` returning a 21 MB unpaginated blob that the UI p
## One gallery entry, several builds
<figure>
<img src="/media/v4-8-0-ui-model-variants.png" alt="The model detail pane listing every variant">
<figcaption>One entry, four builds. LocalAI picks the largest that fits and marks it auto-selected.</figcaption>
</figure>
Installing a model no longer means reading a list of quantizations and guessing which one your card will hold. A gallery entry can now declare `variants:`, a list of references to other entries that are alternative builds of the same weights:
```yaml
@@ -55,12 +60,43 @@ Every surface can override the choice: `variant` on `POST /models/apply`, `local
One gap worth knowing about: in distributed mode `InstallModel` resolves against the frontend rather than the worker that will serve the model, so a cluster with a small frontend and large workers selects conservatively. PRs [#10943](https://github.com/mudler/LocalAI/pull/10943), [#10983](https://github.com/mudler/LocalAI/pull/10983), [#10992](https://github.com/mudler/LocalAI/pull/10992), [#11027](https://github.com/mudler/LocalAI/pull/11027) and [#11139](https://github.com/mudler/LocalAI/pull/11139).
## A new engine: vllm.cpp
## A new engine: vllm.cpp (alpha)
[vllm.cpp](https://github.com/mudler/vllm.cpp) is a from-scratch C++20 port of vLLM, written and maintained by the LocalAI team under Apache-2.0, and it ships here as the `vllm-cpp` backend ([#11100](https://github.com/mudler/LocalAI/pull/11100)). It mirrors vLLM's V1 architecture, so paged KV cache, continuous batching, prefix caching, scheduler and sampler, on a portable tensor runtime with no Python, no PyTorch and no ggml at inference. It loads Hugging Face safetensors and GGUF, enforces structured output inside the engine (JSON schema, regex, choice, GBNF), and builds for CPU amd64 and arm64, CUDA 12 and 13 including Blackwell, L4T for GB10, Vulkan and Darwin Metal.
[vllm.cpp](https://github.com/mudler/vllm.cpp) is Apache-2.0 and maintained by the LocalAI team. We want it community-first rather than a LocalAI-only engine, so it lives in its own repository with its own docs, benchmark record and issue tracker, and it runs without LocalAI anywhere in the picture. It began as a C++20 port of vLLM. It ships here as the `vllm-cpp` backend ([#11100](https://github.com/mudler/LocalAI/pull/11100)). It implements vLLM's V1 architecture, so paged KV cache, continuous batching, prefix caching, scheduler and sampler, on a portable tensor runtime with no Python, no PyTorch and no ggml at inference. vLLM stays its reference implementation: correctness is checked by comparing output against it, and the benchmark scoreboard is kept against it.
It has grown features vLLM does not have, which is most of the reason the port exists. It loads GGUF as well as safetensors, runs on CPU, Apple Metal and Vulkan alongside CUDA 12 and 13 and L4T for GB10, and ships speculative decoding and KV offload. Its benchmark page now measures against llama.cpp, MLX-LM and DwarfStar as well as vLLM, because on that hardware those are the engines it competes with. The project is expected to be renamed, with the new name still to be decided; it is drifting far enough that vllm.cpp will eventually mislead.
Tool calling is at llama.cpp parity by construction, because chat deliberately reuses the same autoparser path: full minja chat templates, `tool_choice: auto` lowered to a lazy structural-tag decode constraint, 30 tool dialects, 7 reasoning parsers, and streamed `ChatDelta` and `ToolCallDelta`.
<figure>
<img src="/media/v4-8-0-vllm-cpp-scoreboard.png" alt="Throughput of vllm.cpp relative to each reference engine, drawn as deviation from parity">
<figcaption>llama.cpp is left out because its 1.18x is a prefill ratio, and putting that on the same axis as throughput would compare two different measurements.</figcaption>
</figure>
Numbers from the project's own [scoreboard](https://github.com/mudler/vllm.cpp/blob/master/docs/BENCHMARKS.md), which calls ties ties and losses losses. Above 1.0 means vllm.cpp is ahead:
<div class="tw">
<table>
<thead><tr><th>Reference</th><th>Workload</th><th>Result</th></tr></thead>
<tbody>
<tr><td>vLLM</td><td>Qwen3.6-27B NVFP4, GB10</td><td>1.045x at concurrency 1, 1.007x to 1.017x from c2 to c32, output token-for-token identical</td></tr>
<tr><td>vLLM</td><td>Qwen3.6-35B-A3B NVFP4, GB10</td><td>1.010x at c16 and 1.013x at c32, behind from c1 to c8 (0.817x at c1)</td></tr>
<tr><td>llama.cpp</td><td>Qwen3.5-2B GGUF, CPU aarch64</td><td>prefill 1.18x, decode a tie, memory parity</td></tr>
<tr><td>MLX-LM</td><td>Qwen3-0.6B, Apple M4</td><td>97.6% of warm total, prefill ahead</td></tr>
<tr><td>DwarfStar (ds4)</td><td>DeepSeek-V4-Flash IQ2_XXS, one DGX Spark</td><td>18.69 vs 16.33 tok/s decode, <b>1.144x</b>, same output</td></tr>
<tr><td>vLLM</td><td>Laguna-XS-2.1 NVFP4, GB10</td><td>44.46 vs 43.10 tok/s, <b>1.03x</b>, same output</td></tr>
</tbody>
</table>
</div>
The upstream page is careful about its own noise: on the 27B grid the run-to-run spread is 0.5% and c2 through c32 land between 0.7% and 1.7%, so it calls those five ties rather than wins. The concurrency-1 result is the one it stands behind.
The DeepSeek-V4-Flash row is the one that shows how far this has moved from being a vLLM port. It runs DeepSeek-V4-Flash at roughly 2-bit (IQ2_XXS mixed, about 80 GB) on a single DGX Spark, decoding at 18.69 tok/s against DwarfStar's 16.33. At 300B+ total parameters even a 4-bit checkpoint is 156 GB or more, so a 2-bit GGUF is what fits inside the Spark's 119 GiB unified pool, and reading GGUF is what makes that possible.
That number moved twice in a week, and the second move came from one lever. The dense Q8_0 projection tower was being read from the GGUF mmap over unified memory, which the GB10 reads about 20% slower per-GEMV than device memory. Staging that 6 GiB tower device-resident once at load, same bytes and same kernels, took decode from 16.23 to 18.69, generating the same tokens and using no more peak memory. The same change took Laguna-XS-2.1 from 87% of vLLM to 1.03x ahead of it.
Speculative decoding is in similar shape: MTP on Qwen3.6-27B NVFP4 generates the same tokens as vLLM's MTP and runs about 4% faster at concurrency 1.
Configuration is a normal backend install:
```yaml
@@ -73,9 +109,24 @@ options:
- max_num_seqs:16 # also: block_size:<n>, num_blocks:<n>
```
The CPU path is verified end to end against `Qwen3.5-2B-UD-Q8_K_XL.gguf` with the full Ginkgo suite, covering blocking and streaming byte-parity, greedy determinism, stop words, GBNF-constrained generation, concurrent streams, reasoning split and both `required` and `auto` tool calls. The maturity statement from the release notes is worth repeating in full:
**Treat these as alpha development builds, not a released backend.** vllm.cpp is early, and shipping it in 4.8 is about getting it in front of people who want to try it, not about recommending it for anything you care about. `llama-cpp` stays the default for real use.
> The GPU images build and ship, but their runtime behavior has not been through the same e2e gate yet. This is a first release of a young engine: no throughput comparison against upstream vLLM is claimed here, and `llama-cpp` remains the default recommendation for general use. Try it, and please report what breaks.
The CPU path is verified end to end against `Qwen3.5-2B-UD-Q8_K_XL.gguf` with the full Ginkgo suite, covering blocking and streaming byte-parity, greedy determinism, stop words, GBNF-constrained generation, concurrent streams, reasoning split and both `required` and `auto` tool calls. The GPU images build and ship, but their runtime behavior has not been through that gate. No throughput comparison against upstream vLLM is claimed. Expect rough edges, and please report what breaks.
On Apple Silicon the image now ships vllm.cpp's MLX GEMM provider ([#11137](https://github.com/mudler/LocalAI/pull/11137)). Upstream keeps it off by default because it adds about 124 MB, so we measured before turning it on. Qwen3-1.7B-bf16 on an M4, p=512 g=128, both arms toggled on one binary so a build difference cannot explain the gap:
<div class="tw">
<table>
<thead><tr><th>Batch</th><th>MLX tok/s</th><th>native tok/s</th><th>speedup</th><th>MLX TTFT</th><th>native TTFT</th></tr></thead>
<tbody>
<tr><td>1</td><td>5.79</td><td>3.08</td><td><b>1.88x</b></td><td>3.32 s</td><td>7.68 s</td></tr>
<tr><td>4</td><td>15.75</td><td>10.24</td><td><b>1.54x</b></td><td>9.63 s</td><td>18.77 s</td></tr>
<tr><td>16</td><td>38.65</td><td>17.69</td><td><b>2.19x</b></td><td>18.33 s</td><td>54.48 s</td></tr>
</tbody>
</table>
</div>
Two reps, with rep spread reaching 9.4%, so treat the multipliers as +/-10%. Time to first token roughly halves across the range.
<figure>
<video src="/media/vllm-race.mp4" muted loop playsinline preload="none" data-lazy aria-label="vllm.cpp generating tokens"></video>
@@ -84,7 +135,7 @@ The CPU path is verified end to end against `Qwen3.5-2B-UD-Q8_K_XL.gguf` with th
## LocalAI generates 3D models now
This is a new modality rather than a new backend under an existing one, so it goes through the whole stack: a `Generate3D` RPC in `backend.proto`, a `FLAG_3D` capability so the loader knows which backends can serve it, and `POST /v1/3d/generations`.
3D generation is a new modality, so it had to be wired through the whole stack: a `Generate3D` RPC in `backend.proto`, a `FLAG_3D` capability so the loader knows which backends can serve it, and `POST /v1/3d/generations`.
The first engine behind it is `trellis2cpp`, an image-to-3D backend over TRELLIS.2. You give it an image, you get a GLB back. The web UI has a page for it with a native GLB viewer, so you can turn the result around in the browser instead of downloading it to find out whether it worked, history kept in IndexedDB so a reload does not lose your generations, and previewable print remeshing for output you actually intend to send to a printer ([#10979](https://github.com/mudler/LocalAI/pull/10979)).
@@ -93,9 +144,23 @@ The first engine behind it is `trellis2cpp`, an image-to-3D backend over TRELLIS
<figcaption>trellis2-4b, 2,502,928 vertices and 5,012,118 triangles, turning in the browser. The remesh slider below it is the print path.</figcaption>
</figure>
## `local-ai chat` stopped being a REPL
`local-ai chat` used to be a chat prompt in a terminal. It is now an agent, and it is the [nib](https://github.com/mudler/nib) harness compiled straight into the binary: tool use behind an approval gate, sub-agents, MCP servers, plugins and skills, auto-configured against your own instance. Nothing extra to install.
```bash
local-ai chat # the agent, pointed at your models
echo "what is 2+2" | local-ai chat --cli
local-ai chat --init zsh # Ctrl+Space from any shell prompt
```
That last one prints a shell integration script (zsh, bash or fish), so you can pull the agent up from wherever you already are instead of opening something else.
It runs shell commands now, so every tool call goes through an approval prompt you control, and read-only ones like `ls` and `cat` run without asking. If you had habits around the old REPL, a few things moved: `/clear` is gone and `/compact` is the closest thing, `/models` and `/model <name>` mean what they always meant, and switching model keeps the conversation instead of starting over ([#11291](https://github.com/mudler/LocalAI/pull/11291)).
## One backend, six audio endpoints
The usual shape for audio is one backend per model family, which means a process per capability and a config file for each. `audio-cpp` wraps [audio.cpp](https://github.com/0xShug0/audio.cpp), a multi-family ggml audio engine, and inverts that: one backend process serves several unrelated families through a single runtime vocabulary, and works out which family a checkpoint belongs to from the GGUF's own `audiocpp.model_spec.family` metadata key. There is nothing backend-specific to write in the model config.
The usual shape for audio is one backend per model family, which means a process per capability and a config file for each. `audio-cpp` wraps [audio.cpp](https://github.com/0xShug0/audio.cpp), a multi-family ggml audio engine. One backend process serves several unrelated families through a single runtime vocabulary, and works out which family a checkpoint belongs to from the GGUF's own `audiocpp.model_spec.family` metadata key. There is nothing backend-specific to write in the model config.
<div class="tw">
<table>
@@ -130,7 +195,12 @@ The `bonsai` backend serves the 1-bit (Q1_0) and ternary (Q2_0) Bonsai quantizat
## The operations bar became a page
The old operations bar rendered one row per in-flight operation above every page. Queue four model installs and a backend and it took most of the viewport, on every route, until the last one finished. Two things were conflated there: a global "something is happening" signal, which needs one line, and the detail of what is happening, which needs somewhere to put it.
<figure>
<img src="/media/v4-8-0-ui-activity.png" alt="The Activity page with four installs running">
<figcaption>Four backend installs in flight, and the record of what already finished.</figcaption>
</figure>
The old operations bar rendered one row per in-flight operation above every page. Queue four model installs and a backend and it took most of the viewport, on every route, until the last one finished. It was doing two jobs at once. A global "something is happening" signal only needs one line, and the detail of what is happening needs a page of its own.
The strip is now one line, permanently, showing a failure first and otherwise the least-advanced running operation, with a `+N more` pill. Its `✕` hides the strip and no longer cancels anything. That is a deliberate behavior change worth knowing about before you click it out of habit: the same glyph used to cancel a 17 GB download in one row and dismiss a message in the next. Cancelling moved to the new page, behind a button that says so.
@@ -179,6 +249,6 @@ Valkey Search joins the vector store options as the `valkey-store` backend ([#11
This is also the release where localai.io split in two: the project site at the root, and the documentation under `/docs/`. Every URL that was published before still resolves, through 214 generated redirect stubs, because GitHub Pages has no server-side rewrites to do it properly ([#11243](https://github.com/mudler/LocalAI/pull/11243)).
Twenty-four people contributed to this release, eleven of them for the first time. The gallery went from 1,221 entries to 1,505.
Twenty-five people contributed to this release, eleven of them for the first time. The gallery went from 1,221 entries to 1,515.
To upgrade, pull `localai/localai:latest` or re-run the install script. The [full changelog](https://github.com/mudler/LocalAI/compare/v4.7.1...v4.8.0) has everything this post left out.

View File

@@ -19,7 +19,7 @@
<div><b class="tnum" data-count="{{ .Site.Data.stats.stars }}">0</b><span>GitHub stars</span></div>
<div><b class="tnum" data-count="73">0</b><span>Backends</span></div>
<div><b class="tnum" data-count="{{ len .Site.Data.engines.engines }}">0</b><span>Engines we wrote</span></div>
<div><b class="tnum" data-count="1585">0</b><span>Models, one click</span></div>
<div><b class="tnum" data-count="1255">0</b><span>Models, one click</span></div>
</div>
</div>
<div class="fd">
@@ -39,7 +39,8 @@
<p class="kicker rv">The runtime</p>
<h2 class="rv mt1" style="max-width:21ch">Everything else plugs into LocalAI.</h2>
<p class="lede rv mt2">One binary with an OpenAI-compatible API in front of it. Point an existing client at it and the calls keep working, except now the model is on your machine. It also speaks the Anthropic, Ollama and ElevenLabs APIs, so most tools need a URL change and nothing else.</p>
<p class="lede rv mt2">Underneath, a small core pulls each engine in as a separate backend, only when a model asks for it. That is why one install covers this much ground without becoming a 9 GB download.</p>
<p class="lede rv mt2">The engine behind that API is swappable. One model can run on llama.cpp while the next loads on vLLM, SGLang or MLX, and the client never notices: same endpoint, same request, different engine underneath. Switching is one line in the model's config.</p>
<p class="lede rv mt2">A small core pulls each engine in as a separate backend, only when a model asks for it. That is why one install covers this much ground without becoming a 9 GB download.</p>
<div class="apis rv">
<span>OpenAI API</span><span>Anthropic API</span><span>Ollama API</span><span>ElevenLabs API</span><span>Realtime over WebRTC</span>
</div>
@@ -57,7 +58,7 @@
</div>
<div class="duo__m rv">
<figure class="screen" style="margin:0">
<figcaption class="screen__bar"><i></i> localai · model gallery <b>1,585 models</b></figcaption>
<figcaption class="screen__bar"><i></i> localai · model gallery <b>1,255 models</b></figcaption>
<video src="/media/gallery.mp4" muted loop playsinline preload="none" data-lazy aria-label="Installing a model from the LocalAI gallery"></video>
</figure>
</div>
@@ -327,7 +328,7 @@
<div class="shell">
<div class="bars rv" aria-hidden="true"><i></i><i></i><i></i><i></i></div>
<p class="kicker rv">The gallery</p>
<h2 class="rv mt1" style="max-width:20ch">1,585 models. No notebook, no conversion script.</h2>
<h2 class="rv mt1" style="max-width:20ch">1,255 models. No notebook, no conversion script.</h2>
<div class="cards">
<a class="cd rv" href="/docs/getting-started/models/"><p class="cd__k">Quantizations</p><h3>201 APEX builds</h3>
<p>Every tier of every model we quantize, ranked against the hardware you actually have and installed with one click.</p><span class="cd__go">Browse the gallery →</span></a>

View File

Binary file not shown.

Before

Width:  |  Height:  |  Size: 64 KiB

After

Width:  |  Height:  |  Size: 75 KiB

View File

Binary file not shown.

After

Width:  |  Height:  |  Size: 646 KiB

View File

Binary file not shown.

View File

Binary file not shown.

View File

Binary file not shown.

After

Width:  |  Height:  |  Size: 263 KiB

View File

Binary file not shown.

After

Width:  |  Height:  |  Size: 197 KiB

View File

Binary file not shown.

After

Width:  |  Height:  |  Size: 316 KiB

View File

@@ -0,0 +1,100 @@
<!doctype html>
<html>
<head>
<meta charset="utf-8">
<style>
/* palette lifted from the two logos:
LocalAI #0E2632 navy, #385360 slate, #469AAF teal, #90A8AE haze
vllm.cpp #3AB4CA teal, #95C4D1 light */
:root{
--bg:#0b1c25; --ink:#e8f1f4; --dim:#90a8ae; --faint:#5d757f;
--teal:#3ab4ca; --teal-hi:#7fd4e2; --amber:#e0a944; --rule:#1d3440;
}
*{margin:0;padding:0;box-sizing:border-box}
html,body{width:1600px;height:900px}
body{
background:radial-gradient(1250px 720px at 80% -12%, #143140 0%, var(--bg) 62%);
color:var(--ink);
font-family:-apple-system,"SF Pro Display","Segoe UI",Helvetica,Arial,sans-serif;
-webkit-font-smoothing:antialiased; padding:58px 84px; position:relative;
}
.eyebrow{display:flex;align-items:center;gap:14px;color:var(--teal);
font-weight:600;font-size:23px;letter-spacing:.14em;text-transform:uppercase}
.eyebrow .dot{width:11px;height:11px;border-radius:50%;background:var(--teal);
box-shadow:0 0 16px 2px var(--teal)}
h1{font-size:56px;line-height:1.06;font-weight:760;margin:16px 0 6px;letter-spacing:-.02em}
h1 .grad{background:linear-gradient(92deg,var(--teal),var(--teal-hi));
-webkit-background-clip:text;background-clip:text;color:transparent}
.sub{color:var(--dim);font-size:23px;margin-bottom:14px}
svg{width:100%;height:auto;display:block}
.foot{position:absolute;left:84px;right:84px;bottom:40px;display:flex;
justify-content:space-between;align-items:center;color:var(--faint);
font-size:21px;border-top:1px solid var(--rule);padding-top:16px}
.foot .link{color:var(--ink);font-weight:600}
</style>
</head>
<body>
<div class="eyebrow"><span class="dot"></span>vllm.cpp &middot; throughput vs the reference engine</div>
<h1>Measured against <span class="grad">what each workload actually runs on</span></h1>
<div class="sub">Throughput relative to the reference. 1.00 is parity, bars run from it. Higher is faster.</div>
<svg id="c" viewBox="0 0 1432 585"></svg>
<div class="foot">
<span class="link">github.com/mudler/vllm.cpp</span>
<span>GB10 unless noted &middot; greedy, reference in its own production config &middot; docs/BENCHMARKS.md</span>
</div>
<script>
const rows = [
{ref:'DwarfStar (ds4)', work:'DeepSeek-V4-Flash IQ2_XXS', v:1.144, note:'18.69 vs 16.33 tok/s'},
{ref:'vLLM', work:'Qwen3.6-27B NVFP4, c1', v:1.045, note:'86.05 vs 82.32 tok/s'},
{ref:'vLLM', work:'Laguna-XS-2.1 NVFP4', v:1.030, note:'44.46 vs 43.10 tok/s'},
{ref:'vLLM', work:'Qwen3.6-35B-A3B, c32', v:1.013, note:'3030.5 vs 2993.0 tok/s'},
{ref:'MLX-LM', work:'Qwen3-0.6B, Apple M4', v:0.976, note:'97.6% of warm total'},
];
const W=1432, H=585;
const AX=64; // axis strip reserved at the bottom
const LBL=470; // left label gutter
const R=150; // right gutter for the value
const lo=-0.055, hi=0.165; // deviation domain around parity
const pw=W-LBL-R;
const x = d => LBL + pw*((d-lo)/(hi-lo));
const zero = x(0);
const rowH = (H-AX)/rows.length;
const barH = 46;
let g='';
// faint engineering grid at 2% steps
for(let d=-0.04; d<=0.16001; d+=0.02){
const gx=x(d), on0=Math.abs(d)<1e-9;
g+=`<line x1="${gx}" y1="4" x2="${gx}" y2="${H-AX+10}" stroke="${on0?'#4a6b78':'#16303c'}" stroke-width="${on0?2:1}"/>`;
g+=`<text x="${gx}" y="${H-22}" fill="${on0?'#90a8ae':'#4d6570'}" font-size="17" text-anchor="middle"
font-weight="${on0?'700':'400'}">${(1+d).toFixed(2)}</text>`;
}
rows.forEach((r,i)=>{
const cy = i*rowH + rowH/2;
const d = r.v-1;
const ahead = d>=0;
const col = ahead ? '#3ab4ca' : '#e0a944';
const x0 = ahead ? zero : x(d);
const w = Math.abs(x(d)-zero);
// reference + workload, two weights on one line
g+=`<text x="${LBL-26}" y="${cy-4}" fill="#e8f1f4" font-size="25" font-weight="670" text-anchor="end">${r.ref}</text>`;
g+=`<text x="${LBL-26}" y="${cy+22}" fill="#5d757f" font-size="19" text-anchor="end">${r.work}</text>`;
g+=`<rect x="${x0}" y="${cy-barH/2}" width="${Math.max(w,2)}" height="${barH}" rx="4" fill="${col}" opacity="0.92"/>`;
// value, then the raw measurement under it
const vx = ahead ? x(d)+18 : zero+18;
g+=`<text x="${vx}" y="${cy+1}" fill="${col}" font-size="27" font-weight="700"
font-variant-numeric="tabular-nums">${r.v.toFixed(3)}&times;</text>`;
g+=`<text x="${vx}" y="${cy+23}" fill="#5d757f" font-size="17">${r.note}</text>`;
});
document.getElementById('c').innerHTML=g;
</script>
</body>
</html>

View File

Binary file not shown.

After

Width:  |  Height:  |  Size: 689 KiB