Compare commits

..
Author SHA1 Message Date
ParthSareen 3612bcb601 cmd: remove built-in agent 2026-09-11 10:37:48 -07:00
Anas Khan 02dc3ea4c3 cmd: guard empty editor before indexing fields (#17067)
Signed-off-by: Anas Khan <83116240+anxkhn@users.noreply.github.com>
2026-08-24 09:50:40 -07:00
Parth Sareen fb30760996 app: prevent Apps title bar overlap (#17925) 2026-08-21 18:48:22 -07:00
Devon Rifkin add1f92bdd launch: disable claude code token countdown to preserve KV cache (#17918)
Claude Code adds a "tokens left" system message after every tool
result. Since ollama moves system messages to the front of the prompt,
this breaks the KV cache on every request.
2026-08-21 17:44:44 -07:00
Parth Sareen 124e9af9d2 app: sign model recommendation endpoint (#17919) 2026-08-21 17:21:00 -07:00
Parth Sareen 2d9622a4d4 app: claude model management (#17915) 2026-08-21 15:04:58 -07:00
Eva H 30019c87c4 app: add Connect your apps experience (#17900) 2026-08-21 12:56:51 -07:00
Jesse Gross c44575ef14 mlxrunner: keep prefill snapshots when a request is cancelled mid-prompt
A long prompt records restore points during prefill, but they only
reached the prefix trie when the prefill completed; a cancelled request
closed and released everything it had captured. Agent clients routinely
cancel long prefills — their timeouts are shorter than the minutes a
40k-token prompt takes — so every retry started the whole prompt over
and never got further than the timeout allowed, which presents as the
model hanging forever.

Closing a session now attaches every snapshot the prefill crossed, so a
retry resumes from the last one and makes progress across timeouts.
Scenario tests cover retries resuming exactly where a cancelled attempt
stopped and cancellations on divergent conversation variants.

Fixes #17839
2026-08-21 09:58:33 -07:00
Jesse Gross 81f9a394e9 mlxrunner: grow the prefix trie by whole child nodes so restore points survive resumed prefills
A prefill that resumes partway into cached history — routine once
client timeouts interrupt long prompts — used to attach its captures
onto a node extended in place, so the stored snapshot spanned only the
tokens the prefill evaluated while the node's edge reached further
back. Restores walk node by node and trust each snapshot to cover its
node's edge; the short snapshot stranded the caches at mismatched
offsets and, on models with recurrent layers, ended up freeing all
cache state — a request matching 46k of a 47k-token prompt reprocessed
from zero.

Growth now never extends a node underneath its snapshots. New tokens
become a child node that carries exactly its own captures, and the
path stays compressed because non-user segments merge back into their
parent through the caches' snapshot Merge. Close already pages out
what it records, so every merge combines adjacent covered snapshots
and every stored snapshot spans exactly its node's edge.
2026-08-21 09:58:33 -07:00
Jesse Gross 30e2891808 mlxrunner: page out generated tokens when close records them
When a session closes, every cache rests exactly at the end of the
segment the trie is about to record. That is the one moment the
segment's state can be captured for every layer, so close now pages
the new segment out itself instead of recording it without snapshots
and leaving the capture to a later path switch.

Path switching then has nothing left to capture and only rewinds and
pages in. The whole-state entry taken at close is released when the
next request grows past the segment; sliding-window layers pay the
same window copy a scheduled capture already costs.
2026-08-21 09:58:33 -07:00
Jesse Gross b315b3ee97 mlxrunner: clip prefill captures to their trie node's edge
Page-in restores a path node by node and trusts each stored snapshot to
cover its node's whole edge. A capture taken during prefill spans from
the previous capture or the prefill base, which need not line up with
the node it lands on: when a prefill resumes partway into cached
history, a capture can reach back before its node's start, and a
capture landing on a node that already has snapshots replaced them
with a shorter span that page-in then could not serve.

Clip each capture to its node's edge on attach, and keep the snapshots
the node already has instead of replacing them.
2026-08-21 09:58:33 -07:00
Jesse Gross c01eafa552 mlxrunner: settle the draft caches when a prefill is cancelled
A prefill settles the drafter with the seed token after its last chunk,
leveling the draft caches with the targets; a cancelled prefill
returned before that, leaving the targets one token past the draft
caches and the recorded keys. The next request then had to move every
cache, and models with recurrent layers, which cannot rewind, fell back
to the last snapshot: a retry after a client timeout lost up to a full
snapshot interval of the prompt it had just evaluated.

Settle with the next prompt token on the cancelled path too. The caches
then rest level with the recorded keys, and a retry resumes exactly
where the prefill stopped.
2026-08-21 09:58:33 -07:00
Parth Sareen 8f912415e8 launch: fall back to npx for DeepSeek Harness (#17758) 2026-08-20 14:04:44 -07:00
Eva H 5ad1681cf1 polish onboarding layout and disable zoom (#17885) 2026-08-20 17:03:58 -04:00
Parth Sareen 30546d1fd4 app: add claude desktop app (#17899) 2026-08-20 14:03:43 -07:00
Daniel Hiltgen 6bba484f1a lint fixes (#17897) 2026-08-20 10:23:25 -07:00
Daniel Hiltgen e92b7855f6 mlx update (#17886) 2026-08-20 10:02:41 -07:00
Daniel Hiltgen 4e13421378 mlx: fix mac assumptions on linux/windows (#17898)
The default packaging was broken due to mac assumptions
leaking into windows
2026-08-20 09:50:10 -07:00
Eva H b7871fc0d1 app: add desktop onboarding flow (#17853) 2026-08-19 15:40:16 -07:00
Daniel Hiltgen e0c95a5ffd server: don't wedge chat and generate on a mid-stream parser error (#17883)
When a builtin parser rejects model output, the completion callback wrote the
error to an unbuffered channel and returned. The callback cannot stop
generation -- it has no error return -- so the next chunk re-entered the
callback, hit the same parse error and blocked writing to a channel the
consumer had already stopped reading after emitting its 500. The completion
never returned, the goroutine leaked and the runner request was never
released, so retrying the same prompt hung with no log output until the client
gave up.

Record the parse error, cancel the completion, and report it once the
completion has returned. Parse failures landing on the final chunk were
already terminal, which is why non-thinking requests and the direct
qwen3-coder parser path failed cleanly and only thinking mode wedged.

ChatHandler and GenerateHandler share the defect: both run the same parser in
the same shape of callback behind a consumer that stops reading at the first
error. GenerateHandler had no cancel func at all, so one is added there.

Fixes #17825
2026-08-19 15:14:50 -07:00
Parth Sareen b8a6272440 qwen3.8: normalize system messages (#17855) 2026-08-19 13:08:09 -07:00
Daniel Hiltgen d1bd15ccce ci: plumb temporary MLX patch through to docker stages (#17874)
Follow up to #17850
2026-08-19 08:33:53 -07:00
Daniel Hiltgen 0bb0925920 mlx update (#17850)
Temporarily carry https://github.com/ml-explore/mlx-c/pull/127
2026-08-19 07:15:11 -07:00
Gaurav Garg a5165c53ac Add a model metadata cache to reduce Ollama’s per-request overhead (#17752) 2026-08-18 12:22:51 -07:00
Daniel Hiltgen cd37044093 llama.cpp update (#17851) 2026-08-18 11:53:00 -07:00
Daniel Hiltgen d67ad83426 mlx update (#17761) 2026-08-15 11:56:40 -07:00
Daniel Hiltgen e5a81899d0 llama.cpp update (#17760) 2026-08-14 19:03:20 -07:00
Parth Sareen 78e818e3ce docs: register DeepSeek Harness (#17751) 2026-08-14 14:30:21 -07:00
Daniel Hiltgen 87abaa019e renderers/qwen: tolerate non-leading system messages (#17757)
Coding clients may insert runtime system messages after the initial user turn. The shared Qwen renderer rejected these transcripts before rendering, turning a potentially usable non-standard request into an HTTP 500.

Pass non-leading system turns through the existing raw ChatML path and warn when qwen3.8 encounters one. Extend the Anthropic tool-route integration scenario to cover this message pattern and remove the obsolete rejection test.
2026-08-14 14:12:50 -07:00
Daniel Hiltgen f427fa0753 llm: transcode WebP images for llama-server (#17755)
llama-server does not currently support WebP image payloads. Detect WebP media before forwarding, and transcode it to PNG. Pass all other media through unchanged.

Replace an existing vision integration image with a lossless WebP version so we now have coverage of JPG/PNG/WebP formats.

Fixes #17753
2026-08-14 13:21:11 -07:00
Daniel Hiltgen 0f25c31bd5 qwen3.8: support developer instructions (#17749)
* qwen3.8: support developer instructions

Qwen3.8 does not define a developer role, while OpenAI-compatible coding agents commonly send developer instructions before user messages. Fold the leading system/developer instruction prefix into a single system turn before Qwen3.8 validation, preserving instruction precedence without changing Qwen3.5 or other renderer behavior.

Add streaming tool-call integration coverage for the native Ollama, OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages request shapes. Each case exercises prior assistant tool calls, tool results, follow-up rendering, and parsed tool-call output. Add Qwen3.8 to the release tools sweep.

Removes an unnecessary unit test that should not have been included in the original 3.8 PR.

* review comments
2026-08-14 11:30:27 -07:00
Daniel Hiltgen 5512797527 qwen3.8: add renderer and MLX import support (#17745)
Qwen3.8 keeps the Qwen3.5 model architecture and parser, but its chat template adds reasoning-effort and preserved-thinking semantics. Detect those template markers during safetensors import, select the qwen3.8 renderer, and cover thinking, tools, continuation, and malformed parser input.

Make indexed safetensors imports use the weight map's shard names instead of independently filtering files by the model-* convention. Reject unsafe shard paths, ignore unindexed tensors, and fail when an indexed weight is missing or stored in a different shard. Retain the conservative model-* scan when no index is present.

Treat Classification.Quantize as the effective tensor format and pass it to the manifest writer. This records file_type for automatic block-FP8-to-MXFP8 conversion and recognized prequantized inputs, preserves requested quantization and base-plus-draft behavior, and avoids claiming one type for mixed or unknown formats.

Normalize both supported convolution weight layouts with an explicit reshape. Add focused unit coverage for renderer selection, parser behavior, shard inventory, manifest metadata, and convolution layout; heavyweight reference-forward and release integration checks remain bring-up artifacts.
2026-08-14 09:31:09 -07:00
Parth Sareen 39df91c982 launch: add DeepSeek Harness integration (#17733) 2026-08-13 15:19:02 -07:00
Daniel Hiltgen 7ce88bd686 model/renderers: match Muse Glimmer reasoning template (#17732)
Updates Muse Glimmer Jinja reference template to the latest publisher version and mirror its explicit-system reasoning handling in the Go renderer.

Explicit system prompts now normalize "Reasoning effort" to "Reasoning strength" and skip adding a renderer-provided reasoning line when the prompt already contains one. This prevents duplicate or conflicting reasoning directives while preserving the default-system behavior.

Add reference tests for both normalization and deduplication, including Jinja-backed validation.
2026-08-13 14:35:50 -07:00
Daniel Hiltgen 01d04d50f8 launch: add Muse Code integration (#17594)
* launch: add Muse Code integration

Add `ollama launch muse` for Meta's Muse Code CLI.

Muse only takes a model catalog from settings.json (normally it fetches one from its provider and refuses to start otherwise), and that file's endpoint_transport is a global provider switch. So the integration writes a settings file under its own config root (~/.ollama/launch/muse-config via XDG_CONFIG_HOME), leaving a Meta-backed muse install untouched, and re-seeds it from muse's own persisted copy on later runs.

The launched model is preloaded so its catalog row carries the context length the server actually allocated, not the trained maximum; the loaded-context helpers move from cmd/agent_tui.go into cmd/launch for reuse.

Muse sends reasoning efforts outside Ollama's scale (minimal, xhigh, ultra), which were hard 400s; clamp them to the nearest tier in one helper shared by the chat and responses converters.

The registry entry stays Hidden (alias "muse-code"), like kimi and vscode.

* review comments

* skip muse test on windows (unsupported platform)
2026-08-13 13:10:29 -07:00
Parth Sareen 9a56a0e845 agent: allow multiple edits per edit tool call (#17711) 2026-08-12 16:45:53 -07:00
Daniel Hiltgen 88313499e0 mlx: avoid pulling MLX models when MLX is missing (#17710)
As we look to bring Linux and Windows MLX support online, instead of blocking
downloads at the registry to avoid users wasting time downloading a model they
can't run, shift the logic to the local side which knows if MLX is present or not.
2026-08-12 14:42:17 -07:00
Jesse Gross 2b4a99376c nn: speed up prefill on double-scale nvfp4 models
ModelOpt checkpoints apply a float32 global scale to every projection
output on top of the per-group quantization scales. Running the
multiply and the cast back to the activation dtype as separate eager
ops costs an extra kernel launch and a materialized intermediate per
projection.

Compile the multiply and cast into one kernel. On an M5 Max (medians
of order-swapped A/B runs against main; greedy outputs byte-identical):

    qwen3.6:27b        prefill  703 -> 769 t/s  +7.9%
    muse-glimmer:30b   prefill  790 -> 843 t/s  +6.7%

Speculative decode is unchanged within noise on both models. Only
checkpoints with a global scale are affected; single-scale nvfp4,
mxfp8, and affine checkpoints take the unchanged path.
2026-08-12 13:25:33 -07:00
Daniel Hiltgen e922bc7125 llama.cpp bump (#17702) 2026-08-12 12:10:18 -07:00
Daniel Hiltgen 950dd9ac67 MLX update (#17704) 2026-08-12 12:09:50 -07:00
Parth Sareen b6b1b258c3 openai: support web search in Responses API (#17686) 2026-08-12 11:51:54 -07:00
VigneshandPatrick Devine 4138e853d5 server/images: prevent skipVerify map collision with duplicate digests (#15504)
When a manifest contains a config and layer with the same digest, the
skipVerify map entry was overwritten by the config's cache-hit value
(true), replacing the layer's non-cache-hit value (false). This caused
verifyBlob to be skipped for the freshly downloaded blob.

A rogue OCI registry could exploit this by serving a manifest with
duplicate digests and redirecting blob downloads to internal endpoints.
The SSRF response would be written to disk, hash verification would be
skipped due to the map collision, and the blob would persist.

The fix uses logical AND when updating skipVerify: once any download of
a digest was not a cache hit, verification is always performed.

Fixes #15485

---------

Co-authored-by: Patrick Devine <patrick@ollama.com>
2026-08-12 11:44:32 -07:00
Daniel Hiltgen 641df5e5ad mlx: enable CUDA backend in CUDA builds (#17688) 2026-08-12 07:31:00 -07:00
Jesse Gross 6a261db7d8 api: stop applying repeat_penalty 1.1 to models that don't set one
Request options are the model's published parameters and the request's
own options layered over the server defaults, so the default
repeat_penalty of 1.1 reaches every model whose parameters leave it
unset. No maker of the library's current models recommends 1.1: their
generation configs either omit the penalty, meaning 1.0, or pin 1.05.
llama.cpp dropped the same 1.1 default in 2024; vLLM, SGLang, and
transformers apply no penalty. An always-on penalty also distorts
output that legitimately repeats tokens, such as code, JSON, and long
reasoning traces.

The penalty is especially costly for speculative decoding, where
drafts are proposed without it: the penalized target rejects drafted
tokens and the depth controller backs off. On muse-glimmer 30B (DFlash
on M5 Max, HumanEval) the 1.1 default costs 13-16% of end-to-end
throughput at greedy and temperature 1 alike, and drops prose
acceptance at temperature 0.8 from 0.44 to 0.30. On qwen3.6-35B it
cuts the mean accepted draft length from 4.3 to 3.5 tokens and makes
the controller stop speculating on prose.

Defaulting to 1.0 disables the penalty unless a model's parameters or
the request set one. Across the library:

- qwen3, qwen3.6, and qwen3-coder pin their own values (1.0, 1.0, and
  Qwen's recommended 1.05) and are unchanged.
- Everything else local now matches its maker's no-penalty
  recommendation, including gemma2 through gemma4, muse-glimmer, both
  laguna 2.1 models, qwen3.5 (previously 1.1 stacked on its
  presence_penalty of 1.5), gpt-oss, deepseek-r1 and v3.1, the
  nemotron family, granite4, the mistral and llama3/llama4 families,
  phi4, glm4, llava, and devstral.
- qwen2.5 recommends 1.05 but ships no parameters, so it moves from
  1.1 to 1.0 and still needs a parameters layer to conform.
- Cloud models (kimi-k3, deepseek-v4-flash) never receive these
  defaults.

Small older models may repeat themselves more without the penalty
masking it; the remedy is a per-model parameter, not a penalty applied
to every model.
2026-08-11 21:47:51 -07:00
Eva H 948f69330a docs: fix broken links (#17676) 2026-08-11 14:36:39 -07:00
Daniel Hiltgen 96fb6d2fa9 nemotron_h: support the Nemotron 3.5 prompt layout (#17672)
Select the 3.5 parser and renderer from its checkpoint template, preserve its prompt semantics, and map medium reasoning effort to the final-user annotation expected by the reference template.

Exercise parser and renderer registration, create-time metadata inference, and exact Jinja parity so created models cannot silently fall back to the Nemotron 3 renderer.
2026-08-11 06:18:51 -07:00
Daniel Hiltgen 400164d47c parsers: recover boundary tokens fumbled into glimmer ATEM invoke names (#17664)
The model occasionally emits a <|message|> boundary token in the invoke
name region, echoing the header form `to=read<|message|>`. The existing
recovery handled the tag inside a terminated name (`name="read<|message|>">`)
but not the fleet-observed shape where the tag replaces the `">` terminator
itself (`name="read<|message|><atem:parameter ...`), which failed the call
with "malformed ATEM parameter".

Replace the strip-after-cut recovery with a single name scan shared by
parseGlimmerATEM and the content fallback: the name ends at the first `">`,
boundary tokens before it are dropped, and a parameter element immediately
after a dropped token means the token replaced the terminator. Well-formed
calls are unaffected — a boundary token is never legitimate before the
terminator, and parameter values (where the literal text is preserved) only
appear after it. Murkier garbles still fail loudly, the recipient
cross-check still applies, and the recovery WARN is retained.
2026-08-10 21:48:17 -07:00
Daniel Hiltgen bb7bba885e mlx: implement Nemotron 3 Nano Omni (#17060)
Add MLX support for Nemotron 3 Nano Omni, including the model implementation, Mamba2/recurrent pieces, MoE routing, and quantized NVFP4/MXFP8 expert paths.

Use a shared mapped MoE GatherQMM fast path under the generic moe_gather_qmm_mapped naming, with Metal-optimized NVFP4/MXFP8 block-mapped kernels and generic fallbacks for unsupported backends.

Serve the model's multi-token prediction head as a self-draft speculator, so speculative decoding needs no separate draft model.

Render the Nemotron prompt from the published chat template. The template the renderer was based on had drifted from the current reference; refreshing it surfaced five mismatches: stray leading newlines, the wrong turn separator and a trailing newline before the generation prompt; /think and /no_think toggles left in user turns; a trimmed system message the template leaves intact; a user block opened by a leading tool message; and Go scalar syntax for schema extras where the template applies Python str(), sending true/false/<nil> in place of True/False/None. Reference tests now render every case through the template itself.

Also harden the Nemotron parser path shared by both backends: while collecting thinking, preserve whitespace before partial </think>, <think>, and <tool_call> fakeouts, with streaming tests covering those cases.
2026-08-10 21:42:34 -07:00
Daniel Hiltgen 4f066a6fb0 llama.cpp update (#17659) 2026-08-10 15:48:32 -07:00
Eva H a836eb8c3c docs: require VS Code 1.127 (#17655) 2026-08-10 11:21:26 -07:00
Eva H 1a9e4235ac docs: add VS Code context length guidance (#17610) 2026-08-10 09:46:25 -07:00
Daniel Hiltgen 43f4eda808 Release v0.32.7 (#17646)
* glimmer: implement the Muse Glimmer model

MLX model (language + vision encoder) with DFlash draft wiring, llama-server DFlash support and rope-interleave fix, renderer and parser, tokenizer fixes, and the import quantization policy.

* mlxrunner: report committed prefill chunks after the sweep and eval

The drafter's flush evaluates its report, and an eval that runs while the chunk's construction handles are still live cannot free any intermediate buffer. On media chunks that retention keeps the whole vision tower resident and grinds the Metal allocator at its limit until the request dies. Pin the report's inputs across the sweep, report after the chunk materializes, and release media items after the report so a drafter can still capture the rows its deferred flush embeds.

* ci: retry CUDA pre-release download
2026-08-10 04:04:56 -07:00
Daniel Hiltgen acdf81510d MLX: version bump (#17637)
Also bring back version tagging the MLX library with our git hash which was
accidentally dropped when imagegen was removed.  Without this, the version
claimed to be the official tagged version, but we're typically using a git hash
with different content.
2026-08-09 10:38:49 -07:00
Jesse Gross 1e85fe8e9a qwen3_5: image input support
One vision path serves every qwen3.5/qwen3.6 registration, dense and
MoE. Rope positions are precomputed at prepare time as the request's
layout — the family uses interleaved M-RoPE — while text-only requests
keep the fused 1D rope path, which is numerically identical for
uniform channels. Image expansions are causal for this family, so
prefill chunks split them. The MTP head embeds prompt tokens, so it
scatters the delivered image features and applies the same position
tables, keeping speculative decoding working on image prompts. The
merger's exact erf GELU adds an Erf op to the MLX bindings.

A checkpoint whose config declares vision must ship its tower: missing
vision weights or a deepstack_visual_indexes request fail the load
rather than silently serving text-only or skipping the injections.
Text-only checkpoints, which carry no vision_config, load as before.

Verified tensor-by-tensor against HF transformers for all eight family
members and live on every published -mlx tag; published towers are
already bf16, so no re-import is needed.
2026-08-09 10:37:05 -07:00
Jesse Gross 5fcf71b8b8 mlxrunner: feed media features to the model during prefill
Each media item's features are encoded lazily when a prefill chunk
first overlaps its expansion and stay pinned until the expansion is
fully evaluated. A chunk never ends strictly inside an atomic
expansion: a bidirectional run's early rows attend its later keys, so
its first evaluation must cover the whole run in one forward. Items
marked Causal are exempt and split at any boundary.

Draft models need the same request state — reference MTP drafters
embed prompt tokens with the image features merged in, and an M-RoPE
drafter cannot compute positions without the request's layout — so the
layout is stamped on every forward, target and draft alike, and the
MTP session holds feature rows across its deferred flush. The dflash
drafter ignores media: its context rows are target hiddens.
2026-08-09 10:37:05 -07:00
Jesse Gross 60bdc23467 mlxrunner: expand image tags into placeholder tokens
A prompt that references media arrives as text containing [img-N] tags
plus the media bytes. Prepare now splits on the tags, tokenizes the text
between them, and hands the model the resulting segments — text runs and
media in stream order — in a single PrepareMedia call. The model returns
the expanded stream with each media segment's placeholder expansion
spliced in place, described per item so the runner can key identity and
schedule encoding, along with any opaque request-scoped layout state it
derives while building the stream. Building the whole stream in one call
is what lets a model derive values that span items, and lets it choose
item granularity (one per image, or one per independently evaluable
tile).

The runner validates the model-authored items before trusting them —
ranges ordered, non-overlapping, in bounds, and covering every media
segment, since prefix-cache identity is keyed on them.

Unknown tag IDs fail the request, media the prompt never references is
ignored with a warning, and duplicate references are allowed, matching
the previous engine. A media request still produces no image output:
nothing feeds the features to the model yet, and no model implements
the media interface.
2026-08-09 10:37:05 -07:00
Jesse Gross 694487c65b mlxrunner: fold media identity into prefix-trie keys
Media placeholders repeat one token ID, so two prompts with different
images would produce identical trie keys and falsely share cached state.
Substitute a per-item hash of the media bytes and preprocessing shape
across each item's expansion range at the key layer; the model still
sees real token IDs. Fold values carry a bit no token ID has, so a media
stream can never alias text, and the bigram packing for draft caches
composes unchanged, so draft restore points inherit the same identity.

Text-only prompts key exactly as before. Nothing records media items yet;
the change is inert until the prompt preparation wires them.
2026-08-09 10:37:05 -07:00
Jesse Gross af5b627672 mlxrunner: reject media requests the model cannot serve
MLX checkpoints that include a vision tower are already tagged with the
vision capability at import, so the server accepts image chats and ships
the image bytes with the completion request. The MLX client dropped the
bytes, and the prompt's image tags were answered as literal text.

Carry the media through to the runner and fail the request with a clear
error when the loaded model has no media support. Nothing implements the
new media interface yet, so every media request now returns the error
rather than a silently wrong answer; later changes build the image path
on top of the same interface.
2026-08-09 10:37:05 -07:00
Jesse Gross 8713570d3c create: keep vision towers at source precision when quantizing
Vision towers are much more sensitive to weight quantization than
language layers: measured against the reference encoder on a real
image, 4-bit types and scale-only mxfp8 distort the projected image
features by 26-34% mean relative error (worst tokens near-orthogonal),
which shows up as degraded image recognition — down to complete
blindness for the small e-series towers under nvfp4. Affine 8-bit was
the only quantized format that matched the bf16 tower.

Keep vision tower tensors at source precision instead, matching the
audio tower's treatment and every vision component Ollama publishes in
GGUF form, including gemma4's own GGUF tags, which ship f16/f32 vision
beside 4-bit language weights. Towers are small and run once per image,
so neither size nor decode bandwidth argues for quantizing. Existing
MLX imports keep their quantized towers until re-imported.
2026-08-09 10:37:05 -07:00
Daniel Hiltgen 5a173edb63 manifests: remove OCI rootfs from the model config (#17619)
rootfs.diff_ids duplicated the manifest's layer digest list into the config blob and nothing ever read it. On per-tensor safetensors models the copy grows past 100KB and create excessively large config blobs with unused redundant data. Model identity is unaffected: it is the digest of the manifest itself, which already commits to every layer hash.
2026-08-08 19:44:57 -07:00
Jesse Gross b880b76c43 laguna: wire the DFlash target side
Add what a DFlash draft borrows from its target: the tapped layer
outputs, the raw embedding lookup, and the undecorated lm_head
projection. The laguna draft architecture (DFlashLagunaForCausalLM) is
registered here, alongside the only wired target.

Matched nvfp4 target+draft pairs, M5 Max, temp 0.8, repeat_penalty 1.1,
adaptive depth; decode tok/s:

                     prose   code   edit
  laguna-xs  plain   139.4  139.7  137.3
             DFlash  142.3  139.2  145.1
  laguna-s   plain    75.4   70.0   72.4
             DFlash   74.6   80.8  115.3
2026-08-07 19:33:35 -07:00
Jesse Gross cf129bbb11 dflash: add the DFlash block-diffusion draft model
Implements the DFlash draft checkpoint format: a few decoder layers
over fused target-layer outputs, which enter every layer as key/value
context while the block being drafted supplies the queries. The draft
has no embedding table or output head of its own; it borrows the
target's.

One model covers the known checkpoints. Attention weights normalize at
load to a q projection plus a fused k|v, stacking split checkpoints and
slicing fused ones, exact for quantized tensors; gate and up fuse the
same way. Optional tensors decide the output gate and per-tap norms,
config decides attention shape, and the architecture name decides only
laguna's context-norm convention.

A manifest can pair any draft with any target, so construction checks
the fit: tap ids inside the target's layers, matching hidden width, and
the target vocabulary covering the mask token. A bad pairing fails at
load.
2026-08-07 19:33:35 -07:00
Jesse Gross c1bf60d7b1 mlxrunner: add a block-diffusion drafting session
A DFlash draft proposes a whole block per forward, which doesn't fit
the MTP session's one-token-per-call chain. Add a second drafting
session for block drafts: committed target features write straight
into the draft's context caches, and each round drafts a block in one
forward and samples it in one batched call, rolling the block's cache
entries back with the same mechanism speculative rounds use on the
target caches. The depth controller's search is capped at the deepest
draft the drafter can produce, since a depth it can never measure
would otherwise always look best.
2026-08-07 19:33:35 -07:00
Jesse Gross 0fcfc99ea0 sample: define multi-row distributions without a draft chain
Distribution aligns its rows with the end of the draft chain, so when
the caller passes no chain, every row sees the slot history unchanged.
That case already worked; only the row-count guard rejected it. The
guard now applies only when a chain is present, which is where more
rows than chain positions would silently drop history. A block drafter
needs the chainless case to sample its whole proposal batch in one
call.
2026-08-07 19:33:35 -07:00
Jesse Gross e7fbd528f7 mlxrunner: let each model declare the cache slots it needs
The runner used to build caches by probing the model for an optional
NewCaches method, with one KV cache per layer as the fallback. A model
with a draft head appended the draft's cache slots to its own list, and
the speculative engine later recovered the two groups by comparing slot
identities, panicking when the lists didn't line up.

NewCaches is now a required method on both the model and the draft, and
each returns only the slots it writes. The runner concatenates the two
lists for the prefix cache and passes them to the speculative engine
separately, so snapshots and rollback apply to the target's slots and
the draft forward receives both groups as arguments. The identity
comparison, its panics, and the per-request rebinding are gone; the two
groups are fixed at load time.
2026-08-07 19:33:35 -07:00
Jesse Gross 2f84872ce0 mlxrunner: return the draft-conditioning state from a model forward
A draft model conditions on state that the target produces during its
own forward pass. For an MTP head or an assistant model that state is
the final hidden state; for a block draft it is the concatenated
outputs of several layers. The choice belongs to the model, so Forward
now returns the conditioning state along with the hidden state to
unembed. Models without a special conditioning state return the final
hidden state for both, and the decode paths hand the value to the
drafter without looking at it.
2026-08-07 19:33:35 -07:00
Parth Sareen f91cb0d6a7 agent/tui: stream thinking traces (#17611) 2026-08-07 13:11:45 -07:00
Eva H 8dd34b77d1 cmd/tui: restore launcher integrations menu (#17595) 2026-08-07 11:11:16 -07:00
Daniel Hiltgen 35f71382de openai: expand namespace tool declarations in the responses API (#17593)
The Responses API groups related tools by domain: a tool with type "namespace" carries the real function definitions in a nested tools array. The conversion dropped that array, leaving the model a single schema-less pseudo-function and making every namespaced call undeclarable.

Expand namespace declarations into their member functions with namespace-qualified names, since api.Tool carries only a flat function name.

Relates to #15921: full Responses API parity also wants the namespace preserved as a separate field on tool calls in the output, which needs new api surface and is not addressed here.
2026-08-07 09:52:37 -07:00
Jesse Gross 144893850f mlxrunner: stop cache rewind refills from corrupting later lazy snapshots
A lazy KV snapshot indexes into the cache's live buffer instead of owning
a copy, so it must be copied out before an append overwrites the slots it
names. appendKV checked for that only on the first append after a rewind,
and only against that append's own range: a still-lazy snapshot further
ahead in the buffer was overwritten without a copy when a later append
reached it. This happens when a request reuses a short prefix of a longer
cached conversation and prefills past one of the old conversation's
snapshots; restoring that snapshot later silently serves the new request's
KV in place of the old conversation's.

Scan every append instead. The overlap test already limits copies to
snapshots the current write clobbers, and appends outside a rewind refill
sit above every snapshot, so the steady-state scan walks a short list and
finds nothing. This restores the invariant Restore's lazy fast path relies
on: a snapshot still in its lazy state has never been overwritten.
2026-08-05 16:31:56 -07:00
Daniel Hiltgen 26936bea45 ci: fix race in darwin build (#17578)
Do vendoring work once at the top level build to avoid 2 nested builds fighting
with eachother.
2026-08-05 11:10:18 -07:00
Daniel Hiltgen 43983edf18 progress: fix data races on ticker, states, spinner, and bar state (#17445)
* progress: fix data races on ticker, states, spinner, and bar state

NewProgress spawned start() which wrote p.ticker while stop() read and
cleared it with no synchronization; stop() and StopAndClear() also read
p.states and p.pos outside p.mu, Spinner's start() goroutine raced
Stop() and String() on s.value/s.stopped/s.ticker, and Bar.Set raced
Bar.String on currentValue/stopped/buckets (callback goroutine vs the
render goroutine). Detected by go test -race across cmd and cmd/launch
(~20 warnings; the Bar race is latent — never flagged because tests
don't interleave it, but real in production pull/push progress).

Create tickers before spawning the render goroutines and pass the
channel in, guard Progress internals with p.mu throughout stop() (via a
renderLocked core), and give Spinner and Bar their own mutexes.

* use a more idiomatic channel based done signal
2026-08-04 15:06:15 -07:00
Daniel Hiltgen c82ebbd5bf llama.cpp update (#17545) 2026-08-04 09:51:52 -07:00
Bruce MacDonald 8edecb5c69 openai: match openai's streaming wire format for chat completions (#17485)
Reworked our /v1/chat/completions streaming to match what api.openai.com actually sends,
chunk-for-chunk, based on captures I took of real OpenAI traffic.

What changed:
 - finish_reason now goes on its own chunk with an empty delta {}, instead of riding on the last content
   chunk. Precedence is length > tool_calls > the response's done reason > stop.
 - role is only sent on the first chunk of a stream, not on every chunk.
 - With stream_options.include_usage, usage goes out on its own chunk with choices: [] after the finish
   chunk.
 - A truncated response keeps finish_reason: "length" even when tool calls were streamed — it used to get
   overwritten with "tool_calls". Fixed in both streaming and non-streaming paths.
 - The metrics-only trailer response (empty message at end of stream) no longer produces a stray
   delta:{"content":""} chunk before the finish chunk. A wholly empty completion still opens with a role
   chunk.
 - Every chunk in a stream shares one timestamp, from the response's CreatedAt.
2026-08-03 15:36:57 -07:00
Devon Rifkin 8d8c701d6a Merge pull request #17483 from ollama/drifkin/suggest-cloud
cmd: suggest :cloud when a model has no default tag
2026-07-31 14:28:04 -07:00
Daniel Hiltgen b63eed94b6 app/updater: drain background update-check goroutine before returning (#17446)
DownloadNewRelease spawned a background checkForUpdate loop that read
package-level knobs (UpdateCheckInterval et al.) and returned without
waiting for it, so under -race the next test rewrote those globals while
the orphaned goroutine was still reading them. waitDownloadIdle (from

Cancel and WaitGroup-drain the loop before DownloadNewRelease returns,
and have TestCancelOngoingDownload join its download goroutine so the
drain is observable before the test exits.
2026-07-31 10:42:40 -07:00
Jesse Gross 4f9d09ef52 qwen3_5: load and run the MTP head as a speculative draft
Load the MTP head from the mtp.* tensors instead of freeing them and implement
Draft to propose one token per step, gated solely on the tensors being
present; a model whose head ships inline is its own draft via base.SelfDraft.
The runtime keeps sole ownership of the +1 RMSNorm shift (conversion passes
tensors through verbatim), and the head's norms shift under the same
original-format detection as the main stack, so nothing shifts twice.
2026-07-31 10:18:54 -07:00
Jesse Gross ba8f2a324d nn/recurrent: run the gated-delta step in one launch
Decode-length scans spend more time in launch gaps than math: the q/k
norms, decay gate, and recurrence each dispatched separately per layer.
Fuse the step into one Metal kernel over the activated conv output,
with per-token boundary states available from the same pass. The graph
implementation remains as the fallback and contract-miss path, and pins
the kernel bit-for-bit in the parity test.
2026-07-31 10:18:54 -07:00
Jesse Gross 721f05049d nn/recurrent: activate the conv output in CausalConv1D
The activation belongs to the conv stage: downstream consumers see
activated values however the conv is computed. WithConvSiLU routes to a
fused depthwise conv+SiLU kernel when the conv fits its contract and
the same computation as graph ops otherwise; cached conv state is the
raw input tail, unaffected by activation placement.
2026-07-31 10:18:54 -07:00
Jesse Gross accd6d656a mlx: factor custom GPU kernel scaffolding into helpers
Each custom kernel repeated the same host-side creation and launch
boilerplate plus a CUDA-then-Metal-then-graph dispatch at every call
site. gpuKernel declares the sources (either backend may be absent) and
a graph fallback; run executes the first that works.
2026-07-31 10:18:54 -07:00
Jesse Gross bd3f22e2f7 qwen3_5: pack GDN input projections into one layout at load
Split checkpoints ran four input projections per recurrent layer, and
native combined checkpoints paid a per-forward slice-and-concat to
rebuild the contiguous qkv rows the causal conv consumes. Normalize
both at load to packed [q|k|v|z] and [beta|alpha] rows: split tensors
concatenate, native interleaved tensors permute once. The forward keeps
a single projection path, and the packed rows are the layout a fused
scan can consume directly.

Pairs with mismatched quantization dequantize before packing rather
than keeping a split fallback path alive.
2026-07-31 10:18:54 -07:00
Jesse Gross 5db07cad71 mlx: apply global scales in Dequantize
The C-level dequantize accepts a global_scale argument but rejects it
on the Metal backend, so dequantize-fallback sites hand-rolled the
same post-multiply. Take the scale in the Go wrapper and apply it on
top of the op, cast back to the output dtype. The quantized embedding
passes its scale; laguna's expert paths keep their own multiplies,
which shape per-expert scales and differ on result dtype.
2026-07-31 10:18:54 -07:00
Jesse Gross acf96e7ab7 mlx: read scalar items at the array's element width
mlx item<T> reinterprets without checking the dtype, so Array.Int's
8-byte read of int32 scalars took in neighboring pool bytes — masked by
Metal's zeroed allocations, corrupting token IDs on CUDA's warm pool.
Read at the element's width.
2026-07-31 10:18:54 -07:00
Daniel Hiltgen a199313eb3 mlx update (#17476) 2026-07-30 10:16:29 -07:00
Daniel Hiltgen b205993ed4 CI: enable lint on the whole tree (#17457)
golangci-lint ran with only-new-issues, which filters findings down to the
lines a PR adds. That silently drops any issue a diff introduces at a
distance, where the report anchors to a line the diff never touched.
CI now is enabled to scan all files.  This PR also fixes the last few
straggler lint glitches outside of integration, which I'll tackle
in a follow up PR.
2026-07-29 16:25:38 -07:00
Daniel Hiltgen 9ea503f505 lint: clean up current tree (#17456) 2026-07-29 15:33:28 -07:00
Jesse Gross 3ff2dcb649 mlxrunner: count every speculative round and log stats at info
The per-request stats are the main diagnostic for speculative
throughput, so log them at info; the controller line stays debug.
Recording chosen depths at the next beginRound dropped rounds with no
successor, so record at endRound and count resume as a depth-0 round.
2026-07-29 14:55:19 -07:00
Jeffrey Morgan 4713800b08 imagegen: remove MLX image generation code (#16615)
Remove the x/imagegen tree (MLX image generation engine, Flux2/zimage
models, cache, C bindings) and all imagegen integration points:

- server: drop imagegen routes, scheduling, and generate handling
- api/cmd/docs: remove image generation API surface and docs
- middleware/openai: remove image endpoint support
- integration: remove imagegen test suites
- x/create: adopt the rewritten create pipeline from main; drop
  imagegen create path (CreateImageGenModel, IsTensorModelDir,
  model_index.json detection, Flux2KleinPipeline vision hack)
- retain x/imagegen/manifest (Ollama-store safetensors manifest
  loader), still used by x/mlxrunner and x/create/client
- fix Windows MLX dl.dll install, MLX CMake version path, and the
  show command after removing safetensors models
2026-07-28 15:35:28 -07:00
Parth Sareen 0e2e34aa86 cmd/tui: improve prompt debug rendering (#17334) 2026-07-27 17:29:18 -07:00
Parth Sareen 76929b0a8a agent: accept file mentions on Enter (#17384) 2026-07-27 17:28:55 -07:00
Parth Sareen bf7be180e3 tui: avoid table detection for pipe prose (#17424) 2026-07-27 13:42:08 -07:00
Daniel Hiltgen eec8e0b945 ci: on release builds dont fail fast (#17413)
If we have one flake, don't stop other jobs that will most likely work so when
we re-run failed jobs, only the flake and dependents need to be run.  This should
help reduce the time it takes to get past a flake and finish a release build.
2026-07-27 08:01:04 -07:00
Daniel Hiltgen be7572e2cf mlx update (#17397) 2026-07-26 16:59:47 -07:00
Daniel Hiltgen 64ee2f9847 model: add Laguna MLX support (#17237)
* model: add Laguna MLX support

Add Laguna XS 2, XS 2.1, and S 2.1 support to the MLX model and create paths.

Read the source config to apply one quantization policy across dense and routed MoE layers. Keep the tied output head and router at source precision, quantize supported attention and expert projections, selectively promote sensitive expert down projections, and emit per-tensor metadata for mixed quantization blobs.

Correct dense expert loading, BF16 source-layout handling, expert global-scale shapes and dtypes, routing-score scaling, and mixed-precision expert dispatch. Gate/up and down projections select quantized or dense execution independently so promoted BF16 down projections do not force quantized gate/up weights through the dense fallback.

Optimize the forward pass with compatible gate/up fusion, sorted standard GatherMM and GatherQMM operations for larger prefills, model-local mlx.Compile closures for elementwise MoE work, and cache-backed 512-token prefill chunks. This keeps the implementation on maintained MLX operations without custom kernels.

Add focused tests for Laguna configuration variants, quantization policy and metadata, dense and routed expert loading, mixed-precision dispatch, compiled-versus-eager parity, fused projections, routing, and prefill chunking.

* review comments and S 2.1 performance fixes

Address renderer/parser selection and mixed-precision expert quantization review feedback.

Keep Laguna weights resident on Metal to prevent repeated paging of its large, sparsely accessed expert buffers. Scope this policy to Laguna GPU execution.

Remove obsolete 512-token prefill chunking now that the runner's 2048-token path is faster.

* review comments addressed

* fix create
2026-07-24 18:24:53 -07:00
Jesse Gross 132e0ca25d x/create: quantize a draft model's output head at the requested type
Draft token embeddings were kept at source precision. A draft that
reuses its embedding as the output projection (the gemma4 assistant)
then reads the whole 537MB bf16 tensor on every draft step — about half
the step's cost. Draft quality only affects how many drafts are
accepted, so the output head now takes the requested type instead of the
8-bit type that protects a target's output quality.

gemma4:26b-mlx, M5 Max: MTP code decode 148 -> 157 tok/s (+26% -> +37%
over plain); prose goes from roughly zero to +2-5%; acceptance unchanged.
2026-07-24 17:26:58 -07:00
Daniel Hiltgen 9eef4a7195 mlx: keep loaded model memory resident (#17367)
Configure Metal residency after the MLX runner materializes model weights.

Wire up to the smaller of active model memory and the recommended working set, leaving pageable headroom for KV caches and request allocations. If residency setup fails, warn and continue with pageable memory.

Expose recoverable MLX C API errors and verify that an oversized wired limit preserves the previous state and leaves subsequent evaluation usable.
2026-07-24 15:34:32 -07:00
Parth Sareen 3f07e022ac cmd/tui: agent system prompt command (#17296) 2026-07-24 14:44:25 -07:00
Parth Sareen 551809688b agent: permission skill loading (#17304) 2026-07-24 14:37:21 -07:00
Jesse Gross 08edcb8f2c qwen3_5: gather packed gate_up experts in one launch
Gathering gate and up separately cost a third expert gather per MoE
layer. Keep gate_up packed as one tensor, joining it at load when the
checkpoint ships the halves separately, and split the gather's output
instead.

Output is byte-identical; decode is 4% faster (7.89 -> 7.58 ms/token on
M5 Max) and prefill 9% faster.
2026-07-24 14:34:56 -07:00
Jesse Gross d6f69da04d qwen3_5: decode each expert tensor with its own quantization format
The expert matmuls decoded with the model-wide format, so models whose
experts are quantized differently from the rest of the weights could
not run.
2026-07-24 14:34:56 -07:00
Daniel Hiltgen 6cd40001a9 server: fix ps data race on scheduler loaded map (#17376)
PsHandler iterated sched.loaded without holding loadedMu, racing with
scheduler goroutines that mutate the map. It also read runnerRef fields
(model, llama, expiresAt) that unload() and the expiration path mutate
under refMu, so a concurrently unloading runner could nil model out from
under the handler.

Instead of adding locking in routes.go, give the scheduler a small
snapshot API: loadedModels() copies the runner list under loadedMu, then
captures each runner's reporting fields under its refMu, respecting the
refMu-before-loadedMu lock ordering used by the expiration path. The
zero-expiresAt estimate for still-loading models moves into the
scheduler too, since it exists because of scheduler behavior.

Also remove the dead code Scheduler.GetRunner
2026-07-24 13:23:49 -07:00
Daniel Hiltgen a84b315e7b test: harden flaky updater and transfer unit tests (#17378)
app/updater: TestBackgoundChecker / TestAutoUpdateDisabledSkipsDownload hit 'TempDir RemoveAll cleanup: directory not empty' on macOS because the background checker goroutine keeps writing staged files into UpdateStageDir while t.TempDir cleanup runs. The checker's context is cancelled by the time cleanup runs, and after cancellation a new download cannot reach the filesystem (DownloadNewRelease aborts at its HEAD request before any write), so it suffices to wait for any in-flight download to drain. Add a test-only waitDownloadIdle helper (polls the existing cancelDownload sentinel under its lock) and register it via t.Cleanup so TempDir cleanup runs after staged-file handles close. No production code changes.

x/transfer: TestDownloadParallelism asserted elapsed <= 1s against 50ms-per-blob delays, too tight for Windows hosted runners' ~15ms timer granularity and shared-runner jitter. Each blob costs two server sleeps (resolve GET + body GET), so model the serial baseline from the deterministic request count, raise per-blob latency to 100ms so timer quantization is a small fraction of each delay, and key the budget to 75% of the serial baseline so the check still proves parallelism while tolerating jitter.
2026-07-24 13:23:30 -07:00
Jesse Gross 83d4311ffe x/create: quantize lm_head at 8-bit in the requested family
The lm_head rule was asymmetric: the fp modes kept an untied head at
source precision (even under mxfp8, leaving it the only bf16 matmul in
the model), while int4 quantized it at 4 bits with no promotion. The
tied-embedding overrides (gemma4, cohere2moe) already resolve the head
to the 8-bit family type and hold quality close to bf16.

Apply the same decision to untied heads: the 8-bit type in the
requested family when it fits the shape, source precision otherwise.
int4 now promotes the head to int8, and the fp modes quantize it to
mxfp8 instead of keeping bf16.
2026-07-23 17:46:09 -07:00
Parth Sareen fce745fe5e agent: import skills from coding agents (#17294) 2026-07-22 22:40:25 -07:00
Daniel Hiltgen 1fd1ccf7ad model: align Laguna with upstream llama.cpp (#17335)
Update llama.cpp to pick up upstream Laguna implementation and remove Ollama's local Laguna implementation. Retain a narrow Metal-only scaling workaround for routed-MoE prompt overflow.

Translate older Ollama GGUF attention-gate and SWA metadata names so existing models continue to load.
2026-07-22 17:09:18 -07:00
Michael Yang efb7e3c55e docs: update retirements (#17289) 2026-07-22 14:24:11 -07:00
Daniel Hiltgen b517b9bd01 model/parsers: finalize incomplete GLM tool calls (#17250)
The GLM parser buffered tool calls until it observed </tool_call>, but ignored the terminal done signal. If the model omitted or partially emitted the outer closing tag, Ollama returned a successful empty response instead of a tool call or an actionable error, leaving coding agents unable to continue.

On end-of-stream, finalize only structurally complete calls for declared tools with all required arguments. Complete calls missing only the outer delimiter now proceed through the existing parser, while genuinely truncated calls return an explicit error rather than being silently dropped.

Fixes #16497
2026-07-22 13:53:47 -07:00
Daniel Hiltgen 479664e7aa mlx update (#17332) 2026-07-22 13:36:49 -07:00
Daniel Hiltgen a51df81573 test: revamp integration test entrpoints (#16560)
This refactors the existing integration tests into 3 priumary groups: fast,
release, and library.  It also refines some of the release tests to drop some
of the older models and pick up newer models, while retaining the broad
coverage in the library group.
2026-07-21 16:06:38 -07:00
Daniel Hiltgen a18c230189 model: add Laguna v8 chat support and fix Metal inference (#17291)
Add a laguna-v8 renderer/parser matching the Laguna XS 2.1 template, and fix v2 handling of embedded thinking and structured tool arguments.

Prevent FP16 overflow in Metal's quantized routed-MoE prefill path by scaling the linear branch and folding the inverse into the routing scale. Other backends and token-generation paths are unchanged.

Add comprehensive v2/v8 Jinja parity and parser tests.
2026-07-21 16:06:29 -07:00
Daniel Hiltgen e21d5327b0 CI: fix missing CUDA v13.4 sub-package (#17288)
Needed for cross-compiling WoA
2026-07-21 12:25:10 -07:00
Jhye 4d1b53e6fb server: detect download stalls before the first byte (#17259)
* server: detect download stalls before the first byte

* server: keep stall timeout out of download API
2026-07-21 11:28:19 -07:00
Daniel Hiltgen 6100aca085 win: support CUDA on Windows ARM64 (#16931) 2026-07-21 10:53:30 -07:00
Daniel Hiltgen 72116bafb3 llama: enable dio on linux CUDA/ROCm iGPUs (#17286)
Avoid double memory consumption by enabling direct IO for iGPUs
2026-07-21 10:53:08 -07:00
Patrick Devine e2c2edcc27 docs: add renderer/parser fields to the API docs (#17275) 2026-07-20 16:22:46 -07:00
Daniel Hiltgen de1ce45913 cuda: add CC 10.0 for linux in CUDA v12 (#17025)
Add compute capability 10.0 to the Linux CUDA v12 preset so B200-class devices can use the cuda_v12 backend with drivers that do not meet the CUDA v13 minimum.

Fixes #12583
2026-07-20 13:09:36 -07:00
Daniel Hiltgen 51fc00122b build: bump Linux toolchain to GCC 13 (#17244)
GCC 11 builds broken AMX code which causes the Sapphire Rapids CPU backend to crash.

Fixes #17006
Fixes #17205
2026-07-20 11:54:39 -07:00
Daniel Hiltgen 445284b428 MLX update (#17189) 2026-07-20 11:54:24 -07:00
Parth Sareen e8f7c93a0b launch: update Hermes integration (#17202) 2026-07-20 11:28:01 -07:00
Parth Sareen 0de38190d7 cmd/tui/chat: render bold emphasis consistently across markdown (#17224) 2026-07-20 11:25:43 -07:00
Parth Sareen 681dfaedcc cmd: remove standalone agent command (#17229) 2026-07-20 11:25:31 -07:00
Parth Sareen 9893d39218 cmd: complete slash commands before submitting (#17230) 2026-07-20 11:25:12 -07:00
Parth Sareen 5ba17e6fdf agent/tui: remove redundant context-window refreshes from event loop (#17241) 2026-07-20 11:25:01 -07:00
Parth Sareen 6f3b997dec cmd: route root command server start through checkServerHeartbeat (#17245)
The bare `ollama` command (and `ollama launch` with no integration) used a
bespoke `ensureServerRunning` that forked `ollama serve` directly and polled
its heartbeat forever (no timeout, no platform-aware launch). Every other
subcommand (`ollama run`, `ollama pull`, `ollama launch <integration>`, ...)
goes through `checkServerHeartbeat` -> `startApp`, so the root command behaved
differently and could hang indefinitely.

Route `runInteractiveTUI` through `checkServerHeartbeat(cmd, nil)` — the same
path `ollama launch <thing>` uses — so the root command is consistent and no
longer runs an unbounded server-spawn loop. `ensureServerRunning` and its
`backgroundServerSysProcAttr` helpers (only it referenced them) are removed,
along with the now-unused `os/exec` import.

The platform `startApp`/`waitForServer` paths are unchanged, so behavior on
macOS/Windows is identical to the other subcommands, and on Linux the root
command now errors the same way the subcommands already do when no server is
running.
2026-07-20 11:24:25 -07:00
Daniel Hiltgen cc62676656 llama.cpp update (#17186) 2026-07-20 11:21:09 -07:00
Parth Sareen 573386c35e agent: skills system (#17203) 2026-07-17 10:32:22 -07:00
Eva H 794a254111 anthropic: close text block before starting thinking block (#17225) 2026-07-17 10:24:38 -04:00
Parth Sareen 714b6fc2a4 agent: allow unlimited tool rounds for cloud models by default (#17217) 2026-07-16 19:07:01 -07:00
Parth Sareen 61e1b1ba5e agent: clean up semantics, UX, DX, and procedural code (#17212) 2026-07-16 19:06:06 -07:00
Parth Sareen 5865a01e48 agent: reorder working directory instruction (#17228) 2026-07-16 19:05:30 -07:00
Parth Sareen e61c1c73fe cmd: remove dead agent prompt wrappers (#17227) 2026-07-16 17:16:14 -07:00
Eva H 03d61e1925 launch: keep Claude Code channels available (#17210) 2026-07-16 13:08:50 -04:00
Parth Sareen 30c390384e cmd: put current working dir in the system prompt (#17188) 2026-07-15 12:10:28 -07:00
Parth Sareen d590830091 agent/tools: isolate web tests from cloud policy (#17208) 2026-07-15 12:07:05 -07:00
Parth Sareen fdcf9efafd fix launch model picker recovery (#17170) 2026-07-15 11:31:22 -07:00
Parth Sareen 76188f60cd docs: add VS Code extension setup (#17158) 2026-07-15 11:30:43 -07:00
Daniel Hiltgen 8a0016f826 model: align gemma4 chat template handling (#17182)
Incorporate the upstream Gemma4 chat template refinements for tool-calling stability, turn closure, and multi-turn reasoning. This updates the native renderer and checked-in HF template fixtures to keep adjacent assistant/tool continuations in the same model turn, add the post-tool thought-channel cue when thinking is enabled, and match Google's default of not replaying historical thinking before a later user turn.

Also preserve null tool arguments through Gemma4 rendering/parsing and extend the Jinja2 parity coverage for these upstream behaviors.
2026-07-14 15:42:04 -07:00
Michael Yang d49b96d9ab docs: collapsed previous retirements (#17167) 2026-07-14 14:00:10 -07:00
Parth Sareen 3bd506bd1c agent/tools: surface actionable web auth error (#17169) 2026-07-14 13:35:16 -07:00
Jesse Gross 123b1f2479 mlxrunner: raise the MTP pending-flush cap to 256 tokens
Per-token cost of the batched head forward keeps falling until the
flush is large enough to reach the fastest kernels: NAX matmul tiles
for dense heads, and the segmented gather path for MoE heads, which
needs tokens*topK/experts >= 4. Measured across the qwen3.6 heads,
256 is the smallest cap past every threshold and within a few percent
of each head's per-token floor. The cost is bounded: up to 2.5 MiB of
pinned hiddens per request and a flush stall under one decode step.
2026-07-14 10:32:04 -07:00
Jesse Gross 556245843a mlxrunner: key the cache trie by token pairs for draft caches
A draft cache pairs each slot with the token that follows it, so the
deepest stored pair always names one token past what a prefix match
can verify - at generation end, the sampled-but-never-committed
final token. Restoring at the match point reuses that pair blind: a
stop token stripped from the next prompt, or any divergence at the
boundary, leaves it stale, and pairing never rewrites below the
resume position, quietly lowering draft acceptance.

Key the trie by token pairs instead: the key for offset i packs
(token i, token i+1), so matching k keys verifies k+1 tokens and
every match is a valid restore point. A pair is reused only if the
token it names matched, and prefill re-evaluates the boundary token,
rebuilding its pair with the token that actually follows. A token
gets a key only once its successor is recorded, so endings record the
final sampled token - never forwarded - and the trie stays level with
the caches. Without a look-ahead the keys are the tokens and behavior
is unchanged.

The recorded tokens' slice bounds used to reject state past them for
free; close now checks the invariant against the stored keys
directly. The test harness rests requests the way the pipeline does -
the deepest recorded token never enters the caches.
2026-07-14 10:32:04 -07:00
Jesse Gross c963822dca mlxrunner: construct per-model state at load
The caches, the speculation binding, and the drafter were each built
lazily inside the first request: begin constructed the caches, and open
bound the cache partition and made a fresh drafter every time. All of
it is a property of the loaded model, so build it once at load.
newPrefixCache replaces the lazy construction in begin, speculation
binds when it is created, and the drafter splits the way speculation
does: a persistent mtpDrafter constructed at load opens each request's
mtpDraftSession, whose constructor syncs the pairing cursor to the
draft caches' restored offset.
2026-07-14 10:32:04 -07:00
Jesse Gross dd49563d55 mlxrunner: rename kvCache to prefixCache
The type coordinates every cache kind — KV, sliding-window, recurrent —
around prefix matching over the trie, so kv was a misnomer.
2026-07-14 10:32:04 -07:00
Jesse Gross 4e96f4dbf2 cache: stop recurrent conv state from pinning the forward buffer
Keeping the recurrent conv state small was handled unevenly: the committed
live state was recopied on every commit — wasted work on single-token
decode, where the window is already tiny — while boundary states captured as
snapshots could still be plain slices of the forward-sized convolution
buffer. A cached slice pins that whole buffer even though the trie's eviction
accounting only counts the slice's bytes, so recurrent cache memory piled up
across requests and eviction could never reclaim it.

Compact each boundary state to its real size once, where it is produced in
the conv wrapper, so live state and snapshots own only their own bytes and
eviction sees the true cost. Single-token decode leaves the already-tiny
window as a slice.

Fixes #16698
2026-07-14 10:32:04 -07:00
Jesse Gross d573a2367b nn/recurrent: derive conv boundary states from a single conv pass
The MTP validation forward schedules a snapshot at every drafted token, which
made CausalConv1D re-run the depthwise conv once per segment to recover each
boundary's conv tail. A conv boundary state is just the trailing convTail
input positions, so run the conv once over the whole window and slice each
boundary tail from the shared buffer, removing the per-token conv launches.
2026-07-14 10:32:04 -07:00
frob 4f7786d0ba mlx: configurable model load timeout (#14796) 2026-07-13 16:06:29 -07:00
Parth Sareen f1a0ffd621 launch: rename Codex App integration to ChatGPT (#17161) 2026-07-13 14:34:17 -07:00
Parth Sareen cd600e19a3 cmd/tui: simplify integration selection and update menu description (#17159) 2026-07-13 12:52:40 -07:00
Daniel Hiltgen 59bd0b49bb mlx: restore NAX in Metal v4 builds (#17160)
MLX now requires a macOS 26.2 deployment target for NAX kernels. Ollama's Metal v4 build still targeted 26.0, so recent MLX bumps silently built mlx_metal_v4 without NAX kernels.
2026-07-13 12:50:38 -07:00
Parth Sareen 82f905cd9c cmd: agent UI (#17017) 2026-07-09 17:27:31 -07:00
Parth Sareen cb3d98ccb2 launch: warn before launching old agent models (#17063) 2026-07-09 16:55:10 -07:00
Jesse Gross d47859ce49 create: select the qwen3.5 parser and renderer for Qwen3.5/Next
Qwen3.5/Qwen3-Next architecture strings contain the substring "qwen3", so the
broad qwen3 match claimed them for the generic parser and qwen3-coder
renderer, whose template doesn't frame the thinking block — an empty
<think></think> leaked into content and think=false was ignored. Match the
family first via isQwen35Family so the parser, renderer, and
thinking-capability checks share one variant list.
2026-07-08 11:12:48 -07:00
Daniel Hiltgen a6293eb516 llm: allow iGPU mmproj offload with fit padding (#16996)
* llm: allow iGPU mmproj offload with fit padding

llama.cpp's fit pass sizes text-model placement before the multimodal projector is loaded. Ollama had been avoiding that risk on non-Metal iGPUs by disabling projector offload entirely, which forces CLIP onto CPU on GB10 and Strix Halo even when the projector has ample memory available.

Let integrated GPUs use the same projector-memory check as other GPUs. When projector offload is enabled, add the estimated projector memory plus the existing 1 GiB headroom to Ollama-owned LLAMA_ARG_FIT_TARGET so fit leaves space for the later projector allocation. If Ollama/device setup already supplied a fit target, add the projector pad to it. If the user set LLAMA_ARG_FIT_TARGET explicitly, leave it exactly as provided.

Fixes #16419

* review comments
2026-07-07 15:28:42 -07:00
Arkadeep Dutta 892e7f6be6 server: apply format constraint for all thinking parsers when think=false (#15901) 2026-07-07 11:54:50 -07:00
Patrick Devine f3d69a3dee server: remove unused internal/ code (#17071) 2026-07-07 11:44:38 -07:00
Daniel Hiltgen 67b6a1c2d4 create: harden GGUF create flows (#17062)
* create: harden GGUF create flows

* lint
2026-07-06 16:20:20 -07:00
Parth Sareen 87b64213b4 launch: disable claude code telemetry by default (#17061) 2026-07-06 15:24:11 -07:00
Daniel Hiltgen f2d069f6df mlx: update to de7b4ed9 (#17056) 2026-07-06 13:31:22 -07:00
Michael Yang 5208ae7500 server: remove OLLAMA_EXPERIMENT=client2 (#16962) 2026-07-06 13:15:39 -07:00
Daniel Hiltgen 9d779572a7 llama.cpp update (#17055)
Bump to b9888.
2026-07-06 12:52:15 -07:00
Patrick Devine 964ea42c09 mlx: x/create rewrite (#16919)
This is a rewrite of the create functionality for the MLX engine.

The core idea behind the create functionality is to break the import/convert into a pipeline of distinct phases:

* Read (scan the safetensors directory for the various bits of metadata)
* Classify (determine what the import type)
* Plan (determine any transforms that need to be done)
* Write (transform any data as necessary and write out the blobs)
* Create the manifest

Each architecture has a "policy" which determines how to convert the model correctly. A number of different formats for safetensors are supported including:

* nvfp4 (two formats: model optimized, torch)
* fp8 datatypes (convert to mxfp8)
* standard bf16 based weights

A number of cleanups/simplifications have been done including:

* using the baked in names for the tensors instead of munging them into something else
* unified 3d expert tensors (instead of separate per expert tensors)
* fewer unnecessary transforms to the various tensors in a model (keep a model as close to the source as possible)
* unified capability checking
* draft model handling (for MTP) is done on the same path

Image generation has been intentionally removed.
2026-07-03 18:30:45 -07:00
Daniel Hiltgen dba1e27fa8 llama: enable FA on CUDA CC 6.x GPUs (#16994)
Recent upstream Pascal kernel fixes let us compile native SM60/SM61 kernels again instead of relying on PTX JIT, so allow Flash Attention auto at runtime for CC 6.x devices.

Fixes #16591

Fixes #16754
2026-07-02 17:11:39 -07:00
Daniel Hiltgen e436db25ff compat: use UTF-8-safe file open (#16999)
Use ggml_fopen for compat tensor reads so Windows paths with Unicode characters are converted through the same UTF-8-to-wide path as llama.cpp model loading.

Fixes #16493
2026-07-02 16:59:23 -07:00
Daniel Hiltgen 26acfa42b5 rocm: remove no longer supported devices (#17010)
The presets and docs had fallen out of sync with what our current ROCm versions on Linux and Windows actually support.  We rely on Vulkan now to cover these older unsupported devices.
2026-07-02 16:59:01 -07:00
Daniel Hiltgen 7b22ac9683 llama: clean up dead code from llama-server work (#17007)
These pieces were missed in the final merge of llama-server and are dead code.
2026-07-02 12:51:54 -07:00
Parth Sareen a2b3a5e9a3 agent: harness core (#16963) 2026-07-02 11:44:31 -07:00
Kevin Park 624cada952 discover: fall back to standard CUDA when the JetPack runner is absent (#16949)
* discover: use the SBSA CUDA build on JetPack 7 (L4T r38+)

JetPack 7 supports SBSA-based CUDA, so the standard cuda_v13 build — shipped
in the base linux-arm64 package, and given the Orin arch (CC 8.7) in #16628
— runs on these devices.

JETSON_JETPACK=7 previously selected a nonexistent jetpack7 runner, so
runner.go skipped every CUDA library and discovery fell back to CPU. The L4T
releases JetPack 7 uses (r38 on Thor, r39 on Orin) also hit the unrecognized
branch, and install.sh warned the version was unsupported. Map JetPack 7+
(L4T r38 and newer) to cuda_v13 (returned as "" from cudaJetpack); no
Jetson-specific download is needed, so install.sh no longer warns.

Fixes #16602

* discover: fall back to standard CUDA when the JetPack runner is absent

Per review, drop the L4T-version mapping (in cudaJetpack and install.sh) and
instead clear the jetpack override in runner.go when the detected cuda_jetpack
runner isn't installed. Normal discovery then selects the standard cuda_v13
build, which supports Orin (CC 8.7) on JetPack 7.
2026-07-02 08:34:35 -07:00
Michael Yang cecd265d3a docs(cloud): update retirement list (#17000) 2026-07-01 19:43:14 -07:00
Mark Ward 2ea95fb059 fix cuda toolkit lookup and parallel (#16613)
* fix cuda toolkit lookup and parallel

* support user override first

* enable control over the nested parallel count
2026-06-30 10:56:54 -07:00
Daniel Hiltgen 8e7be3aed1 ci: avoid unbounded parallelism (#16966)
build-darwin has gotten very slow in the past few releases, most likely due to unbounded parallelism in the MLX build causing the builder to thrash
2026-06-30 10:49:55 -07:00
Patrick Devine 710292ff4f mlx: tighten up gemma4 moe loading code (#16964)
This change allows .experts.gate_proj / .up_proj / .down_proj tensor names to each
be used for both quantized (i.e. nvfp4 and mxfp8) and non-quantized (bf16) models.
Previous to this only non-quantized models used that tensor naming scheme.
2026-06-29 21:15:08 -07:00
Bruce MacDonald ada1eb5163 launch: check for min version for hermes desktop (#16912) 2026-06-29 11:50:11 -07:00
Daniel Hiltgen 1c5ebbf5f4 llama.cpp update (#16960) 2026-06-29 09:43:41 -07:00
Daniel Hiltgen 7926b99e0e mlx: bump dependency (#16935)
Update MLX to 548dd80.

Fix direct MLX tests to run on pinned MLX threads so test execution matches the runner's MLX thread-affinity model.
2026-06-29 09:39:11 -07:00
Aditya Aggarwal 32a97b7493 tools: ignore braces inside JSON strings when detecting tool call end (#16937)
Parser.done() counted the tag's open/close characters ({}, []) without
tracking JSON string context, so a streamed tool call whose string
argument value contained a closing brace or bracket (e.g.
{"code": "if (x) { y }"}) was treated as complete too early and flushed to
the user as plain text instead of being parsed as a tool call.

findArguments() in the same file already tracks string context; apply the
same handling in done() so open/close characters inside string values are
ignored.
2026-06-27 12:00:55 -07:00
Daniel Hiltgen d26a58557d MLX: wire up scheduler selected context size for ps (#16918)
In the PS output, expose the scheduler selected size (clamped by model context size) instead of always reporting the model max context.  This will help provide a hint to clients to keep the context size below this value to avoid paging and poor performance on smaller VRAM systems.
2026-06-26 08:47:03 -07:00
Parth Sareen 2e474c98f9 parser/renderer: add Ornith 9B renderer/parser support (#16920) 2026-06-25 23:18:47 -07:00
Bruce MacDonald 2cb2c5381f launch: update hermes install urls to official (#16913) 2026-06-25 16:22:19 -07:00
Eva H 2a6b50421a fix capability grid dark mode style (#16907) 2026-06-25 13:55:39 -04:00
Daniel Hiltgen f22ec2ec49 CUDA: require driver 550 or newer for v12 (#16895)
Our cuda_v12 build requires nvcc fatbin compression, which in turn requires driver 550 or newer.  This change filters incompatible CUDA devices based on the runtime and driver version.  This allows users to build from source with older toolkits to support older drivers.

Fixes #16449
2026-06-25 08:46:00 -07:00
Eva H d9075caf1a docs: redesign coding integration docs (#16808) 2026-06-25 10:03:59 -04:00
Daniel Hiltgen e11eeb3ba0 llama.cpp version update (#16548) 2026-06-24 14:03:12 -07:00
Daniel Hiltgen 0a408b2225 jetson: add CC 87 for CUDA v13 (#16628)
The new Jetpack 7.2 supports SBSA based CUDA, so we can add the architecture now.
2026-06-24 14:02:41 -07:00
Daniel Hiltgen 16739dee60 server: align generate with native chat templates (#16878)
* server: align generate with native chat templates

/api/generate rebuilt chat-like prompts through the Go template path even when the model selected its native GGUF Jinja chat template, so the same model rendered differently between generate and chat.

Route chat-like generate requests through the shared native chat preparation path, keep deprecated context and image handling working there, and keep explicit OLLAMA_GO_TEMPLATE overrides intact.

Fixes #16792

* review comments

Fall back to "{{ .Prompt }}" when lacking templates
2026-06-24 13:43:56 -07:00
Eva HandParth Sareen d48d790baf docs: redesign docs landing and integrations overview (#16807)
Co-authored-by: Parth Sareen <parth.sareen@ollama.com>
2026-06-24 16:28:28 -04:00
Philip Sinitsin 0463940334 llm: fix ollama ps double-counting mmap'd weights on partial offload (#16709)
* llm: fix ollama ps double-counting mmap'd weights on partial offload

With mmap enabled, llama-server reports each CPU_Mapped model buffer as the
file-offset span of its CPU-resident tensors. During partial offload that span
covers nearly the whole file because the first and last tensors stay on CPU, so
the parsed buffer sizes count the offloaded weights twice and ollama ps shows
roughly 2x the real size with a false CPU/GPU split. Model weights can never
exceed the model file on disk, so trim the excess over the file size from the
mmap-backed portion when computing MemorySize. This makes the reported size
independent of use_mmap; VRAM accounting and scheduler placement are unchanged.

* llm: exclude repacked model buffers from the mmap overlap trim

The trim that corrects mmap double-counting computed the overlap from all
model buffers, including real copies such as CPU_REPACK. On a CPU-only
repacked model that inflated the excess and trimmed the repack out,
undercounting by the repack size (llama3.2 reported ~1918 MiB instead of
~3218 MiB).

Compute the overlap from file-backed buffers only: mmap views and direct
device copies, whose spans can overlap the file on partial offload.
Repacked or host-pinned CPU copies are separate allocations that never
overlap the on-disk weights, so leave them intact. Adds a CPU_Mapped +
CPU_REPACK regression test and corrects the Metal case to the real total.
2026-06-24 11:43:20 -07:00
Daniel Hiltgen 570679c9e0 mlx: update and fix CUDA JIT packaging (#16871)
Bump MLX to the latest selected upstream ref and update the MLX/imagegen
wrappers and tests for the new API behavior.

Fix the CUDA MLX archive so runtime NVRTC kernels work after deployment:
package CUTE/CUTLASS headers, include the CUDA runtime header closure, and
stage a coherent CUDA-toolkit-matched CCCL tree instead of MLX's fetched CCCL
for CUDA payloads. The previous archive could build successfully but crash at
runtime due to missing or incompatible JIT headers.
2026-06-24 10:36:02 -07:00
Daniel Hiltgen 89a171cc70 llm: use host Vulkan loader on Windows (#16869)
Stop bundling the Vulkan loader and resolve the host runtime for Windows Vulkan discovery and backend dependency loading.

Fixes #16677
2026-06-24 10:35:48 -07:00
Daniel Hiltgen 33878e671a llama: default qwen2.5vl window attention metadata (#16868)
Existing qwen2.5vl GGUFs can contain an empty qwen25vl.vision.fullatt_block_indexes array. The compat layer translated the projector metadata but left clip.vision.n_wa_pattern unset, causing llama-server to fail loading the CLIP model.

Default the runtime compat value to the standard Qwen2.5-VL pattern when the key cannot be derived, and make the converter emit the same default for nil or empty fullatt block metadata.

Fixes #16540
2026-06-24 10:35:29 -07:00
Parth SareenandDaniel Hiltgen c191a145bb llm: preserve generation headroom for shifted prompts (#16856)
---------

Co-authored-by: Daniel Hiltgen <daniel@ollama.com>
2026-06-23 15:29:40 -07:00
Parth Sareen 479e1cf94e docs: document max think level (#16877) 2026-06-23 15:29:15 -07:00
Daniel Hiltgen 836507378b llm: size mmproj offload by projector memory (#16866)
* llm: size mmproj offload by projector memory

Replace the blanket 10 GiB VRAM cutoff with a projector tensor-size estimate plus backend headroom, while preserving the existing CPU-only, partial text offload, shared-memory GPU, and startup OOM retry gates.

This is a stopgap until fit accounts for mmproj memory directly.

The same limited-vram path appears in the qwen3.5 vision hang report: the logs show --no-mmproj-offload on a 7.5 GiB RTX 5050 with about 6.4 GiB free while llama-server estimates the inline mmproj at about 962 MiB.

Fixes #16496

Fixes #16570

* review comments
2026-06-23 13:04:02 -07:00
anishandanish 46bc1bcb4c llama: add sm_86 architecture to cuda_v13_windows preset (#16834)
The llama_cuda_v13_windows preset in llama/server/CMakePresets.json was missing sm_86 and sm_80 architectures, causing RTX 3060 laptop and similar mobile RTX 30-series GPUs to be skipped during runtime GPU detection on Windows with CUDA 13. The Linux preset (llama_cuda_v13_linux) included these architectures as "86-virtual" and "80-virtual", but the Windows preset only had "75-virtual;89-virtual;100-virtual;120-virtual", excluding Ampere mobile GPUs.

Signed-off-by: anish <anishesg@users.noreply.github.com>
Co-authored-by: anish <anishesg@users.noreply.github.com>
2026-06-23 07:35:21 -07:00
Bruce MacDonald 2a8b31531e launch/codex: detect model drift when Codex App UI switches away from Ollama (#16864)
ollama launch codex-app sets root-level model_provider = "ollama-launch-codex-app"
in ~/.codex/config.toml to route requests through the local Ollama server.
In Codex, model_provider is a global config key, there is no per-model provider
in the catalog schema (ModelInfo has no model_provider field), so it applies to
every model, not just Ollama ones.

When a user switches to a built-in OpenAI model (e.g. gpt-5.5) in the Codex App
UI, the UI writes model = "gpt-5.5" to config.toml but does NOT update
model_provider. The root model_provider stays "ollama-launch-codex-app", so the
OpenAI model request goes to http://localhost:11434/v1/responses instead of
OpenAI API, resulting in a 404 ("model gpt-5.5 not found"). The user is
stuck: OpenAI models silently route to localhost until they know to run
"ollama launch codex-app --restore".

Fix: CurrentModel() now verifies the configured model appears as a slug in the
Ollama-managed catalog before reporting the integration as active. When the
model has drifted (user selected a non-Ollama model in the UI), CurrentModel()
returns empty, so the launcher accurately shows the integration as inactive and
the user is directed to restore or re-launch.
2026-06-22 15:38:19 -07:00
Jesse Gross 505e35f2b9 mlxrunner: choose the speculative draft length to maximize throughput
The heuristic schedule grew the draft toward a fixed cap on acceptance alone,
maximizing accepted-tokens-per-step rather than throughput, and on a
steep-forward target it regressed below no speculation. Replace it with an
engine-level controller that drafts the depth maximizing
committed-tokens-per-wallclock from live per-position acceptance and persisted
per-width forward cost, with no draft-length cap; the heuristic schedule and
the OLLAMA_MLX_MTP_* env vars go with it.
2026-06-22 15:25:45 -07:00
Jesse Gross 114875133b mlxrunner: resolve each speculative round in one host sync
Acceptance took two blocking evals per round: one to read the accepted mask,
then a second for the bonus or residual token whose graph needed the
host-known rejection point. Sample the residual at every rejection point in
one batched draw alongside the bonus row, so a single eval covers acceptance
and the next token.
2026-06-22 15:25:45 -07:00
Jesse Gross 42c330283b mlxrunner: run one target forward per MTP decode step
Each speculative round ran the target stack twice — once for the current
token's hidden and base logits, once to validate the drafts — capping
throughput below plain decode. Fuse them into one forward over [current,
draft_0..draft_{N-1}], whose hidden rows already line up with the acceptance
math, so the separate base-logits unembed disappears from the drafted path.
2026-06-22 15:25:45 -07:00
Jesse Gross f93efe2809 mlxrunner: apply in-flight drafts to proposal penalty history
Sampler.Distribution built row i as if draftTokens[:i] were appended, leaving
a single-row proposal call with no draft history, so a drafter skipped the
repeat/presence penalties the target's validation applies and re-proposed
penalized tokens. Align rows with the end of the draft chain instead: the
final row sees every draft token, each earlier row one fewer.
2026-06-22 15:25:45 -07:00
Jesse Gross 28fbbb06d5 mlxrunner: support draft heads that maintain draft caches
Generalize the draft path so a head that maintains a KV cache (EAGLE-style)
and Gemma's read-only single-position assistant both fit one drafter
interface with no per-model branches, and make the committed stream the
drafter's maintenance mechanism — every committed run is reported, the
drafter pairs each draft slot with its look-ahead token and flushes completed
pairs to the draft caches. The draft KV thus stays prefix-cached alongside
the target in every session, drafting or not.
2026-06-22 15:25:45 -07:00
Jesse Gross 340c51bbb7 mlxrunner: host speculative decoding in the text generation pipeline
The pipeline and the MTP decoder each owned a decode loop with duplicated
prefill, budget, and emission handling. Split the pipeline into prefill and
decode phases behind a decoder interface, with the decode loop the sole
emitter enforcing the NumPredict budget, and split speculation into a generic
engine that returns the accepted run and a drafter interface that owns only
how proposals are made.
2026-06-22 15:25:45 -07:00
Jesse Gross 2e9d68dc38 mlxrunner: unify the MTP decode paths
Greedy is a special case of sampled decoding — at temperature 0 the sampler
yields a point mass, so rejection-sampling acceptance reduces to argmax-match
— so collapse the separate greedy, sampled, and serial paths into one. MTP
now honors any temperature, penalty, and top-k/p/min-p setting; logprobs
remain the only gated feature.
2026-06-22 15:25:45 -07:00
Sahil Kadadekar fc58544422 discover: fix inverted iGPU/dGPU Vulkan classification on Windows hybrid graphics (#16669)
On Windows hybrid-graphics systems (Intel iGPU + NVIDIA dGPU), discovery
could classify the integrated GPU as discrete and the discrete GPU as
integrated, dropping the dGPU's Vulkan device and scheduling models onto
the iGPU's shared system RAM (#16667). Two index-keyed correlations
between independently-ordered device enumerations caused this:

1. The native probe's stderr was concatenated into the output passed to
   parseVulkanUMA. The probe enumerates Vulkan devices in its own order,
   so its ggml_vulkan uma lines overwrote llama-server's index-keyed UMA
   map with inverted values. Parse UMA metadata only from llama-server's
   own output.

2. applyWindowsVulkanRefinement required the raw vkEnumeratePhysicalDevices
   count to equal llama-server's Vulkan device count. The raw enumeration
   is a superset on real systems (D3D12 mapping-layer devices, Microsoft
   Basic Render Driver), so the refinement that reads the authoritative
   VkPhysicalDeviceType was always skipped. Match devices by name against
   the probed superset instead, bailing only when a device has no match or
   matches conflicting device types.

Verified on the hardware from #16667 (Intel RaptorLake-S + RTX 4080
Laptop): the raw probe returns 5 devices vs llama-server's 2; with this
change the iGPU is dropped as integrated, the dGPU's Vulkan device
dedupes against CUDA0, and the model loads on the dGPU with no
environment overrides.

Fixes #16667
2026-06-22 14:52:03 -07:00
Eva H e434a93884 launch: auto-install opencode when missing (#16806) 2026-06-19 10:12:11 -07:00
Eva H 9c02d8e69d launch: auto-install Claude Code (#16802) 2026-06-19 10:11:50 -07:00
Eva H 07ed752353 launch: add thinking capability detection to opencode (#15434) 2026-06-18 13:45:16 -04:00
Parth Sareen e1f7f9cbdb ci: pin darwin release xcode (#16788) 2026-06-17 13:01:10 -07:00
595 changed files with 59726 additions and 59303 deletions

No files matched your search

+47 -1
View File
@@ -39,11 +39,27 @@ jobs:
APPLE_ID: ${{ vars.APPLE_ID }}
MACOS_SIGNING_KEY: ${{ secrets.MACOS_SIGNING_KEY }}
MACOS_SIGNING_KEY_PASSWORD: ${{ secrets.MACOS_SIGNING_KEY_PASSWORD }}
DEVELOPER_DIR: /Applications/Xcode_26.4.1.app/Contents/Developer
CGO_CFLAGS: '-mmacosx-version-min=14.0 -O3'
CGO_CXXFLAGS: '-mmacosx-version-min=14.0 -O3'
CGO_LDFLAGS: '-mmacosx-version-min=14.0 -O3'
steps:
- uses: actions/checkout@v4
- name: Select Xcode 26.4.1
shell: bash
run: |
set -euo pipefail
if [ ! -d "${DEVELOPER_DIR}" ]; then
echo "Missing ${DEVELOPER_DIR}"
ls -1 /Applications | grep '^Xcode' || true
exit 1
fi
sudo xcode-select -s "${DEVELOPER_DIR}"
sw_vers
xcodebuild -version
xcrun --sdk macosx --show-sdk-version
xcrun --find metal
- run: |
echo $MACOS_SIGNING_KEY | base64 --decode > certificate.p12
security create-keychain -p password build.keychain
@@ -77,6 +93,7 @@ jobs:
windows-depends:
needs: setup-environment
strategy:
fail-fast: false
matrix:
os: [windows]
arch: [amd64]
@@ -108,6 +125,22 @@ jobs:
- '"nvvm"'
- '"nvptxcompiler"'
cuda-version: '13.0'
- os: windows
arch: amd64
preset: 'CUDA 13 ARM64'
build-steps: cuda13Arm64Cross
install: https://packages.nvidia.com/prerelease/cuda/13.4.0/local_installers/cuda_13.4.0_windows_x86_64.exe
cuda-components:
- '"cudart"'
- '"cudart_cross"'
- '"nvcc"'
- '"nvcc_cross"'
- '"cublas_cross"'
- '"cublas_dev"'
- '"crt"'
- '"nvvm"'
- '"nvptxcompiler"'
cuda-version: '13.4'
- os: windows
arch: amd64
preset: 'ROCm 7'
@@ -182,8 +215,18 @@ jobs:
name: Install CUDA ${{ matrix.cuda-version }}
run: |
$ErrorActionPreference = "Stop"
$ProgressPreference = 'SilentlyContinue'
if ("${{ steps.cache-install.outputs.cache-hit }}" -ne 'true') {
Invoke-WebRequest -Uri "${{ matrix.install }}" -OutFile "install.exe"
for ($attempt = 1; $attempt -le 3; $attempt++) {
try {
Invoke-WebRequest -Uri "${{ matrix.install }}" -OutFile "install.exe"
break
} catch {
if ($attempt -eq 3) { throw }
Write-Host "CUDA installer download attempt $attempt failed: $($_.Exception.Message); retrying in 15s"
Start-Sleep -Seconds 15
}
}
$subpackages = @(${{ join(matrix.cuda-components, ', ') }}) | Foreach-Object {"${_}_${{ matrix.cuda-version }}"}
Start-Process -FilePath .\install.exe -ArgumentList (@("-s") + $subpackages) -NoNewWindow -Wait
}
@@ -418,6 +461,7 @@ jobs:
linux-depends:
strategy:
fail-fast: false
matrix:
include:
- arch: amd64
@@ -499,6 +543,7 @@ jobs:
# and just assembles, runs the Go build, pushes the final image, and extracts release bundles.
docker-build-push:
strategy:
fail-fast: false
matrix:
include:
- os: linux
@@ -649,6 +694,7 @@ jobs:
# Merge Docker images for the same flavor into a single multi-arch manifest
docker-merge-push:
strategy:
fail-fast: false
matrix:
suffix: ['', '-rocm']
runs-on: linux
@@ -1,98 +0,0 @@
name: test-darwin-xcode-pin
on:
workflow_dispatch:
push:
branches:
- test/darwin-xcode-pin
pull_request:
paths:
- '.github/workflows/test-darwin-xcode-pin.yaml'
- 'scripts/build_darwin.sh'
- 'MLX_VERSION'
- 'MLX_C_VERSION'
- 'cmake/**'
- 'x/mlxrunner/**'
permissions:
contents: read
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
env:
CGO_CFLAGS: '-O3'
CGO_CXXFLAGS: '-O3'
PINNED_DEVELOPER_DIR: /Applications/Xcode_26.4.1.app/Contents/Developer
jobs:
darwin-build:
runs-on: macos-26-xlarge
env:
CGO_CFLAGS: '-mmacosx-version-min=14.0 -O3'
CGO_CXXFLAGS: '-mmacosx-version-min=14.0 -O3'
CGO_LDFLAGS: '-mmacosx-version-min=14.0 -O3'
steps:
- uses: actions/checkout@v4
- name: Set build environment
shell: bash
run: |
set -euo pipefail
VERSION="0.0.0-xcode-pin-${GITHUB_SHA::7}"
{
echo "VERSION=${VERSION}"
echo "GOFLAGS='-ldflags=-w -s \"-X=github.com/ollama/ollama/version.Version=${VERSION}\" \"-X=github.com/ollama/ollama/server.mode=release\"'"
} >>"${GITHUB_ENV}"
- name: Select Xcode 26.4.1
shell: bash
run: |
set -euo pipefail
if [ ! -d "${PINNED_DEVELOPER_DIR}" ]; then
echo "Missing ${PINNED_DEVELOPER_DIR}"
ls -1 /Applications | grep '^Xcode' || true
exit 1
fi
sudo xcode-select -s "${PINNED_DEVELOPER_DIR}"
echo "DEVELOPER_DIR=${PINNED_DEVELOPER_DIR}" >>"${GITHUB_ENV}"
sw_vers
xcodebuild -version
xcrun --sdk macosx --show-sdk-version
xcrun --find metal
- uses: actions/setup-go@v5
with:
go-version-file: go.mod
cache-dependency-path: |
go.sum
LLAMA_CPP_VERSION
MLX_VERSION
MLX_C_VERSION
- name: Build unsigned Darwin runtime
run: ./scripts/build_darwin.sh build package
- name: Verify MLX payload
shell: bash
run: |
set -euo pipefail
test -f dist/darwin/lib/ollama/mlx_metal_v3/libmlxc.dylib
test -f dist/darwin/lib/ollama/mlx_metal_v3/mlx.metallib
test -f dist/darwin/lib/ollama/mlx_metal_v4/libmlxc.dylib
test -f dist/darwin/lib/ollama/mlx_metal_v4/mlx.metallib
find dist/darwin/lib/ollama -maxdepth 3 -type f \( -name 'libmlx*.dylib' -o -name '*.metallib' \) -print
lipo -archs dist/darwin/lib/ollama/mlx_metal_v3/libmlxc.dylib
lipo -archs dist/darwin/lib/ollama/mlx_metal_v4/libmlxc.dylib
- name: Log build results
run: ls -l dist/
- uses: actions/upload-artifact@v4
with:
name: ollama-darwin-xcode-pin
path: dist/ollama-darwin.tgz
compression-level: 0
+27 -1
View File
@@ -321,6 +321,22 @@ jobs:
- '"nvvm"'
- '"nvptxcompiler"'
cuda-version: '13.0'
- os: windows
arch: amd64
preset: 'CUDA 13 ARM64'
build-steps: cuda13Arm64Cross
install: https://packages.nvidia.com/prerelease/cuda/13.4.0/local_installers/cuda_13.4.0_windows_x86_64.exe
cuda-components:
- '"cudart"'
- '"cudart_cross"'
- '"nvcc"'
- '"nvcc_cross"'
- '"cublas_cross"'
- '"cublas_dev"'
- '"crt"'
- '"nvvm"'
- '"nvptxcompiler"'
cuda-version: '13.4'
- os: windows
arch: amd64
preset: 'ROCm 7'
@@ -365,8 +381,18 @@ jobs:
name: Install CUDA ${{ matrix.cuda-version }}
run: |
$ErrorActionPreference = "Stop"
$ProgressPreference = 'SilentlyContinue'
if ("${{ steps.cache-install.outputs.cache-hit }}" -ne 'true') {
Invoke-WebRequest -Uri "${{ matrix.install }}" -OutFile "install.exe"
for ($attempt = 1; $attempt -le 3; $attempt++) {
try {
Invoke-WebRequest -Uri "${{ matrix.install }}" -OutFile "install.exe"
break
} catch {
if ($attempt -eq 3) { throw }
Write-Host "CUDA installer download attempt $attempt failed: $($_.Exception.Message); retrying in 15s"
Start-Sleep -Seconds 15
}
}
$subpackages = @(${{ join(matrix.cuda-components, ', ') }}) | Foreach-Object {"${_}_${{ matrix.cuda-version }}"}
Start-Process -FilePath .\install.exe -ArgumentList (@("-s") + $subpackages) -NoNewWindow -Wait
}
-2
View File
@@ -416,5 +416,3 @@ jobs:
run: go test -count=1 -tags updater_live ./app/...
- uses: golangci/golangci-lint-action@v9
with:
only-new-issues: true
+4
View File
@@ -45,6 +45,10 @@ if(APPLE)
set(CMAKE_BUILD_RPATH "@loader_path")
set(CMAKE_INSTALL_RPATH "@loader_path")
set(CMAKE_BUILD_WITH_INSTALL_RPATH ON)
elseif(UNIX)
set(CMAKE_BUILD_RPATH "$ORIGIN")
set(CMAKE_INSTALL_RPATH "$ORIGIN")
set(CMAKE_BUILD_WITH_INSTALL_RPATH ON)
endif()
set(OLLAMA_BUILD_DIR ${CMAKE_BINARY_DIR}/lib/ollama)
+9 -8
View File
@@ -15,9 +15,9 @@ FROM scratch AS local-mlx
FROM scratch AS local-mlx-c
FROM --platform=linux/amd64 rocm/dev-almalinux-8:${ROCMVERSION}-complete AS base-amd64
RUN dnf install -y yum-utils ccache gcc-toolset-11-gcc gcc-toolset-11-gcc-c++ gcc-toolset-11-binutils \
RUN dnf install -y yum-utils ccache gcc-toolset-13-gcc gcc-toolset-13-gcc-c++ gcc-toolset-13-binutils \
&& yum-config-manager --add-repo https://developer.download.nvidia.com/compute/cuda/repos/rhel8/x86_64/cuda-rhel8.repo
ENV PATH=/opt/rh/gcc-toolset-11/root/usr/bin:$PATH
ENV PATH=/opt/rh/gcc-toolset-13/root/usr/bin:$PATH
FROM --platform=linux/arm64 almalinux:8 AS base-arm64
# install epel-release for ccache
@@ -42,8 +42,8 @@ ENV LDFLAGS=-s
#
FROM base AS cpu-deps
RUN dnf install -y gcc-toolset-11-gcc gcc-toolset-11-gcc-c++
ENV PATH=/opt/rh/gcc-toolset-11/root/usr/bin:$PATH
RUN dnf install -y gcc-toolset-13-gcc gcc-toolset-13-gcc-c++
ENV PATH=/opt/rh/gcc-toolset-13/root/usr/bin:$PATH
FROM base AS cuda-12-deps
ARG CUDA12VERSION=12.8
@@ -91,8 +91,8 @@ RUN --mount=type=cache,target=/root/.ccache \
&& for lib in \
/usr/lib64/libgomp.so* \
/usr/lib64/libomp.so* \
/opt/rh/gcc-toolset-11/root/usr/lib64/libgomp.so* \
/opt/rh/gcc-toolset-11/root/usr/lib64/libomp.so*; do \
/opt/rh/gcc-toolset-13/root/usr/lib64/libgomp.so* \
/opt/rh/gcc-toolset-13/root/usr/lib64/libomp.so*; do \
[ -e "$lib" ] && cp -a "$lib" dist/lib/ollama/ || true; \
done
@@ -124,7 +124,7 @@ FROM scratch AS publish-llama-server-cuda_v13
COPY --from=llama-server-cuda_v13 dist/lib/ollama /lib/ollama/
FROM rocm-7-deps AS llama-server-rocm_v7_2
ENV CC=clang CXX=clang++
ENV CC=clang CXX=clang++ CXXFLAGS=--gcc-toolchain=/opt/rh/gcc-toolset-13/root/usr
COPY LLAMA_CPP_VERSION .
COPY llama/server llama/server
COPY llama/compat llama/compat
@@ -213,7 +213,8 @@ ENV CGO_LDFLAGS="-L/usr/local/cuda-13/lib64 -L/usr/local/cuda-13/targets/x86_64-
WORKDIR /go/src/github.com/ollama/ollama
COPY CMakeLists.txt CMakePresets.json .
COPY cmake cmake
COPY x/imagegen/mlx x/imagegen/mlx
COPY mlx mlx
COPY x/mlxrunner/mlx x/mlxrunner/mlx
COPY go.mod go.sum .
COPY MLX_VERSION MLX_C_VERSION .
RUN curl -fsSL https://golang.org/dl/go$(awk '/^go/ { print $2 }' go.mod).linux-$(case $(uname -m) in x86_64) echo amd64 ;; aarch64) echo arm64 ;; esac).tar.gz | tar xz -C /usr/local
+1 -1
View File
@@ -1 +1 @@
b9672
b10488
+1 -1
View File
@@ -1 +1 @@
2165dc08d7b33258260aa849d39f087d50e62962
27fec909a3df9e572f5195607a453e273e7d80d0
+1 -1
View File
@@ -65,7 +65,7 @@ To launch a specific integration:
ollama launch claude
```
Supported integrations include [Claude Code](https://docs.ollama.com/integrations/claude-code), [Codex](https://docs.ollama.com/integrations/codex), [Copilot CLI](https://docs.ollama.com/integrations/copilot-cli), [Droid](https://docs.ollama.com/integrations/droid), and [OpenCode](https://docs.ollama.com/integrations/opencode).
Supported integrations include [Claude Code](https://docs.ollama.com/integrations/claude-code), [Codex](https://docs.ollama.com/integrations/codex), [Copilot CLI](https://docs.ollama.com/integrations/copilot-cli), [DeepSeek Harness](https://docs.ollama.com/integrations/deepseek-harness), [Droid](https://docs.ollama.com/integrations/droid), and [OpenCode](https://docs.ollama.com/integrations/opencode).
### AI assistant
+15 -3
View File
@@ -777,6 +777,18 @@ func (c *StreamConverter) Process(r api.ChatResponse) []StreamEvent {
}
if r.Message.Thinking != "" && !c.thinkingDone {
if c.textStarted {
events = append(events, StreamEvent{
Event: "content_block_stop",
Data: ContentBlockStopEvent{
Type: "content_block_stop",
Index: c.contentIndex,
},
})
c.contentIndex++
c.textStarted = false
}
if !c.thinkingStarted {
c.thinkingStarted = true
events = append(events, StreamEvent{
@@ -1063,7 +1075,7 @@ type CountTokensRequest struct {
// EstimateInputTokens estimates input tokens from a MessagesRequest (reuses CountTokensRequest logic)
func EstimateInputTokens(req MessagesRequest) int {
return estimateTokens(CountTokensRequest{
return EstimateCountTokens(CountTokensRequest{
Model: req.Model,
Messages: req.Messages,
System: req.System,
@@ -1077,10 +1089,10 @@ type CountTokensResponse struct {
InputTokens int `json:"input_tokens"`
}
// estimateTokens returns a rough estimate of tokens (len/4).
// EstimateCountTokens returns a rough estimate of tokens (len/4).
// TODO: Replace with actual tokenization via Tokenize API for accuracy.
// Current len/4 heuristic is a rough approximation (~4 chars/token average).
func estimateTokens(req CountTokensRequest) int {
func EstimateCountTokens(req CountTokensRequest) int {
var totalLen int
// Count system prompt
+56 -5
View File
@@ -3,6 +3,7 @@ package anthropic
import (
"encoding/base64"
"encoding/json"
"fmt"
"strings"
"testing"
@@ -1140,6 +1141,56 @@ func TestStreamConverter_ThinkingDirectlyFollowedByToolCall(t *testing.T) {
}
}
func TestStreamConverter_TextBeforeThinking(t *testing.T) {
conv := NewStreamConverter("msg_123", "test-model", 0)
responses := []api.ChatResponse{
{Message: api.Message{Role: "assistant", Content: "---\n"}},
{Message: api.Message{Role: "assistant", Thinking: "Let me think."}},
{
Message: api.Message{Role: "assistant", Content: "The answer."},
Done: true,
DoneReason: "stop",
Metrics: api.Metrics{PromptEvalCount: 10, EvalCount: 5},
},
}
var got []string
for _, response := range responses {
for _, event := range conv.Process(response) {
switch data := event.Data.(type) {
case ContentBlockStartEvent:
got = append(got, fmt.Sprintf("%s:%s:%d", event.Event, data.ContentBlock.Type, data.Index))
case ContentBlockDeltaEvent:
got = append(got, fmt.Sprintf("%s:%s:%d", event.Event, data.Delta.Type, data.Index))
case ContentBlockStopEvent:
got = append(got, fmt.Sprintf("%s:%d", event.Event, data.Index))
default:
got = append(got, event.Event)
}
}
}
want := []string{
"message_start",
"content_block_start:text:0",
"content_block_delta:text_delta:0",
"content_block_stop:0",
"content_block_start:thinking:1",
"content_block_delta:thinking_delta:1",
"content_block_stop:1",
"content_block_start:text:2",
"content_block_delta:text_delta:2",
"content_block_stop:2",
"message_delta",
"message_stop",
}
if diff := cmp.Diff(want, got); diff != "" {
t.Fatalf("unexpected stream events (-want +got):\n%s", diff)
}
}
func TestStreamConverter_ToolCallWithUnmarshalableArgs(t *testing.T) {
// Test that unmarshalable arguments (like channels) are handled gracefully
// and don't cause a panic or corrupt stream
@@ -1495,7 +1546,7 @@ func TestEstimateTokens_SimpleMessage(t *testing.T) {
},
}
tokens := estimateTokens(req)
tokens := EstimateCountTokens(req)
// "user" (4) + "Hello, world!" (13) = 17 chars / 4 = 4 tokens
if tokens < 1 {
@@ -1516,7 +1567,7 @@ func TestEstimateTokens_WithSystemPrompt(t *testing.T) {
},
}
tokens := estimateTokens(req)
tokens := EstimateCountTokens(req)
// System prompt adds to count
if tokens < 5 {
@@ -1539,7 +1590,7 @@ func TestEstimateTokens_WithTools(t *testing.T) {
},
}
tokens := estimateTokens(req)
tokens := EstimateCountTokens(req)
// Tools add significant content
if tokens < 10 {
@@ -1568,7 +1619,7 @@ func TestEstimateTokens_WithThinking(t *testing.T) {
},
}
tokens := estimateTokens(req)
tokens := EstimateCountTokens(req)
// Thinking content should be counted
if tokens < 10 {
@@ -1582,7 +1633,7 @@ func TestEstimateTokens_EmptyContent(t *testing.T) {
Messages: []MessageParam{},
}
tokens := estimateTokens(req)
tokens := EstimateCountTokens(req)
if tokens != 0 {
t.Errorf("expected 0 tokens for empty content, got %d", tokens)
+20
View File
@@ -473,6 +473,26 @@ func (c *Client) CloudStatusExperimental(ctx context.Context) (*StatusResponse,
return &status, nil
}
// WebSearchExperimental searches the web through the local server's
// experimental web search endpoint.
func (c *Client) WebSearchExperimental(ctx context.Context, req *WebSearchRequest) (*WebSearchResponse, error) {
var resp WebSearchResponse
if err := c.do(ctx, http.MethodPost, "/api/experimental/web_search", req, &resp); err != nil {
return nil, err
}
return &resp, nil
}
// WebFetchExperimental fetches web page content through the local server's
// experimental web fetch endpoint.
func (c *Client) WebFetchExperimental(ctx context.Context, req *WebFetchRequest) (*WebFetchResponse, error) {
var resp WebFetchResponse
if err := c.do(ctx, http.MethodPost, "/api/experimental/web_fetch", req, &resp); err != nil {
return nil, err
}
return &resp, nil
}
// Signout will signout a client for a local ollama server.
func (c *Client) Signout(ctx context.Context) error {
return c.do(ctx, http.MethodPost, "/api/signout", nil, nil)
+135
View File
@@ -2,6 +2,7 @@ package api
import (
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
@@ -351,6 +352,140 @@ func TestClientDo(t *testing.T) {
}
}
func TestClientWebSearchExperimentalUsesLocalRoute(t *testing.T) {
var gotPath string
var gotMethod string
var gotRequest WebSearchRequest
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
gotPath = r.URL.Path
gotMethod = r.Method
if err := json.NewDecoder(r.Body).Decode(&gotRequest); err != nil {
t.Fatal(err)
}
if err := json.NewEncoder(w).Encode(WebSearchResponse{
Results: []WebSearchResult{{Title: "Ollama", URL: "https://ollama.com", Content: "models"}},
}); err != nil {
t.Fatal(err)
}
}))
defer ts.Close()
client := NewClient(&url.URL{Scheme: "http", Host: ts.Listener.Addr().String()}, http.DefaultClient)
resp, err := client.WebSearchExperimental(t.Context(), &WebSearchRequest{Query: "ollama", MaxResults: 3})
if err != nil {
t.Fatal(err)
}
if gotMethod != http.MethodPost {
t.Fatalf("method = %q, want POST", gotMethod)
}
if gotPath != "/api/experimental/web_search" {
t.Fatalf("path = %q, want /api/experimental/web_search", gotPath)
}
if gotRequest.Query != "ollama" || gotRequest.MaxResults != 3 {
t.Fatalf("request = %#v", gotRequest)
}
if len(resp.Results) != 1 || resp.Results[0].Title != "Ollama" {
t.Fatalf("response = %#v", resp)
}
}
func TestClientWebSearchExperimentalErrors(t *testing.T) {
tests := []struct {
name string
status int
body string
assertError func(*testing.T, error)
}{
{
name: "unauthorized retains sign in URL",
status: http.StatusUnauthorized,
body: `{"error":"unauthorized","signin_url":"https://ollama.com/signin/example"}`,
assertError: func(t *testing.T, err error) {
t.Helper()
var authErr AuthorizationError
if !errors.As(err, &authErr) {
t.Fatalf("error = %T, want AuthorizationError", err)
}
if authErr.StatusCode != http.StatusUnauthorized || authErr.SigninURL != "https://ollama.com/signin/example" {
t.Fatalf("authorization error = %#v", authErr)
}
},
},
{
name: "rate limit retains status",
status: http.StatusTooManyRequests,
body: `{"error":"rate limit exceeded"}`,
assertError: func(t *testing.T, err error) {
t.Helper()
var statusErr StatusError
if !errors.As(err, &statusErr) {
t.Fatalf("error = %T, want StatusError", err)
}
if statusErr.StatusCode != http.StatusTooManyRequests || statusErr.ErrorMessage != "rate limit exceeded" {
t.Fatalf("status error = %#v", statusErr)
}
},
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/json")
w.WriteHeader(tt.status)
_, _ = w.Write([]byte(tt.body))
}))
defer ts.Close()
client := NewClient(&url.URL{Scheme: "http", Host: ts.Listener.Addr().String()}, http.DefaultClient)
_, err := client.WebSearchExperimental(t.Context(), &WebSearchRequest{Query: "ollama"})
if err == nil {
t.Fatal("expected error")
}
tt.assertError(t, err)
})
}
}
func TestClientWebFetchExperimentalUsesLocalRoute(t *testing.T) {
var gotPath string
var gotMethod string
var gotRequest WebFetchRequest
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
gotPath = r.URL.Path
gotMethod = r.Method
if err := json.NewDecoder(r.Body).Decode(&gotRequest); err != nil {
t.Fatal(err)
}
if err := json.NewEncoder(w).Encode(WebFetchResponse{
Title: "Ollama",
Content: "models",
Links: []string{"https://ollama.com/library"},
}); err != nil {
t.Fatal(err)
}
}))
defer ts.Close()
client := NewClient(&url.URL{Scheme: "http", Host: ts.Listener.Addr().String()}, http.DefaultClient)
resp, err := client.WebFetchExperimental(t.Context(), &WebFetchRequest{URL: "https://ollama.com"})
if err != nil {
t.Fatal(err)
}
if gotMethod != http.MethodPost {
t.Fatalf("method = %q, want POST", gotMethod)
}
if gotPath != "/api/experimental/web_fetch" {
t.Fatalf("path = %q, want /api/experimental/web_fetch", gotPath)
}
if gotRequest.URL != "https://ollama.com" {
t.Fatalf("request = %#v", gotRequest)
}
if resp.Title != "Ollama" || resp.Content != "models" {
t.Fatalf("response = %#v", resp)
}
}
type roundTripFunc func(*http.Request) (*http.Response, error)
func (f roundTripFunc) RoundTrip(req *http.Request) (*http.Response, error) {
+35 -30
View File
@@ -127,20 +127,6 @@ type GenerateRequest struct {
// each with an associated log probability. Only applies when Logprobs is true.
// Valid values are 0-20. Default is 0 (only return the selected token's logprob).
TopLogprobs int `json:"top_logprobs,omitempty"`
// Experimental: Image generation fields (may change or be removed)
// Width is the width of the generated image in pixels.
// Only used for image generation models.
Width int32 `json:"width,omitempty"`
// Height is the height of the generated image in pixels.
// Only used for image generation models.
Height int32 `json:"height,omitempty"`
// Steps is the number of diffusion steps for image generation.
// Only used for image generation models.
Steps int32 `json:"steps,omitempty"`
}
// ChatRequest describes a request sent by [Client.Chat].
@@ -706,8 +692,11 @@ type CreateRequest struct {
// Messages is a list of messages added to the model before chat and generation requests.
Messages []Message `json:"messages,omitempty"`
// Renderer is the name of the renderer used when constructing a request to the model.
Renderer string `json:"renderer,omitempty"`
Parser string `json:"parser,omitempty"`
// Parser is the name of the parser used to parse the output of the request.
Parser string `json:"parser,omitempty"`
// Requires is the minimum version of Ollama required by the model.
Requires string `json:"requires,omitempty"`
@@ -868,6 +857,36 @@ type StatusResponse struct {
Cloud CloudStatus `json:"cloud"`
}
// WebSearchRequest is the request for [Client.WebSearchExperimental].
type WebSearchRequest struct {
Query string `json:"query"`
MaxResults int `json:"max_results,omitempty"`
}
// WebSearchResult is a single result from [Client.WebSearchExperimental].
type WebSearchResult struct {
Title string `json:"title"`
URL string `json:"url"`
Content string `json:"content"`
}
// WebSearchResponse is the response from [Client.WebSearchExperimental].
type WebSearchResponse struct {
Results []WebSearchResult `json:"results"`
}
// WebFetchRequest is the request for [Client.WebFetchExperimental].
type WebFetchRequest struct {
URL string `json:"url"`
}
// WebFetchResponse is the response from [Client.WebFetchExperimental].
type WebFetchResponse struct {
Title string `json:"title"`
Content string `json:"content"`
Links []string `json:"links,omitempty"`
}
// GenerateResponse is the response passed into [GenerateResponseFunc].
type GenerateResponse struct {
// Model is the model name that generated the response.
@@ -908,20 +927,6 @@ type GenerateResponse struct {
// Logprobs contains log probability information for the generated tokens,
// if requested via the Logprobs parameter.
Logprobs []Logprob `json:"logprobs,omitempty"`
// Experimental: Image generation fields (may change or be removed)
// Image contains a base64-encoded generated image.
// Only present for image generation models.
Image string `json:"image,omitempty"`
// Completed is the number of completed steps in image generation.
// Only present for image generation models during streaming.
Completed int64 `json:"completed,omitempty"`
// Total is the total number of steps for image generation.
// Only present for image generation models during streaming.
Total int64 `json:"total,omitempty"`
}
// ModelDetails provides details about a model.
@@ -1100,7 +1105,7 @@ func DefaultOptions() Options {
TopP: 0.9,
TypicalP: 1.0,
RepeatLastN: 64,
RepeatPenalty: 1.1,
RepeatPenalty: 1.0,
PresencePenalty: 0.0,
FrequencyPenalty: 0.0,
Seed: -1,
+64 -24
View File
@@ -146,15 +146,10 @@ func main() {
// Do this after logging is set up so we can debug issues
if runtime.GOOS == "windows" && urlSchemeRequest != "" {
slog.Debug("checking for existing instance", "url", urlSchemeRequest)
if checkAndHandleExistingInstance(urlSchemeRequest) {
// The function will exit if it successfully sends to another instance
// If we reach here, we're the first/only instance
} else {
// No existing instance found, handle the URL scheme in this instance
go func() {
handleURLSchemeInCurrentInstance(urlSchemeRequest)
}()
}
// This exits after forwarding the request when another instance is
// running. First-instance requests are handled later by osRun, after the
// Windows UI dependencies are initialized and from the primary thread.
checkAndHandleExistingInstance(urlSchemeRequest)
}
// Detect if this is a first start after an upgrade, in
@@ -205,6 +200,12 @@ func main() {
uiServerPort = port
st := &store.Store{}
if devMode {
if dbPath := strings.TrimSpace(os.Getenv("OLLAMA_APP_DB_PATH")); dbPath != "" {
st.DBPath = dbPath
slog.Debug("using development app database", "path", dbPath)
}
}
appStore = st
// Enable CORS in development mode
@@ -324,11 +325,11 @@ func main() {
quit()
}()
if urlSchemeRequest != "" {
if urlSchemeRequest != "" && runtime.GOOS != "windows" {
go func() {
handleURLSchemeInCurrentInstance(urlSchemeRequest)
}()
} else {
} else if urlSchemeRequest == "" {
slog.Debug("no URL scheme request to handle")
}
@@ -343,7 +344,13 @@ func main() {
}
}()
osRun(cancel, hasCompletedFirstRun, startHidden)
settings, settingsErr := st.Settings()
showOnboarding := shouldShowOnboarding(settings, settingsErr)
if settingsErr != nil {
slog.Error("failed to load onboarding state", "error", settingsErr)
}
osRun(cancel, hasCompletedFirstRun, startHidden, showOnboarding, urlSchemeRequest)
slog.Info("shutting down desktop server")
if err := srv.Close(); err != nil {
@@ -355,6 +362,33 @@ func main() {
<-done
}
func shouldShowOnboarding(settings store.Settings, err error) bool {
return err != nil || settings.OnboardingVersion < store.CurrentOnboardingVersion
}
func runInitialWindowsUI(
startHidden bool,
showOnboarding bool,
urlSchemeRequest string,
startHiddenFn func(),
handleURLFn func(string),
showUIFn func(string),
) {
if urlSchemeRequest != "" {
handleURLFn(urlSchemeRequest)
return
}
if startHidden {
startHiddenFn()
return
}
if showOnboarding {
showUIFn("/")
return
}
showUIFn("/connect")
}
func startHiddenTasks() {
// If an upgrade is ready and we're in hidden mode, perform it at startup.
// If we're not in hidden mode, we want to start as fast as possible and not
@@ -375,7 +409,7 @@ func startHiddenTasks() {
return
}
if err := updater.DoUpgradeAtStartup(); err != nil {
if err := updater.DoUpgradeAtStartup(); err != nil { //nolint:staticcheck,nolintlint // DoUpgradeAtStartup may always return non-nil on Windows
slog.Info("unable to perform upgrade at startup", "error", err)
// Make sure the restart to upgrade menu shows so we can attempt an interactive upgrade to get authorization
UpdateAvailable("")
@@ -432,7 +466,7 @@ func checkUserLoggedIn(uiServerPort int) bool {
func handleConnectURLScheme() {
if checkUserLoggedIn(uiServerPort) {
slog.Info("user is already logged in, opening app instead")
showWindow(wv.webview.Window())
openUI("/")
return
}
@@ -491,17 +525,23 @@ func parseURLScheme(urlSchemeRequest string) (isConnect bool, err error) {
// handleURLSchemeInCurrentInstance processes URL scheme requests in the current instance
func handleURLSchemeInCurrentInstance(urlSchemeRequest string) {
isConnect, err := parseURLScheme(urlSchemeRequest)
err := dispatchURLSchemeRequest(urlSchemeRequest, handleConnectURLScheme, func() {
openUI("/")
})
if err != nil {
slog.Error("failed to parse URL scheme request", "url", urlSchemeRequest, "error", err)
return
}
if isConnect {
handleConnectURLScheme()
} else {
if wv.webview != nil {
showWindow(wv.webview.Window())
}
}
}
func dispatchURLSchemeRequest(urlSchemeRequest string, connect, open func()) error {
isConnect, err := parseURLScheme(urlSchemeRequest)
if err != nil {
return err
}
if isConnect {
connect()
} else {
open()
}
return nil
}
+1044 -15
View File
File diff suppressed because it is too large. Load diff
+27 -1
View File
@@ -16,8 +16,9 @@ enum AppMove
MoveError,
};
void run(bool firstTimeRun, bool startHidden);
void run(bool showOnboarding, bool startHidden);
void killOtherInstances();
bool otherOllamaInstanceRunning(void);
enum AppMove askToMoveToApplications();
int createSymlinkWithAuthorization();
int installSymlink(const char *cliPath);
@@ -25,6 +26,7 @@ extern void Restart();
// extern void Quit();
void StartUI(const char *path);
void ShowUI();
bool IsOnboardingActive(void);
void StopUI();
void StartUpdate();
void darwinStartHiddenTasks();
@@ -38,6 +40,30 @@ void setWindowDelegate(void *window);
void showWindow(uintptr_t wndPtr);
void hideWindow(uintptr_t wndPtr);
void styleWindow(uintptr_t wndPtr);
void setWindowResizable(uintptr_t wndPtr, bool resizable);
void drag(uintptr_t wndPtr);
void doubleClick(uintptr_t wndPtr);
void handleConnectURL();
bool SetClaudeGatewayInstalled(bool installed, bool restartClaude);
bool HasUsedClaudeDesktopIntegration(void);
bool RestoreClaudeGatewayForShutdown(void);
bool IsClaudeGatewayConfigured(void);
bool IsClaudeDesktopInstalled(void);
bool IsClaudeDesktopRunning(void);
bool ClaudeGatewayStartFailed(void);
bool ClaudeGatewayPortConflict(void);
char *ClaudeGatewayErrorMessage(void);
int ClaudeGatewayPort(void);
void RefreshClaudeProxyMenu(void);
void updateClaudeProxyMenu(unsigned long long routed);
bool ShowAppsInMenu(void);
void SetShowAppsInMenu(bool visible);
enum ClaudeInstallResult
{
ClaudeInstallCancelled,
ClaudeInstallerOpened,
ClaudeInstallFailed,
};
enum ClaudeInstallResult installClaudeDesktop(void);
char *ClaudeDesktopDownloadRequest(char **authorization);
bool InstallClaudeDesktopArchive(const char *archivePath);
+1063 -71
View File
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
+140
View File
@@ -0,0 +1,140 @@
//go:build windows || darwin
package main
import (
"errors"
"testing"
"github.com/ollama/ollama/app/store"
)
func TestShouldShowOnboarding(t *testing.T) {
tests := []struct {
name string
settings store.Settings
err error
want bool
}{
{
name: "fresh install",
settings: store.Settings{OnboardingVersion: 0},
want: true,
},
{
name: "completed onboarding",
settings: store.Settings{OnboardingVersion: store.CurrentOnboardingVersion},
want: false,
},
{
name: "settings failure",
err: errors.New("settings unavailable"),
want: true,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
if got := shouldShowOnboarding(tt.settings, tt.err); got != tt.want {
t.Fatalf("shouldShowOnboarding() = %v, want %v", got, tt.want)
}
})
}
}
func TestDispatchURLSchemeRequest(t *testing.T) {
tests := []struct {
name string
request string
wantConnect bool
wantOpen bool
wantErr bool
}{
{name: "bare URL opens app", request: "ollama://", wantOpen: true},
{name: "connect URL starts connection", request: "ollama://connect", wantConnect: true},
{name: "unsupported URL", request: "ollama://unsupported", wantErr: true},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
connected := false
opened := false
err := dispatchURLSchemeRequest(
tt.request,
func() { connected = true },
func() { opened = true },
)
if (err != nil) != tt.wantErr {
t.Fatalf("dispatchURLSchemeRequest() error = %v, wantErr %v", err, tt.wantErr)
}
if connected != tt.wantConnect {
t.Errorf("connect called = %v, want %v", connected, tt.wantConnect)
}
if opened != tt.wantOpen {
t.Errorf("open called = %v, want %v", opened, tt.wantOpen)
}
})
}
}
func TestRunInitialWindowsUIWithBareURL(t *testing.T) {
hiddenCalls := 0
urlCalls := 0
onboardingCalls := 0
openCalls := 0
runInitialWindowsUI(
false,
true,
"ollama://",
func() { hiddenCalls++ },
func(request string) {
urlCalls++
if err := dispatchURLSchemeRequest(request, func() {}, func() { openCalls++ }); err != nil {
t.Fatalf("dispatchURLSchemeRequest() error = %v", err)
}
},
func(path string) {
onboardingCalls++
},
)
if urlCalls != 1 {
t.Fatalf("URL handled %d times, want 1", urlCalls)
}
if openCalls != 1 {
t.Errorf("app opened %d times, want 1", openCalls)
}
if hiddenCalls != 0 {
t.Errorf("hidden startup called %d times, want 0", hiddenCalls)
}
if onboardingCalls != 0 {
t.Errorf("onboarding opened %d times, want 0", onboardingCalls)
}
}
func TestRunInitialWindowsUIRoutesInteractiveLaunch(t *testing.T) {
for _, tt := range []struct {
name string
showOnboarding bool
wantPath string
}{
{name: "fresh install preserves onboarding", showOnboarding: true, wantPath: "/"},
{name: "returning launch opens apps", wantPath: "/connect"},
} {
t.Run(tt.name, func(t *testing.T) {
var gotPath string
runInitialWindowsUI(
false,
tt.showOnboarding,
"",
func() { t.Fatal("unexpected hidden startup") },
func(string) { t.Fatal("unexpected URL handling") },
func(path string) { gotPath = path },
)
if gotPath != tt.wantPath {
t.Fatalf("initial UI path = %q, want %q", gotPath, tt.wantPath)
}
})
}
}
+21 -29
View File
@@ -95,11 +95,15 @@ func (ac *appCallbacks) UIRun(path string) {
}
func (*appCallbacks) UIShow() {
if wv.webview != nil {
openUI("/")
}
func openUI(path string) {
if wv.IsRunning() && wv.webview != nil {
showWindow(wv.webview.Window())
} else {
wv.Run("/")
return
}
wv.Run(path)
}
func (*appCallbacks) UITerminate() {
@@ -110,6 +114,10 @@ func (*appCallbacks) UIRunning() bool {
return wv.IsRunning()
}
func (*appCallbacks) UIOnboarding() bool {
return wv.OnboardingActive()
}
func (app *appCallbacks) Quit() {
app.t.Quit()
wv.Terminate()
@@ -126,7 +134,7 @@ func (app *appCallbacks) DoUpdate() {
app.shutdown()
if err := updater.DoUpgrade(true); err != nil {
if err := updater.DoUpgrade(true); err != nil { //nolint:staticcheck,nolintlint // DoUpgrade may always return non-nil on Windows
slog.Warn(fmt.Sprintf("upgrade attempt failed: %s", err))
}
}
@@ -138,19 +146,7 @@ func (app *appCallbacks) HandleURLScheme(urlScheme string) {
// handleURLSchemeRequest processes URL scheme requests from other instances
func handleURLSchemeRequest(urlScheme string) {
isConnect, err := parseURLScheme(urlScheme)
if err != nil {
slog.Error("failed to parse URL scheme request", "url", urlScheme, "error", err)
return
}
if isConnect {
handleConnectURLScheme()
} else {
if wv.webview != nil {
showWindow(wv.webview.Window())
}
}
handleURLSchemeInCurrentInstance(urlScheme)
}
func UpdateAvailable(ver string) error {
@@ -161,7 +157,7 @@ func UpdateAvailable(ver string) error {
return app.t.UpdateAvailable(ver)
}
func osRun(shutdown func(), hasCompletedFirstRun, startHidden bool) {
func osRun(shutdown func(), hasCompletedFirstRun, startHidden, showOnboarding bool, urlSchemeRequest string) {
var err error
app.shutdown = shutdown
app.t, err = wintray.NewTray(app)
@@ -205,10 +201,8 @@ func osRun(shutdown func(), hasCompletedFirstRun, startHidden bool) {
}
}
}
if startHidden {
startHiddenTasks()
} else {
ptr := wv.Run("/")
runInitialWindowsUI(startHidden, showOnboarding, urlSchemeRequest, startHiddenTasks, handleURLSchemeInCurrentInstance, func(path string) {
ptr := wv.Run(path)
// Set the window icon using the tray icon
if ptr != nil {
@@ -225,7 +219,7 @@ func osRun(shutdown func(), hasCompletedFirstRun, startHidden bool) {
}
centerWindow(ptr)
}
})
if !hasCompletedFirstRun {
// Only create the login shortcut on first start
@@ -408,6 +402,8 @@ func hideWindow(ptr unsafe.Pointer) {
}
}
func setOnboardingWindowStyle(_ unsafe.Pointer, _ bool) {}
func runInBackground() {
exe, err := os.Executable()
if err != nil {
@@ -432,17 +428,13 @@ func drag(ptr unsafe.Pointer) {}
func doubleClick(ptr unsafe.Pointer) {}
// checkAndHandleExistingInstance checks if another instance is running and sends the URL to it
func checkAndHandleExistingInstance(urlSchemeRequest string) bool {
func checkAndHandleExistingInstance(urlSchemeRequest string) {
if urlSchemeRequest == "" {
return false
return
}
// Try to send URL to existing instance using wintray messaging
if wintray.CheckAndSendToExistingInstance(urlSchemeRequest) {
os.Exit(0)
return true
}
// No existing instance, we'll handle it ourselves
return false
}
@@ -0,0 +1,61 @@
//go:build darwin
package main
import "github.com/ollama/ollama/app/webview"
func bindClaudeDesktop(wv webview.WebView) {
wv.Bind("getClaudeDesktopStatus", func() claudeDesktopStatus {
return getClaudeDesktopConnectionStatus()
})
wv.Bind("setClaudeDesktopConnected", func(enabled bool) claudeDesktopActionResult {
err := setClaudeDesktopConnection(enabled)
result := claudeDesktopActionResult{
Status: getClaudeDesktopConnectionStatus(),
}
if err != nil {
result.Error = err.Error()
}
return result
})
wv.Bind("prepareClaudeDesktopConnection", func() claudeDesktopActionResult {
err := prepareClaudeDesktopConnection()
result := claudeDesktopActionResult{
Status: getClaudeDesktopConnectionStatus(),
}
if err != nil {
result.Error = err.Error()
}
return result
})
wv.Bind("openClaudeDesktop", func() string {
if err := openClaudeDesktopApplication(); err != nil {
return err.Error()
}
return ""
})
wv.Bind("installClaudeDesktop", func() claudeDesktopInstallResult {
return requestClaudeDesktopInstall()
})
wv.Bind("restartClaudeDesktop", func(models []string) claudeDesktopActionResult {
err := restartClaudeDesktopWithModels(models)
result := claudeDesktopActionResult{Status: getClaudeDesktopConnectionStatus()}
if err != nil {
result.Error = err.Error()
}
return result
})
wv.Bind("getShowAppsInMenu", func() bool {
return getShowAppsInMenu()
})
wv.Bind("setShowAppsInMenu", func(visible bool) {
setShowAppsInMenu(visible)
})
}
@@ -0,0 +1,7 @@
//go:build windows
package main
import "github.com/ollama/ollama/app/webview"
func bindClaudeDesktop(_ webview.WebView) {}
@@ -0,0 +1,252 @@
//go:build darwin
package main
import (
"archive/zip"
"errors"
"fmt"
"io"
"os"
"os/exec"
"path/filepath"
"strings"
)
const (
maxClaudeDesktopArchiveBytes = 1 << 30
maxClaudeDesktopExtractBytes = 2 << 30
maxClaudeDesktopArchiveFiles = 100_000
claudeDesktopBundleID = "com.anthropic.claudefordesktop"
claudeDesktopTeamID = "Q6L2SF6YDW"
)
var errClaudeDesktopDestinationExists = errors.New("Claude Desktop installation destination already exists")
func claudeDesktopInstallDestinations() []string {
destinations := []string{"/Applications/Claude.app"}
if home, err := os.UserHomeDir(); err == nil {
destinations = append(destinations, filepath.Join(home, "Applications", "Claude.app"))
}
return destinations
}
func installClaudeDesktopZip(archivePath string, destinations []string, verify func(string) error) (string, error) {
if len(destinations) == 0 {
return "", errors.New("Claude Desktop installation destination is required")
}
if verify == nil {
return "", errors.New("Claude Desktop bundle verifier is required")
}
info, err := os.Stat(archivePath)
if err != nil {
return "", fmt.Errorf("stat Claude Desktop archive: %w", err)
}
if !info.Mode().IsRegular() {
return "", errors.New("Claude Desktop archive is not a regular file")
}
if info.Size() > maxClaudeDesktopArchiveBytes {
return "", fmt.Errorf("Claude Desktop archive exceeds %d bytes", maxClaudeDesktopArchiveBytes)
}
workDir, err := os.MkdirTemp("", "ollama-claude-install-")
if err != nil {
return "", fmt.Errorf("create Claude Desktop installation directory: %w", err)
}
defer os.RemoveAll(workDir)
if err := extractClaudeDesktopZip(archivePath, workDir); err != nil {
return "", err
}
bundlePath := filepath.Join(workDir, "Claude.app")
if err := validateClaudeDesktopBundle(bundlePath); err != nil {
return "", err
}
if err := verify(bundlePath); err != nil {
return "", fmt.Errorf("verify Claude Desktop signature: %w", err)
}
var permissionErr error
for _, destination := range destinations {
if strings.TrimSpace(destination) == "" {
continue
}
if _, err := os.Stat(destination); err == nil {
return "", fmt.Errorf("%w: %s", errClaudeDesktopDestinationExists, destination)
} else if !errors.Is(err, os.ErrNotExist) {
return "", fmt.Errorf("check Claude Desktop destination %s: %w", destination, err)
}
if err := os.MkdirAll(filepath.Dir(destination), 0o755); err != nil {
if errors.Is(err, os.ErrPermission) {
permissionErr = err
continue
}
return "", fmt.Errorf("create Claude Desktop destination: %w", err)
}
if err := os.Rename(bundlePath, destination); err != nil {
if errors.Is(err, os.ErrPermission) {
permissionErr = err
continue
}
return "", fmt.Errorf("move Claude Desktop to %s: %w", destination, err)
}
return destination, nil
}
if permissionErr != nil {
return "", fmt.Errorf("install Claude Desktop in Applications: %w", permissionErr)
}
return "", errors.New("Claude Desktop installation destination is required")
}
func extractClaudeDesktopZip(archivePath, destination string) error {
reader, err := zip.OpenReader(archivePath)
if err != nil {
return fmt.Errorf("open Claude Desktop archive: %w", err)
}
defer reader.Close()
if len(reader.File) == 0 {
return errors.New("Claude Desktop archive is empty")
}
if len(reader.File) > maxClaudeDesktopArchiveFiles {
return fmt.Errorf("Claude Desktop archive contains more than %d files", maxClaudeDesktopArchiveFiles)
}
var expanded uint64
for _, file := range reader.File {
clean, err := safeClaudeDesktopArchivePath(file.Name)
if err != nil {
return err
}
expanded += file.UncompressedSize64
if expanded > maxClaudeDesktopExtractBytes {
return fmt.Errorf("Claude Desktop archive expands beyond %d bytes", maxClaudeDesktopExtractBytes)
}
path := filepath.Join(destination, filepath.FromSlash(clean))
switch {
case file.FileInfo().IsDir():
if err := os.MkdirAll(path, file.Mode().Perm()); err != nil {
return fmt.Errorf("create Claude Desktop archive directory: %w", err)
}
case file.Mode()&os.ModeSymlink != 0:
target, err := readClaudeDesktopZipFile(file, 16<<10)
if err != nil {
return fmt.Errorf("read Claude Desktop archive symlink: %w", err)
}
if err := validateClaudeDesktopSymlink(clean, string(target)); err != nil {
return err
}
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
return fmt.Errorf("create Claude Desktop archive directory: %w", err)
}
if err := os.Symlink(string(target), path); err != nil {
return fmt.Errorf("create Claude Desktop archive symlink: %w", err)
}
case file.Mode().IsRegular():
if err := extractClaudeDesktopZipFile(file, path); err != nil {
return err
}
default:
return fmt.Errorf("Claude Desktop archive contains unsupported file %q", file.Name)
}
}
return nil
}
func safeClaudeDesktopArchivePath(name string) (string, error) {
if strings.ContainsRune(name, '\x00') || filepath.IsAbs(name) {
return "", fmt.Errorf("Claude Desktop archive contains unsafe path %q", name)
}
clean := filepath.ToSlash(filepath.Clean(name))
if clean != "Claude.app" && !strings.HasPrefix(clean, "Claude.app/") {
return "", fmt.Errorf("Claude Desktop archive contains unexpected path %q", name)
}
return clean, nil
}
func validateClaudeDesktopSymlink(name, target string) error {
if target == "" || filepath.IsAbs(target) {
return fmt.Errorf("Claude Desktop archive contains unsafe symlink %q", name)
}
resolved := filepath.Clean(filepath.Join(filepath.Dir(name), target))
resolved = filepath.ToSlash(resolved)
if resolved != "Claude.app" && !strings.HasPrefix(resolved, "Claude.app/") {
return fmt.Errorf("Claude Desktop archive symlink %q escapes Claude.app", name)
}
return nil
}
func extractClaudeDesktopZipFile(file *zip.File, path string) error {
if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil {
return fmt.Errorf("create Claude Desktop archive directory: %w", err)
}
input, err := file.Open()
if err != nil {
return fmt.Errorf("open Claude Desktop archive file: %w", err)
}
output, err := os.OpenFile(path, os.O_CREATE|os.O_EXCL|os.O_WRONLY, file.Mode().Perm())
if err != nil {
input.Close()
return fmt.Errorf("create Claude Desktop archive file: %w", err)
}
_, copyErr := io.Copy(output, input)
inputErr := input.Close()
outputErr := output.Close()
if copyErr != nil {
return fmt.Errorf("extract Claude Desktop archive file: %w", copyErr)
}
if inputErr != nil {
return fmt.Errorf("close Claude Desktop archive file: %w", inputErr)
}
if outputErr != nil {
return fmt.Errorf("close extracted Claude Desktop file: %w", outputErr)
}
return nil
}
func readClaudeDesktopZipFile(file *zip.File, limit int64) ([]byte, error) {
reader, err := file.Open()
if err != nil {
return nil, err
}
defer reader.Close()
data, err := io.ReadAll(io.LimitReader(reader, limit+1))
if err != nil {
return nil, err
}
if int64(len(data)) > limit {
return nil, fmt.Errorf("archive entry exceeds %d bytes", limit)
}
return data, nil
}
func validateClaudeDesktopBundle(bundlePath string) error {
info, err := os.Stat(bundlePath)
if err != nil || !info.IsDir() {
return errors.New("Claude Desktop archive does not contain Claude.app")
}
executable := filepath.Join(bundlePath, "Contents", "MacOS", "Claude")
info, err = os.Stat(executable)
if err != nil {
return fmt.Errorf("Claude Desktop executable is missing: %w", err)
}
if !info.Mode().IsRegular() || info.Mode()&0o111 == 0 {
return errors.New("Claude Desktop executable is not executable")
}
return nil
}
func verifyClaudeDesktopBundle(bundlePath string) error {
if output, err := exec.Command("/usr/bin/codesign", "--verify", "--deep", "--strict", bundlePath).CombinedOutput(); err != nil {
return fmt.Errorf("codesign verification failed: %w: %s", err, strings.TrimSpace(string(output)))
}
output, err := exec.Command("/usr/bin/codesign", "-d", "--verbose=4", bundlePath).CombinedOutput()
if err != nil {
return fmt.Errorf("read code signature: %w: %s", err, strings.TrimSpace(string(output)))
}
details := string(output)
if !strings.Contains(details, "Identifier="+claudeDesktopBundleID) ||
!strings.Contains(details, "TeamIdentifier="+claudeDesktopTeamID) {
return fmt.Errorf("unexpected Claude Desktop signing identity")
}
return nil
}
@@ -0,0 +1,162 @@
//go:build darwin
package main
import (
"archive/zip"
"errors"
"os"
"path/filepath"
"strings"
"testing"
)
func TestInstallClaudeDesktopZip(t *testing.T) {
archive := writeClaudeDesktopTestZip(t, map[string]claudeDesktopTestZipEntry{
"Claude.app/": {directory: true},
"Claude.app/Contents/": {directory: true},
"Claude.app/Contents/MacOS/": {directory: true},
"Claude.app/Contents/MacOS/Claude": {body: "binary", mode: 0o755},
"Claude.app/Contents/Resources/": {directory: true},
"Claude.app/Contents/Resources/link": {body: "../MacOS/Claude", mode: os.ModeSymlink | 0o777},
})
destination := filepath.Join(t.TempDir(), "Applications", "Claude.app")
var verified string
installed, err := installClaudeDesktopZip(archive, []string{destination}, func(bundle string) error {
verified = bundle
return nil
})
if err != nil {
t.Fatal(err)
}
if installed != destination || verified == "" {
t.Fatalf("installed = %q, verified = %q", installed, verified)
}
info, err := os.Stat(filepath.Join(installed, "Contents", "MacOS", "Claude"))
if err != nil {
t.Fatal(err)
}
if info.Mode()&0o111 == 0 {
t.Fatal("installed Claude executable is not executable")
}
if target, err := os.Readlink(filepath.Join(installed, "Contents", "Resources", "link")); err != nil || target != "../MacOS/Claude" {
t.Fatalf("symlink target = %q, err = %v", target, err)
}
}
func TestInstallClaudeDesktopZipRejectsUnsafeArchives(t *testing.T) {
for _, test := range []struct {
name string
entries map[string]claudeDesktopTestZipEntry
}{
{name: "path traversal", entries: map[string]claudeDesktopTestZipEntry{"../Claude.app/Contents/MacOS/Claude": {body: "binary", mode: 0o755}}},
{name: "unexpected root", entries: map[string]claudeDesktopTestZipEntry{"README": {body: "nope", mode: 0o644}}},
{name: "escaping symlink", entries: map[string]claudeDesktopTestZipEntry{
"Claude.app/Contents/MacOS/Claude": {body: "binary", mode: 0o755},
"Claude.app/escape": {body: "../../outside", mode: os.ModeSymlink | 0o777},
}},
} {
t.Run(test.name, func(t *testing.T) {
archive := writeClaudeDesktopTestZip(t, test.entries)
destination := filepath.Join(t.TempDir(), "Claude.app")
if _, err := installClaudeDesktopZip(archive, []string{destination}, func(string) error { return nil }); err == nil {
t.Fatal("installClaudeDesktopZip succeeded")
}
if _, err := os.Stat(destination); !errors.Is(err, os.ErrNotExist) {
t.Fatalf("unsafe archive created destination: %v", err)
}
})
}
}
func TestInstallClaudeDesktopZipVerifiesBeforeMove(t *testing.T) {
archive := writeClaudeDesktopTestZip(t, map[string]claudeDesktopTestZipEntry{
"Claude.app/Contents/MacOS/Claude": {body: "binary", mode: 0o755},
})
destination := filepath.Join(t.TempDir(), "Claude.app")
wantErr := errors.New("invalid signature")
if _, err := installClaudeDesktopZip(archive, []string{destination}, func(string) error { return wantErr }); !errors.Is(err, wantErr) {
t.Fatalf("error = %v, want %v", err, wantErr)
}
if _, err := os.Stat(destination); !errors.Is(err, os.ErrNotExist) {
t.Fatalf("invalid bundle created destination: %v", err)
}
}
func TestInstallClaudeDesktopZipDoesNotOverwrite(t *testing.T) {
archive := writeClaudeDesktopTestZip(t, map[string]claudeDesktopTestZipEntry{
"Claude.app/Contents/MacOS/Claude": {body: "binary", mode: 0o755},
})
destination := filepath.Join(t.TempDir(), "Claude.app")
if err := os.MkdirAll(destination, 0o755); err != nil {
t.Fatal(err)
}
if _, err := installClaudeDesktopZip(archive, []string{destination}, func(string) error { return nil }); !errors.Is(err, errClaudeDesktopDestinationExists) {
t.Fatalf("error = %v, want destination exists", err)
}
}
func TestInstallClaudeDesktopZipRealArchive(t *testing.T) {
archive := os.Getenv("OLLAMA_TEST_CLAUDE_DESKTOP_ZIP")
if archive == "" {
t.Skip("set OLLAMA_TEST_CLAUDE_DESKTOP_ZIP to a downloaded Claude Desktop ZIP")
}
destination := filepath.Join(t.TempDir(), "Applications", "Claude.app")
installed, err := installClaudeDesktopZip(
archive,
[]string{destination},
verifyClaudeDesktopBundle,
)
if err != nil {
t.Fatal(err)
}
if installed != destination {
t.Fatalf("installed = %q, want %q", installed, destination)
}
}
type claudeDesktopTestZipEntry struct {
body string
mode os.FileMode
directory bool
}
func writeClaudeDesktopTestZip(t *testing.T, entries map[string]claudeDesktopTestZipEntry) string {
t.Helper()
path := filepath.Join(t.TempDir(), "Claude.zip")
file, err := os.Create(path)
if err != nil {
t.Fatal(err)
}
writer := zip.NewWriter(file)
for name, entry := range entries {
header := &zip.FileHeader{Name: name, Method: zip.Deflate}
if entry.directory {
header.SetMode(os.ModeDir | 0o755)
} else {
header.SetMode(entry.mode)
}
item, err := writer.CreateHeader(header)
if err != nil {
t.Fatal(err)
}
if _, err := item.Write([]byte(entry.body)); err != nil {
t.Fatal(err)
}
}
if err := writer.Close(); err != nil {
t.Fatal(err)
}
if err := file.Close(); err != nil {
t.Fatal(err)
}
return path
}
func TestSafeClaudeDesktopArchivePath(t *testing.T) {
for _, name := range []string{"Claude.app", "Claude.app/Contents/MacOS/Claude"} {
if got, err := safeClaudeDesktopArchivePath(name); err != nil || got != strings.TrimSuffix(name, "/") {
t.Fatalf("safeClaudeDesktopArchivePath(%q) = %q, %v", name, got, err)
}
}
}
+44
View File
@@ -0,0 +1,44 @@
//go:build darwin
package main
import "github.com/ollama/ollama/internal/proxy"
type claudeDesktopInstallResult string
const (
claudeDesktopInstallCancelled claudeDesktopInstallResult = "cancelled"
claudeDesktopInstallerOpened claudeDesktopInstallResult = "opened"
claudeDesktopInstallFailed claudeDesktopInstallResult = "failed"
)
type claudeDesktopStatus struct {
Supported bool `json:"supported"`
Used bool `json:"used"`
Installed bool `json:"installed"`
Configured bool `json:"configured"`
Connected bool `json:"connected"`
Running bool `json:"running"`
StartFailed bool `json:"startFailed"`
PortConflict bool `json:"portConflict"`
GatewayPort int `json:"gatewayPort,omitempty"`
Error string `json:"error,omitempty"`
ModelSource string `json:"modelSource,omitempty"`
Models []claudeDesktopModelStatus `json:"models,omitempty"`
}
type claudeDesktopModelStatus struct {
Name string `json:"name"`
DisplayName string `json:"displayName"`
Description string `json:"description,omitempty"`
Cloud bool `json:"cloud"`
Selected bool `json:"selected"`
Availability proxy.ClaudeDesktopAvailability `json:"availability"`
Reason proxy.ClaudeDesktopAccessReason `json:"reason,omitempty"`
RequiredPlan string `json:"requiredPlan,omitempty"`
}
type claudeDesktopActionResult struct {
Status claudeDesktopStatus `json:"status"`
Error string `json:"error,omitempty"`
}
+94 -83
View File
@@ -16,6 +16,7 @@ import (
"runtime"
"strings"
"sync"
"sync/atomic"
"time"
"unsafe"
@@ -24,11 +25,21 @@ import (
"github.com/ollama/ollama/app/webview"
)
const (
defaultWindowWidth = 1360
defaultWindowHeight = 960
onboardingWindowWidth = 900
onboardingWindowHeight = 660
minimumWindowWidth = onboardingWindowWidth
minimumWindowHeight = onboardingWindowHeight
)
type Webview struct {
port int
token string
webview webview.WebView
mutex sync.Mutex
port int
token string
webview webview.WebView
mutex sync.Mutex
onboarding atomic.Bool
Store *store.Store
}
@@ -88,85 +99,32 @@ func (w *Webview) Run(path string) unsafe.Pointer {
// Windows-specific scrollbar styling
if runtime.GOOS == "windows" {
init += `
// Fix scrollbar styling for Edge WebView2 on Windows only
// Keep Edge WebView2 scrollbars aligned with the light-only app theme.
function updateScrollbarStyles() {
const isDark = window.matchMedia('(prefers-color-scheme: dark)').matches;
const existingStyle = document.getElementById('scrollbar-style');
if (existingStyle) existingStyle.remove();
const style = document.createElement('style');
style.id = 'scrollbar-style';
if (isDark) {
style.textContent = ` + "`" + `
::-webkit-scrollbar { width: 6px !important; height: 6px !important; }
::-webkit-scrollbar-track { background: #1a1a1a !important; }
::-webkit-scrollbar-thumb { background: #404040 !important; border-radius: 6px !important; }
::-webkit-scrollbar-thumb:hover { background: #505050 !important; }
::-webkit-scrollbar-corner { background: #1a1a1a !important; }
::-webkit-scrollbar-button {
background: transparent !important;
border: none !important;
width: 0px !important;
height: 0px !important;
margin: 0 !important;
padding: 0 !important;
}
::-webkit-scrollbar-button:vertical:start:decrement {
background: transparent !important;
height: 0px !important;
}
::-webkit-scrollbar-button:vertical:end:increment {
background: transparent !important;
height: 0px !important;
}
::-webkit-scrollbar-button:horizontal:start:decrement {
background: transparent !important;
width: 0px !important;
}
::-webkit-scrollbar-button:horizontal:end:increment {
background: transparent !important;
width: 0px !important;
}
` + "`" + `;
} else {
style.textContent = ` + "`" + `
::-webkit-scrollbar { width: 6px !important; height: 6px !important; }
::-webkit-scrollbar-track { background: #f0f0f0 !important; }
::-webkit-scrollbar-thumb { background: #c0c0c0 !important; border-radius: 6px !important; }
::-webkit-scrollbar-thumb:hover { background: #a0a0a0 !important; }
::-webkit-scrollbar-corner { background: #f0f0f0 !important; }
::-webkit-scrollbar-button {
background: transparent !important;
border: none !important;
width: 0px !important;
height: 0px !important;
margin: 0 !important;
padding: 0 !important;
}
::-webkit-scrollbar-button:vertical:start:decrement {
background: transparent !important;
height: 0px !important;
}
::-webkit-scrollbar-button:vertical:end:increment {
background: transparent !important;
height: 0px !important;
}
::-webkit-scrollbar-button:horizontal:start:decrement {
background: transparent !important;
width: 0px !important;
}
::-webkit-scrollbar-button:horizontal:end:increment {
background: transparent !important;
width: 0px !important;
}
` + "`" + `;
}
style.textContent = ` + "`" + `
::-webkit-scrollbar { width: 6px !important; height: 6px !important; }
::-webkit-scrollbar-track { background: #f0f0f0 !important; }
::-webkit-scrollbar-thumb { background: #c0c0c0 !important; border-radius: 6px !important; }
::-webkit-scrollbar-thumb:hover { background: #a0a0a0 !important; }
::-webkit-scrollbar-corner { background: #f0f0f0 !important; }
::-webkit-scrollbar-button {
background: transparent !important;
border: none !important;
width: 0px !important;
height: 0px !important;
margin: 0 !important;
padding: 0 !important;
}
` + "`" + `;
document.head.appendChild(style);
}
window.addEventListener('load', updateScrollbarStyles);
window.matchMedia('(prefers-color-scheme: dark)').addEventListener('change', updateScrollbarStyles);
`
}
// on windows make ctrl+n open new chat
@@ -187,15 +145,32 @@ func (w *Webview) Run(path string) unsafe.Pointer {
`
}
init += `
init += fmt.Sprintf(`
window.OLLAMA_PLATFORM = %q;
window.OLLAMA_WEBSEARCH = true;
`
`, runtime.GOOS)
wv.Init(init)
// Add keyboard handler for zoom
wv.Init(`
window.addEventListener('keydown', function(e) {
const isZoomShortcut = (e.metaKey || e.ctrlKey) && (
e.key === '+' || e.key === '=' || e.key === '-' ||
e.key === '_' || e.key === '0' ||
e.code === 'NumpadAdd' || e.code === 'NumpadSubtract'
);
// Keep fixed-scale onboarding and apps pages at their intended size.
const isFixedScalePage =
window.location.pathname === '/onboarding' ||
window.location.pathname === '/connect';
if (isFixedScalePage && isZoomShortcut) {
e.preventDefault();
e.stopImmediatePropagation();
return false;
}
// CMD/Ctrl + Plus/Equals (zoom in)
if ((e.metaKey || e.ctrlKey) && (e.key === '+' || e.key === '=')) {
e.preventDefault();
@@ -237,10 +212,41 @@ func (w *Webview) Run(path string) unsafe.Pointer {
showWindow(wv.Window())
})
wv.Bind("activateOllama", func() {
showWindow(wv.Window())
})
bindClaudeDesktop(wv)
wv.Bind("close", func() {
hideWindow(wv.Window())
})
wv.Bind("setOnboardingWindow", func(enabled bool) {
w.onboarding.Store(enabled)
wv.Dispatch(func() {
if enabled {
wv.SetSize(onboardingWindowWidth, onboardingWindowHeight, webview.HintFixed)
setOnboardingWindowStyle(wv.Window(), true)
return
}
width, height := defaultWindowWidth, defaultWindowHeight
if w.Store != nil {
storedWidth, storedHeight, err := w.Store.WindowSize()
if err != nil {
slog.Error("failed to restore window size", "error", err)
} else if storedWidth > 0 && storedHeight > 0 {
width, height = storedWidth, storedHeight
}
}
wv.SetSize(width, height, webview.HintNone)
wv.SetSize(minimumWindowWidth, minimumWindowHeight, webview.HintMin)
setOnboardingWindowStyle(wv.Window(), false)
})
})
// Webviews do not allow access to the file system by default, so we need to
// bind file system operations here
wv.Bind("selectModelsDirectory", func() {
@@ -450,18 +456,18 @@ func (w *Webview) Run(path string) unsafe.Pointer {
}()
}
width, height := defaultWindowWidth, defaultWindowHeight
if w.Store != nil {
width, height, err := w.Store.WindowSize()
storedWidth, storedHeight, err := w.Store.WindowSize()
if err != nil {
slog.Error("failed to get window size", "error", err)
}
if width > 0 && height > 0 {
wv.SetSize(width, height, webview.HintNone)
} else {
wv.SetSize(800, 600, webview.HintNone)
if storedWidth > 0 && storedHeight > 0 {
width, height = storedWidth, storedHeight
}
}
wv.SetSize(800, 600, webview.HintMin)
wv.SetSize(width, height, webview.HintNone)
wv.SetSize(minimumWindowWidth, minimumWindowHeight, webview.HintMin)
w.webview = wv
w.webview.Navigate(url)
@@ -476,6 +482,7 @@ func (w *Webview) Run(path string) unsafe.Pointer {
}
func (w *Webview) Terminate() {
w.onboarding.Store(false)
w.mutex.Lock()
if w.webview == nil {
w.mutex.Unlock()
@@ -489,6 +496,10 @@ func (w *Webview) Terminate() {
wv.Destroy()
}
func (w *Webview) OnboardingActive() bool {
return w.onboarding.Load()
}
func (w *Webview) IsRunning() bool {
w.mutex.Lock()
defer w.mutex.Unlock()
@@ -0,0 +1,7 @@
<?xml version="1.0" encoding="UTF-8"?>
<!-- Generated by Pixelmator Pro 3.6.17 -->
<svg width="1200" height="1200" viewBox="0 0 1200 1200" xmlns="http://www.w3.org/2000/svg">
<g id="g314">
<path id="path147" fill="#d97757" stroke="none" d="M 233.959793 800.214905 L 468.644287 668.536987 L 472.590637 657.100647 L 468.644287 650.738403 L 457.208069 650.738403 L 417.986633 648.322144 L 283.892639 644.69812 L 167.597321 639.865845 L 54.926208 633.825623 L 26.577238 627.785339 L 3.3e-05 592.751709 L 2.73832 575.27533 L 26.577238 559.248352 L 60.724873 562.228149 L 136.187973 567.382629 L 249.422867 575.194763 L 331.570496 580.026978 L 453.261841 592.671082 L 472.590637 592.671082 L 475.328857 584.859009 L 468.724915 580.026978 L 463.570557 575.194763 L 346.389313 495.785217 L 219.543671 411.865906 L 153.100723 363.543762 L 117.181267 339.060425 L 99.060455 316.107361 L 91.248367 266.01355 L 123.865784 230.093994 L 167.677887 233.073853 L 178.872513 236.053772 L 223.248367 270.201477 L 318.040283 343.570496 L 441.825592 434.738342 L 459.946411 449.798706 L 467.194672 444.64447 L 468.080597 441.020203 L 459.946411 427.409485 L 392.617493 305.718323 L 320.778564 181.932983 L 288.80542 130.630859 L 280.348999 99.865845 C 277.369171 87.221436 275.194641 76.590698 275.194641 63.624268 L 312.322174 13.20813 L 332.8591 6.604126 L 382.389313 13.20813 L 403.248352 31.328979 L 434.013519 101.71814 L 483.865753 212.537048 L 561.181274 363.221497 L 583.812134 407.919434 L 595.892639 449.315491 L 600.40271 461.959839 L 608.214783 461.959839 L 608.214783 454.711609 L 614.577271 369.825623 L 626.335632 265.61084 L 637.771851 131.516846 L 641.718201 93.745117 L 660.402832 48.483276 L 697.530334 24.000122 L 726.52356 37.852417 L 750.362549 72 L 747.060486 94.067139 L 732.886047 186.201416 L 705.100708 330.52356 L 686.979919 427.167847 L 697.530334 427.167847 L 709.61084 415.087341 L 758.496704 350.174561 L 840.644348 247.490051 L 876.885925 206.738342 L 919.167847 161.71814 L 946.308838 140.29541 L 997.61084 140.29541 L 1035.38269 196.429626 L 1018.469849 254.416199 L 965.637634 321.422852 L 921.825562 378.201538 L 859.006714 462.765259 L 819.785278 530.41626 L 823.409424 535.812073 L 832.75177 534.92627 L 974.657776 504.724915 L 1051.328979 490.872559 L 1142.818848 475.167786 L 1184.214844 494.496582 L 1188.724854 514.147644 L 1172.456421 554.335693 L 1074.604126 578.496765 L 959.838989 601.449829 L 788.939636 641.879272 L 786.845764 643.409485 L 789.261841 646.389343 L 866.255127 653.637634 L 899.194702 655.409424 L 979.812134 655.409424 L 1129.932861 666.604187 L 1169.154419 692.537109 L 1192.671265 724.268677 L 1188.724854 748.429688 L 1128.322144 779.194641 L 1046.818848 759.865845 L 856.590759 714.604126 L 791.355774 698.335754 L 782.335693 698.335754 L 782.335693 703.731567 L 836.69812 756.885986 L 936.322205 846.845581 L 1061.073975 962.81897 L 1067.436279 991.490112 L 1051.409424 1014.120911 L 1034.496704 1011.704712 L 924.885986 929.234924 L 882.604126 892.107544 L 786.845764 811.48999 L 780.483276 811.48999 L 780.483276 819.946289 L 802.550415 852.241699 L 919.087341 1027.409424 L 925.127625 1081.127686 L 916.671204 1098.604126 L 886.469849 1109.154419 L 853.288696 1103.114136 L 785.073914 1007.355835 L 714.684631 899.516785 L 657.906067 802.872498 L 650.979858 806.81897 L 617.476624 1167.704834 L 601.771851 1186.147705 L 565.530212 1200 L 535.328857 1177.046997 L 519.302124 1139.919556 L 535.328857 1066.550537 L 554.657776 970.792053 L 570.362488 894.68457 L 584.536926 800.134277 L 592.993347 768.724976 L 592.429626 766.630859 L 585.503479 767.516968 L 514.22821 865.369263 L 405.825531 1011.865906 L 320.053711 1103.677979 L 299.516815 1111.812256 L 263.919525 1093.369263 L 267.221497 1060.429688 L 287.114136 1031.114136 L 405.825531 880.107361 L 477.422913 786.52356 L 523.651062 732.483276 L 523.328918 724.671265 L 520.590698 724.671265 L 205.288605 929.395935 L 149.154434 936.644409 L 124.993355 914.01355 L 127.973183 876.885986 L 139.409409 864.80542 L 234.201385 799.570435 L 233.879227 799.8927 Z"/>
</g>
</svg>

After

Width:  |  Height:  |  Size: 4.0 KiB

+2 -2
View File
@@ -143,13 +143,13 @@ func utf16ptr(utf16 []uint16) *uint16 {
func utf16slice(ptr *uint16) []uint16 { //nolint:unused
hdr := reflect.SliceHeader{Data: uintptr(unsafe.Pointer(ptr)), Len: 1, Cap: 1}
slice := *((*[]uint16)(unsafe.Pointer(&hdr))) //nolint:govet
slice := *(*[]uint16)(unsafe.Pointer(&hdr)) //nolint:govet
i := 0
for slice[len(slice)-1] != 0 {
i++
}
hdr.Len = i
slice = *((*[]uint16)(unsafe.Pointer(&hdr))) //nolint:govet
slice = *(*[]uint16)(unsafe.Pointer(&hdr)) //nolint:govet
return slice
}
+54 -22
View File
@@ -14,7 +14,7 @@ import (
// currentSchemaVersion defines the current database schema version.
// Increment this when making schema changes that require migrations.
const currentSchemaVersion = 16
const currentSchemaVersion = 18
// database wraps the SQLite connection.
// SQLite handles its own locking for concurrent access:
@@ -82,12 +82,14 @@ func (db *database) init() error {
websearch_enabled BOOLEAN NOT NULL DEFAULT 0,
selected_model TEXT NOT NULL DEFAULT '',
sidebar_open BOOLEAN NOT NULL DEFAULT 0,
last_home_view TEXT NOT NULL DEFAULT 'launch',
last_home_view TEXT NOT NULL DEFAULT 'chat',
onboarding_version INTEGER NOT NULL DEFAULT 0,
think_enabled BOOLEAN NOT NULL DEFAULT 0,
think_level TEXT NOT NULL DEFAULT '',
cloud_setting_migrated BOOLEAN NOT NULL DEFAULT 0,
remote TEXT NOT NULL DEFAULT '', -- deprecated
auto_update_enabled BOOLEAN NOT NULL DEFAULT 1,
claude_desktop_used BOOLEAN NOT NULL DEFAULT 0,
schema_version INTEGER NOT NULL DEFAULT %d
);
@@ -271,6 +273,18 @@ func (db *database) migrate() error {
return fmt.Errorf("migrate v15 to v16: %w", err)
}
version = 16
case 16:
// Existing users should not be shown onboarding after an upgrade.
if err := db.migrateV16ToV17(); err != nil {
return fmt.Errorf("migrate v16 to v17: %w", err)
}
version = 17
case 17:
// Remember that Claude Desktop has been connected at least once.
if err := db.migrateV17ToV18(); err != nil {
return fmt.Errorf("migrate v17 to v18: %w", err)
}
version = 18
default:
// If we have a version we don't recognize, just set it to current
// This might happen during development
@@ -527,7 +541,7 @@ func (db *database) migrateV14ToV15() error {
// migrateV15ToV16 adds the last_home_view column to the settings table
func (db *database) migrateV15ToV16() error {
_, err := db.conn.Exec(`ALTER TABLE settings ADD COLUMN last_home_view TEXT NOT NULL DEFAULT 'launch'`)
_, err := db.conn.Exec(`ALTER TABLE settings ADD COLUMN last_home_view TEXT NOT NULL DEFAULT 'chat'`)
if err != nil && !duplicateColumnError(err) {
return fmt.Errorf("add last_home_view column: %w", err)
}
@@ -540,6 +554,38 @@ func (db *database) migrateV15ToV16() error {
return nil
}
// migrateV16ToV17 adds versioned onboarding state. The schema default stays at
// zero for genuinely new installs, while all existing rows are marked complete
// and moved off the retired launch home view.
func (db *database) migrateV16ToV17() error {
_, err := db.conn.Exec(`ALTER TABLE settings ADD COLUMN onboarding_version INTEGER NOT NULL DEFAULT 0`)
if err != nil && !duplicateColumnError(err) {
return fmt.Errorf("add onboarding_version column: %w", err)
}
_, err = db.conn.Exec(`UPDATE settings SET onboarding_version = 1, last_home_view = 'chat', schema_version = 17`)
if err != nil {
return fmt.Errorf("complete onboarding for existing users: %w", err)
}
return nil
}
// migrateV17ToV18 adds durable Claude Desktop integration history.
func (db *database) migrateV17ToV18() error {
_, err := db.conn.Exec(`ALTER TABLE settings ADD COLUMN claude_desktop_used BOOLEAN NOT NULL DEFAULT 0`)
if err != nil && !duplicateColumnError(err) {
return fmt.Errorf("add claude_desktop_used column: %w", err)
}
_, err = db.conn.Exec(`UPDATE settings SET schema_version = 18`)
if err != nil {
return fmt.Errorf("update schema version: %w", err)
}
return nil
}
// cleanupOrphanedData removes orphaned records that may exist due to the foreign key bug
func (db *database) cleanupOrphanedData() error {
_, err := db.conn.Exec(`
@@ -1188,9 +1234,9 @@ func (db *database) getSettings() (Settings, error) {
var s Settings
err := db.conn.QueryRow(`
SELECT expose, survey, browser, models, agent, tools, working_dir, context_length, turbo_enabled, websearch_enabled, selected_model, sidebar_open, last_home_view, think_enabled, think_level, auto_update_enabled
SELECT expose, survey, browser, models, agent, tools, working_dir, context_length, turbo_enabled, websearch_enabled, selected_model, sidebar_open, last_home_view, onboarding_version, think_enabled, think_level, auto_update_enabled, claude_desktop_used
FROM settings
`).Scan(&s.Expose, &s.Survey, &s.Browser, &s.Models, &s.Agent, &s.Tools, &s.WorkingDir, &s.ContextLength, &s.TurboEnabled, &s.WebSearchEnabled, &s.SelectedModel, &s.SidebarOpen, &s.LastHomeView, &s.ThinkEnabled, &s.ThinkLevel, &s.AutoUpdateEnabled)
`).Scan(&s.Expose, &s.Survey, &s.Browser, &s.Models, &s.Agent, &s.Tools, &s.WorkingDir, &s.ContextLength, &s.TurboEnabled, &s.WebSearchEnabled, &s.SelectedModel, &s.SidebarOpen, &s.LastHomeView, &s.OnboardingVersion, &s.ThinkEnabled, &s.ThinkLevel, &s.AutoUpdateEnabled, &s.ClaudeDesktopUsed)
if err != nil {
return Settings{}, fmt.Errorf("get settings: %w", err)
}
@@ -1200,28 +1246,14 @@ func (db *database) getSettings() (Settings, error) {
func (db *database) setSettings(s Settings) error {
lastHomeView := strings.ToLower(strings.TrimSpace(s.LastHomeView))
validLaunchView := map[string]struct{}{
"launch": {},
"openclaw": {},
"claude": {},
"hermes": {},
"codex": {},
"codex-app": {},
"copilot": {},
"opencode": {},
"droid": {},
"pi": {},
}
if lastHomeView != "chat" {
if _, ok := validLaunchView[lastHomeView]; !ok {
lastHomeView = "launch"
}
lastHomeView = "chat"
}
_, err := db.conn.Exec(`
UPDATE settings
SET expose = ?, survey = ?, browser = ?, models = ?, agent = ?, tools = ?, working_dir = ?, context_length = ?, turbo_enabled = ?, websearch_enabled = ?, selected_model = ?, sidebar_open = ?, last_home_view = ?, think_enabled = ?, think_level = ?, auto_update_enabled = ?
`, s.Expose, s.Survey, s.Browser, s.Models, s.Agent, s.Tools, s.WorkingDir, s.ContextLength, s.TurboEnabled, s.WebSearchEnabled, s.SelectedModel, s.SidebarOpen, lastHomeView, s.ThinkEnabled, s.ThinkLevel, s.AutoUpdateEnabled)
SET expose = ?, survey = ?, browser = ?, models = ?, agent = ?, tools = ?, working_dir = ?, context_length = ?, turbo_enabled = ?, websearch_enabled = ?, selected_model = ?, sidebar_open = ?, last_home_view = ?, onboarding_version = ?, think_enabled = ?, think_level = ?, auto_update_enabled = ?, claude_desktop_used = ?
`, s.Expose, s.Survey, s.Browser, s.Models, s.Agent, s.Tools, s.WorkingDir, s.ContextLength, s.TurboEnabled, s.WebSearchEnabled, s.SelectedModel, s.SidebarOpen, lastHomeView, s.OnboardingVersion, s.ThinkEnabled, s.ThinkLevel, s.AutoUpdateEnabled, s.ClaudeDesktopUsed)
if err != nil {
return fmt.Errorf("set settings: %w", err)
}
+85 -3
View File
@@ -135,7 +135,7 @@ func TestMigrationV13ToV14ContextLength(t *testing.T) {
}
}
func TestMigrationV15ToV16LastHomeViewDefaultsToLaunch(t *testing.T) {
func TestMigrationV15ToV16LastHomeViewMigratesToChat(t *testing.T) {
tmpDir := t.TempDir()
dbPath := filepath.Join(tmpDir, "test.db")
@@ -161,8 +161,8 @@ func TestMigrationV15ToV16LastHomeViewDefaultsToLaunch(t *testing.T) {
t.Fatalf("failed to read last_home_view: %v", err)
}
if lastHomeView != "launch" {
t.Fatalf("expected last_home_view to default to launch after migration, got %q", lastHomeView)
if lastHomeView != "chat" {
t.Fatalf("expected last_home_view to migrate to chat, got %q", lastHomeView)
}
version, err := db.getSchemaVersion()
@@ -174,6 +174,88 @@ func TestMigrationV15ToV16LastHomeViewDefaultsToLaunch(t *testing.T) {
}
}
func TestOnboardingVersionDefaultsAndMigration(t *testing.T) {
t.Run("fresh installs need onboarding", func(t *testing.T) {
dbPath := filepath.Join(t.TempDir(), "fresh.db")
db, err := newDatabase(dbPath)
if err != nil {
t.Fatalf("failed to create database: %v", err)
}
defer db.Close()
settings, err := db.getSettings()
if err != nil {
t.Fatalf("failed to read settings: %v", err)
}
if settings.OnboardingVersion != 0 {
t.Fatalf("expected fresh install onboarding version 0, got %d", settings.OnboardingVersion)
}
})
t.Run("existing installs skip onboarding", func(t *testing.T) {
dbPath := filepath.Join(t.TempDir(), "existing.db")
db, err := newDatabase(dbPath)
if err != nil {
t.Fatalf("failed to create database: %v", err)
}
defer db.Close()
if _, err := db.conn.Exec(`
ALTER TABLE settings DROP COLUMN onboarding_version;
UPDATE settings SET schema_version = 16;
`); err != nil {
t.Fatalf("failed to seed v16 settings row: %v", err)
}
if err := db.migrate(); err != nil {
t.Fatalf("migration from v16 to v17 failed: %v", err)
}
settings, err := db.getSettings()
if err != nil {
t.Fatalf("failed to read settings: %v", err)
}
if settings.OnboardingVersion != 1 {
t.Fatalf("expected existing install onboarding version 1, got %d", settings.OnboardingVersion)
}
})
}
func TestClaudeDesktopUsedDefaultsAndMigration(t *testing.T) {
dbPath := filepath.Join(t.TempDir(), "claude-history.db")
db, err := newDatabase(dbPath)
if err != nil {
t.Fatalf("failed to create database: %v", err)
}
defer db.Close()
settings, err := db.getSettings()
if err != nil {
t.Fatalf("failed to read settings: %v", err)
}
if settings.ClaudeDesktopUsed {
t.Fatal("expected fresh installs to have no Claude Desktop history")
}
if _, err := db.conn.Exec(`
ALTER TABLE settings DROP COLUMN claude_desktop_used;
UPDATE settings SET schema_version = 17;
`); err != nil {
t.Fatalf("failed to seed v17 settings row: %v", err)
}
if err := db.migrate(); err != nil {
t.Fatalf("migration from v17 to v18 failed: %v", err)
}
settings, err = db.getSettings()
if err != nil {
t.Fatalf("failed to read migrated settings: %v", err)
}
if settings.ClaudeDesktopUsed {
t.Fatal("expected existing installs to start with no inferred Claude Desktop history")
}
}
func TestChatDeletionWithCascade(t *testing.T) {
t.Run("chat deletion cascades to related messages", func(t *testing.T) {
tmpDir := t.TempDir()
+8
View File
@@ -57,6 +57,14 @@ func TestConfigMigration(t *testing.T) {
t.Error("expected has completed first run to be true after migration")
}
settings, err := s.Settings()
if err != nil {
t.Fatalf("failed to get settings: %v", err)
}
if settings.OnboardingVersion != CurrentOnboardingVersion {
t.Fatalf("expected migrated user to skip onboarding, got version %d", settings.OnboardingVersion)
}
// Verify migration is marked as complete
migrated, err := s.db.isConfigMigrated()
if err != nil {
+21 -2
View File
@@ -167,13 +167,22 @@ type Settings struct {
// SidebarOpen indicates if the chat sidebar is open
SidebarOpen bool
// LastHomeView stores the preferred home route target ("chat" or integration name)
// LastHomeView is retained for settings compatibility and resolves to chat.
LastHomeView string
// OnboardingVersion stores the latest onboarding flow the user has completed.
OnboardingVersion int
// AutoUpdateEnabled indicates if automatic updates should be downloaded
AutoUpdateEnabled bool
// ClaudeDesktopUsed records whether Claude Desktop has ever been connected through Ollama.
ClaudeDesktopUsed bool
}
// Keep in sync with CURRENT_ONBOARDING_VERSION in app/ui/app/src/lib/onboarding.ts.
const CurrentOnboardingVersion = 1
type Store struct {
// DBPath allows overriding the default database path (mainly for testing)
DBPath string
@@ -334,6 +343,16 @@ func (s *Store) migrateFromConfig(database *database) error {
if err := database.setHasCompletedFirstRun(hasCompleted); err != nil {
return fmt.Errorf("migrate first time run: %w", err)
}
if hasCompleted {
settings, err := database.getSettings()
if err != nil {
return fmt.Errorf("read settings for onboarding migration: %w", err)
}
settings.OnboardingVersion = CurrentOnboardingVersion
if err := database.setSettings(settings); err != nil {
return fmt.Errorf("migrate onboarding completion: %w", err)
}
}
slog.Info("migrated first run status from config.json", "hasCompleted", hasCompleted)
// Mark as migrated
@@ -393,7 +412,7 @@ func (s *Store) Settings() (Settings, error) {
}
if settings.LastHomeView == "" {
settings.LastHomeView = "launch"
settings.LastHomeView = "chat"
}
return settings, nil
+64 -12
View File
@@ -81,18 +81,18 @@ func TestStore(t *testing.T) {
}
})
t.Run("settings default home view is launch", func(t *testing.T) {
t.Run("settings default home view is chat", func(t *testing.T) {
loaded, err := s.Settings()
if err != nil {
t.Fatal(err)
}
if loaded.LastHomeView != "launch" {
t.Fatalf("expected default LastHomeView to be launch, got %q", loaded.LastHomeView)
if loaded.LastHomeView != "chat" {
t.Fatalf("expected default LastHomeView to be chat, got %q", loaded.LastHomeView)
}
})
t.Run("settings empty home view falls back to launch", func(t *testing.T) {
t.Run("settings empty home view falls back to chat", func(t *testing.T) {
if err := s.SetSettings(Settings{LastHomeView: ""}); err != nil {
t.Fatal(err)
}
@@ -102,12 +102,12 @@ func TestStore(t *testing.T) {
t.Fatal(err)
}
if loaded.LastHomeView != "launch" {
t.Fatalf("expected empty LastHomeView to fall back to launch, got %q", loaded.LastHomeView)
if loaded.LastHomeView != "chat" {
t.Fatalf("expected empty LastHomeView to fall back to chat, got %q", loaded.LastHomeView)
}
})
t.Run("settings disabled home view falls back to launch", func(t *testing.T) {
t.Run("settings retired home view falls back to chat", func(t *testing.T) {
if err := s.SetSettings(Settings{LastHomeView: "claude-desktop"}); err != nil {
t.Fatal(err)
}
@@ -117,12 +117,12 @@ func TestStore(t *testing.T) {
t.Fatal(err)
}
if loaded.LastHomeView != "launch" {
t.Fatalf("expected disabled LastHomeView to fall back to launch, got %q", loaded.LastHomeView)
if loaded.LastHomeView != "chat" {
t.Fatalf("expected retired LastHomeView to fall back to chat, got %q", loaded.LastHomeView)
}
})
t.Run("settings codex app home view is accepted", func(t *testing.T) {
t.Run("settings integration home view falls back to chat", func(t *testing.T) {
if err := s.SetSettings(Settings{LastHomeView: "codex-app"}); err != nil {
t.Fatal(err)
}
@@ -132,8 +132,8 @@ func TestStore(t *testing.T) {
t.Fatal(err)
}
if loaded.LastHomeView != "codex-app" {
t.Fatalf("expected codex-app LastHomeView to be preserved, got %q", loaded.LastHomeView)
if loaded.LastHomeView != "chat" {
t.Fatalf("expected integration LastHomeView to fall back to chat, got %q", loaded.LastHomeView)
}
})
@@ -227,6 +227,58 @@ func TestStore(t *testing.T) {
})
}
func TestOnboardingVersionRoundTrip(t *testing.T) {
s, cleanup := setupTestStore(t)
defer cleanup()
settings, err := s.Settings()
if err != nil {
t.Fatal(err)
}
if settings.OnboardingVersion != 0 {
t.Fatalf("expected onboarding version 0 by default, got %d", settings.OnboardingVersion)
}
settings.OnboardingVersion = 1
if err := s.SetSettings(settings); err != nil {
t.Fatal(err)
}
loaded, err := s.Settings()
if err != nil {
t.Fatal(err)
}
if loaded.OnboardingVersion != 1 {
t.Fatalf("expected onboarding version 1, got %d", loaded.OnboardingVersion)
}
}
func TestClaudeDesktopUsedRoundTrip(t *testing.T) {
s, cleanup := setupTestStore(t)
defer cleanup()
settings, err := s.Settings()
if err != nil {
t.Fatal(err)
}
if settings.ClaudeDesktopUsed {
t.Fatal("expected Claude Desktop history to be false by default")
}
settings.ClaudeDesktopUsed = true
if err := s.SetSettings(settings); err != nil {
t.Fatal(err)
}
loaded, err := s.Settings()
if err != nil {
t.Fatal(err)
}
if !loaded.ClaudeDesktopUsed {
t.Fatal("expected Claude Desktop history to persist")
}
}
// setupTestStore creates a temporary store for testing
func setupTestStore(t *testing.T) (*Store, func()) {
t.Helper()
+4
View File
@@ -415,7 +415,9 @@ export class Settings {
SelectedModel: string;
SidebarOpen: boolean;
LastHomeView: string;
OnboardingVersion: number;
AutoUpdateEnabled: boolean;
ClaudeDesktopUsed: boolean;
constructor(source: any = {}) {
if ('string' === typeof source) source = JSON.parse(source);
@@ -434,7 +436,9 @@ export class Settings {
this.SelectedModel = source["SelectedModel"];
this.SidebarOpen = source["SidebarOpen"];
this.LastHomeView = source["LastHomeView"];
this.OnboardingVersion = source["OnboardingVersion"];
this.AutoUpdateEnabled = source["AutoUpdateEnabled"];
this.ClaudeDesktopUsed = source["ClaudeDesktopUsed"];
}
}
export class SettingsResponse {
+3 -2
View File
@@ -1,13 +1,14 @@
<!doctype html>
<html lang="en" style="overflow: hidden">
<html lang="en" style="overflow: hidden; color-scheme: light">
<head>
<meta charset="UTF-8" />
<meta name="color-scheme" content="light" />
<link rel="icon" type="image/svg+xml" href="/vite.svg" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<link rel="stylesheet" href="/src/index.css" />
<title>Ollama</title>
</head>
<body class="dark:bg-neutral-900 select-text">
<body class="bg-white select-text">
<div id="root"></div>
<script type="module" src="/src/main.tsx"></script>
<script>
Binary file not shown.

After

Width:  |  Height:  |  Size: 245 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 21 KiB

After

Width:  |  Height:  |  Size: 10 KiB

+8
View File
@@ -0,0 +1,8 @@
<svg width="92" height="96" viewBox="0 0 92 96" xmlns="http://www.w3.org/2000/svg">
<g fill="#24292F">
<path fill-rule="evenodd" d="M65.45 16.8c10.89 0 19.71 8.86 19.71 19.8v6.6l5.74 11.46a4 4 0 0 1-.01 3.6l-5.73 11.34v6.6c0 10.94-8.82 19.8-19.71 19.8H26.02C15.13 96 6.31 87.14 6.31 76.2v-6.6L.45 58.3a4 4 0 0 1-.01-3.67l5.87-11.43v-6.6c0-10.94 8.82-19.8 19.71-19.8h39.43Zm-2.52 5.7H29.19c-9.32 0-16.87 7.56-16.87 16.88V45L7.44 54.46a4 4 0 0 0 .01 3.68L12.32 67.5v5.63c0 9.32 7.55 16.87 16.87 16.87h33.74c9.32 0 16.87-7.55 16.87-16.87V67.5l4.77-9.39a4 4 0 0 0 .01-3.61L79.8 45v-5.62c0-9.32-7.55-16.88-16.87-16.88Z"/>
<circle cx="45.73" cy="11.5" r="11"/>
<rect x="27" y="41" width="13" height="30" rx="6.5"/>
<rect x="51" y="41" width="13" height="30" rx="6.5"/>
</g>
</svg>

After

Width:  |  Height:  |  Size: 795 B

@@ -0,0 +1,4 @@
<svg xmlns="http://www.w3.org/2000/svg" width="50" height="50" viewBox="0 0 50 50" fill="none">
<style>@media (prefers-color-scheme: dark) { path { fill: #fff; } }</style>
<path d="M48.8354 10.0479C48.3232 9.79199 48.1025 10.2798 47.8032 10.5278C46.7793 11.624 45.9048 12.1597 44.7622 12.0957C43.0923 12 41.666 12.5356 40.4058 13.8398C40.1377 12.2319 39.2476 11.272 37.8926 10.6558C36.4668 10.0156 35.9702 9.31982 35.356 7.72754C35.2456 7.3999 35.1353 7.06396 34.7651 7.00781C34.3633 6.94385 34.2056 7.2876 34.0479 7.57568C33.418 8.75195 33.1733 10.0479 33.1973 11.3599C33.2524 14.312 34.4736 16.6641 36.8999 18.3359C37.1758 18.5278 37.2466 18.7197 37.1597 19C36.9946 19.5757 36.7974 20.1357 36.624 20.7119C36.5137 21.0801 36.3486 21.1597 35.9624 21C32.4092 19.4878 30.0381 16.2319 27.2334 13.52C26.7764 13.1758 26.3193 12.856 25.8467 12.5518C23.8618 10.584 26.1069 8.96777 26.627 8.77588C27.1704 8.57568 26.8159 7.8877 25.0591 7.896C22.8691 7.90381 20.4507 9.06396 18.7095 9.58398C16.8501 9.22363 14.9199 9.14355 12.9033 9.37598C5.30859 10.2397 1.15674 16.4717 1.30664 27.2559C2.11768 31.9521 4.46582 35.8398 8.07373 38.8799C11.8159 42.0322 16.1255 43.5762 21.041 43.2803C24.0269 43.104 27.3516 42.6963 31.1016 39.4561C33.0396 40.1279 37.1758 40.208 38.1211 40.0078C39.6021 39.688 39.4995 38.2881 38.9639 38.0322C34.623 35.9678 35.5762 36.8081 34.71 36.1279C36.9155 33.4639 40.2402 30.6958 41.54 21.728C41.6426 21.0161 41.5557 20.5679 41.54 19.9917C41.5322 19.6396 41.6108 19.5039 42.0049 19.4639C46.6924 18.9116 49.064 15.9038 49.3315 11.2559C49.3711 10.7837 49.3237 10.2959 48.8354 10.0479ZM24.3262 37.8398C20.1196 34.4639 18.0791 33.3521 17.2358 33.3999C16.4482 33.4482 16.5898 34.3682 16.7632 34.9678C16.9443 35.5601 17.1812 35.9683 17.5117 36.4878C17.7402 36.832 17.8979 37.3442 17.2832 37.728C15.9282 38.584 13.5728 37.4399 13.4624 37.3838C7.97949 34.0879 4.48926 28.9282 4.19775 21.3677C4.1582 20.5757 4.38672 20.2959 5.15869 20.1519C11.8945 18.8799 17.165 22.0879 19.2529 25.7759C23.5381 30.104 25.335 35.1523 30.479 39.104C28.8643 39.2881 26.1699 39.3281 24.3262 37.8398ZM26.3433 24.6001C26.3433 24.248 26.6191 23.9678 26.9658 23.9678C27.3042 23.9678 27.5801 24.248 27.5801 24.6001C27.5801 24.9521 27.3042 25.2319 26.9575 25.2319C26.6108 25.2319 26.3433 24.9521 26.3433 24.6001ZM32.6064 27.8799C31.6372 28.2881 30.6289 28.3042 29.8096 27.688C28.6987 26.8555 28.6279 25.7759 28.7305 24.9199C28.8721 24.248 28.7144 23.8159 28.2495 23.4238C27.8716 23.104 27.3911 23.0161 26.8633 23.0161C26.666 23.0161 26.4849 22.9277 26.3511 22.856C25.8467 22.5762 25.9805 22.1758 26.5088 21.688C28.0996 20.7598 29.6362 21.9917 30.834 23.3281C31.6216 24.2559 32.8901 26.312 33.1104 26.9521C33.2446 27.3521 33.0713 27.6802 32.6064 27.8799Z" fill="#000"/>
</svg>

After

Width:  |  Height:  |  Size: 2.7 KiB

@@ -0,0 +1,11 @@
<svg viewBox="0 0 64 64" xmlns="http://www.w3.org/2000/svg">
<defs>
<linearGradient id="omp-gradient" x1="0" y1="0" x2="1" y2="1">
<stop offset="0" stop-color="#ed4abf"/>
<stop offset=".5" stop-color="#9b4dff"/>
<stop offset="1" stop-color="#5ad8e6"/>
</linearGradient>
</defs>
<rect width="64" height="64" rx="12" fill="#0f0a14"/>
<path fill="url(#omp-gradient)" d="M14 16h36v8H40v32h-8V24h-6v22h-8V24h-4z"/>
</svg>

After

Width:  |  Height:  |  Size: 451 B

@@ -0,0 +1,11 @@
<svg viewBox="0 0 64 64" xmlns="http://www.w3.org/2000/svg">
<defs>
<linearGradient id="poolside-gradient" x1="8" y1="5" x2="55" y2="59" gradientUnits="userSpaceOnUse">
<stop stop-color="#6c5cff"/>
<stop offset="1" stop-color="#3c2cff"/>
</linearGradient>
</defs>
<rect width="64" height="64" rx="13" fill="url(#poolside-gradient)"/>
<path d="M13 32c0-10.5 8.5-19 19-19 10.49 0 19 8.5 19 19s-8.51 19-19 19c-10.5 0-19-8.5-19-19Z" fill="none" stroke="#fff" stroke-width="4"/>
<path d="M16 24c8-4.1 17.1-.9 22.6 7.1 4.3-1.2 8.6.5 11 4.1M23.5 47.5 38 17.5" fill="none" stroke="#fff" stroke-linecap="round" stroke-linejoin="round" stroke-width="4"/>
</svg>

After

Width:  |  Height:  |  Size: 682 B

@@ -0,0 +1,3 @@
<svg viewBox="0 0 141.38 140" xmlns="http://www.w3.org/2000/svg">
<path fill="#6D44E8" d="m140.93 85-16.35-28.33-1.93-3.34 8.66-15a3.32 3.32 0 0 0 0-3.34l-9.62-16.67a3.34 3.34 0 0 0-2.89-1.67H82.23l-8.66-15A3.33 3.33 0 0 0 70.68-.02H51.43a3.33 3.33 0 0 0-2.88 1.67L32.19 29.98l-1.92 3.33H12.96a3.34 3.34 0 0 0-2.88 1.67L.45 51.66a3.32 3.32 0 0 0 0 3.34l18.28 31.67-8.66 15a3.32 3.32 0 0 0 0 3.34l9.62 16.67a3.34 3.34 0 0 0 2.89 1.67h36.56l8.66 15a3.35 3.35 0 0 0 2.89 1.67h19.25a3.34 3.34 0 0 0 2.89-1.67l18.28-31.67h17.32a3.34 3.34 0 0 0 2.89-1.67l9.62-16.67a3.32 3.32 0 0 0-.01-3.34ZM51.44 3.33 61.07 20l-9.63 16.66h76.98l-9.62 16.66H45.67l-11.54-20zM57.21 120H22.58l9.63-16.67h19.25l-38.5-66.67h19.25l9.62 16.67L68.78 100l-11.55 20Zm61.59-33.34-9.62-16.67-38.49 66.67-9.63-16.67 9.63-16.66 26.94-46.67h23.1l17.32 30z"/>
</svg>

After

Width:  |  Height:  |  Size: 832 B

+141
View File
@@ -0,0 +1,141 @@
import { afterEach, describe, expect, it, vi } from "vitest";
const { listModels } = vi.hoisted(() => ({ listModels: vi.fn() }));
vi.mock("./lib/ollama-client", () => ({
ollamaClient: { list: listModels },
}));
import {
fetchConnectUrl,
getClaudeDesktopAvailableModels,
getIntegrationStatuses,
} from "./api";
describe("fetchConnectUrl", () => {
afterEach(() => {
vi.unstubAllGlobals();
});
it("requests a desktop handoff after account creation", async () => {
vi.stubGlobal(
"fetch",
vi.fn().mockResolvedValue(
new Response(
JSON.stringify({
signin_url:
"https://ollama.com/connect?name=MacBook&key=public-key",
}),
{ status: 401 },
),
),
);
await expect(fetchConnectUrl()).resolves.toBe(
"https://ollama.com/connect?name=MacBook&key=public-key&launch=true",
);
});
});
describe("getIntegrationStatuses", () => {
afterEach(() => {
vi.unstubAllGlobals();
});
it("returns desktop and launcher integration metadata", async () => {
const fetch = vi.fn().mockResolvedValue(
new Response(
JSON.stringify([
{
id: "claude-desktop",
name: "Claude",
description: "Use Ollama models in Claude Desktop",
installed: true,
},
{
id: "opencode",
name: "OpenCode",
description: "Open-source coding agent",
command: "ollama launch opencode",
},
]),
{ status: 200 },
),
);
vi.stubGlobal("fetch", fetch);
await expect(getIntegrationStatuses()).resolves.toEqual([
{
id: "claude-desktop",
name: "Claude",
description: "Use Ollama models in Claude Desktop",
installed: true,
},
{
id: "opencode",
name: "OpenCode",
description: "Open-source coding agent",
command: "ollama launch opencode",
},
]);
expect(fetch).toHaveBeenCalledWith(
"http://127.0.0.1:3001/api/v1/integrations",
);
});
});
describe("getClaudeDesktopAvailableModels", () => {
afterEach(() => {
listModels.mockReset();
vi.unstubAllGlobals();
vi.restoreAllMocks();
});
it("returns installed local models while pruning remote entries", async () => {
listModels.mockResolvedValue({
models: [
{ name: "llama3.2:latest", digest: "local" },
{
name: "remote-placeholder",
digest: "remote",
remote_host: "https://ollama.com",
},
],
});
const fetch = vi.fn();
vi.stubGlobal("fetch", fetch);
const models = await getClaudeDesktopAvailableModels();
expect(models.map((model) => model.model)).toEqual(["llama3.2"]);
expect(fetch).not.toHaveBeenCalled();
});
it("does not request cloud models when they are unavailable to the user", async () => {
listModels.mockResolvedValue({
models: [
{ name: "qwen3:8b", digest: "local" },
{ name: "deepseek-v4-flash:cloud", digest: "cached-cloud" },
{ name: "gemma4:31b-cloud", digest: "legacy-cached-cloud" },
],
});
const fetch = vi.fn();
vi.stubGlobal("fetch", fetch);
const models = await getClaudeDesktopAvailableModels();
expect(models.map((model) => model.model)).toEqual(["qwen3:8b"]);
expect(fetch).not.toHaveBeenCalled();
});
it("does not use the global cloud model list", async () => {
listModels.mockResolvedValue({
models: [{ name: "qwen3:8b", digest: "local" }],
});
const fetch = vi.fn().mockRejectedValue(new Error("offline"));
vi.stubGlobal("fetch", fetch);
const models = await getClaudeDesktopAvailableModels();
expect(models.map((model) => model.model)).toEqual(["qwen3:8b"]);
expect(fetch).not.toHaveBeenCalled();
});
});
+71 -2
View File
@@ -32,6 +32,24 @@ export interface CloudStatusResponse {
disabled: boolean;
source: CloudStatusSource;
}
export interface IntegrationStatus {
id: string;
name: string;
description: string;
installed?: boolean;
command?: string;
}
export type IntegrationStatuses = IntegrationStatus[];
export async function getIntegrationStatuses(): Promise<IntegrationStatuses> {
const response = await fetch(`${API_BASE}/api/v1/integrations`);
if (!response.ok) {
throw new Error(`Failed to fetch integration statuses: ${response.status}`);
}
return response.json();
}
// Helper function to convert Uint8Array to base64
function uint8ArrayToBase64(uint8Array: Uint8Array): string {
const chunkSize = 0x8000; // 32KB chunks to avoid stack overflow
@@ -81,7 +99,9 @@ export async function fetchConnectUrl(): Promise<string> {
if (response.status === 401) {
const data = await response.json();
if (data.signin_url) {
return data.signin_url;
const connectUrl = new URL(data.signin_url);
connectUrl.searchParams.set("launch", "true");
return connectUrl.toString();
}
}
@@ -176,6 +196,53 @@ export async function getModels(query?: string): Promise<Model[]> {
}
}
export async function getClaudeDesktopAvailableModels(): Promise<Model[]> {
try {
const { models: modelsResponse } = await ollama.list();
const seen = new Set<string>();
return modelsResponse
.filter((model: ModelResponse) => {
const response = model as ModelResponse & {
remote_model?: string;
remote_host?: string;
};
const name = model.name.replace(/:latest$/, "");
return (
!response.remote_model &&
!response.remote_host &&
!name.endsWith("cloud")
);
})
.filter((model: ModelResponse) => {
const base = model.name.replace(/:latest$/, "");
if (!base || seen.has(base)) return false;
const families = model.details?.families;
const supported =
!families ||
families.length === 0 ||
!families.every((family: string) =>
family.toLowerCase().includes("bert"),
);
if (supported) seen.add(base);
return supported;
})
.map(
(model: ModelResponse) =>
new Model({
model: model.name.replace(/:latest$/, ""),
digest: model.digest,
modified_at: model.modified_at
? new Date(model.modified_at)
: undefined,
}),
);
} catch (err) {
throw new Error(`Failed to fetch Ollama models: ${err}`);
}
}
export async function getModelCapabilities(
modelName: string,
): Promise<ModelCapabilitiesResponse> {
@@ -418,7 +485,9 @@ export interface ModelRecommendationsResponse {
recommendations: ModelRecommendation[];
}
export async function getModelRecommendations(): Promise<ModelRecommendation[]> {
export async function getModelRecommendations(): Promise<
ModelRecommendation[]
> {
const response = await fetch(
`${API_BASE}/api/experimental/model-recommendations`,
);
+43
View File
@@ -0,0 +1,43 @@
import { Link } from "@/components/ui/link";
import { ChatIcon } from "@/components/ChatIcon";
import { Cog6ToothIcon, RectangleGroupIcon } from "@heroicons/react/24/outline";
type AppSection = "apps" | "chat" | "settings";
export function AppNavigation({ current }: { current: AppSection }) {
const itemClass = (section: AppSection) =>
`flex w-full items-center gap-3 rounded-lg px-2 py-2 text-left text-sm text-neutral-700 hover:bg-neutral-100 dark:text-neutral-100 dark:hover:bg-neutral-800 ${
current === section ? "bg-neutral-100 dark:bg-neutral-800" : ""
}`;
return (
<div className="flex flex-col gap-0.5">
<Link to="/connect" className={itemClass("apps")} draggable={false}>
<RectangleGroupIcon className="h-5 w-5 stroke-current" />
<span className="truncate">Apps</span>
</Link>
<Link
to="/c/$chatId"
params={{ chatId: "new" }}
mask={{ to: "/" }}
className={itemClass("chat")}
draggable={false}
>
<ChatIcon />
<span className="truncate">Chat</span>
</Link>
<Link to="/settings" className={itemClass("settings")} draggable={false}>
<Cog6ToothIcon className="h-5 w-5 stroke-current" />
<span className="truncate">Settings</span>
</Link>
</div>
);
}
export function AppSidebar({ current }: { current: AppSection }) {
return (
<nav className="flex flex-1 flex-col px-4 pb-4 select-none">
<AppNavigation current={current} />
</nav>
);
}
+14
View File
@@ -0,0 +1,14 @@
export function ChatIcon({ className = "h-5 w-5" }: { className?: string }) {
return (
<svg
aria-hidden="true"
className={`${className} fill-current`}
viewBox="0 0 24 24"
fill="none"
xmlns="http://www.w3.org/2000/svg"
>
<path d="M17.0859 3.39949L15.2135 5.27196H7.27028C5.78649 5.27196 4.94684 6.11336 4.94684 7.59716V16.664C4.94684 18.1558 5.78649 18.9892 7.27028 18.9892H16.3406C17.8324 18.9892 18.6623 18.1558 18.6623 16.664V8.79514L20.5428 6.9115C20.567 7.11532 20.5773 7.33066 20.5773 7.55419V16.7149C20.5773 19.4069 19.0818 20.9024 16.3898 20.9024H7.22107C4.53708 20.9024 3.03357 19.4069 3.03357 16.7149V7.55419C3.03357 4.8622 4.53708 3.35869 7.22107 3.35869H16.3898C16.6329 3.35869 16.8662 3.37094 17.0859 3.39949Z" />
<path d="M9.92714 14.381L11.914 13.5403L20.8312 4.63114L19.3404 3.1581L10.433 12.0655L9.55234 13.9964C9.45664 14.2169 9.70293 14.4714 9.92714 14.381ZM21.5767 3.89364L22.2588 3.19384C22.6347 2.80184 22.6435 2.2663 22.2711 1.90536L22.0148 1.64287C21.6822 1.31377 21.1334 1.36513 20.7689 1.72158L20.0859 2.39833L21.5767 3.89364Z" />
</svg>
);
}
+9 -56
View File
@@ -6,14 +6,12 @@ import { getChat } from "@/api";
import { Link } from "@/components/ui/link";
import { useState, useRef, useEffect, useCallback, useMemo } from "react";
import { ChatsResponse } from "@/gotypes";
import { CogIcon, RocketLaunchIcon } from "@heroicons/react/24/outline";
import { AppNavigation } from "@/components/AppSidebar";
// there's a hidden debug feature to copy a chat's data to the clipboard by
// holding shift and clicking this many times within this many seconds
const DEBUG_SHIFT_CLICKS_REQUIRED = 5;
const DEBUG_SHIFT_CLICK_WINDOW_MS = 7000; // 7 seconds
const launchSidebarRequestedKey = "ollama.launchSidebarRequested";
interface ChatSidebarProps {
currentChatId?: string;
}
@@ -260,56 +258,10 @@ export function ChatSidebar({ currentChatId }: ChatSidebarProps) {
);
}
const isWindows = navigator.platform.toLowerCase().includes("win");
return (
<nav className="flex flex-1 flex-col min-h-0 select-none">
<header className="flex flex-col gap-0.5 px-4 pb-2">
<Link
href="/c/new"
mask={{ to: "/" }}
className={`flex w-full items-center gap-3 rounded-lg px-2 py-2 text-left text-sm text-neutral-700 hover:bg-neutral-100 dark:hover:bg-neutral-800 dark:text-neutral-100 ${currentChatId === "new" ? "bg-neutral-100 dark:bg-neutral-800" : ""
}`}
draggable={false}
>
<svg
className="h-5 w-5 fill-current"
viewBox="0 0 24 24"
fill="none"
xmlns="http://www.w3.org/2000/svg"
>
<path d="M17.0859 3.39949L15.2135 5.27196H7.27028C5.78649 5.27196 4.94684 6.11336 4.94684 7.59716V16.664C4.94684 18.1558 5.78649 18.9892 7.27028 18.9892H16.3406C17.8324 18.9892 18.6623 18.1558 18.6623 16.664V8.79514L20.5428 6.9115C20.567 7.11532 20.5773 7.33066 20.5773 7.55419V16.7149C20.5773 19.4069 19.0818 20.9024 16.3898 20.9024H7.22107C4.53708 20.9024 3.03357 19.4069 3.03357 16.7149V7.55419C3.03357 4.8622 4.53708 3.35869 7.22107 3.35869H16.3898C16.6329 3.35869 16.8662 3.37094 17.0859 3.39949Z" />
<path d="M9.92714 14.381L11.914 13.5403L20.8312 4.63114L19.3404 3.1581L10.433 12.0655L9.55234 13.9964C9.45664 14.2169 9.70293 14.4714 9.92714 14.381ZM21.5767 3.89364L22.2588 3.19384C22.6347 2.80184 22.6435 2.2663 22.2711 1.90536L22.0148 1.64287C21.6822 1.31377 21.1334 1.36513 20.7689 1.72158L20.0859 2.39833L21.5767 3.89364Z" />
</svg>
<span className="truncate">New Chat</span>
</Link>
<Link
to="/c/$chatId"
params={{ chatId: "launch" }}
onClick={() => {
if (currentChatId !== "launch") {
sessionStorage.setItem(launchSidebarRequestedKey, "1");
}
}}
className={`flex w-full items-center gap-3 rounded-lg px-2 py-2 text-left text-sm text-neutral-700 hover:bg-neutral-100 dark:hover:bg-neutral-800 dark:text-neutral-100 cursor-pointer ${currentChatId === "launch"
? "bg-neutral-100 dark:bg-neutral-800"
: ""
}`}
draggable={false}
>
<RocketLaunchIcon className="h-5 w-5 stroke-current" />
<span className="truncate">Launch</span>
</Link>
{isWindows && (
<Link
href="/settings"
className={`flex w-full items-center gap-3 rounded-lg px-2 py-2 text-left text-sm text-neutral-700 hover:bg-neutral-100 dark:hover:bg-neutral-800 dark:text-neutral-300`}
draggable={false}
>
<CogIcon className="h-5 w-5 stroke-current" />
<span className="truncate">Settings</span>
</Link>
)}
<AppNavigation current="chat" />
</header>
<div className="flex flex-1 flex-col px-4 py-1 overflow-y-auto overscroll-auto scrollbar-gutter">
<div className="flex flex-col gap-3 pt-4">
@@ -321,18 +273,19 @@ export function ChatSidebar({ currentChatId }: ChatSidebarProps) {
{group.chats.map((chat) => (
<div
key={chat.id}
className={`allow-context-menu flex items-center relative text-sm text-neutral-800 dark:text-neutral-400 rounded-lg hover:bg-neutral-100 dark:hover:bg-neutral-800 ${chat.id === currentChatId
? "bg-neutral-100 text-black dark:bg-neutral-800"
: ""
}`}
className={`allow-context-menu flex items-center relative text-sm text-neutral-800 dark:text-neutral-400 rounded-lg hover:bg-neutral-100 dark:hover:bg-neutral-800 ${
chat.id === currentChatId
? "bg-neutral-100 text-black dark:bg-neutral-800"
: ""
}`}
onMouseEnter={() => handleMouseEnter(chat.id)}
onContextMenu={(e) =>
handleContextMenu(
e,
chat.id,
chat.title ||
chat.userExcerpt ||
chat.createdAt.toLocaleString(),
chat.userExcerpt ||
chat.createdAt.toLocaleString(),
)
}
>
@@ -0,0 +1,424 @@
import { renderToStaticMarkup } from "react-dom/server";
import { describe, expect, it } from "vitest";
import { ClaudeDesktopModelsSettings } from "./ClaudeDesktopModelsSettings";
describe("ClaudeDesktopModelsSettings", () => {
it("shows Claude recommendations and an installed-model search in Settings", () => {
const html = renderToStaticMarkup(
<ClaudeDesktopModelsSettings
initialLocalModels={["llama3.2", "qwen3:8b"]}
initialStatus={{
supported: true,
used: true,
installed: true,
connected: true,
running: false,
startFailed: false,
portConflict: false,
modelSource: "endpoint",
models: [
{
name: "glm-5.2:cloud",
displayName: "glm-5.2:cloud",
cloud: true,
selected: true,
},
{
name: "deepseek-v4-flash:cloud",
displayName: "deepseek-v4-flash:cloud",
cloud: true,
selected: false,
},
],
}}
/>,
);
expect(html).toContain(">Claude<");
expect(html).toContain(">Apps<");
expect(html).toContain('id="apps-settings-heading"');
expect(html).toContain('src="/launch-icons/claude.svg"');
expect(html).not.toContain("Models in Claude");
expect(html).toContain("glm-5.2:cloud");
expect(html).toContain("deepseek-v4-flash:cloud");
expect(html).toContain("Search Ollama models");
expect(html).not.toContain("Add any Ollama model");
expect(html).toContain("Restart Claude");
expect((html.match(/checked=""/g) ?? []).length).toBe(2);
});
it("does not show the invalid Ollama Cloud sentinel", () => {
const html = renderToStaticMarkup(
<ClaudeDesktopModelsSettings
initialStatus={{
supported: true,
used: true,
installed: true,
connected: true,
running: false,
startFailed: false,
portConflict: false,
modelSource: "user",
models: [
{
name: "Ollama Cloud",
displayName: "Ollama Cloud",
selected: true,
},
{
name: "qwen3:8b",
displayName: "qwen3:8b",
selected: true,
},
],
}}
/>,
);
expect(html).not.toContain("Ollama Cloud");
expect(html).toContain("qwen3:8b");
});
it("labels the built-in fallback without exposing MLX", () => {
const html = renderToStaticMarkup(
<ClaudeDesktopModelsSettings
initialStatus={{
supported: true,
used: true,
installed: true,
connected: true,
running: false,
startFailed: false,
portConflict: false,
modelSource: "fallback",
models: [
{
name: "deepseek-v4-flash:0731:cloud",
displayName: "deepseek-v4-flash:0731:cloud",
cloud: true,
selected: true,
},
],
}}
/>,
);
expect(html).toContain("Built-in defaults");
expect(html).toContain("deepseek-v4-flash:0731:cloud");
expect(html).not.toContain("MLX");
});
it("prevents a sixth selection when the literal Claude slots are full", () => {
const html = renderToStaticMarkup(
<ClaudeDesktopModelsSettings
initialStatus={{
supported: true,
used: true,
installed: true,
connected: true,
running: false,
startFailed: false,
portConflict: false,
modelSource: "endpoint",
maxModels: 5,
models: [
{
name: "glm-5.2:cloud",
displayName: "glm-5.2:cloud",
cloud: true,
selected: true,
},
{
name: "kimi-k3:cloud",
displayName: "kimi-k3:cloud",
cloud: true,
selected: true,
},
{
name: "deepseek-v4-pro:cloud",
displayName: "deepseek-v4-pro:cloud",
cloud: true,
selected: true,
},
{
name: "deepseek-v4-flash:cloud",
displayName: "deepseek-v4-flash:cloud",
cloud: true,
selected: true,
},
{
name: "gemma4:26b:cloud",
displayName: "gemma4:26b:cloud",
cloud: true,
selected: true,
},
{
name: "qwen3:8b",
displayName: "qwen3:8b",
selected: false,
},
],
}}
/>,
);
const index = html.indexOf(">qwen3:8b</span>");
expect(index).toBeGreaterThan(-1);
const label = html.slice(html.lastIndexOf("<label", index), index);
expect(label).toContain("disabled");
expect((html.match(/disabled=""/g) ?? []).length).toBe(1);
expect(html).toContain(
"Claude supports up to 5 models. Deselect one to add another.",
);
});
it("honors a smaller maxModels limit from the status", () => {
const html = renderToStaticMarkup(
<ClaudeDesktopModelsSettings
initialStatus={{
supported: true,
used: true,
installed: true,
connected: true,
running: false,
startFailed: false,
portConflict: false,
modelSource: "endpoint",
maxModels: 1,
models: [
{
name: "kimi-k3:cloud",
displayName: "kimi-k3:cloud",
cloud: true,
selected: true,
},
{
name: "qwen3:8b",
displayName: "qwen3:8b",
selected: false,
},
],
}}
/>,
);
const index = html.indexOf(">qwen3:8b</span>");
expect(index).toBeGreaterThan(-1);
const label = html.slice(html.lastIndexOf("<label", index), index);
expect(label).toContain("disabled");
});
it("hides recommendation and selected cloud models when cloud is disabled", () => {
const html = renderToStaticMarkup(
<ClaudeDesktopModelsSettings
initialLocalModels={["qwen3:8b"]}
initialStatus={{
supported: true,
used: true,
installed: true,
connected: false,
running: false,
startFailed: true,
portConflict: false,
error: "Cloud models are off. Select an installed model in Settings.",
modelSource: "endpoint",
models: [
{
name: "glm-5.2:cloud",
displayName: "glm-5.2:cloud",
cloud: true,
selected: true,
availability: "unavailable",
reason: "cloud_off",
},
],
}}
/>,
);
expect(html).not.toContain("glm-5.2:cloud");
expect(html).toContain("Search Ollama models");
expect(html).toContain(
"Cloud models are off. Select an installed model in Settings.",
);
expect(html).not.toContain(
"These models will be available when Claude starts.",
);
expect(html).not.toContain("text-red");
});
it("shows account requirements and prevents selecting unavailable models", () => {
const html = renderToStaticMarkup(
<ClaudeDesktopModelsSettings
initialStatus={{
supported: true,
used: true,
installed: true,
connected: true,
running: false,
startFailed: false,
portConflict: false,
modelSource: "endpoint",
models: [
{
name: "glm-5.2:cloud",
displayName: "glm-5.2:cloud",
cloud: true,
selected: false,
availability: "unavailable",
reason: "upgrade_required",
requiredPlan: "pro",
},
{
name: "gemma4:31b-cloud",
displayName: "gemma4:31b-cloud",
cloud: true,
selected: true,
availability: "available",
requiredPlan: "free",
},
],
}}
/>,
);
expect(html).toContain("pro plan required");
const index = html.indexOf(">glm-5.2:cloud</span>");
const label = html.slice(html.lastIndexOf("<label", index), index);
expect(label).toContain("disabled");
expect(html).toContain("gemma4:31b-cloud");
});
it("replaces paid defaults with the available free recommendation", () => {
const html = renderToStaticMarkup(
<ClaudeDesktopModelsSettings
initialStatus={{
supported: true,
used: true,
installed: true,
connected: false,
running: false,
startFailed: false,
portConflict: false,
modelSource: "endpoint",
models: [
"glm-5.2:cloud",
"kimi-k3:cloud",
"deepseek-v4-pro:cloud",
"deepseek-v4-flash:cloud",
]
.map((name) => ({
name,
displayName: name,
cloud: true,
selected: true,
availability: "unavailable" as const,
reason: "upgrade_required" as const,
requiredPlan: "pro",
}))
.concat([
{
name: "gemma4:31b-cloud",
displayName: "gemma4:31b-cloud",
cloud: true,
selected: false,
availability: "available" as const,
requiredPlan: "free",
},
]),
}}
/>,
);
expect((html.match(/>pro plan required<\/span>/g) ?? []).length).toBe(4);
expect((html.match(/checked=""/g) ?? []).length).toBe(1);
const gemmaIndex = html.indexOf(">gemma4:31b-cloud</span>");
const gemmaInput = html.slice(
html.lastIndexOf("<input", gemmaIndex),
gemmaIndex,
);
expect(gemmaInput).toContain('checked=""');
expect(html).toContain("Start Claude");
});
it("does not restart Claude when every selected model is unavailable", () => {
const html = renderToStaticMarkup(
<ClaudeDesktopModelsSettings
initialStatus={{
supported: true,
used: true,
installed: true,
connected: true,
running: false,
startFailed: false,
portConflict: false,
modelSource: "endpoint",
models: [
{
name: "glm-5.2:cloud",
displayName: "glm-5.2:cloud",
cloud: true,
selected: true,
availability: "unavailable",
reason: "upgrade_required",
requiredPlan: "pro",
},
],
}}
/>,
);
expect(html).toContain("Select a model available to your account.");
const button = html.slice(html.lastIndexOf("<button"));
expect(button).toContain("disabled");
});
it("remains visible after Claude has been disconnected", () => {
const html = renderToStaticMarkup(
<ClaudeDesktopModelsSettings
initialStatus={{
supported: true,
used: true,
installed: true,
connected: false,
running: false,
startFailed: false,
portConflict: false,
modelSource: "user",
models: [
{
name: "qwen3:8b",
displayName: "qwen3:8b",
selected: true,
},
],
}}
/>,
);
expect(html).toContain(">Claude<");
expect(html).toContain("qwen3:8b");
expect(html).toContain(
"These models will be available when Claude starts.",
);
expect(html).toContain("Start Claude");
});
it("stays hidden until Claude has been enabled once", () => {
const html = renderToStaticMarkup(
<ClaudeDesktopModelsSettings
initialStatus={{
supported: true,
used: false,
installed: true,
connected: false,
running: false,
startFailed: false,
portConflict: false,
}}
/>,
);
expect(html).toBe("");
});
});
@@ -0,0 +1,420 @@
import { getClaudeDesktopAvailableModels } from "@/api";
import { Button } from "@/components/ui/button";
import { Input } from "@/components/ui/input";
import {
addClaudeModelSelection,
claudeDesktopRecoveryMessage,
claudeDesktopMaxModels,
claudeDesktopMaxModelsMessage,
claudeDesktopUsableSelection,
} from "@/lib/claudeDesktop";
import type {
ClaudeDesktopModelStatus,
ClaudeDesktopStatus,
} from "@/types/webview";
import { ArrowPathIcon } from "@heroicons/react/20/solid";
import { useCallback, useEffect, useMemo, useRef, useState } from "react";
interface ClaudeDesktopModelsSettingsProps {
initialStatus?: ClaudeDesktopStatus;
initialLocalModels?: string[];
}
function isInvalidModelName(name: string): boolean {
const normalized = name.trim().toLowerCase().replace(/[-:]+/g, " ");
return normalized === "ollama cloud";
}
function visibleModels(
status: ClaudeDesktopStatus,
): ClaudeDesktopModelStatus[] {
return (status.models ?? []).filter(
(model) => !isInvalidModelName(model.name) && model.reason !== "cloud_off",
);
}
function selectedModelNames(status: ClaudeDesktopStatus): string[] {
return claudeDesktopUsableSelection(
visibleModels(status),
status.modelSource !== "user",
claudeDesktopMaxModels(status),
);
}
function modelAccessLabel(model: ClaudeDesktopModelStatus): string | null {
switch (model.reason) {
case "sign_in_required":
return "Sign in required";
case "upgrade_required":
return model.requiredPlan
? `${model.requiredPlan} plan required`
: "Upgrade required";
case "verification_unavailable":
return "Access unavailable";
case "model_not_installed":
return "Not installed";
default:
return null;
}
}
function explicitCloudName(name: string): string {
return name.endsWith(":cloud") ? name : `${name}:cloud`;
}
export function ClaudeDesktopModelsSettings({
initialStatus,
initialLocalModels,
}: ClaudeDesktopModelsSettingsProps) {
const [status, setStatus] = useState<ClaudeDesktopStatus | null>(
initialStatus ?? null,
);
const [models, setModels] = useState<ClaudeDesktopModelStatus[]>(() =>
initialStatus ? visibleModels(initialStatus) : [],
);
const [selection, setSelection] = useState<string[]>(() =>
initialStatus ? selectedModelNames(initialStatus) : [],
);
const [localModels, setLocalModels] = useState<string[]>(
initialLocalModels ?? [],
);
const [searchQuery, setSearchQuery] = useState("");
const [pickerOpen, setPickerOpen] = useState(false);
const [modelsLoading, setModelsLoading] = useState(false);
const [error, setError] = useState<string | null>(null);
const [restarting, setRestarting] = useState(false);
const pickerRef = useRef<HTMLDivElement>(null);
const applyStatus = useCallback((next: ClaudeDesktopStatus) => {
setStatus(next);
setModels(visibleModels(next));
setSelection(selectedModelNames(next));
setError(null);
}, []);
const refreshStatus = useCallback(async () => {
if (!window.getClaudeDesktopStatus) return;
try {
applyStatus(await window.getClaudeDesktopStatus());
} catch {
setError("Ollama could not read the Claude connection status.");
}
}, [applyStatus]);
useEffect(() => {
if (!initialStatus) void refreshStatus();
const handleFocus = () => void refreshStatus();
window.addEventListener("focus", handleFocus);
return () => window.removeEventListener("focus", handleFocus);
}, [initialStatus, refreshStatus]);
useEffect(() => {
if (!status) return;
setModels(visibleModels(status));
setSelection(selectedModelNames(status));
}, [status]);
useEffect(() => {
if (initialLocalModels || !status?.used) return;
let cancelled = false;
setModelsLoading(true);
void getClaudeDesktopAvailableModels()
.then((installed) => {
if (!cancelled) {
setLocalModels(installed.map((model) => model.model));
}
})
.catch(() => {
if (!cancelled) setError("Ollama could not load your models.");
})
.finally(() => {
if (!cancelled) setModelsLoading(false);
});
return () => {
cancelled = true;
};
}, [initialLocalModels, status?.used]);
useEffect(() => {
const handleClickOutside = (event: MouseEvent) => {
if (
pickerRef.current &&
!pickerRef.current.contains(event.target as Node)
) {
setPickerOpen(false);
}
};
document.addEventListener("mousedown", handleClickOutside);
return () => document.removeEventListener("mousedown", handleClickOutside);
}, []);
const matchingLocalModels = useMemo(() => {
const current = new Set(
models.map((model) =>
model.cloud ? explicitCloudName(model.name) : model.name,
),
);
const query = searchQuery.trim().toLowerCase();
return localModels
.filter(
(name) =>
!current.has(name) &&
!isInvalidModelName(name) &&
(!query || name.toLowerCase().includes(query)),
)
.sort((left, right) => left.localeCompare(right));
}, [localModels, models, searchQuery]);
const toggleModel = (name: string) => {
setError(null);
setSelection((current) => {
if (!current.includes(name)) {
const result = addClaudeModelSelection(
current,
name,
claudeDesktopMaxModels(status),
);
if (result.error) {
setError(result.error);
return current;
}
return result.selection;
}
if (current.length === 1) {
setError("Select at least one model for Claude.");
return current;
}
return current.filter((model) => model !== name);
});
};
const addLocalModel = (name: string) => {
const maxModels = claudeDesktopMaxModels(status);
const result = addClaudeModelSelection(selection, name, maxModels);
if (result.error) {
setError(result.error);
return;
}
const cloud = name.endsWith(":cloud");
setModels((current) => [
...current,
{
name,
displayName: name,
cloud,
selected: true,
availability: "available",
},
]);
setSelection(result.selection);
setSearchQuery("");
setPickerOpen(false);
setError(null);
};
const hasAvailableSelection = selection.some((name) => {
const model = models.find((candidate) => candidate.name === name);
return !model?.availability || model.availability === "available";
});
const restartClaude = async () => {
if (!window.restartClaudeDesktop) {
setError("Claude restart is available in the Ollama macOS app.");
return;
}
if (selection.length === 0) {
setError("Select at least one model for Claude.");
return;
}
if (!hasAvailableSelection) {
setError("Select a model available to your account.");
return;
}
if (
status?.running &&
!window.confirm(
"Restart Claude Desktop to update its models? Any running task will stop.",
)
) {
return;
}
setError(null);
setRestarting(true);
try {
const result = await window.restartClaudeDesktop(selection);
applyStatus(result.status);
if (result.error) setError(result.error);
} catch {
setError("Ollama could not restart Claude.");
} finally {
setRestarting(false);
}
};
if (!status?.supported || !status.used) {
return null;
}
const maxModels = claudeDesktopMaxModels(status);
const selectionFull = selection.length >= maxModels;
const guidance =
claudeDesktopRecoveryMessage(status.error, error) ??
(!hasAvailableSelection && models.length > 0
? "Select a model available to your account."
: selectionFull
? claudeDesktopMaxModelsMessage(maxModels)
: status.connected
? "Restart Claude to refresh its model list."
: "These models will be available when Claude starts.");
return (
<section aria-labelledby="apps-settings-heading" className="space-y-2">
<h2
id="apps-settings-heading"
className="px-1 text-xs font-medium uppercase tracking-wider text-neutral-400 dark:text-neutral-500"
>
Apps
</h2>
<div
aria-labelledby="claude-models-settings-heading"
className="overflow-visible rounded-xl bg-white p-4 dark:bg-neutral-800"
>
<div className="flex items-start space-x-3">
<img
src="/launch-icons/claude.svg"
alt=""
className="mt-0.5 h-5 w-5 flex-shrink-0"
/>
<div className="min-w-0 flex-1">
<div className="flex items-center justify-between gap-3">
<h2
id="claude-models-settings-heading"
className="text-sm font-medium text-neutral-900 dark:text-white"
>
Claude
</h2>
{status.modelSource === "fallback" && models.length > 0 && (
<span className="text-xs text-neutral-400">
Built-in defaults
</span>
)}
</div>
<div className="mt-3 grid grid-cols-2 gap-x-5 gap-y-2 max-[850px]:grid-cols-1">
{models.map((model) => {
const selected = selection.includes(model.name);
const accessLabel = modelAccessLabel(model);
const unavailable =
model.availability !== undefined &&
model.availability !== "available";
return (
<label
key={model.name}
className="flex min-w-0 cursor-pointer items-center gap-2 rounded-md py-1 text-sm text-neutral-600 dark:text-neutral-300"
title={accessLabel ?? model.description}
>
<input
type="checkbox"
checked={selected}
disabled={
restarting ||
(!selected && (selectionFull || unavailable))
}
onChange={() => toggleModel(model.name)}
className="h-4 w-4 rounded border-neutral-300 accent-neutral-900 dark:border-neutral-600 dark:accent-white"
/>
<span className="truncate">{model.displayName}</span>
{accessLabel && (
<span className="flex-shrink-0 text-xs text-neutral-400">
{accessLabel}
</span>
)}
</label>
);
})}
</div>
<div ref={pickerRef} className="relative mt-3">
<Input
type="search"
value={searchQuery}
onFocus={() => setPickerOpen(true)}
onChange={(event) => {
setSearchQuery(event.target.value);
setPickerOpen(true);
}}
placeholder="Search Ollama models"
aria-label="Search Ollama models"
aria-expanded={pickerOpen}
aria-controls="claude-local-models"
disabled={restarting}
autoComplete="off"
/>
{pickerOpen && (
<div
id="claude-local-models"
className="absolute z-20 mt-1 max-h-52 w-full overflow-y-auto rounded-lg border border-neutral-200 bg-white p-1 shadow-lg dark:border-neutral-600 dark:bg-neutral-800"
>
{modelsLoading ? (
<p className="px-3 py-2 text-sm text-neutral-400">
Loading models
</p>
) : selectionFull ? (
<p className="px-3 py-2 text-sm text-neutral-400">
{claudeDesktopMaxModelsMessage(maxModels)}
</p>
) : matchingLocalModels.length > 0 ? (
matchingLocalModels.map((name) => (
<button
key={name}
type="button"
onClick={() => addLocalModel(name)}
className="block w-full rounded-md px-3 py-2 text-left text-sm text-neutral-700 hover:bg-neutral-100 dark:text-neutral-200 dark:hover:bg-neutral-700"
>
{name}
</button>
))
) : (
<p className="px-3 py-2 text-sm text-neutral-400">
No models found.
</p>
)}
</div>
)}
</div>
<div className="mt-3 flex items-center justify-between gap-4 max-sm:flex-col max-sm:items-stretch">
<p
role={error || status.error ? "alert" : undefined}
className="text-xs leading-5 text-neutral-500 dark:text-neutral-400"
>
{guidance}
</p>
<Button
type="button"
color="white"
onClick={restartClaude}
disabled={
restarting || selection.length === 0 || !hasAvailableSelection
}
className="flex-shrink-0 max-sm:w-full"
>
{restarting && (
<ArrowPathIcon data-slot="icon" className="animate-spin" />
)}
{restarting
? status.connected
? "Restarting…"
: "Starting…"
: status.connected
? "Restart Claude"
: "Start Claude"}
</Button>
</div>
</div>
</div>
</div>
</section>
);
}
+1 -1
View File
@@ -68,7 +68,7 @@ const CopyButton: React.FC<CopyButtonProps> = ({
const iconSize = size === "sm" ? "h-3 w-3" : "h-7 w-7";
const baseClasses =
size === "sm"
? `text-xs px-4 py-2 z-10 rounded-lg hover:cursor-pointer ${className}`
? `text-xs px-4 py-2 z-10 cursor-pointer rounded-lg ${className}`
: `${iconSize} px-1 py-0.5 text-xs cursor-pointer rounded-lg hover:bg-neutral-100 dark:hover:bg-neutral-800 flex items-center justify-center ${className}`;
const icon = isCopied ? (
@@ -1,158 +0,0 @@
import { useSettings } from "@/hooks/useSettings";
import CopyButton from "@/components/CopyButton";
interface LaunchCommand {
id: string;
name: string;
command: string;
description: string;
icon: string;
darkIcon?: string;
iconClassName?: string;
borderless?: boolean;
}
const LAUNCH_COMMANDS: LaunchCommand[] = [
{
id: "claude",
name: "Claude Code",
command: "ollama launch claude",
description: "Anthropic's coding tool with subagents",
icon: "/launch-icons/claude-code.svg",
iconClassName: "h-7 w-7",
},
{
id: "codex-app",
name: "Codex App",
command: "ollama launch codex-app",
description: "An AI agent you can delegate real work to, by OpenAI",
icon: "/launch-icons/codex-app.png",
iconClassName: "h-full w-full",
},
{
id: "hermes",
name: "Hermes Agent",
command: "ollama launch hermes",
description: "Self-improving AI agent built by Nous Research",
icon: "/launch-icons/hermes-agent.svg",
iconClassName: "h-7 w-7",
},
{
id: "openclaw",
name: "OpenClaw",
command: "ollama launch openclaw",
description: "Personal AI with 100+ skills",
icon: "/launch-icons/openclaw.svg",
},
{
id: "opencode",
name: "OpenCode",
command: "ollama launch opencode",
description: "Anomaly's open-source coding agent",
icon: "/launch-icons/opencode.svg",
iconClassName: "h-7 w-7 rounded",
},
{
id: "codex",
name: "Codex",
command: "ollama launch codex",
description: "OpenAI's open-source coding agent",
icon: "/launch-icons/codex.svg",
darkIcon: "/launch-icons/codex-dark.svg",
iconClassName: "h-7 w-7",
},
{
id: "copilot",
name: "Copilot CLI",
command: "ollama launch copilot",
description: "GitHub's AI coding agent for the terminal",
icon: "/launch-icons/copilot.svg",
darkIcon: "/launch-icons/copilot-dark.svg",
iconClassName: "h-7 w-7",
},
{
id: "droid",
name: "Droid",
command: "ollama launch droid",
description: "Factory's coding agent across terminal and IDEs",
icon: "/launch-icons/droid.svg",
},
{
id: "pi",
name: "Pi",
command: "ollama launch pi",
description: "Minimal AI agent toolkit with plugin support",
icon: "/launch-icons/pi.svg",
darkIcon: "/launch-icons/pi-dark.svg",
iconClassName: "h-7 w-7",
},
];
export default function LaunchCommands() {
const isWindows = navigator.platform.toLowerCase().includes("win");
const { setSettings } = useSettings();
const renderCommandCard = (item: LaunchCommand) => (
<div key={item.command} className="w-full text-left">
<div className="flex items-start gap-4 sm:gap-5">
<div
aria-hidden="true"
className={`flex h-10 w-10 shrink-0 items-center justify-center rounded-lg overflow-hidden ${item.borderless ? "" : "border border-neutral-200 bg-white dark:border-neutral-700 dark:bg-neutral-900"}`}
>
{item.darkIcon ? (
<picture>
<source srcSet={item.darkIcon} media="(prefers-color-scheme: dark)" />
<img src={item.icon} alt="" className={`${item.iconClassName ?? "h-8 w-8"} rounded-sm`} />
</picture>
) : (
<img src={item.icon} alt="" className={item.borderless ? "h-full w-full rounded-xl" : `${item.iconClassName ?? "h-8 w-8"} rounded-sm`} />
)}
</div>
<div className="min-w-0 flex-1">
<span className="text-sm font-medium text-neutral-900 dark:text-neutral-100">
{item.name}
</span>
<p className="mt-0.5 text-xs text-neutral-500 dark:text-neutral-400">
{item.description}
</p>
<div className="mt-2 flex items-center gap-2 rounded-xl border-neutral-200 dark:border-neutral-700 bg-neutral-50 dark:bg-neutral-800 px-3 py-2">
<code className="min-w-0 flex-1 truncate text-xs text-neutral-600 dark:text-neutral-300">
{item.command}
</code>
<CopyButton
content={item.command}
size="md"
title="Copy command to clipboard"
className="text-neutral-500 dark:text-neutral-400 hover:text-neutral-700 dark:hover:text-neutral-200 hover:bg-neutral-200/60 dark:hover:bg-neutral-700/70"
onCopy={() => {
setSettings({ LastHomeView: item.id }).catch(() => { });
}}
/>
</div>
</div>
</div>
</div>
);
return (
<main className="flex h-screen w-full flex-col relative">
<section
className={`flex-1 overflow-y-auto overscroll-contain relative min-h-0 ${isWindows ? "xl:pt-4" : "xl:pt-8"}`}
>
<div className="max-w-[730px] mx-auto w-full px-4 pt-4 pb-20 sm:px-6 sm:pt-6 sm:pb-24 lg:px-8 lg:pt-8 lg:pb-28">
<h1 className="text-xl font-semibold text-neutral-900 dark:text-neutral-100">
Launch
</h1>
<p className="mt-1 text-sm text-neutral-500 dark:text-neutral-400">
Copy a command and run it in your terminal.
</p>
<div className="mt-6 grid gap-7">
{LAUNCH_COMMANDS.map(renderCommandCard)}
</div>
</div>
</section>
</main>
);
}
File diff suppressed because one or more lines are too long.
@@ -0,0 +1,597 @@
import { renderToStaticMarkup } from "react-dom/server";
import { describe, expect, it, vi } from "vitest";
import {
FIRST_MODEL_COMMAND,
ConnectAppsScreen,
IntroScreen,
default as Onboarding,
RunOllamaScreen,
terminalRowsForWindowHeight,
WelcomeScreen,
} from "./Onboarding";
import {
CLAUDE_INSTALL_TIMEOUT_MS,
isClaudeConnectionComplete,
scheduleClaudeInstallTimeout,
} from "@/lib/claudeDesktop";
import { isWindowsPlatform } from "@/lib/platform";
import {
authenticationTimeoutAction,
nextOnboardingStep,
onboardingConnectUrl,
} from "@/lib/onboarding";
import type { IntegrationStatuses } from "@/api";
describe("Onboarding", () => {
it("explains what Ollama is before asking the user to choose a path", () => {
const html = renderToStaticMarkup(<IntroScreen onContinue={vi.fn()} />);
expect(html).toContain("Welcome to Ollama!");
expect(html.indexOf('alt="Ollama waving"')).toBeLessThan(
html.indexOf("Welcome to Ollama!"),
);
expect(html).toContain(
"Run open models with your coding agents so you can spend less while keeping your data private.",
);
expect(html.indexOf("Connect your apps")).toBeLessThan(
html.indexOf("Easily switch models"),
);
expect(html.indexOf("Easily switch models")).toBeLessThan(
html.indexOf("Your data stays yours"),
);
expect(html).toContain("Power your existing coding apps with open models");
expect(html).toContain("Swap between frontier models in one click.");
expect(html).toContain("Your prompt data is never logged or trained on.");
expect(html).toContain("Continue");
expect(html).not.toContain("Skip");
});
it("shows more terminal integrations as the window gets taller", () => {
expect(terminalRowsForWindowHeight(400)).toBe(1);
expect(terminalRowsForWindowHeight(660)).toBe(4);
expect(terminalRowsForWindowHeight(960)).toBe(8);
});
it("renders the apps screen without browser platform globals", () => {
vi.stubGlobal("navigator", undefined);
try {
expect(() =>
renderToStaticMarkup(<ConnectAppsScreen initialIntegrations={[]} />),
).not.toThrow();
} finally {
vi.unstubAllGlobals();
}
});
it("hides the Claude application on Windows", () => {
vi.stubGlobal("window", {
OLLAMA_PLATFORM: "windows",
innerHeight: 660,
});
vi.stubGlobal("navigator", { platform: "MacIntel" });
try {
expect(isWindowsPlatform()).toBe(true);
const html = renderToStaticMarkup(
<ConnectAppsScreen
initialIntegrations={[
{
id: "claude-desktop",
name: "Claude",
description: "Use Ollama models in Claude Desktop",
installed: true,
},
{
id: "claude",
name: "Claude Code",
description: "Anthropic's coding tool with subagents",
command: "ollama launch claude",
},
]}
/>,
);
expect(html).not.toContain('id="applications-heading"');
expect(html).not.toContain("Use Ollama models in Claude Desktop");
expect(html).toContain('id="terminal-heading"');
expect(html).toContain("ollama launch claude");
} finally {
vi.unstubAllGlobals();
}
});
it("shows the account choice only to signed-out users", () => {
expect(nextOnboardingStep("intro", "continue", false)).toBe("welcome");
expect(nextOnboardingStep("intro", "continue", true)).toBe("apps");
expect(nextOnboardingStep("welcome", "authenticated", true)).toBe("apps");
expect(nextOnboardingStep("apps", "continue", true)).toBe("apps");
expect(nextOnboardingStep("welcome", "local", false)).toBe("run");
});
it("lets an in-flight authentication check finish before timing out", () => {
expect(authenticationTimeoutAction(false, true)).toBe("defer");
expect(authenticationTimeoutAction(false, false)).toBe("fail");
expect(authenticationTimeoutAction(true, true)).toBe("ignore");
});
it("finishes Claude connection states from the native status hook", () => {
const status = {
supported: true,
installed: true,
configured: true,
connected: true,
running: false,
startFailed: false,
portConflict: false,
};
expect(isClaudeConnectionComplete(true, status)).toBe(true);
expect(
isClaudeConnectionComplete(true, { ...status, connected: false }),
).toBe(false);
expect(
isClaudeConnectionComplete(true, { ...status, startFailed: true }),
).toBe(false);
expect(
isClaudeConnectionComplete(false, {
...status,
configured: false,
connected: false,
}),
).toBe(true);
expect(
isClaudeConnectionComplete(false, { ...status, connected: false }),
).toBe(false);
});
it("bounds the Claude installer wait", () => {
vi.useFakeTimers();
vi.stubGlobal("window", { setTimeout: globalThis.setTimeout });
const onTimeout = vi.fn();
try {
scheduleClaudeInstallTimeout(onTimeout);
vi.advanceTimersByTime(CLAUDE_INSTALL_TIMEOUT_MS - 1);
expect(onTimeout).not.toHaveBeenCalled();
vi.advanceTimersByTime(1);
expect(onTimeout).toHaveBeenCalledOnce();
} finally {
vi.useRealTimers();
vi.unstubAllGlobals();
}
});
it("opens the device connection flow without relaunching the app", () => {
expect(
onboardingConnectUrl(
"https://ollama.com/connect?name=MacBook&key=public-key&launch=true",
"signin",
),
).toBe("https://ollama.com/connect?name=MacBook&key=public-key");
expect(
onboardingConnectUrl(
"https://ollama.com/connect?name=MacBook&key=public-key",
"signup",
),
).toBe(
"https://ollama.com/connect?name=MacBook&key=public-key&signup=true",
);
});
it("preserves the intro for a device that is already connected", () => {
const html = renderToStaticMarkup(
<Onboarding
isAuthenticated
isSigningIn={false}
signInError={null}
completionError={null}
onOpenApps={vi.fn().mockResolvedValue(true)}
onSignIn={vi.fn()}
onSignUp={vi.fn()}
onRetryCompletion={vi.fn()}
onUseLocal={vi.fn()}
/>,
);
expect(html).toContain("Welcome to Ollama");
expect(html).not.toContain("Run Ollama");
expect(html).not.toContain("Sign up");
});
it("groups disconnected Claude with applications and terminal separately", () => {
const integrations: IntegrationStatuses = [
{
id: "claude-desktop",
name: "Claude",
description: "Use Ollama models in Claude Desktop",
installed: true,
action: "connect",
},
{
id: "claude",
name: "Claude Code",
description: "Anthropic's coding tool with subagents",
installed: true,
action: "copy",
command: "ollama launch claude",
},
{
id: "codex",
name: "Codex",
description: "OpenAI's open-source coding agent",
installed: true,
action: "copy",
command: "ollama launch codex",
},
{
id: "openclaw",
name: "OpenClaw",
description: "Personal AI with 100+ skills",
installed: true,
action: "copy",
command: "ollama launch openclaw",
},
{
id: "opencode",
name: "OpenCode",
description: "Anomaly's open-source coding agent",
installed: false,
action: "copy",
command: "ollama launch opencode",
},
{
id: "droid",
name: "Droid",
description: "AI software engineering agent",
installed: false,
action: "copy",
command: "ollama launch droid",
},
{
id: "terminal",
name: "Terminal",
description: "Run local models from your terminal",
action: "copy",
command: "ollama",
},
];
const html = renderToStaticMarkup(
<ConnectAppsScreen
completionError={null}
onRetryCompletion={vi.fn()}
initialIntegrations={integrations}
/>,
);
expect(html).not.toContain(
"Connect Claude, or copy a command to run in your terminal.",
);
expect(html).toContain("Claude");
expect(html).toContain("Use Ollama models in Claude Desktop");
expect(html).toContain("Claude Code");
expect(html).not.toContain("Search apps");
expect(html).not.toContain('type="search"');
expect(html).toContain("Application");
expect(html).toContain('id="applications-heading"');
expect(html).toContain('id="terminal-heading"');
expect(html).not.toContain("Ready to launch");
expect(html).not.toContain('id="claude-apps-heading"');
expect(html.indexOf("Application")).toBeLessThan(
html.indexOf("Use Ollama models in Claude Desktop"),
);
expect(html).not.toContain(">Command</th>");
expect(html).toContain("ollama launch claude");
expect(html).not.toContain("Installed");
expect(html).not.toContain("Not installed");
expect(html).toContain('aria-label="Connect Claude"');
expect(html).toContain('role="switch"');
expect(html).toContain('aria-checked="false"');
expect(html).not.toContain("Inactive");
expect(html).not.toContain("Download &amp; connect");
expect(html).not.toContain("Active");
expect(html).toContain("bg-transparent");
expect(html).toContain('aria-label="Copy OpenCode command"');
expect(html).toContain('aria-label="Copy Terminal command"');
expect(html).not.toContain(">Copy command</button>");
expect(html).not.toContain("ChatGPT");
expect(html).toContain("OpenCode");
expect(html).toContain("Terminal");
expect(html).toContain('aria-label="Show more apps"');
expect(html).toContain('aria-expanded="false"');
expect(html).toContain("grid-rows-[0fr]");
expect(html).not.toContain("Collapse");
expect(html).toContain("/launch-icons/claude.svg");
expect(html).toContain("/launch-icons/claude-code.svg");
expect(html).not.toContain("<table");
expect(html).not.toContain("<footer");
expect(html).not.toContain("Command copied. Run it in your terminal.");
expect(html).toContain("Run local models from your terminal");
expect(html).not.toContain("Launch command");
expect(html).not.toContain('aria-pressed="true"');
expect(html).not.toContain("Continue");
expect(html).not.toContain("Run Ollama");
expect(html).not.toContain('viewBox="0 0 3400 3400"');
});
it("keeps connected Claude in Application without an idle status", () => {
const html = renderToStaticMarkup(
<ConnectAppsScreen
completionError={null}
onRetryCompletion={vi.fn()}
initialClaudeStatus={{
supported: true,
used: true,
installed: true,
connected: true,
running: false,
startFailed: false,
portConflict: false,
}}
initialIntegrations={[
{
id: "claude-desktop",
name: "Claude",
description: "Use Ollama models in Claude Desktop",
installed: true,
action: "connect",
},
{
id: "codex",
name: "Codex",
description: "OpenAI's open-source coding agent",
installed: true,
action: "copy",
command: "ollama launch codex",
},
]}
/>,
);
expect(html).toContain('id="applications-heading"');
expect(html).not.toContain('id="claude-apps-heading"');
expect(html).not.toContain("Ready to launch");
expect(html).not.toContain("Active");
expect(html).not.toContain("Inactive");
expect(html).toContain('aria-checked="true"');
expect(html).toContain('aria-label="Disconnect Claude"');
});
it("shows initial Claude recovery guidance without error styling", () => {
const html = renderToStaticMarkup(
<ConnectAppsScreen
completionError={null}
onRetryCompletion={vi.fn()}
initialClaudeStatus={{
supported: true,
used: true,
installed: true,
configured: true,
connected: false,
running: false,
startFailed: true,
portConflict: false,
error: "Cloud models are off. Select an installed model in Settings.",
}}
initialIntegrations={[
{
id: "claude-desktop",
name: "Claude",
description: "Use Ollama models in Claude Desktop",
installed: true,
action: "connect",
},
]}
/>,
);
expect(html).toContain(
"Cloud models are off. Select an installed model in Settings.",
);
expect(html).toContain('role="alert"');
expect(html).not.toContain("text-red");
expect(html).toContain('aria-checked="true"');
expect(html).toContain('aria-label="Disconnect Claude"');
});
it("keeps Claude model management off the Connect Apps page", () => {
const html = renderToStaticMarkup(
<ConnectAppsScreen
completionError={null}
onRetryCompletion={vi.fn()}
initialClaudeStatus={{
supported: true,
used: true,
installed: true,
connected: true,
running: true,
startFailed: false,
portConflict: false,
modelSource: "endpoint",
models: [
{
name: "glm-5.2:cloud",
displayName: "GLM 5.2",
description: "Long-horizon coding",
selected: true,
},
{
name: "qwen3.8:27b",
displayName: "Qwen 3.8 27B",
description: "Local coding",
selected: false,
},
],
}}
initialIntegrations={[
{
id: "claude-desktop",
name: "Claude",
description: "Use Ollama models in Claude Desktop",
installed: true,
action: "connect",
},
]}
/>,
);
expect(html).not.toContain("Models in Claude");
expect(html).not.toContain("GLM 5.2");
expect(html).not.toContain("Qwen 3.8 27B");
expect(html).not.toContain('type="checkbox"');
expect(html).not.toContain("Restart Claude");
expect(html).not.toContain("Built-in defaults");
});
it("keeps Claude available without a separate not-installed group", () => {
const html = renderToStaticMarkup(
<ConnectAppsScreen
completionError={null}
onRetryCompletion={vi.fn()}
initialIntegrations={[
{
id: "claude-desktop",
name: "Claude",
description: "Use Ollama models in Claude Desktop",
installed: false,
action: "connect",
},
]}
/>,
);
expect(html).toContain("Use Ollama models in Claude Desktop");
expect(html).toContain('aria-label="Connect Claude"');
expect(html).toContain("Download &amp; connect");
expect(html).not.toContain("Inactive");
expect(html).not.toContain("Not installed");
expect(html).not.toContain('disabled=""');
});
it("uses branded icons for the remaining launcher integrations", () => {
const html = renderToStaticMarkup(
<ConnectAppsScreen
completionError={null}
onRetryCompletion={vi.fn()}
initialIntegrations={[
{
id: "cline",
name: "Cline",
description: "Autonomous coding agent",
action: "copy",
command: "ollama launch cline",
},
{
id: "omp",
name: "Oh My Pi",
description: "AI coding agent",
action: "copy",
command: "ollama launch omp",
},
{
id: "pool",
name: "Poolside",
description: "Poolside's coding agent",
action: "copy",
command: "ollama launch pool",
},
{
id: "qwen",
name: "Qwen Code",
description: "Qwen's coding agent",
action: "copy",
command: "ollama launch qwen",
},
]}
/>,
);
expect(html).toContain("/launch-icons/cline.svg");
expect(html).toContain("/launch-icons/oh-my-pi.svg");
expect(html).toContain("/launch-icons/poolside.svg");
expect(html).toContain("/launch-icons/qwen-code.svg");
});
it("offers cloud sign-up, local setup, and sign in on the welcome screen", () => {
const html = renderToStaticMarkup(
<WelcomeScreen
isAuthenticated={false}
isSigningIn={false}
signInError={null}
onSignIn={vi.fn()}
onSignUp={vi.fn()}
onLocal={vi.fn()}
/>,
);
expect(html).toContain("Create an account");
expect(html).toContain(
"Create your account for access to faster, larger open models.",
);
expect(html).toContain("Your data is never logged or trained on.");
expect(html).toContain("Sign up");
expect(html).toContain("No thanks, I&#x27;ll use Ollama locally");
expect(html).toContain("Sign in");
expect(html).not.toContain("Skip");
});
it("shows the cloud choice without a sign-in link for authenticated users", () => {
const html = renderToStaticMarkup(
<WelcomeScreen
isAuthenticated
isSigningIn={false}
signInError={null}
onSignIn={vi.fn()}
onSignUp={vi.fn()}
onLocal={vi.fn()}
/>,
);
expect(html).toContain("Create an account");
expect(html).toContain(
"Create your account for access to faster, larger open models.",
);
expect(html).toContain("Your data is never logged or trained on.");
expect(html).not.toContain(">Sign in<");
});
it("shows only the local command on the final page", () => {
const html = renderToStaticMarkup(
<RunOllamaScreen completionError={null} onRetryCompletion={vi.fn()} />,
);
expect(html).toContain("Run Ollama");
expect(html).toContain(FIRST_MODEL_COMMAND);
expect(html).not.toContain("Finish");
expect(html).not.toContain("Sign in");
expect(html).not.toContain("create an account");
});
it("shows the connecting state on the welcome action", () => {
const html = renderToStaticMarkup(
<WelcomeScreen
isAuthenticated={false}
isSigningIn
signInError={null}
onSignIn={vi.fn()}
onSignUp={vi.fn()}
onLocal={vi.fn()}
/>,
);
expect(html).toContain("Finish in your browser…");
expect(html).not.toContain("Waiting for sign in…");
});
it("shows a retryable error when onboarding completion cannot be saved", () => {
const onRetryCompletion = vi.fn();
const html = renderToStaticMarkup(
<RunOllamaScreen
completionError="Unable to save setup. Please try again."
onRetryCompletion={onRetryCompletion}
/>,
);
expect(html).toContain("Unable to save setup. Please try again.");
expect(html).toContain('role="alert"');
expect(html).toContain("Try again");
});
});
File diff suppressed because it is too large. Load diff
+65 -41
View File
@@ -6,19 +6,19 @@ import { Field, Label, Description } from "@/components/ui/fieldset";
import { Badge } from "@/components/ui/badge";
import { Button } from "@/components/ui/button";
import { Slider } from "@/components/ui/slider";
import { ClaudeDesktopModelsSettings } from "@/components/ClaudeDesktopModelsSettings";
import {
WifiIcon,
FolderIcon,
BoltIcon,
WrenchIcon,
CloudIcon,
XMarkIcon,
CogIcon,
ArrowLeftIcon,
ArrowDownTrayIcon,
Squares2X2Icon,
} from "@heroicons/react/20/solid";
import { Settings as SettingsType } from "@/gotypes";
import { useNavigate } from "@tanstack/react-router";
import { isWindowsPlatform } from "@/lib/platform";
import { useUser } from "@/hooks/useUser";
import { useCloudStatus } from "@/hooks/useCloudStatus";
import { useQuery, useMutation, useQueryClient } from "@tanstack/react-query";
@@ -48,6 +48,8 @@ export default function Settings() {
const queryClient = useQueryClient();
const [showSaved, setShowSaved] = useState(false);
const [restartMessage, setRestartMessage] = useState(false);
const [showAppsInMenu, setShowAppsInMenuState] = useState(true);
const [showAppsInMenuPending, setShowAppsInMenuPending] = useState(false);
const {
user,
isAuthenticated,
@@ -61,7 +63,6 @@ export default function Settings() {
const [isAwaitingConnection, setIsAwaitingConnection] = useState(false);
const [connectionError, setConnectionError] = useState<string | null>(null);
const [pollingInterval, setPollingInterval] = useState<number | null>(null);
const navigate = useNavigate();
const {
cloudDisabled,
cloudStatus,
@@ -143,6 +144,15 @@ export default function Settings() {
refetchUser();
}, []); // eslint-disable-line react-hooks/exhaustive-deps
useEffect(() => {
window
.getShowAppsInMenu?.()
.then(setShowAppsInMenuState)
.catch((error) =>
console.error("Failed to load menu app visibility:", error),
);
}, []);
useEffect(() => {
const handleFocus = () => {
if (isAwaitingConnection && pollingInterval) {
@@ -220,6 +230,22 @@ export default function Settings() {
}
};
const handleShowAppsInMenu = async (checked: boolean) => {
const previous = showAppsInMenu;
setShowAppsInMenuState(checked);
setShowAppsInMenuPending(true);
try {
await window.setShowAppsInMenu?.(checked);
setShowSaved(true);
setTimeout(() => setShowSaved(false), 1500);
} catch (error) {
setShowAppsInMenuState(previous);
console.error("Failed to update menu app visibility:", error);
} finally {
setShowAppsInMenuPending(false);
}
};
const cloudOverriddenByEnv =
cloudStatus?.source === "env" || cloudStatus?.source === "both";
const cloudToggleDisabled =
@@ -266,49 +292,18 @@ export default function Settings() {
if (error || !settings) {
return (
<div className="flex min-h-screen items-center justify-center">
<div className="flex flex-1 items-center justify-center">
<div className="text-red-500">Failed to load settings</div>
</div>
);
}
const isWindows = navigator.platform.toLowerCase().includes("win");
const handleCloseSettings = () => {
const chatId = settings.LastHomeView === "chat" ? "new" : "launch";
navigate({ to: "/c/$chatId", params: { chatId } });
};
const isWindows = isWindowsPlatform();
return (
<main className="flex h-screen w-full flex-col select-none dark:bg-neutral-900">
<header
className="w-full flex flex-none justify-between h-[52px] py-2.5 items-center border-b border-neutral-200 dark:border-neutral-800 select-none"
onMouseDown={() => window.drag && window.drag()}
onDoubleClick={() => window.doubleClick && window.doubleClick()}
>
<h1
className={`${isWindows ? "pl-4" : "pl-24"} flex items-center font-rounded text-md font-medium dark:text-white`}
>
{isWindows && (
<button
onClick={handleCloseSettings}
className="hover:bg-neutral-100 mr-3 dark:hover:bg-neutral-800 rounded-full p-1.5"
>
<ArrowLeftIcon className="w-5 h-5 dark:text-white" />
</button>
)}
Settings
</h1>
{!isWindows && (
<button
onClick={handleCloseSettings}
className="p-1 hover:bg-neutral-100 mr-3 dark:hover:bg-neutral-800 rounded-full"
>
<XMarkIcon className="w-6 h-6 dark:text-white" />
</button>
)}
</header>
<main className="flex min-h-0 w-full flex-1 flex-col select-none dark:bg-neutral-900">
<div className="w-full p-6 overflow-y-auto flex-1 overscroll-contain">
<div className="space-y-4 max-w-2xl mx-auto">
<div className="mx-auto max-w-4xl space-y-4">
{/* Connect Ollama Account */}
<div className="overflow-hidden rounded-xl bg-white dark:bg-neutral-800">
<div className="p-4">
@@ -446,6 +441,29 @@ export default function Settings() {
</div>
</Field>
{!isWindows && (
<Field>
<div className="flex items-start justify-between gap-4">
<div className="flex flex-1 items-start space-x-3">
<Squares2X2Icon className="mt-1 h-5 w-5 flex-shrink-0 text-black dark:text-neutral-100" />
<div>
<Label>Show apps in menu</Label>
<Description>
Show connected apps at the top of the Ollama menu.
</Description>
</div>
</div>
<div className="flex-shrink-0">
<Switch
checked={showAppsInMenu}
disabled={showAppsInMenuPending}
onChange={handleShowAppsInMenu}
/>
</div>
</div>
</Field>
)}
{/* Auto Update */}
<Field>
<div className="flex items-start justify-between gap-4">
@@ -463,7 +481,9 @@ export default function Settings() {
<div className="flex-shrink-0">
<Switch
checked={settings.AutoUpdateEnabled}
onChange={(checked) => handleChange("AutoUpdateEnabled", checked)}
onChange={(checked) =>
handleChange("AutoUpdateEnabled", checked)
}
/>
</div>
</div>
@@ -544,7 +564,9 @@ export default function Settings() {
</Description>
<div className="mt-3">
<Slider
value={settings.ContextLength || defaultContextLength || 0}
value={
settings.ContextLength || defaultContextLength || 0
}
onChange={(value) => {
handleChange("ContextLength", value);
}}
@@ -566,6 +588,8 @@ export default function Settings() {
</div>
</div>
<ClaudeDesktopModelsSettings />
{/* Agent Mode */}
{window.OLLAMA_TOOLS && (
<div className="overflow-hidden rounded-xl bg-white dark:bg-neutral-800">
@@ -0,0 +1,23 @@
import { renderToStaticMarkup } from "react-dom/server";
import { afterEach, describe, expect, it, vi } from "vitest";
import { SidebarLayout } from "./layout";
describe("SidebarLayout", () => {
afterEach(() => {
vi.unstubAllGlobals();
});
it("keeps the macOS title offset in step with the sidebar transition", () => {
vi.stubGlobal("window", { OLLAMA_PLATFORM: "darwin" });
const html = renderToStaticMarkup(
<SidebarLayout title="Connect your apps" sidebar={<nav />}>
<div />
</SidebarLayout>,
);
expect(html).toContain("pl-36");
expect(html).toContain("transition-[padding-left]");
expect(html).toContain("duration-300");
});
});
+47 -36
View File
@@ -1,30 +1,39 @@
import { Link } from "@tanstack/react-router";
import { useSettings } from "@/hooks/useSettings";
import { ChatIcon } from "@/components/ChatIcon";
import { isWindowsPlatform } from "@/lib/platform";
import { useState } from "react";
let sessionSidebarOpen = false;
export function SidebarLayout({
sidebar,
title,
children,
}: React.PropsWithChildren<{
sidebar: React.ReactNode;
collapsible?: boolean;
chatId?: string;
title?: string;
}>) {
const { settings, setSettings } = useSettings();
const isWindows = navigator.platform.toLowerCase().includes("win");
const [sidebarOpen, setSidebarOpen] = useState(sessionSidebarOpen);
const isWindows = isWindowsPlatform();
const toggleSidebar = () => {
sessionSidebarOpen = !sidebarOpen;
setSidebarOpen(sessionSidebarOpen);
};
return (
<div className={`flex transition-[width] duration-300 dark:bg-neutral-900`}>
<div className="flex h-screen w-full overflow-hidden dark:bg-neutral-900">
<div
className={`absolute flex mx-2 py-2 z-20 items-center transition-[left] duration-375 text-neutral-500 dark:text-neutral-400 ${settings.sidebarOpen ? (isWindows ? "left-2" : "left-[204px]") : isWindows ? "left-2" : "left-20"}`}
className={`absolute flex mx-2 py-2 z-20 items-center transition-[left] duration-375 text-neutral-500 dark:text-neutral-400 ${sidebarOpen ? (isWindows ? "left-2" : "left-[140px]") : isWindows ? "left-2" : "left-20"}`}
>
<button
onClick={() => setSettings({ SidebarOpen: !settings.sidebarOpen })}
onClick={toggleSidebar}
onMouseDown={(e) => {
e.stopPropagation();
}}
className="h-9 w-9 flex items-center justify-center rounded-full hover:bg-neutral-100 dark:hover:bg-neutral-700/75 cursor-pointer"
aria-label={settings.sidebarOpen ? "Hide sidebar" : "Show sidebar"}
title={settings.sidebarOpen ? "Hide sidebar" : "Show sidebar"}
aria-label={sidebarOpen ? "Hide sidebar" : "Show sidebar"}
title={sidebarOpen ? "Hide sidebar" : "Show sidebar"}
>
<svg
className="h-5 w-5 fill-current"
@@ -35,45 +44,47 @@ export function SidebarLayout({
<path d="M7.76132 16.6344H9.58103V1.59842H7.76132V16.6344ZM4.20898 18.2316H19.124C21.6518 18.2316 23.1293 16.6963 23.1293 14.0209V4.2205C23.1293 1.54512 21.6518 0.00351715 19.124 0.00351715H4.20898C1.54336 0.00351715 0 1.54512 0 4.2205V14.0209C0 16.6963 1.54336 18.2316 4.20898 18.2316ZM4.31191 16.3184C2.79628 16.3184 1.91327 15.4434 1.91327 13.926V4.31542C1.91327 2.79979 2.79628 1.91678 4.31191 1.91678H18.8174C20.333 1.91678 21.216 2.79979 21.216 4.31542V13.926C21.216 15.4434 20.333 16.3184 18.8174 16.3184H4.31191ZM5.85116 5.50038C6.1951 5.50038 6.49217 5.20507 6.49217 4.87968C6.49217 4.54628 6.1951 4.25722 5.85116 4.25722H3.8412C3.49725 4.25722 3.20819 4.54628 3.20819 4.87968C3.20819 5.20507 3.49725 5.50038 3.8412 5.50038H5.85116ZM5.85116 8.1158C6.1951 8.1158 6.49217 7.82049 6.49217 7.4871C6.49217 7.1537 6.1951 6.8744 5.85116 6.8744H3.8412C3.49725 6.8744 3.20819 7.1537 3.20819 7.4871C3.20819 7.82049 3.49725 8.1158 3.8412 8.1158H5.85116ZM5.85116 10.725C6.1951 10.725 6.49217 10.4439 6.49217 10.1105C6.49217 9.77713 6.1951 9.48983 5.85116 9.48983H3.8412C3.49725 9.48983 3.20819 9.77713 3.20819 10.1105C3.20819 10.4439 3.49725 10.725 3.8412 10.725H5.85116Z" />
</svg>
</button>
<Link
to="/c/$chatId"
params={{ chatId: "new" }}
title="New chat"
className={`flex ml-1 items-center justify-center rounded-full transition-opacity duration-375 h-9 w-9 hover:bg-neutral-100 dark:hover:bg-neutral-700 ${
settings.sidebarOpen
? "opacity-0 pointer-events-none"
: "opacity-100"
}`}
>
<svg
className="h-5 w-5 fill-current"
viewBox="0 0 24 24"
fill="none"
xmlns="http://www.w3.org/2000/svg"
{!title && (
<Link
to="/c/$chatId"
params={{ chatId: "new" }}
title="New chat"
className={`flex ml-1 items-center justify-center rounded-full transition-opacity duration-375 h-9 w-9 hover:bg-neutral-100 dark:hover:bg-neutral-700 ${
sidebarOpen ? "opacity-0 pointer-events-none" : "opacity-100"
}`}
>
<path d="M17.0859 3.39949L15.2135 5.27196H7.27028C5.78649 5.27196 4.94684 6.11336 4.94684 7.59716V16.664C4.94684 18.1558 5.78649 18.9892 7.27028 18.9892H16.3406C17.8324 18.9892 18.6623 18.1558 18.6623 16.664V8.79514L20.5428 6.9115C20.567 7.11532 20.5773 7.33066 20.5773 7.55419V16.7149C20.5773 19.4069 19.0818 20.9024 16.3898 20.9024H7.22107C4.53708 20.9024 3.03357 19.4069 3.03357 16.7149V7.55419C3.03357 4.8622 4.53708 3.35869 7.22107 3.35869H16.3898C16.6329 3.35869 16.8662 3.37094 17.0859 3.39949Z" />
<path d="M9.92714 14.381L11.914 13.5403L20.8312 4.63114L19.3404 3.1581L10.433 12.0655L9.55234 13.9964C9.45664 14.2169 9.70293 14.4714 9.92714 14.381ZM21.5767 3.89364L22.2588 3.19384C22.6347 2.80184 22.6435 2.2663 22.2711 1.90536L22.0148 1.64287C21.6822 1.31377 21.1334 1.36513 20.7689 1.72158L20.0859 2.39833L21.5767 3.89364Z" />
</svg>
</Link>
<ChatIcon />
</Link>
)}
</div>
<div
className={`flex flex-col transition-[width] duration-300 max-h-screen ${settings.sidebarOpen ? "w-64" : "w-0"}`}
className={`flex max-h-screen flex-col transition-[width] duration-300 ${
sidebarOpen
? "w-48 border-r border-neutral-200 bg-neutral-50 dark:border-neutral-800 dark:bg-neutral-950/40"
: "w-0"
}`}
>
<div
onDoubleClick={() => window.doubleClick && window.doubleClick()}
onMouseDown={() => window.drag && window.drag()}
className="flex-none h-13 w-full"
></div>
{settings.sidebarOpen && sidebar}
{sidebarOpen && sidebar}
</div>
<main
className={`flex flex-1 flex-col min-w-0 transition-all duration-300`}
>
<main className="flex min-w-0 flex-1 flex-col transition-all duration-300">
<div
className={`h-13 flex-none w-full z-10 flex items-center bg-white dark:bg-neutral-900 ${isWindows ? "xl:hidden" : "xl:fixed xl:bg-transparent xl:dark:bg-transparent"}`}
className={`h-13 z-10 flex w-full flex-none items-center bg-white dark:bg-neutral-900 ${title ? "" : isWindows ? "xl:hidden" : "xl:fixed xl:bg-transparent xl:dark:bg-transparent"}`}
onDoubleClick={() => window.doubleClick && window.doubleClick()}
onMouseDown={() => window.drag && window.drag()}
></div>
>
{title && (
<h1
className={`${sidebarOpen ? "pl-6" : isWindows ? "pl-16" : "pl-36"} transition-[padding-left] duration-300 font-rounded text-md font-medium dark:text-white`}
>
{title}
</h1>
)}
</div>
{children}
</main>
</div>
+4 -1
View File
@@ -10,6 +10,7 @@ interface SettingsState {
selectedModel: string;
sidebarOpen: boolean;
lastHomeView: string;
onboardingVersion: number;
thinkEnabled: boolean;
thinkLevel: string;
}
@@ -23,6 +24,7 @@ type SettingsUpdate = Partial<{
SelectedModel: string;
SidebarOpen: boolean;
LastHomeView: string;
OnboardingVersion: number;
}>;
export function useSettings() {
@@ -52,7 +54,8 @@ export function useSettings() {
thinkLevel: settingsData?.settings?.ThinkLevel ?? "none",
selectedModel: settingsData?.settings?.SelectedModel ?? "",
sidebarOpen: settingsData?.settings?.SidebarOpen ?? false,
lastHomeView: settingsData?.settings?.LastHomeView ?? "launch",
lastHomeView: settingsData?.settings?.LastHomeView ?? "chat",
onboardingVersion: settingsData?.settings?.OnboardingVersion ?? 0,
}),
[settingsData?.settings],
);
+51 -5
View File
@@ -2,18 +2,35 @@
@plugin "@tailwindcss/typography";
@import "katex/dist/katex.min.css";
/* Retain component class names while making the dark variant unreachable. */
@custom-variant dark (@media not all);
@theme {
--font-sans: ui-sans-serif, system-ui, "Segoe UI", sans-serif;
--font-rounded:
"SF Pro Rounded", ui-sans-serif, system-ui, "Segoe UI", sans-serif;
}
@media (prefers-color-scheme: dark) {
/* Dark mode styles go here */
@layer base {
:root {
/* Example dark mode variables */
--bg-color: #1a1a1a;
--text-color: #ffffff;
color-scheme: light;
}
html,
body,
#root {
background-color: #fff;
}
a[href],
button:not(:disabled),
[role="button"]:not([aria-disabled="true"]) {
cursor: pointer;
}
button:disabled,
[role="button"][aria-disabled="true"] {
cursor: not-allowed;
}
}
@@ -28,3 +45,32 @@
opacity: 1;
}
}
@keyframes claude-connected-backdrop-in {
from {
opacity: 0;
}
}
@keyframes claude-connected-dialog-in {
from {
opacity: 0;
transform: translateY(6px) scale(0.98);
}
}
.claude-connected-backdrop {
animation: claude-connected-backdrop-in 280ms ease-in-out both;
}
.claude-connected-dialog {
animation: claude-connected-dialog-in 345ms cubic-bezier(0.4, 0, 0.2, 1) both;
will-change: opacity, transform;
}
@media (prefers-reduced-motion: reduce) {
.claude-connected-backdrop,
.claude-connected-dialog {
animation: none;
}
}
+188
View File
@@ -0,0 +1,188 @@
import { describe, expect, it } from "vitest";
import {
addClaudeModelSelection,
claudeDesktopRecoveryMessage,
claudeDesktopMaxModels,
claudeDesktopMaxModelsMessage,
claudeDesktopUsableSelection,
defaultClaudeDesktopMaxModels,
isClaudeConfigured,
} from "./claudeDesktop";
describe("isClaudeConfigured", () => {
it("keeps a failed configured profile switchable off", () => {
expect(
isClaudeConfigured({
supported: true,
used: true,
installed: true,
configured: true,
connected: false,
running: false,
startFailed: true,
portConflict: true,
}),
).toBe(true);
});
});
describe("claudeDesktopMaxModels", () => {
it("falls back to the five literal Claude slots without a status", () => {
expect(claudeDesktopMaxModels(undefined)).toBe(5);
expect(
claudeDesktopMaxModels({
supported: true,
used: true,
installed: true,
connected: false,
running: false,
startFailed: false,
portConflict: false,
}),
).toBe(defaultClaudeDesktopMaxModels);
});
it("uses the server-provided limit when present", () => {
expect(
claudeDesktopMaxModels({
supported: true,
used: true,
installed: true,
connected: false,
running: false,
startFailed: false,
portConflict: false,
maxModels: 3,
}),
).toBe(3);
});
});
describe("claudeDesktopRecoveryMessage", () => {
it("prefers current native guidance over stale action errors", () => {
expect(
claudeDesktopRecoveryMessage(
"Cloud models are off. Select an installed model in Settings.",
"Ollama could not open Claude.",
),
).toBe("Cloud models are off. Select an installed model in Settings.");
});
it("clears after native recovery when no action error remains", () => {
expect(claudeDesktopRecoveryMessage(undefined, null)).toBeNull();
});
});
describe("addClaudeModelSelection", () => {
it("appends models below the limit", () => {
expect(addClaudeModelSelection(["kimi-k3:cloud"], "qwen3:8b", 5)).toEqual({
selection: ["kimi-k3:cloud", "qwen3:8b"],
});
});
it("rejects a sixth selection with a clear message", () => {
const selection = [
"glm-5.2:cloud",
"kimi-k3:cloud",
"deepseek-v4-pro",
"deepseek-v4-flash",
"gemma4:26b:cloud",
];
const result = addClaudeModelSelection(selection, "qwen3:8b", 5);
expect(result.selection).toBe(selection);
expect(result.error).toBe(
"Claude supports up to 5 models. Deselect one to add another.",
);
});
it("honors a smaller server-provided limit", () => {
const result = addClaudeModelSelection(["qwen3:8b"], "llama3.2", 1);
expect(result.selection).toEqual(["qwen3:8b"]);
expect(result.error).toBe(claudeDesktopMaxModelsMessage(1));
});
it("is a no-op for an already selected model", () => {
expect(addClaudeModelSelection(["qwen3:8b"], "qwen3:8b", 5)).toEqual({
selection: ["qwen3:8b"],
});
});
});
describe("claudeDesktopUsableSelection", () => {
it("replaces unavailable paid selections with an available free model", () => {
expect(
claudeDesktopUsableSelection([
{
name: "glm-5.2:cloud",
displayName: "glm-5.2:cloud",
selected: true,
availability: "unavailable",
reason: "upgrade_required",
requiredPlan: "pro",
},
{
name: "gemma4:31b-cloud",
displayName: "gemma4:31b-cloud",
selected: false,
availability: "available",
requiredPlan: "free",
},
]),
).toEqual(["gemma4:31b-cloud"]);
});
it("preserves selected models that remain available", () => {
expect(
claudeDesktopUsableSelection([
{
name: "gemma4:31b-cloud",
displayName: "gemma4:31b-cloud",
selected: true,
availability: "available",
},
{
name: "qwen3:8b",
displayName: "qwen3:8b",
selected: false,
availability: "available",
},
]),
).toEqual(["gemma4:31b-cloud"]);
});
it("selects all available recommendations for a default catalog", () => {
expect(
claudeDesktopUsableSelection(
[
{
name: "glm-5.2:cloud",
displayName: "glm-5.2:cloud",
selected: false,
availability: "available",
},
{
name: "kimi-k3:cloud",
displayName: "kimi-k3:cloud",
selected: false,
availability: "available",
},
],
true,
5,
),
).toEqual(["glm-5.2:cloud", "kimi-k3:cloud"]);
});
it("returns no selection when no model is available", () => {
expect(
claudeDesktopUsableSelection([
{
name: "glm-5.2:cloud",
displayName: "glm-5.2:cloud",
selected: true,
availability: "unavailable",
},
]),
).toEqual([]);
});
});
+82
View File
@@ -0,0 +1,82 @@
import type {
ClaudeDesktopModelStatus,
ClaudeDesktopStatus,
} from "@/types/webview";
export const CLAUDE_INSTALL_TIMEOUT_MS = 120_000;
export function isClaudeConnectionComplete(
enabled: boolean,
status: ClaudeDesktopStatus,
) {
const configured = isClaudeConfigured(status);
return enabled
? configured && status.connected && !status.startFailed
: !configured;
}
export function isClaudeConfigured(status: ClaudeDesktopStatus): boolean {
return status.configured ?? status.connected;
}
export function claudeDesktopRecoveryMessage(
statusError?: string,
actionError?: string | null,
): string | null {
return statusError || actionError || null;
}
export function scheduleClaudeInstallTimeout(onTimeout: () => void) {
return window.setTimeout(onTimeout, CLAUDE_INSTALL_TIMEOUT_MS);
}
// Claude Desktop has a bounded model list. The app supplies the limit when it
// knows it; retain the existing five-model behavior for older app versions.
export const defaultClaudeDesktopMaxModels = 5;
export function claudeDesktopMaxModels(
status?: ClaudeDesktopStatus | null,
): number {
return status?.maxModels && status.maxModels > 0
? status.maxModels
: defaultClaudeDesktopMaxModels;
}
export function claudeDesktopMaxModelsMessage(maxModels: number): string {
return `Claude supports up to ${maxModels} models. Deselect one to add another.`;
}
// Unavailable models remain visible for account guidance, but cannot remain
// selected. Choose one available model when filtering would empty the list.
export function claudeDesktopUsableSelection(
models: ClaudeDesktopModelStatus[],
selectAllAvailable = false,
maxModels = defaultClaudeDesktopMaxModels,
): string[] {
const available = models.filter(
(model) =>
model.availability === undefined || model.availability === "available",
);
if (selectAllAvailable) {
return available.slice(0, maxModels).map((model) => model.name);
}
const selected = available
.filter((model) => model.selected)
.map((model) => model.name);
if (selected.length > 0) return selected;
return available.length > 0 ? [available[0].name] : [];
}
// addClaudeModelSelection returns the selection with name appended, or an
// unchanged selection plus an error when the Claude model limit is reached.
export function addClaudeModelSelection(
selection: string[],
name: string,
maxModels: number,
): { selection: string[]; error?: string } {
if (selection.includes(name)) return { selection };
if (selection.length >= maxModels) {
return { selection, error: claudeDesktopMaxModelsMessage(maxModels) };
}
return { selection: [...selection, name] };
}
+25
View File
@@ -0,0 +1,25 @@
export interface IntegrationIcon {
src: string;
className?: string;
}
export const INTEGRATION_ICONS: Record<string, IntegrationIcon> = {
"claude-desktop": { src: "/launch-icons/claude.svg" },
claude: { src: "/launch-icons/claude-code.svg" },
hermes: { src: "/launch-icons/hermes-agent.svg" },
"hermes-desktop": { src: "/launch-icons/hermes-agent.svg" },
openclaw: { src: "/launch-icons/openclaw.svg" },
opencode: {
src: "/launch-icons/opencode.svg",
className: "h-7 w-7 rounded",
},
codex: { src: "/launch-icons/codex.svg" },
copilot: { src: "/launch-icons/copilot.svg" },
droid: { src: "/launch-icons/droid.svg" },
dsh: { src: "/launch-icons/deepseek-harness.svg" },
pi: { src: "/launch-icons/pi.svg" },
cline: { src: "/launch-icons/cline.svg" },
omp: { src: "/launch-icons/oh-my-pi.svg" },
pool: { src: "/launch-icons/poolside.svg" },
qwen: { src: "/launch-icons/qwen-code.svg" },
};
+51
View File
@@ -0,0 +1,51 @@
// Keep in sync with store.CurrentOnboardingVersion in app/store/store.go.
export const CURRENT_ONBOARDING_VERSION = 1;
export type OnboardingAuthMode = "signin" | "signup";
export function onboardingConnectUrl(
connectUrl: string,
mode: OnboardingAuthMode,
): string {
const url = new URL(connectUrl);
url.searchParams.delete("launch");
if (mode === "signup") {
url.searchParams.set("signup", "true");
} else {
url.searchParams.delete("signup");
}
return url.toString();
}
export const AUTHENTICATION_TIMEOUT_MS = 5 * 60 * 1000;
export type OnboardingStep = "intro" | "welcome" | "apps" | "run";
export type OnboardingAction = "continue" | "authenticated" | "local";
export type AuthenticationTimeoutAction = "ignore" | "defer" | "fail";
export function nextOnboardingStep(
step: OnboardingStep,
action: OnboardingAction,
isAuthenticated: boolean,
): OnboardingStep {
if (action === "local") return "run";
if (step === "intro" && action === "continue") {
return isAuthenticated ? "apps" : "welcome";
}
if (step === "welcome" && action === "authenticated") return "apps";
return step;
}
export function authenticationTimeoutAction(
settled: boolean,
checking: boolean,
): AuthenticationTimeoutAction {
if (settled) return "ignore";
if (checking) return "defer";
return "fail";
}
export function homeChatId(): "new" {
return "new";
}
+10
View File
@@ -0,0 +1,10 @@
export function isWindowsPlatform(): boolean {
if (typeof window !== "undefined" && window.OLLAMA_PLATFORM) {
return window.OLLAMA_PLATFORM === "windows";
}
return (
typeof navigator !== "undefined" &&
navigator.platform.toLowerCase().includes("win")
);
}
+49 -3
View File
@@ -12,6 +12,8 @@
import { Route as rootRoute } from './routes/__root'
import { Route as SettingsImport } from './routes/settings'
import { Route as OnboardingImport } from './routes/onboarding'
import { Route as ConnectImport } from './routes/connect'
import { Route as IndexImport } from './routes/index'
import { Route as CChatIdImport } from './routes/c.$chatId'
@@ -23,6 +25,18 @@ const SettingsRoute = SettingsImport.update({
getParentRoute: () => rootRoute,
} as any)
const OnboardingRoute = OnboardingImport.update({
id: '/onboarding',
path: '/onboarding',
getParentRoute: () => rootRoute,
} as any)
const ConnectRoute = ConnectImport.update({
id: '/connect',
path: '/connect',
getParentRoute: () => rootRoute,
} as any)
const IndexRoute = IndexImport.update({
id: '/',
path: '/',
@@ -46,6 +60,20 @@ declare module '@tanstack/react-router' {
preLoaderRoute: typeof IndexImport
parentRoute: typeof rootRoute
}
'/connect': {
id: '/connect'
path: '/connect'
fullPath: '/connect'
preLoaderRoute: typeof ConnectImport
parentRoute: typeof rootRoute
}
'/onboarding': {
id: '/onboarding'
path: '/onboarding'
fullPath: '/onboarding'
preLoaderRoute: typeof OnboardingImport
parentRoute: typeof rootRoute
}
'/settings': {
id: '/settings'
path: '/settings'
@@ -67,12 +95,16 @@ declare module '@tanstack/react-router' {
export interface FileRoutesByFullPath {
'/': typeof IndexRoute
'/connect': typeof ConnectRoute
'/onboarding': typeof OnboardingRoute
'/settings': typeof SettingsRoute
'/c/$chatId': typeof CChatIdRoute
}
export interface FileRoutesByTo {
'/': typeof IndexRoute
'/connect': typeof ConnectRoute
'/onboarding': typeof OnboardingRoute
'/settings': typeof SettingsRoute
'/c/$chatId': typeof CChatIdRoute
}
@@ -80,27 +112,33 @@ export interface FileRoutesByTo {
export interface FileRoutesById {
__root__: typeof rootRoute
'/': typeof IndexRoute
'/connect': typeof ConnectRoute
'/onboarding': typeof OnboardingRoute
'/settings': typeof SettingsRoute
'/c/$chatId': typeof CChatIdRoute
}
export interface FileRouteTypes {
fileRoutesByFullPath: FileRoutesByFullPath
fullPaths: '/' | '/settings' | '/c/$chatId'
fullPaths: '/' | '/connect' | '/onboarding' | '/settings' | '/c/$chatId'
fileRoutesByTo: FileRoutesByTo
to: '/' | '/settings' | '/c/$chatId'
id: '__root__' | '/' | '/settings' | '/c/$chatId'
to: '/' | '/connect' | '/onboarding' | '/settings' | '/c/$chatId'
id: '__root__' | '/' | '/connect' | '/onboarding' | '/settings' | '/c/$chatId'
fileRoutesById: FileRoutesById
}
export interface RootRouteChildren {
IndexRoute: typeof IndexRoute
ConnectRoute: typeof ConnectRoute
OnboardingRoute: typeof OnboardingRoute
SettingsRoute: typeof SettingsRoute
CChatIdRoute: typeof CChatIdRoute
}
const rootRouteChildren: RootRouteChildren = {
IndexRoute: IndexRoute,
ConnectRoute: ConnectRoute,
OnboardingRoute: OnboardingRoute,
SettingsRoute: SettingsRoute,
CChatIdRoute: CChatIdRoute,
}
@@ -116,6 +154,8 @@ export const routeTree = rootRoute
"filePath": "__root.tsx",
"children": [
"/",
"/connect",
"/onboarding",
"/settings",
"/c/$chatId"
]
@@ -123,6 +163,12 @@ export const routeTree = rootRoute
"/": {
"filePath": "index.tsx"
},
"/connect": {
"filePath": "connect.tsx"
},
"/onboarding": {
"filePath": "onboarding.tsx"
},
"/settings": {
"filePath": "settings.tsx"
},
+13 -78
View File
@@ -1,40 +1,25 @@
import { createFileRoute } from "@tanstack/react-router";
import { createFileRoute, redirect } from "@tanstack/react-router";
import { useChat } from "@/hooks/useChats";
import Chat from "@/components/Chat";
import { getChat } from "@/api";
import { SidebarLayout } from "@/components/layout/layout";
import { ChatSidebar } from "@/components/ChatSidebar";
import LaunchCommands from "@/components/LaunchCommands";
import { useEffect, useRef } from "react";
import { useEffect } from "react";
import { useSettings } from "@/hooks/useSettings";
const launchSidebarRequestedKey = "ollama.launchSidebarRequested";
const launchSidebarSeenKey = "ollama.launchSidebarSeen";
const fallbackSessionState = new Map<string, string>();
function getSessionState() {
if (typeof sessionStorage !== "undefined") {
return sessionStorage;
}
return {
getItem(key: string) {
return fallbackSessionState.get(key) ?? null;
},
setItem(key: string, value: string) {
fallbackSessionState.set(key, value);
},
removeItem(key: string) {
fallbackSessionState.delete(key);
},
};
}
export const Route = createFileRoute("/c/$chatId")({
component: RouteComponent,
beforeLoad: ({ params }) => {
if (params.chatId === "launch") {
throw redirect({
to: "/c/$chatId",
params: { chatId: "new" },
mask: { to: "/" },
});
}
},
loader: async ({ context, params }) => {
// Skip loading for special non-chat views
if (params.chatId !== "new" && params.chatId !== "launch") {
if (params.chatId !== "new") {
context.queryClient.ensureQueryData({
queryKey: ["chat", params.chatId],
queryFn: () => getChat(params.chatId),
@@ -47,61 +32,19 @@ export const Route = createFileRoute("/c/$chatId")({
function RouteComponent() {
const { chatId } = Route.useParams();
const { settingsData, setSettings } = useSettings();
const previousChatIdRef = useRef<string | null>(null);
// Always call hooks at the top level - use a flag to skip data when chatId is a special view
const {
data: chatData,
isLoading: chatLoading,
error: chatError,
} = useChat(chatId === "new" || chatId === "launch" ? "" : chatId);
} = useChat(chatId === "new" ? "" : chatId);
useEffect(() => {
if (!settingsData) {
return;
}
const previousChatId = previousChatIdRef.current;
previousChatIdRef.current = chatId;
if (chatId === "launch") {
const sessionState = getSessionState();
const shouldOpenSidebar =
previousChatId !== "launch" &&
(() => {
if (sessionState.getItem(launchSidebarRequestedKey) === "1") {
sessionState.removeItem(launchSidebarRequestedKey);
sessionState.setItem(launchSidebarSeenKey, "1");
return true;
}
if (sessionState.getItem(launchSidebarSeenKey) !== "1") {
sessionState.setItem(launchSidebarSeenKey, "1");
return true;
}
return false;
})();
const updates: { LastHomeView?: string; SidebarOpen?: boolean } = {};
if (settingsData.LastHomeView !== "launch") {
updates.LastHomeView = "launch";
}
if (shouldOpenSidebar && !settingsData.SidebarOpen) {
updates.SidebarOpen = true;
}
if (Object.keys(updates).length === 0) {
return;
}
setSettings(updates).catch(() => {
// Best effort persistence for home view preference.
});
return;
}
if (settingsData.LastHomeView === "chat") {
return;
}
@@ -120,14 +63,6 @@ function RouteComponent() {
);
}
if (chatId === "launch") {
return (
<SidebarLayout sidebar={<ChatSidebar currentChatId={chatId} />}>
<LaunchCommands />
</SidebarLayout>
);
}
// Handle existing chat case
if (chatLoading) {
return (
+19
View File
@@ -0,0 +1,19 @@
import { AppSidebar } from "@/components/AppSidebar";
import { ConnectAppsScreen } from "@/components/Onboarding";
import { SidebarLayout } from "@/components/layout/layout";
import { createFileRoute } from "@tanstack/react-router";
export const Route = createFileRoute("/connect")({
component: ConnectRoute,
});
function ConnectRoute() {
return (
<SidebarLayout
title="Connect your apps"
sidebar={<AppSidebar current="apps" />}
>
<ConnectAppsScreen />
</SidebarLayout>
);
}
+6 -2
View File
@@ -1,5 +1,6 @@
import { createFileRoute, redirect } from "@tanstack/react-router";
import { getSettings } from "@/api";
import { CURRENT_ONBOARDING_VERSION, homeChatId } from "@/lib/onboarding";
export const Route = createFileRoute("/")({
beforeLoad: async ({ context }) => {
@@ -7,8 +8,11 @@ export const Route = createFileRoute("/")({
queryKey: ["settings"],
queryFn: getSettings,
});
const chatId =
settingsData?.settings?.LastHomeView === "chat" ? "new" : "launch";
if (settingsData.settings.OnboardingVersion < CURRENT_ONBOARDING_VERSION) {
throw redirect({ to: "/onboarding" });
}
const chatId = homeChatId();
throw redirect({
to: "/c/$chatId",
+198
View File
@@ -0,0 +1,198 @@
import Onboarding from "@/components/Onboarding";
import { getSettings } from "@/api";
import { useSettings } from "@/hooks/useSettings";
import { useUser } from "@/hooks/useUser";
import {
AUTHENTICATION_TIMEOUT_MS,
authenticationTimeoutAction,
CURRENT_ONBOARDING_VERSION,
homeChatId,
onboardingConnectUrl,
type OnboardingAuthMode,
} from "@/lib/onboarding";
import { createFileRoute, redirect, useNavigate } from "@tanstack/react-router";
import { useCallback, useEffect, useRef, useState } from "react";
export const Route = createFileRoute("/onboarding")({
beforeLoad: async ({ context }) => {
// Let developers review onboarding without resetting their local app data.
if (
import.meta.env.DEV &&
new URLSearchParams(window.location.search).get("preview") === "1"
) {
return;
}
const settingsData = await context.queryClient.ensureQueryData({
queryKey: ["settings"],
queryFn: getSettings,
});
if (settingsData.settings.OnboardingVersion >= CURRENT_ONBOARDING_VERSION) {
const chatId = homeChatId();
throw redirect({
to: "/c/$chatId",
params: { chatId },
mask: { to: "/" },
});
}
},
component: OnboardingRoute,
});
function OnboardingRoute() {
const navigate = useNavigate();
const { settingsData, setSettings } = useSettings();
const { fetchConnectUrl, refetchUser, isAuthenticated } = useUser();
const [isAwaitingAuth, setIsAwaitingAuth] = useState(false);
const [signInError, setSignInError] = useState<string | null>(null);
const [completionError, setCompletionError] = useState<string | null>(null);
const authAttemptRef = useRef(0);
const completeOnboarding = useCallback(async (): Promise<boolean> => {
setCompletionError(null);
try {
if (!settingsData) {
throw new Error("Settings are not loaded");
}
await setSettings({
OnboardingVersion: CURRENT_ONBOARDING_VERSION,
});
return true;
} catch (error) {
console.error("Failed to save onboarding state:", error);
setCompletionError("Unable to save setup. Please try again.");
return false;
}
}, [setSettings, settingsData]);
const finishSetup = useCallback(() => {
void completeOnboarding();
}, [completeOnboarding]);
const openApps = useCallback(async (): Promise<boolean> => {
if (!(await completeOnboarding())) return false;
await navigate({ to: "/connect" });
return true;
}, [completeOnboarding, navigate]);
const retryCompletion = useCallback(() => {
void completeOnboarding();
}, [completeOnboarding]);
const authenticate = useCallback(
async (mode: OnboardingAuthMode) => {
setSignInError(null);
if (isAuthenticated) {
return;
}
const authAttempt = ++authAttemptRef.current;
setIsAwaitingAuth(true);
try {
const result = await fetchConnectUrl();
if (authAttempt !== authAttemptRef.current) return;
if (!result.data) {
throw new Error("No sign-in URL was returned");
}
window.open(onboardingConnectUrl(result.data, mode), "_blank");
} catch (error) {
if (authAttempt !== authAttemptRef.current) return;
console.error("Failed to start sign in:", error);
setIsAwaitingAuth(false);
setSignInError("Unable to start sign in. Please try again.");
}
},
[fetchConnectUrl, isAuthenticated],
);
const signIn = useCallback(() => authenticate("signin"), [authenticate]);
const signUp = useCallback(() => authenticate("signup"), [authenticate]);
const useLocal = useCallback(() => {
authAttemptRef.current += 1;
setIsAwaitingAuth(false);
setSignInError(null);
finishSetup();
}, [finishSetup]);
useEffect(() => {
if (!isAwaitingAuth) return;
let checking = false;
let settled = false;
let timeoutPending = false;
const authAttempt = authAttemptRef.current;
const failConnection = () => {
if (settled || authAttempt !== authAttemptRef.current) return;
settled = true;
setIsAwaitingAuth(false);
setSignInError(
"Connection is taking longer than expected. Please try again.",
);
};
const checkConnection = async () => {
if (checking || settled || authAttempt !== authAttemptRef.current) return;
checking = true;
try {
const result = await refetchUser();
if (
!settled &&
authAttempt === authAttemptRef.current &&
result.data?.name
) {
settled = true;
setIsAwaitingAuth(false);
window.activateOllama?.();
}
} catch (error) {
console.error("Failed to check sign-in status:", error);
} finally {
checking = false;
if (timeoutPending) failConnection();
}
};
void checkConnection();
const pollingInterval = window.setInterval(checkConnection, 1000);
const timeout = window.setTimeout(() => {
const action = authenticationTimeoutAction(settled, checking);
if (action === "ignore") return;
if (action === "defer") {
timeoutPending = true;
return;
}
failConnection();
}, AUTHENTICATION_TIMEOUT_MS);
window.addEventListener("focus", checkConnection);
return () => {
settled = true;
window.clearInterval(pollingInterval);
window.clearTimeout(timeout);
window.removeEventListener("focus", checkConnection);
};
}, [isAwaitingAuth, refetchUser]);
return (
<Onboarding
completionError={completionError}
isAuthenticated={isAuthenticated}
isSigningIn={isAwaitingAuth}
signInError={signInError}
onOpenApps={openApps}
onSignIn={signIn}
onSignUp={signUp}
onRetryCompletion={retryCompletion}
onUseLocal={useLocal}
/>
);
}
+11 -1
View File
@@ -1,6 +1,16 @@
import { AppSidebar } from "@/components/AppSidebar";
import { SidebarLayout } from "@/components/layout/layout";
import { createFileRoute } from "@tanstack/react-router";
import Settings from "@/components/Settings";
export const Route = createFileRoute("/settings")({
component: Settings,
component: SettingsRoute,
});
function SettingsRoute() {
return (
<SidebarLayout title="Settings" sidebar={<AppSidebar current="settings" />}>
<Settings />
</SidebarLayout>
);
}
+64 -1
View File
@@ -12,6 +12,45 @@ interface MenuItem {
separator?: boolean;
}
interface ClaudeDesktopStatus {
supported: boolean;
used: boolean;
installed: boolean;
configured?: boolean;
connected: boolean;
running: boolean;
startFailed: boolean;
portConflict: boolean;
gatewayPort?: number;
error?: string;
modelSource?: "user" | "endpoint" | "fallback";
maxModels?: number;
models?: ClaudeDesktopModelStatus[];
}
interface ClaudeDesktopModelStatus {
name: string;
displayName: string;
description?: string;
cloud?: boolean;
selected: boolean;
availability?: "unknown" | "available" | "unavailable";
reason?:
| "cloud_off"
| "sign_in_required"
| "upgrade_required"
| "verification_unavailable"
| "model_not_installed";
requiredPlan?: string;
}
interface ClaudeDesktopActionResult {
status: ClaudeDesktopStatus;
error?: string;
}
type ClaudeDesktopInstallResult = "opened" | "cancelled" | "failed";
interface WebviewAPI {
selectFile: () => Promise<ImageData | null>;
selectMultipleFiles: () => Promise<ImageData[] | null>;
@@ -24,9 +63,24 @@ declare global {
webview?: WebviewAPI;
drag?: () => void;
doubleClick?: () => void;
activateOllama?: () => void;
getClaudeDesktopStatus?: () => Promise<ClaudeDesktopStatus>;
setClaudeDesktopConnected?: (
enabled: boolean,
) => Promise<ClaudeDesktopActionResult>;
prepareClaudeDesktopConnection?: () => Promise<ClaudeDesktopActionResult>;
openClaudeDesktop?: () => Promise<string>;
installClaudeDesktop?: () => Promise<ClaudeDesktopInstallResult>;
getShowAppsInMenu?: () => Promise<boolean>;
setShowAppsInMenu?: (visible: boolean) => Promise<void>;
restartClaudeDesktop?: (
models: string[],
) => Promise<ClaudeDesktopActionResult>;
setOnboardingWindow?: (enabled: boolean) => void;
menu: (items: MenuItem[]) => Promise<string | null>;
OLLAMA_TOOLS?: boolean;
OLLAMA_WEBSEARCH?: boolean;
OLLAMA_PLATFORM?: "darwin" | "windows";
}
namespace JSX {
@@ -46,4 +100,13 @@ declare global {
}
}
export type { ImageData, WebviewAPI, ContextMenuItem, ContextMenuResult };
export type {
ClaudeDesktopActionResult,
ClaudeDesktopInstallResult,
ClaudeDesktopModelStatus,
ClaudeDesktopStatus,
ContextMenuItem,
ContextMenuResult,
ImageData,
WebviewAPI,
};
+145 -4
View File
@@ -31,6 +31,7 @@ import (
"github.com/ollama/ollama/app/updater"
"github.com/ollama/ollama/app/version"
ollamaAuth "github.com/ollama/ollama/auth"
"github.com/ollama/ollama/cmd/launch"
"github.com/ollama/ollama/envconfig"
"github.com/ollama/ollama/manifest"
"github.com/ollama/ollama/types/model"
@@ -110,8 +111,10 @@ type Server struct {
Dev bool
// Updater for checking and downloading updates
Updater *updater.Updater
UpdateAvailableFunc func()
Updater *updater.Updater
UpdateAvailableFunc func()
IntegrationInstalled func(string) bool
ListCloudModels func(context.Context) (*api.ListResponse, error)
}
func (s *Server) log() *slog.Logger {
@@ -292,6 +295,8 @@ func (s *Server) Handler() http.Handler {
mux.Handle("POST /api/v1/settings", handle(s.settings))
mux.Handle("GET /api/v1/cloud", handle(s.getCloudSetting))
mux.Handle("POST /api/v1/cloud", handle(s.cloudSetting))
mux.Handle("GET /api/v1/models/cloud", handle(s.getCloudModels))
mux.Handle("GET /api/v1/integrations", handle(s.getIntegrationStatuses))
// Ollama proxy endpoints
ollamaProxy := s.ollamaProxy()
@@ -314,6 +319,74 @@ func (s *Server) Handler() http.Handler {
return mux
}
func (s *Server) getIntegrationStatuses(w http.ResponseWriter, _ *http.Request) error {
isInstalled := s.IntegrationInstalled
if isInstalled == nil {
isInstalled = launch.IsIntegrationInstalled
}
type integrationStatus struct {
ID string `json:"id"`
Name string `json:"name"`
Description string `json:"description"`
Installed *bool `json:"installed,omitempty"`
Action string `json:"action"`
Command string `json:"command,omitempty"`
}
infos := launch.ListIntegrationInfos()
statuses := make([]integrationStatus, 0, len(infos)+2)
claudeDesktopInstalled := isInstalled("claude-desktop")
statuses = append(statuses, integrationStatus{
ID: "claude-desktop",
Name: "Claude",
Description: "Use Ollama models in Claude Desktop",
Installed: &claudeDesktopInstalled,
Action: "connect",
})
byName := make(map[string]launch.IntegrationInfo, len(infos))
for _, info := range infos {
byName[info.Name] = info
}
seen := map[string]bool{"chatgpt": true}
launcherMenuOrder := []string{"claude", "codex", "openclaw", "opencode", "droid", "pi", "cline"}
orderedInfos := make([]launch.IntegrationInfo, 0, len(infos))
for _, name := range launcherMenuOrder {
if info, ok := byName[name]; ok {
orderedInfos = append(orderedInfos, info)
seen[name] = true
}
}
for _, info := range infos {
if !seen[info.Name] {
orderedInfos = append(orderedInfos, info)
}
}
for _, info := range orderedInfos {
installed := isInstalled(info.Name)
statuses = append(statuses, integrationStatus{
ID: info.Name,
Name: info.DisplayName,
Description: info.Description,
Installed: &installed,
Action: "copy",
Command: "ollama launch " + info.Name,
})
}
statuses = append(statuses, integrationStatus{
ID: "terminal",
Name: "Terminal",
Description: "Run local models from your terminal",
Action: "copy",
Command: "ollama",
})
return json.NewEncoder(w).Encode(statuses)
}
// handleError renders appropriate error responses based on request type
func (s *Server) handleError(w http.ResponseWriter, e error) {
// Preserve CORS headers for API requests
@@ -1464,11 +1537,27 @@ func (s *Server) settings(w http.ResponseWriter, r *http.Request) error {
return fmt.Errorf("failed to load settings: %w", err)
}
var settings store.Settings
if err := json.NewDecoder(r.Body).Decode(&settings); err != nil {
var request struct {
store.Settings
OnboardingVersion *int
ClaudeDesktopUsed *bool
}
if err := json.NewDecoder(r.Body).Decode(&request); err != nil {
return fmt.Errorf("invalid request body: %w", err)
}
settings := request.Settings
if request.OnboardingVersion == nil {
settings.OnboardingVersion = old.OnboardingVersion
} else {
settings.OnboardingVersion = *request.OnboardingVersion
}
if request.ClaudeDesktopUsed == nil {
settings.ClaudeDesktopUsed = old.ClaudeDesktopUsed
} else {
settings.ClaudeDesktopUsed = *request.ClaudeDesktopUsed
}
if err := s.Store.SetSettings(settings); err != nil {
return fmt.Errorf("failed to save settings: %w", err)
}
@@ -1537,6 +1626,58 @@ func (s *Server) writeCloudStatus(w http.ResponseWriter) error {
})
}
func (s *Server) getCloudModels(w http.ResponseWriter, r *http.Request) error {
w.Header().Set("Content-Type", "application/json")
disabled, _, err := s.Store.CloudStatus()
if err != nil {
return fmt.Errorf("failed to load cloud status: %w", err)
}
if disabled {
return json.NewEncoder(w).Encode(api.ListResponse{Models: []api.ListModelResponse{}})
}
list := s.ListCloudModels
if list == nil {
list = s.listCloudModels
}
ctx, cancel := context.WithTimeout(r.Context(), 3*time.Second)
defer cancel()
models, err := list(ctx)
if err != nil {
var authErr api.AuthorizationError
if errors.As(err, &authErr) && (authErr.StatusCode == http.StatusUnauthorized || authErr.StatusCode == http.StatusForbidden) {
return json.NewEncoder(w).Encode(api.ListResponse{Models: []api.ListModelResponse{}})
}
return fmt.Errorf("failed to list cloud models: %w", err)
}
if models == nil {
models = &api.ListResponse{Models: []api.ListModelResponse{}}
}
return json.NewEncoder(w).Encode(models)
}
func (s *Server) listCloudModels(ctx context.Context) (*api.ListResponse, error) {
resp, err := s.doSelfSigned(ctx, http.MethodGet, "/api/tags")
if err != nil {
return nil, err
}
defer resp.Body.Close()
if resp.StatusCode == http.StatusUnauthorized || resp.StatusCode == http.StatusForbidden {
return nil, api.AuthorizationError{StatusCode: resp.StatusCode}
}
if resp.StatusCode != http.StatusOK {
return nil, fmt.Errorf("ollama.com/api/tags returned %s", resp.Status)
}
var models api.ListResponse
if err := json.NewDecoder(resp.Body).Decode(&models); err != nil {
return nil, fmt.Errorf("failed to parse cloud models: %w", err)
}
return &models, nil
}
func (s *Server) getInferenceCompute(w http.ResponseWriter, r *http.Request) error {
ctx, cancel := context.WithTimeout(r.Context(), 500*time.Millisecond)
defer cancel()
+226
View File
@@ -18,6 +18,7 @@ import (
"github.com/ollama/ollama/api"
"github.com/ollama/ollama/app/store"
"github.com/ollama/ollama/app/updater"
"github.com/ollama/ollama/cmd/launch"
)
func TestHandlePostApiSettings(t *testing.T) {
@@ -119,6 +120,80 @@ func TestHandlePostApiSettings(t *testing.T) {
}
}
func TestGetIntegrationStatuses(t *testing.T) {
server := &Server{
IntegrationInstalled: func(name string) bool {
return name == "claude-desktop" || name == "codex"
},
}
req := httptest.NewRequest(http.MethodGet, "/api/v1/integrations", nil)
rr := httptest.NewRecorder()
if err := server.getIntegrationStatuses(rr, req); err != nil {
t.Fatalf("getIntegrationStatuses() error = %v", err)
}
var got []struct {
ID string `json:"id"`
Installed *bool `json:"installed"`
Action string `json:"action"`
Command string `json:"command"`
}
if err := json.NewDecoder(rr.Body).Decode(&got); err != nil {
t.Fatalf("decode response: %v", err)
}
if len(got) < 5 {
t.Fatalf("got %d integrations, want the full registry", len(got))
}
if got[0].ID != "claude-desktop" || got[0].Action != "connect" || got[0].Command != "" {
t.Fatalf("first integration = %+v, want command-free Claude Desktop connect", got[0])
}
wantPrefix := []string{"claude-desktop", "claude", "codex", "openclaw", "opencode", "droid", "pi", "cline"}
for i, want := range wantPrefix {
if got[i].ID != want {
t.Fatalf("integration %d = %q, want launcher menu order entry %q", i, got[i].ID, want)
}
}
byID := make(map[string]struct {
Installed *bool
Action string
Command string
}, len(got))
for _, item := range got {
byID[item.ID] = struct {
Installed *bool
Action string
Command string
}{item.Installed, item.Action, item.Command}
}
for name, want := range map[string]bool{
"claude-desktop": true,
"opencode": false,
"codex": true,
} {
item, ok := byID[name]
if !ok || item.Installed == nil || *item.Installed != want {
t.Errorf("%s installed = %v, want %v", name, item.Installed, want)
}
}
if item, ok := byID["claude"]; !ok || item.Command != "ollama launch claude" {
t.Fatal("Claude Code should follow Claude Desktop with its launch command")
}
if _, ok := byID["chatgpt"]; ok {
t.Fatal("ChatGPT should be excluded from onboarding integrations")
}
wantCount := len(launch.ListIntegrationInfos()) + 1 // Claude Desktop and Terminal replace omitted ChatGPT.
if len(got) != wantCount {
t.Fatalf("got %d integrations, want %d launcher entries", len(got), wantCount)
}
terminal := got[len(got)-1]
if terminal.ID != "terminal" || terminal.Installed != nil || terminal.Command != "ollama" {
t.Fatalf("last integration = %+v, want Terminal without install status", terminal)
}
}
func TestHandlePostApiCloudSetting(t *testing.T) {
tmpHome := t.TempDir()
t.Setenv("HOME", tmpHome)
@@ -527,6 +602,63 @@ func TestUserAgentTransport(t *testing.T) {
t.Logf("User-Agent transport successfully set: %s", receivedUA)
}
func TestGetCloudModels(t *testing.T) {
t.Run("does not call ollama.com when cloud is disabled", func(t *testing.T) {
t.Setenv("HOME", t.TempDir())
t.Setenv("OLLAMA_NO_CLOUD", "1")
testStore := &store.Store{DBPath: filepath.Join(t.TempDir(), "db.sqlite")}
defer testStore.Close()
server := &Server{
Store: testStore,
ListCloudModels: func(context.Context) (*api.ListResponse, error) {
t.Fatal("cloud model list called while cloud was disabled")
return nil, nil
},
}
req := httptest.NewRequest(http.MethodGet, "/api/v1/models/cloud", nil)
rr := httptest.NewRecorder()
if err := server.getCloudModels(rr, req); err != nil {
t.Fatal(err)
}
var got api.ListResponse
if err := json.NewDecoder(rr.Body).Decode(&got); err != nil {
t.Fatal(err)
}
if len(got.Models) != 0 {
t.Fatalf("models = %+v, want none", got.Models)
}
})
t.Run("returns no cloud models when account is unauthorized", func(t *testing.T) {
t.Setenv("HOME", t.TempDir())
t.Setenv("OLLAMA_NO_CLOUD", "")
testStore := &store.Store{DBPath: filepath.Join(t.TempDir(), "db.sqlite")}
defer testStore.Close()
server := &Server{
Store: testStore,
ListCloudModels: func(context.Context) (*api.ListResponse, error) {
return nil, api.AuthorizationError{StatusCode: http.StatusUnauthorized}
},
}
req := httptest.NewRequest(http.MethodGet, "/api/v1/models/cloud", nil)
rr := httptest.NewRecorder()
if err := server.getCloudModels(rr, req); err != nil {
t.Fatal(err)
}
var got api.ListResponse
if err := json.NewDecoder(rr.Body).Decode(&got); err != nil {
t.Fatal(err)
}
if len(got.Models) != 0 {
t.Fatalf("models = %+v, want none", got.Models)
}
})
}
func TestInferenceClientUsesUserAgent(t *testing.T) {
var gotUserAgent atomic.Value
ts := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
@@ -718,6 +850,100 @@ func TestSettingsToggleAutoUpdateOff_CancelsDownload(t *testing.T) {
}
}
func TestSettingsPreservesOnboardingVersionWhenOmitted(t *testing.T) {
testStore := &store.Store{
DBPath: filepath.Join(t.TempDir(), "db.sqlite"),
}
defer testStore.Close()
settings, err := testStore.Settings()
if err != nil {
t.Fatal(err)
}
settings.OnboardingVersion = 1
if err := testStore.SetSettings(settings); err != nil {
t.Fatal(err)
}
payload, err := json.Marshal(settings)
if err != nil {
t.Fatal(err)
}
var fields map[string]any
if err := json.Unmarshal(payload, &fields); err != nil {
t.Fatal(err)
}
delete(fields, "OnboardingVersion")
payload, err = json.Marshal(fields)
if err != nil {
t.Fatal(err)
}
server := &Server{
Store: testStore,
Restart: func() {},
}
req := httptest.NewRequest("POST", "/api/v1/settings", bytes.NewReader(payload))
rr := httptest.NewRecorder()
if err := server.settings(rr, req); err != nil {
t.Fatalf("settings() error = %v", err)
}
saved, err := testStore.Settings()
if err != nil {
t.Fatal(err)
}
if saved.OnboardingVersion != 1 {
t.Fatalf("OnboardingVersion = %d, want 1", saved.OnboardingVersion)
}
}
func TestSettingsPreservesClaudeDesktopUsedWhenOmitted(t *testing.T) {
testStore := &store.Store{
DBPath: filepath.Join(t.TempDir(), "db.sqlite"),
}
defer testStore.Close()
settings, err := testStore.Settings()
if err != nil {
t.Fatal(err)
}
settings.ClaudeDesktopUsed = true
if err := testStore.SetSettings(settings); err != nil {
t.Fatal(err)
}
payload, err := json.Marshal(settings)
if err != nil {
t.Fatal(err)
}
var fields map[string]any
if err := json.Unmarshal(payload, &fields); err != nil {
t.Fatal(err)
}
delete(fields, "ClaudeDesktopUsed")
payload, err = json.Marshal(fields)
if err != nil {
t.Fatal(err)
}
server := &Server{Store: testStore, Restart: func() {}}
req := httptest.NewRequest("POST", "/api/v1/settings", bytes.NewReader(payload))
rr := httptest.NewRecorder()
if err := server.settings(rr, req); err != nil {
t.Fatalf("settings() error = %v", err)
}
saved, err := testStore.Settings()
if err != nil {
t.Fatal(err)
}
if !saved.ClaudeDesktopUsed {
t.Fatal("expected ClaudeDesktopUsed to be preserved")
}
}
func TestSettingsToggleAutoUpdateOn_WithPendingUpdate_ShowsNotification(t *testing.T) {
testStore := &store.Store{
DBPath: filepath.Join(t.TempDir(), "db.sqlite"),
+9 -3
View File
@@ -153,10 +153,12 @@ func (u *Updater) DownloadNewRelease(ctx context.Context, updateResp UpdateRespo
return err
}
// In case of slow downloads, continue the update check in the background
// In case of slow downloads, continue the update check in the background.
// Drain the goroutine before returning: it reads package-level knobs
// (e.g. UpdateCheckInterval), which callers may mutate once we return.
bgctx, bgcancel := context.WithCancel(downloadCtx)
defer bgcancel()
go func() {
var bgwg sync.WaitGroup
bgwg.Go(func() {
for {
select {
case <-bgctx.Done():
@@ -165,6 +167,10 @@ func (u *Updater) DownloadNewRelease(ctx context.Context, updateResp UpdateRespo
u.checkForUpdate(bgctx)
}
}
})
defer func() {
bgcancel()
bgwg.Wait()
}()
resp, err := http.DefaultClient.Do(req)
+1 -1
View File
@@ -434,7 +434,7 @@ func IsUpdatePending() bool {
func chownWithAuthorization(user string) bool {
u := C.CString(user)
defer C.free(unsafe.Pointer(u))
return (bool)(C.chownWithAuthorization(u))
return bool(C.chownWithAuthorization(u))
}
func verifyExtractedBundle(path string) error {
+28
View File
@@ -190,6 +190,23 @@ func TestDownloadNewReleaseDoesNotUseRawETagAsPathComponent(t *testing.T) {
}
}
// waitDownloadIdle blocks until no download is in flight, so staged-file
// handles close before t.TempDir cleanup removes the stage directory. After
// the context is cancelled a new download can't write (it aborts at the HEAD
// request), so reaching idle makes cleanup race-free.
func (u *Updater) waitDownloadIdle() {
deadline := time.Now().Add(2 * time.Second)
for time.Now().Before(deadline) {
u.cancelDownloadLock.Lock()
idle := u.cancelDownload == nil
u.cancelDownloadLock.Unlock()
if idle {
return
}
time.Sleep(time.Millisecond)
}
}
func TestBackgroundCheckerSkipsAlreadyStagedETagDownload(t *testing.T) {
UpdateStageDir = t.TempDir()
oldInstaller := Installer
@@ -276,6 +293,7 @@ func TestBackgroundCheckerSkipsAlreadyStagedETagDownload(t *testing.T) {
callbacks <- ver
return nil
})
t.Cleanup(updater.waitDownloadIdle)
for range 2 {
select {
@@ -364,6 +382,7 @@ func TestBackgoundChecker(t *testing.T) {
}
updater.StartBackgroundUpdaterChecker(ctx, cb)
t.Cleanup(updater.waitDownloadIdle)
select {
case <-stallTimer.C:
t.Fatal("stalled")
@@ -426,6 +445,7 @@ func TestAutoUpdateDisabledSkipsDownload(t *testing.T) {
}
updater.StartBackgroundUpdaterChecker(ctx, cb)
t.Cleanup(updater.waitDownloadIdle)
// Wait enough time for multiple check cycles
time.Sleep(50 * time.Millisecond)
@@ -488,6 +508,7 @@ func TestAutoUpdateReenabledDownloadsUpdate(t *testing.T) {
}
upd.StartBackgroundUpdaterChecker(ctx, cb)
t.Cleanup(upd.waitDownloadIdle)
// Wait for a few cycles with auto-update disabled - no download should happen
time.Sleep(50 * time.Millisecond)
@@ -556,7 +577,9 @@ func TestCancelOngoingDownload(t *testing.T) {
_, resp := updater.checkForUpdate(ctx)
// Start download in goroutine
downloadDone := make(chan struct{})
go func() {
defer close(downloadDone)
_ = updater.DownloadNewRelease(ctx, resp)
}()
@@ -577,6 +600,10 @@ func TestCancelOngoingDownload(t *testing.T) {
case <-time.After(2 * time.Second):
t.Fatal("download cancellation was not received by server")
}
// Wait for the download goroutine to unwind: it drags along a background
// update-check loop that reads package-level knobs the next test rewrites.
<-downloadDone
}
func TestTriggerImmediateCheck(t *testing.T) {
@@ -615,6 +642,7 @@ func TestTriggerImmediateCheck(t *testing.T) {
}
updater.StartBackgroundUpdaterChecker(ctx, cb)
t.Cleanup(updater.waitDownloadIdle)
// Wait for the initial check that fires after the initial delay
select {
+3 -3
View File
@@ -78,7 +78,7 @@ func init() {
func loadOSVersion() {
UserAgentOS = "Windows"
verInfo := OSVERSIONINFOEXW{}
verInfo.dwOSVersionInfoSize = (uint32)(unsafe.Sizeof(verInfo))
verInfo.dwOSVersionInfoSize = uint32(unsafe.Sizeof(verInfo))
ntdll, err := windows.LoadDLL("ntdll.dll")
if err != nil {
slog.Warn("unable to find ntdll", "error", err)
@@ -394,13 +394,13 @@ func IsProcRunning(procName string) []uint32 {
defer windows.CloseHandle(hProcess)
var module windows.Handle
var cbNeeded uint32
cb := (uint32)(unsafe.Sizeof(module))
cb := uint32(unsafe.Sizeof(module))
if err := windows.EnumProcessModules(hProcess, &module, cb, &cbNeeded); err != nil {
continue
}
var sz uint32 = 1024 * 8
moduleName := make([]uint16, sz)
cb = uint32(len(moduleName)) * (uint32)(unsafe.Sizeof(uint16(0)))
cb = uint32(len(moduleName)) * uint32(unsafe.Sizeof(uint16(0)))
if err := windows.GetModuleBaseName(hProcess, module, &moduleName[0], cb); err != nil && err != syscall.ERROR_INSUFFICIENT_BUFFER {
continue
}
+2 -12
View File
@@ -2535,19 +2535,9 @@ inline SIZE make_window_frame_size(HWND window, int width, int height,
return {frame_width, frame_height};
}
inline bool is_dark_theme_enabled() {
constexpr auto *sub_key =
L"SOFTWARE\\Microsoft\\Windows\\CurrentVersion\\Themes\\Personalize";
reg_key key(HKEY_CURRENT_USER, sub_key, 0, KEY_READ);
if (!key.is_open()) {
// Default is light theme
return false;
}
return key.query_uint(L"AppsUseLightTheme", 1) == 0;
}
inline void apply_window_theme(HWND window) {
auto dark_theme_enabled = is_dark_theme_enabled();
// Ollama uses a light-only application appearance.
constexpr bool dark_theme_enabled = false;
// Use "immersive dark mode" on systems that support it.
// Changes the color of the window's title bar (light or dark).
+11 -10
View File
@@ -78,9 +78,9 @@ func (t *winTray) wndProc(hWnd windows.Handle, message uint32, wParam, lParam ui
t.app.Quit()
case updateMenuID:
t.app.DoUpdate()
case openUIMenuID:
case openAppsMenuID:
// UI must be initialized on this thread so don't use the callbacks
t.app.UIShow()
t.app.UIRun("/connect")
case settingsUIMenuID:
// UI must be initialized on this thread so don't use the callbacks
t.app.UIRun("/settings")
@@ -174,14 +174,7 @@ func (t *winTray) wndProc(hWnd windows.Handle, message uint32, wParam, lParam ui
}
}
case uint32(FOCUS_WINDOW_MSG_ID):
// Handle focus window request from another instance
if t.app.UIRunning() {
// If UI is already running, just show it
t.app.UIShow()
} else {
// If UI is not running, start it
t.app.UIRun("/")
}
focusUI(t.app)
lResult = 1 // Return non-zero to indicate success
default:
// Calls the default window procedure to provide default processing for any window messages that an application does not process.
@@ -197,6 +190,14 @@ func (t *winTray) wndProc(hWnd windows.Handle, message uint32, wParam, lParam ui
return
}
func focusUI(app AppCallbacks) {
if app.UIRunning() && app.UIOnboarding() {
app.UIShow()
return
}
app.UIRun("/connect")
}
func (t *winTray) Quit() {
// slog.Debug("XXX in winTray.Quit")
t.quitting = true
+47
View File
@@ -0,0 +1,47 @@
//go:build windows
package wintray
import "testing"
type lifecycleApp struct {
running bool
onboarding bool
runPath string
showCall bool
}
func (a *lifecycleApp) UIRun(path string) { a.runPath = path }
func (a *lifecycleApp) UIShow() { a.showCall = true }
func (a *lifecycleApp) UITerminate() {}
func (a *lifecycleApp) UIRunning() bool { return a.running }
func (a *lifecycleApp) UIOnboarding() bool { return a.onboarding }
func (a *lifecycleApp) Quit() {}
func (a *lifecycleApp) DoUpdate() {}
func TestFocusUICreatesOrShowsWindow(t *testing.T) {
tests := []struct {
name string
running bool
onboarding bool
wantRun string
wantShow bool
}{
{name: "creates Apps window when tray only", wantRun: "/connect"},
{name: "routes existing window to Apps", running: true, wantRun: "/connect"},
{name: "preserves onboarding", running: true, onboarding: true, wantShow: true},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
app := &lifecycleApp{running: tt.running, onboarding: tt.onboarding}
focusUI(app)
if app.runPath != tt.wantRun {
t.Errorf("run path = %q, want %q", app.runPath, tt.wantRun)
}
if app.showCall != tt.wantShow {
t.Errorf("show called = %v, want %v", app.showCall, tt.wantShow)
}
})
}
}
+2 -2
View File
@@ -16,7 +16,7 @@ import (
const (
_ = iota
openUIMenuID
openAppsMenuID
settingsUIMenuID
updateSeparatorMenuID
updateAvailableMenuID
@@ -28,7 +28,7 @@ const (
)
func (t *winTray) initMenus() error {
if err := t.addOrUpdateMenuItem(openUIMenuID, 0, openUIMenuTitle, false); err != nil {
if err := t.addOrUpdateMenuItem(openAppsMenuID, 0, openAppsMenuTitle, false); err != nil {
return fmt.Errorf("unable to create menu entries %w", err)
}
if err := t.addOrUpdateMenuItem(settingsUIMenuID, 0, settingsUIMenuTitle, false); err != nil {
+2 -2
View File
@@ -12,6 +12,6 @@ const (
updateAvailableMenuTitle = "An update is available"
updateMenuTitle = "Restart to update"
diagLogsMenuTitle = "View logs"
openUIMenuTitle = "Open Ollama"
settingsUIMenuTitle = "Settings..."
openAppsMenuTitle = "Open Ollama"
settingsUIMenuTitle = "Settings"
)
+1
View File
@@ -49,6 +49,7 @@ type AppCallbacks interface {
UIShow()
UITerminate()
UIRunning() bool
UIOnboarding() bool
Quit()
DoUpdate()
}
+119 -15
View File
@@ -59,7 +59,7 @@ function(ollama_macos_major_version output)
RESULT_VARIABLE _macos_result
ERROR_QUIET)
if(_macos_result EQUAL 0)
string(REGEX MATCH "^[0-9]+" _macos_major "${_macos_version}")
string(REGEX MATCH "^[0-9]+(\\.[0-9]+)?" _macos_major "${_macos_version}")
endif()
set(${output} "${_macos_major}" PARENT_SCOPE)
endfunction()
@@ -72,7 +72,7 @@ function(ollama_macos_sdk_major_version output)
RESULT_VARIABLE _sdk_result
ERROR_QUIET)
if(_sdk_result EQUAL 0)
string(REGEX MATCH "^[0-9]+" _sdk_major "${_sdk_version}")
string(REGEX MATCH "^[0-9]+(\\.[0-9]+)?" _sdk_major "${_sdk_version}")
endif()
set(${output} "${_sdk_major}" PARENT_SCOPE)
endfunction()
@@ -83,7 +83,9 @@ function(ollama_default_mlx_backends output)
ollama_check_metal_toolchain(_metal_version)
ollama_macos_major_version(_macos_major)
ollama_macos_sdk_major_version(_sdk_major)
if(_macos_major AND _sdk_major AND _macos_major GREATER_EQUAL 26 AND _sdk_major GREATER_EQUAL 26)
if(_macos_major AND _sdk_major
AND _macos_major VERSION_GREATER_EQUAL 26.2
AND _sdk_major VERSION_GREATER_EQUAL 26.2)
set(_backends "metal_v4")
else()
set(_backends "metal_v3")
@@ -167,6 +169,14 @@ if(OLLAMA_MLX_BACKENDS)
list(APPEND _mlx_source_targets ollama-mlx-source)
endif()
# Temporary MLX-C carry patch: regenerated bindings for force_fused and the
# thread-local compile cache, carried until they merge upstream into
# ml-explore/mlx-c. Then bump MLX_C_VERSION and delete mlx/compat/.
find_package(Git REQUIRED)
set(OLLAMA_MLX_C_COMPAT_PATCH_COMMAND
${GIT_EXECUTABLE} apply ${CMAKE_SOURCE_DIR}/mlx/compat/0001-mlx-c-regen-0.32.1.patch
CACHE INTERNAL "MLX-C carry patch")
if(DEFINED "FETCHCONTENT_SOURCE_DIR_MLX-C" AND NOT "${FETCHCONTENT_SOURCE_DIR_MLX-C}" STREQUAL "")
get_filename_component(OLLAMA_MLX_C_SOURCE_DIR
"${FETCHCONTENT_SOURCE_DIR_MLX-C}" ABSOLUTE BASE_DIR "${CMAKE_SOURCE_DIR}")
@@ -186,14 +196,35 @@ if(OLLAMA_MLX_BACKENDS)
CONFIGURE_COMMAND ""
BUILD_COMMAND ""
INSTALL_COMMAND ""
USES_TERMINAL_DOWNLOAD TRUE)
PATCH_COMMAND ${OLLAMA_MLX_C_COMPAT_PATCH_COMMAND}
USES_TERMINAL_DOWNLOAD TRUE
USES_TERMINAL_PATCH TRUE)
list(APPEND _mlx_source_targets ollama-mlx-c-source)
endif()
add_custom_target(ollama-mlx-sources DEPENDS ${_mlx_source_targets})
# Refresh the vendored MLX-C headers once the sources are present. Every MLX
# backend variant shares this destination in the source tree, so the copy has
# to happen here rather than in each variant's build.
add_custom_target(ollama-mlx-vendor-headers
COMMAND ${CMAKE_COMMAND}
-DMLX_C_HEADERS_DIR=${OLLAMA_MLX_C_SOURCE_DIR}/mlx/c
-DMLX_C_HEADERS_DEST=${CMAKE_SOURCE_DIR}/x/mlxrunner/mlx/include/mlx/c
-P "${CMAKE_SOURCE_DIR}/cmake/vendor-mlx-c-headers.cmake"
DEPENDS ${_mlx_source_targets}
COMMENT "Vendoring MLX-C headers"
VERBATIM)
add_custom_target(ollama-mlx-sources DEPENDS ollama-mlx-vendor-headers)
endif()
set(OLLAMA_BUILD_PARALLEL "" CACHE STRING
"Number of parallel jobs for nested native builds (empty = use generator default)")
set(_native_parallel_args --parallel)
if(NOT OLLAMA_BUILD_PARALLEL STREQUAL "")
list(APPEND _native_parallel_args ${OLLAMA_BUILD_PARALLEL})
endif()
set(OLLAMA_NATIVE_BUILD_TOOL_COMMAND
${CMAKE_COMMAND} --build <BINARY_DIR>)
${CMAKE_COMMAND} --build <BINARY_DIR> ${_native_parallel_args})
set(OLLAMA_NATIVE_BUILD_TARGET_ARG --target)
if(CMAKE_GENERATOR MATCHES "Makefiles")
set(OLLAMA_NATIVE_BUILD_TOOL_COMMAND
@@ -236,6 +267,67 @@ function(ollama_cache_arg_is_set name output)
endif()
endfunction()
function(ollama_backend_cuda_major backend output)
if("${backend}" MATCHES "^cuda_v([0-9]+)$")
set(${output} "${CMAKE_MATCH_1}" PARENT_SCOPE)
else()
set(${output} "" PARENT_SCOPE)
endif()
endfunction()
function(ollama_find_windows_cuda_root major output)
if(NOT WIN32 OR "${major}" STREQUAL "")
set(${output} "" PARENT_SCOPE)
return()
endif()
execute_process(
COMMAND ${CMAKE_COMMAND} -E environment
OUTPUT_VARIABLE _environment)
string(REPLACE "\r\n" "\n" _environment "${_environment}")
string(REPLACE "\r" "\n" _environment "${_environment}")
string(REGEX MATCHALL "CUDA_PATH_V${major}_[0-9]+=[^\n]*" _matches "${_environment}")
set(_best_minor -1)
set(_best_root "")
foreach(_entry IN LISTS _matches)
if(_entry MATCHES "^CUDA_PATH_V${major}_([0-9]+)=(.*)$")
set(_minor "${CMAKE_MATCH_1}")
set(_root "${CMAKE_MATCH_2}")
if(_minor GREATER _best_minor)
set(_best_minor ${_minor})
set(_best_root "${_root}")
endif()
endif()
endforeach()
if(_best_root STREQUAL "" AND DEFINED ENV{CUDA_PATH})
set(_cuda_path "$ENV{CUDA_PATH}")
if(EXISTS "${_cuda_path}/version.json")
file(READ "${_cuda_path}/version.json" _version_json)
if(_version_json MATCHES "\"cuda\"[ \t\r\n]*:[ \t\r\n]*\"${major}\\.")
set(_best_root "${_cuda_path}")
endif()
endif()
endif()
set(${output} "${_best_root}" PARENT_SCOPE)
endfunction()
function(ollama_append_cuda_toolkit_args output backend)
# If CUDAToolkit_ROOT is already explicitly set, just forward it.
ollama_append_cache_arg_if_set(${output} CUDAToolkit_ROOT)
if(NOT DEFINED CUDAToolkit_ROOT OR "${CUDAToolkit_ROOT}" STREQUAL "")
# Auto-discover CUDA toolkit for the requested backend version on Windows.
ollama_backend_cuda_major("${backend}" _cuda_major)
ollama_find_windows_cuda_root("${_cuda_major}" _cuda_root)
if(NOT "${_cuda_root}" STREQUAL "")
ollama_escape_cmake_list("${_cuda_root}" _value)
set(${output} ${${output}} "-DCUDAToolkit_ROOT=${_value}" PARENT_SCOPE)
endif()
endif()
endfunction()
function(ollama_llama_cuda_preset backend output)
ollama_cache_arg_is_set(CMAKE_CUDA_ARCHITECTURES _has_cuda_arch)
if(_has_cuda_arch)
@@ -327,12 +419,28 @@ function(ollama_add_llama_server_build name)
-DCMAKE_OSX_DEPLOYMENT_TARGET=${CMAKE_OSX_DEPLOYMENT_TARGET})
endif()
endif()
# Visual Studio requires -T toolset override to select the correct CUDA toolkit.
# MSBuild's CUDA integration ignores -DCUDAToolkit_ROOT for nvcc selection.
# Prefer user-specified CUDAToolkit_ROOT before falling back to auto-discovery.
set(_generator_args)
if(WIN32 AND CMAKE_GENERATOR MATCHES "Visual Studio")
set(_cuda_root "${CUDAToolkit_ROOT}")
if("${_cuda_root}" STREQUAL "")
ollama_backend_cuda_major("${name}" _cuda_major)
ollama_find_windows_cuda_root("${_cuda_major}" _cuda_root)
endif()
if(NOT "${_cuda_root}" STREQUAL "")
list(APPEND _generator_args -T cuda=${_cuda_root})
endif()
endif()
set(_configure_command ${CMAKE_COMMAND}
${_generator_args}
-S ${CMAKE_SOURCE_DIR}/llama/server
-B <BINARY_DIR>
${_cmake_args})
if(ARG_PRESET)
set(_configure_command ${CMAKE_COMMAND}
${_generator_args}
-S ${CMAKE_SOURCE_DIR}/llama/server
--preset ${ARG_PRESET}
-B <BINARY_DIR>
@@ -441,15 +549,8 @@ endfunction()
find_program(GO_EXECUTABLE go)
if(OLLAMA_MLX_BACKENDS)
set(_mlx_c_headers_dir "${OLLAMA_MLX_C_SOURCE_DIR}/mlx/c")
set(_mlx_c_headers_dest "${CMAKE_SOURCE_DIR}/x/mlxrunner/mlx/include/mlx/c")
if(GO_EXECUTABLE AND (NOT APPLE OR CMAKE_SYSTEM_PROCESSOR STREQUAL CMAKE_HOST_SYSTEM_PROCESSOR))
add_custom_target(ollama-mlx-generate-wrappers
COMMAND ${CMAKE_COMMAND}
-DMLX_C_HEADERS_DIR=${_mlx_c_headers_dir}
-DMLX_C_HEADERS_DEST=${_mlx_c_headers_dest}
-P "${CMAKE_SOURCE_DIR}/cmake/vendor-mlx-c-headers.cmake"
COMMAND ${CMAKE_COMMAND} -E env
CC= CGO_CFLAGS= CGO_CXXFLAGS=
${GO_EXECUTABLE} generate ./x/...
@@ -544,6 +645,7 @@ if(OLLAMA_HAVE_LLAMA_SERVER)
set(_cuda_args)
ollama_append_cache_arg_if_set(_cuda_args CMAKE_CUDA_ARCHITECTURES)
ollama_append_cache_arg_if_set(_cuda_args CMAKE_CUDA_FLAGS)
ollama_append_cuda_toolkit_args(_cuda_args ${_backend})
ollama_add_llama_server_build(${_backend}
PRESET ${_cuda_preset}
RUNNER_DIR ${_backend}
@@ -555,6 +657,7 @@ if(OLLAMA_HAVE_LLAMA_SERVER)
set(_cuda_args)
ollama_append_cache_arg_if_set(_cuda_args CMAKE_CUDA_ARCHITECTURES)
ollama_append_cache_arg_if_set(_cuda_args CMAKE_CUDA_FLAGS)
ollama_append_cuda_toolkit_args(_cuda_args ${_backend})
ollama_add_llama_server_build(${_backend}
PRESET ${_cuda_preset}
RUNNER_DIR ${_backend}
@@ -664,14 +767,15 @@ foreach(_backend IN LISTS OLLAMA_MLX_BACKENDS)
endif()
ollama_check_metal_toolchain(_metal_version)
ollama_macos_sdk_major_version(_ollama_mlx_sdk_major)
if(_ollama_mlx_sdk_major AND _ollama_mlx_sdk_major GREATER_EQUAL 26)
if(_ollama_mlx_sdk_major
AND _ollama_mlx_sdk_major VERSION_GREATER_EQUAL 26.2)
ollama_add_mlx_build(metal_v4
PRESET mlx_metal_v4
RUNNER_DIR mlx_metal_v4)
list(APPEND _mlx_targets ollama-mlx-metal_v4)
else()
message(FATAL_ERROR
"OLLAMA_MLX_BACKENDS=metal_v4 requires the macOS 26 SDK. "
"OLLAMA_MLX_BACKENDS=metal_v4 requires the macOS 26.2 SDK. "
"Install a newer Xcode or use OLLAMA_MLX_BACKENDS=metal_v3.")
endif()
else()
+148 -36
View File
@@ -23,6 +23,10 @@ if(APPLE)
set(CMAKE_BUILD_RPATH "@loader_path")
set(CMAKE_INSTALL_RPATH "@loader_path")
set(CMAKE_BUILD_WITH_INSTALL_RPATH ON)
elseif(UNIX)
set(CMAKE_BUILD_RPATH "$ORIGIN")
set(CMAKE_INSTALL_RPATH "$ORIGIN")
set(CMAKE_BUILD_WITH_INSTALL_RPATH ON)
endif()
if(NOT DEFINED OLLAMA_SOURCE_DIR OR "${OLLAMA_SOURCE_DIR}" STREQUAL "")
@@ -50,7 +54,12 @@ endif()
option(OLLAMA_MLX_GENERATE_WRAPPERS "Regenerate MLX Go wrappers" OFF)
message(STATUS "Setting up MLX (this takes a while...)")
add_subdirectory(${OLLAMA_SOURCE_DIR}/x/imagegen/mlx ${CMAKE_BINARY_DIR}/x/imagegen/mlx)
foreach(_cudnn_var CUDNN_INCLUDE_PATH CUDNN_LIBRARY_PATH)
if((NOT DEFINED ${_cudnn_var} OR "${${_cudnn_var}}" STREQUAL "") AND DEFINED ENV{${_cudnn_var}})
set(${_cudnn_var} "$ENV{${_cudnn_var}}" CACHE PATH "")
endif()
endforeach()
add_subdirectory(${OLLAMA_SOURCE_DIR}/x/mlxrunner/mlx ${CMAKE_BINARY_DIR}/x/mlxrunner/mlx)
# Find CUDA toolkit if MLX is built with CUDA support.
find_package(CUDAToolkit)
@@ -65,6 +74,7 @@ elseif(DEFINED ENV{CUDNN_ROOT_DIR})
set(_cudnn_root "$ENV{CUDNN_ROOT_DIR}")
endif()
if(_cudnn_root)
file(TO_CMAKE_PATH "${_cudnn_root}" _cudnn_root)
# cuDNN 9.x has versioned subdirectories under bin/ (e.g., bin/13.0/).
file(GLOB CUDNN_BIN_SUBDIRS "${_cudnn_root}/bin/*")
list(APPEND MLX_RUNTIME_DIRS ${CUDNN_BIN_SUBDIRS})
@@ -102,7 +112,8 @@ install(RUNTIME_DEPENDENCY_SET mlx_runtime_deps
LIBRARY DESTINATION ${OLLAMA_INSTALL_DIR} COMPONENT MLX_VENDOR
)
if(TARGET jaccl)
get_target_property(_MLX_LINK_LIBRARIES mlx LINK_LIBRARIES)
if(TARGET jaccl AND "jaccl" IN_LIST _MLX_LINK_LIBRARIES)
install(TARGETS jaccl
RUNTIME DESTINATION ${OLLAMA_INSTALL_DIR} COMPONENT MLX
LIBRARY DESTINATION ${OLLAMA_INSTALL_DIR} COMPONENT MLX
@@ -123,29 +134,53 @@ endif()
# --component MLX. Headers are installed alongside libmlx in OLLAMA_INSTALL_DIR.
#
# Layout:
# ${OLLAMA_INSTALL_DIR}/include/cccl/{cuda,nv}/ - CCCL headers
# ${OLLAMA_INSTALL_DIR}/include/*.h - CUDA toolkit headers
# ${OLLAMA_INSTALL_DIR}/include/cccl/ - CCCL headers
# ${OLLAMA_INSTALL_DIR}/include/{cute,cutlass}/ - CUTLASS/CUTE headers
# ${OLLAMA_INSTALL_DIR}/include/ - CUDA runtime/core headers
#
# MLX's jit_module.cpp resolves CCCL via
# current_binary_dir()[.parent_path()] / "include" / "cccl"
# On Linux, MLX's jit_module.cpp resolves CCCL via
# current_binary_dir().parent_path() / "include" / "cccl", so we create a
# symlink from lib/ollama/include -> ${OLLAMA_RUNNER_DIR}/include.
# MLX's jit_module.cpp resolves JIT support headers from the backend-local
# include directory. On Linux it also probes current_binary_dir().parent_path()
# / "include", so we create a symlink from lib/ollama/include to the backend
# include directory for archive packaging.
# This will need refinement if we add multiple CUDA versions for MLX in the future.
# CUDA runtime headers are found via CUDA_PATH env var (set by mlxrunner).
if(EXISTS ${CMAKE_BINARY_DIR}/_deps/cccl-src/include/cuda)
install(DIRECTORY ${CMAKE_BINARY_DIR}/_deps/cccl-src/include/cuda
DESTINATION ${OLLAMA_INSTALL_DIR}/include/cccl
set(_mlx_jit_cccl_include_dir "")
if(CUDAToolkit_FOUND)
foreach(_dir ${CUDAToolkit_INCLUDE_DIRS})
if(EXISTS "${_dir}/cccl/cuda/std")
set(_mlx_jit_cccl_include_dir "${_dir}/cccl")
break()
endif()
endforeach()
endif()
if(NOT _mlx_jit_cccl_include_dir AND EXISTS ${CMAKE_BINARY_DIR}/_deps/cccl-src/include/cuda)
set(_mlx_jit_cccl_include_dir "${CMAKE_BINARY_DIR}/_deps/cccl-src/include")
endif()
if(_mlx_jit_cccl_include_dir)
foreach(_cccl_dir cuda nv cub thrust)
if(EXISTS "${_mlx_jit_cccl_include_dir}/${_cccl_dir}")
install(DIRECTORY "${_mlx_jit_cccl_include_dir}/${_cccl_dir}"
DESTINATION ${OLLAMA_INSTALL_DIR}/include/cccl
COMPONENT MLX)
endif()
endforeach()
endif()
if(EXISTS ${CMAKE_BINARY_DIR}/_deps/cutlass-src/include/cute)
install(DIRECTORY ${CMAKE_BINARY_DIR}/_deps/cutlass-src/include/cute
DESTINATION ${OLLAMA_INSTALL_DIR}/include
COMPONENT MLX)
install(DIRECTORY ${CMAKE_BINARY_DIR}/_deps/cccl-src/include/nv
DESTINATION ${OLLAMA_INSTALL_DIR}/include/cccl
install(DIRECTORY ${CMAKE_BINARY_DIR}/_deps/cutlass-src/include/cutlass
DESTINATION ${OLLAMA_INSTALL_DIR}/include
COMPONENT MLX)
endif()
# Install minimal CUDA toolkit headers needed by MLX JIT kernels.
# These are the transitive closure of includes from mlx/backend/cuda/device/*.cuh.
# Install CUDA runtime/core headers needed by MLX JIT kernels.
# NVIDIA's NVRTC bundled-header model is CUDA Runtime + CCCL, not the entire
# toolkit include tree. Keep CCCL coherent above, include CUTLASS/CUTE above,
# and avoid shipping unrelated SDK headers such as NPP, CUPTI, cuRAND, NVML,
# cuBLAS, cuSPARSE, and cuSOLVER.
# The Go mlxrunner sets CUDA_PATH to OLLAMA_INSTALL_DIR so MLX finds them at
# $CUDA_PATH/include/*.h via NVRTC --include-path.
# $CUDA_PATH/include via NVRTC --include-path.
if(CUDAToolkit_FOUND)
# CUDAToolkit_INCLUDE_DIRS may be a semicolon-separated list
# (e.g. ".../include;.../include/cccl"). Find the entry that
@@ -161,39 +196,97 @@ if(CUDAToolkit_FOUND)
message(WARNING "Could not find cuda_runtime_api.h in CUDAToolkit_INCLUDE_DIRS: ${CUDAToolkit_INCLUDE_DIRS}")
else()
set(_dst "${OLLAMA_INSTALL_DIR}/include")
set(_MLX_JIT_CUDA_HEADERS
set(_mlx_jit_cuda_headers
builtin_types.h
channel_descriptor.h
common_functions.h
cooperative_groups.h
cuComplex.h
cuda.h
cudaTypedefs.h
cuda_awbarrier.h
cuda_awbarrier_helpers.h
cuda_awbarrier_primitives.h
cuda_bf16.h
cuda_bf16.hpp
cuda_device_runtime_api.h
cuda_fp16.h
cuda_fp16.hpp
cuda_fp4.h
cuda_fp4.hpp
cuda_fp6.h
cuda_fp6.hpp
cuda_fp8.h
cuda_fp8.hpp
cuda_fp16.h
cuda_fp16.hpp
cuda_occupancy.h
cuda_pipeline.h
cuda_pipeline_helpers.h
cuda_pipeline_primitives.h
cuda_runtime.h
cuda_runtime_api.h
cuda_stdint.h
cudart_platform.h
device_atomic_functions.h
device_atomic_functions.hpp
device_double_functions.h
device_functions.h
device_launch_parameters.h
device_types.h
driver_functions.h
driver_types.h
fatbinary_section.h
host_config.h
host_defines.h
library_types.h
math_constants.h
math_functions.h
mma.h
nvrtc_device_runtime.h
sm_20_atomic_functions.h
sm_20_atomic_functions.hpp
sm_20_intrinsics.h
sm_20_intrinsics.hpp
sm_30_intrinsics.h
sm_30_intrinsics.hpp
sm_32_atomic_functions.h
sm_32_atomic_functions.hpp
sm_32_intrinsics.h
sm_32_intrinsics.hpp
sm_35_atomic_functions.h
sm_35_intrinsics.h
sm_60_atomic_functions.h
sm_60_atomic_functions.hpp
sm_61_intrinsics.h
sm_61_intrinsics.hpp
surface_indirect_functions.h
surface_types.h
target
texture_indirect_functions.h
texture_types.h
vector_functions.h
vector_functions.hpp
vector_types.h
)
foreach(_hdr ${_MLX_JIT_CUDA_HEADERS})
install(FILES "${_cuda_inc}/${_hdr}"
vector_types.h)
set(_mlx_jit_cuda_header_paths "")
foreach(_header IN LISTS _mlx_jit_cuda_headers)
if(EXISTS "${_cuda_inc}/${_header}")
list(APPEND _mlx_jit_cuda_header_paths "${_cuda_inc}/${_header}")
endif()
endforeach()
if(_mlx_jit_cuda_header_paths)
install(FILES ${_mlx_jit_cuda_header_paths}
DESTINATION ${_dst}
COMPONENT MLX)
endif()
foreach(_runtime_dir cooperative_groups crt)
if(EXISTS "${_cuda_inc}/${_runtime_dir}")
install(DIRECTORY "${_cuda_inc}/${_runtime_dir}"
DESTINATION ${_dst}
COMPONENT MLX)
endif()
endforeach()
# Subdirectory headers.
install(DIRECTORY "${_cuda_inc}/cooperative_groups"
DESTINATION ${_dst}
COMPONENT MLX
FILES_MATCHING PATTERN "*.h")
install(FILES "${_cuda_inc}/crt/host_defines.h"
DESTINATION "${_dst}/crt"
COMPONENT MLX)
if(NOT WIN32 AND NOT APPLE)
install(CODE "
set(_link \"${CMAKE_INSTALL_PREFIX}/${OLLAMA_LIB_DIR}/include\")
@@ -210,9 +303,10 @@ endif()
# RUNTIME_DEPENDENCIES auto-excludes it via POST_EXCLUDE_FILES_STRICT because
# dlfcn-win32 is a known CMake target with its own install rules (which install
# to the wrong destination). We must install it explicitly here.
if(WIN32)
install(FILES ${OLLAMA_BUILD_DIR}/dl.dll
DESTINATION ${OLLAMA_INSTALL_DIR}
if(WIN32 AND TARGET dl)
install(TARGETS dl
RUNTIME DESTINATION ${OLLAMA_INSTALL_DIR}
LIBRARY DESTINATION ${OLLAMA_INSTALL_DIR}
COMPONENT MLX)
endif()
@@ -226,7 +320,25 @@ if(CUDAToolkit_FOUND)
"${CUDAToolkit_LIBRARY_DIR}/libnvrtc.so*"
"${CUDAToolkit_LIBRARY_DIR}/libnvrtc-builtins.so*"
"${CUDAToolkit_LIBRARY_DIR}/libcufft.so*"
"${CUDAToolkit_LIBRARY_DIR}/libcudnn.so*")
"${CUDAToolkit_LIBRARY_DIR}/libcudnn*.so*")
if(WIN32)
file(GLOB MLX_CUDA_DLLS
"${CUDAToolkit_BIN_DIR}/nvrtc-builtins64_*.dll"
"${CUDAToolkit_BIN_DIR}/x64/nvrtc-builtins64_*.dll")
list(APPEND MLX_CUDA_LIBS ${MLX_CUDA_DLLS})
endif()
find_library(MLX_CUDNN_LIBRARY NAMES cudnn HINTS "$ENV{CUDNN_LIBRARY_PATH}")
if(MLX_CUDNN_LIBRARY)
get_filename_component(MLX_CUDNN_LIBRARY_DIR "${MLX_CUDNN_LIBRARY}" DIRECTORY)
file(GLOB MLX_CUDNN_LIBS "${MLX_CUDNN_LIBRARY_DIR}/libcudnn*.so*")
list(APPEND MLX_CUDA_LIBS ${MLX_CUDNN_LIBS})
endif()
if(WIN32 AND _cudnn_root)
file(GLOB MLX_CUDNN_DLLS
"${_cudnn_root}/bin/${CUDAToolkit_VERSION_MAJOR}.0/cudnn*.dll"
"${_cudnn_root}/bin/x64/cudnn*.dll")
list(APPEND MLX_CUDA_LIBS ${MLX_CUDNN_DLLS})
endif()
if(MLX_CUDA_LIBS)
install(FILES ${MLX_CUDA_LIBS}
DESTINATION ${OLLAMA_INSTALL_DIR}
+3 -1
View File
@@ -17,6 +17,8 @@
"inherits": [ "default" ],
"cacheVariables": {
"CMAKE_CUDA_FLAGS": "-t 2",
"MLX_BUILD_CUDA": "ON",
"MLX_BUILD_METAL": "OFF",
"OLLAMA_RUNNER_DIR": "mlx_cuda_v13"
}
},
@@ -55,7 +57,7 @@
"inherits": [ "default" ],
"binaryDir": "${sourceDir}/../../build/metal-v4",
"cacheVariables": {
"CMAKE_OSX_DEPLOYMENT_TARGET": "26.0",
"CMAKE_OSX_DEPLOYMENT_TARGET": "26.2",
"OLLAMA_RUNNER_DIR": "mlx_metal_v4"
}
}
-13
View File
@@ -1,13 +0,0 @@
//go:build !windows
package cmd
import "syscall"
// backgroundServerSysProcAttr returns SysProcAttr for running the server in the background on Unix.
// Setpgid prevents the server from being killed when the parent process exits.
func backgroundServerSysProcAttr() *syscall.SysProcAttr {
return &syscall.SysProcAttr{
Setpgid: true,
}
}
-12
View File
@@ -1,12 +0,0 @@
package cmd
import "syscall"
// backgroundServerSysProcAttr returns SysProcAttr for running the server in the background on Windows.
// CREATE_NO_WINDOW (0x08000000) prevents a console window from appearing.
func backgroundServerSysProcAttr() *syscall.SysProcAttr {
return &syscall.SysProcAttr{
CreationFlags: 0x08000000,
HideWindow: true,
}
}
+121
View File
@@ -0,0 +1,121 @@
package cmd
import (
"context"
"fmt"
"os"
"strings"
"golang.org/x/term"
"github.com/ollama/ollama/api"
"github.com/ollama/ollama/cmd/launch"
"github.com/ollama/ollama/internal/modelref"
"github.com/ollama/ollama/types/model"
)
// for testing
var (
isInteractiveTerminal = func() bool {
return term.IsTerminal(int(os.Stdin.Fd())) && term.IsTerminal(int(os.Stdout.Fd()))
}
confirmCloudSuggestion = func(prompt string) (bool, error) {
// Zero-value options default to Yes being preselected.
return launch.ConfirmPromptWithOptions(prompt, launch.ConfirmOptions{})
}
)
// pullModelNotFoundMessage is how a registry 404 during pull surfaces to
// clients: os.ErrNotExist wrapped server-side and flattened into the error
// string of the pull stream.
const pullModelNotFoundMessage = "pull model manifest: file does not exist"
// isPullNotFoundErr reports whether err is a pull failure caused by the
// requested model or tag not existing in the registry.
func isPullNotFoundErr(err error) bool {
return err != nil && strings.Contains(err.Error(), pullModelNotFoundMessage)
}
// cloudSuggestionCandidate reports whether a failed pull of name should
// trigger a ":cloud" suggestion, and if so returns the cloud model name to
// suggest. It only applies to default-tag lookups (e.g. "kimi-k3") against
// the default registry whose pull failed because the tag doesn't exist.
func cloudSuggestionCandidate(name string, pullErr error, insecure bool) (string, bool) {
if !isPullNotFoundErr(pullErr) {
return "", false
}
return cloudSuggestionName(name, insecure)
}
// cloudSuggestionName applies the name-based eligibility checks for the
// ":cloud" suggestion, returning the cloud model name to suggest.
func cloudSuggestionName(name string, insecure bool) (string, bool) {
// --insecure implies a non-default registry, where an ollama.com cloud
// model wouldn't be a meaningful suggestion.
if insecure {
return "", false
}
ref, err := modelref.ParseRef(name)
if err != nil || ref.Source != modelref.ModelSourceUnspecified {
return "", false
}
if modelref.HasExplicitTag(ref.Base) {
return "", false
}
// Only default-registry names qualify: the existence probe forwards the name
// to ollama.com, and custom-registry model names shouldn't be sent there.
if n := model.ParseName(ref.Base); !n.IsValid() || !strings.EqualFold(n.Host, model.DefaultName().Host) {
return "", false
}
return ref.Base + ":cloud", true
}
// pullWithCloudSuggestion pulls `name`, and if the model's default tag
// doesn't exist but a ":cloud" tag does, offers it: either interactively via
// a confirmation prompt, or by augmenting the returned error when not at a
// terminal. It returns the name that was actually pulled. `verb` is the
// user-facing command ("run" or "pull") used in the hint text.
func pullWithCloudSuggestion(ctx context.Context, client *api.Client, name string, insecure bool, verb string) (string, error) {
// If a suggestion prompt may follow a failed pull, erase the failed
// attempt's progress display instead of leaving its "pulling manifest"
// line to stack up against the accepted pull's identical one.
_, eligible := cloudSuggestionName(name, insecure)
clearNotFound := eligible && isInteractiveTerminal()
pullErr := pullModelWithProgress(ctx, client, name, insecure, clearNotFound)
if pullErr == nil {
return name, nil
}
cloudName, ok := cloudSuggestionCandidate(name, pullErr, insecure)
if !ok || ctx.Err() != nil {
return "", pullErr
}
// Showing a ":cloud" model is proxied to ollama.com and mirrors its status,
// so this reliably answers "does a cloud version exist?". Any error (no
// cloud tag, cloud disabled, older server, offline) means no suggestion.
if _, err := client.Show(ctx, &api.ShowRequest{Model: cloudName}); err != nil {
return "", pullErr
}
if !isInteractiveTerminal() {
return "", fmt.Errorf("%w\n\n%q is available as a cloud model. Try:\n ollama %s %s", pullErr, cloudName, verb, cloudName)
}
accepted, err := confirmCloudSuggestion(fmt.Sprintf("Did you mean %q?", cloudName))
if err != nil || !accepted {
// Declining or cancelling falls back to the original error.
return "", pullErr
}
if err := pullModelWithProgress(ctx, client, cloudName, insecure, false); err != nil {
return "", err
}
return cloudName, nil
}
+411
View File
@@ -0,0 +1,411 @@
package cmd
import (
"cmp"
"encoding/json"
"errors"
"net/http"
"net/http/httptest"
"slices"
"strings"
"testing"
"github.com/spf13/cobra"
"github.com/ollama/ollama/api"
"github.com/ollama/ollama/cmd/launch"
"github.com/ollama/ollama/types/model"
)
func TestCloudSuggestionCandidate(t *testing.T) {
notFoundErr := errors.New("pull model manifest: file does not exist")
suggestedErr := errors.New("pull model manifest: file does not exist\n\nTry one of these models:\n some-model:cloud")
tests := []struct {
name string
model string
pullErr error
insecure bool
want string
wantOK bool
}{
{name: "default tag not found", model: "some-model", pullErr: notFoundErr, want: "some-model:cloud", wantOK: true},
{name: "composes with server tag suggestions", model: "some-model", pullErr: suggestedErr, want: "some-model:cloud", wantOK: true},
{name: "namespaced default tag", model: "user/some-model", pullErr: notFoundErr, want: "user/some-model:cloud", wantOK: true},
{name: "nil error", model: "some-model", pullErr: nil},
{name: "unrelated error", model: "some-model", pullErr: errors.New("boom")},
{name: "insecure registry", model: "some-model", pullErr: notFoundErr, insecure: true},
{name: "explicit tag", model: "some-model:9b", pullErr: notFoundErr},
{name: "explicit latest tag", model: "some-model:latest", pullErr: notFoundErr},
{name: "explicit cloud source", model: "some-model:cloud", pullErr: notFoundErr},
{name: "explicit legacy cloud tag", model: "some-model:9b-cloud", pullErr: notFoundErr},
{name: "explicit local source", model: "some-model:local", pullErr: notFoundErr},
{name: "custom registry host", model: "internal.example.com/team/private-model", pullErr: notFoundErr},
{name: "custom registry host with port", model: "registry.example.com:5000/team/private-model", pullErr: notFoundErr},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
got, ok := cloudSuggestionCandidate(tt.model, tt.pullErr, tt.insecure)
if ok != tt.wantOK {
t.Fatalf("cloudSuggestionCandidate(%q) ok = %v, want %v", tt.model, ok, tt.wantOK)
}
if got != tt.want {
t.Fatalf("cloudSuggestionCandidate(%q) = %q, want %q", tt.model, got, tt.want)
}
})
}
}
// stubCloudSuggest replaces the TTY check and confirmation prompt for the
// duration of the test. If confirm is nil, any prompt fails the test.
func stubCloudSuggest(t *testing.T, interactive bool, confirm func(prompt string) (bool, error)) *[]string {
t.Helper()
oldTTY, oldConfirm := isInteractiveTerminal, confirmCloudSuggestion
t.Cleanup(func() {
isInteractiveTerminal, confirmCloudSuggestion = oldTTY, oldConfirm
})
isInteractiveTerminal = func() bool { return interactive }
prompts := &[]string{}
confirmCloudSuggestion = func(prompt string) (bool, error) {
*prompts = append(*prompts, prompt)
if confirm == nil {
t.Errorf("unexpected cloud suggestion prompt: %q", prompt)
return false, nil
}
return confirm(prompt)
}
return prompts
}
type cloudSuggestServer struct {
cloudName string // model name whose show/pull succeeds (e.g. "some-model:cloud")
cloudExists bool // whether showing/pulling cloudName succeeds
pullErr string // error message for failing pulls
showModels []string
pullModels []string
generateModels []string
}
// start serves mock /api/show, /api/pull, /api/tags, and /api/generate
// endpoints: only cloudName is known (when cloudExists), and pulling any other
// model fails with pullErr streamed the way real servers do (an in-band error
// under HTTP 200).
func (s *cloudSuggestServer) start(t *testing.T) {
t.Helper()
mockServer := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
switch {
case r.URL.Path == "/api/show" && r.Method == http.MethodPost:
var req api.ShowRequest
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
name := cmp.Or(req.Model, req.Name)
s.showModels = append(s.showModels, name)
if s.cloudExists && name == s.cloudName {
if err := json.NewEncoder(w).Encode(api.ShowResponse{
Capabilities: []model.Capability{model.CapabilityCompletion},
RemoteModel: strings.TrimSuffix(s.cloudName, ":cloud"),
}); err != nil {
http.Error(w, err.Error(), http.StatusInternalServerError)
}
return
}
w.WriteHeader(http.StatusNotFound)
if err := json.NewEncoder(w).Encode(map[string]string{
"error": "model '" + name + "' not found",
}); err != nil {
http.Error(w, err.Error(), http.StatusInternalServerError)
}
case r.URL.Path == "/api/pull" && r.Method == http.MethodPost:
var req api.PullRequest
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
name := cmp.Or(req.Model, req.Name)
s.pullModels = append(s.pullModels, name)
var body any
if s.cloudExists && name == s.cloudName {
body = api.ProgressResponse{Status: "success"}
} else {
body = map[string]string{"error": s.pullErr}
}
if err := json.NewEncoder(w).Encode(body); err != nil {
http.Error(w, err.Error(), http.StatusInternalServerError)
}
case r.URL.Path == "/api/tags" && r.Method == http.MethodGet:
if err := json.NewEncoder(w).Encode(api.ListResponse{
Models: []api.ListModelResponse{{Name: s.cloudName}},
}); err != nil {
http.Error(w, err.Error(), http.StatusInternalServerError)
}
case r.URL.Path == "/api/generate" && r.Method == http.MethodPost:
var req api.GenerateRequest
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
http.Error(w, err.Error(), http.StatusBadRequest)
return
}
s.generateModels = append(s.generateModels, req.Model)
if err := json.NewEncoder(w).Encode(api.GenerateResponse{Done: true}); err != nil {
http.Error(w, err.Error(), http.StatusInternalServerError)
}
default:
http.NotFound(w, r)
}
}))
t.Setenv("OLLAMA_HOST", mockServer.URL)
t.Cleanup(mockServer.Close)
}
func newCloudSuggestServer(t *testing.T) *cloudSuggestServer {
t.Helper()
s := &cloudSuggestServer{
cloudName: "some-model:cloud",
cloudExists: true,
pullErr: "pull model manifest: file does not exist",
}
s.start(t)
return s
}
func newPullTestCmd(t *testing.T) *cobra.Command {
t.Helper()
cmd := &cobra.Command{}
cmd.SetContext(t.Context())
cmd.Flags().Bool("insecure", false, "")
return cmd
}
func newRunTestCmd(t *testing.T) *cobra.Command {
t.Helper()
cmd := &cobra.Command{}
cmd.SetContext(t.Context())
cmd.Flags().String("keepalive", "", "")
cmd.Flags().Bool("truncate", false, "")
cmd.Flags().Int("dimensions", 0, "")
cmd.Flags().Bool("verbose", false, "")
cmd.Flags().Bool("insecure", false, "")
cmd.Flags().Bool("nowordwrap", false, "")
cmd.Flags().String("format", "", "")
cmd.Flags().String("think", "", "")
cmd.Flags().Bool("hidethinking", false, "")
return cmd
}
func TestPullHandler_SuccessfulPullNoSuggestion(t *testing.T) {
server := newCloudSuggestServer(t)
server.cloudName = "some-model" // the requested model itself pulls fine
stubCloudSuggest(t, true, nil)
if err := PullHandler(newPullTestCmd(t), []string{"some-model"}); err != nil {
t.Fatalf("PullHandler returned error: %v", err)
}
if want := []string{"some-model"}; !slices.Equal(server.pullModels, want) {
t.Fatalf("pulled models = %v, want %v", server.pullModels, want)
}
if len(server.showModels) != 0 {
t.Fatalf("show models = %v, want no probe after a successful pull", server.showModels)
}
}
func TestPullHandler_CloudSuggestionAccepted(t *testing.T) {
server := newCloudSuggestServer(t)
prompts := stubCloudSuggest(t, true, func(string) (bool, error) { return true, nil })
if err := PullHandler(newPullTestCmd(t), []string{"some-model"}); err != nil {
t.Fatalf("PullHandler returned error: %v", err)
}
if want := []string{"some-model", "some-model:cloud"}; !slices.Equal(server.pullModels, want) {
t.Fatalf("pulled models = %v, want %v", server.pullModels, want)
}
if len(*prompts) != 1 || !strings.Contains((*prompts)[0], `"some-model:cloud"`) {
t.Fatalf("prompts = %v, want one prompt mentioning some-model:cloud", *prompts)
}
}
func TestPullHandler_CloudSuggestionDeclined(t *testing.T) {
server := newCloudSuggestServer(t)
stubCloudSuggest(t, true, func(string) (bool, error) { return false, nil })
err := PullHandler(newPullTestCmd(t), []string{"some-model"})
if err == nil {
t.Fatal("PullHandler returned nil, want an error")
}
if !strings.Contains(err.Error(), "pull model manifest: file does not exist") {
t.Fatalf("error = %q, want it to contain the original pull error", err)
}
if strings.Contains(err.Error(), "Try:") {
t.Fatalf("error = %q, want no non-interactive hint after declining", err)
}
if want := []string{"some-model"}; !slices.Equal(server.pullModels, want) {
t.Fatalf("pulled models = %v, want %v", server.pullModels, want)
}
}
func TestPullHandler_CloudSuggestionCancelled(t *testing.T) {
server := newCloudSuggestServer(t)
stubCloudSuggest(t, true, func(string) (bool, error) { return false, launch.ErrCancelled })
err := PullHandler(newPullTestCmd(t), []string{"some-model"})
if err == nil {
t.Fatal("PullHandler returned nil, want an error")
}
if errors.Is(err, launch.ErrCancelled) {
t.Fatalf("error = %v, want the original pull error rather than ErrCancelled", err)
}
if !strings.Contains(err.Error(), "pull model manifest: file does not exist") {
t.Fatalf("error = %q, want it to contain the original pull error", err)
}
if want := []string{"some-model"}; !slices.Equal(server.pullModels, want) {
t.Fatalf("pulled models = %v, want %v", server.pullModels, want)
}
}
func TestPullHandler_CloudSuggestionNonInteractive(t *testing.T) {
server := newCloudSuggestServer(t)
stubCloudSuggest(t, false, nil)
err := PullHandler(newPullTestCmd(t), []string{"some-model"})
if err == nil {
t.Fatal("PullHandler returned nil, want an error")
}
if !strings.Contains(err.Error(), "pull model manifest: file does not exist") {
t.Fatalf("error = %q, want it to contain the original pull error", err)
}
if !strings.Contains(err.Error(), "ollama pull some-model:cloud") {
t.Fatalf("error = %q, want it to hint at 'ollama pull some-model:cloud'", err)
}
if want := []string{"some-model"}; !slices.Equal(server.pullModels, want) {
t.Fatalf("pulled models = %v, want %v", server.pullModels, want)
}
}
func TestPullHandler_CloudSuggestionNoCloudTag(t *testing.T) {
server := newCloudSuggestServer(t)
server.cloudExists = false
stubCloudSuggest(t, true, nil)
err := PullHandler(newPullTestCmd(t), []string{"some-model"})
if err == nil || err.Error() != "pull model manifest: file does not exist" {
t.Fatalf("error = %v, want the unmodified pull error", err)
}
if want := []string{"some-model:cloud"}; !slices.Equal(server.showModels, want) {
t.Fatalf("show models = %v, want the cloud existence probe %v", server.showModels, want)
}
}
func TestPullHandler_CloudSuggestionExplicitTag(t *testing.T) {
server := newCloudSuggestServer(t)
stubCloudSuggest(t, true, nil)
err := PullHandler(newPullTestCmd(t), []string{"some-model:9b"})
if err == nil || err.Error() != "pull model manifest: file does not exist" {
t.Fatalf("error = %v, want the unmodified pull error", err)
}
if len(server.showModels) != 0 {
t.Fatalf("show models = %v, want no cloud probe for explicitly tagged models", server.showModels)
}
}
func TestPullHandler_CloudSuggestionExplicitCloud(t *testing.T) {
server := newCloudSuggestServer(t)
server.cloudExists = false // make the explicit :cloud pull fail too
stubCloudSuggest(t, true, nil)
err := PullHandler(newPullTestCmd(t), []string{"some-model:cloud"})
if err == nil || err.Error() != "pull model manifest: file does not exist" {
t.Fatalf("error = %v, want the unmodified pull error", err)
}
if len(server.showModels) != 0 {
t.Fatalf("show models = %v, want no probe for explicit :cloud requests", server.showModels)
}
}
func TestPullHandler_CloudSuggestionInsecure(t *testing.T) {
server := newCloudSuggestServer(t)
stubCloudSuggest(t, true, nil)
cmd := newPullTestCmd(t)
if err := cmd.Flags().Set("insecure", "true"); err != nil {
t.Fatal(err)
}
err := PullHandler(cmd, []string{"some-model"})
if err == nil || err.Error() != "pull model manifest: file does not exist" {
t.Fatalf("error = %v, want the unmodified pull error", err)
}
if len(server.showModels) != 0 {
t.Fatalf("show models = %v, want no probe for --insecure pulls", server.showModels)
}
}
func TestPullHandler_CloudSuggestionUnrelatedError(t *testing.T) {
server := newCloudSuggestServer(t)
server.pullErr = "boom"
stubCloudSuggest(t, true, nil)
err := PullHandler(newPullTestCmd(t), []string{"some-model"})
if err == nil || err.Error() != "boom" {
t.Fatalf("error = %v, want the unmodified pull error %q", err, "boom")
}
if len(server.showModels) != 0 {
t.Fatalf("show models = %v, want no probe for unrelated pull errors", server.showModels)
}
}
func TestRunHandler_CloudSuggestionAccepted_RunsCloudModel(t *testing.T) {
server := newCloudSuggestServer(t)
stubCloudSuggest(t, true, func(string) (bool, error) { return true, nil })
if err := RunHandler(newRunTestCmd(t), []string{"some-model", "hi"}); err != nil {
t.Fatalf("RunHandler returned error: %v", err)
}
if want := []string{"some-model", "some-model:cloud"}; !slices.Equal(server.pullModels, want) {
t.Fatalf("pulled models = %v, want %v", server.pullModels, want)
}
if want := []string{"some-model:cloud"}; !slices.Equal(server.generateModels, want) {
t.Fatalf("generate models = %v, want %v", server.generateModels, want)
}
}
func TestRunHandler_CloudSuggestionDeclined_ReturnsNotFound(t *testing.T) {
server := newCloudSuggestServer(t)
stubCloudSuggest(t, true, func(string) (bool, error) { return false, nil })
err := RunHandler(newRunTestCmd(t), []string{"some-model", "hi"})
if err == nil {
t.Fatal("RunHandler returned nil, want an error")
}
if !strings.Contains(err.Error(), "pull model manifest: file does not exist") {
t.Fatalf("error = %q, want it to contain the original pull error", err)
}
if len(server.generateModels) != 0 {
t.Fatalf("generate models = %v, want none after declining", server.generateModels)
}
}
func TestRunHandler_CloudSuggestionNonInteractive_Hint(t *testing.T) {
server := newCloudSuggestServer(t)
stubCloudSuggest(t, false, nil)
err := RunHandler(newRunTestCmd(t), []string{"some-model", "hi"})
if err == nil {
t.Fatal("RunHandler returned nil, want an error")
}
if !strings.Contains(err.Error(), "ollama run some-model:cloud") {
t.Fatalf("error = %q, want it to hint at 'ollama run some-model:cloud'", err)
}
if len(server.generateModels) != 0 {
t.Fatalf("generate models = %v, want none in non-interactive mode", server.generateModels)
}
}
+69 -114
View File
@@ -16,7 +16,6 @@ import (
"net"
"net/http"
"os"
"os/exec"
"os/signal"
"path"
"path/filepath"
@@ -55,10 +54,8 @@ import (
"github.com/ollama/ollama/types/model"
"github.com/ollama/ollama/types/syncmap"
"github.com/ollama/ollama/version"
xcmd "github.com/ollama/ollama/x/cmd"
xcreate "github.com/ollama/ollama/x/create"
xcreateclient "github.com/ollama/ollama/x/create/client"
"github.com/ollama/ollama/x/imagegen"
)
func init() {
@@ -96,6 +93,8 @@ func init() {
}
launch.DefaultConfirmPrompt = tui.RunConfirmWithOptions
launch.DefaultSpinner = tui.RunSpinner
}
func runTUISingleSelector(title string, items []launch.SelectionItem, current string, updates <-chan []launch.SelectionItem) (string, error) {
@@ -192,7 +191,7 @@ func resolveExperimentalLocalModelDir(ref, filename string) string {
}
candidate := filepath.Join(filepath.Dir(filename), ref)
if xcreate.IsSafetensorsModelDir(candidate) || xcreate.IsTensorModelDir(candidate) {
if xcreate.IsSafetensorsModelDir(candidate) {
return candidate
}
@@ -230,8 +229,7 @@ func CreateHandler(cmd *cobra.Command, args []string) error {
return fmt.Errorf("invalid model name: %s", modelName)
}
// Check for --experimental flag for safetensors model creation
// This gates both safetensors LLM and imagegen model creation
// Check for --experimental flag for safetensors model creation.
experimental, _ := cmd.Flags().GetBool("experimental")
draftQuantize, _ := cmd.Flags().GetString("draft-quantize")
if experimental {
@@ -709,6 +707,32 @@ func hasListedModelName(models []api.ListModelResponse, name string) bool {
return false
}
// showOrPullModel returns model info for name, pulling the model if it isn't
// available locally. If the pull finds no default tag but a ":cloud" tag
// exists, the user may be offered the cloud model instead (see
// pullWithCloudSuggestion), in which case the returned name is the cloud
// name the caller should continue with. verb is the user-facing command
// ("run" or "pull") used in hint text.
func showOrPullModel(cmd *cobra.Command, client *api.Client, name string, insecure bool, verb string) (*api.ShowResponse, string, error) {
info, err := client.Show(cmd.Context(), &api.ShowRequest{Model: name})
if err == nil {
return info, name, nil
}
var se api.StatusError
if !errors.As(err, &se) || se.StatusCode != http.StatusNotFound || modelref.HasExplicitCloudSource(name) {
return nil, name, err
}
resolved, err := pullWithCloudSuggestion(cmd.Context(), client, name, insecure, verb)
if err != nil {
return nil, name, err
}
info, err = client.Show(cmd.Context(), &api.ShowRequest{Model: resolved})
return info, resolved, err
}
func RunHandler(cmd *cobra.Command, args []string) error {
interactive := true
@@ -804,30 +828,21 @@ func RunHandler(cmd *cobra.Command, args []string) error {
return err
}
name := args[0]
requestedCloud := modelref.HasExplicitCloudSource(name)
insecure, err := cmd.Flags().GetBool("insecure")
if err != nil {
return err
}
info, err := func() (*api.ShowResponse, error) {
showReq := &api.ShowRequest{Name: name}
info, err := client.Show(cmd.Context(), showReq)
var se api.StatusError
if errors.As(err, &se) && se.StatusCode == http.StatusNotFound {
if requestedCloud {
return nil, err
}
if err := PullHandler(cmd, []string{name}); err != nil {
return nil, err
}
return client.Show(cmd.Context(), &api.ShowRequest{Name: name})
}
return info, err
}()
info, name, err := showOrPullModel(cmd, client, args[0], insecure, "run")
if err != nil {
if handleCloudAuthorizationError(err) {
return nil
}
return err
}
// The model may have been resolved to a different name (e.g. its ":cloud"
// variant), so make sure downstream requests use it.
opts.Model = name
ensureCloudStub(cmd.Context(), client, name)
@@ -877,19 +892,10 @@ func RunHandler(cmd *cobra.Command, args []string) error {
return generateEmbedding(cmd, name, opts.Prompt, opts.KeepAlive, truncate, dimensions)
}
// Check if this is an image generation model
if slices.Contains(info.Capabilities, model.CapabilityImage) {
if opts.Prompt == "" && !interactive {
return errors.New("image generation models require a prompt. Usage: ollama run " + name + " \"your prompt here\"")
}
return imagegen.RunCLI(cmd, name, opts.Prompt, interactive, opts.KeepAlive)
return errors.New("image generation models are not currently supported")
}
// Check for experimental flag
isExperimental, _ := cmd.Flags().GetBool("experimental")
yoloMode, _ := cmd.Flags().GetBool("experimental-yolo")
enableWebsearch, _ := cmd.Flags().GetBool("experimental-websearch")
if interactive {
if err := loadOrUnloadModel(cmd, &opts); err != nil {
var sErr api.AuthorizationError
@@ -916,11 +922,6 @@ func RunHandler(cmd *cobra.Command, args []string) error {
}
}
// Use experimental agent loop with tools
if isExperimental {
return xcmd.GenerateInteractive(cmd, opts.Model, opts.WordWrap, opts.Options, opts.Think, opts.HideThinking, opts.KeepAlive, yoloMode, enableWebsearch)
}
return generateInteractive(cmd, opts)
}
if err := generate(cmd, opts); err != nil {
@@ -1257,6 +1258,10 @@ func ShowHandler(cmd *cobra.Command, args []string) error {
return err
}
if slices.Contains(resp.Capabilities, model.CapabilityImage) {
return errors.New("image generation models are not currently supported")
}
if flagsSet == 1 {
switch showType {
case "license":
@@ -1518,6 +1523,15 @@ func PullHandler(cmd *cobra.Command, args []string) error {
return err
}
_, err = pullWithCloudSuggestion(cmd.Context(), client, args[0], insecure, "pull")
return err
}
// pullModelWithProgress pulls name, rendering progress to stderr. When
// clearNotFound is set and the pull fails because the model doesn't exist,
// the progress display is erased rather than left behind; callers set it
// when a ":cloud" suggestion prompt may immediately follow the failure.
func pullModelWithProgress(ctx context.Context, client *api.Client, name string, insecure, clearNotFound bool) error {
p := progress.NewProgress(os.Stderr)
defer p.Stop()
@@ -1578,8 +1592,13 @@ func PullHandler(cmd *cobra.Command, args []string) error {
return nil
}
request := api.PullRequest{Name: args[0], Insecure: insecure}
return client.Pull(cmd.Context(), &request, fn)
request := api.PullRequest{Name: name, Insecure: insecure}
err := client.Pull(ctx, &request, fn)
if clearNotFound && isPullNotFoundErr(err) {
// The deferred Stop becomes a no-op after this.
p.StopAndClear()
}
return err
}
type generateContextKey string
@@ -2066,7 +2085,7 @@ func checkServerHeartbeat(cmd *cobra.Command, _ []string) error {
if !(strings.Contains(err.Error(), " refused") || strings.Contains(err.Error(), "could not connect")) {
return err
}
if err := startApp(cmd.Context(), client); err != nil {
if err := startApp(cmd.Context(), client); err != nil { //nolint:staticcheck,nolintlint // startApp always returns non-nil on Linux (start_default.go) but can return nil on macOS/Windows
return err
}
}
@@ -2108,40 +2127,6 @@ Environment Variables:
cmd.SetUsageTemplate(cmd.UsageTemplate() + envUsage)
}
// ensureServerRunning checks if the ollama server is running and starts it in the background if not.
func ensureServerRunning(ctx context.Context) error {
client, err := api.ClientFromEnvironment()
if err != nil {
return err
}
// Check if server is already running
if err := client.Heartbeat(ctx); err == nil {
return nil // server is already running
}
// Server not running, start it in the background
exe, err := os.Executable()
if err != nil {
return fmt.Errorf("could not find executable: %w", err)
}
serverCmd := exec.CommandContext(ctx, exe, "serve")
serverCmd.Env = os.Environ()
serverCmd.SysProcAttr = backgroundServerSysProcAttr()
if err := serverCmd.Start(); err != nil {
return fmt.Errorf("failed to start server: %w", err)
}
// Wait for the server to be ready
for {
time.Sleep(500 * time.Millisecond)
if err := client.Heartbeat(ctx); err == nil {
return nil // server has started
}
}
}
func launchInteractiveModel(cmd *cobra.Command, modelName string) error {
opts := runOptions{
Model: modelName,
@@ -2155,31 +2140,15 @@ func launchInteractiveModel(cmd *cobra.Command, modelName string) error {
return err
}
requestedCloud := modelref.HasExplicitCloudSource(modelName)
info, err := func() (*api.ShowResponse, error) {
showReq := &api.ShowRequest{Name: modelName}
info, err := client.Show(cmd.Context(), showReq)
var se api.StatusError
if errors.As(err, &se) && se.StatusCode == http.StatusNotFound {
if requestedCloud {
return nil, err
}
if err := PullHandler(cmd, []string{modelName}); err != nil {
return nil, err
}
return client.Show(cmd.Context(), &api.ShowRequest{Name: modelName})
}
return info, err
}()
info, resolvedModel, err := showOrPullModel(cmd, client, modelName, false, "run")
if err != nil {
if handleCloudAuthorizationError(err) {
return nil
}
return err
}
ensureCloudStub(cmd.Context(), client, modelName)
opts.Model = resolvedModel
ensureCloudStub(cmd.Context(), client, opts.Model)
opts.Think, err = inferThinkingOption(&info.Capabilities, &opts, false)
if err != nil {
@@ -2188,15 +2157,11 @@ func launchInteractiveModel(cmd *cobra.Command, modelName string) error {
audioCapable := slices.Contains(info.Capabilities, model.CapabilityAudio)
opts.MultiModal = slices.Contains(info.Capabilities, model.CapabilityVision) || audioCapable
// TODO: remove the projector info and vision info checks below,
// these are left in for backwards compatibility with older servers
// that don't have the capabilities field in the model info
if len(info.ProjectorInfo) != 0 {
opts.MultiModal = true
}
for k := range info.ModelInfo {
if strings.Contains(k, ".vision.") {
for key := range info.ModelInfo {
if strings.Contains(key, ".vision.") {
opts.MultiModal = true
break
}
@@ -2215,9 +2180,9 @@ func launchInteractiveModel(cmd *cobra.Command, modelName string) error {
// runInteractiveTUI runs the main interactive TUI menu.
func runInteractiveTUI(cmd *cobra.Command) {
// Ensure the server is running before showing the TUI
if err := ensureServerRunning(cmd.Context()); err != nil {
fmt.Fprintf(os.Stderr, "Error starting server: %v\n", err)
// Ensure the server is running via the shared checkServerHeartbeat path.
if err := checkServerHeartbeat(cmd, nil); err != nil {
fmt.Fprintf(os.Stderr, "Error: %v\n", err)
return
}
@@ -2324,7 +2289,7 @@ func runLauncherAction(cmd *cobra.Command, action tui.TUIAction, deps launcherDe
func launcherActionExitsLoop(integration string) bool {
switch integration {
case "codex-app", "vscode":
case "chatgpt", "codex-app", "vscode":
return true
default:
return false
@@ -2413,15 +2378,6 @@ func NewCLI() *cobra.Command {
runCmd.Flags().Bool("hidethinking", false, "Hide thinking output (if provided)")
runCmd.Flags().Bool("truncate", false, "For embedding models: truncate inputs exceeding context length (default: true). Set --truncate=false to error instead")
runCmd.Flags().Int("dimensions", 0, "Truncate output embeddings to specified dimension (embedding models only)")
runCmd.Flags().Bool("experimental", false, "Enable experimental agent loop with tools")
runCmd.Flags().Bool("experimental-yolo", false, "Skip all tool approval prompts (use with caution)")
runCmd.Flags().Bool("experimental-websearch", false, "Enable web search tool in experimental mode")
// Image generation flags (width, height, steps, seed, etc.)
imagegen.RegisterFlags(runCmd)
runCmd.Flags().Bool("imagegen", false, "Use the imagegen runner for LLM inference")
runCmd.Flags().MarkHidden("imagegen")
stopCmd := &cobra.Command{
Use: "stop MODEL",
@@ -2564,7 +2520,6 @@ func NewCLI() *cobra.Command {
} {
switch cmd {
case runCmd:
imagegen.AppendFlagsDocs(cmd)
appendEnvDocs(cmd, []envconfig.EnvVar{envVars["OLLAMA_EDITOR"], envVars["OLLAMA_HOST"], envVars["OLLAMA_NOHISTORY"]})
case serveCmd:
appendEnvDocs(cmd, []envconfig.EnvVar{
+1 -1
View File
@@ -249,7 +249,7 @@ func TestRunLauncherAction_GUIAppsExitTUILoop(t *testing.T) {
cmd := &cobra.Command{}
cmd.SetContext(context.Background())
for _, integration := range []string{"codex-app", "vscode"} {
for _, integration := range []string{"chatgpt", "vscode"} {
continueLoop, err := runLauncherAction(cmd, tui.TUIAction{Kind: tui.TUIActionLaunchIntegration, Integration: integration}, launcherDeps{
resolveRunModel: unexpectedRunModelResolution(t),
launchIntegration: func(ctx context.Context, req launch.IntegrationLaunchRequest) error {
+42 -1
View File
@@ -2079,7 +2079,7 @@ func TestRunOptions_Copy_ThinkValueVariants(t *testing.T) {
}
}
func TestShowInfoImageGen(t *testing.T) {
func TestShowInfoImageCapability(t *testing.T) {
var b bytes.Buffer
err := showInfo(&api.ShowResponse{
Details: api.ModelDetails{
@@ -2390,3 +2390,44 @@ func TestIsLocalhost(t *testing.T) {
})
}
}
func TestRunCommandHasNoAgentFlags(t *testing.T) {
root := NewCLI()
run, _, err := root.Find([]string{"run"})
if err != nil {
t.Fatal(err)
}
for _, name := range []string{"resume", "headless", "auto-approve-tools", "skill", "experimental", "experimental-yolo", "experimental-websearch"} {
if flag := run.Flags().Lookup(name); flag != nil {
t.Errorf("run command still exposes former agent flag --%s", name)
}
}
}
func TestFormerAgentEntryPointsAreRejected(t *testing.T) {
tests := [][]string{
{"run", "llama3", "--resume"},
{"run", "llama3", "--headless"},
{"run", "llama3", "--auto-approve-tools"},
{"run", "llama3", "--skill", "release-notes"},
{"run", "llama3", "--experimental"},
{"run", "llama3", "--experimental-yolo"},
{"run", "llama3", "--experimental-websearch"},
{"agent"},
}
for _, args := range tests {
t.Run(strings.Join(args, " "), func(t *testing.T) {
root := NewCLI()
root.SetArgs(args)
err := root.Execute()
if err == nil {
t.Fatalf("former agent entry point %q succeeded", args)
}
if !strings.Contains(err.Error(), "unknown") {
t.Fatalf("former agent entry point %q returned %v, want unknown command or flag", args, err)
}
})
}
}
+6 -4
View File
@@ -653,9 +653,12 @@ func editInExternalEditor(content string) (string, error) {
}
// Check that the editor binary exists
name := strings.Fields(editor)[0]
if _, err := exec.LookPath(name); err != nil {
return "", fmt.Errorf("editor %q not found, set OLLAMA_EDITOR to the path of your preferred editor", name)
args := strings.Fields(editor)
if len(args) == 0 {
return "", fmt.Errorf("no editor configured, set OLLAMA_EDITOR to the path of your preferred editor")
}
if _, err := exec.LookPath(args[0]); err != nil {
return "", fmt.Errorf("editor %q not found, set OLLAMA_EDITOR to the path of your preferred editor", args[0])
}
tmpFile, err := os.CreateTemp("", "ollama-prompt-*.txt")
@@ -672,7 +675,6 @@ func editInExternalEditor(content string) (string, error) {
}
tmpFile.Close()
args := strings.Fields(editor)
args = append(args, tmpFile.Name())
cmd := exec.Command(args[0], args[1:]...)
cmd.Stdin = os.Stdin
+33
View File
@@ -85,6 +85,39 @@ func TestExtractFileDataRemovesQuotedFilepath(t *testing.T) {
assert.Equal(t, cleaned, "before after")
}
func TestEditInExternalEditorWhitespaceOnly(t *testing.T) {
// A whitespace-only VISUAL or EDITOR is non-empty, so it bypasses the
// editor == "" fallbacks, but strings.Fields collapses it to an empty
// slice. Indexing that slice must not panic; it must return an error.
cases := []struct {
name string
visual string
editor string
}{
{name: "VISUAL whitespace", visual: "\t "},
{name: "EDITOR whitespace", editor: "\t "},
}
for _, tt := range cases {
t.Run(tt.name, func(t *testing.T) {
t.Setenv("OLLAMA_EDITOR", "")
t.Setenv("VISUAL", tt.visual)
t.Setenv("EDITOR", tt.editor)
_, err := editInExternalEditor("content")
assert.Error(t, err)
})
}
}
func TestEditInExternalEditorParsesEditorWithArgs(t *testing.T) {
// A well-formed editor command with arguments must still be parsed so
// its binary is looked up (guards the normal path from regressing).
t.Setenv("OLLAMA_EDITOR", "definitely-not-a-real-editor arg1")
t.Setenv("VISUAL", "")
t.Setenv("EDITOR", "")
_, err := editInExternalEditor("content")
assert.ErrorContains(t, err, "definitely-not-a-real-editor")
}
func TestExtractFileDataWAV(t *testing.T) {
dir := t.TempDir()
fp := filepath.Join(dir, "sample.wav")
+3 -3
View File
@@ -20,7 +20,7 @@ const (
)
var (
ErrPlanVerificationUnavailable = errors.New("Could not verify your plan. Try again in a moment.")
ErrPlanVerificationUnavailable = errors.New("Could not verify Ollama plan. Try again in a moment or use a local model.")
errUpgradeCancelled = errors.New("upgrade cancelled")
)
@@ -247,7 +247,7 @@ func (c *launcherClient) ensureCloudModelAccess(ctx context.Context, model strin
c.accountState = &state
}
if state.Status == accountStateUnknown {
return ErrPlanVerificationUnavailable
return nil
}
if state.Status == accountStateSignedOut {
@@ -259,7 +259,7 @@ func (c *launcherClient) ensureCloudModelAccess(ctx context.Context, model strin
c.accountState = &state
}
if state.Status == accountStateUnknown {
return ErrPlanVerificationUnavailable
return nil
}
}
+104 -11
View File
@@ -7,6 +7,7 @@ import (
"path/filepath"
"runtime"
"strconv"
"strings"
"github.com/ollama/ollama/envconfig"
)
@@ -37,17 +38,21 @@ func (c *Claude) findPath() (string, error) {
if runtime.GOOS == "windows" {
name = "claude.exe"
}
fallback := filepath.Join(home, ".claude", "local", name)
if _, err := os.Stat(fallback); err != nil {
return "", err
for _, fallback := range []string{
filepath.Join(home, ".local", "bin", name),
filepath.Join(home, ".claude", "local", name),
} {
if _, err := os.Stat(fallback); err == nil {
return fallback, nil
}
}
return fallback, nil
return "", fmt.Errorf("claude binary not found")
}
func (c *Claude) Run(model string, _ []LaunchModel, args []string) error {
claudePath, err := c.findPath()
claudePath, err := ensureClaudeInstalled()
if err != nil {
return fmt.Errorf("claude is not installed, install from https://code.claude.com/docs/en/quickstart")
return err
}
cmd := exec.Command(claudePath, c.args(model, args)...)
@@ -55,17 +60,105 @@ func (c *Claude) Run(model string, _ []LaunchModel, args []string) error {
cmd.Stdout = os.Stdout
cmd.Stderr = os.Stderr
env := append(os.Environ(),
"ANTHROPIC_BASE_URL="+envconfig.Host().String(),
cmd.Env = append(os.Environ(), c.envVars(model)...)
return cmd.Run()
}
func (c *Claude) envVars(model string) []string {
env := []string{
"ANTHROPIC_BASE_URL=" + envconfig.Host().String(),
"ANTHROPIC_API_KEY=",
"ANTHROPIC_AUTH_TOKEN=ollama",
"CLAUDE_CODE_ATTRIBUTION_HEADER=0",
)
"CLAUDE_CODE_TOTAL_TOKENS_REMINDER=off",
"DISABLE_ERROR_REPORTING=1",
"DISABLE_FEEDBACK_COMMAND=1",
"CLAUDE_CODE_DISABLE_FEEDBACK_SURVEY=1",
}
env = append(env, c.modelEnvVars(model)...)
return env
}
cmd.Env = env
return cmd.Run()
func ensureClaudeInstalled() (string, error) {
if path, err := (&Claude{}).findPath(); err == nil {
return path, nil
}
if err := checkClaudeInstallerDependencies(); err != nil {
return "", err
}
ok, err := ConfirmPrompt("Claude Code is not installed. Install now?")
if err != nil {
return "", err
}
if !ok {
return "", fmt.Errorf("claude installation cancelled")
}
bin, args, err := claudeInstallerCommand(runtime.GOOS)
if err != nil {
return "", err
}
fmt.Fprintf(os.Stderr, "\nInstalling Claude Code...\n")
cmd := exec.Command(bin, args...)
cmd.Stdin = os.Stdin
cmd.Stdout = os.Stdout
cmd.Stderr = os.Stderr
if err := cmd.Run(); err != nil {
return "", fmt.Errorf("failed to install claude: %w", err)
}
path, err := (&Claude{}).findPath()
if err != nil {
return "", fmt.Errorf("claude was installed but the binary was not found on PATH\n\nYou may need to restart your shell")
}
fmt.Fprintf(os.Stderr, "%sClaude Code installed successfully%s\n\n", ansiGreen, ansiReset)
return path, nil
}
func checkClaudeInstallerDependencies() error {
switch runtime.GOOS {
case "windows":
if _, err := exec.LookPath("powershell"); err != nil {
return fmt.Errorf("claude is not installed and required dependencies are missing\n\nInstall the following first:\n PowerShell: https://learn.microsoft.com/powershell/\n\nThen re-run:\n ollama launch claude")
}
default:
var missing []string
if _, err := exec.LookPath("curl"); err != nil {
missing = append(missing, "curl: https://curl.se/")
}
if _, err := exec.LookPath("bash"); err != nil {
missing = append(missing, "bash: https://www.gnu.org/software/bash/")
}
if len(missing) > 0 {
return fmt.Errorf("claude is not installed and required dependencies are missing\n\nInstall the following first:\n %s\n\nThen re-run:\n ollama launch claude", strings.Join(missing, "\n "))
}
}
return nil
}
func claudeInstallerCommand(goos string) (string, []string, error) {
switch goos {
case "windows":
return "powershell", []string{
"-NoProfile",
"-ExecutionPolicy",
"Bypass",
"-Command",
"irm https://claude.ai/install.ps1 | iex",
}, nil
case "darwin", "linux":
return "bash", []string{
"-c",
"curl -fsSL https://claude.ai/install.sh | bash",
}, nil
default:
return "", nil, fmt.Errorf("unsupported platform for claude install: %s", goos)
}
}
// modelEnvVars returns Claude Code env vars that route all model tiers through Ollama.
+364 -232
View File
@@ -5,8 +5,7 @@ import (
"encoding/json"
"errors"
"fmt"
"io"
"net/http"
"net/url"
"os"
"os/exec"
"path/filepath"
@@ -16,22 +15,25 @@ import (
"github.com/ollama/ollama/cmd/config"
"github.com/ollama/ollama/cmd/internal/fileutil"
"golang.org/x/term"
"github.com/ollama/ollama/internal/proxy"
)
const (
claudeDesktopIntegrationName = "claude-desktop"
claudeDesktopProfileName = "Ollama"
claudeDesktopProfileID = "00000000-0000-4000-8000-000000000114"
claudeDesktopGatewayBaseURL = "https://ollama.com"
claudeDesktopAPIKeyURL = "https://ollama.com/settings/keys"
claudeDesktopModelLabel = "Ollama Cloud"
claudeDesktopUnsupported = "Claude Desktop is no longer supported. Existing installations can be restored with 'ollama launch claude-desktop --restore'."
claudeDesktopSuccessMessage = "Claude Desktop profile changed to Ollama Cloud."
claudeDesktopGatewayBaseURL = "http://" + proxy.DefaultClaudeDesktopListenAddr
claudeDesktopProbeTimeout = 2 * time.Second
claudeDesktopModelLabel = "Default Ollama model"
claudeDesktopSuccessMessage = "Claude Desktop profile changed to Ollama."
claudeDesktopRestoreMessage = "To restore the usual Claude profile, run: ollama launch claude-desktop --restore"
claudeDesktopRestoredMessage = "Claude Desktop restored to the usual Claude profile."
)
// Cowork needs unrestricted egress for user-configured plugins and MCP servers.
// Restore removes this override with the rest of the Ollama profile settings.
var claudeDesktopEgressHosts = []string{"*"}
var (
claudeDesktopGOOS = runtime.GOOS
claudeDesktopUserHome = os.UserHomeDir
@@ -39,17 +41,14 @@ var (
claudeDesktopOpenApp = defaultClaudeDesktopOpenApp
claudeDesktopOpenAppPath = defaultClaudeDesktopOpenAppPath
claudeDesktopQuitApp = defaultClaudeDesktopQuitApp
claudeDesktopIsRunning = defaultClaudeDesktopIsRunning
claudeDesktopIsRunning = defaultClaudeDesktopRunning
claudeDesktopRunningAppPath = defaultClaudeDesktopRunningAppPath
claudeDesktopGlob = filepath.Glob
claudeDesktopSleep = time.Sleep
claudeDesktopHTTPClient = http.DefaultClient
claudeDesktopPromptAPIKey = promptClaudeDesktopAPIKey
claudeDesktopValidateAPIKey = validateClaudeDesktopAPIKey
claudeDesktopProbeGateway = proxy.ProbeClaudeDesktop
)
// ClaudeDesktop configures and launches Claude Desktop in third-party
// inference mode using Ollama Cloud as the gateway.
// inference mode using the Ollama app's local gateway.
type ClaudeDesktop struct{}
func (c *ClaudeDesktop) String() string { return "Claude Desktop" }
@@ -64,38 +63,21 @@ func (c *ClaudeDesktop) AutodiscoveredModel() string {
return claudeDesktopModelLabel
}
// ConfigureAutodiscovery points Claude Desktop at Ollama's local gateway
// without pinning a model list, so Claude discovers the selected catalog and
// exact Ollama route names the gateway advertises.
func (c *ClaudeDesktop) ConfigureAutodiscovery() error {
if err := claudeDesktopSupported(); err != nil {
return err
}
if err := ensureClaudeDesktopGateway(); err != nil {
return err
}
targets, err := claudeDesktopTargetPaths()
if err != nil {
return err
}
key, err := claudeDesktopValidatedAPIKey(context.Background(), claudeDesktopTargetProfilePaths(targets))
if err != nil {
return err
}
for _, path := range targets.normalConfigs {
if err := writeClaudeDesktopDeploymentMode(path, "3p"); err != nil {
return err
}
}
for _, target := range targets.thirdPartyProfiles {
if err := writeClaudeDesktopDeploymentMode(target.desktopConfig, "3p"); err != nil {
return err
}
if err := writeClaudeDesktopMeta(target.meta, claudeDesktopProfileID, claudeDesktopProfileName); err != nil {
return err
}
if err := writeClaudeDesktopGatewayProfile(target.profile, key, true); err != nil {
return err
}
}
return nil
return configureClaudeDesktopTargets(targets, claudeDesktopGatewayBaseURL, "ollama")
}
func (c *ClaudeDesktop) RestoreHint() string {
@@ -118,10 +100,120 @@ func (c *ClaudeDesktop) AutodiscoveryConfigured() bool {
return claudeDesktopTargetsConfigured(targets)
}
// UsesOllamaGateway reports whether Claude Desktop is currently routed through
// Ollama's local gateway. It intentionally ignores auxiliary profile settings
// so the gateway can keep serving while those settings are repaired.
func (c *ClaudeDesktop) UsesOllamaGateway() bool {
targets, err := claudeDesktopTargetPaths()
if err != nil {
return false
}
return claudeDesktopTargetsUseOllamaGateway(targets)
}
// SetInstalledFromDesktop changes the Claude profile from the native Ollama app.
func (c *ClaudeDesktop) SetInstalledFromDesktop(installed, restart bool) error {
if err := claudeDesktopSupported(); err != nil {
return err
}
applyProfile := restoreClaudeDesktopProfile
if installed {
applyProfile = c.ConfigureAutodiscovery
}
running, err := claudeDesktopIsRunning(context.Background())
if err != nil {
return fmt.Errorf("check whether Claude Desktop is running: %w", err)
}
if !running {
if err := applyProfile(); err != nil {
return err
}
if installed {
return claudeDesktopOpenApp()
}
return nil
}
if !restart {
return errors.New("Claude Desktop restart confirmation is required before changing its profile")
}
return restartClaudeDesktop(applyProfile)
}
// RestartWithProfileChange stops Claude before applying a profile-dependent
// change, then reopens it after the change is complete.
func (c *ClaudeDesktop) RestartWithProfileChange(change func() error) error {
if err := claudeDesktopSupported(); err != nil {
return err
}
running, err := claudeDesktopIsRunning(context.Background())
if err != nil {
return fmt.Errorf("check whether Claude Desktop is running: %w", err)
}
if !running {
if err := change(); err != nil {
return err
}
return claudeDesktopOpenApp()
}
return restartClaudeDesktop(change)
}
// RestoreForShutdown restores Claude's usual profile without reopening the app.
func (c *ClaudeDesktop) RestoreForShutdown(ctx context.Context) error {
if err := claudeDesktopSupported(); err != nil {
return err
}
running, err := claudeDesktopIsRunning(ctx)
if err != nil {
return fmt.Errorf("check whether Claude Desktop is running: %w", err)
}
if !running {
return restoreClaudeDesktopProfile()
}
if err := claudeDesktopQuitApp(ctx); err != nil {
return fmt.Errorf("quit Claude Desktop: %w", err)
}
if err := waitForClaudeDesktopExit(ctx); err != nil {
return err
}
return restoreClaudeDesktopProfile()
}
func restoreClaudeDesktopProfile() error {
targets, err := claudeDesktopTargetPaths()
if err != nil {
return err
}
return restoreClaudeDesktopTargets(targets)
}
func (c *ClaudeDesktop) Onboard() error {
return config.MarkIntegrationOnboarded(claudeDesktopIntegrationName)
}
// ClaudeDesktopModels returns the user's explicitly saved Claude Desktop
// model subset. A nil result means the recommendation source should decide.
func ClaudeDesktopModels() []string {
return config.IntegrationModels(claudeDesktopIntegrationName)
}
// SaveClaudeDesktopModels persists the user's explicit Claude Desktop model
// subset in the shared launcher configuration.
func SaveClaudeDesktopModels(models []string) error {
if len(models) == 0 {
return errors.New("select at least one Claude Desktop model")
}
return config.SaveIntegration(claudeDesktopIntegrationName, models)
}
// RestoreClaudeDesktopModels restores a previously captured selection. A nil
// selection restores the implicit recommendation defaults.
func RestoreClaudeDesktopModels(models []string) error {
return config.SaveIntegration(claudeDesktopIntegrationName, models)
}
func (c *ClaudeDesktop) RequiresInteractiveOnboarding() bool {
return false
}
@@ -130,12 +222,21 @@ func (c *ClaudeDesktop) SkipModelReadiness() bool {
return true
}
func (c *ClaudeDesktop) Run(_ string, _ []LaunchModel, _ []string) error {
return errClaudeDesktopUnsupported()
func (c *ClaudeDesktop) Run(_ string, _ []LaunchModel, args []string) error {
if err := claudeDesktopSupported(); err != nil {
return err
}
if len(args) > 0 {
return errors.New("claude-desktop does not accept extra arguments")
}
if err := ensureClaudeDesktopGateway(); err != nil {
return err
}
return claudeDesktopLaunchOrRestart("Restart Claude Desktop to use Ollama?", c.ConfigureAutodiscovery)
}
func (c *ClaudeDesktop) Restore() error {
if err := claudeDesktopSupported(); err != nil {
if err := claudeDesktopRestoreSupported(); err != nil {
return err
}
targets, err := claudeDesktopTargetPaths()
@@ -143,6 +244,35 @@ func (c *ClaudeDesktop) Restore() error {
return err
}
if err := restoreClaudeDesktopTargets(targets); err != nil {
return err
}
return claudeDesktopLaunchOrRestart("Restart Claude Desktop to use the usual Claude profile?", func() error {
return restoreClaudeDesktopTargets(targets)
})
}
func configureClaudeDesktopTargets(targets claudeDesktopTargets, baseURL, apiKey string) error {
for _, target := range targets.thirdPartyProfiles {
if err := writeClaudeDesktopGatewayProfile(target.profile, baseURL, apiKey, true); err != nil {
return err
}
if err := writeClaudeDesktopMeta(target.meta, claudeDesktopProfileID, claudeDesktopProfileName); err != nil {
return err
}
if err := writeClaudeDesktopDeploymentMode(target.desktopConfig, "3p"); err != nil {
return err
}
}
for _, path := range targets.normalConfigs {
if err := writeClaudeDesktopDeploymentMode(path, "3p"); err != nil {
return err
}
}
return nil
}
func restoreClaudeDesktopTargets(targets claudeDesktopTargets) error {
for _, path := range targets.normalConfigs {
if err := writeClaudeDesktopDeploymentMode(path, "1p"); err != nil {
return err
@@ -159,27 +289,42 @@ func (c *ClaudeDesktop) Restore() error {
return err
}
}
return claudeDesktopLaunchOrRestart("Restart Claude Desktop to use the usual Claude profile?")
}
func errClaudeDesktopUnsupported() error {
return errors.New(claudeDesktopUnsupported)
return nil
}
func claudeDesktopSupported() error {
switch claudeDesktopGOOS {
case "darwin", "windows":
if claudeDesktopGOOS == "darwin" {
return nil
default:
return fmt.Errorf("Claude Desktop launch is only supported on macOS and Windows")
}
return errors.New("Claude Desktop launch is only supported on macOS")
}
func claudeDesktopInstalled() bool {
if claudeDesktopAppPath() != "" {
return true
func claudeDesktopRestoreSupported() error {
if claudeDesktopGOOS == "darwin" || claudeDesktopGOOS == "windows" {
return nil
}
if claudeDesktopGOOS == "windows" && claudeDesktopIsRunning() {
return errors.New("Claude Desktop restore is only supported on macOS and Windows")
}
func ensureClaudeDesktopGateway() error {
ctx, cancel := context.WithTimeout(context.Background(), claudeDesktopProbeTimeout)
defer cancel()
if err := claudeDesktopProbeGateway(ctx, claudeDesktopGatewayBaseURL); err != nil {
return fmt.Errorf("Claude gateway is unavailable at %s: %w; restart Ollama and try again", claudeDesktopGatewayBaseURL, err)
}
return nil
}
// ClaudeDesktopInstalled reports whether Claude Desktop is installed.
func ClaudeDesktopInstalled() bool {
if claudeDesktopGOOS == "darwin" {
return claudeDesktopAppPath() != ""
}
if claudeDesktopGOOS != "windows" {
return false
}
running, _ := claudeDesktopIsRunning(context.Background())
if claudeDesktopAppPath() != "" || running {
return true
}
for _, dir := range claudeDesktopProfileDirCandidates(false) {
@@ -327,7 +472,7 @@ func claudeDesktopWindowsConfigPaths() (claudeDesktopPaths, error) {
func claudeDesktopProfileDir(normal bool) (string, error) {
candidates := claudeDesktopProfileDirCandidates(normal)
if len(candidates) == 0 {
return "", fmt.Errorf("Claude Desktop profile directory could not be resolved")
return "", errors.New("Claude Desktop profile directory could not be resolved")
}
for _, candidate := range candidates {
if _, err := claudeDesktopStat(candidate); err == nil {
@@ -413,14 +558,6 @@ func newClaudeDesktopTargets(normalRoots, thirdPartyRoots []string) claudeDeskto
return targets
}
func claudeDesktopTargetProfilePaths(targets claudeDesktopTargets) []string {
paths := make([]string, 0, len(targets.thirdPartyProfiles))
for _, target := range targets.thirdPartyProfiles {
paths = append(paths, target.profile)
}
return paths
}
func claudeDesktopLocalAppData() (string, error) {
if local := strings.TrimSpace(os.Getenv("LOCALAPPDATA")); local != "" {
return local, nil
@@ -435,139 +572,6 @@ func claudeDesktopLocalAppData() (string, error) {
return filepath.Join(home, "AppData", "Local"), nil
}
type claudeDesktopAPIKeySource int
const (
claudeDesktopAPIKeySourceNone claudeDesktopAPIKeySource = iota
claudeDesktopAPIKeySourceEnv
claudeDesktopAPIKeySourceProfile
)
func claudeDesktopValidatedAPIKey(ctx context.Context, profilePaths []string) (string, error) {
key, source, err := claudeDesktopAPIKey(profilePaths)
if err != nil {
return "", err
}
if err := claudeDesktopValidateAPIKey(ctx, key); err == nil {
return key, nil
} else if source != claudeDesktopAPIKeySourceProfile || !canPromptClaudeDesktopAPIKey() {
return "", err
}
return promptValidClaudeDesktopAPIKey(ctx)
}
func claudeDesktopAPIKey(profilePaths []string) (string, claudeDesktopAPIKeySource, error) {
if key := strings.TrimSpace(os.Getenv("OLLAMA_API_KEY")); key != "" {
return key, claudeDesktopAPIKeySourceEnv, nil
}
for _, profilePath := range profilePaths {
if key := readClaudeDesktopGatewayAPIKey(profilePath); key != "" {
return key, claudeDesktopAPIKeySourceProfile, nil
}
}
key, err := promptClaudeDesktopAPIKeyValue()
return key, claudeDesktopAPIKeySourceNone, err
}
func canPromptClaudeDesktopAPIKey() bool {
return isInteractiveSession() && !currentLaunchConfirmPolicy.requireYesMessage
}
func promptValidClaudeDesktopAPIKey(ctx context.Context) (string, error) {
key, err := promptClaudeDesktopAPIKeyValue()
if err != nil {
return "", err
}
if err := claudeDesktopValidateAPIKey(ctx, key); err != nil {
return "", err
}
return key, nil
}
func promptClaudeDesktopAPIKeyValue() (string, error) {
if !canPromptClaudeDesktopAPIKey() {
return "", missingClaudeDesktopAPIKeyError()
}
key, err := claudeDesktopPromptAPIKey()
if err != nil {
return "", err
}
key = strings.TrimSpace(key)
if key == "" {
return "", missingClaudeDesktopAPIKeyError()
}
return key, nil
}
func missingClaudeDesktopAPIKeyError() error {
return fmt.Errorf("OLLAMA_API_KEY is required for Claude Desktop. Create an API key at %s, then re-run with OLLAMA_API_KEY set", claudeDesktopAPIKeyURL)
}
func promptClaudeDesktopAPIKey() (string, error) {
fmt.Fprint(os.Stderr, claudeDesktopAPIKeyPrompt())
key, err := term.ReadPassword(int(os.Stdin.Fd()))
fmt.Fprintln(os.Stderr)
if err != nil {
return "", err
}
return string(key), nil
}
func claudeDesktopAPIKeyPrompt() string {
return fmt.Sprintf("Create an Ollama API key at %s\nEnter Ollama API key (input hidden): ", claudeDesktopAPIKeyURL)
}
func readClaudeDesktopGatewayAPIKey(path string) string {
cfg, err := readClaudeDesktopJSON(path)
if err != nil {
return ""
}
key, _ := cfg["inferenceGatewayApiKey"].(string)
return strings.TrimSpace(key)
}
func validateClaudeDesktopAPIKey(ctx context.Context, key string) error {
ctx, cancel := context.WithTimeout(ctx, 15*time.Second)
defer cancel()
if claudeDesktopAPIKeyHasInvalidHeaderChars(key) {
return claudeDesktopAPIKeyVerificationError()
}
req, err := http.NewRequestWithContext(ctx, http.MethodGet, claudeDesktopGatewayBaseURL+"/v1/models", nil)
if err != nil {
return claudeDesktopAPIKeyVerificationError()
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Accept", "application/json")
resp, err := claudeDesktopHTTPClient.Do(req)
if err != nil {
return claudeDesktopAPIKeyVerificationError()
}
defer resp.Body.Close()
_, _ = io.Copy(io.Discard, io.LimitReader(resp.Body, 4<<10))
switch {
case resp.StatusCode == http.StatusUnauthorized || resp.StatusCode == http.StatusForbidden:
return fmt.Errorf("Ollama API key was rejected; create a valid key at %s", claudeDesktopAPIKeyURL)
case resp.StatusCode >= 200 && resp.StatusCode < 300:
return nil
default:
return fmt.Errorf("could not verify Ollama API key; ollama.com returned status %d, try again later", resp.StatusCode)
}
}
func claudeDesktopAPIKeyHasInvalidHeaderChars(key string) bool {
return strings.ContainsFunc(key, func(r rune) bool {
return r < ' ' || r == 0x7f
})
}
func claudeDesktopAPIKeyVerificationError() error {
return fmt.Errorf("could not verify Ollama API key; copy a key from %s and try again", claudeDesktopAPIKeyURL)
}
func writeClaudeDesktopDeploymentMode(path, mode string) error {
cfg, err := readClaudeDesktopJSONAllowMissing(path)
if err != nil {
@@ -604,17 +608,26 @@ func writeClaudeDesktopMeta(path, id, name string) error {
return writeClaudeDesktopJSON(path, meta)
}
func writeClaudeDesktopGatewayProfile(path string, apiKey string, forceChooser bool) error {
func writeClaudeDesktopGatewayProfile(path, baseURL, apiKey string, forceChooser bool) error {
cfg, err := readClaudeDesktopJSONAllowMissing(path)
if err != nil {
return fmt.Errorf("parse Claude Desktop Ollama profile: %w", err)
}
cfg["inferenceProvider"] = "gateway"
cfg["inferenceGatewayBaseUrl"] = claudeDesktopGatewayBaseURL
cfg["inferenceGatewayBaseUrl"] = baseURL
cfg["inferenceGatewayApiKey"] = apiKey
cfg["inferenceGatewayAuthScheme"] = "bearer"
cfg["deploymentDisplayName"] = claudeDesktopProfileName
cfg["chatTabEnabled"] = true
delete(cfg, "inferenceModels")
cfg["disableDeploymentModeChooser"] = forceChooser
cfg["coworkEgressAllowedHosts"] = claudeDesktopEgressHosts
cfg["disableEssentialTelemetry"] = true
cfg["disableNonessentialTelemetry"] = true
// Auto mode sends separate classifier requests through the configured
// inference provider. Keep it disabled until the mapped models are tested
// for that classifier contract.
cfg["autoModeEnabled"] = false
return writeClaudeDesktopJSON(path, cfg)
}
@@ -665,7 +678,12 @@ func restoreClaudeDesktopOllamaProfile(path string) error {
delete(cfg, "inferenceProvider")
delete(cfg, "inferenceGatewayBaseUrl")
delete(cfg, "inferenceGatewayAuthScheme")
delete(cfg, "deploymentDisplayName")
delete(cfg, "inferenceModels")
delete(cfg, "coworkEgressAllowedHosts")
delete(cfg, "autoModeEnabled")
delete(cfg, "disableEssentialTelemetry")
delete(cfg, "disableNonessentialTelemetry")
return writeClaudeDesktopJSON(path, cfg)
}
@@ -688,6 +706,18 @@ func readClaudeDesktopDeploymentMode(path string) string {
}
func claudeDesktopTargetsConfigured(targets claudeDesktopTargets) bool {
if !claudeDesktopTargetsUseOllamaGateway(targets) {
return false
}
for _, target := range targets.thirdPartyProfiles {
if !claudeDesktopThirdPartyProfileConfigured(target) {
return false
}
}
return true
}
func claudeDesktopTargetsUseOllamaGateway(targets claudeDesktopTargets) bool {
if len(targets.normalConfigs) == 0 || len(targets.thirdPartyProfiles) == 0 {
return false
}
@@ -700,7 +730,7 @@ func claudeDesktopTargetsConfigured(targets claudeDesktopTargets) bool {
if readClaudeDesktopDeploymentMode(target.desktopConfig) != "3p" {
return false
}
if !claudeDesktopThirdPartyProfileConfigured(target) {
if !claudeDesktopThirdPartyProfileUsesOllamaGateway(target) {
return false
}
}
@@ -708,6 +738,39 @@ func claudeDesktopTargetsConfigured(targets claudeDesktopTargets) bool {
}
func claudeDesktopThirdPartyProfileConfigured(target claudeDesktopThirdPartyPaths) bool {
if !claudeDesktopThirdPartyProfileUsesOllamaGateway(target) {
return false
}
cfg, err := readClaudeDesktopJSON(target.profile)
if err != nil {
return false
}
if s, _ := cfg["inferenceGatewayApiKey"].(string); strings.TrimSpace(s) == "" {
return false
}
if s, _ := cfg["deploymentDisplayName"].(string); s != claudeDesktopProfileName {
return false
}
egressHosts := claudeDesktopAnySlice(cfg["coworkEgressAllowedHosts"])
if len(egressHosts) != len(claudeDesktopEgressHosts) {
return false
}
for i, host := range egressHosts {
if host != claudeDesktopEgressHosts[i] {
return false
}
}
if disabled, _ := cfg["disableEssentialTelemetry"].(bool); !disabled {
return false
}
if disabled, _ := cfg["disableNonessentialTelemetry"].(bool); !disabled {
return false
}
return true
}
func claudeDesktopThirdPartyProfileUsesOllamaGateway(target claudeDesktopThirdPartyPaths) bool {
if readClaudeDesktopAppliedID(target.meta) != claudeDesktopProfileID {
return false
}
@@ -719,15 +782,23 @@ func claudeDesktopThirdPartyProfileConfigured(target claudeDesktopThirdPartyPath
if s, _ := cfg["inferenceProvider"].(string); s != "gateway" {
return false
}
if s, _ := cfg["inferenceGatewayBaseUrl"].(string); strings.TrimRight(s, "/") != claudeDesktopGatewayBaseURL {
return false
}
if s, _ := cfg["inferenceGatewayApiKey"].(string); strings.TrimSpace(s) == "" {
if s, _ := cfg["inferenceGatewayBaseUrl"].(string); !claudeDesktopGatewayURLMatches(s) {
return false
}
return true
}
func claudeDesktopGatewayURLMatches(value string) bool {
parsed, err := url.Parse(strings.TrimSpace(value))
if err != nil || parsed.User != nil || parsed.RawQuery != "" || parsed.Fragment != "" {
return false
}
if strings.TrimRight(parsed.EscapedPath(), "/") != "" {
return false
}
return strings.EqualFold(parsed.Scheme, "http") && strings.EqualFold(parsed.Host, proxy.DefaultClaudeDesktopListenAddr)
}
func readClaudeDesktopJSONAllowMissing(path string) (map[string]any, error) {
cfg, err := readClaudeDesktopJSON(path)
if errors.Is(err, os.ErrNotExist) {
@@ -774,15 +845,14 @@ func claudeDesktopAnySlice(value any) []any {
}
}
func claudeDesktopLaunchOrRestart(prompt string) error {
if !claudeDesktopIsRunning() {
func claudeDesktopLaunchOrRestart(prompt string, reapplyProfile func() error) error {
running, err := claudeDesktopIsRunning(context.Background())
if err != nil {
return fmt.Errorf("check whether Claude Desktop is running: %w", err)
}
if !running {
return claudeDesktopOpenApp()
}
restartAppPath := ""
if claudeDesktopGOOS == "windows" {
restartAppPath = claudeDesktopRunningAppPath()
}
restart, err := ConfirmPrompt(prompt)
if err != nil {
return err
@@ -791,41 +861,95 @@ func claudeDesktopLaunchOrRestart(prompt string) error {
fmt.Fprintln(os.Stderr, "\nQuit and reopen Claude Desktop when you're ready for the profile change to take effect.")
return nil
}
return restartClaudeDesktop(reapplyProfile)
}
if err := claudeDesktopQuitApp(); err != nil {
func restartClaudeDesktop(reapplyProfile func() error) error {
restartAppPath := ""
if claudeDesktopGOOS == "windows" {
restartAppPath = claudeDesktopRunningAppPath()
}
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
if err := claudeDesktopQuitApp(ctx); err != nil {
return fmt.Errorf("quit Claude Desktop: %w", err)
}
if err := waitForClaudeDesktopExit(30 * time.Second); err != nil {
if err := waitForClaudeDesktopExit(ctx); err != nil {
return err
}
// Claude persists settings while shutting down. Reapply the profile after
// exit so its last write cannot restore stale surface or gateway values.
if err := reapplyProfile(); err != nil {
reapplyErr := fmt.Errorf("reapply Claude Desktop profile: %w", err)
if openErr := openClaudeDesktopAfterRestart(restartAppPath); openErr != nil {
return errors.Join(reapplyErr, fmt.Errorf("reopen Claude Desktop after profile failure: %w", openErr))
}
return reapplyErr
}
return openClaudeDesktopAfterRestart(restartAppPath)
}
func openClaudeDesktopAfterRestart(restartAppPath string) error {
if restartAppPath != "" {
return claudeDesktopOpenAppPath(restartAppPath)
}
return claudeDesktopOpenApp()
}
func waitForClaudeDesktopExit(timeout time.Duration) error {
deadline := time.Now().Add(timeout)
for time.Now().Before(deadline) {
if !claudeDesktopIsRunning() {
func waitForClaudeDesktopExit(ctx context.Context) error {
for {
running, err := claudeDesktopIsRunning(ctx)
if err != nil {
return fmt.Errorf("check whether Claude Desktop is running: %w", err)
}
if !running {
return nil
}
claudeDesktopSleep(200 * time.Millisecond)
select {
case <-ctx.Done():
return errors.New("Claude Desktop did not quit; quit it manually and re-run the command")
case <-time.After(200 * time.Millisecond):
}
}
return fmt.Errorf("Claude Desktop did not quit; quit it manually and re-run the command")
}
func defaultClaudeDesktopIsRunning() bool {
// ClaudeDesktopRunning reports whether Claude Desktop is open.
func ClaudeDesktopRunning() bool {
running, _ := claudeDesktopIsRunning(context.Background())
return running
}
// OpenClaudeDesktop brings the installed Claude Desktop app to the foreground.
func OpenClaudeDesktop() error {
if err := claudeDesktopSupported(); err != nil {
return err
}
return claudeDesktopOpenApp()
}
func defaultClaudeDesktopRunning(ctx context.Context) (bool, error) {
var (
out []byte
err error
)
switch claudeDesktopGOOS {
case "darwin":
out, err := exec.Command("pgrep", "-f", "Claude.app/Contents/MacOS/Claude").Output()
return err == nil && strings.TrimSpace(string(out)) != ""
out, err = exec.CommandContext(ctx, "pgrep", "-f", "Claude.app/Contents/MacOS/Claude").Output()
if exitErr := (*exec.ExitError)(nil); errors.As(err, &exitErr) && exitErr.ExitCode() == 1 && ctx.Err() == nil {
return false, nil
}
case "windows":
out, err := exec.Command("powershell.exe", "-NoProfile", "-Command", `(Get-Process claude -ErrorAction SilentlyContinue | Where-Object { $_.MainWindowHandle -ne 0 } | Select-Object -First 1).Id`).Output()
return err == nil && strings.TrimSpace(string(out)) != ""
out, err = exec.CommandContext(ctx, "powershell.exe", "-NoProfile", "-Command", `(Get-Process claude -ErrorAction SilentlyContinue | Where-Object { $_.MainWindowHandle -ne 0 } | Select-Object -First 1).Id`).Output()
default:
return false
return false, nil
}
if ctxErr := ctx.Err(); ctxErr != nil {
return false, ctxErr
}
if err != nil {
return false, err
}
return strings.TrimSpace(string(out)) != "", nil
}
func defaultClaudeDesktopOpenApp() error {
@@ -837,9 +961,13 @@ func defaultClaudeDesktopOpenApp() error {
if path := claudeDesktopRunningAppPath(); path != "" {
return claudeDesktopOpenAppPath(path)
}
return fmt.Errorf("Claude Desktop executable was not found; open Claude Desktop manually once and re-run 'ollama launch claude-desktop --restore'")
return errors.New("Claude Desktop executable was not found; open Claude Desktop manually once and re-run 'ollama launch claude-desktop --restore'")
case "darwin":
return openClaudeDesktopDarwin()
path := claudeDesktopAppPath()
if path == "" {
return errors.New("Claude Desktop app was not found")
}
return openClaudeDesktopDarwin(path)
default:
return claudeDesktopSupported()
}
@@ -850,14 +978,18 @@ func defaultClaudeDesktopOpenAppPath(path string) error {
case "windows":
return exec.Command("powershell.exe", "-NoProfile", "-Command", "Start-Process -FilePath "+quotePowerShellString(path)).Run()
case "darwin":
return openClaudeDesktopDarwin()
return openClaudeDesktopDarwin(path)
default:
return claudeDesktopSupported()
}
}
func openClaudeDesktopDarwin() error {
cmd := exec.Command("open", "-a", "Claude")
func claudeDesktopDarwinOpenArgs(path string) []string {
return []string{path}
}
func openClaudeDesktopDarwin(path string) error {
cmd := exec.Command("/usr/bin/open", claudeDesktopDarwinOpenArgs(path)...)
cmd.Stdout = os.Stdout
cmd.Stderr = os.Stderr
return cmd.Run()
@@ -875,12 +1007,12 @@ func defaultClaudeDesktopRunningAppPath() string {
return strings.TrimSpace(string(out))
}
func defaultClaudeDesktopQuitApp() error {
func defaultClaudeDesktopQuitApp(ctx context.Context) error {
if claudeDesktopGOOS == "windows" {
script := `Get-Process claude -ErrorAction SilentlyContinue | Where-Object { $_.MainWindowHandle -ne 0 } | ForEach-Object { [void]$_.CloseMainWindow() }`
return exec.Command("powershell.exe", "-NoProfile", "-Command", script).Run()
return exec.CommandContext(ctx, "powershell.exe", "-NoProfile", "-Command", script).Run()
}
return exec.Command("osascript", "-e", `tell application "Claude" to quit`).Run()
return exec.CommandContext(ctx, "osascript", "-e", `tell application "Claude" to quit`).Run()
}
func quotePowerShellString(s string) string {
File diff suppressed because it is too large. Load diff
Loaded 100 of 595 files, more files were not shown because too many files have changed in this diff. Show more