Compare commits

...
Author SHA1 Message Date
Andrei Cravtov 0abd91c6ca moved to using NewPy<T> 2026-06-22 18:25:16 +01:00
Andrei Cravtov 6440046af1 removed write effects 2026-06-22 16:01:25 +01:00
Andrei Cravtov 7f0ba1628a added config.toml 2026-06-15 21:42:26 +01:00
Andrei Cravtov c2efbfd35d made new_py work 2026-06-14 01:45:15 +01:00
Andrei Cravtov 9fb66e313b moved newpy 2026-06-13 22:51:23 +01:00
Andrei Cravtov 22adc34311 deprecated opts 2026-06-13 21:45:51 +01:00
Andrei Cravtov 04a688083f next targets defined 2026-06-13 19:59:20 +01:00
Andrei Cravtov 9e1a41b687 re-org constants 2026-06-13 19:49:54 +01:00
Andrei Cravtov 87657cf061 EXO_MAX_CONCURRENT_REQUESTS removed 2026-06-13 19:48:57 +01:00
Andrei Cravtov 7258bdbfd1 ENABLE_DISAGGREGATION migrated 2026-06-13 19:33:59 +01:00
Andrei Cravtov 41afeb7c5e EXO_TRACING_ENABLED removed 2026-06-13 19:18:26 +01:00
Andrei Cravtov dc8e3a97fa EXO_ENABLE_IMAGE_MODELS removed 2026-06-13 19:06:05 +01:00
Andrei Cravtov f3f126b69e moved 2026-06-12 21:18:25 +01:00
Andrei Cravtov 8dd3715d13 --bootstrap-peers removed; Args removed 2026-06-12 21:05:20 +01:00
Andrei Cravtov 3b49c1b493 deprecated 2026-06-12 19:47:30 +01:00
Andrei Cravtov b6b9761d5f sample implenentation of CliPy a newtype wrapper around python object references that is supposed to replace Py<T> such that it is parse-able from the CLI and Serde - and hence we can have NORMAL getters/setters without weird copy behaviour 2026-06-12 18:33:08 +01:00
Andrei Cravtov de7454e873 --namespace migrated 2026-06-12 14:54:15 +01:00
Andrei Cravtov 5a2a96b8c5 --zenoh-port; --discovery-port migrated 2026-06-12 14:40:59 +01:00
Andrei Cravtov 4f2427bd22 --api-port migrated 2026-06-12 14:34:17 +01:00
Andrei Cravtov fa6faced53 -q and -v migrated 2026-06-12 14:23:34 +01:00
Andrei Cravtov 98901d3234 verbosity + quet in rust 2026-06-11 22:48:35 +01:00
Andrei Cravtov 21236dd34e fmt 2026-06-11 22:22:55 +01:00
Andrei Cravtov b7e374699c add comment 2026-06-11 22:22:46 +01:00
Andrei Cravtov 8f80dc23d1 migrated --offline 2026-06-11 22:10:18 +01:00
Andrei Cravtov fa4ec5a979 migrated --no-batch 2026-06-11 21:38:37 +01:00
Andrei Cravtov a434944921 move continuous_batching_enabled to AppArgs and AppSettings 2026-06-11 21:34:20 +01:00
Andrei Cravtov 8688ff1855 --force-master; --no-api; --no-worker; --no-downloads; all removed 2026-06-11 21:03:40 +01:00
Andrei Cravtov c405d89a26 migrate --legacy-daemon 2026-06-11 20:54:46 +01:00
Andrei Cravtov dbfc7f8ba6 restart 2026-06-11 20:46:13 +01:00
Andrei Cravtov c5d8f8c331 migrate --fast-sync 2026-06-11 20:36:59 +01:00
Andrei Cravtov 91c4d2b5fc fix all tests and usages locator -> bootstrap 2026-06-11 20:06:00 +01:00
Andrei Cravtov adec95a0a3 config initialization 2026-06-11 19:48:45 +01:00
Andrei Cravtov 9d0d4e878c change Config to AppSettings 2026-06-11 19:37:14 +01:00
Andrei Cravtov ced27f050d rename to bootstrap 2026-06-11 19:28:51 +01:00
Andrei Cravtov 6bc00555ac persist locator configuration to the runner subprocesses 2026-06-11 18:40:19 +01:00
Andrei Cravtov 8f97457a0d fix tests now all is updated to use config.locator 2026-06-11 18:02:09 +01:00
Andrei Cravtov d395c50fba EXO_DATA_HOME, EXO_CONFIG_HOME, _EXO_HOME_ENV removed 2026-06-11 17:45:56 +01:00
Andrei Cravtov 2a507509fa EXO_CACHE_HOME removed 2026-06-11 17:41:54 +01:00
Andrei Cravtov 2a230be446 EXO_DEFAULT_MODELS_DIR, EXO_MODELS_DIRS, EXO_MODELS_READ_ONLY_DIRS removed 2026-06-11 17:37:44 +01:00
Andrei Cravtov d5ab01d9a1 updated EXO_LIBP2P_NAMESPACE -> EXO_ZENOH_NAMESPACE in many places 2026-06-11 17:19:44 +01:00
Andrei Cravtov dc30ca7fa6 all tests updated 2026-06-11 17:11:57 +01:00
Andrei Cravtov 04fec92987 EXO_DEFAULT_MODELS_DIR 2026-06-11 16:54:52 +01:00
Andrei Cravtov 6035f66b61 fixed many tests 2026-06-11 16:29:01 +01:00
Andrei Cravtov 733cd1b9c2 lint 2026-06-11 14:22:59 +01:00
Andrei Cravtov 462a72e22d fixed more tests 2026-06-11 14:17:07 +01:00
Andrei Cravtov b2c8dc9c49 fixed more tests 2026-06-11 14:07:36 +01:00
Andrei Cravtov d5f1d0792f fix more tests 2026-06-11 13:53:06 +01:00
Andrei Cravtov 683e1b1a29 lint 2026-06-11 12:39:22 +01:00
Andrei Cravtov d4c86b68b1 fix test 2026-06-11 03:39:15 +01:00
Andrei Cravtov 8c3e360bd1 make mutable 2026-06-11 01:37:19 +01:00
Andrei Cravtov b00b1882da made locator load from env where possible 2026-06-10 23:06:34 +01:00
Andrei Cravtov 22228c5d02 EXO_LOG EXO_LOG_DIR removed 2026-06-10 22:13:21 +01:00
Andrei Cravtov 37440ce6a0 EXO_RUNNER_LOG_DIR EXO_RUNNER_STDOUT_LOG EXO_RUNNER_STDERR_LOG removed 2026-06-10 22:09:28 +01:00
Andrei Cravtov 8c9b000ec3 EXO_PID_FILE removed 2026-06-10 22:02:56 +01:00
Andrei Cravtov f1f393bfdf EXO_NODE_ZID removed 2026-06-10 21:53:26 +01:00
Andrei Cravtov 131e3af4ff EXO_CONFIG_FILE removed 2026-06-10 21:45:52 +01:00
Andrei Cravtov f4a2ffa577 EXO_CUSTOM_MODEL_CARDS_DIR removed 2026-06-10 21:44:33 +01:00
Andrei Cravtov 817c556851 EXO_EVENT_LOG_DIR removed 2026-06-10 21:37:53 +01:00
Andrei Cravtov 1e8d4abe94 EXO_IMAGE_CACHE_DIR removed 2026-06-10 21:30:02 +01:00
Andrei Cravtov c2ecc8b59e EXO_TRACING_CACHE_DIR removed 2026-06-10 21:26:28 +01:00
Andrei Cravtov a5ec6f783f rename to "locator" 2026-06-10 21:19:44 +01:00
Andrei Cravtov e04208605e whoops fixed bug 2026-06-10 21:17:25 +01:00
Andrei Cravtov 06fa9c3fee all locator configs successfully saved 2026-06-10 21:03:59 +01:00
Andrei Cravtov f54a701979 added the rest of exo paths in constants 2026-06-10 20:26:28 +01:00
Andrei Cravtov 24cab4799c implement log dirs + creation of config file 2026-06-10 20:10:22 +01:00
Andrei Cravtov fdf5f0c00b implement models dirs 2026-06-10 19:05:00 +01:00
Andrei Cravtov 9604a1a18c implemented policy for parsing paths propperly 2026-06-09 20:32:23 +01:00
Andrei Cravtov 13e5bf8c16 parsers 2026-06-09 19:53:22 +01:00
Andrei Cravtov ae3b195868 testing 2026-06-09 19:25:32 +01:00
Andrei Cravtov c9c6b59562 path parser 2026-06-09 19:13:25 +01:00
Andrei Cravtov 72d3bfc088 consolidate into ExoHome 2026-06-09 14:19:13 +01:00
Andrei Cravtov 2db2abbb1e made pickling work:
added to/from bytes + reduce + set module so it isn't builtins.<Class>
2026-06-09 14:00:41 +01:00
Andrei Cravtov 2cacfb5a9b namespace 2026-06-08 22:31:37 +01:00
Andrei Cravtov 051563a303 config 2026-06-08 19:34:16 +01:00
Andrei Cravtov d3d680f569 no more NodeConfig event in info_gather 2026-06-08 19:33:19 +01:00
Andrei Cravtov c7c449f550 remove A 2026-06-08 19:22:54 +01:00
Andrei Cravtov df2925ce15 remove unused constants 2026-06-08 19:22:12 +01:00
Andrei Cravtov 13b4ac4162 expose locator config 2026-06-08 18:45:00 +01:00
Andrei Cravtov 4883bcd3a9 add version from python package and expose CliArgs to python 2026-06-08 18:08:07 +01:00
Andrei Cravtov d4a61620d2 tweak 2026-06-04 21:28:32 +01:00
Andrei Cravtov 6649ce7f0c wrote locator logic 2026-06-04 20:31:05 +01:00
Andrei Cravtov bc06e029be locator args 2026-06-04 18:52:34 +01:00
Andrei Cravtov 92e9c9f8c2 made preparations for config args 2026-06-04 17:56:51 +01:00
Andrei Cravtov 5abd06735b deprecated error validation 2026-06-04 17:47:27 +01:00
Andrei Cravtov 7186ec2423 started working on deprecating: added version from ENV
need to update build scripts for this
2026-06-04 17:28:56 +01:00
Andrei Cravtov f9fda49ae8 added revisions to cargo.toml to prevent recompilation 2026-06-04 15:51:27 +01:00
Andrei Cravtov e892f7fb8a Merge branch 'main' into andrei/rust-settings
# Conflicts:
#	Cargo.lock
#	rust/exo_rs/Cargo.toml
#	rust/exo_rs/exo_rs.pyi
#	rust/exo_rs/src/lib.rs
#	rust/exo_rs/src/networking.rs
#	src/exo/main.py
2026-06-04 15:11:58 +01:00
Andrei Cravtov b7730f743d implemented the cli args propperly - this is step 1: 2026-06-03 17:08:30 +01:00
Evan QuineyandAndrei Cravtov 09f9ea313f libp2p -> zenoh (#2132)
supercedes #2076 and #2073

---------

Co-authored-by: Andrei Cravtov <the.andrei.cravtov@gmail.com>
2026-06-03 16:31:56 +01:00
Andrei Cravtov e12744edd6 initial 2026-06-02 18:06:39 +01:00
Andrei Cravtov 8506e7a4dc async with tokio 2026-06-02 18:04:50 +01:00
Sakutaro 81d7cb0fcd docs: add Homebrew cask install instructions (#2140)
## Motivation

exo is now available as a Homebrew cask, so the README should show the
simplest macOS installation path alongside the existing DMG download.

Fixes https://github.com/exo-explore/exo/issues/2105
https://github.com/exo-explore/exo/issues/176

## Changes

- Added `brew install --cask exo` to the macOS App section of
`README.md`
- Kept the existing DMG download link as the first installation option

## Why It Works

Adding the Homebrew cask command gives macOS users a
package-manager-managed installation path while preserving the existing
DMG download option.

## Test Plan

### Manual Testing

- Reviewed the rendered Markdown structure in `README.md`

### Automated Testing

- Not run. Documentation-only change.

## Related

- https://github.com/Homebrew/homebrew-cask/pull/265956
2026-06-02 15:35:18 +00:00
Andrei Cravtov 439f59924a init 2026-06-02 01:47:31 +01:00
Andrei Cravtov 629c55d6ba Rename exo_pyo3_bindings to exo_rs (#2131)
## Motivation

(I think it) Makes Evan's massive PR easier to merge later on

## Changes

- Renamed exo_pyo3_bindings to exo_rs
- Upgraded versions of pyo3-based dependencies
- Renamed PyFromSwarm to just FromSwarm, and PyNetworkingHandle to just
NetworkingHandle
2026-05-31 19:23:41 +01:00
Andrei Cravtov f9f8cbb3c3 fix: make app builds work again (#2127)
## Motivation

They didn't

## Changes

They now do

## Why It Works

I changed an env flag, and added a keyword
2026-05-29 18:37:47 +01:00
ciaranbor 051a64e3b4 Capture energy in prefill and ageneration separately (#2124)
## Motivation

Energy was reported as a single aggregate. Split into prefill vs.
generation so each phase can be analysed independently.

## Changes

- `PowerSampler`: `mark_prefill_done()` + `trapezoidal_energy_range()`
helper; `result()` now emits per-phase splits.
- `PowerUsage` / `NodePowerStats`: optional `prefill_*` / `generation_*`
fields (back-compat: `None` if unmarked).
- API marks the boundary on the first non-`PrefillProgressChunk`.
- `bench/exo_bench.py` surfaces the split in the log line and persists
`power_usage` to JSON.
- METHODOLOGY: one sentence + one bullet.

## Why It Works

First non-prefill chunk *is* the boundary. Anchoring a sample there and
interpolating power at the boundary makes phase energies sum exactly to
the unsplit total.

## Test Plan

### Manual Testing

`eco`-reserved nodes:
- M3 Ultra, Qwen3-VL-4B, pp=8192/tg=1024: server 1940 J vs client 1931 J
(+0.5 %)
- M4 Pro, Qwen3.6-27B, pp=16384/tg=2048: server 20,292 J vs client
20,221 J (+0.35 %)

### Automated Testing

5 new tests in `test_power_sampler.py` (range integrator,
splits-sum-to-total, `None`-when-unmarked, idempotency). 14/14 pass.
2026-05-28 14:42:36 -07:00
Andrei Cravtov a8602ea6d5 fix(bug): no longer repeated _trigger_notify_user_to_download_model (#2114)
## Motivation

Partially fixes [this](https://github.com/exo-explore/exo/issues/2098)
issue. Removed erroneous logic for telling user to download when they
already downloaded.

Could not figure out about the "spontaneous crashes" in that issue,
author should consolidate more logs and open a new issue dedicated to
that. I believe
[this](https://github.com/exo-explore/exo/commit/74e9fe15e62fe189dc7e019db86e75c83eca2721)
commit solved some EventRouter-related crashes, which was mentioned in
[this](https://github.com/exo-explore/exo/issues/2098) issue, so it may
have already been solved. If not, should be re-submitted as a new issue.

## Changes

- Consolidated _resolve_and_validate_text_model and
_validate_image_model into one function: _validate_model_has_instance;
- + They already had virtually identical logic, it being different seems
to be an artifact of history
- + Added logic to ensure that _trigger_notify_user_to_download_model is
only called when no such model is downloaded, not just if there is no
instance of it
- Added a new `/instance/await` SSE streaming endpoint to wait for when
a model has an instance available. Complements instance-placement API,
so we can wait till that is done without client-side polling.
- Updated docs and a /tmp script to reflect some of the changes
- Updated dashboard `getModelForRequest` to only return model ID if an
instance exists for it, and updated bits to use `handleChatSend` instead
of `sendMessage` because that checks for if a model instance exists
first.

## Why It Works

The problem was that there was erroneous logging for model not
downloaded. I fixed that logic. The rest is extra.
2026-05-26 14:42:39 +01:00
Andrei Cravtov a1a22b5f38 feat: added background/daemon support (#2106)
## Motivation

Addresses [this](https://github.com/exo-explore/exo/issues/1931) issue.

## Changes

You can now launch Exo as a legacy SysV-style daemin (in the background)
with `--legacy-daemon` flag.
NOTE: don't use it if you're managing Exo with systemd or launchd

SIDE FIX: the macmon process not found trace is no longer displayed on
process shutdown via ctrl+c, that error is supressed.

## Why It Works

Because I used a daemonization library and tweaked it not to break
multiprocessing.

## Test Plan

I ran it in daemon mode, non daemon mode, etc., and pid locking +
inference + everything else works just fine.

Also ran it `ssh user@host -t 'cd exo && nohup nix run .#exo --
--legacy-daemon'` on a 4-node TB mac-mini cluster and the mDNS didn't
die
2026-05-25 20:42:47 +01:00
Andrei Cravtov 74e9fe15e6 fix(bug): EventRouter lifetime-handling fixed, no more process crashes (#2102)
## Motivation

Trying to (partially) fix
[this](https://github.com/exo-explore/exo/issues/2101) issue.

## Changes

Changed channels (in channels.py) to support exception overriding.

Made EventRouter channels throw a subclass of the resource closed/broken
errors.

The current lifetime logic of EventRouter in event loop no longer blows
up because components that use channels from EventRouter now catch the
subclass exceptions in the run method: Worker, Master,
DownloadCoordinator, RunnerSupervisor.

Added logic to throw when API server exits without being asked to shut
down - this kill the sleep-forever in the task-group.
2026-05-22 14:20:04 +01:00
Evan Quiney 90f24bef30 fix model cards not validating properly after #2071 (#2096) 2026-05-15 15:17:35 +00:00
Andrei Cravtov 5097b2665d Tweaked workspace settings (#2095)
workspace settings
2026-05-15 13:04:50 +00:00
rltakashigeandEvan bc6661e6aa Add node backends to model cards (#2071)
Co-authored-by: Evan <evanev7@gmail.com>
2026-05-15 12:52:12 +00:00
Andrei CravtovandEvan Quiney 14aab35688 Runner error handling (#2093)
# Runner error handling

## Motivation

Runner failures were mostly surfaced as plain shutdown messages, which
made root cause hard to spot from API errors or runner status.

This adds a MVP path for preserving runner crash context and attaching
known stderr diagnostics to failure reports.

## Changes

- Added `RunnerTerminationError` for Python exceptions raised inside
runner bootstrap
- Changed runner bootstrap to send `Event | RunnerTerminationError` over
the private runner channel
- Moved public `RunnerFailed` emission back into supervisor
- Added stderr-only `RunnerDiagnosticCollector`
- + Added known diagnostics for Metal GPU timeout, ring socket receive
errno, and ring transport abort
- Added diagnostics to `RunnerFailed` and `ErrorChunk`
- Tweaked async process termination to join briefly before
terminate/kill
- Updated tests/fixtures for new failure payload shape
- Added Ruff VS Code formatter settings

## Why It Works

Runner child now reports raw-ish failure context to supervisor instead
of publishing failed status directly.

Supervisor still owns process lifecycle, exit code/signal handling,
in-flight task error chunks, and final runner status. Stderr diagnostics
stay best effort and only known root-cause variants are surfaced.

## Test Plan

### Manual Testing

Hardware: remote runner logs from e16/e11/e4/e2

What you did:
- inspected live runner stderr logs
- used observed Metal GPU timeout and ring socket errors as initial
diagnostic targets

### Automated Testing

- `nix flake check`
- supervisor test covers error chunk + failed status emission
- plan lifecycle test updated for failed runner diagnostics
- type/lint checks cover new runner channel union

---------

Co-authored-by: Evan Quiney <evanev7@gmail.com>
2026-05-15 12:40:59 +00:00
Heidar 88d46d46fd fix: omit null delta fields in streaming chat completions (issue #2082) (#2092)
## Motivation

Streaming /v1/chat/completions responses emitted null for tool_calls,
function_call, name, and tool_call_id in every delta chunk. The OpenAI
streaming spec marks these fields as non-nullable — they must either
carry a
  real value or be absent entirely. Spec-correct clients doing
delta.get("tool_calls", []) receive None and crash with 'NoneType'
object is
  not iterable.

Root cause: the streaming serialisation path called model_dump_json()
without
exclude_none=True, while the request-parsing path already used it
correctly.
Three call sites in chat_completions.py and two in responses.py were
affected.

## Testing

Before — every delta carries explicit nulls:

  $ curl -sN -X POST http://localhost:52415/v1/chat/completions \
    -H 'Content-Type: application/json' \
-d
'{"model":"mlx-community/Qwen3.5-2B-MLX-8bit","messages":[{"role":"user","
  content":"hi"}],"max_tokens":3,"stream":true}' \
    | grep "^data: "
data:
{"id":"7c4dae10-...","choices":[{"index":0,"delta":{"role":"assistant","c

ontent":null,"reasoning_content":"Okay","name":null,"tool_calls":null,"tool_cal

l_id":null,"function_call":null},"logprobs":null,"finish_reason":null,"usage":n
  ull}],"usage":null,"service_tier":null}
data:
{"id":"7c4dae10-...","choices":[{"index":0,"delta":{"role":"assistant","c

ontent":null,"reasoning_content":",","name":null,"tool_calls":null,"tool_call_i

d":null,"function_call":null},"logprobs":null,"finish_reason":null,"usage":null
  }],"usage":null,"service_tier":null}
data:
{"id":"7c4dae10-...","choices":[{"index":0,"delta":{"role":"assistant","c
ontent":"
the","reasoning_content":null,"name":null,"tool_calls":null,"tool_cal

l_id":null,"function_call":null},"logprobs":null,"finish_reason":"length","usag
  e":{"prompt_tokens":11,...}}],"usage":null,"service_tier":null}
  data: [DONE]

  After — only populated fields are emitted:
data:
{"id":"demo","object":"chat.completion","created":...,"model":"mlx-commun

ity/Qwen3.5-2B-MLX-8bit","choices":[{"index":0,"delta":{"role":"assistant","rea
  soning_content":"Okay"}}]}
data:
{"id":"demo","object":"chat.completion","created":...,"model":"mlx-commun

ity/Qwen3.5-2B-MLX-8bit","choices":[{"index":0,"delta":{"role":"assistant","rea
  soning_content":","}}]}
data:
{"id":"demo","object":"chat.completion","created":...,"model":"mlx-commun

ity/Qwen3.5-2B-MLX-8bit","choices":[{"index":0,"delta":{"role":"assistant","con
tent":"
the"},"finish_reason":"length"}],"usage":{"prompt_tokens":11,"completio
  n_tokens":3,"total_tokens":14,...}}
  data: [DONE]
2026-05-14 16:32:54 +00:00
HeidarandClaude Opus 4.7 e8ec8d5010 fix ollama API compatibility for VS Code Copilot (#2091)
Ollama adapter fixes for VS Code Copilot (#2042):

  - /api/version: bare semver "1.0.0" - Copilot parseInts each segment.
- /api/show: populate model_info + capabilities - Copilot crashes on
null model_info and filters by `tools`.
- Add POST /ollama/v1/chat/completions - ollama serves the OpenAI-compat
route here, BYOK clients 405 without it.


Before:
<img width="1380" height="144" alt="image"
src="https://github.com/user-attachments/assets/99d5464f-187d-4432-9a31-8229c55aa209"
/>

After:
<img width="1362" height="181" alt="image"
src="https://github.com/user-attachments/assets/361dc006-d8df-435f-8d8b-4fa4f44a8c23"
/>
<img width="279" height="909" alt="image"
src="https://github.com/user-attachments/assets/4621aba7-bd57-4762-8568-34a3383a6025"
/>

Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
2026-05-14 16:12:58 +00:00
Heidar 1fd15d59fc create directory on startup (#2089)
## Motivation

<!-- Why is this change needed? What problem does it solve? -->
<!-- If it fixes an open issue, please link to the issue here -->

When you first run `uv run exo` you get an error like :

`FileNotFoundError: [Errno 2] No such file or directory:
'/Users/heidar/.exo/models'`

Manually tested on Macbook Pro M1 32GB

Fixes issue - https://github.com/exo-explore/exo/issues/2090
2026-05-14 16:03:51 +00:00
Evan Quiney 4466cd5323 use custom mlx sources for linux (#2087)
switch to hosting mlx sources on github & cachix instead of using a
broken version of mlx. closes #2043.
2026-05-13 10:45:11 +01:00
Andrei Cravtov ed2d10bdc6 Redirect runner stdout/stderr to file logs (#2084)
## Motivation

We want to use log mining tools like
[Drain3](https://github.com/logpai/Drain3) to get standardized error
formats, but for that we should record runner stdout/stderr in a massive
append-only log to gather training data for such tools. Also useful for
future opt-in telemetry.

## Changes

The stdout/stderr from runner now splits into 3 tasks: 
1) raw write to dedicated runner logs 
2) sanitized line-by-line logging with log-guru 
3) stub for further error-processing (i.e. turning lines into errors)

### Manual Testing
Works on 4x mac mini clusted connected as TB4 ring.
2026-05-12 11:48:08 +01:00
Andrei Cravtov 87c72fc1fd Fixes issue #2068 (#2083)
## Motivation

To fix https://github.com/exo-explore/exo/issues/2068

## Changes

Adds queue shutdown logic & hard-timeouts for closing server.

## Why It Works

Prevents API from hanging more than 5 seconds.
2026-05-11 12:15:22 +00:00
Evan Quiney b76bc30107 bump rust versions (#2081) 2026-05-10 17:11:46 +00:00
08ffa5f637 Map GLM 4.7 stop tokens to GLM 4 IDs (#2061)
## Motivation

GLM 4.7 reuses the GLM 4 chat-template tokenizer, but the model card and
EOS-detection path didn't have an explicit mapping for it, so
OpenAI-compatible clients didn't see a clean stop and the runner emitted
follow-on role turns (e.g. \`<|user|>\` continuations after
\`<|assistant|>\`'s output).

## Changes

\`src/exo/worker/engines/mlx/utils_mlx.py\` — add the GLM 4 stop-token
IDs as the EOS set when the loaded model's tokenizer matches GLM 4 / 4.7
chat templates.

## Why It Works

The GLM 4 tokenizer's \`<|user|>\`, \`<|observation|>\`, and
\`<|endoftext|>\` IDs are stable across the GLM 4 / 4.7 line; treating
any of them as EOS lets the runner stop at the assistant turn boundary
the same way it stops at \`</s>\` for Llama-style models. No
prompt-template changes — only the stop set widens.

## Test Plan

### Automated Testing

New unit test
\`src/exo/worker/tests/unittests/test_mlx/test_eos_token_ids.py\`
covering: GLM 4 / 4.7 path returns the expected stop ID set; non-GLM
path returns the standard EOS only.

\`\`\`
src/exo/worker/tests/unittests/test_mlx/test_eos_token_ids.py ..
=== 2 passed in 0.01s ===
\`\`\`

\`uv run basedpyright\` and \`uv run ruff check\` both clean.

### Manual Testing

Hardware: 4-node Apple Silicon cluster, M5 Max master.

- Loaded \`mlx-community/GLM-4.7-Air-mlx-4bit\`, ran chat completion via
\`/v1/chat/completions\`. Before this fix the assistant turn ran on into
a synthetic \`<|user|>\` continuation; after the fix the response stops
cleanly at the assistant boundary.

---------

Co-authored-by: jw-wcv <101585096+jw-wcv@users.noreply.github.com>
Co-authored-by: Evan Quiney <evanev7@gmail.com>
2026-05-10 17:02:22 +00:00
Andrei Cravtov 45df74ba98 Andrei/mp capture stdio (#2056)
## Motivation

Process-isolated runner crashes and C-extension failures can write
directly to fd-level stdout/stderr, bypassing Python/loguru. We need to
capture that output per runner process without polluting the main
process or other workers, and without breaking operation when the parent
stdio is detached.

## Changes

- Added `AsyncProcess`, a spawn-only multiprocessing wrapper that
redirects child stdout/stderr to pipes and exposes them as in-memory
`Receiver[bytes]`s
- Replaced runner-supervisor's raw `multiprocessing.Process` usage with
`AsyncProcess`
- Added `--no-stdio`, redirecting stdin/stdout/stderr to `/dev/null`
after logging is configured
- Disabled verbose MLX
- Added tests covering stdio capture, child crashes, repeated bad
children, SIGTERM/SIGKILL shutdown escalation, stdio detachment, and
spawning captured children from a stdio-detached parent

## Why It Works

The parent can redirect its own stdio fds to `/dev/null`, while
`AsyncProcess` installs fresh pipe fds over fd 1 and 2 inside each
spawned child. That keeps stdio-detached parents quiet while preserving
per-runner stdout/stderr capture. Runner shutdown is still bounded:
SIGTERM grace first, then SIGKILL escalation if needed.

Next direction: the runner supervisor currently drains captured output
and logs it as stdout/debug and stderr/warning. This should be split
into more useful process-isolated error reporting instead of just log
forwarding (regex match on errors to obtain "reason" string, best
effort).

## Test Plan

### Manual Testing

Ran on 4 Mac Minis in a Thunderbolt 4 ring, can see that runner's
stdout/stderr contents are being captured.

### Automated Testing

- Added async-process tests for fd-level stdout/stderr capture, Python
traceback capture, bounded-buffer output, child `exit`/abort, parent
stdio preservation, fd leak checks, spawn-context mp channels, and
SIGTERM/SIGKILL shutdown behavior
- Added stdio-detach tests proving stdio detaches to `/dev/null`, a
stdio-detached parent can still spawn and capture a child, and the same
stdio-detached parent can spawn/capture multiple children sequentially
- Updated runner-supervisor tests for the new `AsyncProcess.exitcode`
path
2026-05-09 22:45:14 +01:00
Kerollos Magdy ce37bdceb6 fix: Create directory for PID file if it doesn't exist (#2075)
Ensure the directory for the PID file exists before creating it.

## Motivation

Fixes https://github.com/exo-explore/exo/issues/2074

## Changes

<!-- Describe what you changed in detail -->

## Why It Works

<!-- Explain why your approach solves the problem -->

## Test Plan

### Manual Testing
<!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB,
connected via Thunderbolt 4) -->
<!-- What you did: -->
<!-- - -->

### Automated Testing
<!-- Describe changes to automated tests, or how existing tests cover
this change -->
<!-- - -->
2026-05-09 12:10:22 +00:00
Andrei Cravtov e5a1e5dadb Create PID file locking for EXO (#2072)
## Motivation

EXO should be PID file locked, to prevent duplicate processes from
clobbering the log, right now this isn't the case.

## Changes

I added a wrapper around a Rust PID file lock library, and used it to
implement PID locking for EXO, with the PID file being in exo cache
directory.

## Test Plan

### Manual Testing
Tested on e11, trying to spawn duplicate EXO processes prevented.
2026-05-08 18:50:18 +01:00
ciaranbor fa57131374 Integration tests infra (#1995)
## Motivation

No automated integration tests exist for exo. Manual testing against
real hardware clusters is slow and error-prone. We need a pytest
framework that deploys clusters via `eco`, runs inference scenarios, and
tears down cleanly.

## Changes

- **`tools/src/exo_tools/`** — New workspace member shared by bench,
eval, and tests:
- `client.py` — `ExoClient` HTTP client (extracted from
`bench/harness.py`)
- `harness.py` — instance lifecycle helpers (placement, wait-for-ready,
etc.)
- `cluster.py` — `EcoSession` for eco cluster lifecycle
(deploy/stop/start/release/logs/exec) with unique `USER=<prefix>-<uuid>`
per session and atexit/signal cleanup
- **`tests/integration/`** — 17 pytest tests across 5 files:
- `test_1node.py` — place, chat, multi-turn, delete, state/models
endpoints, cluster snapshot, download-from-scratch
- `test_2node.py` — parametrized tensor/jaccl + pipeline/ring inference
and multi-turn
- `test_4node.py` — parametrized 4-node pipeline/ring inference, cluster
state
- `test_resilience.py` — full disconnect/reconnect cycle (2-node →
disconnect → 1-node → reconnect → 2-node)
- `test_dashboard.py` — Playwright: dashboard loads, shows node info,
chat flow
- `helpers.py` — placement/inference helpers, re-exports from
`exo_tools`
- `conftest.py` — session-scoped cluster fixtures with constraint-based
eco reservations; `--hosts` override; `EXO_REF` env var for CI
deployments from a GitHub branch
- **`bench/`** — Updated imports from `exo_tools.client` /
`exo_tools.harness`
- **`pyproject.toml`** — Added `tools` workspace member, `playwright`
dev dep, `--ignore=tests/integration`

## Why It Works

Tests use `eco` for cluster lifecycle and `ExoClient` for API
interactions — same tools humans use. Session-scoped fixtures deploy
once per file. Unique eco users prevent test runs from interfering with
each other or manual usage.

## Test Plan

### Automated Testing

- `uv run pytest tests/integration/ -v -s` — full suite (~4-5 min, 17/17
passing)
- `uv run pytest tests/integration/ -v -s --hosts s4,s9,s10,s22` — pin
specific hosts
- `EXO_REF=main uv run pytest tests/integration/ -v` — deploy from a
GitHub branch (CI)
- `uv run pytest` — confirms integration tests are excluded from default
runs
2026-05-08 17:15:08 +01:00
Alex Cheema 414132ae9c Use time-weighted power sampling (#2038)
## Why

The power sampler currently averages sampled wattage values
arithmetically. That can be materially wrong when sample intervals are
uneven: a short high-power spike gets the same weight as a long steady
interval. Energy should be computed by integrating power over time, and
average power should be derived from energy / elapsed time.

## How

- Store each power sample with its relative timestamp.
- Anchor the first sample at `t=0` and take a final sample at `elapsed`
when producing results.
- Integrate per-node power using the trapezoidal rule.
- Sum node energy for total cluster energy, then derive total average
system power from total energy / elapsed.
- Add focused unit tests for uneven sample intervals and the
single-sample fallback.

## Tests

- `uv run pytest src/exo/utils/tests/test_power_sampler.py`
- `uv run basedpyright`
- `uv run ruff check src/exo/utils/power_sampler.py
src/exo/utils/tests/test_power_sampler.py`
- `nix fmt`
2026-05-07 10:42:14 +00:00
589 changed files with 15693 additions and 6198 deletions

No files matched your search

-7
View File
@@ -1,8 +1 @@
use flake
# creates .venv if doesn't exist and loads its environment
export VIRTUAL_ENV=".venv"
if ! [ -d "./$VIRTUAL_ENV" ]; then
uv venv
fi
layout python
+1 -1
View File
@@ -34,7 +34,7 @@ jobs:
SPARKLE_S3_PREFIX: ${{ secrets.SPARKLE_S3_PREFIX }}
AWS_REGION: ${{ secrets.AWS_REGION }}
EXO_BUILD_NUMBER: ${{ github.run_number }}
EXO_LIBP2P_NAMESPACE: ${{ github.ref_name }}
EXO_NAMESPACE: ${{ github.ref_name }}
steps:
# ============================================================
+4 -1
View File
@@ -18,7 +18,6 @@ digest.txt
app/EXO/build/
dist/
# rust
target/
**/*.rs.bk
@@ -39,4 +38,8 @@ bench/**/*.json
# tmp
tmp/models
/build/exo
/.agents
/.claude/skills
/.claude
/.codex
skills-lock.json
+3
View File
@@ -4,4 +4,7 @@
<option name="sdkName" value="Python 3.13 (exo)" />
</component>
<component name="ProjectRootManager" version="2" project-jdk-name="Python 3.13 (exo)" project-jdk-type="Python SDK" />
<component name="RuffConfiguration">
<option name="enabled" value="true" />
</component>
</project>
File renamed without changes.
File renamed without changes.
File renamed without changes.
File renamed without changes.
Loaded 100 of 589 files, more files were not shown because too many files have changed in this diff. Show more