Merge remote-tracking branch 'origin/main' into alexcheema/prefill-progress

chore: fix formatting after merge
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-02-18 14:55:13 -05:00 · 2026-02-13 09:37:55 -08:00 · 2026-02-13 06:54:18 -08:00 · 2026-02-13 05:58:45 -08:00 · 2026-02-05 05:38:14 -08:00 · 2026-02-05 05:30:22 -08:00
44 changed files with 1159 additions and 1624 deletions
--- a/.github/workflows/pipeline.yml
+++ b/.github/workflows/pipeline.yml
@@ -8,6 +8,33 @@ on:
      - main

 jobs:
+  typecheck:
+    runs-on: ubuntu-latest
+    steps:
+      - name: Checkout repository
+        uses: actions/checkout@v4
+        with:
+          lfs: false
+
+      - uses: cachix/install-nix-action@v31
+        with:
+          nix_path: nixpkgs=channel:nixos-unstable
+
+      - uses: cachix/cachix-action@v14
+        name: Configure Cachix
+        with:
+          name: exo
+          authToken: "${{ secrets.CACHIX_AUTH_TOKEN }}"
+
+      - name: Load nix develop environment
+        run: nix run github:nicknovitski/nix-develop/v1
+
+      - name: Sync dependencies
+        run: uv sync --all-packages
+
+      - name: Run type checker
+        run: uv run basedpyright --project pyproject.toml
+
  nix:
    name: Build and check (${{ matrix.system }})
    runs-on: ${{ matrix.runner }}
--- a/MISSED_THINGS.md
+++ b/MISSED_THINGS.md
@@ -5,21 +5,21 @@
 [X] Fetching download status of all models on start
 [X] Deduplication of tasks in plan_step.
 [X] resolve_allow_patterns should just be wildcard now.
-[X] no mx_barrier in genreate.py mlx_generate at the end.
+[] no mx_barrier in genreate.py mlx_generate at the end.
 [] cache assertion not needed in auto_parallel.py PipelineLastLayer.
-[X] GPTOSS support dropped in auto_parallel.py.
-[X] sharding changed "all-to-sharded" became _all_to_sharded in auto_parallel.py.
-[X] same as above with "sharded-to-all" became _sharded_to_all in auto_parallel.py.
-[X] Dropped support for Ministral3Model, DeepseekV32Model, Glm4MoeModel, Qwen3NextModel, GptOssMode in auto_parallel.py.
+[] GPTOSS support dropped in auto_parallel.py.
+[] sharding changed "all-to-sharded" became _all_to_sharded in auto_parallel.py.
+[] same as above with "sharded-to-all" became _sharded_to_all in auto_parallel.py.
+[] Dropped support for Ministral3Model, DeepseekV32Model, Glm4MoeModel, Qwen3NextModel, GptOssMode in auto_parallel.py.
 [] Dropped prefill/decode code in auto_parallel.py and utils_mlx.py.
 [X] KV_CACHE_BITS should be None to disable quantized KV cache.
-[X] Dropped _set_nofile_limit in utils_mlx.py.
-[X] We have group optional in load_mlx_items in utils_mlx.py.
-[X] Dropped add_missing_chat_templates for GptOss in load_mlx_items in utils_mlx.py.
-[X] Dropped model.make_cache in make_kv_cache in utils_mlx.py.
+[] Dropped _set_nofile_limit in utils_mlx.py.
+[] We have group optional in load_mlx_items in utils_mlx.py.
+[] Dropped add_missing_chat_templates for GptOss in load_mlx_items in utils_mlx.py.
+[] Dropped model.make_cache in make_kv_cache in utils_mlx.py.
 [X] We put cache limit back in utils_mlx.py.
-[X] topology.py remove_node removes the connections after checking if node is is in self._node_id_to_rx_id_map. on beta_1 it checks after, so would remove stale connections I guess?
-[X] Missing Glm 4.7 model cards (this isn't ready yet but should be picked up, probably create an issue... the blocker is transforemrs version doesn't support the tokenizer for Glm 4.7. rc-1 does but we can't upgrade as it breaks other things.)
+[] topology.py remove_node removes the connections after checking if node is is in self._node_id_to_rx_id_map. on beta_1 it checks after, so would remove stale connections I guess?
+[] Missing Glm 4.7 model cards (this isn't ready yet but should be picked up, probably create an issue... the blocker is transforemrs version doesn't support the tokenizer for Glm 4.7. rc-1 does but we can't upgrade as it breaks other things.)
 [] try-except in _command_processor only excepts ValueError. This was silently failing leading to un-debuggable errors (we had a KeyError that was happening ). Changed this to catch Exception instead of ValueError. See exo-v2 89ae38405e0052e3c22405daf094b065878aa873 and fb99fea69b5a39017efc90c5dad0072e677455f0.
 [X] In placement.py, place_instance no longer looks at model_meta.supports_tensor and check if this tensor parallel number of nodes is supported by the model's tensor dimensions.
 [X] In placement.py, place_instanec, we no longer have the special case to exclude DeepSeek v3.1 pipeline parallel (it doesn't work).
--- a/app/EXO/EXO/ExoProcessController.swift
+++ b/app/EXO/EXO/ExoProcessController.swift
@@ -126,37 +126,11 @@ final class ExoProcessController: ObservableObject {
            return
        }
        process.terminationHandler = nil
-        status = .stopped
-
-        guard process.isRunning else {
-            self.process = nil
-            return
+        if process.isRunning {
+            process.terminate()
        }
-
-        let proc = process
        self.process = nil
-
-        Task.detached {
-            proc.interrupt()
-
-            for _ in 0..<50 {
-                if !proc.isRunning { return }
-                try? await Task.sleep(nanoseconds: 100_000_000)
-            }
-
-            if proc.isRunning {
-                proc.terminate()
-            }
-
-            for _ in 0..<30 {
-                if !proc.isRunning { return }
-                try? await Task.sleep(nanoseconds: 100_000_000)
-            }
-
-            if proc.isRunning {
-                kill(proc.processIdentifier, SIGKILL)
-            }
-        }
+        status = .stopped
    }

    func restart() {
--- a/bench/bench.toml
+++ b/bench/bench.toml
@@ -1,7 +0,0 @@
-# Canary benchmark manifest
-#
-# Lists the suite files to include. Each file defines benchmarks
-# with shared constraints, topology, and default args.
-include = [
-    "single-m3-ultra.toml",
-]
--- a/bench/exo_bench.py
+++ b/bench/exo_bench.py
@@ -288,151 +288,6 @@ def resolve_model_short_id(client: ExoClient, model_arg: str) -> tuple[str, str]
    raise ValueError(f"Model not found in /models: {model_arg}")


-def run_planning_phase(
-    client: ExoClient,
-    full_model_id: str,
-    preview: dict[str, Any],
-    danger_delete: bool,
-    timeout: float,
-    settle_deadline: float | None,
-) -> None:
-    """Check disk space and ensure model is downloaded before benchmarking."""
-    # Get model size from /models
-    models = client.request_json("GET", "/models") or {}
-    model_bytes = 0
-    for m in models.get("data", []):
-        if m.get("hugging_face_id") == full_model_id:
-            model_bytes = m.get("storage_size_megabytes", 0) * 1024 * 1024
-            break
-
-    if not model_bytes:
-        logger.warning(
-            f"Could not determine size for {full_model_id}, skipping disk check"
-        )
-        return
-
-    # Get nodes from preview
-    inner = unwrap_instance(preview["instance"])
-    node_ids = list(inner["shardAssignments"]["nodeToRunner"].keys())
-    runner_to_shard = inner["shardAssignments"]["runnerToShard"]
-
-    state = client.request_json("GET", "/state")
-    downloads = state.get("downloads", {})
-    node_disk = state.get("nodeDisk", {})
-
-    for node_id in node_ids:
-        node_downloads = downloads.get(node_id, [])
-
-        # Check if model already downloaded on this node
-        already_downloaded = any(
-            "DownloadCompleted" in p
-            and unwrap_instance(p["DownloadCompleted"]["shardMetadata"])["modelCard"][
-                "modelId"
-            ]
-            == full_model_id
-            for p in node_downloads
-        )
-        if already_downloaded:
-            continue
-
-        # Wait for disk info if settle_deadline is set
-        disk_info = node_disk.get(node_id, {})
-        backoff = _SETTLE_INITIAL_BACKOFF_S
-        while not disk_info and settle_deadline and time.monotonic() < settle_deadline:
-            remaining = settle_deadline - time.monotonic()
-            logger.info(
-                f"Waiting for disk info on {node_id} ({remaining:.0f}s remaining)..."
-            )
-            time.sleep(min(backoff, remaining))
-            backoff = min(backoff * _SETTLE_BACKOFF_MULTIPLIER, _SETTLE_MAX_BACKOFF_S)
-            state = client.request_json("GET", "/state")
-            node_disk = state.get("nodeDisk", {})
-            disk_info = node_disk.get(node_id, {})
-
-        if not disk_info:
-            logger.warning(f"No disk info for {node_id}, skipping space check")
-            continue
-
-        avail = disk_info.get("available", {}).get("inBytes", 0)
-        if avail >= model_bytes:
-            continue
-
-        if not danger_delete:
-            raise RuntimeError(
-                f"Insufficient disk on {node_id}: need {model_bytes // (1024**3)}GB, "
-                f"have {avail // (1024**3)}GB. Use --danger-delete-downloads to free space."
-            )
-
-        # Delete from smallest to largest
-        completed = [
-            (
-                unwrap_instance(p["DownloadCompleted"]["shardMetadata"])["modelCard"][
-                    "modelId"
-                ],
-                p["DownloadCompleted"]["totalBytes"]["inBytes"],
-            )
-            for p in node_downloads
-            if "DownloadCompleted" in p
-        ]
-        for del_model, size in sorted(completed, key=lambda x: x[1]):
-            logger.info(f"Deleting {del_model} from {node_id} ({size // (1024**2)}MB)")
-            client.request_json("DELETE", f"/download/{node_id}/{del_model}")
-            avail += size
-            if avail >= model_bytes:
-                break
-
-        if avail < model_bytes:
-            raise RuntimeError(f"Could not free enough space on {node_id}")
-
-    # Start downloads (idempotent)
-    for node_id in node_ids:
-        runner_id = inner["shardAssignments"]["nodeToRunner"][node_id]
-        shard = runner_to_shard[runner_id]
-        client.request_json(
-            "POST",
-            "/download/start",
-            body={
-                "targetNodeId": node_id,
-                "shardMetadata": shard,
-            },
-        )
-        logger.info(f"Started download on {node_id}")
-
-    # Wait for downloads
-    start = time.time()
-    while time.time() - start < timeout:
-        state = client.request_json("GET", "/state")
-        downloads = state.get("downloads", {})
-        all_done = True
-        for node_id in node_ids:
-            done = any(
-                "DownloadCompleted" in p
-                and unwrap_instance(p["DownloadCompleted"]["shardMetadata"])[
-                    "modelCard"
-                ]["modelId"]
-                == full_model_id
-                for p in downloads.get(node_id, [])
-            )
-            failed = [
-                p["DownloadFailed"]["errorMessage"]
-                for p in downloads.get(node_id, [])
-                if "DownloadFailed" in p
-                and unwrap_instance(p["DownloadFailed"]["shardMetadata"])["modelCard"][
-                    "modelId"
-                ]
-                == full_model_id
-            ]
-            if failed:
-                raise RuntimeError(f"Download failed on {node_id}: {failed[0]}")
-            if not done:
-                all_done = False
-        if all_done:
-            return
-        time.sleep(1)
-
-    raise TimeoutError("Downloads did not complete in time")
-
-
 def placement_filter(instance_meta: str, wanted: str) -> bool:
    s = (instance_meta or "").lower()
    if wanted == "both":
@@ -680,11 +535,6 @@ def main() -> int:
        default=0,
        help="Max seconds to wait for the cluster to produce valid placements (0 = try once).",
    )
-    ap.add_argument(
-        "--danger-delete-downloads",
-        action="store_true",
-        help="Delete existing models from smallest to largest to make room for benchmark model.",
-    )
    args = ap.parse_args()

    pp_list = parse_int_list(args.pp)
@@ -719,16 +569,13 @@ def main() -> int:
        logger.error("[exo-bench] tokenizer usable but prompt sizing failed")
        raise

-    settle_deadline = (
-        time.monotonic() + args.settle_timeout if args.settle_timeout > 0 else None
-    )
-
    selected = fetch_and_filter_placements(client, full_model_id, args)

-    if not selected and settle_deadline:
+    if not selected and args.settle_timeout > 0:
        backoff = _SETTLE_INITIAL_BACKOFF_S
-        while not selected and time.monotonic() < settle_deadline:
-            remaining = settle_deadline - time.monotonic()
+        deadline = time.monotonic() + args.settle_timeout
+        while not selected and time.monotonic() < deadline:
+            remaining = deadline - time.monotonic()
            logger.warning(
                f"No valid placements yet (cluster may still be settling). "
                f"Retrying in {backoff:.1f}s ({remaining:.0f}s remaining)..."
@@ -760,16 +607,6 @@ def main() -> int:
    if args.dry_run:
        return 0

-    logger.info("Planning phase: checking downloads...")
-    run_planning_phase(
-        client,
-        full_model_id,
-        selected[0],
-        args.danger_delete_downloads,
-        args.timeout,
-        settle_deadline,
-    )
-
    all_rows: list[dict[str, Any]] = []

    for preview in selected:
--- a/bench/single-m3-ultra.toml
+++ b/bench/single-m3-ultra.toml
@@ -1,189 +0,0 @@
-# Single-node M3 Ultra benchmarks
-#
-# Shared constraints applied to ALL benchmarks in this file.
-constraints = [
-    "All(MacOsBuild(=25D125))",
-    "Hosts(=1)",
-    "All(Chip(m3_ultra))",
-    "All(GpuCores(=80))",
-]
-
-[topology]
-type = "none"
-
-# Default args merged into each benchmark's args (benchmark-level args win).
-[defaults]
-pp = [512, 2048, 8192, 16384]
-tg = 128
-
-[[benchmark]]
-model = "mlx-community/Meta-Llama-3.1-70B-Instruct-4bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/gpt-oss-120b-MXFP4-Q8"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/GLM-4.7-Flash-8bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-Coder-Next-6bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-30B-A3B-8bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-0.6B-4bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-0.6B-8bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Llama-3.2-1B-Instruct-4bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Llama-3.2-3B-Instruct-4bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Llama-3.2-3B-Instruct-8bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Meta-Llama-3.1-8B-Instruct-4bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Meta-Llama-3.1-8B-Instruct-8bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Meta-Llama-3.1-8B-Instruct-bf16"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/gpt-oss-20b-MXFP4-Q8"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-30B-A3B-4bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/GLM-4.7-Flash-4bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/GLM-4.7-Flash-5bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/GLM-4.7-Flash-6bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Llama-3.3-70B-Instruct-4bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-Coder-Next-4bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-Coder-Next-5bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-Coder-Next-8bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-Next-80B-A3B-Instruct-4bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-Next-80B-A3B-Instruct-8bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-Next-80B-A3B-Thinking-4bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-Next-80B-A3B-Thinking-8bit"
-extra_constraints = ["All(Memory(>=96GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Llama-3.3-70B-Instruct-8bit"
-extra_constraints = ["All(Memory(>=256GiB))"]
-
-[[benchmark]]
-model = "mlx-community/llama-3.3-70b-instruct-fp16"
-extra_constraints = ["All(Memory(>=256GiB))"]
-
-[[benchmark]]
-model = "mlx-community/GLM-4.5-Air-8bit"
-extra_constraints = ["All(Memory(>=256GiB))"]
-
-[[benchmark]]
-model = "mlx-community/GLM-4.5-Air-bf16"
-extra_constraints = ["All(Memory(>=256GiB))"]
-
-[[benchmark]]
-model = "mlx-community/GLM-4.7-4bit"
-extra_constraints = ["All(Memory(>=256GiB))"]
-
-[[benchmark]]
-model = "mlx-community/MiniMax-M2.1-3bit"
-extra_constraints = ["All(Memory(>=256GiB))"]
-
-[[benchmark]]
-model = "mlx-community/MiniMax-M2.1-8bit"
-extra_constraints = ["All(Memory(>=256GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-235B-A22B-Instruct-2507-4bit"
-extra_constraints = ["All(Memory(>=256GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-Coder-Next-bf16"
-extra_constraints = ["All(Memory(>=256GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Step-3.5-Flash-4bit"
-extra_constraints = ["All(Memory(>=256GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Step-3.5-Flash-6bit"
-extra_constraints = ["All(Memory(>=256GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Step-3.5-Flash-8Bit"
-extra_constraints = ["All(Memory(>=256GiB))"]
-
-[[benchmark]]
-model = "mlx-community/DeepSeek-V3.1-4bit"
-extra_constraints = ["All(Memory(>=512GiB))"]
-
-[[benchmark]]
-model = "mlx-community/GLM-4.7-6bit"
-extra_constraints = ["All(Memory(>=512GiB))"]
-
-[[benchmark]]
-model = "mlx-community/GLM-4.7-8bit-gs32"
-extra_constraints = ["All(Memory(>=512GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-235B-A22B-Instruct-2507-8bit"
-extra_constraints = ["All(Memory(>=512GiB))"]
-
-[[benchmark]]
-model = "mlx-community/Qwen3-Coder-480B-A35B-Instruct-4bit"
-extra_constraints = ["All(Memory(>=512GiB))"]
--- a/dashboard/src/lib/components/ChatForm.svelte
+++ b/dashboard/src/lib/components/ChatForm.svelte
@@ -265,7 +265,6 @@

  function handleSubmit() {
    if ((!message.trim() && uploadedFiles.length === 0) || loading) return;
-    if (isEditOnlyWithoutImage) return;

    const content = message.trim();
    const files = [...uploadedFiles];
@@ -290,11 +289,7 @@
      if (imageFile.preview) {
        editImage(content, imageFile.preview);
      }
-    } else if (
-      currentModel &&
-      modelSupportsTextToImage(currentModel) &&
-      content
-    ) {
+    } else if (isImageModel() && content) {
      // Use image generation for text-to-image models
      generateImage(content);
    } else {
--- a/dashboard/src/lib/components/ChatMessages.svelte
+++ b/dashboard/src/lib/components/ChatMessages.svelte
@@ -9,7 +9,6 @@
    regenerateFromToken,
    setEditingImage,
  } from "$lib/stores/app.svelte";
-  import type { Message } from "$lib/stores/app.svelte";
  import type { MessageAttachment } from "$lib/stores/app.svelte";
  import MarkdownContent from "./MarkdownContent.svelte";
  import TokenHeatmap from "./TokenHeatmap.svelte";
@@ -225,7 +224,6 @@
  }

  function handleDeleteClick(messageId: string) {
-    if (loading) return;
    deleteConfirmId = messageId;
  }

@@ -256,7 +254,7 @@
 </script>

 <div class="flex flex-col gap-4 sm:gap-6 {className}">
-  {#each messageList as message, i (message.id)}
+  {#each messageList as message (message.id)}
    <div
      class="group flex {message.role === 'user'
        ? 'justify-end'
@@ -318,11 +316,9 @@
          <!-- Delete confirmation -->
          <div class="bg-red-500/10 border border-red-500/30 rounded-lg p-3">
            <p class="text-xs text-red-400 mb-3">
-              {#if i === messageList.length - 1}
-                Delete this message?
-              {:else}
-                Delete this message and all messages after it?
-              {/if}
+              Delete this message{message.role === "user"
+                ? " and all responses after it"
+                : ""}?
            </p>
            <div class="flex gap-2 justify-end">
              <button
@@ -754,13 +750,8 @@
            <!-- Delete button -->
            <button
              onclick={() => handleDeleteClick(message.id)}
-              disabled={loading}
-              class="p-1.5 transition-colors rounded {loading
-                ? 'text-exo-light-gray/30 cursor-not-allowed'
-                : 'text-exo-light-gray hover:text-red-400 hover:bg-red-500/10 cursor-pointer'}"
-              title={loading
-                ? "Cannot delete while generating"
-                : "Delete message"}
+              class="p-1.5 text-exo-light-gray hover:text-red-400 transition-colors rounded hover:bg-red-500/10 cursor-pointer"
+              title="Delete message"
            >
              <svg
                class="w-3.5 h-3.5"
--- a/dashboard/src/lib/components/PrefillProgressBar.svelte
+++ b/dashboard/src/lib/components/PrefillProgressBar.svelte
@@ -0,0 +1,51 @@
+<script lang="ts">
+  import type { PrefillProgress } from "$lib/stores/app.svelte";
+
+  interface Props {
+    progress: PrefillProgress;
+    class?: string;
+  }
+
+  let { progress, class: className = "" }: Props = $props();
+
+  const percentage = $derived(
+    progress.total > 0
+      ? Math.round((progress.processed / progress.total) * 100)
+      : 0,
+  );
+
+  function formatTokenCount(count: number): string {
+    if (count >= 1000) {
+      return `${(count / 1000).toFixed(1)}k`;
+    }
+    return count.toString();
+  }
+</script>
+
+<div class="prefill-progress {className}">
+  <div
+    class="flex items-center justify-between text-xs text-exo-light-gray mb-1"
+  >
+    <span>Processing prompt</span>
+    <span class="font-mono">
+      {formatTokenCount(progress.processed)} / {formatTokenCount(
+        progress.total,
+      )} tokens
+    </span>
+  </div>
+  <div class="h-1.5 bg-exo-black/60 rounded-full overflow-hidden">
+    <div
+      class="h-full bg-exo-yellow rounded-full transition-all duration-150 ease-out"
+      style="width: {percentage}%"
+    ></div>
+  </div>
+  <div class="text-right text-xs text-exo-light-gray/70 mt-0.5 font-mono">
+    {percentage}%
+  </div>
+</div>
+
+<style>
+  .prefill-progress {
+    width: 100%;
+  }
+</style>
--- a/dashboard/src/lib/stores/app.svelte.ts
+++ b/dashboard/src/lib/stores/app.svelte.ts
@@ -273,6 +273,11 @@ export interface TokenData {
  topLogprobs: TopLogprob[];
 }

+export interface PrefillProgress {
+  processed: number;
+  total: number;
+}
+
 export interface Message {
  id: string;
  role: "user" | "assistant" | "system";
--- a/dashboard/src/routes/downloads/+page.svelte
+++ b/dashboard/src/routes/downloads/+page.svelte
--- a/pyproject.toml
+++ b/pyproject.toml
@@ -132,7 +132,7 @@ markers = [
 env = [
  "EXO_TESTS=1"
 ]
-addopts = "-m 'not slow' --ignore=tests/start_distributed_test.py"
+addopts = "-m 'not slow'"
 filterwarnings = [
    "ignore:builtin type Swig:DeprecationWarning",
 ]
--- a/python/parts.nix
+++ b/python/parts.nix
@@ -14,9 +14,7 @@

      # Override overlay to inject Nix-built components
      exoOverlay = final: prev: {
-        # Replace workspace exo_pyo3_bindings with Nix-built wheel.
-        # Preserve passthru so mkVirtualEnv can resolve dependency groups.
-        # Copy .pyi stub + py.typed marker so basedpyright can find the types.
+        # Replace workspace exo_pyo3_bindings with Nix-built wheel
        exo-pyo3-bindings = pkgs.stdenv.mkDerivation {
          pname = "exo-pyo3-bindings";
          version = "0.1.0";
@@ -24,12 +22,6 @@
          # Install from pre-built wheel
          nativeBuildInputs = [ final.pyprojectWheelHook ];
          dontStrip = true;
-          passthru = prev.exo-pyo3-bindings.passthru or { };
-          postInstall = ''
-            local siteDir=$out/${final.python.sitePackages}/exo_pyo3_bindings
-            cp ${inputs.self}/rust/exo_pyo3_bindings/exo_pyo3_bindings.pyi $siteDir/
-            touch $siteDir/py.typed
-          '';
        };
      };

@@ -37,32 +29,17 @@

      # Overlay to provide build systems and custom packages
      buildSystemsOverlay = final: prev: {
+        # Use our pure Nix-built MLX with Metal support
+        mlx = self'.packages.mlx;
+
        # mlx-lm is a git dependency that needs setuptools
        mlx-lm = prev.mlx-lm.overrideAttrs (old: {
          nativeBuildInputs = (old.nativeBuildInputs or [ ]) ++ [
            final.setuptools
          ];
        });
-      } // lib.optionalAttrs pkgs.stdenv.hostPlatform.isDarwin {
-        # Use our pure Nix-built MLX with Metal support (macOS only)
-        mlx = self'.packages.mlx;
      };

-      # Additional overlay for Linux-specific fixes (type checking env).
-      # Native wheels have shared lib dependencies we don't need at type-check time.
-      linuxOverlay = final: prev:
-        let
-          ignoreMissing = drv: drv.overrideAttrs { autoPatchelfIgnoreMissingDeps = [ "*" ]; };
-          nvidiaPackages = lib.filterAttrs (name: _: lib.hasPrefix "nvidia-" name) prev;
-        in
-        lib.optionalAttrs pkgs.stdenv.hostPlatform.isLinux (
-          (lib.mapAttrs (_: ignoreMissing) nvidiaPackages) // {
-            mlx = ignoreMissing prev.mlx;
-            torch = ignoreMissing prev.torch;
-            triton = ignoreMissing prev.triton;
-          }
-        );
-
      pythonSet = (pkgs.callPackage inputs.pyproject-nix.build.packages {
        inherit python;
      }).overrideScope (
@@ -71,7 +48,6 @@
          overlay
          exoOverlay
          buildSystemsOverlay
-          linuxOverlay
        ]
      );
      exoVenv = pythonSet.mkVirtualEnv "exo-env" workspace.deps.default;
@@ -142,21 +118,6 @@
          ${pkgs.ruff}/bin/ruff check ${inputs.self}
          touch $out
        '';
-
-        # Hermetic basedpyright type checking
-        typecheck = pkgs.runCommand "typecheck"
-          {
-            nativeBuildInputs = [
-              testVenv
-              pkgs.basedpyright
-            ];
-          }
-          ''
-            cd ${inputs.self}
-            export HOME=$TMPDIR
-            basedpyright --pythonpath ${testVenv}/bin/python
-            touch $out
-          '';
      };
    };
 }
--- a/src/exo/download/coordinator.py
+++ b/src/exo/download/coordinator.py
@@ -14,7 +14,6 @@ from exo.download.download_utils import (
    map_repo_download_progress_to_download_progress_data,
 )
 from exo.download.shard_downloader import ShardDownloader
-from exo.shared.constants import EXO_MODELS_DIR
 from exo.shared.models.model_cards import ModelId
 from exo.shared.types.commands import (
    CancelDownload,
@@ -64,9 +63,6 @@ class DownloadCoordinator:
        self.event_sender, self.event_receiver = channel[Event]()
        self.shard_downloader.on_progress(self._download_progress_callback)

-    def _model_dir(self, model_id: ModelId) -> str:
-        return str(EXO_MODELS_DIR / model_id.normalize())
-
    async def _download_progress_callback(
        self, callback_shard: ShardMetadata, progress: RepoDownloadProgress
    ) -> None:
@@ -78,7 +74,6 @@ class DownloadCoordinator:
                shard_metadata=callback_shard,
                node_id=self.node_id,
                total_bytes=progress.total_bytes,
-                model_directory=self._model_dir(model_id),
            )
            self.download_status[model_id] = completed
            await self.event_sender.send(
@@ -98,7 +93,6 @@ class DownloadCoordinator:
                download_progress=map_repo_download_progress_to_download_progress_data(
                    progress
                ),
-                model_directory=self._model_dir(model_id),
            )
            self.download_status[model_id] = ongoing
            await self.event_sender.send(
@@ -176,11 +170,7 @@ class DownloadCoordinator:
                return

        # Emit pending status
-        progress = DownloadPending(
-            shard_metadata=shard,
-            node_id=self.node_id,
-            model_directory=self._model_dir(model_id),
-        )
+        progress = DownloadPending(shard_metadata=shard, node_id=self.node_id)
        self.download_status[model_id] = progress
        await self.event_sender.send(NodeDownloadProgress(download_progress=progress))

@@ -194,7 +184,6 @@ class DownloadCoordinator:
                shard_metadata=shard,
                node_id=self.node_id,
                total_bytes=initial_progress.total_bytes,
-                model_directory=self._model_dir(model_id),
            )
            self.download_status[model_id] = completed
            await self.event_sender.send(
@@ -217,7 +206,6 @@ class DownloadCoordinator:
            download_progress=map_repo_download_progress_to_download_progress_data(
                initial_progress
            ),
-            model_directory=self._model_dir(model_id),
        )
        self.download_status[model_id] = status
        self.event_sender.send_nowait(NodeDownloadProgress(download_progress=status))
@@ -231,7 +219,6 @@ class DownloadCoordinator:
                    shard_metadata=shard,
                    node_id=self.node_id,
                    error_message=str(e),
-                    model_directory=self._model_dir(model_id),
                )
                self.download_status[model_id] = failed
                await self.event_sender.send(
@@ -266,7 +253,6 @@ class DownloadCoordinator:
            pending = DownloadPending(
                shard_metadata=current_status.shard_metadata,
                node_id=self.node_id,
-                model_directory=self._model_dir(model_id),
            )
            await self.event_sender.send(
                NodeDownloadProgress(download_progress=pending)
@@ -309,18 +295,11 @@ class DownloadCoordinator:
                            node_id=self.node_id,
                            shard_metadata=progress.shard,
                            total_bytes=progress.total_bytes,
-                            model_directory=self._model_dir(
-                                progress.shard.model_card.model_id
-                            ),
                        )
                    elif progress.status in ["in_progress", "not_started"]:
                        if progress.downloaded_bytes_this_session.in_bytes == 0:
                            status = DownloadPending(
-                                node_id=self.node_id,
-                                shard_metadata=progress.shard,
-                                model_directory=self._model_dir(
-                                    progress.shard.model_card.model_id
-                                ),
+                                node_id=self.node_id, shard_metadata=progress.shard
                            )
                        else:
                            status = DownloadOngoing(
@@ -329,9 +308,6 @@ class DownloadCoordinator:
                                download_progress=map_repo_download_progress_to_download_progress_data(
                                    progress
                                ),
-                                model_directory=self._model_dir(
-                                    progress.shard.model_card.model_id
-                                ),
                            )
                    else:
                        continue
--- a/src/exo/main.py
+++ b/src/exo/main.py
@@ -136,8 +136,6 @@ class Node:

    async def run(self):
        async with self._tg as tg:
-            signal.signal(signal.SIGINT, lambda _, __: self.shutdown())
-            signal.signal(signal.SIGTERM, lambda _, __: self.shutdown())
            tg.start_soon(self.router.run)
            tg.start_soon(self.election.run)
            if self.download_coordinator:
@@ -149,6 +147,8 @@ class Node:
            if self.api:
                tg.start_soon(self.api.run)
            tg.start_soon(self._elect_loop)
+            signal.signal(signal.SIGINT, lambda _, __: self.shutdown())
+            signal.signal(signal.SIGTERM, lambda _, __: self.shutdown())

    def shutdown(self):
        # if this is our second call to shutdown, just sys.exit
--- a/src/exo/master/adapters/chat_completions.py
+++ b/src/exo/master/adapters/chat_completions.py
@@ -17,9 +17,13 @@ from exo.shared.types.api import (
    LogprobsContentItem,
    StreamingChoiceResponse,
    ToolCall,
-    Usage,
 )
-from exo.shared.types.chunks import ErrorChunk, TokenChunk, ToolCallChunk
+from exo.shared.types.chunks import (
+    ErrorChunk,
+    PrefillProgressData,
+    TokenChunk,
+    ToolCallChunk,
+)
 from exo.shared.types.common import CommandId
 from exo.shared.types.text_generation import InputMessage, TextGenerationTaskParams

@@ -123,70 +127,71 @@ def chunk_to_response(

 async def generate_chat_stream(
    command_id: CommandId,
-    chunk_stream: AsyncGenerator[ErrorChunk | ToolCallChunk | TokenChunk, None],
+    event_stream: AsyncGenerator[
+        PrefillProgressData | ErrorChunk | ToolCallChunk | TokenChunk, None
+    ],
 ) -> AsyncGenerator[str, None]:
-    """Generate Chat Completions API streaming events from chunks."""
-    last_usage: Usage | None = None
+    """Generate Chat Completions API streaming events from StreamEvents.

-    async for chunk in chunk_stream:
-        if isinstance(chunk, ErrorChunk):
-            error_response = ErrorResponse(
-                error=ErrorInfo(
-                    message=chunk.error_message or "Internal server error",
-                    type="InternalServerError",
-                    code=500,
-                )
-            )
-            yield f"data: {error_response.model_dump_json()}\n\n"
-            yield "data: [DONE]\n\n"
-            return
+    Handles PrefillProgressData, ErrorChunk, ToolCallChunk, and TokenChunk.
+    """
+    async for event in event_stream:
+        match event:
+            case PrefillProgressData():
+                yield f"event: prefill_progress\ndata: {event.model_dump_json()}\n\n"

-        last_usage = chunk.usage or last_usage
-
-        if isinstance(chunk, ToolCallChunk):
-            tool_call_deltas = [
-                ToolCall(
-                    id=tool.id,
-                    index=i,
-                    function=tool,
-                )
-                for i, tool in enumerate(chunk.tool_calls)
-            ]
-            tool_response = ChatCompletionResponse(
-                id=command_id,
-                created=int(time.time()),
-                model=chunk.model,
-                choices=[
-                    StreamingChoiceResponse(
-                        index=0,
-                        delta=ChatCompletionMessage(
-                            role="assistant",
-                            tool_calls=tool_call_deltas,
-                        ),
-                        finish_reason="tool_calls",
+            case ErrorChunk():
+                error_response = ErrorResponse(
+                    error=ErrorInfo(
+                        message=event.error_message or "Internal server error",
+                        type="InternalServerError",
+                        code=500,
                    )
-                ],
-                usage=last_usage,
-            )
-            yield f"data: {tool_response.model_dump_json()}\n\n"
-            yield "data: [DONE]\n\n"
-            return
+                )
+                yield f"data: {error_response.model_dump_json()}\n\n"
+                yield "data: [DONE]\n\n"
+                return

-        chunk_response = chunk_to_response(chunk, command_id)
-        if chunk.finish_reason is not None:
-            chunk_response = chunk_response.model_copy(update={"usage": last_usage})
-        yield f"data: {chunk_response.model_dump_json()}\n\n"
+            case ToolCallChunk():
+                tool_call_deltas = [
+                    ToolCall(
+                        id=tool.id,
+                        index=i,
+                        function=tool,
+                    )
+                    for i, tool in enumerate(event.tool_calls)
+                ]
+                tool_response = ChatCompletionResponse(
+                    id=command_id,
+                    created=int(time.time()),
+                    model=event.model,
+                    choices=[
+                        StreamingChoiceResponse(
+                            index=0,
+                            delta=ChatCompletionMessage(
+                                role="assistant",
+                                tool_calls=tool_call_deltas,
+                            ),
+                            finish_reason="tool_calls",
+                        )
+                    ],
+                )
+                yield f"data: {tool_response.model_dump_json()}\n\n"
+                yield "data: [DONE]\n\n"
+                return

-        if chunk.finish_reason is not None:
-            yield "data: [DONE]\n\n"
+            case TokenChunk():
+                chunk_response = chunk_to_response(event, command_id)
+                yield f"data: {chunk_response.model_dump_json()}\n\n"
+
+                if event.finish_reason is not None:
+                    yield "data: [DONE]\n\n"


 async def collect_chat_response(
    command_id: CommandId,
    chunk_stream: AsyncGenerator[ErrorChunk | ToolCallChunk | TokenChunk, None],
-) -> AsyncGenerator[str]:
-    # This is an AsyncGenerator[str] rather than returning a ChatCompletionReponse because
-    # FastAPI handles the cancellation better but wouldn't auto-serialize for some reason
+) -> ChatCompletionResponse:
    """Collect all token chunks and return a single ChatCompletionResponse."""
    text_parts: list[str] = []
    tool_calls: list[ToolCall] = []
@@ -194,7 +199,6 @@ async def collect_chat_response(
    model: str | None = None
    finish_reason: FinishReason | None = None
    error_message: str | None = None
-    last_usage: Usage | None = None

    async for chunk in chunk_stream:
        if isinstance(chunk, ErrorChunk):
@@ -204,8 +208,6 @@ async def collect_chat_response(
        if model is None:
            model = chunk.model

-        last_usage = chunk.usage or last_usage
-
        if isinstance(chunk, TokenChunk):
            text_parts.append(chunk.text)
            if chunk.logprob is not None:
@@ -236,7 +238,7 @@ async def collect_chat_response(
    combined_text = "".join(text_parts)
    assert model is not None

-    yield ChatCompletionResponse(
+    return ChatCompletionResponse(
        id=command_id,
        created=int(time.time()),
        model=model,
@@ -254,6 +256,4 @@ async def collect_chat_response(
                finish_reason=finish_reason,
            )
        ],
-        usage=last_usage,
-    ).model_dump_json()
-    return
+    )
--- a/src/exo/master/adapters/claude.py
+++ b/src/exo/master/adapters/claude.py
@@ -4,7 +4,7 @@ import json
 from collections.abc import AsyncGenerator
 from typing import Any

-from exo.shared.types.api import FinishReason, Usage
+from exo.shared.types.api import FinishReason
 from exo.shared.types.chunks import ErrorChunk, TokenChunk, ToolCallChunk
 from exo.shared.types.claude_api import (
    ClaudeContentBlock,
@@ -161,14 +161,12 @@ async def collect_claude_response(
    command_id: CommandId,
    model: str,
    chunk_stream: AsyncGenerator[ErrorChunk | ToolCallChunk | TokenChunk, None],
-) -> AsyncGenerator[str]:
-    # This is an AsyncGenerator[str] rather than returning a ChatCompletionReponse because
-    # FastAPI handles the cancellation better but wouldn't auto-serialize for some reason
+) -> ClaudeMessagesResponse:
    """Collect all token chunks and return a single ClaudeMessagesResponse."""
    text_parts: list[str] = []
    tool_use_blocks: list[ClaudeToolUseBlock] = []
    stop_reason: ClaudeStopReason | None = None
-    last_usage: Usage | None = None
+    last_stats = None
    error_message: str | None = None

    async for chunk in chunk_stream:
@@ -176,8 +174,6 @@ async def collect_claude_response(
            error_message = chunk.error_message or "Internal server error"
            break

-        last_usage = chunk.usage or last_usage
-
        if isinstance(chunk, ToolCallChunk):
            for tool in chunk.tool_calls:
                tool_use_blocks.append(
@@ -187,10 +183,12 @@ async def collect_claude_response(
                        input=json.loads(tool.arguments),  # pyright: ignore[reportAny]
                    )
                )
+            last_stats = chunk.stats or last_stats
            stop_reason = "tool_use"
            continue

        text_parts.append(chunk.text)
+        last_stats = chunk.stats or last_stats

        if chunk.finish_reason is not None:
            stop_reason = finish_reason_to_claude_stop_reason(chunk.finish_reason)
@@ -210,11 +208,11 @@ async def collect_claude_response(
    if not content:
        content.append(ClaudeTextBlock(text=""))

-    # Use actual usage data if available
-    input_tokens = last_usage.prompt_tokens if last_usage else 0
-    output_tokens = last_usage.completion_tokens if last_usage else 0
+    # Use actual usage data from stats if available
+    input_tokens = last_stats.prompt_tokens if last_stats else 0
+    output_tokens = last_stats.generation_tokens if last_stats else 0

-    yield ClaudeMessagesResponse(
+    return ClaudeMessagesResponse(
        id=f"msg_{command_id}",
        model=model,
        content=content,
@@ -223,8 +221,7 @@ async def collect_claude_response(
            input_tokens=input_tokens,
            output_tokens=output_tokens,
        ),
-    ).model_dump_json()
-    return
+    )


 async def generate_claude_stream(
@@ -252,7 +249,7 @@ async def generate_claude_stream(

    output_tokens = 0
    stop_reason: ClaudeStopReason | None = None
-    last_usage: Usage | None = None
+    last_stats = None
    next_block_index = 1  # text block is 0, tool blocks start at 1

    async for chunk in chunk_stream:
@@ -260,9 +257,8 @@ async def generate_claude_stream(
            # Close text block and bail
            break

-        last_usage = chunk.usage or last_usage
-
        if isinstance(chunk, ToolCallChunk):
+            last_stats = chunk.stats or last_stats
            stop_reason = "tool_use"

            # Emit tool_use content blocks
@@ -294,6 +290,7 @@ async def generate_claude_stream(
            continue

        output_tokens += 1  # Count each chunk as one token
+        last_stats = chunk.stats or last_stats

        # content_block_delta
        delta_event = ClaudeContentBlockDeltaEvent(
@@ -305,9 +302,9 @@ async def generate_claude_stream(
        if chunk.finish_reason is not None:
            stop_reason = finish_reason_to_claude_stop_reason(chunk.finish_reason)

-    # Use actual token count from usage if available
-    if last_usage is not None:
-        output_tokens = last_usage.completion_tokens
+    # Use actual token count from stats if available
+    if last_stats is not None:
+        output_tokens = last_stats.generation_tokens

    # content_block_stop for text block
    block_stop = ClaudeContentBlockStopEvent(index=0)
--- a/src/exo/master/adapters/responses.py
+++ b/src/exo/master/adapters/responses.py
@@ -4,7 +4,6 @@ from collections.abc import AsyncGenerator
 from itertools import count
 from typing import Any

-from exo.shared.types.api import Usage
 from exo.shared.types.chunks import ErrorChunk, TokenChunk, ToolCallChunk
 from exo.shared.types.common import CommandId
 from exo.shared.types.openai_responses import (
@@ -122,15 +121,13 @@ async def collect_responses_response(
    command_id: CommandId,
    model: str,
    chunk_stream: AsyncGenerator[ErrorChunk | ToolCallChunk | TokenChunk, None],
-) -> AsyncGenerator[str]:
-    # This is an AsyncGenerator[str] rather than returning a ChatCompletionReponse because
-    # FastAPI handles the cancellation better but wouldn't auto-serialize for some reason
+) -> ResponsesResponse:
    """Collect all token chunks and return a single ResponsesResponse."""
    response_id = f"resp_{command_id}"
    item_id = f"item_{command_id}"
    accumulated_text = ""
    function_call_items: list[ResponseFunctionCallItem] = []
-    last_usage: Usage | None = None
+    last_stats = None
    error_message: str | None = None

    async for chunk in chunk_stream:
@@ -138,32 +135,32 @@ async def collect_responses_response(
            error_message = chunk.error_message or "Internal server error"
            break

-        last_usage = chunk.usage or last_usage
-
        if isinstance(chunk, ToolCallChunk):
            for tool in chunk.tool_calls:
                function_call_items.append(
                    ResponseFunctionCallItem(
-                        id=tool.id,
-                        call_id=tool.id,
+                        id=f"fc_{tool.id}",
+                        call_id=f"call_{tool.id}",
                        name=tool.name,
                        arguments=tool.arguments,
                    )
                )
+            last_stats = chunk.stats or last_stats
            continue

        accumulated_text += chunk.text
+        last_stats = chunk.stats or last_stats

    if error_message is not None:
        raise ValueError(error_message)

-    # Create usage from usage data if available
+    # Create usage from stats if available
    usage = None
-    if last_usage is not None:
+    if last_stats is not None:
        usage = ResponseUsage(
-            input_tokens=last_usage.prompt_tokens,
-            output_tokens=last_usage.completion_tokens,
-            total_tokens=last_usage.total_tokens,
+            input_tokens=last_stats.prompt_tokens,
+            output_tokens=last_stats.generation_tokens,
+            total_tokens=last_stats.prompt_tokens + last_stats.generation_tokens,
        )

    output: list[ResponseItem] = [
@@ -175,15 +172,14 @@ async def collect_responses_response(
    ]
    output.extend(function_call_items)

-    yield ResponsesResponse(
+    return ResponsesResponse(
        id=response_id,
        model=model,
        status="completed",
        output=output,
        output_text=accumulated_text,
        usage=usage,
-    ).model_dump_json()
-    return
+    )


 async def generate_responses_stream(
@@ -239,16 +235,15 @@ async def generate_responses_stream(

    accumulated_text = ""
    function_call_items: list[ResponseFunctionCallItem] = []
-    last_usage: Usage | None = None
+    last_stats = None
    next_output_index = 1  # message item is at 0

    async for chunk in chunk_stream:
        if isinstance(chunk, ErrorChunk):
            break

-        last_usage = chunk.usage or last_usage
-
        if isinstance(chunk, ToolCallChunk):
+            last_stats = chunk.stats or last_stats
            for tool in chunk.tool_calls:
                fc_id = f"fc_{tool.id}"
                call_id = f"call_{tool.id}"
@@ -307,6 +302,7 @@ async def generate_responses_stream(
            continue

        accumulated_text += chunk.text
+        last_stats = chunk.stats or last_stats

        # response.output_text.delta
        delta_event = ResponseTextDeltaEvent(
@@ -350,13 +346,13 @@ async def generate_responses_stream(
    )
    yield f"event: response.output_item.done\ndata: {item_done.model_dump_json()}\n\n"

-    # Create usage from usage data if available
+    # Create usage from stats if available
    usage = None
-    if last_usage is not None:
+    if last_stats is not None:
        usage = ResponseUsage(
-            input_tokens=last_usage.prompt_tokens,
-            output_tokens=last_usage.completion_tokens,
-            total_tokens=last_usage.total_tokens,
+            input_tokens=last_stats.prompt_tokens,
+            output_tokens=last_stats.generation_tokens,
+            total_tokens=last_stats.prompt_tokens + last_stats.generation_tokens,
        )

    # response.completed
--- a/src/exo/master/api.py
+++ b/src/exo/master/api.py
@@ -105,6 +105,7 @@ from exo.shared.types.chunks import (
    ErrorChunk,
    ImageChunk,
    InputImageChunk,
+    PrefillProgressData,
    TokenChunk,
    ToolCallChunk,
 )
@@ -125,7 +126,6 @@ from exo.shared.types.commands import (
    PlaceInstance,
    SendInputChunk,
    StartDownload,
-    TaskCancelled,
    TaskFinished,
    TextGeneration,
 )
@@ -135,6 +135,7 @@ from exo.shared.types.events import (
    Event,
    ForwarderEvent,
    IndexedEvent,
+    PrefillProgress,
    TracesMerged,
 )
 from exo.shared.types.memory import Memory
@@ -218,7 +219,8 @@ class API:
        )

        self._text_generation_queues: dict[
-            CommandId, Sender[TokenChunk | ErrorChunk | ToolCallChunk]
+            CommandId,
+            Sender[TokenChunk | ErrorChunk | ToolCallChunk | PrefillProgressData],
        ] = {}
        self._image_generation_queues: dict[
            CommandId, Sender[ImageChunk | ErrorChunk]
@@ -522,36 +524,51 @@ class API:
            instance_id=instance_id,
        )

-    async def _token_chunk_stream(
+    async def _stream_events(
        self, command_id: CommandId
-    ) -> AsyncGenerator[ErrorChunk | ToolCallChunk | TokenChunk, None]:
-        """Yield chunks for a given command until completion.
+    ) -> AsyncGenerator[
+        TokenChunk | ErrorChunk | ToolCallChunk | PrefillProgressData, None
+    ]:
+        """Yield stream events for a command.

        This is the internal low-level stream used by all API adapters.
        """
        try:
            self._text_generation_queues[command_id], recv = channel[
-                ErrorChunk | ToolCallChunk | TokenChunk
+                TokenChunk | ErrorChunk | ToolCallChunk | PrefillProgressData
            ]()

-            with recv as token_chunks:
-                async for chunk in token_chunks:
-                    yield chunk
-                    if chunk.finish_reason is not None:
+            with recv as events:
+                async for event in events:
+                    yield event
+                    if (
+                        isinstance(event, TokenChunk)
+                        and event.finish_reason is not None
+                    ):
                        break

        except anyio.get_cancelled_exc_class():
-            command = TaskCancelled(cancelled_command_id=command_id)
-            with anyio.CancelScope(shield=True):
-                await self.command_sender.send(
-                    ForwarderCommand(origin=self.node_id, command=command)
-                )
+            # TODO: TaskCancelled
+            """
+            self.command_sender.send_nowait(
+                ForwarderCommand(origin=self.node_id, command=command)
+            )
+            """
            raise
        finally:
-            await self._send(TaskFinished(finished_command_id=command_id))
+            command = TaskFinished(finished_command_id=command_id)
+            await self._send(command)
            if command_id in self._text_generation_queues:
                del self._text_generation_queues[command_id]

+    async def _chunk_stream(
+        self, command_id: CommandId
+    ) -> AsyncGenerator[ErrorChunk | ToolCallChunk | TokenChunk, None]:
+        """Yield chunks, filtering out prefill progress events."""
+        async for event in self._stream_events(command_id):
+            if not isinstance(event, PrefillProgressData):
+                yield event
+
    async def _collect_text_generation_with_stats(
        self, command_id: CommandId
    ) -> BenchChatCompletionResponse:
@@ -562,7 +579,7 @@ class API:

        stats: GenerationStats | None = None

-        async for chunk in self._token_chunk_stream(command_id):
+        async for chunk in self._chunk_stream(command_id):
            if chunk.finish_reason == "error":
                raise HTTPException(
                    status_code=500,
@@ -634,7 +651,7 @@ class API:
            return StreamingResponse(
                generate_chat_stream(
                    command.command_id,
-                    self._token_chunk_stream(command.command_id),
+                    self._stream_events(command.command_id),
                ),
                media_type="text/event-stream",
                headers={
@@ -643,14 +660,14 @@ class API:
                    "X-Accel-Buffering": "no",
                },
            )
-        else:
-            return StreamingResponse(
-                collect_chat_response(
-                    command.command_id,
-                    self._token_chunk_stream(command.command_id),
-                ),
-                media_type="application/json",
+
+        try:
+            return await collect_chat_response(
+                command.command_id,
+                self._chunk_stream(command.command_id),
            )
+        except ValueError as e:
+            raise HTTPException(status_code=500, detail=str(e)) from e

    async def bench_chat_completions(
        self, payload: BenchChatCompletionRequest
@@ -666,7 +683,8 @@ class API:
        command = TextGeneration(task_params=task_params)
        await self._send(command)

-        return await self._collect_text_generation_with_stats(command.command_id)
+        response = await self._collect_text_generation_with_stats(command.command_id)
+        return response

    async def _resolve_and_validate_text_model(self, model_id: ModelId) -> ModelId:
        """Validate a text model exists and return the resolved model ID.
@@ -884,11 +902,6 @@ class API:
                        del image_metadata[key]

        except anyio.get_cancelled_exc_class():
-            command = TaskCancelled(cancelled_command_id=command_id)
-            with anyio.CancelScope(shield=True):
-                await self.command_sender.send(
-                    ForwarderCommand(origin=self.node_id, command=command)
-                )
            raise
        finally:
            await self._send(TaskFinished(finished_command_id=command_id))
@@ -970,11 +983,6 @@ class API:

            return (images, stats if capture_stats else None)
        except anyio.get_cancelled_exc_class():
-            command = TaskCancelled(cancelled_command_id=command_id)
-            with anyio.CancelScope(shield=True):
-                await self.command_sender.send(
-                    ForwarderCommand(origin=self.node_id, command=command)
-                )
            raise
        finally:
            await self._send(TaskFinished(finished_command_id=command_id))
@@ -1223,7 +1231,7 @@ class API:
                generate_claude_stream(
                    command.command_id,
                    payload.model,
-                    self._token_chunk_stream(command.command_id),
+                    self._chunk_stream(command.command_id),
                ),
                media_type="text/event-stream",
                headers={
@@ -1232,15 +1240,15 @@ class API:
                    "X-Accel-Buffering": "no",
                },
            )
-        else:
-            return StreamingResponse(
-                collect_claude_response(
-                    command.command_id,
-                    payload.model,
-                    self._token_chunk_stream(command.command_id),
-                ),
-                media_type="application/json",
+
+        try:
+            return await collect_claude_response(
+                command.command_id,
+                payload.model,
+                self._chunk_stream(command.command_id),
            )
+        except ValueError as e:
+            raise HTTPException(status_code=500, detail=str(e)) from e

    async def openai_responses(
        self, payload: ResponsesRequest
@@ -1258,7 +1266,7 @@ class API:
                generate_responses_stream(
                    command.command_id,
                    payload.model,
-                    self._token_chunk_stream(command.command_id),
+                    self._chunk_stream(command.command_id),
                ),
                media_type="text/event-stream",
                headers={
@@ -1268,15 +1276,14 @@ class API:
                },
            )

-        else:
-            return StreamingResponse(
-                collect_responses_response(
-                    command.command_id,
-                    payload.model,
-                    self._token_chunk_stream(command.command_id),
-                ),
-                media_type="application/json",
+        try:
+            return await collect_responses_response(
+                command.command_id,
+                payload.model,
+                self._chunk_stream(command.command_id),
            )
+        except ValueError as e:
+            raise HTTPException(status_code=500, detail=str(e)) from e

    def _calculate_total_available_memory(self) -> Memory:
        """Calculate total available memory across all nodes in bytes."""
@@ -1430,6 +1437,20 @@ class API:
                            except BrokenResourceError:
                                self._text_generation_queues.pop(event.command_id, None)

+                    elif isinstance(event, PrefillProgress):
+                        if queue := self._text_generation_queues.get(
+                            event.command_id, None
+                        ):
+                            try:
+                                await queue.send(
+                                    PrefillProgressData(
+                                        processed_tokens=event.processed_tokens,
+                                        total_tokens=event.total_tokens,
+                                    )
+                                )
+                            except BrokenResourceError:
+                                self._text_generation_queues.pop(event.command_id, None)
+
                    if isinstance(event, TracesMerged):
                        self._save_merged_trace(event)

--- a/src/exo/master/main.py
+++ b/src/exo/master/main.py
@@ -24,7 +24,6 @@ from exo.shared.types.commands import (
    PlaceInstance,
    RequestEventLog,
    SendInputChunk,
-    TaskCancelled,
    TaskFinished,
    TestCommand,
    TextGeneration,
@@ -40,7 +39,6 @@ from exo.shared.types.events import (
    NodeTimedOut,
    TaskCreated,
    TaskDeleted,
-    TaskStatusUpdated,
    TraceEventData,
    TracesCollected,
    TracesMerged,
@@ -281,7 +279,7 @@ class Master:
                        case DeleteInstance():
                            placement = delete_instance(command, self.state.instances)
                            transition_events = get_transition_events(
-                                self.state.instances, placement, self.state.tasks
+                                self.state.instances, placement
                            )
                            for cmd in cancel_unnecessary_downloads(
                                placement, self.state.downloads
@@ -301,7 +299,7 @@ class Master:
                                self.state.node_network,
                            )
                            transition_events = get_transition_events(
-                                self.state.instances, placement, self.state.tasks
+                                self.state.instances, placement
                            )
                            generated_events.extend(transition_events)
                        case CreateInstance():
@@ -311,7 +309,7 @@ class Master:
                                self.state.instances,
                            )
                            transition_events = get_transition_events(
-                                self.state.instances, placement, self.state.tasks
+                                self.state.instances, placement
                            )
                            generated_events.extend(transition_events)
                        case SendInputChunk(chunk=chunk):
@@ -321,18 +319,6 @@ class Master:
                                    chunk=chunk,
                                )
                            )
-                        case TaskCancelled():
-                            if (
-                                task_id := self.command_task_mapping.get(
-                                    command.cancelled_command_id
-                                )
-                            ) is not None:
-                                generated_events.append(
-                                    TaskStatusUpdated(
-                                        task_status=TaskStatus.Cancelled,
-                                        task_id=task_id,
-                                    )
-                                )
                        case TaskFinished():
                            generated_events.append(
                                TaskDeleted(
@@ -341,9 +327,10 @@ class Master:
                                    ]
                                )
                            )
-                            self.command_task_mapping.pop(
-                                command.finished_command_id, None
-                            )
+                            if command.finished_command_id in self.command_task_mapping:
+                                del self.command_task_mapping[
+                                    command.finished_command_id
+                                ]
                        case RequestEventLog():
                            # We should just be able to send everything, since other buffers will ignore old messages
                            # rate limit to 1000 at a time
--- a/src/exo/master/placement.py
+++ b/src/exo/master/placement.py
@@ -22,15 +22,9 @@ from exo.shared.types.commands import (
    PlaceInstance,
 )
 from exo.shared.types.common import NodeId
-from exo.shared.types.events import (
-    Event,
-    InstanceCreated,
-    InstanceDeleted,
-    TaskStatusUpdated,
-)
+from exo.shared.types.events import Event, InstanceCreated, InstanceDeleted
 from exo.shared.types.memory import Memory
 from exo.shared.types.profiling import MemoryUsage, NodeNetworkInfo
-from exo.shared.types.tasks import Task, TaskId, TaskStatus
 from exo.shared.types.worker.downloads import (
    DownloadOngoing,
    DownloadProgress,
@@ -192,7 +186,6 @@ def delete_instance(
 def get_transition_events(
    current_instances: Mapping[InstanceId, Instance],
    target_instances: Mapping[InstanceId, Instance],
-    tasks: Mapping[TaskId, Task],
 ) -> Sequence[Event]:
    events: list[Event] = []

@@ -208,18 +201,6 @@ def get_transition_events(
    # find instances to delete
    for instance_id in current_instances:
        if instance_id not in target_instances:
-            for task in tasks.values():
-                if task.instance_id == instance_id and task.task_status in [
-                    TaskStatus.Pending,
-                    TaskStatus.Running,
-                ]:
-                    events.append(
-                        TaskStatusUpdated(
-                            task_status=TaskStatus.Cancelled,
-                            task_id=task.task_id,
-                        )
-                    )
-
            events.append(
                InstanceDeleted(
                    instance_id=instance_id,
--- a/src/exo/master/tests/test_claude_tool_use.py
+++ b/src/exo/master/tests/test_claude_tool_use.py
@@ -4,11 +4,7 @@ import json
 from collections.abc import AsyncGenerator
 from typing import Any, cast

-from exo.master.adapters.claude import (
-    ClaudeMessagesResponse,
-    collect_claude_response,
-    generate_claude_stream,
-)
+from exo.master.adapters.claude import collect_claude_response, generate_claude_stream
 from exo.shared.types.api import ToolCallItem
 from exo.shared.types.chunks import ErrorChunk, TokenChunk, ToolCallChunk
 from exo.shared.types.common import CommandId, ModelId
@@ -21,18 +17,6 @@ async def _chunks_to_stream(
        yield chunk


-async def _collect_response(
-    command_id: CommandId,
-    model: str,
-    chunk_stream: AsyncGenerator[ErrorChunk | ToolCallChunk | TokenChunk, None],
-) -> ClaudeMessagesResponse:
-    """Helper to consume the async generator and parse the JSON response."""
-    parts: list[str] = []
-    async for part in collect_claude_response(command_id, model, chunk_stream):
-        parts.append(part)
-    return ClaudeMessagesResponse.model_validate_json("".join(parts))
-
-
 MODEL = ModelId("test-model")
 COMMAND_ID = CommandId("cmd_test123")

@@ -63,7 +47,7 @@ class TestCollectClaudeResponseToolUse:
                ],
            ),
        ]
-        response = await _collect_response(
+        response = await collect_claude_response(
            COMMAND_ID, "test-model", _chunks_to_stream(chunks)
        )

@@ -93,7 +77,7 @@ class TestCollectClaudeResponseToolUse:
                ],
            ),
        ]
-        response = await _collect_response(
+        response = await collect_claude_response(
            COMMAND_ID, "test-model", _chunks_to_stream(chunks)
        )

@@ -118,7 +102,7 @@ class TestCollectClaudeResponseToolUse:
                ],
            ),
        ]
-        response = await _collect_response(
+        response = await collect_claude_response(
            COMMAND_ID, "test-model", _chunks_to_stream(chunks)
        )

@@ -132,7 +116,7 @@ class TestCollectClaudeResponseToolUse:

    async def test_no_content_produces_empty_text_block(self):
        chunks: list[ErrorChunk | ToolCallChunk | TokenChunk] = []
-        response = await _collect_response(
+        response = await collect_claude_response(
            COMMAND_ID, "test-model", _chunks_to_stream(chunks)
        )
        assert len(response.content) == 1
--- a/src/exo/master/tests/test_placement.py
+++ b/src/exo/master/tests/test_placement.py
@@ -239,7 +239,7 @@ def test_get_transition_events_no_change(instance: Instance):
    target_instances = {instance_id: instance}

    # act
-    events = get_transition_events(current_instances, target_instances, {})
+    events = get_transition_events(current_instances, target_instances)

    # assert
    assert len(events) == 0
@@ -252,7 +252,7 @@ def test_get_transition_events_create_instance(instance: Instance):
    target_instances: dict[InstanceId, Instance] = {instance_id: instance}

    # act
-    events = get_transition_events(current_instances, target_instances, {})
+    events = get_transition_events(current_instances, target_instances)

    # assert
    assert len(events) == 1
@@ -266,7 +266,7 @@ def test_get_transition_events_delete_instance(instance: Instance):
    target_instances: dict[InstanceId, Instance] = {}

    # act
-    events = get_transition_events(current_instances, target_instances, {})
+    events = get_transition_events(current_instances, target_instances)

    # assert
    assert len(events) == 1
--- a/src/exo/routing/router.py
+++ b/src/exo/routing/router.py
@@ -211,14 +211,6 @@ class Router:
                    pass
                except AllQueuesFullError:
                    logger.warning(f"All peer queues full, dropping message on {topic}")
-                except RuntimeError as e:
-                    if "MessageTooLarge" in str(e):
-                        logger.error(
-                            f"Message too large for gossipsub on topic {topic} "
-                            f"({len(data)} bytes), dropping message"
-                        )
-                    else:
-                        raise


 def get_node_id_keypair(
--- a/src/exo/shared/apply.py
+++ b/src/exo/shared/apply.py
@@ -15,6 +15,7 @@ from exo.shared.types.events import (
    NodeDownloadProgress,
    NodeGatheredInfo,
    NodeTimedOut,
+    PrefillProgress,
    RunnerDeleted,
    RunnerStatusUpdated,
    TaskAcknowledged,
@@ -64,6 +65,7 @@ def event_apply(event: Event, state: State) -> State:
            | ChunkGenerated()
            | TaskAcknowledged()
            | InputChunkReceived()
+            | PrefillProgress()
            | TracesCollected()
            | TracesMerged()
        ):  # Pass-through events that don't modify state
@@ -218,6 +220,11 @@ def apply_node_timed_out(event: NodeTimedOut, state: State) -> State:
        key: value for key, value in state.downloads.items() if key != event.node_id
    }
    # Clean up all granular node mappings
+    node_identities = {
+        key: value
+        for key, value in state.node_identities.items()
+        if key != event.node_id
+    }
    node_memory = {
        key: value for key, value in state.node_memory.items() if key != event.node_id
    }
@@ -258,6 +265,7 @@ def apply_node_timed_out(event: NodeTimedOut, state: State) -> State:
            "downloads": downloads,
            "topology": topology,
            "last_seen": last_seen,
+            "node_identities": node_identities,
            "node_memory": node_memory,
            "node_disk": node_disk,
            "node_system": node_system,
--- a/src/exo/shared/types/api.py
+++ b/src/exo/shared/types/api.py
@@ -3,7 +3,8 @@ from collections.abc import Generator
 from typing import Annotated, Any, Literal
 from uuid import uuid4

-from pydantic import BaseModel, Field
+from pydantic import BaseModel, Field, field_validator
+from pydantic_core import PydanticUseDefault

 from exo.shared.models.model_cards import ModelCard, ModelId
 from exo.shared.types.common import CommandId, NodeId
@@ -227,6 +228,13 @@ class PlaceInstanceParams(BaseModel):
    instance_meta: InstanceMeta = InstanceMeta.MlxRing
    min_nodes: int = 1

+    @field_validator("sharding", "instance_meta", mode="plain")
+    @classmethod
+    def use_default(cls, v: object):
+        if not v or not isinstance(v, (Sharding, InstanceMeta)):
+            raise PydanticUseDefault()
+        return v
+

 class CreateInstanceParams(BaseModel):
    instance: Instance
--- a/src/exo/shared/types/chunks.py
+++ b/src/exo/shared/types/chunks.py
@@ -77,3 +77,13 @@ class InputImageChunk(BaseChunk):


 GenerationChunk = TokenChunk | ImageChunk | ToolCallChunk | ErrorChunk
+
+
+class PrefillProgressData(TaggedModel):
+    """Data class for prefill progress events during streaming."""
+
+    processed_tokens: int
+    total_tokens: int
+
+
+StreamEvent = TokenChunk | PrefillProgressData
--- a/src/exo/shared/types/commands.py
+++ b/src/exo/shared/types/commands.py
@@ -48,10 +48,6 @@ class DeleteInstance(BaseCommand):
    instance_id: InstanceId


-class TaskCancelled(BaseCommand):
-    cancelled_command_id: CommandId
-
-
 class TaskFinished(BaseCommand):
    finished_command_id: CommandId

@@ -93,7 +89,6 @@ Command = (
    | PlaceInstance
    | CreateInstance
    | DeleteInstance
-    | TaskCancelled
    | TaskFinished
    | SendInputChunk
 )
--- a/src/exo/shared/types/events.py
+++ b/src/exo/shared/types/events.py
@@ -102,6 +102,12 @@ class InputChunkReceived(BaseEvent):
    chunk: InputImageChunk


+class PrefillProgress(BaseEvent):
+    command_id: CommandId
+    processed_tokens: int
+    total_tokens: int
+
+
 class TopologyEdgeCreated(BaseEvent):
    conn: Connection

@@ -148,6 +154,7 @@ Event = (
    | NodeDownloadProgress
    | ChunkGenerated
    | InputChunkReceived
+    | PrefillProgress
    | TopologyEdgeCreated
    | TopologyEdgeDeleted
    | TracesCollected
--- a/src/exo/shared/types/tasks.py
+++ b/src/exo/shared/types/tasks.py
@@ -24,7 +24,6 @@ class TaskStatus(str, Enum):
    Complete = "Complete"
    TimedOut = "TimedOut"
    Failed = "Failed"
-    Cancelled = "Cancelled"


 class BaseTask(TaggedModel):
@@ -61,11 +60,6 @@ class TextGeneration(BaseTask):  # emitted by Master
    error_message: str | None = Field(default=None)


-class CancelTask(BaseTask):
-    cancelled_task_id: TaskId
-    runner_id: RunnerId
-
-
 class ImageGeneration(BaseTask):  # emitted by Master
    command_id: CommandId
    task_params: ImageGenerationTaskParams
@@ -93,7 +87,6 @@ Task = (
    | LoadModel
    | StartWarmup
    | TextGeneration
-    | CancelTask
    | ImageGeneration
    | ImageEdits
    | Shutdown
--- a/src/exo/shared/types/worker/downloads.py
+++ b/src/exo/shared/types/worker/downloads.py
@@ -26,7 +26,6 @@ class DownloadProgressData(CamelCaseModel):
 class BaseDownloadProgress(TaggedModel):
    node_id: NodeId
    shard_metadata: ShardMetadata
-    model_directory: str = ""


 class DownloadPending(BaseDownloadProgress):
--- a/src/exo/shared/types/worker/runner_response.py
+++ b/src/exo/shared/types/worker/runner_response.py
@@ -62,8 +62,12 @@ class PartialImageResponse(BaseRunnerResponse):
 class ToolCallResponse(BaseRunnerResponse):
    tool_calls: list[ToolCallItem]
    usage: Usage | None
-    stats: GenerationStats | None = None


 class FinishedResponse(BaseRunnerResponse):
    pass
+
+
+class PrefillProgressResponse(BaseRunnerResponse):
+    processed_tokens: int
+    total_tokens: int
--- a/src/exo/utils/banner.py
+++ b/src/exo/utils/banner.py
@@ -1,7 +1,5 @@
-import sys
-
-
 def print_startup_banner(port: int) -> None:
+    """Print a prominent startup banner with API endpoint information."""
    dashboard_url = f"http://localhost:{port}"
    banner = f"""
 ╔═══════════════════════════════════════════════════════════════════════╗
@@ -29,4 +27,4 @@ def print_startup_banner(port: int) -> None:

 """

-    print(banner, file=sys.stderr)
+    print(banner)
--- a/src/exo/utils/channels.py
+++ b/src/exo/utils/channels.py
@@ -125,9 +125,7 @@ class MpSender[T]:
            self._state.buffer.put(item, block=True)

    async def send_async(self, item: T) -> None:
-        await to_thread.run_sync(
-            self.send, item, limiter=CapacityLimiter(1), abandon_on_cancel=True
-        )
+        await to_thread.run_sync(self.send, item, limiter=CapacityLimiter(1))

    def close(self) -> None:
        if not self._state.closed.is_set():
--- a/src/exo/worker/engines/mlx/generator/generate.py
+++ b/src/exo/worker/engines/mlx/generator/generate.py
@@ -160,7 +160,7 @@ def warmup_inference(
        max_tokens=50,
        sampler=sampler,
        prompt_cache=cache,
-        prefill_step_size=2048,
+        prefill_step_size=1024,
        kv_group_size=KV_GROUP_SIZE,
        kv_bits=KV_BITS,
    ):
@@ -252,6 +252,7 @@ def mlx_generate(
    task: TextGenerationTaskParams,
    prompt: str,
    kv_prefix_cache: KVPrefixCache | None = None,
+    on_prefill_progress: Callable[[int, int], None] | None = None,
    group: mx.distributed.Group | None = None,
 ) -> Generator[GenerationResponse]:
    # Ensure that generation stats only contains peak memory for this generation
@@ -306,7 +307,7 @@ def mlx_generate(
    max_stop_len = max((len(s) for s in stop_sequences), default=0)

    mx_barrier(group)
-    logger.info("Starting prefill")
+    logger.info("Ready to prefill")

    # Prefill cache with all tokens except the last one
    prefill_tps, prefill_tokens, ssm_snapshots_list = prefill(
@@ -345,6 +346,7 @@ def mlx_generate(
            prefill_step_size=1,
            kv_group_size=KV_GROUP_SIZE,
            kv_bits=KV_BITS,
+            prompt_progress_callback=on_prefill_progress,
        ),
        start=1,
    ):
@@ -393,11 +395,10 @@ def mlx_generate(
                    f"Model generated unexpected finish_reason: {out.finish_reason}"
                )

-            total_prompt_tokens = len(all_prompt_tokens)
            usage = Usage(
-                prompt_tokens=total_prompt_tokens,
+                prompt_tokens=int(out.prompt_tokens),
                completion_tokens=completion_tokens,
-                total_tokens=total_prompt_tokens + completion_tokens,
+                total_tokens=int(out.prompt_tokens) + completion_tokens,
                prompt_tokens_details=PromptTokensDetails(
                    cached_tokens=prefix_hit_length
                ),
--- a/src/exo/worker/engines/mlx/utils_mlx.py
+++ b/src/exo/worker/engines/mlx/utils_mlx.py
@@ -64,6 +64,8 @@ from exo.worker.runner.bootstrap import logger
 Group = mx.distributed.Group


+# TODO: Test this
+#  ALSO https://github.com/exo-explore/exo/pull/233#discussion_r2549683673
 def get_weights_size(model_shard_meta: ShardMetadata) -> Memory:
    return Memory.from_float_kb(
        (model_shard_meta.end_layer - model_shard_meta.start_layer)
@@ -81,6 +83,30 @@ class ModelLoadingTimeoutError(Exception):
    pass


+def mx_barrier(group: Group | None = None):
+    mx.eval(
+        mx.distributed.all_sum(
+            mx.array(1.0),
+            stream=mx.default_stream(mx.Device(mx.cpu)),
+            group=group,
+        )
+    )
+
+
+def broadcast_from_zero(value: int, group: Group | None = None):
+    if group is None:
+        return value
+
+    if group.rank() == 0:
+        a = mx.array([value], dtype=mx.int32)
+    else:
+        a = mx.array([0], dtype=mx.int32)
+
+    m = mx.distributed.all_sum(a, stream=mx.Device(mx.DeviceType.cpu), group=group)
+    mx.eval(m)
+    return int(m.item())
+
+
 class HostList(RootModel[list[str]]):
    @classmethod
    def from_hosts(cls, hosts: list[Host]) -> "HostList":
@@ -353,13 +379,7 @@ def load_tokenizer_for_model_id(
            return list(hf_tokenizer.model.encode(text, allowed_special="all"))  # pyright: ignore[reportUnknownMemberType,reportUnknownArgumentType]

        hf_tokenizer.encode = _patched_encode
-        return TokenizerWrapper(
-            hf_tokenizer,
-            eos_token_ids=eos_token_ids,
-            tool_call_start="<|tool_calls_section_begin|>",
-            tool_call_end="<|tool_calls_section_end|>",
-            tool_parser=_parse_kimi_tool_calls,
-        )
+        return TokenizerWrapper(hf_tokenizer, eos_token_ids=eos_token_ids)

    tokenizer = load_tokenizer(
        model_path,
@@ -571,61 +591,3 @@ def mlx_cleanup(
    import gc

    gc.collect()
-
-
-def mx_any(bool_: bool, group: Group | None) -> bool:
-    if group is None:
-        return bool_
-    num_true = mx.distributed.all_sum(
-        mx.array(bool_), group=group, stream=mx.default_stream(mx.Device(mx.cpu))
-    )
-    mx.eval(num_true)
-    return num_true.item() > 0
-
-
-def mx_barrier(group: Group | None):
-    if group is None:
-        return
-    mx.eval(
-        mx.distributed.all_sum(
-            mx.array(1.0), group=group, stream=mx.default_stream(mx.Device(mx.cpu))
-        )
-    )
-
-
-def _parse_kimi_tool_calls(text: str):
-    import regex as re
-
-    # kimi has a fixed function naming scheme, with a json formatted arg
-    #   functions.multiply:0<|tool_call_argument_begin|>{"a": 2, "b": 3}
-    _func_name_regex = re.compile(
-        r"^\s*((?:functions\.)?(.+?):\d+)\s*<\|tool_call_argument_begin\|>", re.DOTALL
-    )
-    _func_arg_regex = re.compile(r"<\|tool_call_argument_begin\|>\s*(.*)\s*", re.DOTALL)
-    _tool_call_split_regex = re.compile(
-        r"<\|tool_call_begin\|>(.*?)<\|tool_call_end\|>", re.DOTALL
-    )
-
-    def _parse_single_tool(text: str) -> dict[str, Any]:
-        func_name_match = _func_name_regex.search(text)
-        if func_name_match is None:
-            raise ValueError("No tool call found.")
-        tool_call_id = func_name_match.group(1)  # e.g. "functions.get_weather:0"
-        func_name = func_name_match.group(2)  # e.g. "get_weather"
-
-        func_args_match = _func_arg_regex.search(text)
-        if func_args_match is None:
-            raise ValueError("No tool call arguments found.")
-        func_args = func_args_match.group(1)
-        try:
-            arg_dct = json.loads(func_args)  # pyright: ignore[reportAny]
-        except Exception:
-            arg_dct = None
-
-        return dict(id=tool_call_id, name=func_name, arguments=arg_dct)
-
-    tool_matches = _tool_call_split_regex.findall(text)
-    if tool_matches:
-        return [_parse_single_tool(match) for match in tool_matches]  # pyright: ignore[reportAny]
-    else:
-        return [_parse_single_tool(text)]
--- a/src/exo/worker/main.py
+++ b/src/exo/worker/main.py
@@ -33,7 +33,6 @@ from exo.shared.types.events import (
 from exo.shared.types.multiaddr import Multiaddr
 from exo.shared.types.state import State
 from exo.shared.types.tasks import (
-    CancelTask,
    CreateRunner,
    DownloadModel,
    ImageEdits,
@@ -225,22 +224,15 @@ class Worker:
                        )
                    )
                case Shutdown(runner_id=runner_id):
-                    runner = self.runners.pop(runner_id)
                    try:
                        with fail_after(3):
-                            await runner.start_task(task)
+                            await self.runners.pop(runner_id).start_task(task)
                    except TimeoutError:
                        await self.event_sender.send(
                            TaskStatusUpdated(
                                task_id=task.task_id, task_status=TaskStatus.TimedOut
                            )
                        )
-                    finally:
-                        runner.shutdown()
-                case CancelTask(
-                    cancelled_task_id=cancelled_task_id, runner_id=runner_id
-                ):
-                    await self.runners[runner_id].cancel_task(cancelled_task_id)
                case ImageEdits() if task.task_params.total_input_chunks > 0:
                    # Assemble image from chunks and inject into task
                    cmd_id = task.command_id
@@ -278,18 +270,18 @@ class Worker:
                        del self.input_chunk_buffer[cmd_id]
                    if cmd_id in self.input_chunk_counts:
                        del self.input_chunk_counts[cmd_id]
-                    await self._start_runner_task(modified_task)
+                    await self.runners[self._task_to_runner_id(task)].start_task(
+                        modified_task
+                    )
                case task:
-                    await self._start_runner_task(task)
+                    await self.runners[self._task_to_runner_id(task)].start_task(task)

    def shutdown(self):
        self._tg.cancel_scope.cancel()

-    async def _start_runner_task(self, task: Task):
-        if (instance := self.state.instances.get(task.instance_id)) is not None:
-            await self.runners[
-                instance.shard_assignments.node_to_runner[self.node_id]
-            ].start_task(task)
+    def _task_to_runner_id(self, task: Task):
+        instance = self.state.instances[task.instance_id]
+        return instance.shard_assignments.node_to_runner[self.node_id]

    async def _nack_request(self, since_idx: int) -> None:
        # We request all events after (and including) the missing index.
@@ -328,6 +320,8 @@ class Worker:
            for event in self.out_for_delivery.copy().values():
                await self.local_event_sender.send(event)

+    ## Op Executors
+
    def _create_supervisor(self, task: CreateRunner) -> RunnerSupervisor:
        """Creates and stores a new AssignedRunner with initial downloading status."""
        runner = RunnerSupervisor.create(
--- a/src/exo/worker/plan.py
+++ b/src/exo/worker/plan.py
@@ -4,7 +4,6 @@ from collections.abc import Mapping, Sequence

 from exo.shared.types.common import CommandId, NodeId
 from exo.shared.types.tasks import (
-    CancelTask,
    ConnectToGroup,
    CreateRunner,
    DownloadModel,
@@ -54,14 +53,13 @@ def plan(
 ) -> Task | None:
    # Python short circuiting OR logic should evaluate these sequentially.
    return (
-        _cancel_tasks(runners, tasks)
-        or _kill_runner(runners, all_runners, instances)
+        _kill_runner(runners, all_runners, instances)
        or _create_runner(node_id, runners, instances)
        or _model_needs_download(node_id, runners, global_download_status)
        or _init_distributed_backend(runners, all_runners)
        or _load_model(runners, all_runners, global_download_status)
        or _ready_to_warmup(runners, all_runners)
-        or _pending_tasks(runners, tasks, all_runners, input_chunk_buffer or {})
+        or _pending_tasks(runners, tasks, all_runners, input_chunk_buffer)
    )


@@ -272,7 +270,7 @@ def _pending_tasks(
    runners: Mapping[RunnerId, RunnerSupervisor],
    tasks: Mapping[TaskId, Task],
    all_runners: Mapping[RunnerId, RunnerStatus],
-    input_chunk_buffer: Mapping[CommandId, dict[int, str]],
+    input_chunk_buffer: Mapping[CommandId, dict[int, str]] | None = None,
 ) -> Task | None:
    for task in tasks.values():
        # for now, just forward chat completions
@@ -286,7 +284,7 @@ def _pending_tasks(
        if isinstance(task, ImageEdits) and task.task_params.total_input_chunks > 0:
            cmd_id = task.command_id
            expected = task.task_params.total_input_chunks
-            received = len(input_chunk_buffer.get(cmd_id, {}))
+            received = len((input_chunk_buffer or {}).get(cmd_id, {}))
            if received < expected:
                continue  # Wait for all chunks to arrive

@@ -294,33 +292,16 @@ def _pending_tasks(
            if task.instance_id != runner.bound_instance.instance.instance_id:
                continue

-            # the task status _should_ be set to completed by the LAST runner
-            # it is currently set by the first
-            # this is definitely a hack
+            # I have a design point here; this is a state race in disguise as the task status doesn't get updated to completed fast enough
+            # however, realistically the task status should be set to completed by the LAST runner, so this is a true race
+            # the actual solution is somewhat deeper than this bypass - TODO!
            if task.task_id in runner.completed:
                continue

+            # TODO: Check ordering aligns with MLX distributeds expectations.
+
            if isinstance(runner.status, RunnerReady) and all(
                isinstance(all_runners[global_runner_id], (RunnerReady, RunnerRunning))
                for global_runner_id in runner.bound_instance.instance.shard_assignments.runner_to_shard
            ):
                return task
-
-
-def _cancel_tasks(
-    runners: Mapping[RunnerId, RunnerSupervisor],
-    tasks: Mapping[TaskId, Task],
-) -> Task | None:
-    for task in tasks.values():
-        if task.task_status != TaskStatus.Cancelled:
-            continue
-        for runner_id, runner in runners.items():
-            if task.instance_id != runner.bound_instance.instance.instance_id:
-                continue
-            if task.task_id in runner.cancelled:
-                continue
-            return CancelTask(
-                instance_id=task.instance_id,
-                cancelled_task_id=task.task_id,
-                runner_id=runner_id,
-            )
--- a/src/exo/worker/runner/bootstrap.py
+++ b/src/exo/worker/runner/bootstrap.py
@@ -3,7 +3,7 @@ import os
 import loguru

 from exo.shared.types.events import Event, RunnerStatusUpdated
-from exo.shared.types.tasks import Task, TaskId
+from exo.shared.types.tasks import Task
 from exo.shared.types.worker.instances import BoundInstance, MlxJacclInstance
 from exo.shared.types.worker.runners import RunnerFailed
 from exo.utils.channels import ClosedResourceError, MpReceiver, MpSender
@@ -15,7 +15,6 @@ def entrypoint(
    bound_instance: BoundInstance,
    event_sender: MpSender[Event],
    task_receiver: MpReceiver[Task],
-    cancel_receiver: MpReceiver[TaskId],
    _logger: "loguru.Logger",
 ) -> None:
    fast_synch_override = os.environ.get("EXO_FAST_SYNCH")
@@ -39,7 +38,7 @@ def entrypoint(
    try:
        from exo.worker.runner.runner import main

-        main(bound_instance, event_sender, task_receiver, cancel_receiver)
+        main(bound_instance, event_sender, task_receiver)
    except ClosedResourceError:
        logger.warning("Runner communication closed unexpectedly")
    except Exception as e:
--- a/src/exo/worker/runner/runner.py
+++ b/src/exo/worker/runner/runner.py
@@ -1,10 +1,10 @@
 import base64
-import math
+import json
 import resource
 import time
 from collections.abc import Generator
 from functools import cache
-from typing import Literal
+from typing import Any, Callable, Literal

 import mlx.core as mx
 from mlx_lm.models.gpt_oss import Model as GptOssModel
@@ -15,6 +15,7 @@ from openai_harmony import (  # pyright: ignore[reportMissingTypeStubs]
    StreamableParser,
    load_harmony_encoding,
 )
+from pydantic import ValidationError

 from exo.shared.constants import EXO_MAX_CHUNK_SIZE, EXO_TRACING_ENABLED
 from exo.shared.models.model_cards import ModelId, ModelTask
@@ -25,6 +26,7 @@ from exo.shared.types.common import CommandId
 from exo.shared.types.events import (
    ChunkGenerated,
    Event,
+    PrefillProgress,
    RunnerStatusUpdated,
    TaskAcknowledged,
    TaskStatusUpdated,
@@ -87,12 +89,9 @@ from exo.worker.engines.mlx.utils_mlx import (
    initialize_mlx,
    load_mlx_items,
    mlx_force_oom,
-    mx_any,
 )
 from exo.worker.runner.bootstrap import logger

-from .tool_parsers import ToolParser, make_mlx_parser
-

 def _is_primary_output_node(shard_metadata: ShardMetadata) -> bool:
    """Check if this node is the primary output node for image generation.
@@ -114,7 +113,6 @@ def main(
    bound_instance: BoundInstance,
    event_sender: MpSender[Event],
    task_receiver: MpReceiver[Task],
-    cancel_receiver: MpReceiver[TaskId],
 ):
    soft, hard = resource.getrlimit(resource.RLIMIT_NOFILE)
    resource.setrlimit(resource.RLIMIT_NOFILE, (min(max(soft, 2048), hard), hard))
@@ -132,16 +130,11 @@ def main(
        time.sleep(timeout)

    setup_start_time = time.time()
-    cancelled_tasks = set[TaskId]()

-    # type checker was unhappy with me - splitting these fixed it
-    inference_model: Model | None = None
-    image_model: DistributedImageModel | None = None
+    model: Model | DistributedImageModel | None = None
    tokenizer = None
-    tool_parser: ToolParser | None = None
    group = None
    kv_prefix_cache: KVPrefixCache | None = None
-    check_for_cancel_every: int | None = None

    current_status: RunnerStatus = RunnerIdle()
    logger.info("runner created")
@@ -154,7 +147,6 @@ def main(
            if task.task_id in seen:
                logger.warning("repeat task - potential error")
            seen.add(task.task_id)
-            cancelled_tasks.discard(TaskId("CANCEL_CURRENT_TASK"))
            event_sender.send(
                TaskStatusUpdated(task_id=task.task_id, task_status=TaskStatus.Running)
            )
@@ -200,28 +192,19 @@ def main(
                        time.sleep(0.5)

                    if ModelTask.TextGeneration in shard_metadata.model_card.tasks:
-                        inference_model, tokenizer = load_mlx_items(
+                        model, tokenizer = load_mlx_items(
                            bound_instance, group, on_timeout=on_model_load_timeout
                        )
                        logger.info(
-                            f"model has_tool_calling={tokenizer.has_tool_calling} using tokens {tokenizer.tool_call_start}, {tokenizer.tool_call_end}"
+                            f"model has_tool_calling={tokenizer.has_tool_calling}"
                        )
-                        if tokenizer.has_tool_calling:
-                            assert tokenizer.tool_call_start
-                            assert tokenizer.tool_call_end
-                            assert tokenizer.tool_parser  # pyright: ignore[reportAny]
-                            tool_parser = make_mlx_parser(
-                                tokenizer.tool_call_start,
-                                tokenizer.tool_call_end,
-                                tokenizer.tool_parser,  # pyright: ignore[reportAny]
-                            )
                        kv_prefix_cache = KVPrefixCache(group)

                    elif (
                        ModelTask.TextToImage in shard_metadata.model_card.tasks
                        or ModelTask.ImageToImage in shard_metadata.model_card.tasks
                    ):
-                        image_model = initialize_image_model(bound_instance)
+                        model = initialize_image_model(bound_instance)
                    else:
                        raise ValueError(
                            f"Unknown model task(s): {shard_metadata.model_card.tasks}"
@@ -229,6 +212,8 @@ def main(
                    current_status = RunnerLoaded()
                    logger.info("runner loaded")
                case StartWarmup() if isinstance(current_status, RunnerLoaded):
+                    assert model
+
                    current_status = RunnerWarmingUp()
                    logger.info("runner warming up")
                    event_sender.send(
@@ -240,31 +225,16 @@ def main(

                    logger.info(f"warming up inference for instance: {instance}")
                    if ModelTask.TextGeneration in shard_metadata.model_card.tasks:
-                        assert inference_model
+                        assert not isinstance(model, DistributedImageModel)
                        assert tokenizer

-                        t = time.monotonic()
                        toks = warmup_inference(
-                            model=inference_model,
+                            model=model,
                            tokenizer=tokenizer,
                            group=group,
+                            # kv_prefix_cache=kv_prefix_cache,  # supply for warmup-time prefix caching
                        )
                        logger.info(f"warmed up by generating {toks} tokens")
-                        check_for_cancel_every = min(
-                            math.ceil(toks / min(time.monotonic() - t, 0.001)), 100
-                        )
-                        if group is not None:
-                            check_for_cancel_every = int(
-                                mx.max(
-                                    mx.distributed.all_gather(
-                                        mx.array([check_for_cancel_every]), group=group
-                                    )
-                                ).item()
-                            )
-
-                        logger.info(
-                            f"runner checking for cancellation every {check_for_cancel_every} tokens"
-                        )
                        logger.info(
                            f"runner initialized in {time.time() - setup_start_time} seconds"
                        )
@@ -272,8 +242,8 @@ def main(
                        ModelTask.TextToImage in shard_metadata.model_card.tasks
                        or ModelTask.ImageToImage in shard_metadata.model_card.tasks
                    ):
-                        assert image_model
-                        image = warmup_image_generator(model=image_model)
+                        assert isinstance(model, DistributedImageModel)
+                        image = warmup_image_generator(model=model)
                        if image is not None:
                            logger.info(f"warmed up by generating {image.size} image")
                        else:
@@ -293,9 +263,20 @@ def main(
                        )
                    )
                    event_sender.send(TaskAcknowledged(task_id=task.task_id))
-                    assert inference_model
+
+                    assert model and not isinstance(model, DistributedImageModel)
                    assert tokenizer
-                    assert check_for_cancel_every
+
+                    # Define callback to send prefill progress events directly
+                    def on_prefill_progress(processed: int, total: int) -> None:
+                        if device_rank == 0:
+                            event_sender.send(
+                                PrefillProgress(
+                                    command_id=command_id,
+                                    processed_tokens=processed,
+                                    total_tokens=total,
+                                )
+                            )

                    try:
                        _check_for_debug_prompts(task_params)
@@ -305,11 +286,12 @@ def main(

                        # Generate responses using the actual MLX generation
                        mlx_generator = mlx_generate(
-                            model=inference_model,
+                            model=model,
                            tokenizer=tokenizer,
                            task=task_params,
                            prompt=prompt,
                            kv_prefix_cache=kv_prefix_cache,
+                            on_prefill_progress=on_prefill_progress,
                            group=group,
                        )

@@ -320,25 +302,34 @@ def main(
                                mlx_generator, tokenizer
                            )

+                        # Kimi-K2 has tool call sections - we don't care about them
+                        if "kimi" in shard_metadata.model_card.model_id.lower():
+                            mlx_generator = filter_kimi_tokens(mlx_generator)
+                            patch_kimi_tokenizer(tokenizer)
+
+                        # GLM models need patched parser (upstream has bug with None regex match)
+                        elif "glm" in shard_metadata.model_card.model_id.lower():
+                            patch_glm_tokenizer(tokenizer)
+
                        # GPT-OSS specific parsing to match other model formats.
-                        if isinstance(inference_model, GptOssModel):
+                        elif isinstance(model, GptOssModel):
                            mlx_generator = parse_gpt_oss(mlx_generator)
-                        elif tool_parser:
-                            mlx_generator = parse_tool_calls(mlx_generator, tool_parser)
+
+                        if tokenizer.has_tool_calling and not isinstance(
+                            model, GptOssModel
+                        ):
+                            assert tokenizer.tool_call_start
+                            assert tokenizer.tool_call_end
+                            assert tokenizer.tool_parser  # pyright: ignore[reportAny]
+                            mlx_generator = parse_tool_calls(
+                                mlx_generator,
+                                tokenizer.tool_call_start,
+                                tokenizer.tool_call_end,
+                                tokenizer.tool_parser,  # pyright: ignore[reportAny]
+                            )

                        completion_tokens = 0
-                        tokens_since_last_cancel_check = 0
                        for response in mlx_generator:
-                            tokens_since_last_cancel_check += 1
-                            if tokens_since_last_cancel_check >= check_for_cancel_every:
-                                tokens_since_last_cancel_check = 0
-                                cancelled_tasks.update(cancel_receiver.collect())
-                                want_to_cancel = (task.task_id in cancelled_tasks) or (
-                                    TaskId("CANCEL_CURRENT_TASK") in cancelled_tasks
-                                )
-                                if mx_any(want_to_cancel, group):
-                                    break
-
                            match response:
                                case GenerationResponse():
                                    completion_tokens += 1
@@ -386,7 +377,6 @@ def main(
                                                    tool_calls=response.tool_calls,
                                                    model=shard_metadata.model_card.model_id,
                                                    usage=response.usage,
-                                                    stats=response.stats,
                                                ),
                                            )
                                        )
@@ -411,7 +401,7 @@ def main(
                case ImageGeneration(
                    task_params=task_params, command_id=command_id
                ) if isinstance(current_status, RunnerReady):
-                    assert image_model
+                    assert isinstance(model, DistributedImageModel)
                    logger.info(f"received image generation request: {str(task)[:500]}")
                    current_status = RunnerRunning()
                    logger.info("runner running")
@@ -424,9 +414,7 @@ def main(

                    try:
                        image_index = 0
-                        for response in generate_image(
-                            model=image_model, task=task_params
-                        ):
+                        for response in generate_image(model=model, task=task_params):
                            is_primary_output = _is_primary_output_node(shard_metadata)

                            if is_primary_output:
@@ -476,7 +464,7 @@ def main(
                case ImageEdits(task_params=task_params, command_id=command_id) if (
                    isinstance(current_status, RunnerReady)
                ):
-                    assert image_model
+                    assert isinstance(model, DistributedImageModel)
                    logger.info(f"received image edits request: {str(task)[:500]}")
                    current_status = RunnerRunning()
                    logger.info("runner running")
@@ -489,9 +477,7 @@ def main(

                    try:
                        image_index = 0
-                        for response in generate_image(
-                            model=image_model, task=task_params
-                        ):
+                        for response in generate_image(model=model, task=task_params):
                            if _is_primary_output_node(shard_metadata):
                                match response:
                                    case PartialImageResponse():
@@ -550,20 +536,14 @@ def main(
                    raise ValueError(
                        f"Received {task.__class__.__name__} outside of state machine in {current_status=}"
                    )
-            was_cancelled = (task.task_id in cancelled_tasks) or (
-                TaskId("CANCEL_CURRENT_TASK") in cancelled_tasks
+            event_sender.send(
+                TaskStatusUpdated(task_id=task.task_id, task_status=TaskStatus.Complete)
            )
-            if not was_cancelled:
-                event_sender.send(
-                    TaskStatusUpdated(
-                        task_id=task.task_id, task_status=TaskStatus.Complete
-                    )
-                )
            event_sender.send(
                RunnerStatusUpdated(runner_id=runner_id, runner_status=current_status)
            )
            if isinstance(current_status, RunnerShutdown):
-                del inference_model, image_model, tokenizer, group
+                del model, tokenizer, group
                mx.clear_cache()
                import gc

@@ -577,8 +557,21 @@ def get_gpt_oss_encoding():
    return encoding


+def filter_kimi_tokens(
+    responses: Generator[GenerationResponse | ToolCallResponse],
+) -> Generator[GenerationResponse]:
+    for resp in responses:
+        assert isinstance(resp, GenerationResponse)
+        if (
+            resp.text == "<|tool_calls_section_begin|>"
+            or resp.text == "<|tool_calls_section_end|>"
+        ):
+            continue
+        yield resp
+
+
 def parse_gpt_oss(
-    responses: Generator[GenerationResponse],
+    responses: Generator[GenerationResponse | ToolCallResponse],
 ) -> Generator[GenerationResponse | ToolCallResponse]:
    encoding = get_gpt_oss_encoding()
    stream = StreamableParser(encoding, role=Role.ASSISTANT)
@@ -635,9 +628,9 @@ def parse_gpt_oss(


 def parse_thinking_models(
-    responses: Generator[GenerationResponse],
+    responses: Generator[GenerationResponse | ToolCallResponse],
    tokenizer: TokenizerWrapper,
-) -> Generator[GenerationResponse]:
+) -> Generator[GenerationResponse | ToolCallResponse]:
    """
    For models that inject thinking tags in the prompt (like GLM-4.7),
    prepend the thinking tag to the output stream so the frontend
@@ -758,55 +751,218 @@ def _process_image_response(


 def parse_tool_calls(
-    responses: Generator[GenerationResponse], tool_parser: ToolParser
+    responses: Generator[GenerationResponse | ToolCallResponse],
+    tool_call_start: str,
+    tool_call_end: str,
+    tool_parser: Callable[[str], dict[str, Any] | list[dict[str, Any]]],
 ) -> Generator[GenerationResponse | ToolCallResponse]:
    in_tool_call = False
    tool_call_text_parts: list[str] = []
    for response in responses:
-        if response.text.startswith(tool_parser.start_parsing):
+        assert isinstance(response, GenerationResponse)
+        # assumption: the tool call start is one token
+        if response.text == tool_call_start:
            in_tool_call = True
-
-        if in_tool_call:
-            tool_call_text_parts.append(response.text)
-            if response.text.endswith(tool_parser.end_parsing):
-                # parse the actual tool calls from the tool call text
-                parsed = tool_parser.parse_tool_calls(
-                    "".join(tool_call_text_parts).strip()
-                )
+            continue
+        # assumption: the tool call end is one token
+        if in_tool_call and response.text == tool_call_end:
+            try:
+                # tool_parser returns an arbitrarily nested python dictionary
+                # we actually don't want the python dictionary, we just want to
+                # parse the top level { function: ..., arguments: ... } structure
+                # as we're just gonna hand it back to the api anyway
+                parsed = tool_parser("".join(tool_call_text_parts).strip())
                logger.info(f"parsed {tool_call_text_parts=} into {parsed=}")
-                if parsed is not None:
-                    yield ToolCallResponse(
-                        tool_calls=parsed, usage=response.usage, stats=response.stats
-                    )
+                if isinstance(parsed, list):
+                    tools = [_validate_single_tool(tool) for tool in parsed]
                else:
-                    logger.warning(
-                        f"tool call parsing failed for text {''.join(tool_call_text_parts)}"
-                    )
-                    response.text = "".join(tool_call_text_parts)
-                    yield response
+                    tools = [_validate_single_tool(parsed)]
+                yield ToolCallResponse(tool_calls=tools, usage=response.usage)

-                in_tool_call = False
-                tool_call_text_parts = []
-                continue
-
-            if response.finish_reason is not None:
-                logger.info(
-                    "tool call parsing interrupted, yield partial tool call as text"
-                )
-                response = response.model_copy(
-                    update={
-                        "text": "".join(tool_call_text_parts),
-                        "token": 0,
-                    }
+            except (
+                json.JSONDecodeError,
+                ValidationError,
+                ValueError,
+                AttributeError,
+            ) as e:
+                # ValueError: our parsers raise this for malformed tool calls
+                # AttributeError: upstream parsers (e.g. glm47) may raise this when regex doesn't match
+                logger.opt(exception=e).warning("tool call parsing failed")
+                # assumption: talking about tool calls, not making a tool call
+                response.text = (
+                    tool_call_start + "".join(tool_call_text_parts) + tool_call_end
                )
                yield response

+            in_tool_call = False
+            tool_call_text_parts = []
            continue

+        if in_tool_call:
+            tool_call_text_parts.append(response.text)
+            if response.finish_reason is not None:
+                logger.info(
+                    "toll call parsing interrupted, yield partial tool call as text"
+                )
+                yield GenerationResponse(
+                    text=tool_call_start + "".join(tool_call_text_parts),
+                    token=0,
+                    finish_reason=response.finish_reason,
+                    usage=None,
+                )
+            continue
        # fallthrough
        yield response


+def patch_kimi_tokenizer(tokenizer: TokenizerWrapper):
+    """
+    Version of to-be-upstreamed kimi-k2 tool parser
+    """
+    import ast
+    import json
+    from typing import Any
+
+    import regex as re
+
+    # kimi has a fixed function naming scheme, with a json formatted arg
+    #   functions.multiply:0 <|tool_call_argument_begin|> {"a": 2, "b": 3}
+    #   Also needs to handle tools like call_0<|tool_call_argument_begin|>{"filePath": "..."}
+    _func_name_regex = re.compile(
+        r"^\s*(.+)[:](\d+)\s*<\|tool_call_argument_begin\|>", re.DOTALL
+    )
+    _func_arg_regex = re.compile(r"<\|tool_call_argument_begin\|>\s*(.*)\s*", re.DOTALL)
+
+    # kimi has a tool_calls_section - we're leaving this up to the caller to handle
+    tool_call_start = "<|tool_call_begin|>"
+    tool_call_end = "<|tool_call_end|>"
+
+    def _deserialize(value: str) -> Any:  # pyright: ignore[reportAny]
+        try:
+            return json.loads(value)  # pyright: ignore[reportAny]
+        except Exception:
+            pass
+
+        try:
+            return ast.literal_eval(value)  # pyright: ignore[reportAny]
+        except Exception:
+            pass
+        return value
+
+    def parse_tool_call(text: str, tools: Any | None = None):
+        func_name_match = _func_name_regex.search(text)
+        if func_name_match is None:
+            raise ValueError(f"Could not parse function name from tool call: {text!r}")
+        original_func_name = func_name_match.group(1)
+        tool_id = func_name_match.group(2)
+        # strip off the `functions.` prefix, if it exists.
+        func_name = original_func_name[original_func_name.find(".") + 1 :]
+
+        func_args_match = _func_arg_regex.search(text)
+        if func_args_match is None:
+            raise ValueError(f"Could not parse function args from tool call: {text!r}")
+        func_args = func_args_match.group(1)
+        # the args should be valid json - no need to check against our tools to deserialize
+        arg_dct = _deserialize(func_args)  # pyright: ignore[reportAny]
+
+        return dict(
+            id=f"{original_func_name}:{tool_id}",
+            name=func_name,
+            arguments=arg_dct,  # pyright: ignore[reportAny]
+        )
+
+    tokenizer._tool_call_start = tool_call_start
+    tokenizer._tool_call_end = tool_call_end
+    tokenizer._tool_parser = parse_tool_call
+
+
+def patch_glm_tokenizer(tokenizer: TokenizerWrapper):
+    """
+    Fixed version of mlx_lm's glm47 tool parser that handles regex match failures.
+    """
+    import ast
+    import json
+    from typing import Any
+
+    import regex as re
+
+    _func_name_regex = re.compile(r"^(.*?)<arg_key>", re.DOTALL)
+    _func_arg_regex = re.compile(
+        r"<arg_key>(.*?)</arg_key>(?:\n|\s)*<arg_value>(.*?)(?:</arg_value>|(?=<arg_key>)|$)",
+        re.DOTALL,
+    )
+
+    tool_call_start = "<tool_call>"
+    tool_call_end = "</tool_call>"
+
+    def _is_string_type(
+        tool_name: str,
+        arg_name: str,
+        tools: list[Any] | None,
+    ) -> bool:
+        if tools is None:
+            return False
+        for tool in tools:  # pyright: ignore[reportAny]
+            func = tool["function"]  # pyright: ignore[reportAny]
+            if func["name"] == tool_name:
+                params = func["parameters"]  # pyright: ignore[reportAny]
+                if params is None:
+                    return False
+                props = params.get("properties", {})  # pyright: ignore[reportAny]
+                arg_props = props.get(arg_name, {})  # pyright: ignore[reportAny]
+                arg_type = arg_props.get("type", None)  # pyright: ignore[reportAny]
+                return arg_type == "string"  # pyright: ignore[reportAny]
+        return False
+
+    def _deserialize(value: str) -> Any:  # pyright: ignore[reportAny]
+        try:
+            return json.loads(value)  # pyright: ignore[reportAny]
+        except Exception:
+            pass
+        try:
+            return ast.literal_eval(value)  # pyright: ignore[reportAny]
+        except Exception:
+            pass
+        return value
+
+    def parse_tool_call(text: str, tools: list[Any] | None = None):
+        func_name_match = _func_name_regex.search(text)
+        if func_name_match is None:
+            raise ValueError(f"Could not parse function name from tool call: {text!r}")
+        func_name = func_name_match.group(1)
+
+        pairs = _func_arg_regex.findall(text)
+        arg_dct: dict[str, Any] = {}
+        for key, value in pairs:  # pyright: ignore[reportAny]
+            arg_key = key.strip()  # pyright: ignore[reportAny]
+            arg_val = value.strip()  # pyright: ignore[reportAny]
+            if not _is_string_type(func_name, arg_key, tools):  # pyright: ignore[reportAny]
+                arg_val = _deserialize(arg_val)  # pyright: ignore[reportAny]
+            arg_dct[arg_key] = arg_val
+        return dict(name=func_name, arguments=arg_dct)
+
+    tokenizer._tool_call_start = tool_call_start
+    tokenizer._tool_call_end = tool_call_end
+    tokenizer._tool_parser = parse_tool_call
+
+
+def _validate_single_tool(obj: dict[str, Any]) -> ToolCallItem:
+    if (
+        ((name := obj.get("name")) is not None)
+        and ((args := obj.get("arguments")) is not None)
+        and isinstance(name, str)
+    ):
+        raw_id: object = obj.get("id")
+        extra = {"id": str(raw_id)} if raw_id is not None else {}
+        return ToolCallItem(
+            **extra,
+            name=name,
+            arguments=json.dumps(args),
+        )
+    else:
+        raise ValidationError
+
+
 EXO_RUNNER_MUST_FAIL = "EXO RUNNER MUST FAIL"
 EXO_RUNNER_MUST_OOM = "EXO RUNNER MUST OOM"
 EXO_RUNNER_MUST_TIMEOUT = "EXO RUNNER MUST TIMEOUT"
--- a/src/exo/worker/runner/runner_supervisor.py
+++ b/src/exo/worker/runner/runner_supervisor.py
@@ -47,11 +47,9 @@ class RunnerSupervisor:
    _ev_recv: MpReceiver[Event]
    _task_sender: MpSender[Task]
    _event_sender: Sender[Event]
-    _cancel_sender: MpSender[TaskId]
    status: RunnerStatus = field(default_factory=RunnerIdle, init=False)
    pending: dict[TaskId, anyio.Event] = field(default_factory=dict, init=False)
    completed: set[TaskId] = field(default_factory=set, init=False)
-    cancelled: set[TaskId] = field(default_factory=set, init=False)

    @classmethod
    def create(
@@ -62,8 +60,8 @@ class RunnerSupervisor:
        initialize_timeout: float = 400,
    ) -> Self:
        ev_send, ev_recv = mp_channel[Event]()
+        # A task is kind of a runner command
        task_sender, task_recv = mp_channel[Task]()
-        cancel_sender, cancel_recv = mp_channel[TaskId]()

        runner_process = Process(
            target=entrypoint,
@@ -71,7 +69,6 @@ class RunnerSupervisor:
                bound_instance,
                ev_send,
                task_recv,
-                cancel_recv,
                logger,
            ),
            daemon=True,
@@ -86,7 +83,6 @@ class RunnerSupervisor:
            initialize_timeout=initialize_timeout,
            _ev_recv=ev_recv,
            _task_sender=task_sender,
-            _cancel_sender=cancel_sender,
            _event_sender=event_sender,
        )

@@ -101,8 +97,6 @@ class RunnerSupervisor:
        self._ev_recv.close()
        self._task_sender.close()
        self._event_sender.close()
-        self._cancel_sender.send(TaskId("CANCEL_CURRENT_TASK"))
-        self._cancel_sender.close()
        self.runner_process.join(1)
        if not self.runner_process.is_alive():
            logger.info("Runner process succesfully terminated")
@@ -118,6 +112,14 @@ class RunnerSupervisor:
        logger.critical("Runner process didn't respond to SIGTERM, killing")
        self.runner_process.kill()

+        self.runner_process.join(1)
+        if not self.runner_process.is_alive():
+            return
+
+        logger.critical(
+            "Runner process didn't respond to SIGKILL. System resources may have leaked"
+        )
+
    async def start_task(self, task: Task):
        if task.task_id in self.pending:
            logger.warning(
@@ -139,17 +141,6 @@ class RunnerSupervisor:
            return
        await event.wait()

-    async def cancel_task(self, task_id: TaskId):
-        if task_id in self.completed:
-            logger.info(f"Unable to cancel {task_id} as it has been completed")
-            return
-        self.cancelled.add(task_id)
-        with anyio.move_on_after(0.5) as scope:
-            await self._cancel_sender.send_async(task_id)
-        if scope.cancel_called:
-            logger.error("RunnerSupervisor cancel pipe blocked")
-            await self._check_runner(TimeoutError("cancel pipe blocked"))
-
    async def _forward_events(self):
        with self._ev_recv as events:
            try:
--- a/src/exo/worker/runner/tool_parsers.py
+++ b/src/exo/worker/runner/tool_parsers.py
@@ -1,72 +0,0 @@
-import json
-from dataclasses import dataclass
-from typing import Any, Callable
-
-from exo.shared.types.api import ToolCallItem
-
-
-@dataclass
-class ToolParser:
-    start_parsing: str
-    end_parsing: str
-    parse_tool_calls: Callable[[str], list[ToolCallItem] | None]
-
-
-def make_mlx_parser(
-    tool_call_start: str,
-    tool_call_end: str,
-    tool_parser: Callable[[str], dict[str, Any] | list[dict[str, Any]]],
-) -> ToolParser:
-    def parse_tool_calls(text: str) -> list[ToolCallItem] | None:
-        try:
-            text = text.removeprefix(tool_call_start)
-            text = text.removesuffix(tool_call_end)
-            parsed = tool_parser(text)
-            if isinstance(parsed, list):
-                return [ToolCallItem.model_validate(_flatten(p)) for p in parsed]
-            else:
-                return [ToolCallItem.model_validate(_flatten(parsed))]
-
-        except Exception:
-            return None
-
-    return ToolParser(
-        start_parsing=tool_call_start,
-        end_parsing=tool_call_end,
-        parse_tool_calls=parse_tool_calls,
-    )
-
-
-# TODO / example code:
-def _parse_json_calls(text: str) -> list[ToolCallItem] | None:
-    try:
-        text = text.removeprefix("<tool_call>")
-        text = text.removesuffix("</tool_call>")
-        top_level = {
-            k: json.dumps(v) if isinstance(v, (dict, list)) else v
-            for k, v in json.loads(text).items()  # pyright: ignore[reportAny]
-        }
-        return [ToolCallItem.model_validate(top_level)]
-    except Exception:
-        return None
-
-
-def _flatten(p: dict[str, Any]) -> dict[str, str]:
-    return {
-        k: json.dumps(v) if isinstance(v, (dict, list)) else str(v)  # pyright: ignore[reportAny]
-        for k, v in p.items()  # pyright: ignore[reportAny]
-    }
-
-
-json_tool_parser = ToolParser(
-    start_parsing="<tool_call>",
-    end_parsing="</tool_call>",
-    parse_tool_calls=_parse_json_calls,
-)
-
-
-def infer_tool_parser(chat_template: str) -> ToolParser | None:
-    """Attempt to auto-infer a tool parser from the chat template."""
-    if "<tool_call>" in chat_template and "tool_call.name" in chat_template:
-        return json_tool_parser
-    return None
--- a/src/exo/worker/tests/unittests/test_runner/test_event_ordering.py
+++ b/src/exo/worker/tests/unittests/test_runner/test_event_ordering.py
@@ -1,9 +1,7 @@
 # Check tasks are complete before runner is ever ready.
-import unittest.mock
 from collections.abc import Iterable
 from typing import Callable

-import mlx.core as mx
 import pytest

 import exo.worker.runner.runner as mlx_runner
@@ -21,7 +19,6 @@ from exo.shared.types.tasks import (
    Shutdown,
    StartWarmup,
    Task,
-    TaskId,
    TaskStatus,
    TextGeneration,
 )
@@ -116,7 +113,6 @@ def patch_out_mlx(monkeypatch: pytest.MonkeyPatch):
    monkeypatch.setattr(mlx_runner, "load_mlx_items", make_nothin((1, MockTokenizer)))
    monkeypatch.setattr(mlx_runner, "warmup_inference", make_nothin(1))
    monkeypatch.setattr(mlx_runner, "_check_for_debug_prompts", nothin)
-    monkeypatch.setattr(mlx_runner, "mx_any", make_nothin(False))
    # Mock apply_chat_template since we're using a fake tokenizer (integer 1).
    # Returns a prompt without thinking tag so detect_thinking_prompt_suffix returns None.
    monkeypatch.setattr(mlx_runner, "apply_chat_template", make_nothin("test prompt"))
@@ -167,7 +163,6 @@ def _run(tasks: Iterable[Task]):
    )

    task_sender, task_receiver = mp_channel[Task]()
-    _cancel_sender, cancel_receiver = mp_channel[TaskId]()
    event_sender = EventCollector()

    with task_sender:
@@ -178,16 +173,8 @@ def _run(tasks: Iterable[Task]):
        # this is some c++ nonsense
        task_receiver.close = nothin
        task_receiver.join = nothin
-        with unittest.mock.patch(
-            "exo.worker.runner.runner.mx.distributed.all_gather",
-            make_nothin(mx.array([1])),
-        ):
-            mlx_runner.main(
-                bound_instance,
-                event_sender,  # pyright: ignore[reportArgumentType]
-                task_receiver,
-                cancel_receiver,
-            )
+
+        mlx_runner.main(bound_instance, event_sender, task_receiver)  # type: ignore[arg-type]

        return event_sender.events

--- a/src/exo/worker/tests/unittests/test_runner/test_parse_tool_calls.py
+++ b/src/exo/worker/tests/unittests/test_runner/test_parse_tool_calls.py
@@ -5,13 +5,12 @@ from typing import Any

 from exo.shared.types.worker.runner_response import GenerationResponse, ToolCallResponse
 from exo.worker.runner.runner import parse_tool_calls
-from exo.worker.runner.tool_parsers import make_mlx_parser


 def _make_responses(
    texts: list[str],
    finish_on_last: bool = True,
-) -> Generator[GenerationResponse]:
+) -> Generator[GenerationResponse | ToolCallResponse]:
    """Create a sequence of GenerationResponses from text strings."""
    for i, text in enumerate(texts):
        is_last = i == len(texts) - 1
@@ -23,13 +22,10 @@ def _make_responses(
        )


-def _dummier_parser(text: str) -> dict[str, Any]:
+def _dummy_parser(text: str) -> dict[str, Any]:
    return {"name": "test_fn", "arguments": {"arg": text}}


-_dummy_parser = make_mlx_parser("<tool_call>", "</tool_call>", _dummier_parser)
-
-
 class TestParseToolCalls:
    """Tests for parse_tool_calls generator."""

@@ -39,6 +35,8 @@ class TestParseToolCalls:
        results = list(
            parse_tool_calls(
                _make_responses(texts, finish_on_last=False),
+                "<tool_call>",
+                "</tool_call>",
                _dummy_parser,
            )
        )
@@ -52,6 +50,8 @@ class TestParseToolCalls:
        results = list(
            parse_tool_calls(
                _make_responses(texts),
+                "<tool_call>",
+                "</tool_call>",
                _dummy_parser,
            )
        )
@@ -76,7 +76,9 @@ class TestParseToolCalls:
        results = list(
            parse_tool_calls(
                _make_responses(texts, finish_on_last=False),
-                make_mlx_parser("<tool_call>", "</tool_call>", _failing_parser),
+                "<tool_call>",
+                "</tool_call>",
+                _failing_parser,
            )
        )
Author	SHA1	Message	Date
Alex Cheema	03f5dafa2a	Merge remote-tracking branch 'origin/main' into alexcheema/prefill-progress	2026-02-13 09:37:55 -08:00
Alex Cheema	92401ab7f8	chore: fix formatting after merge Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>	2026-02-13 06:54:18 -08:00
Alex Cheema	d0c8501cb0	Merge remote-tracking branch 'origin/main' into alexcheema/prefill-progress # Conflicts: # src/exo/master/adapters/chat_completions.py # src/exo/worker/engines/mlx/generator/generate.py # src/exo/worker/runner/runner.py	2026-02-13 05:58:45 -08:00
Alex Cheema	f94870f2c0	Merge remote-tracking branch 'origin/main' into alexcheema/prefill-progress # Conflicts: # dashboard/src/lib/stores/app.svelte.ts	2026-02-05 05:38:14 -08:00
Alex Cheema	18928d4d9e	Merge remote-tracking branch 'origin/main' into alexcheema/prefill-progress	2026-02-05 05:30:22 -08:00
Alex Cheema	937da476b0	address PR review comments: use model_dump_json, match/case, prefill_step_size=1024 - Use model_dump_json() for prefill progress SSE instead of hand-crafted JSON - Convert isinstance chain to match/case in generate_chat_stream - Set all prefill_step_size to 1024 (was 256 for testing, 2048 default) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-02-03 09:04:25 -08:00
Alex Cheema	1c1c286127	Merge uncertainty-visualization into prefill-progress Combines prefill progress bar (PrefillProgressData, SSE named events) with uncertainty visualization features (logprobs, tool calls, error chunks, traces, image generation). Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-02-03 06:18:54 -08:00
Alex Cheema	258785be84	Merge remote-tracking branch 'origin/main' into alexcheema/uncertainty-visualization	2026-02-03 06:03:01 -08:00
Alex Cheema	13a6b9819a	fix: assistant prefilling for regenerate-from-token and tooltip UX Support assistant message continuation by popping the last assistant message before template formatting and appending its content raw, keeping the turn open without a closing token. Improve tooltip hover UX: use getClientRects() for correct multi-line token positioning, add padding to bridge the hover gap, and increase the hide delay. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-02-03 06:00:28 -08:00
Alex Cheema	1733d07cb3	fix: enable uncertainty visualization for regular chat messages The sendMessage method was missing logprobs request params and token collection, so the heatmap toggle never appeared. Also rename the top_k parameter to top_logprobs in extract_top_logprobs to avoid confusion with the sampling top_k parameter. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-02-03 05:08:49 -08:00
Alex Cheema	b3e4c9b1e5	fix: populate logprobs in non-streaming chat completions responses collect_chat_response() was dropping logprobs data from TokenChunks, so non-streaming requests never returned logprobs even when requested. Accumulate LogprobsContentItems and attach them to the ChatCompletionChoice. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-02-03 04:45:39 -08:00
Alex Cheema	4c74792373	Merge branch 'main' into alexcheema/uncertainty-visualization	2026-02-03 04:44:14 -08:00
Alex Cheema	eadb6de1f7	Merge main into uncertainty-visualization branch Resolve conflicts by keeping main's structure (TextGenerationTaskParams, tool calling, KV prefix cache, Claude/OpenAI APIs) and surgically adding the uncertainty visualization features on top: - Add logprob/top_logprobs fields to GenerationResponse and TokenChunk - Add extract_top_logprobs() to MLX generator for per-token logprob extraction - Build Logprobs in chat completions adapter for streaming responses - Add SSE headers (Cache-Control, Connection, X-Accel-Buffering) to streaming endpoints - Add TokenHeatmap component and uncertainty toggle in dashboard - Add logprobs collection in streaming response handler - Add regenerateFromToken method for re-generation from specific tokens - Strip token data from localStorage to avoid storage bloat Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-02-02 11:33:24 -08:00
Alex Cheema	7ba2408eed	fix: restore dashboard build by using main's app.svelte.ts The prefill-progress branch's app.svelte.ts was missing image generation features from main. To fix the dashboard build, restored main's app.svelte.ts and removed the uncertainty visualization and prefill progress bar features from ChatMessages.svelte that depended on the missing exports. Note: TokenHeatmap and PrefillProgressBar components still exist but are not currently used. The prefill progress backend code is still in place and can be re-enabled in the dashboard once app.svelte.ts is properly updated to include both image generation and prefill/uncertainty features. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 12:07:16 +00:00
Alex Cheema	ce4d7f4d43	style: simplify prefill progress bar and use exo color palette - Remove spinner (progress bar is dynamic enough) - Use exo-yellow for progress bar fill - Use exo-black/60 for progress bar background - Use exo-light-gray for text Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 12:00:55 +00:00
Alex Cheema	6727523eab	fix: wire prefill progress events to chat completions stream - Move PrefillProgressData to shared types (chunks.py) to avoid circular imports - Update generate_chat_stream adapter to handle both TokenChunk and PrefillProgressData - Use _stream_events instead of _chat_chunk_stream for streaming endpoint - Prefill progress now properly sent as SSE 'event: prefill_progress' to frontend Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 12:00:55 +00:00
Alex Cheema	a8f81e0495	feat: add prefill progress bar for long prompts Shows real-time progress during prompt processing (prefill phase). Progress is sent via SSE named events that maintain OpenAI API compatibility. - Add PrefillProgress event type and PrefillProgressData dataclass - Wire prompt_progress_callback through MLX stream_generate - Send progress events directly from callback for real-time updates - Add PrefillProgressBar.svelte component - Parse event: prefill_progress SSE events in dashboard Note: prefill_step_size temporarily set to 256 for testing (normally 2048) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:59:52 +00:00
Alex Cheema	ba7148ccec	style: format app.svelte.ts with nix fmt	2026-01-22 11:53:43 +00:00
Alex Cheema	a64b8addc6	Fix localStorage quota issues by stripping tokens and auto-pruning - Strip tokens (logprobs data) from messages before saving to localStorage since they're large and not essential for persistence - Add pruneOldConversations() to automatically remove oldest conversations when quota is exceeded - This prevents QuotaExceededError from crashing the app Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:53:43 +00:00
Alex Cheema	e6599a9408	Fix ReferenceError: controller undefined in sendMessage finally block Move AbortController creation before the try block in both sendMessageWithLogprobs and regenerateFromToken functions. Previously, controller was defined inside the try block but referenced in the finally block, causing a ReferenceError if an exception was thrown before the controller was created. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:53:43 +00:00
Alex Cheema	93f4753598	Add SSE headers to properly close streaming connections Add Cache-Control, Connection: close, and X-Accel-Buffering headers to all SSE streaming responses. These headers help ensure: - No caching of streaming responses - Connection closes when stream ends (instead of keep-alive) - No proxy buffering that could delay stream closure This should fix the issue where the frontend stays on "PROCESSING" even after receiving the complete response. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:53:43 +00:00
Alex Cheema	75fe505275	Add debug logging to generate_chat_stream Add logging to help diagnose why streaming might not be ending properly. This will show when [DONE] is yielded, when return is called, and when the finally block runs. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:53:43 +00:00
Alex Cheema	d7c044e349	Fix streaming not ending after [DONE] is yielded Add missing return statement after yielding [DONE] in generate_chat_stream. Without this, the async generator continues waiting for more chunks from chunk_stream even though generation is complete, causing the stream to hang indefinitely. The frontend waits for the stream to close (reader.done) which never happens, resulting in the chat button staying on "PROCESSING" forever. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:53:43 +00:00
Alex Cheema	53b6d56e9f	fix: restore extract_top_logprobs function for uncertainty visualization The extract_top_logprobs function was lost during rebases. This function processes the out.logprobs array (full vocabulary logprobs from MLX) to extract the selected token's logprob and top-k alternatives. The previous code tried to use getattr(out, "logprob", None) which doesn't exist - mlx_lm returns logprobs as an mx.array, not individual values. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:53:43 +00:00
Alex Cheema	7fe0a61230	fix: remove unsupported logprob params from stream_generate The mlx_lm.stream_generate already returns logprobs in its output - we don't need to pass return_logprob or return_top_logprobs kwargs. The uncertainty visualization feature extracts logprobs from the existing out.logprobs field. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:53:43 +00:00
Alex Cheema	5a36542631	feat: add uncertainty visualization with token-level logprobs - Add TokenHeatmap component for visualizing token confidence - Collect and stream logprobs in generation pipeline - Add regenerate-from-token feature with continue_from_prefix - Add AbortController for request cancellation - Support continue_final_message for seamless prefix continuation Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:53:43 +00:00
Alex Cheema	955e0105b3	fix: resolve import and type errors from rebase - Use claude_request_to_internal instead of old function name - Fix ModelId imports in runner.py and test files - Update test_mlx/conftest.py to use ResponsesRequest format - Remove unused imports Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:36:11 +00:00
Evan	4d1eb1d9bd	fix: rebase fix	2026-01-22 11:32:46 +00:00
Alex Cheema	365416c65e	style: move inline imports to top of file in api.py Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:32:26 +00:00
Alex Cheema	04af76e10f	fix: restore try/except structure in runner.py Replace non-existent context manager with proper try/except block and remove unused ModelId import. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:32:04 +00:00
Alex Cheema	a84c3431cd	style: fix formatting issues caught by treefmt Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:31:45 +00:00
Alex Cheema	52445b21f6	refactor: use ResponsesRequest as canonical internal type - Extend ResponsesRequest with fields: top_k, seed, stop, tools - Remove redundant InternalTaskParams and InputMessage types - Update all adapters to convert to ResponsesRequest - Simplify Responses API (no conversion needed - native passthrough) - Update all imports across codebase and tests This eliminates type duplication and makes the Responses API relationship explicit throughout the codebase. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:31:44 +00:00
Alex Cheema	435bd7f6fa	refactor: make Responses API the canonical internal format Restructure the API layer so that OpenAI Responses API is the native format, with Chat Completions and Claude Messages as adapters on top. Changes: - Add new chat_completions.py adapter with streaming/non-streaming support - Update responses.py with collect_responses_response() for non-streaming - Update claude.py with collect_claude_response() for non-streaming - Refactor api.py so all endpoints use adapters uniformly - Rename _chat_chunk_stream to _token_chunk_stream (generic internal format) - Remove unused chat_response_to_* converter functions - Update tests to remove tests for deleted functions Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:30:27 +00:00
Alex Cheema	dd25b5b90e	feat: add Claude Messages API and OpenAI Responses API support Adds two new API endpoints that wrap the existing chat completions: - /v1/messages - Claude Messages API compatible endpoint - /v1/responses - OpenAI Responses API compatible endpoint Both support streaming (SSE) and non-streaming modes with proper token usage reporting from actual inference stats. Also adds top_k sampling parameter and stop sequence support to the MLX inference engine. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-22 11:28:49 +00:00