style: simplify prefill progress bar and use exo color palette

- Remove spinner (progress bar is dynamic enough) - Use exo-yellow for progress bar fill - Use exo-black/60 for progress bar background - Use exo-light-gray for text Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
fix: wire prefill progress events to chat completions stream
2026-01-20 03:51:14 -05:00 · 2026-01-19 16:57:58 +00:00 · 2026-01-19 16:41:16 +00:00 · 2026-01-19 14:56:53 +00:00 · 2026-01-19 14:55:42 +00:00 · 2026-01-19 14:55:42 +00:00
40 changed files with 3217 additions and 241 deletions
--- a/AGENTS.md
+++ b/AGENTS.md
@@ -40,6 +40,31 @@ uv run ruff check
 nix fmt
 ```

+## Pre-Commit Checks (REQUIRED)
+
+**IMPORTANT: Always run these checks before committing code. CI will fail if these don't pass.**
+
+```bash
+# 1. Type checking - MUST pass with 0 errors
+uv run basedpyright
+
+# 2. Linting - MUST pass
+uv run ruff check
+
+# 3. Formatting - MUST be applied
+nix fmt
+
+# 4. Tests - MUST pass
+uv run pytest
+```
+
+Run all checks in sequence:
+```bash
+uv run basedpyright && uv run ruff check && nix fmt && uv run pytest
+```
+
+If `nix fmt` changes any files, stage them before committing. The CI runs `nix flake check` which verifies formatting, linting, and runs Rust tests.
+
 ## Architecture

 ### Node Composition
--- a/README.md
+++ b/README.md
@@ -27,6 +27,15 @@ exo connects all your devices into an AI cluster. Not only does exo enable runni
 - **Tensor Parallelism**: exo supports sharding models, for up to 1.8x speedup on 2 devices and 3.2x speedup on 4 devices.
 - **MLX Support**: exo uses [MLX](https://github.com/ml-explore/mlx) as an inference backend and [MLX distributed](https://ml-explore.github.io/mlx/build/html/usage/distributed.html) for distributed communication.

+## Dashboard
+
+exo includes a built-in dashboard for managing your cluster and chatting with models.
+
+<p align="center">
+  <img src="docs/imgs/dashboard-cluster-view.png" alt="exo dashboard - cluster view showing 4 x M3 Ultra Mac Studio with DeepSeek v3.1 and Kimi-K2-Thinking loaded" width="80%" />
+</p>
+<p align="center"><em>4 × 512GB M3 Ultra Mac Studio running DeepSeek v3.1 (8-bit) and Kimi-K2-Thinking (4-bit)</em></p>
+
 ## Benchmarks

 <details>
--- a/bench/exo_bench.py
+++ b/bench/exo_bench.py
@@ -3,6 +3,7 @@
 from __future__ import annotations

 import argparse
+import contextlib
 import http.client
 import json
 import os
@@ -26,7 +27,7 @@ class ExoHttpError(RuntimeError):


 class ExoClient:
-    def __init__(self, host: str, port: int, timeout_s: float = 2400.0):
+    def __init__(self, host: str, port: int, timeout_s: float = 600.0):
        self.host = host
        self.port = port
        self.timeout_s = timeout_s
@@ -104,22 +105,46 @@ def runner_ready(runner: dict[str, Any]) -> bool:
    return "RunnerReady" in runner


+def runner_failed(runner: dict[str, Any]) -> bool:
+    return "RunnerFailed" in runner
+
+
+def get_runner_failed_message(runner: dict[str, Any]) -> str | None:
+    if "RunnerFailed" in runner:
+        return runner["RunnerFailed"].get("errorMessage")
+    return None
+
+
 def wait_for_instance_ready(
    client: ExoClient, instance_id: str, timeout: float = 24000.0
 ) -> None:
    start_time = time.time()
+    instance_existed = False
    while time.time() - start_time < timeout:
        state = client.request_json("GET", "/state")
        instances = state.get("instances", {})

        if instance_id not in instances:
+            if instance_existed:
+                # Instance was deleted after being created - likely due to runner failure
+                raise RuntimeError(
+                    f"Instance {instance_id} was deleted (runner may have failed)"
+                )
            time.sleep(0.1)
            continue

+        instance_existed = True
        instance = instances[instance_id]
        runner_ids = runner_ids_from_instance(instance)
        runners = state.get("runners", {})

+        # Check for failed runners first
+        for rid in runner_ids:
+            runner = runners.get(rid, {})
+            if runner_failed(runner):
+                error_msg = get_runner_failed_message(runner) or "Unknown error"
+                raise RuntimeError(f"Runner {rid} failed: {error_msg}")
+
        if all(runner_ready(runners.get(rid, {})) for rid in runner_ids):
            return

@@ -299,6 +324,12 @@ def main() -> int:
        default=4,
        help="Only consider placements using <= this many nodes.",
    )
+    ap.add_argument(
+        "--min-nodes",
+        type=int,
+        default=1,
+        help="Only consider placements using >= this many nodes.",
+    )
    ap.add_argument(
        "--instance-meta", choices=["ring", "jaccl", "both"], default="both"
    )
@@ -320,7 +351,7 @@ def main() -> int:
        help="Warmup runs per placement (uses first pp/tg).",
    )
    ap.add_argument(
-        "--timeout", type=float, default=2400.0, help="HTTP timeout (seconds)."
+        "--timeout", type=float, default=600.0, help="HTTP timeout (seconds)."
    )
    ap.add_argument(
        "--json-out",
@@ -399,7 +430,7 @@ def main() -> int:
        ):
            continue

-        if 0 < n <= args.max_nodes:
+        if args.min_nodes <= n <= args.max_nodes:
            selected.append(p)

    if not selected:
@@ -441,7 +472,13 @@ def main() -> int:
        )

        client.request_json("POST", "/instance", body={"instance": instance})
-        wait_for_instance_ready(client, instance_id)
+        try:
+            wait_for_instance_ready(client, instance_id)
+        except (RuntimeError, TimeoutError) as e:
+            logger.error(f"Failed to initialize placement: {e}")
+            with contextlib.suppress(ExoHttpError):
+                client.request_json("DELETE", f"/instance/{instance_id}")
+            continue

        time.sleep(1)

--- a/dashboard/src/lib/components/ChatMessages.svelte
+++ b/dashboard/src/lib/components/ChatMessages.svelte
@@ -1,14 +1,17 @@
 <script lang="ts">
-	import { 
-		messages, 
-		currentResponse, 
+	import {
+		messages,
+		currentResponse,
 		isLoading,
 		deleteMessage,
 		editAndRegenerate,
-		regenerateLastResponse
+		regenerateLastResponse,
+		regenerateFromToken
 	} from '$lib/stores/app.svelte';
 	import type { MessageAttachment } from '$lib/stores/app.svelte';
 	import MarkdownContent from './MarkdownContent.svelte';
+	import TokenHeatmap from './TokenHeatmap.svelte';
+	import PrefillProgressBar from './PrefillProgressBar.svelte';

 	interface Props {
 		class?: string;
@@ -95,6 +98,23 @@
 let copiedMessageId = $state<string | null>(null);
 let expandedThinkingMessageIds = $state<Set<string>>(new Set());

+// Uncertainty view state - tracks which messages show token heatmap
+let uncertaintyViewMessageIds = $state<Set<string>>(new Set());
+
+function toggleUncertaintyView(messageId: string) {
+	const newSet = new Set(uncertaintyViewMessageIds);
+	if (newSet.has(messageId)) {
+		newSet.delete(messageId);
+	} else {
+		newSet.add(messageId);
+	}
+	uncertaintyViewMessageIds = newSet;
+}
+
+function isUncertaintyViewEnabled(messageId: string): boolean {
+	return uncertaintyViewMessageIds.has(messageId);
+}
+
 	function formatTimestamp(timestamp: number): string {
 		return new Date(timestamp).toLocaleTimeString('en-US', { 
 			hour12: false,
@@ -330,6 +350,10 @@ function isThinkingExpanded(messageId: string): boolean {
 						{:else}
 							<!-- Assistant message styling -->
 							<div class="p-3 sm:p-4">
+								{#if message.prefillProgress}
+									<!-- Prefill progress bar -->
+									<PrefillProgressBar progress={message.prefillProgress} class="mb-3" />
+								{/if}
 								{#if message.thinking && message.thinking.trim().length > 0}
 									<div class="mb-3 rounded border border-exo-yellow/20 bg-exo-black/40">
 										<button
@@ -366,7 +390,17 @@ function isThinkingExpanded(messageId: string): boolean {
 									</div>
 								{/if}
 								<div class="text-xs text-foreground">
-									<MarkdownContent content={message.content || (loading ? response : '')} />
+									{#if message.role === 'assistant' && isUncertaintyViewEnabled(message.id) && message.tokens && message.tokens.length > 0}
+										<!-- Uncertainty heatmap view -->
+										<TokenHeatmap
+											tokens={message.tokens}
+											isGenerating={loading}
+											onRegenerateFrom={(tokenIndex) => regenerateFromToken(message.id, tokenIndex)}
+										/>
+									{:else}
+										<!-- Normal markdown view -->
+										<MarkdownContent content={message.content || (loading ? response : '')} />
+									{/if}
 									{#if loading && !message.content}
 										<span class="inline-block w-2 h-4 bg-exo-yellow/70 ml-1 cursor-blink"></span>
 									{/if}
@@ -419,7 +453,20 @@ function isThinkingExpanded(messageId: string): boolean {
 								</svg>
 							</button>
 						{/if}
-						
+
+						<!-- Uncertainty view toggle (assistant messages with tokens only) -->
+						{#if message.role === 'assistant' && message.tokens && message.tokens.length > 0}
+							<button
+								onclick={() => toggleUncertaintyView(message.id)}
+								class="p-1.5 transition-colors rounded cursor-pointer {isUncertaintyViewEnabled(message.id) ? 'text-exo-yellow' : 'text-exo-light-gray hover:text-exo-yellow'}"
+								title={isUncertaintyViewEnabled(message.id) ? 'Hide uncertainty' : 'Show uncertainty'}
+							>
+								<svg class="w-3.5 h-3.5" fill="none" viewBox="0 0 24 24" stroke="currentColor">
+									<path stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M9 19v-6a2 2 0 00-2-2H5a2 2 0 00-2 2v6a2 2 0 002 2h2a2 2 0 002-2zm0 0V9a2 2 0 012-2h2a2 2 0 012 2v10m-6 0a2 2 0 002 2h2a2 2 0 002-2m0 0V5a2 2 0 012-2h2a2 2 0 012 2v14a2 2 0 01-2 2h-2a2 2 0 01-2-2z" />
+								</svg>
+							</button>
+						{/if}
+
 						<!-- Delete button -->
 						<button
 							onclick={() => handleDeleteClick(message.id)}
--- a/dashboard/src/lib/components/PrefillProgressBar.svelte
+++ b/dashboard/src/lib/components/PrefillProgressBar.svelte
@@ -0,0 +1,45 @@
+<script lang="ts">
+	import type { PrefillProgress } from '$lib/stores/app.svelte';
+
+	interface Props {
+		progress: PrefillProgress;
+		class?: string;
+	}
+
+	let { progress, class: className = '' }: Props = $props();
+
+	const percentage = $derived(
+		progress.total > 0 ? Math.round((progress.processed / progress.total) * 100) : 0
+	);
+
+	function formatTokenCount(count: number): string {
+		if (count >= 1000) {
+			return `${(count / 1000).toFixed(1)}k`;
+		}
+		return count.toString();
+	}
+</script>
+
+<div class="prefill-progress {className}">
+	<div class="flex items-center justify-between text-xs text-exo-light-gray mb-1">
+		<span>Processing prompt</span>
+		<span class="font-mono">
+			{formatTokenCount(progress.processed)} / {formatTokenCount(progress.total)} tokens
+		</span>
+	</div>
+	<div class="h-1.5 bg-exo-black/60 rounded-full overflow-hidden">
+		<div
+			class="h-full bg-exo-yellow rounded-full transition-all duration-150 ease-out"
+			style="width: {percentage}%"
+		></div>
+	</div>
+	<div class="text-right text-xs text-exo-light-gray/70 mt-0.5 font-mono">
+		{percentage}%
+	</div>
+</div>
+
+<style>
+	.prefill-progress {
+		width: 100%;
+	}
+</style>
--- a/dashboard/src/lib/components/TokenHeatmap.svelte
+++ b/dashboard/src/lib/components/TokenHeatmap.svelte
@@ -0,0 +1,192 @@
+<script lang="ts">
+	import type { TokenData } from '$lib/stores/app.svelte';
+
+	interface Props {
+		tokens: TokenData[];
+		class?: string;
+		isGenerating?: boolean;
+		onRegenerateFrom?: (tokenIndex: number) => void;
+	}
+
+	let { tokens, class: className = '', isGenerating = false, onRegenerateFrom }: Props = $props();
+
+	// Tooltip state - track both token data and index
+	let hoveredTokenIndex = $state<number | null>(null);
+	let hoveredPosition = $state<{ x: number; y: number } | null>(null);
+	let isTooltipHovered = $state(false);
+	let hideTimeoutId: ReturnType<typeof setTimeout> | null = null;
+
+	// Derive the hovered token from the index (stable across re-renders)
+	const hoveredToken = $derived(
+		hoveredTokenIndex !== null && hoveredPosition && tokens[hoveredTokenIndex]
+			? { token: tokens[hoveredTokenIndex], index: hoveredTokenIndex, ...hoveredPosition }
+			: null
+	);
+
+	/**
+	 * Get confidence styling based on probability.
+	 * Following Apple design principles: high confidence tokens blend in,
+	 * only uncertainty draws attention.
+	 */
+	function getConfidenceClass(probability: number): string {
+		if (probability > 0.8) return 'text-inherit'; // Expected tokens - blend in
+		if (probability > 0.5) return 'bg-gray-500/10 text-inherit'; // Slight hint
+		if (probability > 0.2) return 'bg-amber-500/15 text-amber-200/90'; // Subtle warmth
+		return 'bg-red-500/20 text-red-200/90'; // Draws attention
+	}
+
+	/**
+	 * Get border/underline styling for uncertain tokens
+	 */
+	function getBorderClass(probability: number): string {
+		if (probability > 0.8) return 'border-transparent'; // No border for expected
+		if (probability > 0.5) return 'border-gray-500/20';
+		if (probability > 0.2) return 'border-amber-500/30';
+		return 'border-red-500/40';
+	}
+
+	function clearHideTimeout() {
+		if (hideTimeoutId) {
+			clearTimeout(hideTimeoutId);
+			hideTimeoutId = null;
+		}
+	}
+
+	function handleMouseEnter(event: MouseEvent, token: TokenData, index: number) {
+		clearHideTimeout();
+		const rect = (event.target as HTMLElement).getBoundingClientRect();
+		hoveredTokenIndex = index;
+		hoveredPosition = {
+			x: rect.left + rect.width / 2,
+			y: rect.top - 10
+		};
+	}
+
+	function handleMouseLeave() {
+		clearHideTimeout();
+		// Use longer delay during generation to account for re-renders
+		const delay = isGenerating ? 300 : 100;
+		hideTimeoutId = setTimeout(() => {
+			if (!isTooltipHovered) {
+				hoveredTokenIndex = null;
+				hoveredPosition = null;
+			}
+		}, delay);
+	}
+
+	function handleTooltipEnter() {
+		clearHideTimeout();
+		isTooltipHovered = true;
+	}
+
+	function handleTooltipLeave() {
+		isTooltipHovered = false;
+		hoveredTokenIndex = null;
+		hoveredPosition = null;
+	}
+
+	function handleRegenerate() {
+		if (hoveredToken && onRegenerateFrom) {
+			const indexToRegenerate = hoveredToken.index;
+			// Clear hover state immediately
+			hoveredTokenIndex = null;
+			hoveredPosition = null;
+			isTooltipHovered = false;
+			// Call regenerate
+			onRegenerateFrom(indexToRegenerate);
+		}
+	}
+
+	function formatProbability(prob: number): string {
+		return (prob * 100).toFixed(1) + '%';
+	}
+
+	function formatLogprob(logprob: number): string {
+		return logprob.toFixed(3);
+	}
+
+	function getProbabilityColor(probability: number): string {
+		if (probability > 0.8) return 'text-gray-300';
+		if (probability > 0.5) return 'text-gray-400';
+		if (probability > 0.2) return 'text-amber-400';
+		return 'text-red-400';
+	}
+</script>
+
+<div class="token-heatmap leading-relaxed {className}">
+	{#each tokens as tokenData, i (i)}
+		<span
+			role="button"
+			tabindex="0"
+			class="token-span inline rounded px-0.5 py-0.5 cursor-pointer transition-all duration-150 border {getConfidenceClass(tokenData.probability)} {getBorderClass(tokenData.probability)} hover:opacity-80"
+			onmouseenter={(e) => handleMouseEnter(e, tokenData, i)}
+			onmouseleave={handleMouseLeave}
+		>{tokenData.token}</span>
+	{/each}
+</div>
+
+<!-- Tooltip -->
+{#if hoveredToken}
+	<div
+		class="fixed z-50"
+		style="left: {hoveredToken.x}px; top: {hoveredToken.y}px; transform: translate(-50%, -100%);"
+		onmouseenter={handleTooltipEnter}
+		onmouseleave={handleTooltipLeave}
+	>
+		<div class="bg-gray-900/95 backdrop-blur-sm border border-gray-700/50 rounded-xl shadow-xl p-3 text-sm min-w-48">
+			<!-- Token info -->
+			<div class="mb-2">
+				<span class="text-gray-500 text-xs">Token:</span>
+				<span class="text-white font-mono ml-1">"{hoveredToken.token.token}"</span>
+				<span class="{getProbabilityColor(hoveredToken.token.probability)} ml-2">{formatProbability(hoveredToken.token.probability)}</span>
+			</div>
+
+			<div class="text-gray-400 text-xs mb-1">
+				logprob: <span class="text-gray-300 font-mono">{formatLogprob(hoveredToken.token.logprob)}</span>
+			</div>
+
+			<!-- Top alternatives -->
+			{#if hoveredToken.token.topLogprobs.length > 0}
+				<div class="border-t border-gray-700/50 mt-2 pt-2">
+					<div class="text-gray-500 text-xs mb-1">Alternatives:</div>
+					{#each hoveredToken.token.topLogprobs.slice(0, 5) as alt, idx (idx)}
+						{@const altProb = Math.exp(alt.logprob)}
+						<div class="flex justify-between items-center text-xs py-0.5">
+							<span class="text-gray-300 font-mono truncate max-w-24">"{alt.token}"</span>
+							<span class="text-gray-400 ml-2">{formatProbability(altProb)}</span>
+						</div>
+					{/each}
+				</div>
+			{/if}
+
+			<!-- Regenerate button -->
+			{#if onRegenerateFrom}
+				<button
+					onclick={handleRegenerate}
+					class="w-full mt-2 pt-2 border-t border-gray-700/50 flex items-center justify-center gap-1.5 text-xs text-gray-400 hover:text-white transition-colors cursor-pointer"
+				>
+					<svg class="w-3 h-3" fill="none" viewBox="0 0 24 24" stroke="currentColor">
+						<path stroke-linecap="round" stroke-linejoin="round" stroke-width="2" d="M4 4v5h.582m15.356 2A8.001 8.001 0 004.582 9m0 0H9m11 11v-5h-.581m0 0a8.003 8.003 0 01-15.357-2m15.357 2H15" />
+					</svg>
+					Regenerate from here
+				</button>
+			{/if}
+		</div>
+		<!-- Arrow -->
+		<div class="absolute left-1/2 -translate-x-1/2 top-full">
+			<div class="border-8 border-transparent border-t-gray-900"></div>
+		</div>
+	</div>
+{/if}
+
+<style>
+	.token-heatmap {
+		word-wrap: break-word;
+		white-space: pre-wrap;
+	}
+
+	.token-span {
+		margin: 0;
+		border-width: 1px;
+	}
+</style>
--- a/dashboard/src/lib/stores/app.svelte.ts
+++ b/dashboard/src/lib/stores/app.svelte.ts
@@ -182,6 +182,26 @@ export interface MessageAttachment {
 	mimeType?: string;
 }

+// Token-level data for uncertainty visualization
+export interface TopLogprob {
+	token: string;
+	logprob: number;
+	bytes?: number[];
+}
+
+export interface TokenData {
+	token: string;
+	logprob: number;
+	probability: number; // exp(logprob)
+	topLogprobs: TopLogprob[];
+}
+
+// Prefill progress data for long prompts
+export interface PrefillProgress {
+	processed: number;
+	total: number;
+}
+
 export interface Message {
 	id: string;
 	role: "user" | "assistant" | "system";
@@ -191,6 +211,8 @@ export interface Message {
 	attachments?: MessageAttachment[];
 	ttftMs?: number; // Time to first token in ms (for assistant messages)
 	tps?: number; // Tokens per second (for assistant messages)
+	tokens?: TokenData[]; // Token-level data for uncertainty visualization
+	prefillProgress?: PrefillProgress | null; // Prefill progress for long prompts
 }

 export interface Conversation {
@@ -368,6 +390,21 @@ class AppStore {
 	private fetchInterval: ReturnType<typeof setInterval> | null = null;
 	private previewsInterval: ReturnType<typeof setInterval> | null = null;
 	private lastConversationPersistTs = 0;
+	private currentRequestController: AbortController | null = null;
+
+	/**
+	 * Abort any in-flight generation request
+	 */
+	abortCurrentRequest(): boolean {
+		if (this.currentRequestController) {
+			this.currentRequestController.abort();
+			this.currentRequestController = null;
+			this.isLoading = false;
+			this.currentResponse = "";
+			return true;
+		}
+		return false;
+	}

 	constructor() {
 		if (browser) {
@@ -405,12 +442,61 @@ class AppStore {

 	/**
 	 * Save conversations to localStorage
+	 * Note: We strip tokens (logprobs data) to save space - they're large and not essential for persistence
 	 */
 	private saveConversationsToStorage() {
 		try {
-			localStorage.setItem(STORAGE_KEY, JSON.stringify(this.conversations));
+			// Strip tokens from messages to save localStorage space
+			const conversationsToSave = this.conversations.map((conv) => ({
+				...conv,
+				messages: conv.messages.map((msg) => {
+					// eslint-disable-next-line @typescript-eslint/no-unused-vars
+					const { tokens, ...msgWithoutTokens } = msg;
+					return msgWithoutTokens;
+				}),
+			}));
+			localStorage.setItem(STORAGE_KEY, JSON.stringify(conversationsToSave));
 		} catch (error) {
 			console.error("Failed to save conversations:", error);
+			// If quota exceeded, try to clear old conversations and retry
+			if (
+				error instanceof DOMException &&
+				error.name === "QuotaExceededError"
+			) {
+				console.warn(
+					"Storage quota exceeded, clearing oldest conversations...",
+				);
+				this.pruneOldConversations();
+			}
+		}
+	}
+
+	/**
+	 * Remove oldest conversations to free up storage space
+	 */
+	private pruneOldConversations() {
+		if (this.conversations.length <= 1) return;
+
+		// Sort by updatedAt and remove oldest half
+		const sorted = [...this.conversations].sort(
+			(a, b) => (b.updatedAt || 0) - (a.updatedAt || 0),
+		);
+		const keepCount = Math.max(1, Math.ceil(sorted.length / 2));
+		this.conversations = sorted.slice(0, keepCount);
+
+		// Try saving again
+		try {
+			const conversationsToSave = this.conversations.map((conv) => ({
+				...conv,
+				messages: conv.messages.map((msg) => {
+					// eslint-disable-next-line @typescript-eslint/no-unused-vars
+					const { tokens, ...msgWithoutTokens } = msg;
+					return msgWithoutTokens;
+				}),
+			}));
+			localStorage.setItem(STORAGE_KEY, JSON.stringify(conversationsToSave));
+		} catch {
+			console.error("Still failed to save after pruning");
 		}
 	}

@@ -1331,6 +1417,11 @@ class AppStore {
 		const assistantMessage = this.addMessage("assistant", "");
 		this.updateActiveConversation();

+		// Create abort controller for this request - must be defined before try block
+		// so it's available in the finally block
+		const controller = new AbortController();
+		this.currentRequestController = controller;
+
 		try {
 			// Build the messages array for the API with system prompt
 			const systemPrompt = {
@@ -1408,7 +1499,10 @@ class AppStore {
 					messages: apiMessages,
 					temperature: 0.7,
 					stream: true,
+					logprobs: true,
+					top_logprobs: 5,
 				}),
+				signal: controller.signal,
 			});

 			if (!response.ok) {
@@ -1424,6 +1518,8 @@ class AppStore {
 			const decoder = new TextDecoder();
 			let fullContent = "";
 			let buffer = "";
+			const collectedTokens: TokenData[] = [];
+			let currentEventType = ""; // Track SSE event type

 			while (true) {
 				const { done, value } = await reader.read();
@@ -1437,22 +1533,59 @@ class AppStore {

 				for (const line of lines) {
 					const trimmed = line.trim();
-					if (!trimmed) continue;
+					if (!trimmed) {
+						// Empty line resets event type
+						currentEventType = "";
+						continue;
+					}
+
+					// Handle event type declaration
+					if (trimmed.startsWith("event: ")) {
+						currentEventType = trimmed.slice(7);
+						continue;
+					}

 					if (trimmed.startsWith("data: ")) {
 						const data = trimmed.slice(6);
-						if (data === "[DONE]") continue;
+						if (data === "[DONE]") {
+							currentEventType = "";
+							continue;
+						}

 						try {
 							const parsed = JSON.parse(data);
-							const tokenContent = parsed.choices?.[0]?.delta?.content;
-							if (tokenContent) {
+
+							// Handle prefill progress events
+							if (currentEventType === "prefill_progress") {
+								const idx = this.messages.findIndex(
+									(m) => m.id === assistantMessage.id,
+								);
+								if (idx !== -1) {
+									this.messages[idx].prefillProgress = {
+										processed: parsed.processed,
+										total: parsed.total,
+									};
+								}
+								continue;
+							}
+
+							// Handle regular token data
+							const delta = parsed.choices?.[0]?.delta?.content;
+							if (delta) {
 								// Track first token for TTFT
 								if (firstTokenTime === null) {
 									firstTokenTime = performance.now();
 									this.ttftMs = firstTokenTime - requestStartTime;
 								}

+								// Clear prefill progress when first token arrives
+								const msgIdx = this.messages.findIndex(
+									(m) => m.id === assistantMessage.id,
+								);
+								if (msgIdx !== -1 && this.messages[msgIdx].prefillProgress) {
+									this.messages[msgIdx].prefillProgress = null;
+								}
+
 								// Count tokens (each SSE chunk is typically one token)
 								tokenCount += 1;
 								this.totalTokens = tokenCount;
@@ -1463,7 +1596,30 @@ class AppStore {
 									this.tps = (tokenCount / elapsed) * 1000;
 								}

-								fullContent += tokenContent;
+								// Extract logprobs for uncertainty visualization
+								const logprobsData = parsed.choices?.[0]?.logprobs;
+								if (logprobsData?.content?.[0]) {
+									const logprobItem = logprobsData.content[0];
+									const tokenData: TokenData = {
+										token: logprobItem.token || delta,
+										logprob: logprobItem.logprob ?? 0,
+										probability: Math.exp(logprobItem.logprob ?? 0),
+										topLogprobs: (logprobItem.top_logprobs || []).map(
+											(item: {
+												token: string;
+												logprob: number;
+												bytes?: number[];
+											}) => ({
+												token: item.token,
+												logprob: item.logprob,
+												bytes: item.bytes,
+											}),
+										),
+									};
+									collectedTokens.push(tokenData);
+								}
+
+								fullContent += delta;

 								// Strip thinking tags for display and extract thinking content
 								const { displayContent, thinkingContent } =
@@ -1477,6 +1633,7 @@ class AppStore {
 								if (idx !== -1) {
 									this.messages[idx].content = displayContent;
 									this.messages[idx].thinking = thinkingContent || undefined;
+									this.messages[idx].tokens = [...collectedTokens];
 								}
 								this.persistActiveConversation();
 							}
@@ -1524,9 +1681,16 @@ class AppStore {
 				if (this.tps !== null) {
 					this.messages[idx].tps = this.tps;
 				}
+				if (collectedTokens.length > 0) {
+					this.messages[idx].tokens = collectedTokens;
+				}
 			}
 			this.persistActiveConversation();
 		} catch (error) {
+			// Don't show error for aborted requests (user cancelled)
+			if (error instanceof Error && error.name === "AbortError") {
+				return;
+			}
 			console.error("Error sending message:", error);
 			// Update the assistant message with error
 			const idx = this.messages.findIndex((m) => m.id === assistantMessage.id);
@@ -1536,6 +1700,237 @@ class AppStore {
 			}
 			this.persistActiveConversation();
 		} finally {
+			// Clean up controller if this is still the active request
+			if (this.currentRequestController === controller) {
+				this.currentRequestController = null;
+			}
+			this.isLoading = false;
+			this.currentResponse = "";
+			this.updateActiveConversation();
+		}
+	}
+
+	/**
+	 * Regenerate from a specific token in an assistant message.
+	 * Keeps content up to and including the specified token, then continues generation.
+	 * If a generation is already in progress, it will be aborted first.
+	 */
+	async regenerateFromToken(
+		messageId: string,
+		tokenIndex: number,
+	): Promise<void> {
+		// Abort any in-flight request first
+		this.abortCurrentRequest();
+
+		const messageIdx = this.messages.findIndex((m) => m.id === messageId);
+		if (messageIdx === -1) return;
+
+		const message = this.messages[messageIdx];
+		if (message.role !== "assistant" || !message.tokens) return;
+
+		// Get tokens up to and including the specified index
+		const tokensToKeep = message.tokens.slice(0, tokenIndex + 1);
+		const prefixText = tokensToKeep.map((t) => t.token).join("");
+
+		// Remove all messages after this assistant message
+		this.messages = this.messages.slice(0, messageIdx + 1);
+
+		// Update the message to show the prefix
+		this.messages[messageIdx].content = prefixText;
+		this.messages[messageIdx].tokens = tokensToKeep;
+
+		// Set up for continuation
+		this.isLoading = true;
+		this.currentResponse = prefixText;
+		this.ttftMs = null;
+		this.tps = null;
+		this.totalTokens = tokensToKeep.length;
+
+		// Create abort controller before try block so it's available in finally
+		const controller = new AbortController();
+		this.currentRequestController = controller;
+
+		try {
+			// Build messages for API - include the partial assistant message
+			const systemPrompt = {
+				role: "system" as const,
+				content:
+					"You are a helpful AI assistant. Respond directly and concisely. Do not show your reasoning or thought process.",
+			};
+
+			// Get all messages up to and including the one we're regenerating from
+			const apiMessages = [
+				systemPrompt,
+				...this.messages.map((m) => {
+					let msgContent = m.content;
+					if (m.attachments) {
+						for (const attachment of m.attachments) {
+							if (attachment.type === "text" && attachment.content) {
+								msgContent += `\n\n[File: ${attachment.name}]\n\`\`\`\n${attachment.content}\n\`\`\``;
+							}
+						}
+					}
+					return { role: m.role, content: msgContent };
+				}),
+			];
+
+			// Determine model
+			let modelToUse = this.selectedChatModel;
+			if (!modelToUse) {
+				for (const [, instanceWrapper] of Object.entries(this.instances)) {
+					if (instanceWrapper && typeof instanceWrapper === "object") {
+						const keys = Object.keys(
+							instanceWrapper as Record<string, unknown>,
+						);
+						if (keys.length === 1) {
+							const instance = (instanceWrapper as Record<string, unknown>)[
+								keys[0]
+							] as { shardAssignments?: { modelId?: string } };
+							if (instance?.shardAssignments?.modelId) {
+								modelToUse = instance.shardAssignments.modelId;
+								break;
+							}
+						}
+					}
+				}
+			}
+
+			if (!modelToUse) {
+				throw new Error("No model available");
+			}
+
+			// Start timing
+			const requestStartTime = performance.now();
+			let firstTokenTime: number | null = null;
+			let tokenCount = tokensToKeep.length;
+
+			const response = await fetch("/v1/chat/completions", {
+				method: "POST",
+				headers: { "Content-Type": "application/json" },
+				body: JSON.stringify({
+					model: modelToUse,
+					messages: apiMessages,
+					stream: true,
+					logprobs: true,
+					top_logprobs: 5,
+					continue_from_prefix: true,
+				}),
+				signal: controller.signal,
+			});
+
+			if (!response.ok) {
+				const errorText = await response.text();
+				throw new Error(`API error: ${response.status} - ${errorText}`);
+			}
+
+			const reader = response.body?.getReader();
+			if (!reader) throw new Error("No response body");
+
+			const decoder = new TextDecoder();
+			let fullContent = prefixText;
+			let buffer = "";
+			const collectedTokens: TokenData[] = [...tokensToKeep];
+
+			while (true) {
+				const { done, value } = await reader.read();
+				if (done) break;
+
+				buffer += decoder.decode(value, { stream: true });
+				const lines = buffer.split("\n");
+				buffer = lines.pop() || "";
+
+				for (const line of lines) {
+					const trimmed = line.trim();
+					if (!trimmed || trimmed === "data: [DONE]") continue;
+
+					if (trimmed.startsWith("data: ")) {
+						try {
+							const json = JSON.parse(trimmed.slice(6));
+							const delta = json.choices?.[0]?.delta?.content;
+							if (delta) {
+								if (firstTokenTime === null) {
+									firstTokenTime = performance.now();
+									this.ttftMs = firstTokenTime - requestStartTime;
+								}
+
+								tokenCount += 1;
+								this.totalTokens = tokenCount;
+
+								if (
+									firstTokenTime !== null &&
+									tokenCount > tokensToKeep.length
+								) {
+									const elapsed = performance.now() - firstTokenTime;
+									this.tps =
+										((tokenCount - tokensToKeep.length) / elapsed) * 1000;
+								}
+
+								// Extract logprobs
+								const logprobsData = json.choices?.[0]?.logprobs;
+								if (logprobsData?.content?.[0]) {
+									const logprobItem = logprobsData.content[0];
+									collectedTokens.push({
+										token: logprobItem.token || delta,
+										logprob: logprobItem.logprob ?? 0,
+										probability: Math.exp(logprobItem.logprob ?? 0),
+										topLogprobs: (logprobItem.top_logprobs || []).map(
+											(item: {
+												token: string;
+												logprob: number;
+												bytes?: number[];
+											}) => ({
+												token: item.token,
+												logprob: item.logprob,
+												bytes: item.bytes,
+											}),
+										),
+									});
+								}
+
+								fullContent += delta;
+								const { displayContent, thinkingContent } =
+									this.stripThinkingTags(fullContent);
+								this.currentResponse = displayContent;
+
+								this.messages[messageIdx].content = displayContent;
+								this.messages[messageIdx].thinking =
+									thinkingContent || undefined;
+								this.messages[messageIdx].tokens = [...collectedTokens];
+								this.persistActiveConversation();
+							}
+						} catch {
+							// Skip malformed JSON
+						}
+					}
+				}
+			}
+
+			// Final update
+			const { displayContent, thinkingContent } =
+				this.stripThinkingTags(fullContent);
+			this.messages[messageIdx].content = displayContent;
+			this.messages[messageIdx].thinking = thinkingContent || undefined;
+			this.messages[messageIdx].tokens = collectedTokens;
+
+			if (this.ttftMs !== null) {
+				this.messages[messageIdx].ttftMs = this.ttftMs;
+			}
+			if (this.tps !== null) {
+				this.messages[messageIdx].tps = this.tps;
+			}
+			this.persistActiveConversation();
+		} catch (error) {
+			if (error instanceof Error && error.name === "AbortError") {
+				return;
+			}
+			console.error("Error regenerating from token:", error);
+			this.messages[messageIdx].content =
+				`${prefixText}\n\nError: ${error instanceof Error ? error.message : "Unknown error"}`;
+			this.persistActiveConversation();
+		} finally {
+			if (this.currentRequestController === controller) {
+				this.currentRequestController = null;
+			}
 			this.isLoading = false;
 			this.currentResponse = "";
 			this.updateActiveConversation();
@@ -1615,6 +2010,8 @@ export const editMessage = (messageId: string, newContent: string) =>
 export const editAndRegenerate = (messageId: string, newContent: string) =>
 	appStore.editAndRegenerate(messageId, newContent);
 export const regenerateLastResponse = () => appStore.regenerateLastResponse();
+export const regenerateFromToken = (messageId: string, tokenIndex: number) =>
+	appStore.regenerateFromToken(messageId, tokenIndex);

 // Conversation actions
 export const conversations = () => appStore.conversations;
--- a/docs/imgs/dashboard-cluster-view.png
+++ b/docs/imgs/dashboard-cluster-view.png
--- a/pyproject.toml
+++ b/pyproject.toml
@@ -126,3 +126,6 @@ env = [
  "EXO_TESTS=1"
 ]
 addopts = "-m 'not slow'"
+filterwarnings = [
+    "ignore:builtin type Swig:DeprecationWarning",
+]
--- a/src/exo/main.py
+++ b/src/exo/main.py
@@ -205,6 +205,14 @@ def main():
    logger.info("Starting EXO")
    logger.info(f"EXO_LIBP2P_NAMESPACE: {os.getenv('EXO_LIBP2P_NAMESPACE')}")

+    # Set FAST_SYNCH override env var for runner subprocesses
+    if args.fast_synch is True:
+        os.environ["EXO_FAST_SYNCH"] = "on"
+        logger.info("FAST_SYNCH forced ON")
+    elif args.fast_synch is False:
+        os.environ["EXO_FAST_SYNCH"] = "off"
+        logger.info("FAST_SYNCH forced OFF")
+
    node = anyio.run(Node.create, args)
    anyio.run(node.run)
    logger.info("EXO Shutdown complete")
@@ -218,6 +226,7 @@ class Args(CamelCaseModel):
    api_port: PositiveInt = 52415
    tb_only: bool = False
    no_worker: bool = False
+    fast_synch: bool | None = None  # None = auto, True = force on, False = force off

    @classmethod
    def parse(cls) -> Self:
@@ -259,6 +268,20 @@ class Args(CamelCaseModel):
            "--no-worker",
            action="store_true",
        )
+        fast_synch_group = parser.add_mutually_exclusive_group()
+        fast_synch_group.add_argument(
+            "--fast-synch",
+            action="store_true",
+            dest="fast_synch",
+            default=None,
+            help="Force MLX FAST_SYNCH on (for JACCL backend)",
+        )
+        fast_synch_group.add_argument(
+            "--no-fast-synch",
+            action="store_false",
+            dest="fast_synch",
+            help="Force MLX FAST_SYNCH off",
+        )

        args = parser.parse_args()
        return cls(**vars(args))  # pyright: ignore[reportAny] - We are intentionally validating here, we can't do it statically
--- a/src/exo/master/adapters/init.py
+++ b/src/exo/master/adapters/init.py
@@ -0,0 +1 @@
+"""API adapters for different API formats (Claude, OpenAI Responses, etc.)."""
--- a/src/exo/master/adapters/chat_completions.py
+++ b/src/exo/master/adapters/chat_completions.py
@@ -0,0 +1,197 @@
+"""OpenAI Chat Completions API adapter for converting requests/responses."""
+
+import time
+from collections.abc import AsyncGenerator
+
+from loguru import logger
+
+from exo.shared.types.api import (
+    ChatCompletionChoice,
+    ChatCompletionMessage,
+    ChatCompletionMessageText,
+    ChatCompletionResponse,
+    ChatCompletionTaskParams,
+    ErrorInfo,
+    ErrorResponse,
+    FinishReason,
+    Logprobs,
+    LogprobsContentItem,
+    StreamingChoiceResponse,
+)
+from exo.shared.types.chunks import PrefillProgressData, StreamEvent, TokenChunk
+from exo.shared.types.common import CommandId
+from exo.shared.types.openai_responses import ResponseInputMessage, ResponsesRequest
+
+
+def chat_request_to_internal(request: ChatCompletionTaskParams) -> ResponsesRequest:
+    """Convert Chat Completions API request to ResponsesRequest (canonical internal format).
+
+    Extracts system message as instructions, converts messages to input.
+    """
+    instructions: str | None = None
+    input_messages: list[ResponseInputMessage] = []
+
+    for msg in request.messages:
+        # Normalize content to string
+        content: str
+        if msg.content is None:
+            content = ""
+        elif isinstance(msg.content, str):
+            content = msg.content
+        elif isinstance(msg.content, ChatCompletionMessageText):
+            content = msg.content.text
+        else:
+            # List of ChatCompletionMessageText
+            content = "\n".join(item.text for item in msg.content)
+
+        # Extract system message as instructions
+        if msg.role == "system":
+            if instructions is None:
+                instructions = content
+            else:
+                # Append additional system messages
+                instructions = f"{instructions}\n{content}"
+        else:
+            # Convert to ResponseInputMessage (only user, assistant, developer roles)
+            if msg.role in ("user", "assistant", "developer"):
+                input_messages.append(
+                    ResponseInputMessage(role=msg.role, content=content)
+                )
+
+    return ResponsesRequest(
+        model=request.model,
+        input=input_messages if input_messages else "",
+        instructions=instructions,
+        max_output_tokens=request.max_tokens,
+        temperature=request.temperature,
+        top_p=request.top_p,
+        top_k=request.top_k,
+        stop=request.stop,
+        seed=request.seed,
+        stream=request.stream,
+        tools=request.tools,
+        continue_from_prefix=request.continue_from_prefix,
+    )
+
+
+def chunk_to_response(
+    chunk: TokenChunk, command_id: CommandId
+) -> ChatCompletionResponse:
+    """Convert a TokenChunk to a streaming ChatCompletionResponse."""
+    # Build logprobs if available
+    logprobs: Logprobs | None = None
+    if chunk.logprob is not None:
+        logprobs = Logprobs(
+            content=[
+                LogprobsContentItem(
+                    token=chunk.text,
+                    logprob=chunk.logprob,
+                    top_logprobs=chunk.top_logprobs or [],
+                )
+            ]
+        )
+
+    return ChatCompletionResponse(
+        id=command_id,
+        created=int(time.time()),
+        model=chunk.model,
+        choices=[
+            StreamingChoiceResponse(
+                index=0,
+                delta=ChatCompletionMessage(role="assistant", content=chunk.text),
+                logprobs=logprobs,
+                finish_reason=chunk.finish_reason,
+            )
+        ],
+    )
+
+
+async def generate_chat_stream(
+    command_id: CommandId,
+    event_stream: AsyncGenerator[StreamEvent, None],
+) -> AsyncGenerator[str, None]:
+    """Generate Chat Completions API streaming events from StreamEvents.
+
+    Handles both TokenChunks (token generation) and PrefillProgressData (prefill progress).
+    """
+    try:
+        async for event in event_stream:
+            if isinstance(event, PrefillProgressData):
+                # Send prefill progress as a named SSE event
+                progress_json = f'{{"processed":{event.processed_tokens},"total":{event.total_tokens}}}'
+                yield f"event: prefill_progress\ndata: {progress_json}\n\n"
+                continue
+
+            # TokenChunk handling
+            chunk = event
+            if chunk.finish_reason == "error":
+                error_response = ErrorResponse(
+                    error=ErrorInfo(
+                        message=chunk.error_message or "Internal server error",
+                        type="InternalServerError",
+                        code=500,
+                    )
+                )
+                yield f"data: {error_response.model_dump_json()}\n\n"
+                yield "data: [DONE]\n\n"
+                logger.info(f"generate_chat_stream ending (error): {command_id}")
+                return
+
+            chunk_response = chunk_to_response(chunk, command_id)
+            yield f"data: {chunk_response.model_dump_json()}\n\n"
+
+            if chunk.finish_reason is not None:
+                logger.info(
+                    f"generate_chat_stream yielding [DONE] for finish_reason={chunk.finish_reason}: {command_id}"
+                )
+                yield "data: [DONE]\n\n"
+                logger.info(f"generate_chat_stream returning: {command_id}")
+                return
+    finally:
+        logger.info(f"generate_chat_stream finally block: {command_id}")
+
+
+async def collect_chat_response(
+    command_id: CommandId,
+    chunk_stream: AsyncGenerator[TokenChunk, None],
+) -> ChatCompletionResponse:
+    """Collect all token chunks and return a single ChatCompletionResponse."""
+    text_parts: list[str] = []
+    model: str | None = None
+    finish_reason: FinishReason | None = None
+    error_message: str | None = None
+
+    async for chunk in chunk_stream:
+        if chunk.finish_reason == "error":
+            error_message = chunk.error_message or "Internal server error"
+            break
+
+        if model is None:
+            model = chunk.model
+
+        text_parts.append(chunk.text)
+
+        if chunk.finish_reason is not None:
+            finish_reason = chunk.finish_reason
+
+    if error_message is not None:
+        raise ValueError(error_message)
+
+    combined_text = "".join(text_parts)
+    assert model is not None
+
+    return ChatCompletionResponse(
+        id=command_id,
+        created=int(time.time()),
+        model=model,
+        choices=[
+            ChatCompletionChoice(
+                index=0,
+                message=ChatCompletionMessage(
+                    role="assistant",
+                    content=combined_text,
+                ),
+                finish_reason=finish_reason,
+            )
+        ],
+    )
--- a/src/exo/master/adapters/claude.py
+++ b/src/exo/master/adapters/claude.py
@@ -0,0 +1,190 @@
+"""Claude Messages API adapter for converting requests/responses."""
+
+from collections.abc import AsyncGenerator
+
+from exo.shared.types.api import FinishReason
+from exo.shared.types.chunks import TokenChunk
+from exo.shared.types.claude_api import (
+    ClaudeContentBlockDeltaEvent,
+    ClaudeContentBlockStartEvent,
+    ClaudeContentBlockStopEvent,
+    ClaudeMessageDelta,
+    ClaudeMessageDeltaEvent,
+    ClaudeMessageDeltaUsage,
+    ClaudeMessagesRequest,
+    ClaudeMessagesResponse,
+    ClaudeMessageStart,
+    ClaudeMessageStartEvent,
+    ClaudeMessageStopEvent,
+    ClaudeStopReason,
+    ClaudeTextBlock,
+    ClaudeTextDelta,
+    ClaudeUsage,
+)
+from exo.shared.types.common import CommandId
+from exo.shared.types.openai_responses import ResponseInputMessage, ResponsesRequest
+
+
+def finish_reason_to_claude_stop_reason(
+    finish_reason: FinishReason | None,
+) -> ClaudeStopReason | None:
+    """Map OpenAI finish_reason to Claude stop_reason."""
+    if finish_reason is None:
+        return None
+    mapping: dict[FinishReason, ClaudeStopReason] = {
+        "stop": "end_turn",
+        "length": "max_tokens",
+        "tool_calls": "tool_use",
+        "content_filter": "end_turn",
+        "function_call": "tool_use",
+    }
+    return mapping.get(finish_reason, "end_turn")
+
+
+def claude_request_to_internal(request: ClaudeMessagesRequest) -> ResponsesRequest:
+    """Convert Claude Messages API request to ResponsesRequest (canonical internal format).
+
+    Converts Claude's system parameter to instructions,
+    and messages to input.
+    """
+    # Handle system message
+    instructions: str | None = None
+    if request.system:
+        if isinstance(request.system, str):
+            instructions = request.system
+        else:
+            # List of text blocks
+            instructions = "".join(block.text for block in request.system)
+
+    # Convert messages to input
+    input_messages: list[ResponseInputMessage] = []
+    for msg in request.messages:
+        content: str
+        if isinstance(msg.content, str):
+            content = msg.content
+        else:
+            # Concatenate text blocks (images not supported for MVP)
+            text_parts: list[str] = []
+            for block in msg.content:
+                if isinstance(block, ClaudeTextBlock):
+                    text_parts.append(block.text)
+            content = "".join(text_parts)
+
+        # Claude uses "user" and "assistant" roles
+        input_messages.append(ResponseInputMessage(role=msg.role, content=content))
+
+    return ResponsesRequest(
+        model=request.model,
+        input=input_messages if input_messages else "",
+        instructions=instructions,
+        max_output_tokens=request.max_tokens,
+        temperature=request.temperature,
+        top_p=request.top_p,
+        top_k=request.top_k,
+        stop=request.stop_sequences,
+        stream=request.stream,
+    )
+
+
+async def collect_claude_response(
+    command_id: CommandId,
+    model: str,
+    chunk_stream: AsyncGenerator[TokenChunk, None],
+) -> ClaudeMessagesResponse:
+    """Collect all token chunks and return a single ClaudeMessagesResponse."""
+    text_parts: list[str] = []
+    stop_reason: ClaudeStopReason | None = None
+    last_stats = None
+    error_message: str | None = None
+
+    async for chunk in chunk_stream:
+        if chunk.finish_reason == "error":
+            error_message = chunk.error_message or "Internal server error"
+            break
+
+        text_parts.append(chunk.text)
+        last_stats = chunk.stats or last_stats
+
+        if chunk.finish_reason is not None:
+            stop_reason = finish_reason_to_claude_stop_reason(chunk.finish_reason)
+
+    if error_message is not None:
+        raise ValueError(error_message)
+
+    combined_text = "".join(text_parts)
+
+    # Use actual usage data from stats if available
+    input_tokens = last_stats.prompt_tokens if last_stats else 0
+    output_tokens = last_stats.generation_tokens if last_stats else 0
+
+    return ClaudeMessagesResponse(
+        id=f"msg_{command_id}",
+        model=model,
+        content=[ClaudeTextBlock(text=combined_text)],
+        stop_reason=stop_reason,
+        usage=ClaudeUsage(
+            input_tokens=input_tokens,
+            output_tokens=output_tokens,
+        ),
+    )
+
+
+async def generate_claude_stream(
+    command_id: CommandId,
+    model: str,
+    chunk_stream: AsyncGenerator[TokenChunk, None],
+) -> AsyncGenerator[str, None]:
+    """Generate Claude Messages API streaming events from TokenChunks."""
+    # Initial message_start event
+    initial_message = ClaudeMessageStart(
+        id=f"msg_{command_id}",
+        model=model,
+        content=[],
+        stop_reason=None,
+        usage=ClaudeUsage(input_tokens=0, output_tokens=0),
+    )
+    start_event = ClaudeMessageStartEvent(message=initial_message)
+    yield f"event: message_start\ndata: {start_event.model_dump_json()}\n\n"
+
+    # content_block_start
+    block_start = ClaudeContentBlockStartEvent(
+        index=0, content_block=ClaudeTextBlock(text="")
+    )
+    yield f"event: content_block_start\ndata: {block_start.model_dump_json()}\n\n"
+
+    output_tokens = 0
+    stop_reason: ClaudeStopReason | None = None
+    last_stats = None
+
+    async for chunk in chunk_stream:
+        output_tokens += 1  # Count each chunk as one token
+        last_stats = chunk.stats or last_stats
+
+        # content_block_delta
+        delta_event = ClaudeContentBlockDeltaEvent(
+            index=0,
+            delta=ClaudeTextDelta(text=chunk.text),
+        )
+        yield f"event: content_block_delta\ndata: {delta_event.model_dump_json()}\n\n"
+
+        if chunk.finish_reason is not None:
+            stop_reason = finish_reason_to_claude_stop_reason(chunk.finish_reason)
+
+    # Use actual token count from stats if available
+    if last_stats is not None:
+        output_tokens = last_stats.generation_tokens
+
+    # content_block_stop
+    block_stop = ClaudeContentBlockStopEvent(index=0)
+    yield f"event: content_block_stop\ndata: {block_stop.model_dump_json()}\n\n"
+
+    # message_delta
+    message_delta = ClaudeMessageDeltaEvent(
+        delta=ClaudeMessageDelta(stop_reason=stop_reason),
+        usage=ClaudeMessageDeltaUsage(output_tokens=output_tokens),
+    )
+    yield f"event: message_delta\ndata: {message_delta.model_dump_json()}\n\n"
+
+    # message_stop
+    message_stop = ClaudeMessageStopEvent()
+    yield f"event: message_stop\ndata: {message_stop.model_dump_json()}\n\n"
--- a/src/exo/master/adapters/responses.py
+++ b/src/exo/master/adapters/responses.py
@@ -0,0 +1,173 @@
+"""OpenAI Responses API adapter for converting requests/responses.
+
+ResponsesRequest is the canonical internal format. Responses API is the most featureful,
+making it the natural choice for the internal format. All other API formats (Chat
+Completions, Claude) are converted TO ResponsesRequest.
+"""
+
+from collections.abc import AsyncGenerator
+
+from exo.shared.types.chunks import TokenChunk
+from exo.shared.types.common import CommandId
+from exo.shared.types.openai_responses import (
+    ResponseCompletedEvent,
+    ResponseContentPartAddedEvent,
+    ResponseContentPartDoneEvent,
+    ResponseCreatedEvent,
+    ResponseInProgressEvent,
+    ResponseMessageItem,
+    ResponseOutputItemAddedEvent,
+    ResponseOutputItemDoneEvent,
+    ResponseOutputText,
+    ResponsesResponse,
+    ResponseTextDeltaEvent,
+    ResponseTextDoneEvent,
+    ResponseUsage,
+)
+
+
+async def collect_responses_response(
+    command_id: CommandId,
+    model: str,
+    chunk_stream: AsyncGenerator[TokenChunk, None],
+) -> ResponsesResponse:
+    """Collect all token chunks and return a single ResponsesResponse."""
+    response_id = f"resp_{command_id}"
+    item_id = f"item_{command_id}"
+    accumulated_text = ""
+    last_stats = None
+    error_message: str | None = None
+
+    async for chunk in chunk_stream:
+        if chunk.finish_reason == "error":
+            error_message = chunk.error_message or "Internal server error"
+            break
+
+        accumulated_text += chunk.text
+        last_stats = chunk.stats or last_stats
+
+    if error_message is not None:
+        raise ValueError(error_message)
+
+    # Create usage from stats if available
+    usage = None
+    if last_stats is not None:
+        usage = ResponseUsage(
+            input_tokens=last_stats.prompt_tokens,
+            output_tokens=last_stats.generation_tokens,
+            total_tokens=last_stats.prompt_tokens + last_stats.generation_tokens,
+        )
+
+    output_item = ResponseMessageItem(
+        id=item_id,
+        content=[ResponseOutputText(text=accumulated_text)],
+        status="completed",
+    )
+
+    return ResponsesResponse(
+        id=response_id,
+        model=model,
+        status="completed",
+        output=[output_item],
+        output_text=accumulated_text,
+        usage=usage,
+    )
+
+
+async def generate_responses_stream(
+    command_id: CommandId,
+    model: str,
+    chunk_stream: AsyncGenerator[TokenChunk, None],
+) -> AsyncGenerator[str, None]:
+    """Generate OpenAI Responses API streaming events from TokenChunks."""
+    response_id = f"resp_{command_id}"
+    item_id = f"item_{command_id}"
+
+    # response.created
+    initial_response = ResponsesResponse(
+        id=response_id,
+        model=model,
+        status="in_progress",
+        output=[],
+        output_text="",
+    )
+    created_event = ResponseCreatedEvent(response=initial_response)
+    yield f"event: response.created\ndata: {created_event.model_dump_json()}\n\n"
+
+    # response.in_progress
+    in_progress_event = ResponseInProgressEvent(response=initial_response)
+    yield f"event: response.in_progress\ndata: {in_progress_event.model_dump_json()}\n\n"
+
+    # response.output_item.added
+    initial_item = ResponseMessageItem(
+        id=item_id,
+        content=[ResponseOutputText(text="")],
+        status="in_progress",
+    )
+    item_added = ResponseOutputItemAddedEvent(output_index=0, item=initial_item)
+    yield f"event: response.output_item.added\ndata: {item_added.model_dump_json()}\n\n"
+
+    # response.content_part.added
+    initial_part = ResponseOutputText(text="")
+    part_added = ResponseContentPartAddedEvent(
+        output_index=0, content_index=0, part=initial_part
+    )
+    yield f"event: response.content_part.added\ndata: {part_added.model_dump_json()}\n\n"
+
+    accumulated_text = ""
+    last_stats = None
+
+    async for chunk in chunk_stream:
+        accumulated_text += chunk.text
+        last_stats = chunk.stats or last_stats
+
+        # response.output_text.delta
+        delta_event = ResponseTextDeltaEvent(
+            output_index=0,
+            content_index=0,
+            delta=chunk.text,
+        )
+        yield f"event: response.output_text.delta\ndata: {delta_event.model_dump_json()}\n\n"
+
+    # response.output_text.done
+    text_done = ResponseTextDoneEvent(
+        output_index=0, content_index=0, text=accumulated_text
+    )
+    yield f"event: response.output_text.done\ndata: {text_done.model_dump_json()}\n\n"
+
+    # response.content_part.done
+    final_part = ResponseOutputText(text=accumulated_text)
+    part_done = ResponseContentPartDoneEvent(
+        output_index=0, content_index=0, part=final_part
+    )
+    yield f"event: response.content_part.done\ndata: {part_done.model_dump_json()}\n\n"
+
+    # response.output_item.done
+    final_item = ResponseMessageItem(
+        id=item_id,
+        content=[ResponseOutputText(text=accumulated_text)],
+        status="completed",
+    )
+    item_done = ResponseOutputItemDoneEvent(output_index=0, item=final_item)
+    yield f"event: response.output_item.done\ndata: {item_done.model_dump_json()}\n\n"
+
+    # Create usage from stats if available
+    usage = None
+    if last_stats is not None:
+        usage = ResponseUsage(
+            input_tokens=last_stats.prompt_tokens,
+            output_tokens=last_stats.generation_tokens,
+            total_tokens=last_stats.prompt_tokens + last_stats.generation_tokens,
+        )
+
+    # response.completed
+    final_response = ResponsesResponse(
+        id=response_id,
+        model=model,
+        status="completed",
+        output=[final_item],
+        output_text=accumulated_text,
+        usage=usage,
+    )
+    completed_event = ResponseCompletedEvent(response=final_response)
+    yield f"event: response.completed\ndata: {completed_event.model_dump_json()}\n\n"
--- a/src/exo/master/api.py
+++ b/src/exo/master/api.py
@@ -1,19 +1,35 @@
 import time
 from collections.abc import AsyncGenerator
+from http import HTTPStatus
 from typing import cast

 import anyio
-from anyio import create_task_group
+from anyio import BrokenResourceError, create_task_group
 from anyio.abc import TaskGroup
-from fastapi import FastAPI, HTTPException
+from fastapi import FastAPI, HTTPException, Request
 from fastapi.middleware.cors import CORSMiddleware
-from fastapi.responses import StreamingResponse
+from fastapi.responses import JSONResponse, StreamingResponse
 from fastapi.staticfiles import StaticFiles
 from hypercorn.asyncio import serve  # pyright: ignore[reportUnknownVariableType]
 from hypercorn.config import Config
 from hypercorn.typing import ASGIFramework
 from loguru import logger

+from exo.master.adapters.chat_completions import (
+    chat_request_to_internal,
+    chunk_to_response,
+    collect_chat_response,
+    generate_chat_stream,
+)
+from exo.master.adapters.claude import (
+    claude_request_to_internal,
+    collect_claude_response,
+    generate_claude_stream,
+)
+from exo.master.adapters.responses import (
+    collect_responses_response,
+    generate_responses_stream,
+)
 from exo.master.placement import place_instance as get_instance_placements
 from exo.shared.apply import apply
 from exo.shared.election import ElectionMessage
@@ -26,9 +42,12 @@ from exo.shared.types.api import (
    ChatCompletionChoice,
    ChatCompletionMessage,
    ChatCompletionResponse,
+    ChatCompletionTaskParams,
    CreateInstanceParams,
    CreateInstanceResponse,
    DeleteInstanceResponse,
+    ErrorInfo,
+    ErrorResponse,
    FinishReason,
    GenerationStats,
    ModelList,
@@ -36,9 +55,12 @@ from exo.shared.types.api import (
    PlaceInstanceParams,
    PlacementPreview,
    PlacementPreviewResponse,
-    StreamingChoiceResponse,
 )
-from exo.shared.types.chunks import TokenChunk
+from exo.shared.types.chunks import PrefillProgressData, StreamEvent, TokenChunk
+from exo.shared.types.claude_api import (
+    ClaudeMessagesRequest,
+    ClaudeMessagesResponse,
+)
 from exo.shared.types.commands import (
    ChatCompletion,
    Command,
@@ -49,11 +71,20 @@ from exo.shared.types.commands import (
    TaskFinished,
 )
 from exo.shared.types.common import CommandId, NodeId, SessionId
-from exo.shared.types.events import ChunkGenerated, Event, ForwarderEvent, IndexedEvent
+from exo.shared.types.events import (
+    ChunkGenerated,
+    Event,
+    ForwarderEvent,
+    IndexedEvent,
+    PrefillProgress,
+)
 from exo.shared.types.memory import Memory
 from exo.shared.types.models import ModelId, ModelMetadata
+from exo.shared.types.openai_responses import (
+    ResponsesRequest,
+    ResponsesResponse,
+)
 from exo.shared.types.state import State
-from exo.shared.types.tasks import ChatCompletionTaskParams
 from exo.shared.types.worker.instances import Instance, InstanceId, InstanceMeta
 from exo.shared.types.worker.shards import Sharding
 from exo.utils.banner import print_startup_banner
@@ -62,23 +93,6 @@ from exo.utils.dashboard_path import find_dashboard
 from exo.utils.event_buffer import OrderedBuffer


-def chunk_to_response(
-    chunk: TokenChunk, command_id: CommandId
-) -> ChatCompletionResponse:
-    return ChatCompletionResponse(
-        id=command_id,
-        created=int(time.time()),
-        model=chunk.model,
-        choices=[
-            StreamingChoiceResponse(
-                index=0,
-                delta=ChatCompletionMessage(role="assistant", content=chunk.text),
-                finish_reason=chunk.finish_reason,
-            )
-        ],
-    )
-
-
 async def resolve_model_meta(model_id: str) -> ModelMetadata:
    if model_id in MODEL_CARDS:
        model_card = MODEL_CARDS[model_id]
@@ -115,6 +129,7 @@ class API:
        self.paused_ev: anyio.Event = anyio.Event()

        self.app = FastAPI()
+        self._setup_exception_handlers()
        self._setup_cors()
        self._setup_routes()

@@ -127,7 +142,7 @@ class API:
            name="dashboard",
        )

-        self._chat_completion_queues: dict[CommandId, Sender[TokenChunk]] = {}
+        self._chat_completion_queues: dict[CommandId, Sender[StreamEvent]] = {}
        self._tg: TaskGroup | None = None

    def reset(self, new_session_id: SessionId, result_clock: int):
@@ -145,6 +160,21 @@ class API:
        self.paused_ev.set()
        self.paused_ev = anyio.Event()

+    def _setup_exception_handlers(self) -> None:
+        self.app.exception_handler(HTTPException)(self.http_exception_handler)
+
+    async def http_exception_handler(
+        self, _: Request, exc: HTTPException
+    ) -> JSONResponse:
+        err = ErrorResponse(
+            error=ErrorInfo(
+                message=exc.detail,
+                type=HTTPStatus(exc.status_code).phrase,
+                code=exc.status_code,
+            )
+        )
+        return JSONResponse(err.model_dump(), status_code=exc.status_code)
+
    def _setup_cors(self) -> None:
        self.app.add_middleware(
            CORSMiddleware,
@@ -168,6 +198,8 @@ class API:
            self.chat_completions
        )
        self.app.post("/bench/chat/completions")(self.bench_chat_completions)
+        self.app.post("/v1/messages", response_model=None)(self.claude_messages)
+        self.app.post("/v1/responses", response_model=None)(self.openai_responses)
        self.app.get("/state")(lambda: self.state)
        self.app.get("/events")(lambda: self._event_log)

@@ -373,18 +405,23 @@ class API:
            instance_id=instance_id,
        )

-    async def _chat_chunk_stream(
+    async def _stream_events(
        self, command_id: CommandId
-    ) -> AsyncGenerator[TokenChunk, None]:
-        """Yield `TokenChunk`s for a given command until completion."""
+    ) -> AsyncGenerator[StreamEvent, None]:
+        """Yield stream events (TokenChunks or PrefillProgressData) for a command.

+        This is the internal low-level stream used by all API adapters.
+        """
        try:
-            self._chat_completion_queues[command_id], recv = channel[TokenChunk]()
+            self._chat_completion_queues[command_id], recv = channel[StreamEvent]()

-            with recv as token_chunks:
-                async for chunk in token_chunks:
-                    yield chunk
-                    if chunk.finish_reason is not None:
+            with recv as events:
+                async for event in events:
+                    yield event
+                    if (
+                        isinstance(event, TokenChunk)
+                        and event.finish_reason is not None
+                    ):
                        break

        except anyio.get_cancelled_exc_class():
@@ -400,21 +437,36 @@ class API:
            await self._send(command)
            del self._chat_completion_queues[command_id]

+    async def _chat_chunk_stream(
+        self, command_id: CommandId
+    ) -> AsyncGenerator[TokenChunk, None]:
+        """Yield only TokenChunks, filtering out progress events."""
+
+        async for event in self._stream_events(command_id):
+            if isinstance(event, TokenChunk):
+                yield event
+
    async def _generate_chat_stream(
        self, command_id: CommandId
    ) -> AsyncGenerator[str, None]:
        """Generate chat completion stream as JSON strings."""

-        async for chunk in self._chat_chunk_stream(command_id):
-            chunk_response: ChatCompletionResponse = chunk_to_response(
-                chunk, command_id
-            )
-            logger.debug(f"chunk_response: {chunk_response}")
+        async for event in self._stream_events(command_id):
+            if isinstance(event, PrefillProgressData):
+                # Send prefill progress as a named SSE event
+                progress_json = f'{{"processed":{event.processed_tokens},"total":{event.total_tokens}}}'
+                yield f"event: prefill_progress\ndata: {progress_json}\n\n"
+            else:
+                # TokenChunk - regular token generation
+                chunk_response: ChatCompletionResponse = chunk_to_response(
+                    event, command_id
+                )
+                logger.debug(f"chunk_response: {chunk_response}")

-            yield f"data: {chunk_response.model_dump_json()}\n\n"
+                yield f"data: {chunk_response.model_dump_json()}\n\n"

-            if chunk.finish_reason is not None:
-                yield "data: [DONE]\n\n"
+                if event.finish_reason is not None:
+                    yield "data: [DONE]\n\n"

    async def _collect_chat_completion(
        self, command_id: CommandId
@@ -463,6 +515,12 @@ class API:
        stats: GenerationStats | None = None

        async for chunk in self._chat_chunk_stream(command_id):
+            if chunk.finish_reason == "error":
+                raise HTTPException(
+                    status_code=500,
+                    detail=chunk.error_message or "Internal server error",
+                )
+
            if model is None:
                model = chunk.model

@@ -500,54 +558,162 @@ class API:
    async def chat_completions(
        self, payload: ChatCompletionTaskParams
    ) -> ChatCompletionResponse | StreamingResponse:
-        """Handle chat completions, supporting both streaming and non-streaming responses."""
-        model_meta = await resolve_model_meta(payload.model)
-        payload.model = model_meta.model_id
+        """OpenAI Chat Completions API - adapter."""
+        internal_params = chat_request_to_internal(payload)
+        model_meta = await resolve_model_meta(internal_params.model)
+        internal_params.model = model_meta.model_id

        if not any(
-            instance.shard_assignments.model_id == payload.model
+            instance.shard_assignments.model_id == internal_params.model
            for instance in self.state.instances.values()
        ):
-            await self._trigger_notify_user_to_download_model(payload.model)
+            await self._trigger_notify_user_to_download_model(internal_params.model)
            raise HTTPException(
-                status_code=404, detail=f"No instance found for model {payload.model}"
+                status_code=404,
+                detail=f"No instance found for model {internal_params.model}",
            )

-        command = ChatCompletion(
-            request_params=payload,
-        )
+        command = ChatCompletion(request_params=internal_params)
        await self._send(command)
+
        if payload.stream:
            return StreamingResponse(
-                self._generate_chat_stream(command.command_id),
+                generate_chat_stream(
+                    command.command_id,
+                    self._stream_events(command.command_id),
+                ),
                media_type="text/event-stream",
+                headers={
+                    "Cache-Control": "no-cache",
+                    "Connection": "close",
+                    "X-Accel-Buffering": "no",
+                },
            )

-        return await self._collect_chat_completion(command.command_id)
+        try:
+            return await collect_chat_response(
+                command.command_id,
+                self._chat_chunk_stream(command.command_id),
+            )
+        except ValueError as e:
+            raise HTTPException(status_code=500, detail=str(e)) from e

    async def bench_chat_completions(
        self, payload: BenchChatCompletionTaskParams
    ) -> BenchChatCompletionResponse:
-        model_meta = await resolve_model_meta(payload.model)
-        payload.model = model_meta.model_id
+        # Convert to internal format (BenchChatCompletionTaskParams extends ChatCompletionTaskParams)
+        internal_params = chat_request_to_internal(payload)
+        model_meta = await resolve_model_meta(internal_params.model)
+        internal_params.model = model_meta.model_id

        if not any(
-            instance.shard_assignments.model_id == payload.model
+            instance.shard_assignments.model_id == internal_params.model
            for instance in self.state.instances.values()
        ):
-            await self._trigger_notify_user_to_download_model(payload.model)
+            await self._trigger_notify_user_to_download_model(internal_params.model)
            raise HTTPException(
-                status_code=404, detail=f"No instance found for model {payload.model}"
+                status_code=404,
+                detail=f"No instance found for model {internal_params.model}",
            )

-        payload.stream = False
+        internal_params.stream = False

-        command = ChatCompletion(request_params=payload)
+        command = ChatCompletion(request_params=internal_params)
        await self._send(command)

        response = await self._collect_chat_completion_with_stats(command.command_id)
        return response

+    async def claude_messages(
+        self, payload: ClaudeMessagesRequest
+    ) -> ClaudeMessagesResponse | StreamingResponse:
+        """Claude Messages API - adapter."""
+        internal_params = claude_request_to_internal(payload)
+        model_meta = await resolve_model_meta(internal_params.model)
+        internal_params.model = model_meta.model_id
+
+        if not any(
+            instance.shard_assignments.model_id == internal_params.model
+            for instance in self.state.instances.values()
+        ):
+            await self._trigger_notify_user_to_download_model(internal_params.model)
+            raise HTTPException(
+                status_code=404,
+                detail=f"No instance found for model {internal_params.model}",
+            )
+
+        command = ChatCompletion(request_params=internal_params)
+        await self._send(command)
+
+        if payload.stream:
+            return StreamingResponse(
+                generate_claude_stream(
+                    command.command_id,
+                    payload.model,
+                    self._chat_chunk_stream(command.command_id),
+                ),
+                media_type="text/event-stream",
+                headers={
+                    "Cache-Control": "no-cache",
+                    "Connection": "close",
+                    "X-Accel-Buffering": "no",
+                },
+            )
+
+        try:
+            return await collect_claude_response(
+                command.command_id,
+                payload.model,
+                self._chat_chunk_stream(command.command_id),
+            )
+        except ValueError as e:
+            raise HTTPException(status_code=500, detail=str(e)) from e
+
+    async def openai_responses(
+        self, payload: ResponsesRequest
+    ) -> ResponsesResponse | StreamingResponse:
+        """OpenAI Responses API - native format (no conversion needed)."""
+        model_meta = await resolve_model_meta(payload.model)
+        # Update model to resolved model_id
+        request_params = payload.model_copy(update={"model": model_meta.model_id})
+
+        if not any(
+            instance.shard_assignments.model_id == request_params.model
+            for instance in self.state.instances.values()
+        ):
+            await self._trigger_notify_user_to_download_model(request_params.model)
+            raise HTTPException(
+                status_code=404,
+                detail=f"No instance found for model {request_params.model}",
+            )
+
+        command = ChatCompletion(request_params=request_params)
+        await self._send(command)
+
+        if payload.stream:
+            return StreamingResponse(
+                generate_responses_stream(
+                    command.command_id,
+                    payload.model,
+                    self._chat_chunk_stream(command.command_id),
+                ),
+                media_type="text/event-stream",
+                headers={
+                    "Cache-Control": "no-cache",
+                    "Connection": "close",
+                    "X-Accel-Buffering": "no",
+                },
+            )
+
+        try:
+            return await collect_responses_response(
+                command.command_id,
+                payload.model,
+                self._chat_chunk_stream(command.command_id),
+            )
+        except ValueError as e:
+            raise HTTPException(status_code=500, detail=str(e)) from e
+
    def _calculate_total_available_memory(self) -> Memory:
        """Calculate total available memory across all nodes in bytes."""
        total_available = Memory()
@@ -607,14 +773,26 @@ class API:
                for idx, event in self.event_buffer.drain_indexed():
                    self._event_log.append(event)
                    self.state = apply(self.state, IndexedEvent(event=event, idx=idx))
-                    if (
-                        isinstance(event, ChunkGenerated)
-                        and event.command_id in self._chat_completion_queues
-                    ):
+                    if isinstance(event, ChunkGenerated):
                        assert isinstance(event.chunk, TokenChunk)
-                        await self._chat_completion_queues[event.command_id].send(
-                            event.chunk
-                        )
+                        queue = self._chat_completion_queues.get(event.command_id)
+                        if queue is not None:
+                            try:
+                                await queue.send(event.chunk)
+                            except BrokenResourceError:
+                                self._chat_completion_queues.pop(event.command_id, None)
+                    elif isinstance(event, PrefillProgress):
+                        queue = self._chat_completion_queues.get(event.command_id)
+                        if queue is not None:
+                            try:
+                                await queue.send(
+                                    PrefillProgressData(
+                                        processed_tokens=event.processed_tokens,
+                                        total_tokens=event.total_tokens,
+                                    )
+                                )
+                            except BrokenResourceError:
+                                self._chat_completion_queues.pop(event.command_id, None)

    async def _pause_on_new_election(self):
        with self.election_receiver as ems:
--- a/src/exo/master/placement_utils.py
+++ b/src/exo/master/placement_utils.py
@@ -49,33 +49,83 @@ def get_smallest_cycles(cycles: list[list[NodeInfo]]) -> list[list[NodeInfo]]:
    return [cycle for cycle in cycles if len(cycle) == min_nodes]


+def allocate_layers_proportionally(
+    total_layers: int,
+    memory_fractions: list[float],
+) -> list[int]:
+    n = len(memory_fractions)
+    if n == 0:
+        raise ValueError("Cannot allocate layers to an empty node list")
+    if total_layers < n:
+        raise ValueError(
+            f"Cannot distribute {total_layers} layers across {n} nodes "
+            "(need at least 1 layer per node)"
+        )
+
+    # Largest remainder: floor each, then distribute remainder by fractional part
+    raw = [f * total_layers for f in memory_fractions]
+    result = [int(r) for r in raw]
+    by_remainder = sorted(range(n), key=lambda i: raw[i] - result[i], reverse=True)
+    for i in range(total_layers - sum(result)):
+        result[by_remainder[i]] += 1
+
+    # Ensure minimum 1 per node by taking from the largest
+    for i in range(n):
+        if result[i] == 0:
+            max_idx = max(range(n), key=lambda j: result[j])
+            assert result[max_idx] > 1
+            result[max_idx] -= 1
+            result[i] = 1
+
+    return result
+
+
 def get_shard_assignments_for_pipeline_parallel(
    model_meta: ModelMetadata,
    selected_cycle: list[NodeWithProfile],
 ):
+    if not selected_cycle:
+        raise ValueError("Cannot create shard assignments for empty node cycle")
+
    cycle_memory = sum(
        (node.node_profile.memory.ram_available for node in selected_cycle),
        start=Memory(),
    )
+
+    if cycle_memory.in_bytes == 0:
+        raise ValueError("Cannot create shard assignments: total available memory is 0")
+
    total_layers = model_meta.n_layers
    world_size = len(selected_cycle)
    runner_to_shard: dict[RunnerId, ShardMetadata] = {}
    node_to_runner: dict[NodeId, RunnerId] = {}

-    layers_assigned = 0
-    for i, node in enumerate(selected_cycle):
-        if i == len(selected_cycle) - 1:
-            node_layers = total_layers - layers_assigned
-        else:
-            node_layers = round(
-                total_layers
-                * (
-                    node.node_profile.memory.ram_available.in_bytes
-                    / cycle_memory.in_bytes
-                )
-            )
-            node_layers = max(1, node_layers)
+    layer_allocations = allocate_layers_proportionally(
+        total_layers=total_layers,
+        memory_fractions=[
+            node.node_profile.memory.ram_available.in_bytes / cycle_memory.in_bytes
+            for node in selected_cycle
+        ],
+    )

+    # Validate each node has sufficient memory for its assigned layers
+    memory_per_layer = model_meta.storage_size.in_bytes / total_layers
+    for i, (node, node_layers) in enumerate(
+        zip(selected_cycle, layer_allocations, strict=True)
+    ):
+        required_memory = node_layers * memory_per_layer
+        available_memory = node.node_profile.memory.ram_available.in_bytes
+        if required_memory > available_memory:
+            raise ValueError(
+                f"Node {i} ({node.node_id}) has insufficient memory: "
+                f"requires {required_memory / (1024**3):.2f} GB for {node_layers} layers, "
+                f"but only has {available_memory / (1024**3):.2f} GB available"
+            )
+
+    layers_assigned = 0
+    for i, (node, node_layers) in enumerate(
+        zip(selected_cycle, layer_allocations, strict=True)
+    ):
        runner_id = RunnerId()

        shard = PipelineShardMetadata(
--- a/src/exo/master/tests/test_api_error_handling.py
+++ b/src/exo/master/tests/test_api_error_handling.py
@@ -0,0 +1,107 @@
+# pyright: reportUnusedFunction=false, reportAny=false
+from typing import Any, get_args
+
+from fastapi import FastAPI, HTTPException
+from fastapi.testclient import TestClient
+
+from exo.shared.types.api import ErrorInfo, ErrorResponse, FinishReason
+from exo.shared.types.chunks import TokenChunk
+from exo.worker.tests.constants import MODEL_A_ID
+
+
+def test_http_exception_handler_formats_openai_style() -> None:
+    """Test that HTTPException is converted to OpenAI-style error format."""
+    from exo.master.api import API
+
+    app = FastAPI()
+
+    # Setup exception handler
+    api = object.__new__(API)
+    api.app = app
+    api._setup_exception_handlers()  # pyright: ignore[reportPrivateUsage]
+
+    # Add test routes that raise HTTPException
+    @app.get("/test-error")
+    async def _test_error() -> None:
+        raise HTTPException(status_code=500, detail="Test error message")
+
+    @app.get("/test-not-found")
+    async def _test_not_found() -> None:
+        raise HTTPException(status_code=404, detail="Resource not found")
+
+    client = TestClient(app)
+
+    # Test 500 error
+    response = client.get("/test-error")
+    assert response.status_code == 500
+    data: dict[str, Any] = response.json()
+    assert "error" in data
+    assert data["error"]["message"] == "Test error message"
+    assert data["error"]["type"] == "Internal Server Error"
+    assert data["error"]["code"] == 500
+
+    # Test 404 error
+    response = client.get("/test-not-found")
+    assert response.status_code == 404
+    data = response.json()
+    assert "error" in data
+    assert data["error"]["message"] == "Resource not found"
+    assert data["error"]["type"] == "Not Found"
+    assert data["error"]["code"] == 404
+
+
+def test_finish_reason_includes_error() -> None:
+    valid_reasons = get_args(FinishReason)
+    assert "error" in valid_reasons
+
+
+def test_token_chunk_with_error_fields() -> None:
+    chunk = TokenChunk(
+        idx=0,
+        model=MODEL_A_ID,
+        text="",
+        token_id=0,
+        finish_reason="error",
+        error_message="Something went wrong",
+    )
+
+    assert chunk.finish_reason == "error"
+    assert chunk.error_message == "Something went wrong"
+
+
+def test_token_chunk_without_error() -> None:
+    chunk = TokenChunk(
+        idx=1,
+        model=MODEL_A_ID,
+        text="Hello",
+        token_id=42,
+        finish_reason=None,
+    )
+
+    assert chunk.finish_reason is None
+    assert chunk.error_message is None
+
+
+def test_error_response_construction() -> None:
+    error_response = ErrorResponse(
+        error=ErrorInfo(
+            message="Generation failed",
+            type="InternalServerError",
+            code=500,
+        )
+    )
+
+    assert error_response.error.message == "Generation failed"
+    assert error_response.error.code == 500
+
+
+def test_normal_finish_reasons_still_work() -> None:
+    for reason in ["stop", "length", "tool_calls", "content_filter", "function_call"]:
+        chunk = TokenChunk(
+            idx=0,
+            model=MODEL_A_ID,
+            text="done",
+            token_id=100,
+            finish_reason=reason,  # type: ignore[arg-type]
+        )
+        assert chunk.finish_reason == reason
--- a/src/exo/master/tests/test_claude_api.py
+++ b/src/exo/master/tests/test_claude_api.py
@@ -0,0 +1,283 @@
+"""Tests for Claude Messages API conversion functions and types."""
+
+import json
+from typing import Any, cast
+
+import pydantic
+import pytest
+
+from exo.master.adapters.claude import (
+    claude_request_to_internal,
+    finish_reason_to_claude_stop_reason,
+)
+from exo.shared.types.claude_api import (
+    ClaudeContentBlockDeltaEvent,
+    ClaudeContentBlockStartEvent,
+    ClaudeContentBlockStopEvent,
+    ClaudeMessage,
+    ClaudeMessageDelta,
+    ClaudeMessageDeltaEvent,
+    ClaudeMessageDeltaUsage,
+    ClaudeMessagesRequest,
+    ClaudeMessageStart,
+    ClaudeMessageStartEvent,
+    ClaudeMessageStopEvent,
+    ClaudeTextBlock,
+    ClaudeTextDelta,
+    ClaudeUsage,
+)
+
+
+class TestFinishReasonToClaudeStopReason:
+    """Tests for finish_reason to Claude stop_reason mapping."""
+
+    def test_stop_maps_to_end_turn(self):
+        assert finish_reason_to_claude_stop_reason("stop") == "end_turn"
+
+    def test_length_maps_to_max_tokens(self):
+        assert finish_reason_to_claude_stop_reason("length") == "max_tokens"
+
+    def test_tool_calls_maps_to_tool_use(self):
+        assert finish_reason_to_claude_stop_reason("tool_calls") == "tool_use"
+
+    def test_function_call_maps_to_tool_use(self):
+        assert finish_reason_to_claude_stop_reason("function_call") == "tool_use"
+
+    def test_content_filter_maps_to_end_turn(self):
+        assert finish_reason_to_claude_stop_reason("content_filter") == "end_turn"
+
+    def test_none_returns_none(self):
+        assert finish_reason_to_claude_stop_reason(None) is None
+
+
+class TestClaudeRequestToInternal:
+    """Tests for converting Claude Messages API requests to ResponsesRequest."""
+
+    def test_basic_request_conversion(self):
+        request = ClaudeMessagesRequest(
+            model="claude-3-opus",
+            max_tokens=100,
+            messages=[
+                ClaudeMessage(role="user", content="Hello"),
+            ],
+        )
+        params = claude_request_to_internal(request)
+
+        assert params.model == "claude-3-opus"
+        assert params.max_output_tokens == 100
+        assert isinstance(params.input, list)
+        assert len(params.input) == 1
+        assert params.input[0].role == "user"
+        assert params.input[0].content == "Hello"
+        assert params.instructions is None
+
+    def test_request_with_system_string(self):
+        request = ClaudeMessagesRequest(
+            model="claude-3-opus",
+            max_tokens=100,
+            system="You are a helpful assistant.",
+            messages=[
+                ClaudeMessage(role="user", content="Hello"),
+            ],
+        )
+        params = claude_request_to_internal(request)
+
+        assert params.instructions == "You are a helpful assistant."
+        assert isinstance(params.input, list)
+        assert len(params.input) == 1
+        assert params.input[0].role == "user"
+        assert params.input[0].content == "Hello"
+
+    def test_request_with_system_text_blocks(self):
+        request = ClaudeMessagesRequest(
+            model="claude-3-opus",
+            max_tokens=100,
+            system=[
+                ClaudeTextBlock(text="You are helpful. "),
+                ClaudeTextBlock(text="Be concise."),
+            ],
+            messages=[
+                ClaudeMessage(role="user", content="Hello"),
+            ],
+        )
+        params = claude_request_to_internal(request)
+
+        assert params.instructions == "You are helpful. Be concise."
+        assert isinstance(params.input, list)
+        assert len(params.input) == 1
+
+    def test_request_with_content_blocks(self):
+        request = ClaudeMessagesRequest(
+            model="claude-3-opus",
+            max_tokens=100,
+            messages=[
+                ClaudeMessage(
+                    role="user",
+                    content=[
+                        ClaudeTextBlock(text="First part. "),
+                        ClaudeTextBlock(text="Second part."),
+                    ],
+                ),
+            ],
+        )
+        params = claude_request_to_internal(request)
+
+        assert isinstance(params.input, list)
+        assert len(params.input) == 1
+        assert params.input[0].content == "First part. Second part."
+
+    def test_request_with_multi_turn_conversation(self):
+        request = ClaudeMessagesRequest(
+            model="claude-3-opus",
+            max_tokens=100,
+            messages=[
+                ClaudeMessage(role="user", content="Hello"),
+                ClaudeMessage(role="assistant", content="Hi there!"),
+                ClaudeMessage(role="user", content="How are you?"),
+            ],
+        )
+        params = claude_request_to_internal(request)
+
+        assert isinstance(params.input, list)
+        assert len(params.input) == 3
+        assert params.input[0].role == "user"
+        assert params.input[1].role == "assistant"
+        assert params.input[2].role == "user"
+
+    def test_request_with_optional_parameters(self):
+        request = ClaudeMessagesRequest(
+            model="claude-3-opus",
+            max_tokens=100,
+            messages=[ClaudeMessage(role="user", content="Hello")],
+            temperature=0.7,
+            top_p=0.9,
+            top_k=40,
+            stop_sequences=["STOP", "END"],
+            stream=True,
+        )
+        params = claude_request_to_internal(request)
+
+        assert params.temperature == 0.7
+        assert params.top_p == 0.9
+        assert params.top_k == 40
+        assert params.stop == ["STOP", "END"]
+        assert params.stream is True
+
+
+class TestClaudeMessagesRequestValidation:
+    """Tests for Claude Messages API request validation."""
+
+    def test_request_requires_model(self):
+        with pytest.raises(pydantic.ValidationError):
+            ClaudeMessagesRequest.model_validate(
+                {
+                    "max_tokens": 100,
+                    "messages": [{"role": "user", "content": "Hello"}],
+                }
+            )
+
+    def test_request_requires_max_tokens(self):
+        with pytest.raises(pydantic.ValidationError):
+            ClaudeMessagesRequest.model_validate(
+                {
+                    "model": "claude-3-opus",
+                    "messages": [{"role": "user", "content": "Hello"}],
+                }
+            )
+
+    def test_request_requires_messages(self):
+        with pytest.raises(pydantic.ValidationError):
+            ClaudeMessagesRequest.model_validate(
+                {
+                    "model": "claude-3-opus",
+                    "max_tokens": 100,
+                }
+            )
+
+
+class TestClaudeStreamingEvents:
+    """Tests for Claude Messages API streaming event serialization."""
+
+    def test_message_start_event_format(self):
+        message = ClaudeMessageStart(
+            id="msg_123",
+            model="claude-3-opus",
+            content=[],
+            stop_reason=None,
+            usage=ClaudeUsage(input_tokens=10, output_tokens=0),
+        )
+        event = ClaudeMessageStartEvent(message=message)
+        json_str = event.model_dump_json()
+        parsed = cast(dict[str, Any], json.loads(json_str))
+
+        assert parsed["type"] == "message_start"
+        assert parsed["message"]["id"] == "msg_123"
+        assert parsed["message"]["type"] == "message"
+        assert parsed["message"]["role"] == "assistant"
+        assert parsed["message"]["model"] == "claude-3-opus"
+
+    def test_content_block_start_event_format(self):
+        event = ClaudeContentBlockStartEvent(
+            index=0,
+            content_block=ClaudeTextBlock(text=""),
+        )
+        json_str = event.model_dump_json()
+        parsed = cast(dict[str, Any], json.loads(json_str))
+
+        assert parsed["type"] == "content_block_start"
+        assert parsed["index"] == 0
+        assert parsed["content_block"]["type"] == "text"
+        assert parsed["content_block"]["text"] == ""
+
+    def test_content_block_delta_event_format(self):
+        event = ClaudeContentBlockDeltaEvent(
+            index=0,
+            delta=ClaudeTextDelta(text="Hello"),
+        )
+        json_str = event.model_dump_json()
+        parsed = cast(dict[str, Any], json.loads(json_str))
+
+        assert parsed["type"] == "content_block_delta"
+        assert parsed["index"] == 0
+        assert parsed["delta"]["type"] == "text_delta"
+        assert parsed["delta"]["text"] == "Hello"
+
+    def test_content_block_stop_event_format(self):
+        event = ClaudeContentBlockStopEvent(index=0)
+        json_str = event.model_dump_json()
+        parsed = cast(dict[str, Any], json.loads(json_str))
+
+        assert parsed["type"] == "content_block_stop"
+        assert parsed["index"] == 0
+
+    def test_message_delta_event_format(self):
+        event = ClaudeMessageDeltaEvent(
+            delta=ClaudeMessageDelta(stop_reason="end_turn"),
+            usage=ClaudeMessageDeltaUsage(output_tokens=25),
+        )
+        json_str = event.model_dump_json()
+        parsed = cast(dict[str, Any], json.loads(json_str))
+
+        assert parsed["type"] == "message_delta"
+        assert parsed["delta"]["stop_reason"] == "end_turn"
+        assert parsed["usage"]["output_tokens"] == 25
+
+    def test_message_stop_event_format(self):
+        event = ClaudeMessageStopEvent()
+        json_str = event.model_dump_json()
+        parsed = cast(dict[str, Any], json.loads(json_str))
+
+        assert parsed["type"] == "message_stop"
+
+    def test_sse_format(self):
+        """Test that SSE format is correctly generated."""
+        event = ClaudeContentBlockDeltaEvent(
+            index=0,
+            delta=ClaudeTextDelta(text="Hello"),
+        )
+        # Simulate the SSE format used in the streaming generator
+        sse_line = f"event: content_block_delta\ndata: {event.model_dump_json()}\n\n"
+
+        assert sse_line.startswith("event: content_block_delta\n")
+        assert "data: " in sse_line
+        assert sse_line.endswith("\n\n")
--- a/src/exo/master/tests/test_master.py
+++ b/src/exo/master/tests/test_master.py
@@ -7,7 +7,6 @@ from loguru import logger

 from exo.master.main import Master
 from exo.routing.router import get_node_id_keypair
-from exo.shared.types.api import ChatCompletionMessage, ChatCompletionTaskParams
 from exo.shared.types.commands import (
    ChatCompletion,
    CommandId,
@@ -24,6 +23,7 @@ from exo.shared.types.events import (
 )
 from exo.shared.types.memory import Memory
 from exo.shared.types.models import ModelId, ModelMetadata
+from exo.shared.types.openai_responses import ResponsesRequest
 from exo.shared.types.profiling import (
    MemoryPerformanceProfile,
    NodePerformanceProfile,
@@ -143,13 +143,9 @@ async def test_master():
                command=(
                    ChatCompletion(
                        command_id=CommandId(),
-                        request_params=ChatCompletionTaskParams(
+                        request_params=ResponsesRequest(
                            model="llama-3.2-1b",
-                            messages=[
-                                ChatCompletionMessage(
-                                    role="user", content="Hello, how are you?"
-                                )
-                            ],
+                            input="Hello, how are you?",
                        ),
                    )
                ),
@@ -200,11 +196,9 @@ async def test_master():
        assert isinstance(events[2].event, TaskCreated)
        assert events[2].event.task.task_status == TaskStatus.Pending
        assert isinstance(events[2].event.task, ChatCompletionTask)
-        assert events[2].event.task.task_params == ChatCompletionTaskParams(
+        assert events[2].event.task.task_params == ResponsesRequest(
            model="llama-3.2-1b",
-            messages=[
-                ChatCompletionMessage(role="user", content="Hello, how are you?")
-            ],
+            input="Hello, how are you?",
        )

        await master.shutdown()
--- a/src/exo/master/tests/test_openai_responses_api.py
+++ b/src/exo/master/tests/test_openai_responses_api.py
@@ -0,0 +1,293 @@
+"""Tests for OpenAI Responses API types.
+
+ResponsesRequest is the canonical internal type used throughout the pipeline.
+No conversion is needed for Responses API requests.
+"""
+
+import json
+from typing import Any, cast
+
+import pydantic
+import pytest
+
+from exo.shared.types.openai_responses import (
+    ResponseCompletedEvent,
+    ResponseContentPartAddedEvent,
+    ResponseCreatedEvent,
+    ResponseInputMessage,
+    ResponseMessageItem,
+    ResponseOutputItemAddedEvent,
+    ResponseOutputItemDoneEvent,
+    ResponseOutputText,
+    ResponsesRequest,
+    ResponsesResponse,
+    ResponseTextDeltaEvent,
+    ResponseTextDoneEvent,
+    ResponseUsage,
+)
+
+
+class TestResponsesRequestAsCanonicalType:
+    """Tests for ResponsesRequest as the canonical internal type."""
+
+    def test_string_input(self):
+        request = ResponsesRequest(
+            model="gpt-4o",
+            input="Hello, how are you?",
+        )
+
+        assert request.model == "gpt-4o"
+        assert request.input == "Hello, how are you?"
+        assert request.instructions is None
+
+    def test_message_array_input(self):
+        request = ResponsesRequest(
+            model="gpt-4o",
+            input=[
+                ResponseInputMessage(role="user", content="Hello"),
+                ResponseInputMessage(role="assistant", content="Hi there!"),
+                ResponseInputMessage(role="user", content="How are you?"),
+            ],
+        )
+
+        assert isinstance(request.input, list)
+        assert len(request.input) == 3
+        assert request.input[0].role == "user"
+        assert request.input[0].content == "Hello"
+        assert request.input[1].role == "assistant"
+        assert request.input[1].content == "Hi there!"
+        assert request.input[2].role == "user"
+        assert request.input[2].content == "How are you?"
+
+    def test_request_with_instructions(self):
+        request = ResponsesRequest(
+            model="gpt-4o",
+            input="Hello",
+            instructions="You are a helpful assistant. Be concise.",
+        )
+
+        assert request.input == "Hello"
+        assert request.instructions == "You are a helpful assistant. Be concise."
+
+    def test_request_with_optional_parameters(self):
+        request = ResponsesRequest(
+            model="gpt-4o",
+            input="Hello",
+            max_output_tokens=500,
+            temperature=0.8,
+            top_p=0.95,
+            stream=True,
+        )
+
+        assert request.max_output_tokens == 500
+        assert request.temperature == 0.8
+        assert request.top_p == 0.95
+        assert request.stream is True
+
+    def test_request_with_new_fields(self):
+        """Test the additional fields added for internal use."""
+        request = ResponsesRequest(
+            model="gpt-4o",
+            input="Hello",
+            top_k=40,
+            seed=42,
+            stop=["STOP", "END"],
+            tools=[{"type": "function", "function": {"name": "test"}}],
+        )
+
+        assert request.top_k == 40
+        assert request.seed == 42
+        assert request.stop == ["STOP", "END"]
+        assert request.tools == [{"type": "function", "function": {"name": "test"}}]
+
+    def test_request_with_system_role_in_messages(self):
+        request = ResponsesRequest(
+            model="gpt-4o",
+            input=[
+                ResponseInputMessage(role="system", content="Be helpful"),
+                ResponseInputMessage(role="user", content="Hello"),
+            ],
+        )
+
+        assert isinstance(request.input, list)
+        assert len(request.input) == 2
+        assert request.input[0].role == "system"
+        assert request.input[1].role == "user"
+
+    def test_request_with_developer_role(self):
+        request = ResponsesRequest(
+            model="gpt-4o",
+            input=[
+                ResponseInputMessage(role="developer", content="Internal note"),
+                ResponseInputMessage(role="user", content="Hello"),
+            ],
+        )
+
+        assert isinstance(request.input, list)
+        assert len(request.input) == 2
+        assert request.input[0].role == "developer"
+
+
+class TestResponsesRequestValidation:
+    """Tests for OpenAI Responses API request validation."""
+
+    def test_request_requires_model(self):
+        with pytest.raises(pydantic.ValidationError):
+            ResponsesRequest.model_validate(
+                {
+                    "input": "Hello",
+                }
+            )
+
+    def test_request_requires_input(self):
+        with pytest.raises(pydantic.ValidationError):
+            ResponsesRequest.model_validate(
+                {
+                    "model": "gpt-4o",
+                }
+            )
+
+    def test_request_accepts_string_input(self):
+        request = ResponsesRequest(
+            model="gpt-4o",
+            input="Hello",
+        )
+        assert request.input == "Hello"
+
+    def test_request_accepts_message_array_input(self):
+        request = ResponsesRequest(
+            model="gpt-4o",
+            input=[ResponseInputMessage(role="user", content="Hello")],
+        )
+        assert len(request.input) == 1
+
+
+class TestResponsesStreamingEvents:
+    """Tests for OpenAI Responses API streaming event serialization."""
+
+    def test_response_created_event_format(self):
+        response = ResponsesResponse(
+            id="resp_123",
+            model="gpt-4o",
+            status="in_progress",
+            output=[],
+            output_text="",
+        )
+        event = ResponseCreatedEvent(response=response)
+        json_str = event.model_dump_json()
+        parsed = cast(dict[str, Any], json.loads(json_str))
+
+        assert parsed["type"] == "response.created"
+        assert parsed["response"]["id"] == "resp_123"
+        assert parsed["response"]["object"] == "response"
+        assert parsed["response"]["status"] == "in_progress"
+
+    def test_output_item_added_event_format(self):
+        item = ResponseMessageItem(
+            id="item_123",
+            content=[ResponseOutputText(text="")],
+            status="in_progress",
+        )
+        event = ResponseOutputItemAddedEvent(output_index=0, item=item)
+        json_str = event.model_dump_json()
+        parsed = cast(dict[str, Any], json.loads(json_str))
+
+        assert parsed["type"] == "response.output_item.added"
+        assert parsed["output_index"] == 0
+        assert parsed["item"]["type"] == "message"
+        assert parsed["item"]["id"] == "item_123"
+        assert parsed["item"]["role"] == "assistant"
+
+    def test_content_part_added_event_format(self):
+        part = ResponseOutputText(text="")
+        event = ResponseContentPartAddedEvent(
+            output_index=0,
+            content_index=0,
+            part=part,
+        )
+        json_str = event.model_dump_json()
+        parsed = cast(dict[str, Any], json.loads(json_str))
+
+        assert parsed["type"] == "response.content_part.added"
+        assert parsed["output_index"] == 0
+        assert parsed["content_index"] == 0
+        assert parsed["part"]["type"] == "output_text"
+
+    def test_text_delta_event_format(self):
+        event = ResponseTextDeltaEvent(
+            output_index=0,
+            content_index=0,
+            delta="Hello",
+        )
+        json_str = event.model_dump_json()
+        parsed = cast(dict[str, Any], json.loads(json_str))
+
+        assert parsed["type"] == "response.output_text.delta"
+        assert parsed["output_index"] == 0
+        assert parsed["content_index"] == 0
+        assert parsed["delta"] == "Hello"
+
+    def test_text_done_event_format(self):
+        event = ResponseTextDoneEvent(
+            output_index=0,
+            content_index=0,
+            text="Hello, world!",
+        )
+        json_str = event.model_dump_json()
+        parsed = cast(dict[str, Any], json.loads(json_str))
+
+        assert parsed["type"] == "response.output_text.done"
+        assert parsed["text"] == "Hello, world!"
+
+    def test_output_item_done_event_format(self):
+        item = ResponseMessageItem(
+            id="item_123",
+            content=[ResponseOutputText(text="Hello, world!")],
+            status="completed",
+        )
+        event = ResponseOutputItemDoneEvent(output_index=0, item=item)
+        json_str = event.model_dump_json()
+        parsed = cast(dict[str, Any], json.loads(json_str))
+
+        assert parsed["type"] == "response.output_item.done"
+        assert parsed["item"]["status"] == "completed"
+        assert parsed["item"]["content"][0]["text"] == "Hello, world!"
+
+    def test_response_completed_event_format(self):
+        item = ResponseMessageItem(
+            id="item_123",
+            content=[ResponseOutputText(text="Hello!")],
+            status="completed",
+        )
+        response = ResponsesResponse(
+            id="resp_123",
+            model="gpt-4o",
+            status="completed",
+            output=[item],
+            output_text="Hello!",
+            usage=ResponseUsage(input_tokens=10, output_tokens=5, total_tokens=15),
+        )
+        event = ResponseCompletedEvent(response=response)
+        json_str = event.model_dump_json()
+        parsed = cast(dict[str, Any], json.loads(json_str))
+
+        assert parsed["type"] == "response.completed"
+        assert parsed["response"]["status"] == "completed"
+        assert parsed["response"]["output_text"] == "Hello!"
+        assert parsed["response"]["usage"]["total_tokens"] == 15
+
+    def test_sse_format(self):
+        """Test that SSE format is correctly generated."""
+        event = ResponseTextDeltaEvent(
+            output_index=0,
+            content_index=0,
+            delta="Hello",
+        )
+        # Simulate the SSE format used in the streaming generator
+        sse_line = (
+            f"event: response.output_text.delta\ndata: {event.model_dump_json()}\n\n"
+        )
+
+        assert sse_line.startswith("event: response.output_text.delta\n")
+        assert "data: " in sse_line
+        assert sse_line.endswith("\n\n")
--- a/src/exo/master/tests/test_placement.py
+++ b/src/exo/master/tests/test_placement.py
@@ -70,7 +70,7 @@ def place_instance_command(model_meta: ModelMetadata) -> PlaceInstance:
    [
        ((500, 500, 1000), 12, (3, 3, 6)),
        ((500, 500, 500), 12, (4, 4, 4)),
-        ((312, 518, 1024), 12, (2, 3, 7)),
+        ((312, 468, 1092), 12, (2, 3, 7)),
    ],
 )
 def test_get_instance_placements_create_instance(
--- a/src/exo/master/tests/test_placement_utils.py
+++ b/src/exo/master/tests/test_placement_utils.py
@@ -3,6 +3,7 @@ from typing import Callable
 import pytest

 from exo.master.placement_utils import (
+    allocate_layers_proportionally,
    filter_cycles_by_memory,
    get_hosts_from_subgraph,
    get_mlx_jaccl_coordinators,
@@ -165,6 +166,9 @@ def test_get_smallest_cycles(
        ((500, 500, 1000), 12, (3, 3, 6)),
        ((500, 500, 500), 12, (4, 4, 4)),
        ((312, 518, 1024), 12, (2, 3, 7)),
+        # Edge case: one node has ~90% of memory - should not over-allocate.
+        # Each node must have enough memory for at least 1 layer (50 KB = 1000/20).
+        ((900, 50, 50), 20, (18, 1, 1)),
    ],
 )
 def test_get_shard_assignments(
@@ -397,3 +401,96 @@ def test_get_mlx_jaccl_coordinators(
    assert coordinators[node_c_id] == (
        f"{conn_c_a.send_back_multiaddr.ip_address}:5000"
    ), "node_c should use the IP from conn_c_a"
+
+
+class TestAllocateLayersProportionally:
+    def test_empty_node_list_raises(self):
+        with pytest.raises(ValueError, match="empty node list"):
+            allocate_layers_proportionally(total_layers=10, memory_fractions=[])
+
+    def test_zero_layers_raises(self):
+        with pytest.raises(ValueError, match="need at least 1 layer per node"):
+            allocate_layers_proportionally(total_layers=0, memory_fractions=[0.5, 0.5])
+
+    def test_negative_layers_raises(self):
+        with pytest.raises(ValueError, match="need at least 1 layer per node"):
+            allocate_layers_proportionally(total_layers=-1, memory_fractions=[0.5, 0.5])
+
+    def test_fewer_layers_than_nodes_raises(self):
+        with pytest.raises(ValueError, match="need at least 1 layer per node"):
+            allocate_layers_proportionally(
+                total_layers=2, memory_fractions=[0.33, 0.33, 0.34]
+            )
+
+    def test_equal_distribution(self):
+        result = allocate_layers_proportionally(
+            total_layers=12, memory_fractions=[0.25, 0.25, 0.25, 0.25]
+        )
+        assert result == [3, 3, 3, 3]
+        assert sum(result) == 12
+
+    def test_proportional_distribution(self):
+        result = allocate_layers_proportionally(
+            total_layers=12, memory_fractions=[0.25, 0.25, 0.50]
+        )
+        assert result == [3, 3, 6]
+        assert sum(result) == 12
+
+    def test_extreme_imbalance_ensures_minimum(self):
+        result = allocate_layers_proportionally(
+            total_layers=20, memory_fractions=[0.975, 0.0125, 0.0125]
+        )
+        assert all(layers >= 1 for layers in result)
+        assert sum(result) == 20
+        # Small nodes get minimum 1 layer
+        assert result == [18, 1, 1]
+
+    def test_single_node_gets_all_layers(self):
+        result = allocate_layers_proportionally(total_layers=10, memory_fractions=[1.0])
+        assert result == [10]
+
+    def test_minimum_viable_allocation(self):
+        result = allocate_layers_proportionally(
+            total_layers=3, memory_fractions=[0.33, 0.33, 0.34]
+        )
+        assert result == [1, 1, 1]
+        assert sum(result) == 3
+
+
+def test_get_shard_assignments_insufficient_memory_raises(
+    topology: Topology,
+    create_node: Callable[[int, NodeId | None], NodeInfo],
+    create_connection: Callable[[NodeId, NodeId], Connection],
+):
+    """Test that ValueError is raised when a node has insufficient memory for its layers."""
+    node_a_id = NodeId()
+    node_b_id = NodeId()
+    node_c_id = NodeId()
+
+    # Node C has only 10 KB but would need 50 KB for 1 layer (1000 KB / 20 layers)
+    node_a = create_node(900 * 1024, node_a_id)
+    node_b = create_node(50 * 1024, node_b_id)
+    node_c = create_node(10 * 1024, node_c_id)  # Insufficient memory
+
+    topology.add_node(node_a)
+    topology.add_node(node_b)
+    topology.add_node(node_c)
+
+    topology.add_connection(create_connection(node_a_id, node_b_id))
+    topology.add_connection(create_connection(node_b_id, node_c_id))
+    topology.add_connection(create_connection(node_c_id, node_a_id))
+    topology.add_connection(create_connection(node_b_id, node_a_id))
+
+    model_meta = ModelMetadata(
+        model_id=ModelId("test-model"),
+        pretty_name="Test Model",
+        n_layers=20,
+        storage_size=Memory.from_kb(1000),
+        hidden_size=1000,
+        supports_tensor=True,
+    )
+    cycles = topology.get_cycles()
+    selected_cycle = cycles[0]
+
+    with pytest.raises(ValueError, match="insufficient memory"):
+        get_shard_assignments(model_meta, selected_cycle, Sharding.Pipeline)
--- a/src/exo/shared/apply.py
+++ b/src/exo/shared/apply.py
@@ -16,6 +16,7 @@ from exo.shared.types.events import (
    NodeMemoryMeasured,
    NodePerformanceMeasured,
    NodeTimedOut,
+    PrefillProgress,
    RunnerDeleted,
    RunnerStatusUpdated,
    TaskAcknowledged,
@@ -40,7 +41,7 @@ def event_apply(event: Event, state: State) -> State:
    """Apply an event to state."""
    match event:
        case (
-            TestEvent() | ChunkGenerated() | TaskAcknowledged()
+            TestEvent() | ChunkGenerated() | TaskAcknowledged() | PrefillProgress()
        ):  # TaskAcknowledged should never be sent by a worker but i dont mind if it just gets ignored
            return state
        case InstanceCreated():
--- a/src/exo/shared/types/api.py
+++ b/src/exo/shared/types/api.py
@@ -11,10 +11,21 @@ from exo.shared.types.worker.instances import Instance, InstanceId, InstanceMeta
 from exo.shared.types.worker.shards import Sharding

 FinishReason = Literal[
-    "stop", "length", "tool_calls", "content_filter", "function_call"
+    "stop", "length", "tool_calls", "content_filter", "function_call", "error"
 ]


+class ErrorInfo(BaseModel):
+    message: str
+    type: str
+    param: str | None = None
+    code: int
+
+
+class ErrorResponse(BaseModel):
+    error: ErrorInfo
+
+
 class ModelListModel(BaseModel):
    id: str
    object: str = "model"
@@ -146,10 +157,13 @@ class ChatCompletionTaskParams(BaseModel):
    stream: bool = False
    temperature: float | None = None
    top_p: float | None = None
+    top_k: int | None = None
    tools: list[dict[str, Any]] | None = None
    tool_choice: str | dict[str, Any] | None = None
    parallel_tool_calls: bool | None = None
    user: str | None = None
+    # When True, continue the last assistant message without EOS tokens
+    continue_from_prefix: bool = False


 class BenchChatCompletionTaskParams(ChatCompletionTaskParams):
--- a/src/exo/shared/types/chunks.py
+++ b/src/exo/shared/types/chunks.py
@@ -1,6 +1,6 @@
 from enum import Enum

-from exo.shared.types.api import GenerationStats
+from exo.shared.types.api import GenerationStats, TopLogprobItem
 from exo.utils.pydantic_ext import TaggedModel

 from .api import FinishReason
@@ -20,8 +20,11 @@ class BaseChunk(TaggedModel):
 class TokenChunk(BaseChunk):
    text: str
    token_id: int
+    logprob: float | None = None  # Log probability of the selected token
+    top_logprobs: list[TopLogprobItem] | None = None  # Top-k alternative tokens
    finish_reason: FinishReason | None = None
    stats: GenerationStats | None = None
+    error_message: str | None = None


 class ImageChunk(BaseChunk):
@@ -29,3 +32,14 @@ class ImageChunk(BaseChunk):


 GenerationChunk = TokenChunk | ImageChunk
+
+
+class PrefillProgressData(TaggedModel):
+    """Data class for prefill progress events during streaming."""
+
+    processed_tokens: int
+    total_tokens: int
+
+
+# Stream events can be either token chunks or prefill progress
+StreamEvent = TokenChunk | PrefillProgressData
--- a/src/exo/shared/types/claude_api.py
+++ b/src/exo/shared/types/claude_api.py
@@ -0,0 +1,168 @@
+"""Claude Messages API types for request/response conversion."""
+
+from typing import Literal
+
+from pydantic import BaseModel, Field
+
+# Type aliases
+ClaudeRole = Literal["user", "assistant"]
+ClaudeStopReason = Literal["end_turn", "max_tokens", "stop_sequence", "tool_use"]
+
+
+# Content block types
+class ClaudeTextBlock(BaseModel, frozen=True):
+    """Text content block in Claude Messages API."""
+
+    type: Literal["text"] = "text"
+    text: str
+
+
+class ClaudeImageSource(BaseModel, frozen=True):
+    """Image source for Claude image blocks."""
+
+    type: Literal["base64", "url"]
+    media_type: str | None = None
+    data: str | None = None
+    url: str | None = None
+
+
+class ClaudeImageBlock(BaseModel, frozen=True):
+    """Image content block in Claude Messages API."""
+
+    type: Literal["image"] = "image"
+    source: ClaudeImageSource
+
+
+ClaudeContentBlock = ClaudeTextBlock | ClaudeImageBlock
+
+
+# Request types
+class ClaudeMessage(BaseModel, frozen=True):
+    """Message in Claude Messages API request."""
+
+    role: ClaudeRole
+    content: str | list[ClaudeContentBlock]
+
+
+class ClaudeMessagesRequest(BaseModel):
+    """Request body for Claude Messages API."""
+
+    model: str
+    max_tokens: int
+    messages: list[ClaudeMessage]
+    system: str | list[ClaudeTextBlock] | None = None
+    stop_sequences: list[str] | None = None
+    stream: bool = False
+    temperature: float | None = None
+    top_p: float | None = None
+    top_k: int | None = None
+    metadata: dict[str, str] | None = None
+
+
+# Response types
+class ClaudeUsage(BaseModel, frozen=True):
+    """Token usage in Claude Messages API response."""
+
+    input_tokens: int
+    output_tokens: int
+
+
+class ClaudeMessagesResponse(BaseModel, frozen=True):
+    """Response body for Claude Messages API."""
+
+    id: str
+    type: Literal["message"] = "message"
+    role: Literal["assistant"] = "assistant"
+    content: list[ClaudeTextBlock]
+    model: str
+    stop_reason: ClaudeStopReason | None = None
+    stop_sequence: str | None = None
+    usage: ClaudeUsage
+
+
+# Streaming event types
+class ClaudeMessageStart(BaseModel, frozen=True):
+    """Partial message in message_start event."""
+
+    id: str
+    type: Literal["message"] = "message"
+    role: Literal["assistant"] = "assistant"
+    content: list[ClaudeTextBlock] = Field(default_factory=list)
+    model: str
+    stop_reason: ClaudeStopReason | None = None
+    stop_sequence: str | None = None
+    usage: ClaudeUsage
+
+
+class ClaudeMessageStartEvent(BaseModel, frozen=True):
+    """Event sent at start of message stream."""
+
+    type: Literal["message_start"] = "message_start"
+    message: ClaudeMessageStart
+
+
+class ClaudeContentBlockStartEvent(BaseModel, frozen=True):
+    """Event sent at start of a content block."""
+
+    type: Literal["content_block_start"] = "content_block_start"
+    index: int
+    content_block: ClaudeTextBlock
+
+
+class ClaudeTextDelta(BaseModel, frozen=True):
+    """Delta for text content block."""
+
+    type: Literal["text_delta"] = "text_delta"
+    text: str
+
+
+class ClaudeContentBlockDeltaEvent(BaseModel, frozen=True):
+    """Event sent for content block delta."""
+
+    type: Literal["content_block_delta"] = "content_block_delta"
+    index: int
+    delta: ClaudeTextDelta
+
+
+class ClaudeContentBlockStopEvent(BaseModel, frozen=True):
+    """Event sent at end of a content block."""
+
+    type: Literal["content_block_stop"] = "content_block_stop"
+    index: int
+
+
+class ClaudeMessageDeltaUsage(BaseModel, frozen=True):
+    """Usage in message_delta event."""
+
+    output_tokens: int
+
+
+class ClaudeMessageDelta(BaseModel, frozen=True):
+    """Delta in message_delta event."""
+
+    stop_reason: ClaudeStopReason | None = None
+    stop_sequence: str | None = None
+
+
+class ClaudeMessageDeltaEvent(BaseModel, frozen=True):
+    """Event sent with final message delta."""
+
+    type: Literal["message_delta"] = "message_delta"
+    delta: ClaudeMessageDelta
+    usage: ClaudeMessageDeltaUsage
+
+
+class ClaudeMessageStopEvent(BaseModel, frozen=True):
+    """Event sent at end of message stream."""
+
+    type: Literal["message_stop"] = "message_stop"
+
+
+ClaudeStreamEvent = (
+    ClaudeMessageStartEvent
+    | ClaudeContentBlockStartEvent
+    | ClaudeContentBlockDeltaEvent
+    | ClaudeContentBlockStopEvent
+    | ClaudeMessageDeltaEvent
+    | ClaudeMessageStopEvent
+)
--- a/src/exo/shared/types/commands.py
+++ b/src/exo/shared/types/commands.py
@@ -1,8 +1,8 @@
 from pydantic import Field

-from exo.shared.types.api import ChatCompletionTaskParams
 from exo.shared.types.common import CommandId, NodeId
 from exo.shared.types.models import ModelMetadata
+from exo.shared.types.openai_responses import ResponsesRequest
 from exo.shared.types.worker.instances import Instance, InstanceId, InstanceMeta
 from exo.shared.types.worker.shards import Sharding
 from exo.utils.pydantic_ext import CamelCaseModel, TaggedModel
@@ -17,7 +17,7 @@ class TestCommand(BaseCommand):


 class ChatCompletion(BaseCommand):
-    request_params: ChatCompletionTaskParams
+    request_params: ResponsesRequest


 class PlaceInstance(BaseCommand):
--- a/src/exo/shared/types/events.py
+++ b/src/exo/shared/types/events.py
@@ -106,6 +106,12 @@ class ChunkGenerated(BaseEvent):
    chunk: GenerationChunk


+class PrefillProgress(BaseEvent):
+    command_id: CommandId
+    processed_tokens: int
+    total_tokens: int
+
+
 class TopologyEdgeCreated(BaseEvent):
    edge: Connection

@@ -131,6 +137,7 @@ Event = (
    | NodeMemoryMeasured
    | NodeDownloadProgress
    | ChunkGenerated
+    | PrefillProgress
    | TopologyEdgeCreated
    | TopologyEdgeDeleted
 )
--- a/src/exo/shared/types/openai_responses.py
+++ b/src/exo/shared/types/openai_responses.py
@@ -0,0 +1,190 @@
+"""OpenAI Responses API types for request/response conversion.
+
+ResponsesRequest serves as both:
+1. The external API request type for /v1/responses
+2. The canonical internal type used throughout the inference pipeline
+
+All external API formats (Chat Completions, Claude) are converted to
+ResponsesRequest at the API boundary.
+"""
+
+import time
+from typing import Any, Literal
+
+from pydantic import BaseModel, Field
+
+# Type aliases
+ResponseStatus = Literal["completed", "failed", "in_progress", "incomplete"]
+ResponseRole = Literal["user", "assistant", "system", "developer"]
+
+
+# Request types
+class ResponseInputMessage(BaseModel, frozen=True):
+    """Input message for Responses API.
+
+    This is also used as the internal message format throughout the pipeline.
+    """
+
+    role: ResponseRole
+    content: str
+
+
+class ResponsesRequest(BaseModel):
+    """Request body for OpenAI Responses API.
+
+    This is also the canonical internal task params format used throughout
+    the inference pipeline. All external API formats are converted to this
+    format at the API boundary.
+
+    Field mapping from other APIs:
+    - input: Replaces 'messages' from Chat Completions
+    - instructions: System message, extracted from messages or Claude's 'system'
+    - max_output_tokens: Replaces 'max_tokens' from Chat Completions
+    """
+
+    model: str
+    input: str | list[ResponseInputMessage]
+    instructions: str | None = None
+    max_output_tokens: int | None = None
+    temperature: float | None = None
+    top_p: float | None = None
+    top_k: int | None = None
+    stop: str | list[str] | None = None
+    seed: int | None = None
+    stream: bool = False
+    # Tools support
+    tools: list[dict[str, Any]] | None = None
+    # previous_response_id not supported in MVP
+    metadata: dict[str, str] | None = None
+    # When True, continue the last assistant message without EOS tokens
+    continue_from_prefix: bool = False
+
+
+# Response types
+class ResponseOutputText(BaseModel, frozen=True):
+    """Text content in response output."""
+
+    type: Literal["output_text"] = "output_text"
+    text: str
+    annotations: list[dict[str, str]] = Field(default_factory=list)
+
+
+class ResponseMessageItem(BaseModel, frozen=True):
+    """Message item in response output array."""
+
+    type: Literal["message"] = "message"
+    id: str
+    role: Literal["assistant"] = "assistant"
+    content: list[ResponseOutputText]
+    status: ResponseStatus = "completed"
+
+
+ResponseItem = ResponseMessageItem  # Can expand for function_call, reasoning, etc.
+
+
+class ResponseUsage(BaseModel, frozen=True):
+    """Token usage in Responses API response."""
+
+    input_tokens: int
+    output_tokens: int
+    total_tokens: int
+
+
+class ResponsesResponse(BaseModel, frozen=True):
+    """Response body for OpenAI Responses API."""
+
+    id: str
+    object: Literal["response"] = "response"
+    created_at: int = Field(default_factory=lambda: int(time.time()))
+    status: ResponseStatus = "completed"
+    model: str
+    output: list[ResponseItem]
+    output_text: str
+    usage: ResponseUsage | None = None
+
+
+# Streaming event types
+class ResponseCreatedEvent(BaseModel, frozen=True):
+    """Event sent when response is created."""
+
+    type: Literal["response.created"] = "response.created"
+    response: ResponsesResponse
+
+
+class ResponseInProgressEvent(BaseModel, frozen=True):
+    """Event sent when response starts processing."""
+
+    type: Literal["response.in_progress"] = "response.in_progress"
+    response: ResponsesResponse
+
+
+class ResponseOutputItemAddedEvent(BaseModel, frozen=True):
+    """Event sent when an output item is added."""
+
+    type: Literal["response.output_item.added"] = "response.output_item.added"
+    output_index: int
+    item: ResponseItem
+
+
+class ResponseContentPartAddedEvent(BaseModel, frozen=True):
+    """Event sent when a content part is added."""
+
+    type: Literal["response.content_part.added"] = "response.content_part.added"
+    output_index: int
+    content_index: int
+    part: ResponseOutputText
+
+
+class ResponseTextDeltaEvent(BaseModel, frozen=True):
+    """Event sent for text delta during streaming."""
+
+    type: Literal["response.output_text.delta"] = "response.output_text.delta"
+    output_index: int
+    content_index: int
+    delta: str
+
+
+class ResponseTextDoneEvent(BaseModel, frozen=True):
+    """Event sent when text content is done."""
+
+    type: Literal["response.output_text.done"] = "response.output_text.done"
+    output_index: int
+    content_index: int
+    text: str
+
+
+class ResponseContentPartDoneEvent(BaseModel, frozen=True):
+    """Event sent when a content part is done."""
+
+    type: Literal["response.content_part.done"] = "response.content_part.done"
+    output_index: int
+    content_index: int
+    part: ResponseOutputText
+
+
+class ResponseOutputItemDoneEvent(BaseModel, frozen=True):
+    """Event sent when an output item is done."""
+
+    type: Literal["response.output_item.done"] = "response.output_item.done"
+    output_index: int
+    item: ResponseItem
+
+
+class ResponseCompletedEvent(BaseModel, frozen=True):
+    """Event sent when response is completed."""
+
+    type: Literal["response.completed"] = "response.completed"
+    response: ResponsesResponse
+
+
+ResponsesStreamEvent = (
+    ResponseCreatedEvent
+    | ResponseInProgressEvent
+    | ResponseOutputItemAddedEvent
+    | ResponseContentPartAddedEvent
+    | ResponseTextDeltaEvent
+    | ResponseTextDoneEvent
+    | ResponseContentPartDoneEvent
+    | ResponseOutputItemDoneEvent
+    | ResponseCompletedEvent
+)
--- a/src/exo/shared/types/tasks.py
+++ b/src/exo/shared/types/tasks.py
@@ -2,8 +2,8 @@ from enum import Enum

 from pydantic import Field

-from exo.shared.types.api import ChatCompletionTaskParams
 from exo.shared.types.common import CommandId, Id
+from exo.shared.types.openai_responses import ResponsesRequest
 from exo.shared.types.worker.instances import BoundInstance, InstanceId
 from exo.shared.types.worker.runners import RunnerId
 from exo.shared.types.worker.shards import ShardMetadata
@@ -50,7 +50,7 @@ class StartWarmup(BaseTask):  # emitted by Worker

 class ChatCompletion(BaseTask):  # emitted by Master
    command_id: CommandId
-    task_params: ChatCompletionTaskParams
+    task_params: ResponsesRequest

    error_type: str | None = Field(default=None)
    error_message: str | None = Field(default=None)
--- a/src/exo/shared/types/worker/runner_response.py
+++ b/src/exo/shared/types/worker/runner_response.py
@@ -1,4 +1,4 @@
-from exo.shared.types.api import FinishReason, GenerationStats
+from exo.shared.types.api import FinishReason, GenerationStats, TopLogprobItem
 from exo.utils.pydantic_ext import TaggedModel


@@ -13,10 +13,16 @@ class TokenizedResponse(BaseRunnerResponse):
 class GenerationResponse(BaseRunnerResponse):
    text: str
    token: int
-    # logprobs: list[float] | None = None # too big. we can change to be top-k
+    logprob: float | None = None  # Log probability of the selected token
+    top_logprobs: list[TopLogprobItem] | None = None  # Top-k alternative tokens
    finish_reason: FinishReason | None = None
    stats: GenerationStats | None = None


 class FinishedResponse(BaseRunnerResponse):
    pass
+
+
+class PrefillProgressResponse(BaseRunnerResponse):
+    processed_tokens: int
+    total_tokens: int
--- a/src/exo/worker/download/download_utils.py
+++ b/src/exo/worker/download/download_utils.py
@@ -245,12 +245,15 @@ def create_http_session(
        sock_read_timeout = 1800
        sock_connect_timeout = 60

-    ssl_context = ssl.create_default_context(cafile=certifi.where())
+    ssl_context = ssl.create_default_context(
+        cafile=os.getenv("SSL_CERT_FILE") or certifi.where()
+    )
    connector = aiohttp.TCPConnector(ssl=ssl_context)

    return aiohttp.ClientSession(
        auto_decompress=auto_decompress,
        connector=connector,
+        proxy=os.getenv("HTTPS_PROXY") or os.getenv("HTTP_PROXY") or None,
        timeout=aiohttp.ClientTimeout(
            total=total_timeout,
            connect=connect_timeout,
--- a/src/exo/worker/engines/mlx/init.py
+++ b/src/exo/worker/engines/mlx/init.py
@@ -40,4 +40,6 @@ class TokenizerWrapper:
        messages_dicts: list[dict[str, Any]],
        tokenize: bool = False,
        add_generation_prompt: bool = True,
+        continue_final_message: bool = False,
+        tools: list[dict[str, Any]] | None = None,
    ) -> str: ...
--- a/src/exo/worker/engines/mlx/generator/generate.py
+++ b/src/exo/worker/engines/mlx/generator/generate.py
@@ -8,13 +8,12 @@ from mlx_lm.tokenizer_utils import TokenizerWrapper

 # from exo.engines.mlx.cache import KVPrefixCache
 from exo.shared.types.api import (
-    BenchChatCompletionTaskParams,
-    ChatCompletionMessage,
    FinishReason,
    GenerationStats,
+    TopLogprobItem,
 )
 from exo.shared.types.memory import Memory
-from exo.shared.types.tasks import ChatCompletionTaskParams
+from exo.shared.types.openai_responses import ResponsesRequest
 from exo.shared.types.worker.runner_response import (
    GenerationResponse,
 )
@@ -53,14 +52,9 @@ def warmup_inference(

    warmup_prompt = apply_chat_template(
        tokenizer=tokenizer,
-        chat_task_data=ChatCompletionTaskParams(
+        task_params=ResponsesRequest(
            model="",
-            messages=[
-                ChatCompletionMessage(
-                    role="user",
-                    content=content,
-                )
-            ],
+            input=content,
        ),
    )

@@ -81,7 +75,7 @@ def warmup_inference(
        max_tokens=50,
        sampler=sampler,
        prompt_cache=cache,
-        prefill_step_size=2048,
+        prefill_step_size=256,  # Temporarily reduced from 2048 for testing progress bar
        kv_group_size=KV_GROUP_SIZE,
        kv_bits=KV_BITS,
    ):
@@ -115,14 +109,69 @@ def eos_ids_from_tokenizer(tokenizer: TokenizerWrapper) -> list[int]:
    return eos


+def extract_top_logprobs(
+    logprobs: mx.array,
+    tokenizer: TokenizerWrapper,
+    top_k: int,
+    selected_token: int,
+) -> tuple[float, list[TopLogprobItem]]:
+    """Extract the selected token's logprob and top-k alternative tokens.
+
+    Args:
+        logprobs: Full vocabulary logprobs array from MLX
+        tokenizer: Tokenizer for decoding token IDs to strings
+        top_k: Number of top alternatives to return
+        selected_token: The token ID that was actually sampled
+
+    Returns:
+        Tuple of (selected_token_logprob, list of TopLogprobItem for top-k tokens)
+    """
+    # Get the logprob of the selected token
+    selected_logprob = float(logprobs[selected_token].item())
+
+    # Get top-k indices (most probable tokens)
+    # mx.argpartition gives indices that would partition the array
+    # We negate logprobs since argpartition finds smallest, and we want largest
+    top_k = min(top_k, logprobs.shape[0])  # Don't exceed vocab size
+    top_indices = mx.argpartition(-logprobs, top_k)[:top_k]
+
+    # Get the actual logprob values for these indices
+    top_values = logprobs[top_indices]
+
+    # Sort by logprob (descending) for consistent ordering
+    sort_order = mx.argsort(-top_values)
+    top_indices = top_indices[sort_order]
+    top_values = top_values[sort_order]
+
+    # Convert to list of TopLogprobItem
+    top_logprob_items: list[TopLogprobItem] = []
+    for i in range(top_k):
+        token_id = int(top_indices[i].item())
+        token_logprob = float(top_values[i].item())
+        # Decode token ID to string
+        token_str = tokenizer.decode([token_id])
+        # Get byte representation
+        token_bytes = list(token_str.encode("utf-8"))
+        top_logprob_items.append(
+            TopLogprobItem(
+                token=token_str,
+                logprob=token_logprob,
+                bytes=token_bytes,
+            )
+        )
+
+    return selected_logprob, top_logprob_items
+
+
 def mlx_generate(
    model: Model,
    tokenizer: TokenizerWrapper,
-    task: ChatCompletionTaskParams,
+    task: ResponsesRequest,
+    is_bench: bool = False,
+    on_prefill_progress: Callable[[int, int], None] | None = None,
 ) -> Generator[GenerationResponse]:
    # Ensure that generation stats only contains peak memory for this generation
    mx.reset_peak_memory()
-    is_bench: bool = isinstance(task, BenchChatCompletionTaskParams)

    # Currently we support chat-completion tasks only.
    logger.info(f"task_params: {task}")
@@ -132,7 +181,7 @@ def mlx_generate(

    prompt = apply_chat_template(
        tokenizer=tokenizer,
-        chat_task_data=task,
+        task_params=task,
    )

    caches = make_kv_cache(model=model)
@@ -146,9 +195,20 @@ def mlx_generate(
    sampler = make_sampler(
        temp=task.temperature if task.temperature is not None else 0.7,
        top_p=task.top_p if task.top_p is not None else 1.0,
+        top_k=task.top_k if task.top_k is not None else 0,
    )

-    max_tokens = task.max_tokens or MAX_TOKENS
+    # Normalize stop sequences to a list
+    stop_sequences: list[str] = (
+        ([task.stop] if isinstance(task.stop, str) else task.stop)
+        if task.stop is not None
+        else []
+    )
+    max_stop_len = max((len(s) for s in stop_sequences), default=0)
+
+    max_tokens = task.max_output_tokens or MAX_TOKENS
+    accumulated_text = ""
+
    for out in stream_generate(
        model=model,
        tokenizer=tokenizer,
@@ -158,14 +218,36 @@ def mlx_generate(
        logits_processors=logits_processors,
        prompt_cache=caches,
        # TODO: Dynamically change prefill step size to be the maximum possible without timing out.
-        prefill_step_size=2048,
+        prefill_step_size=256,  # Temporarily reduced from 2048 for testing progress bar
        kv_group_size=KV_GROUP_SIZE,
        kv_bits=KV_BITS,
+        prompt_progress_callback=on_prefill_progress,
    ):
        logger.info(out.text)
+        accumulated_text += out.text

+        # Check for stop sequences
+        text = out.text
+        finish_reason: FinishReason | None = cast(
+            FinishReason | None, out.finish_reason
+        )
+        stop_matched = False
+
+        if stop_sequences:
+            for stop_seq in stop_sequences:
+                if stop_seq in accumulated_text:
+                    # Trim text to just before the stop sequence
+                    stop_index = accumulated_text.find(stop_seq)
+                    text_before_stop = accumulated_text[:stop_index]
+                    chunk_start = len(accumulated_text) - len(out.text)
+                    text = text_before_stop[chunk_start:]
+                    finish_reason = "stop"
+                    stop_matched = True
+                    break
+
+        is_done = finish_reason is not None
        stats: GenerationStats | None = None
-        if out.finish_reason is not None:
+        if is_done:
            stats = GenerationStats(
                prompt_tps=float(out.prompt_tps),
                generation_tps=float(out.generation_tps),
@@ -173,22 +255,33 @@ def mlx_generate(
                generation_tokens=int(out.generation_tokens),
                peak_memory_usage=Memory.from_gb(out.peak_memory),
            )
-
-            if out.finish_reason not in get_args(FinishReason):
-                # We don't throw here as this failure case is really not all that bad
-                # Just log the error and move on
+            if not stop_matched and out.finish_reason not in get_args(FinishReason):
                logger.warning(
                    f"Model generated unexpected finish_reason: {out.finish_reason}"
                )

+        # Extract logprobs from the full vocabulary logprobs array
+        logprob, top_logprobs = extract_top_logprobs(
+            logprobs=out.logprobs,
+            tokenizer=tokenizer,
+            top_k=5,
+            selected_token=out.token,
+        )
+
        yield GenerationResponse(
-            text=out.text,
+            text=text,
            token=out.token,
-            finish_reason=cast(FinishReason | None, out.finish_reason),
+            logprob=logprob,
+            top_logprobs=top_logprobs,
+            finish_reason=finish_reason,
            stats=stats,
        )

-        if out.finish_reason is not None:
+        if is_done:
            break

+        # Limit accumulated_text to what's needed for stop sequence detection
+        if max_stop_len > 0 and len(accumulated_text) > max_stop_len:
+            accumulated_text = accumulated_text[-max_stop_len:]
+
        # TODO: Do we want an mx_barrier?
--- a/src/exo/worker/engines/mlx/utils_mlx.py
+++ b/src/exo/worker/engines/mlx/utils_mlx.py
@@ -2,7 +2,9 @@ import json
 import os
 import resource
 import sys
+import threading
 import time
+from collections.abc import Callable
 from pathlib import Path
 from typing import Any, cast

@@ -40,10 +42,9 @@ import mlx.nn as nn
 from mlx_lm.utils import load_model
 from pydantic import RootModel

-from exo.shared.types.api import ChatCompletionMessageText
 from exo.shared.types.common import Host
 from exo.shared.types.memory import Memory
-from exo.shared.types.tasks import ChatCompletionTaskParams
+from exo.shared.types.openai_responses import ResponsesRequest
 from exo.shared.types.worker.instances import (
    BoundInstance,
    MlxJacclInstance,
@@ -82,6 +83,45 @@ def get_weights_size(model_shard_meta: ShardMetadata) -> Memory:
    )


+class ModelLoadingTimeoutError(Exception):
+    pass
+
+
+TimeoutCallback = Callable[[], None]
+
+
+def eval_with_timeout(
+    mlx_item: Any,  # pyright: ignore[reportAny]
+    timeout_seconds: float = 60.0,
+    on_timeout: TimeoutCallback | None = None,
+) -> None:
+    """Evaluate MLX item with a hard timeout.
+
+    If on_timeout callback is provided, it will be called before terminating
+    the process. This allows the runner to send a failure event before exit.
+    """
+    completed = threading.Event()
+
+    def watchdog() -> None:
+        if not completed.wait(timeout=timeout_seconds):
+            logger.error(
+                f"mlx_item evaluation timed out after {timeout_seconds:.0f}s. "
+                "This may indicate an issue with FAST_SYNCH and tensor parallel sharding. "
+                "Terminating process."
+            )
+            if on_timeout is not None:
+                on_timeout()
+            os._exit(1)
+
+    watchdog_thread = threading.Thread(target=watchdog, daemon=True)
+    watchdog_thread.start()
+
+    try:
+        mx.eval(mlx_item)  # pyright: ignore[reportAny]
+    finally:
+        completed.set()
+
+
 def mx_barrier(group: Group | None = None):
    mx.eval(
        mx.distributed.all_sum(
@@ -188,7 +228,9 @@ def initialize_mlx(


 def load_mlx_items(
-    bound_instance: BoundInstance, group: Group | None
+    bound_instance: BoundInstance,
+    group: Group | None,
+    on_timeout: TimeoutCallback | None = None,
 ) -> tuple[Model, TokenizerWrapper]:
    if group is None:
        logger.info(f"Single device used for {bound_instance.instance}")
@@ -202,7 +244,9 @@ def load_mlx_items(
    else:
        logger.info("Starting distributed init")
        start_time = time.perf_counter()
-        model, tokenizer = shard_and_load(bound_instance.bound_shard, group=group)
+        model, tokenizer = shard_and_load(
+            bound_instance.bound_shard, group=group, on_timeout=on_timeout
+        )
        end_time = time.perf_counter()
        logger.info(
            f"Time taken to shard and load model: {(end_time - start_time):.2f}s"
@@ -216,6 +260,7 @@ def load_mlx_items(
 def shard_and_load(
    shard_metadata: ShardMetadata,
    group: Group,
+    on_timeout: TimeoutCallback | None = None,
 ) -> tuple[nn.Module, TokenizerWrapper]:
    model_path = build_model_path(shard_metadata.model_meta.model_id)

@@ -252,7 +297,15 @@ def shard_and_load(
            logger.info(f"loading model from {model_path} with pipeline parallelism")
            model = pipeline_auto_parallel(model, group, shard_metadata)

-    mx.eval(model.parameters())
+    # Estimate timeout based on model size
+    base_timeout = float(os.environ.get("EXO_MODEL_LOAD_TIMEOUT", "60"))
+    model_size_gb = get_weights_size(shard_metadata).in_bytes / (1024**3)
+    timeout_seconds = base_timeout + model_size_gb / 5
+    logger.info(
+        f"Evaluating model parameters with timeout of {timeout_seconds:.0f}s "
+        f"(model size: {model_size_gb:.1f}GB)"
+    )
+    eval_with_timeout(model.parameters(), timeout_seconds, on_timeout)

    # TODO: Do we need this?
    mx.eval(model)
@@ -336,35 +389,53 @@ def load_tokenizer_for_model_id(model_id: str, model_path: Path) -> TokenizerWra

 def apply_chat_template(
    tokenizer: TokenizerWrapper,
-    chat_task_data: ChatCompletionTaskParams,
+    task_params: ResponsesRequest,
 ) -> str:
-    # Now we can properly access the messages
-    messages = chat_task_data.messages
+    """Convert ResponsesRequest to a chat template prompt.

+    Converts the internal format (input + instructions) to a messages list
+    that can be processed by the tokenizer's chat template.
+    """
    formatted_messages: list[dict[str, Any]] = []
-    for message in messages:
-        if isinstance(message.content, ChatCompletionMessageText):
-            message.content = message.content.text
-        if isinstance(message.content, list):
-            if len(message.content) == 0:
-                logger.warning("Received prompt with no content, skipping")
-                continue

-            message.content = "\n".join(c.text for c in message.content).strip()
-        if message.content is None and message.thinking is None:
-            continue
-
-        # Null values are not valid when applying templates in tokenizer
+    # Add system message (instructions) if present
+    if task_params.instructions:
        formatted_messages.append(
-            {k: v for k, v in message.model_dump().items() if v is not None}  # type: ignore
+            {"role": "system", "content": task_params.instructions}
        )

-    prompt: str = tokenizer.apply_chat_template(
-        formatted_messages,
-        tokenize=False,
-        add_generation_prompt=True,
-        tools=chat_task_data.tools,
-    )
+    # Convert input to messages
+    if isinstance(task_params.input, str):
+        # Simple string input becomes a single user message
+        formatted_messages.append({"role": "user", "content": task_params.input})
+    else:
+        # List of InputMessage
+        for msg in task_params.input:
+            if not msg.content:
+                logger.warning("Received message with empty content, skipping")
+                continue
+            formatted_messages.append({"role": msg.role, "content": msg.content})
+
+    # Use continue_final_message when continuing from prefix (e.g., regenerate from token)
+    # This keeps the final assistant message open without EOS tokens
+    # Note: explicitly set add_generation_prompt=False when using continue_final_message
+    # because some tokenizers (e.g., Kimi) default add_generation_prompt=True
+    prompt: str
+    if task_params.continue_from_prefix:
+        prompt = tokenizer.apply_chat_template(
+            formatted_messages,
+            tokenize=False,
+            continue_final_message=True,
+            add_generation_prompt=False,
+            tools=task_params.tools,
+        )
+    else:
+        prompt = tokenizer.apply_chat_template(
+            formatted_messages,
+            tokenize=False,
+            add_generation_prompt=True,
+            tools=task_params.tools,
+        )

    logger.info(prompt)

--- a/src/exo/worker/runner/bootstrap.py
+++ b/src/exo/worker/runner/bootstrap.py
@@ -17,15 +17,23 @@ def entrypoint(
    task_receiver: MpReceiver[Task],
    _logger: "loguru.Logger",
 ) -> None:
-    if (
-        isinstance(bound_instance.instance, MlxJacclInstance)
-        and len(bound_instance.instance.ibv_devices) >= 2
+    fast_synch_override = os.environ.get("EXO_FAST_SYNCH")
+    if fast_synch_override == "on" or (
+        fast_synch_override != "off"
+        and (
+            isinstance(bound_instance.instance, MlxJacclInstance)
+            and len(bound_instance.instance.ibv_devices) >= 2
+        )
    ):
        os.environ["MLX_METAL_FAST_SYNCH"] = "1"
+    else:
+        os.environ["MLX_METAL_FAST_SYNCH"] = "0"

    global logger
    logger = _logger

+    logger.info(f"Fast synch flag: {os.environ['MLX_METAL_FAST_SYNCH']}")
+
    # Import main after setting global logger - this lets us just import logger from this module
    try:
        from exo.worker.runner.runner import main
--- a/src/exo/worker/runner/runner.py
+++ b/src/exo/worker/runner/runner.py
@@ -11,15 +11,16 @@ from openai_harmony import (  # pyright: ignore[reportMissingTypeStubs]
    load_harmony_encoding,
 )

-from exo.shared.types.api import ChatCompletionMessageText
 from exo.shared.types.chunks import TokenChunk
 from exo.shared.types.events import (
    ChunkGenerated,
    Event,
+    PrefillProgress,
    RunnerStatusUpdated,
    TaskAcknowledged,
    TaskStatusUpdated,
 )
+from exo.shared.types.openai_responses import ResponsesRequest
 from exo.shared.types.tasks import (
    ChatCompletion,
    ConnectToGroup,
@@ -67,6 +68,7 @@ def main(
        bound_instance.bound_runner_id,
        bound_instance.bound_shard,
    )
+    device_rank = shard_metadata.device_rank
    logger.info("hello from the runner")
    if getattr(shard_metadata, "immediate_exception", False):
        raise Exception("Fake exception - runner failed to spin up.")
@@ -118,7 +120,20 @@ def main(
                        )
                    )

-                    model, tokenizer = load_mlx_items(bound_instance, group)
+                    def on_model_load_timeout() -> None:
+                        event_sender.send(
+                            RunnerStatusUpdated(
+                                runner_id=runner_id,
+                                runner_status=RunnerFailed(
+                                    error_message="Model loading timed out"
+                                ),
+                            )
+                        )
+                        time.sleep(0.5)
+
+                    model, tokenizer = load_mlx_items(
+                        bound_instance, group, on_timeout=on_model_load_timeout
+                    )

                    current_status = RunnerLoaded()
                    logger.info("runner loaded")
@@ -148,8 +163,6 @@ def main(
                case ChatCompletion(task_params=task_params, command_id=command_id) if (
                    isinstance(current_status, RunnerReady)
                ):
-                    assert model
-                    assert tokenizer
                    logger.info(f"received chat request: {str(task)[:500]}")
                    current_status = RunnerRunning()
                    logger.info("runner running")
@@ -158,41 +171,74 @@ def main(
                            runner_id=runner_id, runner_status=current_status
                        )
                    )
-                    assert task_params.messages[0].content is not None
-                    _check_for_debug_prompts(task_params.messages[0].content)
+                    assert model
+                    assert tokenizer

-                    # Generate responses using the actual MLX generation
-                    mlx_generator = mlx_generate(
-                        model=model,
-                        tokenizer=tokenizer,
-                        task=task_params,
-                    )
+                    # Define callback to send prefill progress events directly
+                    def on_prefill_progress(processed: int, total: int) -> None:
+                        if device_rank == 0:
+                            event_sender.send(
+                                PrefillProgress(
+                                    command_id=command_id,
+                                    processed_tokens=processed,
+                                    total_tokens=total,
+                                )
+                            )

-                    # GPT-OSS specific parsing to match other model formats.
-                    if isinstance(model, GptOssModel):
-                        mlx_generator = parse_gpt_oss(mlx_generator)
+                    try:
+                        _check_for_debug_prompts(task_params)

-                    # TODO: Add tool call parser here
+                        # Generate responses using the actual MLX generation
+                        mlx_generator = mlx_generate(
+                            model=model,
+                            tokenizer=tokenizer,
+                            task=task_params,
+                            on_prefill_progress=on_prefill_progress,
+                        )

-                    for response in mlx_generator:
-                        match response:
-                            case GenerationResponse():
-                                if shard_metadata.device_rank == 0:
-                                    event_sender.send(
-                                        ChunkGenerated(
-                                            command_id=command_id,
-                                            chunk=TokenChunk(
-                                                idx=response.token,
-                                                model=shard_metadata.model_meta.model_id,
-                                                text=response.text,
-                                                token_id=response.token,
-                                                finish_reason=response.finish_reason,
-                                                stats=response.stats,
-                                            ),
+                        # GPT-OSS specific parsing to match other model formats.
+                        if isinstance(model, GptOssModel):
+                            mlx_generator = parse_gpt_oss(mlx_generator)
+
+                        # TODO: Add tool call parser here
+
+                        for response in mlx_generator:
+                            match response:
+                                case GenerationResponse():
+                                    if device_rank == 0:
+                                        event_sender.send(
+                                            ChunkGenerated(
+                                                command_id=command_id,
+                                                chunk=TokenChunk(
+                                                    idx=response.token,
+                                                    model=shard_metadata.model_meta.model_id,
+                                                    text=response.text,
+                                                    token_id=response.token,
+                                                    logprob=response.logprob,
+                                                    top_logprobs=response.top_logprobs,
+                                                    finish_reason=response.finish_reason,
+                                                    stats=response.stats,
+                                                ),
+                                            )
                                        )
-                                    )
-                                # case TokenizedResponse():
-                                # TODO: something here ig
+
+                    # can we make this more explicit?
+                    except Exception as e:
+                        if device_rank == 0:
+                            event_sender.send(
+                                ChunkGenerated(
+                                    command_id=command_id,
+                                    chunk=TokenChunk(
+                                        idx=0,
+                                        model=shard_metadata.model_meta.model_id,
+                                        text="",
+                                        token_id=0,
+                                        finish_reason="error",
+                                        error_message=str(e),
+                                    ),
+                                )
+                            )
+                        raise

                    current_status = RunnerReady()
                    logger.info("runner ready")
@@ -266,17 +312,23 @@ EXO_RUNNER_MUST_OOM = "EXO RUNNER MUST OOM"
 EXO_RUNNER_MUST_TIMEOUT = "EXO RUNNER MUST TIMEOUT"


-def _check_for_debug_prompts(
-    prompt: str | ChatCompletionMessageText | list[ChatCompletionMessageText],
-):
-    if isinstance(prompt, list):
-        if len(prompt) == 0:
-            logger.debug("Empty message prompt received in debug prompt")
-            return
-        prompt = prompt[0]
+def _check_for_debug_prompts(task_params: ResponsesRequest) -> None:
+    """Check for debug prompt triggers in the input.

-    if isinstance(prompt, ChatCompletionMessageText):
-        prompt = prompt.text
+    Extracts the first user input text and checks for debug triggers.
+    """
+    prompt: str
+    if isinstance(task_params.input, str):
+        prompt = task_params.input
+    else:
+        # List of InputMessage - get first message content
+        if len(task_params.input) == 0:
+            logger.debug("Empty message list in debug prompt check")
+            return
+        prompt = task_params.input[0].content
+
+    if not prompt:
+        return

    if EXO_RUNNER_MUST_FAIL in prompt:
        logger.info("raising exception")
--- a/src/exo/worker/tests/unittests/test_plan/test_task_forwarding.py
+++ b/src/exo/worker/tests/unittests/test_plan/test_task_forwarding.py
@@ -1,7 +1,7 @@
 from typing import cast

 import exo.worker.plan as plan_mod
-from exo.shared.types.api import ChatCompletionTaskParams
+from exo.shared.types.openai_responses import ResponsesRequest
 from exo.shared.types.tasks import ChatCompletion, Task, TaskId, TaskStatus
 from exo.shared.types.worker.instances import BoundInstance, InstanceId
 from exo.shared.types.worker.runners import (
@@ -59,7 +59,7 @@ def test_plan_forwards_pending_chat_completion_when_runner_ready():
        instance_id=INSTANCE_1_ID,
        task_status=TaskStatus.Pending,
        command_id=COMMAND_1_ID,
-        task_params=ChatCompletionTaskParams(model=MODEL_A_ID, messages=[]),
+        task_params=ResponsesRequest(model=MODEL_A_ID, input=""),
    )

    result = plan_mod.plan(
@@ -107,7 +107,7 @@ def test_plan_does_not_forward_chat_completion_if_any_runner_not_ready():
        instance_id=INSTANCE_1_ID,
        task_status=TaskStatus.Pending,
        command_id=COMMAND_1_ID,
-        task_params=ChatCompletionTaskParams(model=MODEL_A_ID, messages=[]),
+        task_params=ResponsesRequest(model=MODEL_A_ID, input=""),
    )

    result = plan_mod.plan(
@@ -152,7 +152,7 @@ def test_plan_does_not_forward_tasks_for_other_instances():
        instance_id=other_instance_id,
        task_status=TaskStatus.Pending,
        command_id=COMMAND_1_ID,
-        task_params=ChatCompletionTaskParams(model=MODEL_A_ID, messages=[]),
+        task_params=ResponsesRequest(model=MODEL_A_ID, input=""),
    )

    result = plan_mod.plan(
@@ -201,7 +201,7 @@ def test_plan_ignores_non_pending_or_non_chat_tasks():
        instance_id=INSTANCE_1_ID,
        task_status=TaskStatus.Complete,
        command_id=COMMAND_1_ID,
-        task_params=ChatCompletionTaskParams(model=MODEL_A_ID, messages=[]),
+        task_params=ResponsesRequest(model=MODEL_A_ID, input=""),
    )

    other_task_id = TaskId("other-task")
--- a/src/exo/worker/tests/unittests/test_runner/test_event_ordering.py
+++ b/src/exo/worker/tests/unittests/test_runner/test_event_ordering.py
@@ -5,7 +5,6 @@ from typing import Callable
 import pytest

 import exo.worker.runner.runner as mlx_runner
-from exo.shared.types.api import ChatCompletionMessage
 from exo.shared.types.chunks import TokenChunk
 from exo.shared.types.events import (
    ChunkGenerated,
@@ -14,9 +13,9 @@ from exo.shared.types.events import (
    TaskAcknowledged,
    TaskStatusUpdated,
 )
+from exo.shared.types.openai_responses import ResponsesRequest
 from exo.shared.types.tasks import (
    ChatCompletion,
-    ChatCompletionTaskParams,
    ConnectToGroup,
    LoadModel,
    Shutdown,
@@ -85,11 +84,11 @@ SHUTDOWN_TASK = Shutdown(
    runner_id=RUNNER_1_ID,
 )

-CHAT_PARAMS = ChatCompletionTaskParams(
+CHAT_PARAMS = ResponsesRequest(
    model=str(MODEL_A_ID),
-    messages=[ChatCompletionMessage(role="user", content="hello")],
+    input="hello",
    stream=True,
-    max_tokens=4,
+    max_output_tokens=4,
    temperature=0.0,
 )

@@ -121,6 +120,21 @@ def patch_out_mlx(monkeypatch: pytest.MonkeyPatch):
    monkeypatch.setattr(mlx_runner, "mlx_generate", fake_generate)


+# Use a fake event_sender to remove test flakiness.
+class EventCollector:
+    def __init__(self) -> None:
+        self.events: list[Event] = []
+
+    def send(self, event: Event) -> None:
+        self.events.append(event)
+
+    def close(self) -> None:
+        pass
+
+    def join(self) -> None:
+        pass
+
+
 def _run(tasks: Iterable[Task]):
    bound_instance = get_bound_mlx_ring_instance(
        instance_id=INSTANCE_1_ID,
@@ -130,22 +144,20 @@ def _run(tasks: Iterable[Task]):
    )

    task_sender, task_receiver = mp_channel[Task]()
-    event_sender, event_receiver = mp_channel[Event]()
+    event_sender = EventCollector()

-    with task_sender, event_receiver:
+    with task_sender:
        for t in tasks:
            task_sender.send(t)

        # worst monkeypatch known to man
        # this is some c++ nonsense
-        event_sender.close = nothin
-        event_sender.join = nothin
        task_receiver.close = nothin
        task_receiver.join = nothin

-        mlx_runner.main(bound_instance, event_sender, task_receiver)
+        mlx_runner.main(bound_instance, event_sender, task_receiver)  # type: ignore[arg-type]

-        return event_receiver.collect()
+        return event_sender.events


 def test_events_processed_in_correct_order(patch_out_mlx: pytest.MonkeyPatch):
--- a/tests/headless_runner.py
+++ b/tests/headless_runner.py
@@ -13,10 +13,10 @@ from pydantic import BaseModel

 from exo.shared.logging import InterceptLogger, logger_setup
 from exo.shared.models.model_cards import MODEL_CARDS, ModelId
-from exo.shared.types.api import ChatCompletionMessage, ChatCompletionTaskParams
 from exo.shared.types.commands import CommandId
 from exo.shared.types.common import Host, NodeId
 from exo.shared.types.events import Event
+from exo.shared.types.openai_responses import ResponsesRequest
 from exo.shared.types.tasks import (
    ChatCompletion,
    ConnectToGroup,
@@ -169,16 +169,10 @@ async def execute_test(test: Tests, instance: Instance, hn: str):
            send.send(StartWarmup(instance_id=iid))
            send.send(
                ChatCompletion(
-                    task_params=ChatCompletionTaskParams(
+                    task_params=ResponsesRequest(
                        model=test.model_id,
-                        messages=[
-                            ChatCompletionMessage(
-                                role="system", content="You are a helpful assistant"
-                            ),
-                            ChatCompletionMessage(
-                                role="user", content="What is the capital of France?"
-                            ),
-                        ],
+                        instructions="You are a helpful assistant",
+                        input="What is the capital of France?",
                    ),
                    command_id=CommandId("yo"),
                    instance_id=iid,
Author	SHA1	Message	Date
Alex Cheema	efc5baa6e5	style: simplify prefill progress bar and use exo color palette - Remove spinner (progress bar is dynamic enough) - Use exo-yellow for progress bar fill - Use exo-black/60 for progress bar background - Use exo-light-gray for text Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 16:57:58 +00:00
Alex Cheema	75e3634880	fix: wire prefill progress events to chat completions stream - Move PrefillProgressData to shared types (chunks.py) to avoid circular imports - Update generate_chat_stream adapter to handle both TokenChunk and PrefillProgressData - Use _stream_events instead of _chat_chunk_stream for streaming endpoint - Prefill progress now properly sent as SSE 'event: prefill_progress' to frontend Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 16:41:16 +00:00
Alex Cheema	43a73e3c1f	feat: add prefill progress bar for long prompts Shows real-time progress during prompt processing (prefill phase). Progress is sent via SSE named events that maintain OpenAI API compatibility. - Add PrefillProgress event type and PrefillProgressData dataclass - Wire prompt_progress_callback through MLX stream_generate - Send progress events directly from callback for real-time updates - Add PrefillProgressBar.svelte component - Parse event: prefill_progress SSE events in dashboard Note: prefill_step_size temporarily set to 256 for testing (normally 2048) Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 14:56:53 +00:00
Alex Cheema	3a181fbf33	style: format app.svelte.ts with nix fmt	2026-01-19 14:55:42 +00:00
Alex Cheema	67d4f23c61	Fix localStorage quota issues by stripping tokens and auto-pruning - Strip tokens (logprobs data) from messages before saving to localStorage since they're large and not essential for persistence - Add pruneOldConversations() to automatically remove oldest conversations when quota is exceeded - This prevents QuotaExceededError from crashing the app Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 14:55:42 +00:00
Alex Cheema	0b0d0f7faf	Fix ReferenceError: controller undefined in sendMessage finally block Move AbortController creation before the try block in both sendMessageWithLogprobs and regenerateFromToken functions. Previously, controller was defined inside the try block but referenced in the finally block, causing a ReferenceError if an exception was thrown before the controller was created. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 14:55:42 +00:00
Alex Cheema	28f7521540	Add SSE headers to properly close streaming connections Add Cache-Control, Connection: close, and X-Accel-Buffering headers to all SSE streaming responses. These headers help ensure: - No caching of streaming responses - Connection closes when stream ends (instead of keep-alive) - No proxy buffering that could delay stream closure This should fix the issue where the frontend stays on "PROCESSING" even after receiving the complete response. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 14:55:42 +00:00
Alex Cheema	94f9a09f24	Add debug logging to generate_chat_stream Add logging to help diagnose why streaming might not be ending properly. This will show when [DONE] is yielded, when return is called, and when the finally block runs. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 14:55:42 +00:00
Alex Cheema	2bf64ffd47	Fix streaming not ending after [DONE] is yielded Add missing return statement after yielding [DONE] in generate_chat_stream. Without this, the async generator continues waiting for more chunks from chunk_stream even though generation is complete, causing the stream to hang indefinitely. The frontend waits for the stream to close (reader.done) which never happens, resulting in the chat button staying on "PROCESSING" forever. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 14:55:42 +00:00
Alex Cheema	d091c84dc5	fix: restore extract_top_logprobs function for uncertainty visualization The extract_top_logprobs function was lost during rebases. This function processes the out.logprobs array (full vocabulary logprobs from MLX) to extract the selected token's logprob and top-k alternatives. The previous code tried to use getattr(out, "logprob", None) which doesn't exist - mlx_lm returns logprobs as an mx.array, not individual values. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 14:55:42 +00:00
Alex Cheema	e2c15f76b0	fix: remove unsupported logprob params from stream_generate The mlx_lm.stream_generate already returns logprobs in its output - we don't need to pass return_logprob or return_top_logprobs kwargs. The uncertainty visualization feature extracts logprobs from the existing out.logprobs field. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 14:55:42 +00:00
Alex Cheema	7d77043217	feat: add uncertainty visualization with token-level logprobs - Add TokenHeatmap component for visualizing token confidence - Collect and stream logprobs in generation pipeline - Add regenerate-from-token feature with continue_from_prefix - Add AbortController for request cancellation - Support continue_final_message for seamless prefix continuation Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 14:55:42 +00:00
Alex Cheema	e1e4516a8f	style: move inline imports to top of file in api.py Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 14:55:21 +00:00
Alex Cheema	8a67e949d1	fix: restore try/except structure in runner.py Replace non-existent context manager with proper try/except block and remove unused ModelId import. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 12:08:04 +00:00
Alex Cheema	fbf58bebd2	style: fix formatting issues caught by treefmt Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 12:06:15 +00:00
Alex Cheema	71b8e88d4b	refactor: use ResponsesRequest as canonical internal type - Extend ResponsesRequest with fields: top_k, seed, stop, tools - Remove redundant InternalTaskParams and InputMessage types - Update all adapters to convert to ResponsesRequest - Simplify Responses API (no conversion needed - native passthrough) - Update all imports across codebase and tests This eliminates type duplication and makes the Responses API relationship explicit throughout the codebase. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 12:06:14 +00:00
Alex Cheema	4b0ebb8ae4	refactor: make Responses API the canonical internal format Restructure the API layer so that OpenAI Responses API is the native format, with Chat Completions and Claude Messages as adapters on top. Changes: - Add new chat_completions.py adapter with streaming/non-streaming support - Update responses.py with collect_responses_response() for non-streaming - Update claude.py with collect_claude_response() for non-streaming - Refactor api.py so all endpoints use adapters uniformly - Rename _chat_chunk_stream to _token_chunk_stream (generic internal format) - Remove unused chat_response_to_* converter functions - Update tests to remove tests for deleted functions Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 12:06:14 +00:00
Alex Cheema	4df036d796	feat: add Claude Messages API and OpenAI Responses API support Adds two new API endpoints that wrap the existing chat completions: - /v1/messages - Claude Messages API compatible endpoint - /v1/responses - OpenAI Responses API compatible endpoint Both support streaming (SSE) and non-streaming modes with proper token usage reporting from actual inference stats. Also adds top_k sampling parameter and stop sequence support to the MLX inference engine. Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 12:06:14 +00:00
Evan Quiney	746589ba6b	tidy: remove context manager from api (#1199 )	2026-01-19 11:58:13 +00:00
rltakashige	f82f862fd7	Fix several issues with placement (#1200 ) ## Motivation Uneven placements were causing issues for some users with lopsided setups. While fixing, I ran into another issue with impossible allocation of memory. ## Changes - Allocate at least 1 layer per device. - Catch overallocation of memory with an error. ## Why It Works <!-- Explain why your approach solves the problem --> ## Test Plan ### Manual Testing Tested that GPT OSS is placed correctly. ### Automated Testing Added breaking tests in the first commit. Resolved with new placement algorithm in the second one.	2026-01-19 11:52:35 +00:00
Alex Cheema	7ff937d8a1	Add dashboard screenshots to README (#1185 ) ## Motivation The README showcases exo's features and benchmarks but doesn't show what the dashboard actually looks like. Adding a screenshot helps users understand what they'll get when they run exo. ## Changes - Added dashboard screenshot to `docs/imgs/dashboard-cluster-view.png`: Shows the cluster topology view with 4 × 512GB M3 Ultra Mac Studio running DeepSeek v3.1 (8-bit) and Kimi-K2-Thinking (4-bit) - Added a new "Dashboard" section to README.md below Features, displaying the screenshot with caption ## Why It Works Visual documentation helps users understand what exo offers before they install it. The screenshot demonstrates the cluster management capabilities. ## Test Plan ### Manual Testing - Verified image renders correctly in GitHub markdown preview ### Automated Testing - N/A - documentation only change Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-19 10:43:27 +00:00
Evan Quiney	d19bf02404	re-raise exceptions in the runner (#1198 ) ## Motivation Runners that crash can swallow errors - we should re-raise. Also the exception handler annoyed me. ## Changes The try: except in the runner's chat now re-raises.	2026-01-19 10:35:23 +00:00
rltakashige	618cee5223	Resolve test event ordering flakiness (#1194 ) ## Motivation mp sender occasionally does not have time to flush its events before collect() is called, making the event ordering test fail. ## Changes - Replace mp_channel with simple collector for event ordering test - Also suppress warning for <frozen importlib._bootstrap>:488 <frozen importlib._bootstrap>:488: DeprecationWarning: builtin type SwigPyObject has no __module__ attribute ## Why It Works <!-- Explain why your approach solves the problem --> ## Test Plan ### Manual Testing <!-- Hardware: (e.g., MacBook Pro M1 Max 32GB, Mac Mini M2 16GB, connected via Thunderbolt 4) --> <!-- What you did: --> <!-- - --> ### Automated Testing Ran the test 100 times without it failing.	2026-01-18 20:33:20 +00:00
Antonio Lujano Luna	9c29eb7d48	Add proxy and custom SSL certificate support for corporate networks (#1189 ) Support HTTPS_PROXY/HTTP_PROXY environment variables for proxy configuration and SSL_CERT_FILE for custom CA certificates, enabling use in corporate environments with SSL inspection. ## Motivation Users in corporate environments often need to route traffic through HTTP proxies and use custom CA certificates for SSL inspection. Without this support, exo cannot download models in these network configurations. ## Changes - Added `HTTPS_PROXY`/`HTTP_PROXY` environment variable support to `create_http_session()` in `download_utils.py` - Added `SSL_CERT_FILE` environment variable support for custom CA certificate bundles, falling back to certifi's default bundle ## Why It Works - `aiohttp.ClientSession` natively supports the `proxy` parameter for routing requests through HTTP proxies - `ssl.create_default_context(cafile=...)` accepts a custom CA bundle path, allowing corporate CAs to be trusted - Using environment variables is consistent with the codebase's existing configuration patterns (e.g., `EXO_HOME`, `HF_ENDPOINT`) ## Test Plan ### Manual Testing - Set `HTTPS_PROXY` environment variable and verified model downloads route through proxy - Set `SSL_CERT_FILE` to custom CA bundle and verified SSL verification succeeds with corporate SSL inspection ### Automated Testing - No automated tests added; this change is configuration-only and does not alter existing behavior when environment variables are unset	2026-01-18 12:05:50 +00:00
Alex Cheema	c5158bee53	Add pre-commit checks documentation to AGENTS.md (#1184 ) ## Motivation CI failures can be avoided by running checks locally before committing. This adds clear documentation to AGENTS.md so that AI agents (and humans) know exactly which checks must pass before pushing code. ## Changes Added a new "Pre-Commit Checks (REQUIRED)" section to AGENTS.md that: - Lists all 4 required checks (basedpyright, ruff, nix fmt, pytest) - Provides a one-liner to run all checks in sequence - Notes that `nix fmt` changes must be staged before committing - Explains that CI runs `nix flake check` which verifies everything ## Why It Works Clear documentation prevents CI failures by ensuring contributors run checks locally first. The one-liner command makes it easy to run all checks before committing. ## Test Plan ### Manual Testing - Verified the documented commands work correctly ### Automated Testing - N/A - documentation only change Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com>	2026-01-17 21:50:24 +00:00
rltakashige	5c8a237940	Handle model timeouts (#1177 ) - Add eval with a timeout. - Add fast synch flag ## Motivation Because of the experimental FAST SYNCH flag, some models may not work. This PR catches when this occurs and allows users to specify a run without fast synch ## Changes - Adds a flag to enable or disable fast synch (--fast-synch and --no-fast-synch) - Adds a heuristic timeout - Reduces exo_bench default timeout to 10 minutes. ## Why It Works Heuristic timeout assumes normal loading times on Mac devices (60 + model size in gb / 5: e.g. DeepSeek takes up to 120 seconds to load on tensor parallel, and timeout is set to 60 + 120 = 180s. We could raise this value if necessary. ## Test Plan ### Manual Testing Catches that GPT OSS fails to load in Tensor RDMA Can launch with --no-fast-synch flag to launch GPT OSS. GPT OSS 20B TP with fast synch <img width="3064" height="456" alt="image" src="https://github.com/user-attachments/assets/f6e25cd8-8621-4e99-99fe-292ee05c4035" /> TP without fast synch <img width="3098" height="496" alt="image" src="https://github.com/user-attachments/assets/d36453d9-6686-4cfe-aa7c-a7d458369d4d" /> [Note: the performance is really not great as fast synch is off] (As a sanity check) PP with fast synch <img width="3124" height="496" alt="image" src="https://github.com/user-attachments/assets/e97d4547-c6fa-483d-badb-4b371b900b4c" /> PP without fast synch <img width="3078" height="508" alt="image" src="https://github.com/user-attachments/assets/b2e20dfd-4b0e-4295-8a92-417dfe745c28" /> PP without RDMA <img width="3070" height="498" alt="image" src="https://github.com/user-attachments/assets/a8509d68-0aef-4cda-bca5-a67d39a0801e" /> TP without RDMA <img width="3068" height="496" alt="image" src="https://github.com/user-attachments/assets/b5691429-89f4-4369-bcf2-8fde2ad7154a" />	2026-01-16 20:25:12 +00:00
rltakashige	745343c705	Return error responses for Chat Completions (#1173 ) - Error chunks - Use error handling in exo_bench.py ## Motivation Return when an error occurs so that generation stops. Adding timeouts is a separate TODO for model loading and chat completions. ## Changes - Return HTTP exceptions as JSON responses in an OpenAI compatible format. - Context manager for generation to catch and return error messages. - Use error handling in exo_bench.py. ## Test Plan ### Manual Testing Manually tested that exo_bench returns on failures within and outside generation ### Automated Testing <!-- Describe changes to automated tests, or how existing tests cover this change --> <!-- - -->	2026-01-16 19:24:37 +00:00
				`@@ -0,0 +1 @@`
				`"""API adapters for different API formats (Claude, OpenAI Responses, etc.)."""`