Files
Daniel Hiltgen 855f4bf989 Report cached prompt tokens (#17943)
* Report cached prompt tokens

Add prompt_eval_cached_count to native responses and expose equivalent cached-token fields through the OpenAI- and Anthropic-compatible APIs. Keep prompt_eval_count as the logical input total while excluding cache hits from CLI and benchmark prefill rates. Surface processed and cached prompt counts in benchmark output.

Collect cache counts from llama-server and MLX, preserve coherent metrics across two-pass structured generation.

Fixes #8008

Related to #15758

* review comments
2026-09-02 09:30:44 -07:00
..
2026-09-02 09:30:44 -07:00
2026-08-25 19:34:13 -07:00
2026-09-02 09:30:44 -07:00
2026-08-11 14:36:39 -07:00
2026-08-11 14:36:39 -07:00
2026-08-11 14:36:39 -07:00
2026-08-11 14:36:39 -07:00
2026-08-11 14:36:39 -07:00
2026-08-11 14:36:39 -07:00
2026-08-11 14:36:39 -07:00