# opencoti-llamafile — usage guide How this engine diverges from upstream [Mozilla-Ocho llamafile](https://github.com/Mozilla-Ocho/llamafile), what the added features are, how each is gated, its knobs and defaults, its limitations, and which features are meant to be used together. Audience: anyone running the packaged `opencoti-llamafile---.llamafile` artifact as a local inference server. This is the **narrative** guide; the two companion references are [`docs/llamafile-flags.md`](llamafile-flags.md) (every flag / env / JSON knob, with defaults, reconciled against the engine source) and [`docs/llamafile-artifacts.md`](llamafile-artifacts.md) (downloads, GPU payloads, side-load paths, verification). Deep-dive design docs live in [`docs/features/`](features/), measured evaluations in [`docs/evaluations/`](evaluations/). --- ## Supported / target model families — read this first opencoti-llamafile loads **any GGUF that upstream llama.cpp loads** — that part is inherited unchanged. But the opencoti feature set (KV tiers, rolling-KV, DCA, MTP, sparse attention, RYS) is developed, tuned, and correctness-gated on **two model families**, in a deliberate primary/secondary split: ### Gemma-4 — PRIMARY target | Model | Kind | Notes | |---|---|---| | Gemma-4 26B-A4B-128e ("A4B") | MoE | the flagship serving target; MTP-validated with its native gemma4-assistant drafter | | Gemma-4 12B / 31B | dense | full feature validation incl. RYS + DCA | | Gemma-4 E2B / E4B | elastic (E-series) | shared-KV elastic layers supported; MTP drafters available | Gemma-4 is what the engine is *for*: its unusual head dims (256 and 512), the iSWA sliding/global dual KV cache, and the per-size [gemma4-assistant MTP drafters](https://huggingface.co/ManniX-ITA) all have dedicated kernels and graph paths here that upstream lacks or handles slowly. `--spec-type draft-assistant`, D256/D512 FA-VEC + scalar-MMA decode, iSWA-aware rolling-KV/SharedKVPool/DCA wiring are all Gemma-4-first features. ### Qwen — SECONDARY target / verification family | Model | Kind | Notes | |---|---|---| | Qwen3.5 / Qwen3.6 (e.g. 35B-A3B) | dense / MoE / hybrid (gated-delta-net) | NextN self-spec MTP (`--spec-type draft-mtp`, no external drafter needed) | | Qwen2.5-14B-1M | dense, native-1M | the long-context/DCA validation vehicle | Qwen is the standard-architecture (head_dim 128) counterweight: every feature that ships is verified on it too, and it carries one feature Gemma doesn't — **NextN self-speculation** (the model's own MTP head drafts; fused multi-step, at/above upstream parity). ### Everything else Other architectures run with upstream behavior and safe fallbacks, but opencoti features are **unvalidated** there, and some are arch-gated: MTP needs NextN tensors (Qwen-style) or a gemma4-assistant drafter; RYS `--repeat-layers` supports the qwen2/qwen3(+MoE)/qwen3.5/qwen3next/gemma-4 forward loops; DCA is validated on Gemma-4 and Qwen2.5-1M. Quality gates (KLD, RULER-niah) were run on the two families above — re-gate before trusting aggressive KV tiers on anything else. --- ## 1. Relationship to upstream llamafile opencoti-llamafile is **upstream llamafile plus an additive patch series** (`patches/` in the HF repo, `vendors/patches/llamafile/` in the git repo; the upstream base version and the exact patch list for a given cut are recorded in its `RELEASE_NOTES` and every artifact's `MANIFEST.json`). Three properties are contractual: 1. **Off means off.** Every opencoti feature is opt-in behind a flag, env var, or per-request JSON field. With no opencoti flags set, the engine's compute path is **byte-identical to upstream** — this is a regression gate on every patch, not an aspiration. 2. **Lossless by proof, not vibes.** Features that touch the forward pass are gated by logit-equivalence / KLD / RULER-retrieval against vanilla, never by "the output looks fine". Speculative decode is verified-lossless (the output *is* the target model's). 3. **Single file, zero dependencies.** The artifact is a Cosmopolitan APE: one file runs on Linux/macOS/Windows/BSD, x86_64 and aarch64. In the full x86_64 artifact the CUDA and Vulkan backends are embedded and self-extract to `~/.llamafile/v/opencoti--/` on first GPU run; the `-win` variant ships without them (see [llamafile-artifacts.md](llamafile-artifacts.md)). TCQ codebooks and quantization tables are compiled in. No installer, no downloads. What upstream gives you is unchanged: the server API (`/completion`, `/v1/chat/completions`, `/props`, `/slots`, …), GGUF loading, sampling, chat templates. opencoti adds serving-efficiency machinery on top, aimed at **multi-session agentic serving on a fixed VRAM budget**: more concurrent sessions per card, longer usable context, faster decode. ```bash chmod +x opencoti-llamafile---x86_64.llamafile sh ./opencoti-llamafile---x86_64.llamafile --server --port 8080 \ -m model.gguf -ngl 99 # --version → opencoti-- ; without --server you get the chat CLI ``` > **Note (Linux):** launch via `sh ./file.llamafile` if your kernel > lacks binfmt_misc APE registration. ### 1.1 Artifact variants, GPU payloads, side-load Moved to [`docs/llamafile-artifacts.md`](llamafile-artifacts.md): the full artifact matrix (Linux / Windows / aarch64 / universal, all with compressed GPU payloads), the published side-load DSO set, the versioned `~/.llamafile/v/opencoti--/` runtime path, and byte-level verification. Short version: the host APE is identical in every variant; pick the file whose embedded payloads match your OS, or take the small bare `-win` file and side-load a DSO. ### 1.2 Backends — CUDA, Vulkan, CPU One artifact serves three compute backends; selection is `--gpu {auto,nvidia,vulkan,amd,apple,disable}` (default `auto`, which probes CUDA first, then Vulkan): - **CUDA** (`--gpu nvidia`) — the primary, fully-validated backend; every opencoti kernel family (turbo/TCQ in-register FA, DCA, sparse attention, MTP verify) has CUDA instances. sm_75→120f on x86_64, sbsa Blackwell-class on aarch64. - **Vulkan** (`--gpu vulkan`, alias `vk`; payloads embedded since c7) — AMD, Intel and NVIDIA through one DSO. The opencoti Vulkan port carries the sparse-attention parity gate, streaming flash-attention, LSE emission and coopmat FA shaders; on NVIDIA it is the fallback when CUDA can't load, on AMD/Intel it is the GPU story. Feature parity is asserted knob-for-knob against CUDA where ported (vk==cuda gates), but CUDA remains the performance reference. - **CPU** — always available, no payload needed; iqk flash-attention kernels (`--iqk-flash-attn`, default auto) accelerate it. Deleting the side-loaded/extracted DSO forces guaranteed CPU serving. **AMD Radeon on Windows: AMD Software 26.9.2 or later for Vulkan** (display driver **32.0.32015.2008**, 2026-09-28). This is the minimum driver; older ones are not supported for `--gpu vulkan`. Measured on Windows 11 with an RX 9070 XT (2026-10-04, one machine, one card; other Radeon cards on the old driver untested): | driver | a model loaded through Vulkan and left idle | a clean stop, then the next start | |---|---|---| | AMD Software 26.8.1 — display driver 32.0.31041.1004 | **fails**: within 15 s to 2 min Windows logs `0x141 VIDEO_ENGINE_TIMEOUT_DETECTED` in `amdkmdag.sys`, the next request gets `VK_ERROR_DEVICE_LOST`, the card can stay disabled (code 22 / 31) until a reboot; the machine froze or bug-checked three times | **fails** 4 of 4 | | AMD Software 26.9.2 — display driver 32.0.32015.2008 | clean, 9 idle periods of 9 (stock llama.cpp and this engine) | clean 10 of 10 | The fault is in the driver, not in the engine: it reproduces with a plain Vulkan program that leaves several GiB of device-local, host-visible memory mapped while the queue is idle, and with stock llama.cpp. The engine does not work around it (it sets no `GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM`); update the driver. The chipset package updates only the integrated GPU's driver and does not fix the discrete card. A host that cannot update can at least keep the engine's exit from touching the card with `OPENCOTI_SHUTDOWN_GPU=off` (§ "Stopping the server" in `docs/features/server_shutdown.md`); the idle fault remains. **Stopping the server.** `POST /shutdown` (loopback only), SIGTERM / SIGINT and — with `OPENCOTI_PARENT_PID=` — the death of the supervising process all take one path: the models are freed, every Vulkan device is made idle and destroyed, exit 0. `OPENCOTI_SHUTDOWN_GPU=auto|off|free|device|full` selects how far that goes (`auto` = `full`); `OPENCOTI_SHUTDOWN_DEADLINE_MS` (default 10000) bounds it. `OPENCOTI_LIST_DEVICE_IDS=1` makes `--list-devices` append a stable ` id=` (PCI address, or `uuid:` on Windows) to each line. **Flash attention is a tri-state** (`-fa {on,off,auto}`, default `auto`): `auto` resolves per device at graph-reserve time and prints its decision (`resolve_fused_ops: Flash Attention enabled` on 0.10.5-based builds). Explicit `on`/`off` remain authoritative. The opencoti split-attention paths (head-split, NEO, GPU_STREAM) are resolver-aware — non-split layers still resolve FA normally. --- ## 2. Feature map — what exists and how it's gated | Feature | Default | Turn on with | Class | |---|---|---|---| | Session-keyed KV reuse | off (per request) | `session_id` JSON field | latency | | ReST-KV retention eviction | **off** | `--rest-kv-eviction` | quality-under-overflow | | KV quantization (scalar) | f16 | `-ctk` / `-ctv` | capacity | | TurboQuant / TCQ KV tiers | off | `-ctk`/`-ctv turbo*` | capacity | | Auto KV-tier policy | **off** | `OPENCOTI_KV_AUTO_TIER=1` | capacity (policy) | | PolyKV pool (SharedKVPool) | off (per request) | `shared_pool_slot` JSON field | multi-agent capacity | | PolyKV control-plane API | **off** | `--polykv-max-pools N` + `/polykv/*` routes | multi-agent orchestration | | Rolling-KV window / spill | engages only when the KV cache does not fit VRAM (§3.3) | `--vram-target`, `--kv-rolling-window on` | capacity | | Mixed-KV spilled tail | off | `-ctkt` / `-ctvt` | capacity | | DCA long-context | **off** | `--dca on` | context extension | | Sparse attention (block-selector) | **off** | `--sparse-attn on` | long-ctx decode speed | | Sparse-V | auto on iSWA+quant-V, else off | `TURBO_SPARSE_V_TAU` | decode speed | | MTP speculative decode | **off** | `--spec-type` + drafter | decode speed | | RYS layer duplication | **off** | `--repeat-layers` | quality | | RYS probe | off | `--rys-probe` | tooling | | Lazy slot context | off | `--slot-initial-ctx`, `--slot-shrink-idle-ms` | embedder memory | | GPU-share TDM pacing | **auto** (engages only with ≥2 opencoti servers on one GPU) | `--gpu-share-weight N`, `GET /gpu/peers` | multi-server fairness | | Elastic parallel slots | **off** | `--max-parallel N` (+ tps floor / VRAM reserve) | serving capacity | | iSWA SWA-pool seq budget | off (= n_seq_max) | `--swa-seq-budget B` | serving capacity (VRAM) | | KV admission gate | **on (`enforced`)** | `--admission-poolless {off,warn,enforced}` | serving capacity | | Sampling placement | **`auto` (routes)** | `--sampling-placement {device,cpu,auto}` | decode speed | | Auto-MTP depth policy | **`taper`** | `--auto-mtp-policy {allocator,taper,always,off}` | decode speed | | Introspection API | **always on** | `GET /props`, `GET /slots` | observability | Every boot flag also has an env twin (`OPENCOTI_LLAMAFILE_` for adapter-typed fields, `LLAMA_ARG_*` for llama.cpp-registered ones). The per-flag reference — every knob, value set, default, and env name — is [`docs/llamafile-flags.md`](llamafile-flags.md); this table is the map, not the contract. ### 2.1 Sharing one GPU between servers (gpu-share, c5/c6) Multiple opencoti-llamafile servers on the SAME physical GPU discover each other zero-conf (shared-memory peer registry keyed on the PCI bus id, so `CUDA_VISIBLE_DEVICES` ordering doesn't matter) and pace themselves by **deterministic weighted time-division**: a repeating period is carved into one contiguous window per active peer, proportional to `--gpu-share-weight` (default 1). The split is exact by construction and work-conserving — an idle peer's time flows to the busy ones, and a **single server pays zero overhead** (pacing only engages with ≥2 active peers). The period adapts for interactive latency (80–400 ms; worst token stall at 2:1 ≈ 80 ms). Measured on a 3090: 2:1 → 1.97:1 at ~solo aggregate; on Windows (5080): 2.00:1. Watch peers live via `GET /gpu/peers`. Full design + measurements: `docs/features/gpu_share.md`. Crashed peers age out ≤3 s; no daemon, no MPS, works on Windows. --- ## 3. KV capacity stack — PolyKV **PolyKV** is the umbrella name for this whole stack: the compressed shared KV pool. Concretely it is the KV quantization tiers of §3.1 plus the multi-agent SharedKVPool of §3.4, stacking with the auto-tier policy (§3.2) and the rolling-KV window (§3.3), and orchestrated at runtime through the control-plane API of §3.5 (pools/fork/admission/capacity/tps REST surface + the `opencoti-langgraph` package). If you arrived here looking for "PolyKV" from an announcement: §3.4 is the shared-prefix pool itself; §3.1 is what the pooled cells are made of; §3.5 is how an agent framework drives it. These four features share one goal — **fit more context / more sessions in fixed VRAM** — and are designed to stack. Recommended order of adoption: scalar quant → auto-tier → rolling-KV → turbo tiers → SharedKVPool. ### 3.0 `--kv-unified` — one shared cell pool vs per-slot reservations **This flag decides what `-c` and `--parallel` actually mean.** Get it wrong and every capacity number in this document is wrong too. The two layouts are not a tuning preference; they are different allocation models. #### With `--kv-unified` (recommended for agentic / multi-session) `n_ctx_seq = n_ctx` (`src/llama-context.cpp:456`) and the cache runs as a **single pool of `-c` cells with one stream** (`n_stream = 1`, `src/llama-kv-cache.cpp:701`). - **`-c` is the whole pool, shared.** It is NOT divided by `--parallel`. - **Every slot advertises the full `-c` as its context.** The server sets `slot.n_ctx = llama_n_ctx_seq(ctx)` with no division (`tools/server/server-context.cpp:1894`). One long session may legally consume the entire pool. - **`--parallel` costs no base-KV VRAM.** Raising it does not enlarge the cache. (The one exception is iSWA — see below.) - **There is no isolation.** Cells are first-come-first-served, so a greedy session can starve its neighbours. This is exactly why the PolyKV admission gate (§3.5) exists: with a shared pool, admission control is not optional, it is the isolation mechanism. - Pool prefixes are shared **zero-copy** (§3.4). Provisioning rule of thumb (bug-2287) — size the pool so the expected working set fits: ``` -c ≈ tokens_per_session × (--parallel + --polykv-max-pools) ``` This is a *sizing guideline you apply*, not something the engine enforces. PolyKV pools occupy reserved seq-ids **above** the slot range (`seq = n_parallel + k`), so they need their own share of the pool. #### Without `--kv-unified` (per-slot reservation) `n_ctx_seq = n_ctx / n_seq_max`, and the cache splits into `n_stream = --parallel` independent per-slot caches. - **`-c` is divided.** Each slot gets a hard-reserved `-c / --parallel` and can never exceed it, no matter how idle the others are. - **`--parallel` directly divides usable context per session.** - **Isolation is free** — no session can affect another's capacity. - Idle slots waste their whole reservation. - Pool prefixes are **copy-shared**, not zero-copy. #### Choosing | Want | Use | |---|---| | Many sessions of *varying, unpredictable* length | `--kv-unified` + admission control | | Long single sessions that should use all VRAM | `--kv-unified` | | Hard per-tenant context guarantees | omit `--kv-unified` | | Benchmarks where per-slot capacity must be identical and fixed | omit `--kv-unified` | **The iSWA exception (Gemma-4).** On sliding-window models the *SWA* cache is the one allocation that scales with the seq-id count, **even under `--kv-unified`**, and it is allocated eagerly in VRAM at boot (`src/llama-kv-cache-iswa.cpp:71`): ``` n_seq_max = --parallel + --polykv-max-pools size_swa = n_swa × n_seq_max + n_ubatch (padded to 256) ``` **`--polykv-max-pools` counts.** A pool occupies a reserved seq-id above the slot range, so it costs a full SWA window exactly like a slot does. Budgeting against `--parallel` alone silently under-counts. Measured on bs2 GPU0, `gemma-4-31B-it-Q6_K`, `n_swa = 1024`, `n_ubatch = 512`, `-ctk/-ctv q8_0` (boot-only probe, 5/5 exact): | `--parallel` | `--polykv-max-pools` | `n_seq_max` | `-c` | base cells | SWA cells | |---:|---:|---:|---:|---:|---:| | 64 | 4 | 68 | 417792 | 417792 | 70144 | | 16 | 4 | 20 | 417792 | 417792 | 20992 | | 8 | 4 | 12 | 417792 | 417792 | 12800 | | 8 | 0 | 8 | 417792 | 417792 | 8704 | | 8 | 4 | 12 | 131072 | 131072 | 12800 | Read the two invariants off the table: the **base cache equals `-c` exactly and is identical at `n_seq_max` 8 and 68**, and the **SWA cache is unchanged when `-c` drops 3.2×**. They are independent axes. At 31B's ~425 KiB per SWA cell (50 SWA layers × 16 kv-heads × 256 head_dim × K+V at q8_0) the `--parallel 64 --polykv-max-pools 4` row is **~28.4 GiB of SWA alone** — which is what forces the context budget down on large-slot Gemma runs. `--parallel 8 --polykv-max-pools 4` is ~5.2 GiB. Qwen-class models have no SWA split and pay nothing here. Plan work to make this budget independent of the slot count: `docs/features/elastic_parallel.md`. ### 3.1 KV quantization: scalar types + TurboQuant/TCQ tiers (PolyKV M6) The KV cache type is set per-tensor-half: `-ctk ` (keys) and `-ctv ` (values), independently — **asymmetric pairs are first-class** (e.g. `-ctk q8_0 -ctv q4_0`). Supported types: `f16`, `bf16`, `q8_0`, `q6_0`, `q5_1`, `q5_0`, `q4_0` (scalar), `turbo2`, `turbo3`, `turbo4`, `turbo8` (TurboQuant, MSE-optimal with Walsh-Hadamard rotation + InnerQ), `turbo2_tcq`, `turbo3_tcq` (trellis-coded, Viterbi-encoded). **Which to pick (measured):** - **8-bit / 4-bit: use `q8_0` / `q4_0`.** The native scalar types dominate turbo8/turbo4 at equal width — turbo earns nothing there. - **`-ctk q8_0 -ctv q4_0`** is the workhorse asymmetric pair: keys keep 8-bit fidelity (attention logits are K-sensitive), values take the compression. - **3 bits and below is TurboQuant territory:** `turbo3` Pareto-beats q4_0 (90 vs 129 MiB KV at equal quality, teacher-forced TV 0.0067 vs 0.0094); `turbo2` is the smallest logit-equivalent KV that exists (~2 bit) — the 256k-context play. TCQ variants trade encode cost for a further fidelity step at the same width. - All shipped tiers pass logit-equivalence gates; decode runs the quantized data **in-register** in the flash-attention kernel (no f16 materialization) for turbo2/3/4 and TCQ. Limitations: turbo8 uses a materialize fallback (not fused); at Gemma-4's head_dim 512 only turbo2/turbo3 have fused D=512 instances; prefill on very long prompts uses a hybrid path automatically. Quality validation on Gemma franken-merges must use retrieval (niah), not perplexity. **KVarN (`kvarn2`…`kvarn8`) is the compressed KV tier from c8** and replaces the turbo/TCQ tiers above, which are deprecated and frozen ([ADR 0006](decisions/0006-kvarn-integration.md), §9d). Select it on both sides: `-ctk kvarn3 -ctv kvarn3`; on iSWA models (Gemma-4) the sliding-window ring takes its own type with `--cache-type-k-swa` / `--cache-type-v-swa` (flag reference: `docs/llamafile-flags.md`). At 2–4 bits it is 2–13× better on teacher-forced KLD than the matched turbo/tcq/q4_0 tier and ties q6_0/q8_0 (`docs/evaluations/kvarn_vs_turbo.md`). **KVarN output depends on `-ub` (by design).** KVarN stores the cache in 128-token groups. The groups the current ubatch is writing are held **exact (f16)** in a stage window of `ceil(n_ubatch / 128) + 1` groups per sequence; a group is quantized only when it leaves that window. A larger `-ub` therefore keeps more of the most recent context exact and reads less of it back quantized, so the logits — and any KLD/PPL measured against an f16 reference — move with `-ub`. This is not a defect and not noise: Gemma-4 E4B, `kvarn3`, 4 sequences on a unified cache, KLD vs f16 = 0.0098 / 0.0073 / 0.0053 at `-ub` 512 / 1024 / 2048. The scalar types (`f16`, `q8_0`, `q4_0`, …) have no stage window and are `-ub`-invariant. Consequences: - compare KVarN runs **only at the same `-ub`** (a quality number without its `-ub` is not reproducible); - `--kvarn-stage-window N` sets the window in groups instead of deriving it from `-ub` (non-SWA layers; the Gemma-4 sliding-window ring takes `OPENCOTI_KVARN_SWA_WINDOW=N`). Use it to hold the exact window fixed across a `-ub` sweep; keep `N ≥ ceil(n_ubatch / 128) + 1` or part of the in-flight ubatch is read back quantized. Texts are still not guaranteed bit-identical across `-ub`; - the window costs f16 cells: `(ceil(n_ubatch / 128) + 1) × 128` per sequence and per cache half. On a unified multi-slot iSWA server the sliding-window pool grows to hold it (`0538`); a budgeted or deliberately short pool keeps a smaller window and says so in a startup WARN (`cells hold a stage of N group(s) per sequence, a M-token ubatch spans K`). ### 3.2 Auto KV-tier (`OPENCOTI_KV_AUTO_TIER=1`) Boot policy: pick the **least-compressing scalar pair that keeps the whole KV resident** in the VRAM budget; if even that spills, the T\* model decides between "small f16 spill" and "quantize one tier down" by predicted tokens/s drop. Knobs (env): `OPENCOTI_KV_AUTO_TIER=1` (master), `OPENCOTI_KV_TSTAR_DROP` (target drop, default 20%), `OPENCOTI_KV_TSTAR_MAX_SPILL_MIB` (default 800), `OPENCOTI_KV_AUTO_TIER_TAIL=1` (also auto-pick a q4_0 spilled tail). Explicit `-ctv` disables auto entirely; explicit `-ctk` holds K and walks only V. Dense full-attention models only (iSWA models keep f16). Read back what it decided: `GET /props → .opencoti.kv.effective` — the *configured vs effective* split exists exactly because auto-tier may override you. ### 3.3 KV rolling window (a KV cache larger than VRAM) "KV doesn't have to fit." Since c10 the window is **one sliding window per KV layer, adaptive in size, and it engages only when the KV cache does not fit VRAM**: - Under a VRAM budget each KV layer keeps `C` cells resident on the device; `C` is printed at load (`C= cells resident`, and `position window = / cells resident`). - **Below the limit** (context ≤ `C`) a capped boot decodes with exactly the kernels of an uncapped boot and produces the same output, bit for bit — f16, q4_0, q8_0 and KVarN caches, one slot and `--kv-unified`, CUDA and Vulkan (gated by comparing the 256 log-probabilities of capped and uncapped boots). The rows marked "below the limit" in the tables are that case: decode within 2 % of uncapped (one run per row). Prompt processing is the one difference: a capped boot reads the quantized cache directly instead of converting it to f16 first (the conversion buffer would take VRAM from `C`), so prefill can be slower even below the limit — 0 to 14 % in the tables below. - **Past the limit** the cells `[C, depth)` live in pinned host RAM. Every token, each KV layer copies its whole host part into a device staging slot (two slots, so the copy for the next layer overlaps this layer's compute) and runs ONE attention launch over resident + staged keys. The host part grows with the depth; nothing is evicted, nothing is approximated. Flags (the tables were measured with exactly these): `--vram-target ` (TOTAL VRAM cap — weights + KV + compute; set it to your card's size, or use it to simulate a smaller card) and `--kv-rolling-window on`. `-ctk` / `-ctv` choose the KV type: **q4_0 is the type the window is built for** (the host part is small and the read is the plain q4_0 attention path); KVarN 4-bit works under a cap too (its records are staged and prefetched layer by layer); f16 works but carries four times the bytes over the link. The link: the cost past the limit is the host-to-device copy, so the PCIe link decides it. The engine reads the link from the system on Linux and Windows (`pcie profile: … source=system` in the log); `OPENCOTI_LINK_GBPS=` forces it (or `OPENCOTI_LINK_GBPS==,…` per device), and `--server --link-probe` measures it and exits. Composes with `--parallel N`, `--kv-unified` (one window over the shared pool), the MTP drafter and a Gemma assistant draft (their own caches stay resident; the cap leaves less for the target's), two GPUs, and context shift on a capped cache. **Measured on the published c10 bytes** (2026-10-07; development engine `2610071819001` with the published libraries for bs2, the c10 release engine `2610072136001` for the RX 9070 XT — the same source but for the version line). One request per row, a fresh prompt of the given length (no prompt cache; the same prompts on every arm), then 256 greedy tokens; decode = tok/s over those 256 tokens; prompt = prefill tok/s. "uncapped" = the same model with all VRAM (the big-card reference); "layers in RAM" = no cap, but enough layers on the CPU (`-ngl`) to fit the same card the old way. ### RTX PRO 6000 (bs2, PCIe 5.0), Linux — CUDA and Vulkan The cap simulates a smaller card: 11.5 GB (a 12 GB card) for OmniMerge v6, 14.5 GB for the 14B, 21 GB for Gemma 4 31B. #### OmniMerge v6 IQ2_M (27B hybrid), q4_0 KV, CUDA, cap 11.5 GB Resident under the cap: C = 58,112 cells; deeper than C the rest of each layer is read from host RAM every token. | context (tokens) | below the limit? | decode tok/s: window | uncapped | layers in RAM | prompt tok/s: window | uncapped | layers in RAM | |---|---|---|---|---|---|---|---| | 8,185 | yes | 90.4 | 90.7 | 6.8 | 3030 | 3057 | 1730 | | 30,668 | yes | 86.8 | 87.1 | 6.0 | 2872 | 2978 | 1817 | | 59,162 | no | 82.5 | 83.1 | 5.3 | 2554 | 2719 | 1711 | | 88,243 | no | 73.3 | 78.8 | 4.7 | 2174 | 2516 | 1627 | | 126,373 | no | 40.5 | 73.8 | 4.1 | 1826 | 2278 | 1508 | #### OmniMerge v6 IQ2_M (27B hybrid), q4_0 KV, Vulkan, cap 11.5 GB Resident under the cap: C = 39,936 cells; deeper than C the rest of each layer is read from host RAM every token. | context (tokens) | below the limit? | decode tok/s: window | uncapped | prompt tok/s: window | uncapped | |---|---|---|---|---|---| | 8,185 | yes | 67.4 | 67.3 | 2168 | 2174 | | 30,668 | yes | 63.5 | 63.0 | 1850 | 1853 | | 59,162 | no | 55.7 | 58.7 | 1441 | 1525 | | 88,243 | no | 51.2 | 55.3 | 1227 | 1292 | | 126,373 | no | 32.7 | 50.7 | 1025 | 1077 | #### OmniMerge v6 IQ2_M, KVarN 4-bit KV, CUDA, cap 11.5 GB Resident under the cap: C = 61,056 cells; deeper than C the rest of each layer is read from host RAM every token. KVarN keeps its records in groups of 128 cells. Its window is flexible: 399 groups (51,072 cells) are the window, and while no cell lies past them the 78 groups the read slots will need stay on the device too — C above is that larger figure (477 groups; Vulkan: 482 + 34 = 516 groups = 66,048 cells). Past it those groups move to host RAM and their memory becomes the read slots. | context (tokens) | below the limit? | decode tok/s: window | uncapped | prompt tok/s: window | uncapped | |---|---|---|---|---|---| | 8,185 | yes | 90.4 | 90.5 | 2995 | 2992 | | 30,668 | yes | 82.3 | 82.3 | 2902 | 2890 | | 59,162 | yes | 76.5 | 76.6 | 2585 | 2575 | | 88,243 | no | 56.7 | 71.7 | 2226 | 2304 | | 126,373 | no | 34.3 | 64.0 | 1865 | 2033 | #### OmniMerge v6 IQ2_M, KVarN 4-bit KV, Vulkan, cap 11.5 GB Resident under the cap: C = 66,048 cells; deeper than C the rest of each layer is read from host RAM every token. | context (tokens) | below the limit? | decode tok/s: window | uncapped | prompt tok/s: window | uncapped | |---|---|---|---|---|---| | 8,185 | yes | 58.5 | 57.9 | 1538 | 1534 | | 30,668 | yes | 41.9 | 41.8 | 1459 | 1457 | | 59,162 | yes | 30.8 | 30.8 | 1342 | 1341 | | 88,243 | no | 20.6 | 24.2 | 1236 | 1241 | | 126,373 | no | 14.5 | 19.1 | 1115 | 1127 | #### OmniMerge v6 IQ2_M, q4_0 KV, drafter (MTP) ON, CUDA, cap 11.5 GB Resident under the cap: C = 1,536 cells; deeper than C the rest of each layer is read from host RAM every token. With the drafter on, decode speed is set by how many drafted tokens are accepted, and that depends on the text generated, which can differ between two boots of the same engine — at 59k the window run produced a continuation that repeats the prompt, the lookup drafter copied it (mean accepted length 12.75 against 2.50), and the row shows the text, not the engine. Compare rows only where the acceptance columns agree. The drafter's own cache stays resident, so the cap leaves the target 1,536 cells. | context (tokens) | below the limit? | decode tok/s: window | uncapped | prompt tok/s: window | uncapped | draft acceptance (mean accepted length): window | uncapped | |---|---|---|---|---|---|---|---| | 8,185 | no | 100.3 | 129.9 | 2148 | 2638 | 0.55 (2.63) | 0.52 (2.55) | | 30,668 | no | 110.7 | 189.2 | 2086 | 2573 | 0.97 (3.91) | 0.96 (3.86) | | 59,162 | no | 210.8 | 120.0 | 1911 | 2319 | 0.97 (12.75) | 0.50 (2.50) | | 88,243 | no | 40.5 | 107.5 | 1757 | 2115 | 0.48 (2.44) | 0.42 (2.85) | | 126,373 | no | 31.6 | 116.2 | 1584 | 1896 | 0.48 (2.43) | 0.56 (3.06) | #### OmniMerge v6 IQ2_M, q4_0 KV, drafter (MTP) ON, Vulkan, cap 11.5 GB Resident under the cap: C = 256 cells (the drafter's resident cache takes the budget: 0 MiB left for the target's); deeper than C the rest of each layer is read from host RAM every token. | context (tokens) | below the limit? | decode tok/s: window | uncapped | prompt tok/s: window | uncapped | draft acceptance (mean accepted length): window | uncapped | |---|---|---|---|---|---|---|---| | 8,185 | no | 56.1 | 70.5 | 1716 | 1992 | 0.54 (2.62) | 0.55 (2.63) | | 30,668 | no | 68.1 | 94.2 | 1563 | 1715 | 0.96 (3.86) | 0.96 (3.86) | | 59,162 | no | 34.9 | 59.3 | 1347 | 1409 | 0.49 (2.47) | 0.56 (2.68) | | 88,243 | no | 28.6 | 51.8 | 1155 | 1193 | 0.48 (2.44) | 0.52 (2.68) | | 126,373 | no | 23.3 | 42.7 | 941 | 991 | 0.38 (2.62) | 0.43 (2.54) | #### OmniMerge v6 IQ2_M, f16 KV (control), CUDA, cap 11.5 GB Resident under the cap: C = 7,680 cells; deeper than C the rest of each layer is read from host RAM every token. The control: an f16 cache is four times the bytes of q4_0, so under the same cap only 7,680 cells stay resident and every token moves four times as much over the link. Use q4_0 (or KVarN) with the window. | context (tokens) | below the limit? | decode tok/s: window | uncapped | prompt tok/s: window | uncapped | |---|---|---|---|---|---| | 8,185 | no | 91.2 | 92.1 | 3018 | 3149 | | 30,668 | no | 33.6 | 84.7 | 2669 | 3101 | | 59,162 | no | 15.3 | 76.3 | 2354 | 2839 | | 88,243 | no | 9.9 | 69.4 | 2051 | 2566 | | 126,373 | no | 6.7 | 62.0 | 1735 | 2241 | #### Qwen2.5-14B-Instruct-1M Q6_K, q4_0 KV, CUDA, cap 14.5 GB Resident under the cap: C = 41,984 cells; deeper than C the rest of each layer is read from host RAM every token. | context (tokens) | below the limit? | decode tok/s: window | uncapped | layers in RAM | prompt tok/s: window | uncapped | layers in RAM | |---|---|---|---|---|---|---|---| | 8,058 | yes | 91.4 | 91.3 | 12.9 | 5325 | 5595 | 2169 | | 30,144 | yes | 80.2 | 80.1 | 5.4 | 3616 | 4203 | 1895 | | 49,913 | no | 70.0 | 73.2 | 3.7 | 2702 | 3261 | 1648 | | 70,099 | no | 35.9 | 66.6 | 2.6 | 2125 | 2664 | 1468 | | 88,206 | no | 21.3 | 61.6 | 1.9 | 1792 | 2290 | 1330 | #### Qwen2.5-14B-Instruct-1M Q6_K, q4_0 KV, Vulkan, cap 14.5 GB Resident under the cap: C = 36,608 cells; deeper than C the rest of each layer is read from host RAM every token. | context (tokens) | below the limit? | decode tok/s: window | uncapped | prompt tok/s: window | uncapped | |---|---|---|---|---|---| | 8,058 | yes | 84.6 | 86.1 | 4988 | 5039 | | 30,144 | yes | 72.7 | 73.2 | 3722 | 3729 | | 49,913 | no | 57.6 | 64.9 | 2649 | 2998 | | 70,099 | no | 29.1 | 58.0 | 2013 | 2493 | | 88,206 | no | 18.8 | 53.0 | 1674 | 2167 | #### Gemma 4 31B Q4_K_M (iSWA), q4_0 KV, CUDA, cap 21 GB Resident under the cap: C = 16,640 cells; deeper than C the rest of each layer is read from host RAM every token. | context (tokens) | below the limit? | decode tok/s: window | uncapped | layers in RAM | prompt tok/s: window | uncapped | layers in RAM | |---|---|---|---|---|---|---|---| | 7,216 | yes | 52.6 | 52.6 | 10.5 | 3104 | 3122 | 1097 | | 27,075 | no | 48.7 | 50.4 | 8.2 | 2651 | 2885 | 1103 | | 45,032 | no | 45.1 | 48.6 | 6.9 | 2312 | 2610 | 1066 | | 63,017 | no | 41.7 | 46.9 | 6.0 | 2030 | 2393 | 973 | | 79,275 | no | 34.5 | 45.7 | 5.3 | 1822 | 2223 | 987 | #### Gemma 4 31B Q4_K_M (iSWA), q4_0 KV, Vulkan, cap 21 GB Resident under the cap: C = 256 cells; deeper than C the rest of each layer is read from host RAM every token. Gemma 4 has two caches: the sliding-window layers stay fully resident; C is the global-attention layers', and under this cap on Vulkan their budget holds only 256 cells, so the window is in use from the first token. | context (tokens) | below the limit? | decode tok/s: window | uncapped | prompt tok/s: window | uncapped | |---|---|---|---|---|---| | 7,216 | no | 50.4 | 51.9 | 1978 | 2258 | | 27,075 | no | 47.9 | 50.2 | 1517 | 1597 | | 45,032 | no | 45.2 | 48.7 | 1173 | 1245 | | 63,017 | no | 36.0 | 47.0 | 952 | 1015 | | 79,275 | no | 28.9 | 46.0 | 811 | 869 | ### RX 9070 XT (16 GB, PCIe 5.0 x16), Windows 11, Vulkan — OmniMerge v6 IQ2_M, cap 11.5 GB The same model and cap as above, on an AMD card under Windows (eleven2go, the c10 release engine and the published Vulkan library). Under the same cap less is left for the cache here (1,293 MiB free after loading, against 1,788 MiB on bs2 CUDA), so far less of it stays resident. #### q4_0 KV Resident under the cap: C = 9,728 cells. | context (tokens) | below the limit? | decode tok/s: window | uncapped | layers in RAM | prompt tok/s: window | uncapped | layers in RAM | |---|---|---|---|---|---|---|---| | 8,185 | yes | 36.5 | 36.4 | 6.8 | 747 | 744 | 581 | | 30,668 | no | 33.6 | 34.6 | 6.1 | 611 | 640 | 530 | | 59,162 | no | 31.6 | 32.7 | 5.4 | 508 | 541 | 460 | | 88,243 | no | 29.3 | 30.7 | 4.8 | 433 | 467 | 406 | | 126,373 | no | 23.8 | 28.6 | 4.2 | 364 | 395 | 350 | #### KVarN 4-bit KV Resident under the cap: C = 29,696 cells (the flexible window: 179 groups in the window + 53 for the read slots while no cell lies past them). | context (tokens) | below the limit? | decode tok/s: window | uncapped | prompt tok/s: window | uncapped | |---|---|---|---|---|---| | 8,185 | yes | 33.8 | 33.8 | 721 | 722 | | 30,668 | no | 23.3 | 24.7 | 666 | 671 | | 59,162 | no | 15.8 | 19.4 | 603 | 609 | | 88,243 | no | 11.9 | 15.9 | 550 | 557 | | 126,373 | no | 9.1 | 13.0 | 491 | 498 | ### 3.4 PolyKV SharedKVPool (multi-agent shared prefix) N agents attending **one physical copy** of a common prefix (system prompt + tool defs). Per-request JSON, no CLI flag: ```jsonc { "shared_pool_slot": 0, "shared_prefix_n_tokens": 4096, … } ``` Server must run with `--no-cache-idle-slots` (mandatory — the default idle-slot save/clear would evict the pooled prefix). Since c6 pools work on BOTH cache layouts: with `--kv-unified` the prefix is shared zero-copy (cell set-membership, one physical copy); on the split-stream cache (e.g. rolling-KV window mode) the share is a physical per-stream copy of the prefix — same semantics, higher footprint (device use stays window-capped in window mode). **What the pool speeds up — measured** (Gemma-4-26B-A4B Q4_K_M, RTX 3090, q8/q8 unified KV, P≈5k-token shared prefix, greedy fixed-length decode, 8 concurrent sessions unless noted): | axis | naive (N private copies) | shared pool | gain | |---|---|---|---| | KV cells (N=8) | ~8·P | P + suffixes | **6.9×** (~306 vs ~9 agents on a fixed buffer) | | prefill, 8 sessions joining | 22.8 s | 4.6 s | **~5×** (prefix enters KV once per pool) | | steady-state batched decode (N=8) | ~190 tok/s | 217 tok/s | **+14%** (8 queries read one physical prefix — L2 reuse, smaller cell span) | | multi-turn re-query (N=8) | 99 tok/s | 225 tok/s | **2.3×** (see note) | | iso-speed capacity | 8 sessions @ 24.0 tok/s each | ≥12 sessions @ ≥26.9 tok/s each | **≥1.5×** sessions (crossover not reached at N=12; aggregate 315 tok/s) | The multi-turn row is iSWA-specific and easy to miss: a private slot that has decoded past its prompt cannot partially rewind (upstream SWA-checkpoint semantics, llama.cpp PR #13194), so re-querying it pays a checkpoint restore or a full re-prefill every turn. The pool slot never decodes, so its prefix never slides — every re-attach is free. Note the capacity row is about per-session speed, not just aggregate: 12 pooled sessions each decode faster than 8 private ones. **Cross-architecture results** (RTX 6000 96GB, P≈5073, GEN=256, A/B/A naive/shared/naive): the pool is validated on all three attention architectures, and the memory axis is architecture-independent (~6.8–6.9× at N=8 — it counts cells, not attention math). | axis | Qwen2.5-14B-1M Q8_0 (pure full attention) | Qwen3.6-27B-Omnimerge-v4 Q4_K_M (hybrid GDN + NextN MTP n=3) | |---|---|---| | KV cells (N=8) | **6.78×** (~406 vs ~9 agents on a fixed 8192-cell buffer) | **6.89×** (~304 vs ~9 agents) | | batched decode N=8 | 358 → 403 tok/s (**+12.5%**) | 106 → 123 tok/s (**+15.6%**) | | batched decode N=24 | 412 → 742 tok/s (**+80%**; 17.2 → 30.9 tok/s per session) | 90 → 125 tok/s (**+39.7%**; 3.7 → 5.2 per session) | | shared-only sweep N=32/48/64 | 813 / 871 / 865 tok/s (plateau ~870 near N=48) | 126 / 124 / 124 tok/s (saturates by N≈24–32) | The shared-vs-naive decode gain **grows with N** on both. On hybrid/recurrent models (delta-net, mamba) the *absolute* aggregate saturates much earlier than on pure attention — the recurrent layers batch worse — so there the pool buys **concurrency capacity and memory**, not aggregate throughput past N≈24. **End-to-end agentic A/B — measured** (package-courier harness, `opencoti/langgraph/bench/package_courier/drivers/courier-bs2-ab-polykv.sh`; bs2 RTX 6000 96 GB GPU 0; Gemma-4-26B-A4B-it Q4_K_M + assistant-MTP; q8_0/q8_0 unified KV; 100 packages × 10 chained steps, seed 1, temperature 0). Everything that could compensate is switched off: **8 workers / 8 slots frozen** (`--tps-floor 0`, so `scheduler()` never starts), and **per-slot context held equal at 6144 tok** — `-c 73728/(8+4 pools)` vs `-c 49152/8`, not equal `-c`, which per bug-2287 would have silently handed the naive arm 50% more context per slot. The only difference is the pool. | axis | naive (`--no-polykv`) | shared pool (P=1177) | gain | |---|---|---|---| | wall clock, 100 packages | 1125.5 s | 956.6 s | **−15.0%** | | delivery latency, mean | 86.6 s | 72.7 s | **−16.1%** | | delivery latency, p90 / p99 | 92.7 / 95.1 s | 78.6 / 81.0 s | −15.2% / −14.8% | | per-turn gen tps, p50 | 11.7 | 14.1 | **+21%** | | per-turn gen tps, p90 | 18.3 | 27.5 | **+50%** | | generated tok/s over the run | 49.3 | 58.1 | **+18%** | | prompt tokens processed | 2,708,083 | 2,444,410 | **−9.7%** | | task score | 100/100 | 100/100 | — | Both arms ran exactly 2200 turns and generated within 0.3% of the same token count (55,469 vs 55,625) — the same work, done two ways. **Why 58 tok/s here and 315 tok/s three tables up — they measure different things.** This is the single most common misreading, so state it plainly: `generated tok/s over the run` is a **workload** number, not a decode-speed number. It is answer tokens divided by *wall clock*, and the wall clock of an agentic run is dominated by reading, not writing. Decomposed from this run's own turn events: | | fan-out sweep (315) | single-stream MTP (320) | this A/B (58.1) | |---|---|---|---| | GPU | RTX 3090 | RTX 6000 | RTX 6000 | | sessions decoding | 12, all at once | 1 | **4.14 on average** | | generated per turn | 256 fixed | 256 fixed | **25 tokens** | | prompt : generated | prefill timed separately | n/a | **44 : 1** | | clock includes prefill? | **no** — decode-only rounds | no | **yes — 40% of engine time** | The agentic engine processed **2,444,410 prompt tokens** to produce 55,625 answer tokens, in 956.6 s — i.e. **~2,610 tok/s of total token throughput**, of which only the 58.1 are answer tokens. It spent 3,960 GPU-seconds decoding and 2,667 prefilling (86.6% busy across 8 slots). Per session while decoding it ran 14.0 tok/s, not because decode got slower but because 2.8 other sessions were prefilling into the same GPU at the same time — which the decode-only synthetic rounds deliberately exclude. So all three numbers are true simultaneously: **320 tok/s** is what one session sustains with nothing else running, **315 tok/s** is what twelve concurrent sessions aggregate to in steady-state decode on a 3090, and **58.1 tok/s** is how fast a real agent fleet emits answer tokens while the same GPU also chews through 2.4M tokens of tool output and conversation history. Only compare a number to one measured the same way. Read this table *next to* the synthetic ones, not instead of them. It is the first PolyKV measurement on a real agentic workload — variable-length tool-calling turns, MTP speculation live, and a modest 1177-token shared prefix (system + tool schemas) rather than the 5k synthetic one — so the gain is smaller than the +80% the N=24 fan-out sweep shows. It is also the harder number to argue with: with fan-out frozen, no scheduler behaviour can flatter either arm, and the −9.7% prompt-token count is the shared prefix simply not being prefilled eight times. Expect the gain to *grow* with a longer shared prefix (agent framework + large tool schemas) and with N; expect it to shrink toward zero as the private suffix comes to dominate the prompt. **How to prefill the pool — use the common-prefix token array, not the document text.** Tokenizers merge across the document/suffix boundary (on the Qwen tokenizer the last prefix token fuses with the suffix start), so `tok(DOC)` can be one token longer than the common prefix the agents actually share — and a pool that is even one token longer than `shared_prefix_n_tokens` cannot be shared exactly. The correct client sequence: ```jsonc // 1. tokenize the FULL agent prompts and compute // P = min over agents of commonPrefixLen(tok(DOC), tok(DOC+suffix_i)) // 2. prefill the pool slot with the token array itself (llama.cpp // /completion accepts token arrays) — pool state == P on ANY tokenizer: { "prompt": [/* tok(DOC+suffix_1)[:P] */], "id_slot": 0, "cache_prompt": true, "n_predict": 1 } // 3. agents attach with prompts STRICTLY longer than P: { "prompt": "", "id_slot": 1, "shared_pool_slot": 0, "shared_prefix_n_tokens": P, … } ``` On attention-only models a text prefill happens to work (the ranged cell copy tolerates the extra token); on hybrid/recurrent targets it silently disables every share — see the gotcha below. The token-array prefill is correct everywhere. **Hybrid/recurrent gotcha (GDN / mamba / `qwen35moe`-class models).** A recurrent cache has one rolling state per sequence, not per-position cells, so a pool share is only possible as an **exact full-state** share. The server enforces this: the share engages only when `shared_prefix_n_tokens == pool state length` *and* the request prompt is strictly longer than the shared prefix; anything else logs `poly-kv-pool: hybrid/recurrent target needs exact full-state share … skipping share, full reprocess (bug-2203)` and falls back to a full (correct, slower) reprocess. If you see zero speedup on a hybrid model — or mass HTTP 500s at high N because N unshared full prompt copies overflow the unified KV — grep the server log for that WARN: it almost always means the pool was prefilled with text instead of the token array. (Older builds crashed outright here — `failed to remove sequence N with p0=…` — fixed by patch `0135`.) Sizing note: the pooled prefix pins P cells in **both** iSWA caches (global + SWA) for the pool's lifetime. Budget `-c` for pool prefix + N session windows + generation headroom, or long-running sessions can exhaust slot allocation mid-round. Composes with KV quantization (the pool holds quantized cells) and with session KV-reuse. The pool is read-only for consumers; each agent's divergent suffix is private. **Tiering is pinned per-pool, never per-session.** The K/V tiers — including the mixed-KV recent-window ⊕ compressed-tail pair — are properties of the boot-allocated cache tensors, chosen once at boot (by you or by auto-tier) before any session exists. The window/tail boundary is a per-layer residency budget over the *physical cell axis*, so a prefix cell is in the resident window or evicted (and quantized exactly once, on eviction) for **all** sequences simultaneously. Sharing itself is not copy-on-write: a sharer joins the prefix by adding its sequence bit to the existing cells, and a diverging session just appends private suffix cells — there is no per-session copy that could be re-quantized, and no way for two sessions to see the same prefix at different tiers. The flip side: you cannot give one session a higher-precision read of a shared prefix than another; that would require forking the prefix into a private copy, which is exactly the O(N) memory cost the pool exists to avoid. ### 3.5 PolyKV control-plane API (pools, fork, admission, capacity, tps) Everything in §3.4 manages the pool *by hand* (you own the slot, the token array, the prefix length). The **control-plane API** makes pools first-class server objects with a REST surface, built for agentic orchestrators (LangGraph & co.) that spawn/retire sub-agents at runtime. Enable it at boot: ``` --parallel N --polykv-max-pools M # + --kv-unified for zero-copy shares ``` With `--kv-unified` pool prefixes are shared zero-copy; without it (split cache — required by the rolling-KV window) each attach physically copies the prefix into the session's stream (c6). `--polykv-max-pools M` reserves M pool sequence-ids beyond the slot range (default 0 = off; every route below 404s when off — the feature is fully additive). A reserved pool costs no KV cells until it is materialized. **Pool objects.** A pool is a pinned, immutable token prefix living in the unified KV, addressed by `pool_id`, arranged in a **tree**: children extend (or copy-on-write branch) their parent and share its cells. The pool json: ```jsonc { "pool_id": 1, "parent": 0, "branch_pos": 65, "prefix_len": 82, "prefix_hash": "…", "pinned": false, "children": [], "orphaned_pin": false, "admission": { "mode": "advisory", "target_tps_per_session": 0, … } } ``` **Routes** (all JSON; also see the Python client below): | Route | What it does | |---|---| | `POST /polykv/pools` | Create a root pool. Body: one of `prompt` (text), `tokens` (array), `from_session` / `from_slot` (capture a live session's prefix); plus `pin: true|false`, `ephemeral` (default **true** for `from_session`/`from_slot` — 60 s idle-leaf sweep once attached), `expect_len` (declared prefix length hint). | | `GET /polykv/pools` | List all pools + `pools_max`, `seq_ids_used`, `tree_depth`. | | `GET /polykv/pools/{id}` | One pool, incl. `orphaned_pin` (pinned leaf that nothing references — a compaction leak tell). | | `POST /polykv/pools/{id}/fork` | Child pool. Body: `prompt` or `tokens` = the child's **full prefix** (see contract below), optional `branch_pos: D` for a copy-on-write branch sharing only `[0, D)`, `pin`. | | `POST /polykv/pools/{id}/pin` / `…/unpin` | Pin = survive even with zero attached sessions/children. | | `POST /polykv/pools/{id}/release` | Release; cells reclaim when the subtree refcount hits zero (ancestor cells survive while descendants live). | | `POST /polykv/pools/{id}/sampling` | Per-pool sampling-placement override (see `docs/features/polykv_api.md` §16). | | `POST /polykv/pools/{id}/admission` | Set per-pool admission: `{ "target_tps_per_session": T, "mode": "advisory"\|"enforced", "on_saturation": "reject"\|"warn", "guarantee_min_sessions": G, "settle_tokens": S, "settle_max_ms": M }` (last three = P7, defaults 1 / 48 / 5000). | | `GET /polykv/pools/{id}/capacity[?expected_tokens=N]` | Pre-spawn gate: `can_admit`, `reason`, `headroom_sessions`, `mean_active_tps`, `compaction_pressure` ∈ [0,1], free cells; P7 adds `settling` + `settle_remaining_ms`, `n_warming`, `n_pool_sessions`, `guaranteed`, `projected_mean_tps_model` / `projected_mean_tps_measured` / `drop_per_admit_ewma`, and `projected_idle_estimate` (all sessions between turns ⇒ projection from the last settled mean, < 30 s fresh — "no measurement" is NOT "free capacity"). The `expected_tokens` arm is checked **first**: not enough KV headroom ⇒ `reason: "context headroom exhausted"`. | | `GET /polykv/tps` | Telemetry: aggregate + `n_warming` + per-session `{session_id, tps_ewma, active, ctx_used, ctx_total, ctx_headroom_tokens}`. `?once=1` for a snapshot, default is an SSE stream. | **Admission pacing — P7 settled admission (c5).** Admitting a new agent distorts the very measurement the next admission decides on (the tps EWMA needs ~2 s to absorb it), so `/capacity` runs a **settle window** per pool: after each new-session admit, until that session has decoded `settle_tokens` (48) or `settle_max_ms` (5000) elapsed, the pool reports `settling: true` and withholds `can_admit`. Advisory orchestrators should treat a settling tick as a no-op (neither spawn nor shrink); the enforced gate **holds** the attach briefly instead of 429ing. Sessions still warming up (`tps_ewma` = 0) are excluded from `mean_active_tps` and counted in `n_warming`. The projection improves as the pool runs: each admit's observed mean-tps drop folds into `drop_per_admit_ewma`, and `projected_mean_tps_if_admitted` switches from the `n/(n+1)` model to `mean − drop_ewma` once a sample exists — gate spawns on the projection, not the raw mean. Two escape valves prevent starvation: `guarantee_min_sessions` (default 1) always admits a pool's first N agents even past the floor (only the physical context arm can refuse), so a fresh or nested pool can never deadlock; and `"overcommit": true` on a request explicitly bypasses the enforced gate. Continuing sessions (an existing session→slot affinity) are never gated — only genuinely new sessions count as admissions. **Attaching sessions.** A completion/chat request joins a pool with a single JSON field: ```jsonc { "pool_id": 1, "session_id": "agent-7", … } ``` The server computes the shared prefix P **automatically** — the longest token-exact common prefix between the request and the pool (no `shared_prefix_n_tokens` bookkeeping; that manual field remains for the §3.4 low-level path). Proof of attach: the response `timings.cache_n` ≥ the shared pool prefix on a cold session. Responses carry `X-Session-TPS`; under enforced admission a rejected request gets **HTTP 429 + `Retry-After`**, and admitted ones may carry `X-Sessions-Remaining` backpressure. **The fork/prefix contract (the #1 integration mistake).** `prompt`/`tokens` in fork is the child's **FULL prefix `[0, L)`** — the parent's exact prefix followed by the new suffix — validated token-exact against the parent (a suffix-only prompt 400s with `child prefix shorter than branch_pos`). Text concatenation can shift tokenizer boundaries, so the robust recipe is the token path: ```jsonc // child tokens = parent tokens + suffix tokenized WITHOUT specials POST /tokenize { "content": "", "add_special": false } POST /polykv/pools/{parent}/fork { "tokens": [/* parent ⧺ suffix */] } ``` And if sessions attach through `/v1/chat/completions`, the pool prefix must be the **chat-templated** system block (e.g. `<|im_start|>system\n…<|im_end|>\n` for Qwen) — a raw-text prefix shares zero tokens with templated requests (`cache_n: 0`). **Context exhaustion & compaction (see `docs/features/polykv_api.md` §13 for the full design).** Pool-attached sessions **never context-shift**: a shift rewrites shared cell positions and would corrupt every other reader of the prefix. With `--context-shift` on, the server WARNs at boot and pooled sessions stop gracefully at capacity with `truncated: true` instead of shifting (non-pooled sessions still shift normally). Compaction is a **prompt rewrite, never an in-place KV operation**: watch `compaction_pressure`, then re-root — fork the stable ancestor with `ancestor prefix + summary` as the new full prefix, migrate sessions to the new `pool_id`, unpin + release the old working pool. A pinned pool left behind shows up as `orphaned_pin: true`. **Python / LangGraph.** The `opencoti-langgraph` package (`opencoti/langgraph/` in the repo) wraps all of the above: `PolykvClient`/`AsyncPolykvClient` (incl. `tokenize()` and `compact_by_refork()`), `OpencotiChatModel` (a LangChain `BaseChatModel` with `pool_id`/`session_id` attach, per-response `session_tps`/`cache_n` metadata and 429 → `PoolSaturatedError`), pool-control tools, a capacity-gate node and a CompactionNode. See its README for graph-level usage. ### 3.6 Serving capacity — elastic slots, SWA budget, the admission gate Three c7-era levers turn "how many sessions fit" from a boot-time guess into something measured and enforced at runtime. They matter most on iSWA models (Gemma-4) with `--kv-unified`, where §3.0 showed the SWA cache scaling with the slot count and the base pool being first-come-first-served. **Elastic slots (`--max-parallel N`).** `--parallel` becomes the *initial* live slot count, not the ceiling: N slots are allocated but parked, and the elastic tick admits one at a time while two guards hold — the projected per-slot decode rate stays above `--max-parallel-tps-floor` (default 0 = arm off) and free VRAM stays above `--max-parallel-vram-reserve` (default 512 MiB; an external-pressure guard — growing a live slot allocates nothing). Requires `--kv-unified`. **SWA sequence budget (`--swa-seq-budget B`).** On iSWA models every slot *and* every reserved PolyKV pool pays a full sliding window up front (`n_swa × n_seq_max + n_ubatch` cells, §3.0). This caps the SWA pool at B sequences instead of `n_seq_max` — measured tens of GiB on large-slot Gemma configurations. Sizing **and admission** use the same model, so an over-budget configuration is refused coherently rather than failing later. It can only shrink the pool; inert on non-iSWA models and with `--swa-full`. **The KV admission gate (`--admission-poolless`, default `enforced`).** Without admission, an over-committed pool makes a request pay with an eviction or a failure *after* prefilling. The gate prices **every** completion request (the name is historical — it began with pool-less requests and now covers pooled ones too, after their PolyKV policy gate) against every KV pool the server has — the base pool (under `--kv-unified`) and, on iSWA models, the SWA pool. Demand is **measured, not guessed**: the tokenised prompt plus the declared generation bound (an undeclared bound contributes 0 — this prevents admission-time over-commitment, it does not predict unbounded growth). The reservation is taken **atomically on the server thread** in the same operation that decides it fits, so simultaneous arrivals cannot all be admitted against the same free cells; it is released when the request binds to a slot. - `enforced` (default): refused requests get **HTTP 429 + `Retry-After` before any prefill** — the client-visible contract is a fast, cheap refusal (measured: ~0.4 s vs ~11 s for the prefill-then-fail path it replaces). - `warn`: admit anyway, name the saturated pool in an `X-KV-Pool-Saturated` response header. - `off`: pre-admission behaviour, byte-identical. - Per-request `"overcommit": true` bypasses the enforced gate explicitly (§3.5); pooled sessions additionally get the per-pool settled-admission machinery of §3.5. **Pre-spawn capacity checks.** An orchestrator that spawns agents should ask before spawning, not catch 429s after: `GET /polykv/pools/{id}/capacity?expected_tokens=N` (§3.5) answers `can_admit` from live tps + KV headroom. The 429 path is the backstop for the requests that arrive anyway. --- ## 4. Long context ### 4.1 DCA — Dual Chunk Attention (training-free context extension) Splits attention into intra-chunk / successive / inter-chunk position regimes and merges them exactly by LSE, so a model trained at `n_ctx_train` serves multiples of it **without retraining**. Flags: `--dca on` (default **off**), `--dca-chunk-size N` (default derives from the model's training context; explicit 8192 is the validated recipe), `--dca-yarn-factor F` (default 1.0; measured neutral for retrieval — leave it). Serve beyond the GGUF's declared context with `--override-kv .context_length=int:1048576`. Validated recipe (Gemma-4-A4B, n_ctx_train 256k): ```bash --dca on --dca-chunk-size 8192 -fa on --parallel 1 \ --override-kv gemma4.context_length=int:1048576 ``` Measured retrieval (RULER-VT, n=50): **256k 0.964 · 512k 0.996 · 768k 0.984 · 1M 0.916** — a gentle −7 pp at 4× native, no cliff. Counter-proof on Qwen3-8B (native 41k): plain attention collapses at 128k (PPL 19.2) while DCA holds PPL 7.3. Works on Gemma-4 (its 5 global layers; SWA layers untouched) and Qwen2.5/3/3.5 (all layers). Composes with quantized KV (scalar pairs all pass; q8-DCA decode costs ~2× vs f16-DCA), sparse attention, and rolling-KV. Limitations: DCA caches K un-rope'd → **launch-time toggle only** (a server booted DCA-on can't switch off per request); expect approximation, not identity, past one chunk. On models that are already *native* long-context (e.g. Qwen2.5-1M), DCA can only approximate down — don't use it there. ### 4.2 Sparse attention (Quest block-selector) + sparse-V Two independent decode-bandwidth levers: - **Block-selector** (`--sparse-attn on`): per-block min/max key bounds give an upper bound on each block's attention mass; decode visits only the top-K blocks (+ sinks + recent). Flags: `--sparse-attn-block-size` (128), `--sparse-attn-topk` (default 0 = visit **all** blocks, i.e. no skipping; pass `auto` for adaptive max(64, n_blocks/4), or an explicit block count), `--sparse-attn-recent`, `--sparse-attn-sink` (1), `--sparse-attn-refresh` (8 — re-select every N decode steps), `--sparse-attn-mode` (0). Default **off**. - **Sparse-V**: skips V-dequant for negligible-weight positions inside visited blocks. **Self-configuring**: on iSWA models with quantized V it auto-sets τ=0.05; elsewhere it stays off. Manual override: `TURBO_SPARSE_V_TAU=`. When to use: **long context on quantized KV.** The win grows with context (selectivity 0.91@16k → 0.99@40k and climbing) and lives on quantized KV: q8_0 ⊕ sparse at 50% coverage measured **1.34× decode at niah 100**. Both levers stack (1.31× combined measured). When *not* to use: short contexts or f16 KV on mid-size models — the decode isn't KV-bandwidth-bound there and the selector overhead can make it *slower* than dense. τ values don't transfer across models; retune if you override manually. --- ## 5. Decode speed — MTP speculative decoding Lossless speculative decode; the emitted text is the target model's own (verified). Two flavours, chosen by `--spec-type`: ### 5.1 `--spec-type draft-assistant` (external drafter — Gemma-4) A small `gemma4-assistant` drafter GGUF rides the target's embeddings: ```bash --spec-type draft-assistant --mtp-head gemma4-assistant-A4B-Q8_0.gguf \ -ngld 99 --spec-draft-n-max 2 ``` `--mtp-head` (alias `-md`) names the drafter; **`-ngld 99` matters** (a CPU-resident draft head erases the win). Drafters for A4B/12B/27B/E2B/E4B are published per-size. Setting `mtpHead` in the TS adapter auto-derives the rest. ### 5.2 `--spec-type draft-mtp` (NextN self-spec — Qwen) Qwen 3.5/3.6 GGUFs that embed a NextN/MTP head self-speculate — no second file: ```bash --spec-type draft-mtp --spec-draft-n-max 3 ``` Runs per-slot under `--parallel` (multi-session capable). **Measured (RTX 3090 + upstream-parity campaign):** A4B assistant decode beats upstream llama.cpp b9859 at every depth (+6.6/+12.9/+7.6% at n_max 1/2/3); combined with `turbo3_tcq` KV it reaches **~89 tok/s vs 52.9 plain (+69%)**. Qwen-35B NextN sits at parity with upstream. Recommended depth: `--spec-draft-n-max 2–3` (A4B), `3` (Qwen NextN). Notes/limits: acceptance dips a few pp at depth ≥2 (chained-draft numerics — expected); since c6 assistant-MTP runs on BOTH cache layouts under `--parallel` (split-vs-unified outputs are token-identical — the old `--kv-unified` requirement is retired). Composes with turbo/TCQ KV tiers (its biggest lever), DCA, and quantized KV. Watch live acceptance per slot via `GET /slots` (§7). ### 5.3 Depth policy and sampler placement (auto-selected speculation) Two policies sit on top of speculation and sampling; both shipped defaults were **selected by measurement**, so override them only with your own numbers: - **`--auto-mtp-policy` (default `taper`)** governs draft depth when the speculation type was AUTO-selected (an explicit `--spec-type` always drafts at full depth): `taper` drafts at 1 live slot and switches off from 2 — speculation pays single-stream and costs batched throughput; `allocator` is the measured marginal-value allocator; `always`/`off` are the escape hatches. - **`--sampling-placement` (default `auto`)** picks where the sampler chain runs. `auto` **routes** between device and CPU from a measured crossover (and is MTP-policy-aware: it accounts for whether a slot is drafting); `device`/`cpu` are authoritative when stated. This is distinct from `--backend-sampling` (on by default), which enables/disables the device sampling *path* itself, also per-request. Samplers the device chain cannot express decline to CPU correctly — placement is observable per-slot, not assumed. --- ## 6. Quality — RYS layer duplication `--repeat-layers` re-runs a contiguous block of **middle** layers, weight-shared: zero extra parameter VRAM, no new GGUF, quant-agnostic. You pay in KV cache and tokens/s proportional to the extra effective layers; you buy quality-per-token. ```bash --repeat-layers 33,34 # +1 layer (RYS-S) --repeat-layers 26-34 # +8 layers ([26,34) half-open, RYS-XL) --repeat-layers 8-12;20-24 # disjoint blocks ``` Rules that matter: - **Middle layers only.** Duplicating first/last layers reliably produces incoherent output on merge-fragile models — this is a model property, not an engine bug; the engine prints a boot advisory when a plan touches the boundary band. - Absent flag = identity = byte-identical to stock. - Composes with the full stack: quantized KV, DCA, rolling-KV window/spill, sparse-attn (the residency/DCA/sparse sizing paths are effective-plan-aware), and MTP — where the draft context deliberately runs the un-duplicated base stack while the target keeps RYS (still lossless: the target verifies every drafted token). Wired across all text archs (dense, MoE, Gemma-4 iSWA dual-cache, Qwen 3.5/3.6 recurrent-hybrid); unsupported archs fail loudly at load rather than silently ignoring the plan. **Finding a good plan:** `--rys-probe` enumerates safe-band blocks, scores each by ΔPPL + a task-probe battery, and prints two ready-to-paste templates (most-efficient and max-gain): ```bash sh ./opencoti-llamafile … --rys-probe -m model.gguf -f corpus.txt \ --rys-probe-widths auto --rys-probe-topk 10 ``` Treat its output as a shortlist to verify with your own eval, not a verdict. --- ## 7. Instrumentation — monitor & control API Three planes (full reference: `docs/features/introspection.md`): ### Boot knobs Everything in §§3–6 is a boot flag: set at launch, echoed back at runtime. By design, tier/residency/DCA/retention **cannot** change per request (KV layout would differ). ### Per-request control (JSON body fields) | Field | Default | Effect | |---|---|---| | `session_id` | `""` | Session→slot affinity: the same session returns to the slot holding its KV (prevents cross-session eviction at `--parallel > 1`). Pair with `cache_prompt: true`. | | `shared_pool_slot` | `-1` | Attach this request to SharedKVPool slot N (read-only prefix share, manual path §3.4). | | `shared_prefix_n_tokens` | `0` | Length of the shared prefix (manual path §3.4). | | `pool_id` | unset | Attach to a control-plane pool (§3.5); shared prefix P computed automatically, token-exact. Needs `--polykv-max-pools`. | | `overcommit` | `false` | Skip the enforced admission gate for THIS request (§3.5 P7) — explicit caller-controlled oversubscription past the pool's floor/target. | ### Runtime introspection **`GET /props` → `"opencoti"` object** — boot-state echo plus the *effective* KV state read back from the live cache: ```jsonc "opencoti": { "kv": { "cache_type_k": "q8_0", "cache_type_v": "q4_0", "auto_tier": false, "effective": { "type_k": "q8_0", "type_v": "q4_0", "n_cells": 524288, "n_cells_resident": 524288, "n_layers_spilling": 0, "fully_resident": true, "is_iswa": true } }, "residency": { "kv_residency_mode": 0, "vram_target_mib": 0 }, "dca": { "enabled": true, "chunk_size": 0, "yarn_factor": 1.0 }, "sparse_attn":{ "enabled": false, "block_size": 128, "topk": 0 }, "speculative":{ "types": ["none","draft-assistant"], "n_max": 3 }, "kv_reuse": { "n_parallel": 4, "kv_unified": true, "cache_ram_mib": 8192 }, "rest_kv": { "eviction": false, "recent": 256, "layer": -1 }, "repeat_layers": null } ``` `kv.effective` is the only authoritative record of the auto-tier decision — `configured != effective` is expected when auto-tier engaged. `fully_resident` / `n_layers_spilling` tell you whether rolling-KV is streaming. **`GET /slots` → per-slot `"opencoti"` object** (requires `--slots`): lifetime `draft_n_total` / `draft_n_accepted` / `draft_acceptance` per slot, plus the slot's current `session_id` and pool binding. Operational tell: **sustained draft_acceptance ≳ 0.95 at turn end usually means the model is looping/ruminating** (healthy agentic decode sits ~0.4–0.9) — pollable, no log-scraping. **Per-completion `timings`**: `cache_n` (prefix-reuse hits), `draft_n` / `draft_n_accepted` for that response. **Quick recipes** ```bash curl -s :8080/props | jq .opencoti # what is this server running? curl -s :8080/props | jq .opencoti.kv.effective # did auto-tier/spill engage? curl -s :8080/slots | jq '.[] | {id, acc: .opencoti.draft_acceptance}' ``` For embedders/tools linking the C API: `llama_memory_opencoti_kv_info()` (in `llama.h`) returns the same effective-KV struct. **PolyKV control-plane telemetry** (§3.5, requires `--polykv-max-pools`): `GET /polykv/tps` (SSE or `?once=1`) for per-session tps + context budget; `GET /polykv/pools/{id}/capacity` for admission headroom + `compaction_pressure`; per-slot `n_pool_shared` in `/slots`; PolyKV gauges in `/metrics`. **GPU-share — cooperative peers on one GPU** (c5, `docs/features/gpu_share.md`): multiple opencoti-llamafile processes on the same physical GPU auto-discover each other through a named shared-memory registry (zero-conf, no ports; Linux/macOS/Windows). `--gpu-share-weight W` (ratio, default 1.0) sets this instance's compute share — e.g. weights `2, 1, 1` resolve to 50%/25%/25% — and each instance duty-cycle paces its decode loop toward that share **only while another peer is actively decoding**; a solo or idle-peers instance always runs at full speed. Crashes age out in 3 s. `GET /gpu/peers` lists the live registry (pid, name, weight, resolved `share_pct`, measured `busy_pct`, `active`, `heartbeat_age_ms`); the same snapshot rides in `/props` under `opencoti.gpu_share`. Caveat: the registry is keyed by device description + ordinal, so peers must see the GPU under the same device view (same `CUDA_VISIBLE_DEVICES` ordering). ### Still log-only SharedKVPool share/reject events, retention-eviction discards, rolling-KV tactic selection detail, and the auto-tier WARN line currently appear only in the server log. --- ## 8. Composition matrix | | quant-KV | auto-tier | rolling-KV | PolyKV pool | DCA | sparse-attn | MTP | RYS | |---|---|---|---|---|---|---|---|---| | **quant-KV** | — | K-only honors | ✅ (tiles dequant-on-lift) | ✅ | ✅ | ✅ (the win case) | ✅ (turbo+MTP is the top decode combo) | ✅ | | **auto-tier** | | — | ✅ (it *manages* spill) | ✅ | ✅ (probes in DCA state) | ✅ | ✅ | ✅ (sizing is eff-plan-aware) | | **rolling-KV** | | | — | ✅ | ✅ | ✅ | ✅ | ✅ (validated: window spill × RYS on hybrid) | | **PolyKV pool (SharedKVPool)** | | | | — | ✅ | ✅ | ✅ | ✅ (validated: 2-agent share gate × `--repeat-layers` on A4B; hybrid-GDN omnimerge × NextN MTP full gate, patch `0135`) | | **DCA** | | | | | — | ✅ | ✅ (dual-ctx) | ✅ (eff→src mapped) | | **sparse-attn** | | | | | | — | ✅ | ✅ | | **MTP** | | | | | | | — | ✅ (draft runs base stack; target keeps RYS) | One guard worth restating: SharedKVPool requires `--no-cache-idle-slots` (zero-copy with `--kv-unified`, copy-share on the split cache since c6). Assistant-MTP works on both cache layouts since c6 (the old `--kv-unified` forcing is retired), so MTP + rolling window + pools compose on the split cache. The rolling window by itself also runs under `--kv-unified` (§3.3, validated 2026-09-19 with two concurrent sessions); the three-way combination has only been gated on the split layout. **Reference "agentic serving" launch** (Gemma-4-A4B on a 24 GB card — quantized KV + MTP + introspection): ```bash sh ./opencoti-llamafile---x86_64.llamafile --server --port 8080 \ -m gemma4-A4B-Q4_K_M.gguf -ngl 99 --flash-attn on \ -c 262144 --parallel 4 --kv-unified \ -ctk q8_0 -ctv q4_0 \ --spec-type draft-assistant --mtp-head gemma4-assistant-A4B-Q8_0.gguf \ -ngld 99 --spec-draft-n-max 2 \ --slots ``` --- ## 9. Internal / superseded machinery (so you don't chase ghosts) Present in the patch series but **not** user-facing knobs anymore: - **HeadInfer head-split** (`--headinfer-gpu-heads-frac`): retired as a manual knob; it survives as one tactic inside rolling-KV's auto ladder (`auto` is the only value you should pass, and the adapter does it for you). - **NEO GPU/CPU FA pipelining** (`--neo-pipeline`): structurally shipped, default off; no measurable win on single-GPU consumer hardware. Leave off. - **Fused-MoE up-gate** (`--fused-moe-up-gate`): niche (+2.4% decode on OLMoE-class MoE; Gemma-4 already fuses). Default off. - **Fused-NextN draft graph** (`OPENCOTI_MTP_FUSED_NEXTN=1`): built and shipped (patch 0093) but default off for a measured reason — on CUDA it decodes **−7.0 to −11.8% slower on 4/4 NextN models** than the default autoregressive draft loop (which, post-0128, is at upstream parity or better). Leave off. *Corrected 2026-08-12*: this entry used to blame a non-shape-invariant graph that "rebuilds every cycle". The `OPENCOTI_NEXTN_REUSE_TRACE` counter measures **1.7–1.8% miss** (e.g. hit=1257 miss=23), i.e. the graph is reused ~98% of the time — the default-OFF verdict stands on the throughput measurement, but that mechanism claim was wrong and the real cause is still open. On Vulkan the picture splits (35B **+11.7%**, qwopus27 −3.3%), so it stays off there too pending a fuller grid. See `docs/features/fused_nextn_mtp.md`. - **ScoutAttention, LMCache**: design-only / deferred — the flags don't exist. --- ## 10. Verifying an artifact Moved to [`docs/llamafile-artifacts.md`](llamafile-artifacts.md) (hashes vs `MANIFEST.json`/`SHA256SUMS`, embedded-DSO verification without execution, version-string check, the glibc floor contract). --- ## 11. Migrating between cuts ### From c6 (or earlier) to the current cut — behaviour flips The host binary surface is compatible, but three defaults changed and one path moved. All are visible at boot (`--version`, `--help`, the log) and reversible by flag: - **KV admission is ON by default** — `--admission-poolless` ships `enforced` (§3.6): a request that cannot fit its measured KV demand now gets a fast HTTP 429 + `Retry-After` *before* prefill instead of failing or evicting after it. Clients that never handled 429s should either handle them (the header tells you when to retry) or launch with `--admission-poolless warn` / `off`. - **Sampler placement defaults to `auto` and routes** (§5.3) — earlier cuts defaulted to `device` (and for one cut `auto` only *observed*). Pass `--sampling-placement device` to pin the old behaviour. - **Auto-selected speculation tapers** — `--auto-mtp-policy` ships `taper` (draft at 1 live slot, off from 2). An explicit `--spec-type` is unaffected. Pass `--auto-mtp-policy always` for the old always-draft behaviour. - **The GPU-payload directory is versioned by the full engine string** — `~/.llamafile/v/opencoti--/`, not `~/.llamafile/v//`. Scripts that pre-seed or clean the old path must update (see [llamafile-artifacts.md](llamafile-artifacts.md)). ### To a 0.10.5-based cut (upstream base bump) The first cut on the upstream llamafile 0.10.5 base carries the same opencoti feature set (the whole patch series is ported), with these migration-relevant deltas: - **TurboQuant GGUF type IDs are renumbered.** Upstream 0.10.5 claimed numeric type slot 42, so the turbo family moved (TURBO2_0=43 … TURBO8_0=48). **Any GGUF quantized as `TQ3_1S`/`TQ4_1S` under the old IDs will not load on a 0.10.5-based build — re-quantize it.** KV-cache turbo tiers (`-ctk/-ctv turbo*`) are runtime-only and unaffected; saved session/state files that embed type ids are also invalidated. - **`-fa auto` resolution changed internally** (a fused-ops registry replaced tensor-name parsing). Same flag surface, same tri-state; the boot log line to look for is now `resolve_fused_ops: Flash Attention enabled`. - Everything else is surface-compatible; per-cut specifics live in that release's `RELEASE_NOTES.md`. ## 12. Media engines and their sidecars (c8) From c8 the same server binary also serves images, speech and video. Each engine is opt-in by its model flag and announces itself in `/health` → `features` (gate a client on the feature row, never on a flag or a file): | engine | start it with | routes | feature rows | |---|---|---|---| | images (stable-diffusion.cpp) | `--diffusion-model` (+ `--diffusion-vae`, `--diffusion-llm`, `--diffusion-edit`, `--diffusion-device`) | `/v1/images/generations`, `/v1/images/edits` | `images_generate_v1`, `images_edit_v1` | | speech-to-text | `--stt-model` (+ `--stt-device`, `--stt-threads`) | `/v1/audio/transcriptions`, `/v1/audio/translations` | `audio_transcriptions_v1`, `audio_translations_v1` | | text-to-speech | `--tts-model` (+ `--tts-vocoder` for OuteTTS, `--tts-voices DIR`, `--tts-voice NAME=PATH`, `--tts-device`) | `/v1/audio/speech` | `audio_speech_v1`, `audio_speech_content_format_v1`, `audio_speech_voice_files_v1` (OuteTTS), `audio_speech_audiocpp_v1` (Kokoro / Supertonic / KittenTTS) | | video | `--video-model` (+ `--video-vae`, `--video-t5xxl`, `--video-device`, `--media-job-ttl`) | `POST /v1/videos`, `GET /v1/videos/{id}` (job queue) | `videos_generate_v1` | Model, vocoder and voice files are identified **by their bytes**, never by their name or extension: a blob without an extension loads, a file of the wrong kind is refused at boot with the reason. Design, request shapes and gates: `docs/features/media_engines.md`; every flag: `docs/llamafile-flags.md` "Media engines". ### 12.1 The three sidecars Three capabilities live in separate shared objects that the engine loads at boot, so that their code (and licences) stay out of the engine binary: | sidecar | published name | gives | feature row | |---|---|---|---| | codec | `oc-codec--.` | mp3 speech output, m4a / aac / mp4 transcription input, mp4 (h264 + aac) video | `media_codec_v1` | | audio.cpp | `oc-audiocpp--.` | the Kokoro, Supertonic and KittenTTS speech models | `audio_speech_audiocpp_v1` (when such a model is loaded) | | eSpeak-ng | `oc-espeak--.` | the phonemiser Kokoro and KittenTTS need (Supertonic needs none). One self-contained file: the library with its compiled data inside | `audio_speech_espeak_v1` (when a Kokoro / KittenTTS model is loaded) | **Where they must be.** Next to the engine binary, under the published name. The engine looks, in this order, in: the directory of the executable, the build's payload directory, the files bundled in the binary, and the app directory (`~/.llamafile/v//`). `OPENCOTI_CODEC_SIDECAR=`, `OPENCOTI_AUDIOCPP_SIDECAR=` and `OPENCOTI_ESPEAK_SIDECAR=` name a file directly; `=0` disables that sidecar. Every published build lists the sidecars with their sha256 in `SHA256SUMS`, `BUILD_INFO.md` and the pin (comment rows `#! sidecar `). **Without a sidecar the engine still runs, and says what is missing:** - no codec: an unstated speech format is `wav` (with the codec: `mp3`), an unstated video container is `avi` (with the codec: `mp4`). An **explicit** `response_format: mp3` or `output_format: mp4` is refused with HTTP 501 (`not_supported_error`) — nothing is substituted for a format the client asked for. Transcription accepts wav, mp3 and flac only. - no audio.cpp sidecar: `--tts-model` with a Kokoro / Supertonic / KittenTTS model fails the boot; OuteTTS is unaffected. - no eSpeak-ng sidecar: Kokoro and KittenTTS fall back to a system `libespeak-ng` if the host has one, and fail the boot if it has none. **No system eSpeak-ng is required when `oc-espeak-*` is beside the engine** — that is the shipped configuration. Order: `--tts-espeak-library` / `--tts-espeak-data`, then the sidecar, then the system. `/props` → `media.tts.espeak.source` says which one is in use (`flag`, `sidecar`, `system`). The engine writes the sidecar's data block once to its app directory and audio.cpp unpacks it into its cache (`$XDG_CACHE_HOME` or `~/.cache/audio.cpp/`, `%LOCALAPPDATA%` on Windows): the process needs a writable home. `oc-espeak-*` needs the `oc-audiocpp-*` of the same publication (older ones ignore the path they are given). - from patch `0539` every such message names the exact file and the directories searched, e.g. `the codec sidecar oc-codec-linux-x86_64.so is not loaded: not found in /opt/opencoti/ nor in /root/.llamafile/v/…/ (also tried as oc-codec.so) - put oc-codec-linux-x86_64.so next to the engine binary or set OPENCOTI_CODEC_SIDECAR=`. **Licences.** The codec sidecar links ffmpeg n9.0.2 (LGPL 2.1+; built `--disable-everything` with mp4/mov, mp3, aac, wav and h264 decode only, no GPL and no non-free parts), libmp3lame 3.100 (LGPL) and openh264 v2.6.0 (BSD-2, built from source; the only h264 encoder). Exact sources: `vendors/pin/codec-sidecar.txt`; build configuration: `scripts/codec-sidecar-build.sh`. The eSpeak-ng sidecar is eSpeak-ng 1.52.0, **GPL-3.0-or-later**, built from the unmodified source (`vendors/pin/espeak-sidecar.txt`, `scripts/espeak-sidecar-build.sh`): it is a separate file, loaded at run time by path, never linked into the engine or another sidecar, and published with its licence text (`COPYING.espeak-ng`). A redistributor ships that text and the source reference with it.