opencoti-llamafile / USAGE.md
ManniX-ITA's picture
c7 r2: release notes, SHA256SUMS, USAGE, patch 0253 (the c7 chain is 0001-0244 + 0253), flags availability table
c650fc6 verified
|
Raw History Blame Contribute Delete
62.5 kB

opencoti-llamafile β€” usage guide

How this engine diverges from upstream Mozilla-Ocho llamafile, what the added features are, how each is gated, its knobs and defaults, its limitations, and which features are meant to be used together.

Audience: anyone running the packaged opencoti-llamafile-<ver>-<tag>-<arch>.llamafile artifact as a local inference server. This is the narrative guide; the two companion references are docs/llamafile-flags.md (every flag / env / JSON knob, with defaults, reconciled against the engine source) and docs/llamafile-artifacts.md (downloads, GPU payloads, side-load paths, verification). Deep-dive design docs live in docs/features/, measured evaluations in docs/evaluations/.


Supported / target model families β€” read this first

opencoti-llamafile loads any GGUF that upstream llama.cpp loads β€” that part is inherited unchanged. But the opencoti feature set (KV tiers, rolling-KV, DCA, MTP, sparse attention, RYS) is developed, tuned, and correctness-gated on two model families, in a deliberate primary/secondary split:

Gemma-4 β€” PRIMARY target

Model Kind Notes
Gemma-4 26B-A4B-128e ("A4B") MoE the flagship serving target; MTP-validated with its native gemma4-assistant drafter
Gemma-4 12B / 31B dense full feature validation incl. RYS + DCA
Gemma-4 E2B / E4B elastic (E-series) shared-KV elastic layers supported; MTP drafters available

Gemma-4 is what the engine is for: its unusual head dims (256 and 512), the iSWA sliding/global dual KV cache, and the per-size gemma4-assistant MTP drafters all have dedicated kernels and graph paths here that upstream lacks or handles slowly. --spec-type draft-assistant, D256/D512 FA-VEC + scalar-MMA decode, iSWA-aware rolling-KV/SharedKVPool/DCA wiring are all Gemma-4-first features.

Qwen β€” SECONDARY target / verification family

Model Kind Notes
Qwen3.5 / Qwen3.6 (e.g. 35B-A3B) dense / MoE / hybrid (gated-delta-net) NextN self-spec MTP (--spec-type draft-mtp, no external drafter needed)
Qwen2.5-14B-1M dense, native-1M the long-context/DCA validation vehicle

Qwen is the standard-architecture (head_dim 128) counterweight: every feature that ships is verified on it too, and it carries one feature Gemma doesn't β€” NextN self-speculation (the model's own MTP head drafts; fused multi-step, at/above upstream parity).

Everything else

Other architectures run with upstream behavior and safe fallbacks, but opencoti features are unvalidated there, and some are arch-gated: MTP needs NextN tensors (Qwen-style) or a gemma4-assistant drafter; RYS --repeat-layers supports the qwen2/qwen3(+MoE)/qwen3.5/qwen3next/gemma-4 forward loops; DCA is validated on Gemma-4 and Qwen2.5-1M. Quality gates (KLD, RULER-niah) were run on the two families above β€” re-gate before trusting aggressive KV tiers on anything else.


1. Relationship to upstream llamafile

opencoti-llamafile is upstream llamafile plus an additive patch series (patches/ in the HF repo, vendors/patches/llamafile/ in the git repo; the upstream base version and the exact patch list for a given cut are recorded in its RELEASE_NOTES and every artifact's MANIFEST.json). Three properties are contractual:

  1. Off means off. Every opencoti feature is opt-in behind a flag, env var, or per-request JSON field. With no opencoti flags set, the engine's compute path is byte-identical to upstream β€” this is a regression gate on every patch, not an aspiration.
  2. Lossless by proof, not vibes. Features that touch the forward pass are gated by logit-equivalence / KLD / RULER-retrieval against vanilla, never by "the output looks fine". Speculative decode is verified-lossless (the output is the target model's).
  3. Single file, zero dependencies. The artifact is a Cosmopolitan APE: one file runs on Linux/macOS/Windows/BSD, x86_64 and aarch64. In the full x86_64 artifact the CUDA and Vulkan backends are embedded and self-extract to ~/.llamafile/v/opencoti-<ver>-<tag>/ on first GPU run; the -win variant ships without them (see llamafile-artifacts.md). TCQ codebooks and quantization tables are compiled in. No installer, no downloads.

What upstream gives you is unchanged: the server API (/completion, /v1/chat/completions, /props, /slots, …), GGUF loading, sampling, chat templates. opencoti adds serving-efficiency machinery on top, aimed at multi-session agentic serving on a fixed VRAM budget: more concurrent sessions per card, longer usable context, faster decode.

chmod +x opencoti-llamafile-<ver>-<tag>-x86_64.llamafile
sh ./opencoti-llamafile-<ver>-<tag>-x86_64.llamafile --server --port 8080 \
    -m model.gguf -ngl 99
# --version β†’ opencoti-<ver>-<tag> ; without --server you get the chat CLI

Note (Linux): launch via sh ./file.llamafile if your kernel lacks binfmt_misc APE registration.

1.1 Artifact variants, GPU payloads, side-load

Moved to docs/llamafile-artifacts.md: the full artifact matrix (Linux / Windows / aarch64 / universal, all with compressed GPU payloads), the published side-load DSO set, the versioned ~/.llamafile/v/opencoti-<ver>-<tag>/ runtime path, and byte-level verification. Short version: the host APE is identical in every variant; pick the file whose embedded payloads match your OS, or take the small bare -win file and side-load a DSO.

1.2 Backends β€” CUDA, Vulkan, CPU

One artifact serves three compute backends; selection is --gpu {auto,nvidia,vulkan,amd,apple,disable} (default auto, which probes CUDA first, then Vulkan):

  • CUDA (--gpu nvidia) β€” the primary, fully-validated backend; every opencoti kernel family (turbo/TCQ in-register FA, DCA, sparse attention, MTP verify) has CUDA instances. sm_75β†’120f on x86_64, sbsa Blackwell-class on aarch64.
  • Vulkan (--gpu vulkan, alias vk; payloads embedded since c7) β€” AMD, Intel and NVIDIA through one DSO. The opencoti Vulkan port carries the sparse-attention parity gate, streaming flash-attention, LSE emission and coopmat FA shaders; on NVIDIA it is the fallback when CUDA can't load, on AMD/Intel it is the GPU story. Feature parity is asserted knob-for-knob against CUDA where ported (vk==cuda gates), but CUDA remains the performance reference.
  • CPU β€” always available, no payload needed; iqk flash-attention kernels (--iqk-flash-attn, default auto) accelerate it. Deleting the side-loaded/extracted DSO forces guaranteed CPU serving.

Flash attention is a tri-state (-fa {on,off,auto}, default auto): auto resolves per device at graph-reserve time and prints its decision (resolve_fused_ops: Flash Attention enabled on 0.10.5-based builds). Explicit on/off remain authoritative. The opencoti split-attention paths (head-split, NEO, GPU_STREAM) are resolver-aware β€” non-split layers still resolve FA normally.


2. Feature map β€” what exists and how it's gated

Feature Default Turn on with Class
Session-keyed KV reuse off (per request) session_id JSON field latency
ReST-KV retention eviction off --rest-kv-eviction quality-under-overflow
KV quantization (scalar) f16 -ctk / -ctv capacity
TurboQuant / TCQ KV tiers off -ctk/-ctv turbo* capacity
Auto KV-tier policy off OPENCOTI_KV_AUTO_TIER=1 capacity (policy)
PolyKV pool (SharedKVPool) off (per request) shared_pool_slot JSON field multi-agent capacity
PolyKV control-plane API off --polykv-max-pools N + /polykv/* routes multi-agent orchestration
Rolling-KV window / spill auto (engages only under pressure) --vram-target, --kv-residency-mode capacity
Mixed-KV spilled tail off -ctkt / -ctvt capacity
DCA long-context off --dca on context extension
Sparse attention (block-selector) off --sparse-attn on long-ctx decode speed
Sparse-V auto on iSWA+quant-V, else off TURBO_SPARSE_V_TAU decode speed
MTP speculative decode off --spec-type + drafter decode speed
RYS layer duplication off --repeat-layers quality
RYS probe off --rys-probe tooling
Lazy slot context off --slot-initial-ctx, --slot-shrink-idle-ms embedder memory
GPU-share TDM pacing auto (engages only with β‰₯2 opencoti servers on one GPU) --gpu-share-weight N, GET /gpu/peers multi-server fairness
Elastic parallel slots off --max-parallel N (+ tps floor / VRAM reserve) serving capacity
iSWA SWA-pool seq budget off (= n_seq_max) --swa-seq-budget B serving capacity (VRAM)
KV admission gate on (enforced) --admission-poolless {off,warn,enforced} serving capacity
Sampling placement auto (routes) --sampling-placement {device,cpu,auto} decode speed
Auto-MTP depth policy taper --auto-mtp-policy {allocator,taper,always,off} decode speed
Introspection API always on GET /props, GET /slots observability

Every boot flag also has an env twin (OPENCOTI_LLAMAFILE_<SNAKE_CASE> for adapter-typed fields, LLAMA_ARG_* for llama.cpp-registered ones). The per-flag reference β€” every knob, value set, default, and env name β€” is docs/llamafile-flags.md; this table is the map, not the contract.

2.1 Sharing one GPU between servers (gpu-share, c5/c6)

Multiple opencoti-llamafile servers on the SAME physical GPU discover each other zero-conf (shared-memory peer registry keyed on the PCI bus id, so CUDA_VISIBLE_DEVICES ordering doesn't matter) and pace themselves by deterministic weighted time-division: a repeating period is carved into one contiguous window per active peer, proportional to --gpu-share-weight (default 1). The split is exact by construction and work-conserving β€” an idle peer's time flows to the busy ones, and a single server pays zero overhead (pacing only engages with β‰₯2 active peers). The period adapts for interactive latency (80–400 ms; worst token stall at 2:1 β‰ˆ 80 ms). Measured on a 3090: 2:1 β†’ 1.97:1 at ~solo aggregate; on Windows (5080): 2.00:1. Watch peers live via GET /gpu/peers. Full design + measurements: docs/features/gpu_share.md. Crashed peers age out ≀3 s; no daemon, no MPS, works on Windows.


3. KV capacity stack β€” PolyKV

PolyKV is the umbrella name for this whole stack: the compressed shared KV pool. Concretely it is the KV quantization tiers of Β§3.1 plus the multi-agent SharedKVPool of Β§3.4, stacking with the auto-tier policy (Β§3.2) and the rolling-KV window (Β§3.3), and orchestrated at runtime through the control-plane API of Β§3.5 (pools/fork/admission/capacity/tps REST surface + the opencoti-langgraph package). If you arrived here looking for "PolyKV" from an announcement: Β§3.4 is the shared-prefix pool itself; Β§3.1 is what the pooled cells are made of; Β§3.5 is how an agent framework drives it.

These four features share one goal β€” fit more context / more sessions in fixed VRAM β€” and are designed to stack. Recommended order of adoption: scalar quant β†’ auto-tier β†’ rolling-KV β†’ turbo tiers β†’ SharedKVPool.

3.0 --kv-unified β€” one shared cell pool vs per-slot reservations

This flag decides what -c and --parallel actually mean. Get it wrong and every capacity number in this document is wrong too. The two layouts are not a tuning preference; they are different allocation models.

With --kv-unified (recommended for agentic / multi-session)

n_ctx_seq = n_ctx (src/llama-context.cpp:456) and the cache runs as a single pool of -c cells with one stream (n_stream = 1, src/llama-kv-cache.cpp:701).

  • -c is the whole pool, shared. It is NOT divided by --parallel.
  • Every slot advertises the full -c as its context. The server sets slot.n_ctx = llama_n_ctx_seq(ctx) with no division (tools/server/server-context.cpp:1894). One long session may legally consume the entire pool.
  • --parallel costs no base-KV VRAM. Raising it does not enlarge the cache. (The one exception is iSWA β€” see below.)
  • There is no isolation. Cells are first-come-first-served, so a greedy session can starve its neighbours. This is exactly why the PolyKV admission gate (Β§3.5) exists: with a shared pool, admission control is not optional, it is the isolation mechanism.
  • Pool prefixes are shared zero-copy (Β§3.4).

Provisioning rule of thumb (bug-2287) β€” size the pool so the expected working set fits:

-c  β‰ˆ  tokens_per_session Γ— (--parallel + --polykv-max-pools)

This is a sizing guideline you apply, not something the engine enforces. PolyKV pools occupy reserved seq-ids above the slot range (seq = n_parallel + k), so they need their own share of the pool.

Without --kv-unified (per-slot reservation)

n_ctx_seq = n_ctx / n_seq_max, and the cache splits into n_stream = --parallel independent per-slot caches.

  • -c is divided. Each slot gets a hard-reserved -c / --parallel and can never exceed it, no matter how idle the others are.
  • --parallel directly divides usable context per session.
  • Isolation is free β€” no session can affect another's capacity.
  • Idle slots waste their whole reservation.
  • Pool prefixes are copy-shared, not zero-copy.

Choosing

Want Use
Many sessions of varying, unpredictable length --kv-unified + admission control
Long single sessions that should use all VRAM --kv-unified
Hard per-tenant context guarantees omit --kv-unified
Benchmarks where per-slot capacity must be identical and fixed omit --kv-unified

The iSWA exception (Gemma-4). On sliding-window models the SWA cache is the one allocation that scales with the seq-id count, even under --kv-unified, and it is allocated eagerly in VRAM at boot (src/llama-kv-cache-iswa.cpp:71):

n_seq_max = --parallel + --polykv-max-pools
size_swa  = n_swa Γ— n_seq_max + n_ubatch      (padded to 256)

--polykv-max-pools counts. A pool occupies a reserved seq-id above the slot range, so it costs a full SWA window exactly like a slot does. Budgeting against --parallel alone silently under-counts.

Measured on bs2 GPU0, gemma-4-31B-it-Q6_K, n_swa = 1024, n_ubatch = 512, -ctk/-ctv q8_0 (boot-only probe, 5/5 exact):

--parallel --polykv-max-pools n_seq_max -c base cells SWA cells
64 4 68 417792 417792 70144
16 4 20 417792 417792 20992
8 4 12 417792 417792 12800
8 0 8 417792 417792 8704
8 4 12 131072 131072 12800

Read the two invariants off the table: the base cache equals -c exactly and is identical at n_seq_max 8 and 68, and the SWA cache is unchanged when -c drops 3.2Γ—. They are independent axes.

At 31B's 425 KiB per SWA cell (50 SWA layers Γ— 16 kv-heads Γ— 256 head_dim Γ— K+V at q8_0) the --parallel 64 --polykv-max-pools 4 row is **28.4 GiB of SWA alone** β€” which is what forces the context budget down on large-slot Gemma runs. --parallel 8 --polykv-max-pools 4 is ~5.2 GiB. Qwen-class models have no SWA split and pay nothing here. Plan work to make this budget independent of the slot count: docs/features/elastic_parallel.md.

3.1 KV quantization: scalar types + TurboQuant/TCQ tiers (PolyKV M6)

The KV cache type is set per-tensor-half: -ctk <type> (keys) and -ctv <type> (values), independently β€” asymmetric pairs are first-class (e.g. -ctk q8_0 -ctv q4_0).

Supported types: f16, bf16, q8_0, q6_0, q5_1, q5_0, q4_0 (scalar), turbo2, turbo3, turbo4, turbo8 (TurboQuant, MSE-optimal with Walsh-Hadamard rotation + InnerQ), turbo2_tcq, turbo3_tcq (trellis-coded, Viterbi-encoded).

Which to pick (measured):

  • 8-bit / 4-bit: use q8_0 / q4_0. The native scalar types dominate turbo8/turbo4 at equal width β€” turbo earns nothing there.
  • -ctk q8_0 -ctv q4_0 is the workhorse asymmetric pair: keys keep 8-bit fidelity (attention logits are K-sensitive), values take the compression.
  • 3 bits and below is TurboQuant territory: turbo3 Pareto-beats q4_0 (90 vs 129 MiB KV at equal quality, teacher-forced TV 0.0067 vs 0.0094); turbo2 is the smallest logit-equivalent KV that exists (~2 bit) β€” the 256k-context play. TCQ variants trade encode cost for a further fidelity step at the same width.
  • All shipped tiers pass logit-equivalence gates; decode runs the quantized data in-register in the flash-attention kernel (no f16 materialization) for turbo2/3/4 and TCQ.

Limitations: turbo8 uses a materialize fallback (not fused); at Gemma-4's head_dim 512 only turbo2/turbo3 have fused D=512 instances; prefill on very long prompts uses a hybrid path automatically. Quality validation on Gemma franken-merges must use retrieval (niah), not perplexity.

KVarN (beellama) is NOT a shipped type. It was evaluated against the tiers above in c8 stage 9 (2026-09-10): at 2–4 bits it is 2–13Γ— better on teacher-forced KLD than the matched turbo/tcq/q4_0 tier and ties q6_0/q8_0, but it is a 128-token tile format with its own cache and kernels (not a -ctk type that can be added here), and on an RTX 3090 it decodes β‰ˆ21 % and prefills β‰ˆ27 % slower than f16 at every width (our TCQ tiers decode slower still, plain turbo faster). Decision 2026-09-10: KVarN is being integrated (c8 stage 9b) with a prefill optimization pass to follow; the turbo tiers stay available until that lands and the retirement decision is taken. Numbers and status: docs/evaluations/kvarn_vs_turbo.md.

3.2 Auto KV-tier (OPENCOTI_KV_AUTO_TIER=1)

Boot policy: pick the least-compressing scalar pair that keeps the whole KV resident in the VRAM budget; if even that spills, the T* model decides between "small f16 spill" and "quantize one tier down" by predicted tokens/s drop.

Knobs (env): OPENCOTI_KV_AUTO_TIER=1 (master), OPENCOTI_KV_TSTAR_DROP (target drop, default 20%), OPENCOTI_KV_TSTAR_MAX_SPILL_MIB (default 800), OPENCOTI_KV_AUTO_TIER_TAIL=1 (also auto-pick a q4_0 spilled tail). Explicit -ctv disables auto entirely; explicit -ctk holds K and walks only V. Dense full-attention models only (iSWA models keep f16).

Read back what it decided: GET /props β†’ .opencoti.kv.effective β€” the configured vs effective split exists exactly because auto-tier may override you.

3.3 Rolling-KV window (residency / spill)

"KV doesn't have to fit." Each layer keeps a device-resident window of recent positions; the tail [0 … window_start) lives in pinned host RAM and is streamed through the attention kernel per-tile, merged exactly via online-softmax (LSE). When everything fits, every layer is GPU_RESIDENT and the path is byte-identical to vanilla β€” the feature only engages under memory pressure.

Flags: --vram-target <MiB> (budget cap; 0 = all free VRAM minus reserve), --kv-residency-mode {auto,head,window} (default auto; leave it), -ctkt / -ctvt (distinct, more-compressed types for the spilled tail β€” "mixed-KV": f16 recent window βŠ• q4_0 tail).

Performance model (RTX 3090, PCIe ~6.5 GB/s): spill decode sits at the PCIe floor, t(token) β‰ˆ fixed + tail_bytes / link_bw β€” linear, no cliff. On a fast-link host (RTX 6000, ~50 GB/s) window-mode spill is genuinely usable; on consumer PCIe it's a last resort β€” prefer quantizing (that's what auto-tier automates).

Limitations: while a window is spilled, context-shift and prompt-cache-reuse are guarded off (requests bounded at n_ctx); the compute-buffer reserve for long contexts is measured automatically at boot (two-pass reserve β€” no knob).

Multi-session: since c6 the rolling window works with --parallel N at full speed β€” each session gets its own windowed KV stream, and multi-stream windowed decode scales positively (measured: 2 concurrent windowed sessions aggregate above the solo rate, per-session β‰ˆ 60% of solo on a PCIe-bound 3090).

Fixed in the c7 re-cut (r2, 2026-09-20) β€” check which c7 you have. The FIRST c7 byte set (2026-09-03) aborts as soon as a position window engages (--vram-target / --kv-residency-mode window with a context that overflows it), on both cache layouts: fattn-common.cuh:87: GGML_ASSERT(dst->op == GGML_OP_FLASH_ATTN_EXT). It is a 0.10.5-port regression (bug-3369): the streaming-attention op's plain-FA fallbacks hand the new base a node it now asserts on. Patch 0253 fixes it, and c7 was re-cut with it β€” same tag, same file names, new bytes. If your x86_64 artifact hashes to 3c907bc7… you have the broken set: re-download (the r2 x86_64 sha256 starts 4f4102d6…; full table in the release notes). An r2 binary replaces an already-extracted r1 CUDA DSO by itself β€” no cache clearing. If you cannot update, keep the KV resident on the old bytes: lower -c, or compress it (-ctk/-ctv, Β§3.1).

The window also works under --kv-unified β€” so it composes with the dynamic-slot serving mode (--kv-unified --max-parallel N). Earlier cuts of this guide called the pair "inherently incompatible"; that is not true of the current engine, which never refuses the combination. Validated 2026-09-19 on the development build (Qwen3-8B Q8_0, f16 KV, --parallel 2 --kv-unified --vram-target, two concurrent ~11k-token sessions = 22 240 cells against a 7 424-cell resident window, so ~two thirds of the live cells sat in the host tail): head / middle / tail needles all retrieved in both sessions, and every answer byte-identical to the fully-resident run and to the split-cache run. What differs between the layouts is sizing, not correctness: unified has ONE window over the shared pool (-c cells), split has one window per stream (-c / --parallel cells each). PolyKV pools (Β§3.4) work on both layouts too. Plain (non-windowed) split-KV --parallel serving is also full-speed as of c6 (a 30–60Γ— multi-slot slowdown in earlier cuts was fixed). Quantized spilled tails (-ctk/-ctv q4_0/q8_0 with an active host tail) keep full-bandwidth bulk staging with multiple concurrent sessions too (c6 r2): on a ~50 GB/s PCIe host, two concurrent q4_0-tail sessions aggregate within ~5% of the solo rate (earlier c6 bytes collapsed ~6Γ— in exactly this combination).

3.4 PolyKV SharedKVPool (multi-agent shared prefix)

N agents attending one physical copy of a common prefix (system prompt + tool defs). Per-request JSON, no CLI flag:

{ "shared_pool_slot": 0, "shared_prefix_n_tokens": 4096, … }

Server must run with --no-cache-idle-slots (mandatory β€” the default idle-slot save/clear would evict the pooled prefix). Since c6 pools work on BOTH cache layouts: with --kv-unified the prefix is shared zero-copy (cell set-membership, one physical copy); on the split-stream cache (e.g. rolling-KV window mode) the share is a physical per-stream copy of the prefix β€” same semantics, higher footprint (device use stays window-capped in window mode).

What the pool speeds up β€” measured (Gemma-4-26B-A4B Q4_K_M, RTX 3090, q8/q8 unified KV, Pβ‰ˆ5k-token shared prefix, greedy fixed-length decode, 8 concurrent sessions unless noted):

axis naive (N private copies) shared pool gain
KV cells (N=8) ~8Β·P P + suffixes 6.9Γ— (~306 vs ~9 agents on a fixed buffer)
prefill, 8 sessions joining 22.8 s 4.6 s ~5Γ— (prefix enters KV once per pool)
steady-state batched decode (N=8) ~190 tok/s 217 tok/s +14% (8 queries read one physical prefix β€” L2 reuse, smaller cell span)
multi-turn re-query (N=8) 99 tok/s 225 tok/s 2.3Γ— (see note)
iso-speed capacity 8 sessions @ 24.0 tok/s each β‰₯12 sessions @ β‰₯26.9 tok/s each β‰₯1.5Γ— sessions (crossover not reached at N=12; aggregate 315 tok/s)

The multi-turn row is iSWA-specific and easy to miss: a private slot that has decoded past its prompt cannot partially rewind (upstream SWA-checkpoint semantics, llama.cpp PR #13194), so re-querying it pays a checkpoint restore or a full re-prefill every turn. The pool slot never decodes, so its prefix never slides β€” every re-attach is free. Note the capacity row is about per-session speed, not just aggregate: 12 pooled sessions each decode faster than 8 private ones.

Cross-architecture results (RTX 6000 96GB, Pβ‰ˆ5073, GEN=256, A/B/A naive/shared/naive): the pool is validated on all three attention architectures, and the memory axis is architecture-independent (~6.8–6.9Γ— at N=8 β€” it counts cells, not attention math).

axis Qwen2.5-14B-1M Q8_0 (pure full attention) Qwen3.6-27B-Omnimerge-v4 Q4_K_M (hybrid GDN + NextN MTP n=3)
KV cells (N=8) 6.78Γ— (~406 vs ~9 agents on a fixed 8192-cell buffer) 6.89Γ— (~304 vs ~9 agents)
batched decode N=8 358 β†’ 403 tok/s (+12.5%) 106 β†’ 123 tok/s (+15.6%)
batched decode N=24 412 β†’ 742 tok/s (+80%; 17.2 β†’ 30.9 tok/s per session) 90 β†’ 125 tok/s (+39.7%; 3.7 β†’ 5.2 per session)
shared-only sweep N=32/48/64 813 / 871 / 865 tok/s (plateau ~870 near N=48) 126 / 124 / 124 tok/s (saturates by Nβ‰ˆ24–32)

The shared-vs-naive decode gain grows with N on both. On hybrid/recurrent models (delta-net, mamba) the absolute aggregate saturates much earlier than on pure attention β€” the recurrent layers batch worse β€” so there the pool buys concurrency capacity and memory, not aggregate throughput past Nβ‰ˆ24.

End-to-end agentic A/B β€” measured (package-courier harness, opencoti/langgraph/bench/package_courier/drivers/courier-bs2-ab-polykv.sh; bs2 RTX 6000 96 GB GPU 0; Gemma-4-26B-A4B-it Q4_K_M + assistant-MTP; q8_0/q8_0 unified KV; 100 packages Γ— 10 chained steps, seed 1, temperature 0). Everything that could compensate is switched off: 8 workers / 8 slots frozen (--tps-floor 0, so scheduler() never starts), and per-slot context held equal at 6144 tok β€” -c 73728/(8+4 pools) vs -c 49152/8, not equal -c, which per bug-2287 would have silently handed the naive arm 50% more context per slot. The only difference is the pool.

axis naive (--no-polykv) shared pool (P=1177) gain
wall clock, 100 packages 1125.5 s 956.6 s βˆ’15.0%
delivery latency, mean 86.6 s 72.7 s βˆ’16.1%
delivery latency, p90 / p99 92.7 / 95.1 s 78.6 / 81.0 s βˆ’15.2% / βˆ’14.8%
per-turn gen tps, p50 11.7 14.1 +21%
per-turn gen tps, p90 18.3 27.5 +50%
generated tok/s over the run 49.3 58.1 +18%
prompt tokens processed 2,708,083 2,444,410 βˆ’9.7%
task score 100/100 100/100 β€”

Both arms ran exactly 2200 turns and generated within 0.3% of the same token count (55,469 vs 55,625) β€” the same work, done two ways.

Why 58 tok/s here and 315 tok/s three tables up β€” they measure different things. This is the single most common misreading, so state it plainly: generated tok/s over the run is a workload number, not a decode-speed number. It is answer tokens divided by wall clock, and the wall clock of an agentic run is dominated by reading, not writing. Decomposed from this run's own turn events:

fan-out sweep (315) single-stream MTP (320) this A/B (58.1)
GPU RTX 3090 RTX 6000 RTX 6000
sessions decoding 12, all at once 1 4.14 on average
generated per turn 256 fixed 256 fixed 25 tokens
prompt : generated prefill timed separately n/a 44 : 1
clock includes prefill? no β€” decode-only rounds no yes β€” 40% of engine time

The agentic engine processed 2,444,410 prompt tokens to produce 55,625 answer tokens, in 956.6 s β€” i.e. ~2,610 tok/s of total token throughput, of which only the 58.1 are answer tokens. It spent 3,960 GPU-seconds decoding and 2,667 prefilling (86.6% busy across 8 slots). Per session while decoding it ran 14.0 tok/s, not because decode got slower but because 2.8 other sessions were prefilling into the same GPU at the same time β€” which the decode-only synthetic rounds deliberately exclude.

So all three numbers are true simultaneously: 320 tok/s is what one session sustains with nothing else running, 315 tok/s is what twelve concurrent sessions aggregate to in steady-state decode on a 3090, and 58.1 tok/s is how fast a real agent fleet emits answer tokens while the same GPU also chews through 2.4M tokens of tool output and conversation history. Only compare a number to one measured the same way.

Read this table next to the synthetic ones, not instead of them. It is the first PolyKV measurement on a real agentic workload β€” variable-length tool-calling turns, MTP speculation live, and a modest 1177-token shared prefix (system + tool schemas) rather than the 5k synthetic one β€” so the gain is smaller than the +80% the N=24 fan-out sweep shows. It is also the harder number to argue with: with fan-out frozen, no scheduler behaviour can flatter either arm, and the βˆ’9.7% prompt-token count is the shared prefix simply not being prefilled eight times. Expect the gain to grow with a longer shared prefix (agent framework + large tool schemas) and with N; expect it to shrink toward zero as the private suffix comes to dominate the prompt.

How to prefill the pool β€” use the common-prefix token array, not the document text. Tokenizers merge across the document/suffix boundary (on the Qwen tokenizer the last prefix token fuses with the suffix start), so tok(DOC) can be one token longer than the common prefix the agents actually share β€” and a pool that is even one token longer than shared_prefix_n_tokens cannot be shared exactly. The correct client sequence:

// 1. tokenize the FULL agent prompts and compute
//    P = min over agents of commonPrefixLen(tok(DOC), tok(DOC+suffix_i))
// 2. prefill the pool slot with the token array itself (llama.cpp
//    /completion accepts token arrays) β€” pool state == P on ANY tokenizer:
{ "prompt": [/* tok(DOC+suffix_1)[:P] */], "id_slot": 0,
  "cache_prompt": true, "n_predict": 1 }
// 3. agents attach with prompts STRICTLY longer than P:
{ "prompt": "<DOC + private suffix>", "id_slot": 1,
  "shared_pool_slot": 0, "shared_prefix_n_tokens": P, … }

On attention-only models a text prefill happens to work (the ranged cell copy tolerates the extra token); on hybrid/recurrent targets it silently disables every share β€” see the gotcha below. The token-array prefill is correct everywhere.

Hybrid/recurrent gotcha (GDN / mamba / qwen35moe-class models). A recurrent cache has one rolling state per sequence, not per-position cells, so a pool share is only possible as an exact full-state share. The server enforces this: the share engages only when shared_prefix_n_tokens == pool state length and the request prompt is strictly longer than the shared prefix; anything else logs poly-kv-pool: hybrid/recurrent target needs exact full-state share … skipping share, full reprocess (bug-2203) and falls back to a full (correct, slower) reprocess. If you see zero speedup on a hybrid model β€” or mass HTTP 500s at high N because N unshared full prompt copies overflow the unified KV β€” grep the server log for that WARN: it almost always means the pool was prefilled with text instead of the token array. (Older builds crashed outright here β€” failed to remove sequence N with p0=… β€” fixed by patch 0135.)

Sizing note: the pooled prefix pins P cells in both iSWA caches (global + SWA) for the pool's lifetime. Budget -c for pool prefix

  • N session windows + generation headroom, or long-running sessions can exhaust slot allocation mid-round.

Composes with KV quantization (the pool holds quantized cells) and with session KV-reuse. The pool is read-only for consumers; each agent's divergent suffix is private.

Tiering is pinned per-pool, never per-session. The K/V tiers β€” including the mixed-KV recent-window βŠ• compressed-tail pair β€” are properties of the boot-allocated cache tensors, chosen once at boot (by you or by auto-tier) before any session exists. The window/tail boundary is a per-layer residency budget over the physical cell axis, so a prefix cell is in the resident window or evicted (and quantized exactly once, on eviction) for all sequences simultaneously. Sharing itself is not copy-on-write: a sharer joins the prefix by adding its sequence bit to the existing cells, and a diverging session just appends private suffix cells β€” there is no per-session copy that could be re-quantized, and no way for two sessions to see the same prefix at different tiers. The flip side: you cannot give one session a higher-precision read of a shared prefix than another; that would require forking the prefix into a private copy, which is exactly the O(N) memory cost the pool exists to avoid.

3.5 PolyKV control-plane API (pools, fork, admission, capacity, tps)

Everything in Β§3.4 manages the pool by hand (you own the slot, the token array, the prefix length). The control-plane API makes pools first-class server objects with a REST surface, built for agentic orchestrators (LangGraph & co.) that spawn/retire sub-agents at runtime. Enable it at boot:

--parallel N --polykv-max-pools M          # + --kv-unified for zero-copy shares

With --kv-unified pool prefixes are shared zero-copy; without it (split cache β€” required by the rolling-KV window) each attach physically copies the prefix into the session's stream (c6).

--polykv-max-pools M reserves M pool sequence-ids beyond the slot range (default 0 = off; every route below 404s when off β€” the feature is fully additive). A reserved pool costs no KV cells until it is materialized.

Pool objects. A pool is a pinned, immutable token prefix living in the unified KV, addressed by pool_id, arranged in a tree: children extend (or copy-on-write branch) their parent and share its cells. The pool json:

{ "pool_id": 1, "parent": 0, "branch_pos": 65, "prefix_len": 82,
  "prefix_hash": "…", "pinned": false, "children": [],
  "orphaned_pin": false,
  "admission": { "mode": "advisory", "target_tps_per_session": 0, … } }

Routes (all JSON; also see the Python client below):

Route What it does
POST /polykv/pools Create a root pool. Body: one of prompt (text), tokens (array), from_session / from_slot (capture a live session's prefix); plus `pin: true
GET /polykv/pools List all pools + pools_max, seq_ids_used, tree_depth.
GET /polykv/pools/{id} One pool, incl. orphaned_pin (pinned leaf that nothing references β€” a compaction leak tell).
POST /polykv/pools/{id}/fork Child pool. Body: prompt or tokens = the child's full prefix (see contract below), optional branch_pos: D for a copy-on-write branch sharing only [0, D), pin.
POST /polykv/pools/{id}/pin / …/unpin Pin = survive even with zero attached sessions/children.
POST /polykv/pools/{id}/release Release; cells reclaim when the subtree refcount hits zero (ancestor cells survive while descendants live).
POST /polykv/pools/{id}/sampling Per-pool sampling-placement override (see docs/features/polykv_api.md Β§16).
POST /polykv/pools/{id}/admission Set per-pool admission: { "target_tps_per_session": T, "mode": "advisory"|"enforced", "on_saturation": "reject"|"warn", "guarantee_min_sessions": G, "settle_tokens": S, "settle_max_ms": M } (last three = P7, defaults 1 / 48 / 5000).
GET /polykv/pools/{id}/capacity[?expected_tokens=N] Pre-spawn gate: can_admit, reason, headroom_sessions, mean_active_tps, compaction_pressure ∈ [0,1], free cells; P7 adds settling + settle_remaining_ms, n_warming, n_pool_sessions, guaranteed, projected_mean_tps_model / projected_mean_tps_measured / drop_per_admit_ewma, and projected_idle_estimate (all sessions between turns β‡’ projection from the last settled mean, < 30 s fresh β€” "no measurement" is NOT "free capacity"). The expected_tokens arm is checked first: not enough KV headroom β‡’ reason: "context headroom exhausted".
GET /polykv/tps Telemetry: aggregate + n_warming + per-session {session_id, tps_ewma, active, ctx_used, ctx_total, ctx_headroom_tokens}. ?once=1 for a snapshot, default is an SSE stream.

Admission pacing β€” P7 settled admission (c5). Admitting a new agent distorts the very measurement the next admission decides on (the tps EWMA needs ~2 s to absorb it), so /capacity runs a settle window per pool: after each new-session admit, until that session has decoded settle_tokens (48) or settle_max_ms (5000) elapsed, the pool reports settling: true and withholds can_admit. Advisory orchestrators should treat a settling tick as a no-op (neither spawn nor shrink); the enforced gate holds the attach briefly instead of 429ing. Sessions still warming up (tps_ewma = 0) are excluded from mean_active_tps and counted in n_warming. The projection improves as the pool runs: each admit's observed mean-tps drop folds into drop_per_admit_ewma, and projected_mean_tps_if_admitted switches from the n/(n+1) model to mean βˆ’ drop_ewma once a sample exists β€” gate spawns on the projection, not the raw mean. Two escape valves prevent starvation: guarantee_min_sessions (default 1) always admits a pool's first N agents even past the floor (only the physical context arm can refuse), so a fresh or nested pool can never deadlock; and "overcommit": true on a request explicitly bypasses the enforced gate. Continuing sessions (an existing sessionβ†’slot affinity) are never gated β€” only genuinely new sessions count as admissions.

Attaching sessions. A completion/chat request joins a pool with a single JSON field:

{ "pool_id": 1, "session_id": "agent-7", … }

The server computes the shared prefix P automatically β€” the longest token-exact common prefix between the request and the pool (no shared_prefix_n_tokens bookkeeping; that manual field remains for the Β§3.4 low-level path). Proof of attach: the response timings.cache_n β‰₯ the shared pool prefix on a cold session. Responses carry X-Session-TPS; under enforced admission a rejected request gets HTTP 429 + Retry-After, and admitted ones may carry X-Sessions-Remaining backpressure.

The fork/prefix contract (the #1 integration mistake). prompt/tokens in fork is the child's FULL prefix [0, L) β€” the parent's exact prefix followed by the new suffix β€” validated token-exact against the parent (a suffix-only prompt 400s with child prefix shorter than branch_pos). Text concatenation can shift tokenizer boundaries, so the robust recipe is the token path:

// child tokens = parent tokens + suffix tokenized WITHOUT specials
POST /tokenize { "content": "<suffix>", "add_special": false }
POST /polykv/pools/{parent}/fork { "tokens": [/* parent β§Ί suffix */] }

And if sessions attach through /v1/chat/completions, the pool prefix must be the chat-templated system block (e.g. <|im_start|>system\n…<|im_end|>\n for Qwen) β€” a raw-text prefix shares zero tokens with templated requests (cache_n: 0).

Context exhaustion & compaction (see docs/features/polykv_api.md Β§13 for the full design). Pool-attached sessions never context-shift: a shift rewrites shared cell positions and would corrupt every other reader of the prefix. With --context-shift on, the server WARNs at boot and pooled sessions stop gracefully at capacity with truncated: true instead of shifting (non-pooled sessions still shift normally). Compaction is a prompt rewrite, never an in-place KV operation: watch compaction_pressure, then re-root β€” fork the stable ancestor with ancestor prefix + summary as the new full prefix, migrate sessions to the new pool_id, unpin + release the old working pool. A pinned pool left behind shows up as orphaned_pin: true.

Python / LangGraph. The opencoti-langgraph package (opencoti/langgraph/ in the repo) wraps all of the above: PolykvClient/AsyncPolykvClient (incl. tokenize() and compact_by_refork()), OpencotiChatModel (a LangChain BaseChatModel with pool_id/session_id attach, per-response session_tps/cache_n metadata and 429 β†’ PoolSaturatedError), pool-control tools, a capacity-gate node and a CompactionNode. See its README for graph-level usage.

3.6 Serving capacity β€” elastic slots, SWA budget, the admission gate

Three c7-era levers turn "how many sessions fit" from a boot-time guess into something measured and enforced at runtime. They matter most on iSWA models (Gemma-4) with --kv-unified, where Β§3.0 showed the SWA cache scaling with the slot count and the base pool being first-come-first-served.

Elastic slots (--max-parallel N). --parallel becomes the initial live slot count, not the ceiling: N slots are allocated but parked, and the elastic tick admits one at a time while two guards hold β€” the projected per-slot decode rate stays above --max-parallel-tps-floor (default 0 = arm off) and free VRAM stays above --max-parallel-vram-reserve (default 512 MiB; an external-pressure guard β€” growing a live slot allocates nothing). Requires --kv-unified.

SWA sequence budget (--swa-seq-budget B). On iSWA models every slot and every reserved PolyKV pool pays a full sliding window up front (n_swa Γ— n_seq_max + n_ubatch cells, Β§3.0). This caps the SWA pool at B sequences instead of n_seq_max β€” measured tens of GiB on large-slot Gemma configurations. Sizing and admission use the same model, so an over-budget configuration is refused coherently rather than failing later. It can only shrink the pool; inert on non-iSWA models and with --swa-full.

The KV admission gate (--admission-poolless, default enforced). Without admission, an over-committed pool makes a request pay with an eviction or a failure after prefilling. The gate prices every completion request (the name is historical β€” it began with pool-less requests and now covers pooled ones too, after their PolyKV policy gate) against every KV pool the server has β€” the base pool (under --kv-unified) and, on iSWA models, the SWA pool. Demand is measured, not guessed: the tokenised prompt plus the declared generation bound (an undeclared bound contributes 0 β€” this prevents admission-time over-commitment, it does not predict unbounded growth). The reservation is taken atomically on the server thread in the same operation that decides it fits, so simultaneous arrivals cannot all be admitted against the same free cells; it is released when the request binds to a slot.

  • enforced (default): refused requests get HTTP 429 + Retry-After before any prefill β€” the client-visible contract is a fast, cheap refusal (measured: ~0.4 s vs ~11 s for the prefill-then-fail path it replaces).
  • warn: admit anyway, name the saturated pool in an X-KV-Pool-Saturated response header.
  • off: pre-admission behaviour, byte-identical.
  • Per-request "overcommit": true bypasses the enforced gate explicitly (Β§3.5); pooled sessions additionally get the per-pool settled-admission machinery of Β§3.5.

Pre-spawn capacity checks. An orchestrator that spawns agents should ask before spawning, not catch 429s after: GET /polykv/pools/{id}/capacity?expected_tokens=N (Β§3.5) answers can_admit from live tps + KV headroom. The 429 path is the backstop for the requests that arrive anyway.


4. Long context

4.1 DCA β€” Dual Chunk Attention (training-free context extension)

Splits attention into intra-chunk / successive / inter-chunk position regimes and merges them exactly by LSE, so a model trained at n_ctx_train serves multiples of it without retraining.

Flags: --dca on (default off), --dca-chunk-size N (default derives from the model's training context; explicit 8192 is the validated recipe), --dca-yarn-factor F (default 1.0; measured neutral for retrieval β€” leave it). Serve beyond the GGUF's declared context with --override-kv <arch>.context_length=int:1048576.

Validated recipe (Gemma-4-A4B, n_ctx_train 256k):

--dca on --dca-chunk-size 8192 -fa on --parallel 1 \
--override-kv gemma4.context_length=int:1048576

Measured retrieval (RULER-VT, n=50): 256k 0.964 Β· 512k 0.996 Β· 768k 0.984 Β· 1M 0.916 β€” a gentle βˆ’7 pp at 4Γ— native, no cliff. Counter-proof on Qwen3-8B (native 41k): plain attention collapses at 128k (PPL 19.2) while DCA holds PPL 7.3.

Works on Gemma-4 (its 5 global layers; SWA layers untouched) and Qwen2.5/3/3.5 (all layers). Composes with quantized KV (scalar pairs all pass; q8-DCA decode costs ~2Γ— vs f16-DCA), sparse attention, and rolling-KV.

Limitations: DCA caches K un-rope'd β†’ launch-time toggle only (a server booted DCA-on can't switch off per request); expect approximation, not identity, past one chunk. On models that are already native long-context (e.g. Qwen2.5-1M), DCA can only approximate down β€” don't use it there.

4.2 Sparse attention (Quest block-selector) + sparse-V

Two independent decode-bandwidth levers:

  • Block-selector (--sparse-attn on): per-block min/max key bounds give an upper bound on each block's attention mass; decode visits only the top-K blocks (+ sinks + recent). Flags: --sparse-attn-block-size (128), --sparse-attn-topk (default 0 = visit all blocks, i.e. no skipping; pass auto for adaptive max(64, n_blocks/4), or an explicit block count), --sparse-attn-recent, --sparse-attn-sink (1), --sparse-attn-refresh (8 β€” re-select every N decode steps), --sparse-attn-mode (0). Default off.
  • Sparse-V: skips V-dequant for negligible-weight positions inside visited blocks. Self-configuring: on iSWA models with quantized V it auto-sets Ο„=0.05; elsewhere it stays off. Manual override: TURBO_SPARSE_V_TAU=<float>.

When to use: long context on quantized KV. The win grows with context (selectivity 0.91@16k β†’ 0.99@40k and climbing) and lives on quantized KV: q8_0 βŠ• sparse at 50% coverage measured 1.34Γ— decode at niah 100. Both levers stack (1.31Γ— combined measured).

When not to use: short contexts or f16 KV on mid-size models β€” the decode isn't KV-bandwidth-bound there and the selector overhead can make it slower than dense. Ο„ values don't transfer across models; retune if you override manually.


5. Decode speed β€” MTP speculative decoding

Lossless speculative decode; the emitted text is the target model's own (verified). Two flavours, chosen by --spec-type:

5.1 --spec-type draft-assistant (external drafter β€” Gemma-4)

A small gemma4-assistant drafter GGUF rides the target's embeddings:

--spec-type draft-assistant --mtp-head gemma4-assistant-A4B-Q8_0.gguf \
-ngld 99 --spec-draft-n-max 2

--mtp-head (alias -md) names the drafter; -ngld 99 matters (a CPU-resident draft head erases the win). Drafters for A4B/12B/27B/E2B/E4B are published per-size. Setting mtpHead in the TS adapter auto-derives the rest.

5.2 --spec-type draft-mtp (NextN self-spec β€” Qwen)

Qwen 3.5/3.6 GGUFs that embed a NextN/MTP head self-speculate β€” no second file:

--spec-type draft-mtp --spec-draft-n-max 3

Runs per-slot under --parallel (multi-session capable).

Measured (RTX 3090 + upstream-parity campaign): A4B assistant decode beats upstream llama.cpp b9859 at every depth (+6.6/+12.9/+7.6% at n_max 1/2/3); combined with turbo3_tcq KV it reaches ~89 tok/s vs 52.9 plain (+69%). Qwen-35B NextN sits at parity with upstream. Recommended depth: --spec-draft-n-max 2–3 (A4B), 3 (Qwen NextN).

Notes/limits: acceptance dips a few pp at depth β‰₯2 (chained-draft numerics β€” expected); since c6 assistant-MTP runs on BOTH cache layouts under --parallel (split-vs-unified outputs are token-identical β€” the old --kv-unified requirement is retired). Composes with turbo/TCQ KV tiers (its biggest lever), DCA, and quantized KV. Watch live acceptance per slot via GET /slots (Β§7).

5.3 Depth policy and sampler placement (auto-selected speculation)

Two policies sit on top of speculation and sampling; both shipped defaults were selected by measurement, so override them only with your own numbers:

  • --auto-mtp-policy (default taper) governs draft depth when the speculation type was AUTO-selected (an explicit --spec-type always drafts at full depth): taper drafts at 1 live slot and switches off from 2 β€” speculation pays single-stream and costs batched throughput; allocator is the measured marginal-value allocator; always/off are the escape hatches.
  • --sampling-placement (default auto) picks where the sampler chain runs. auto routes between device and CPU from a measured crossover (and is MTP-policy-aware: it accounts for whether a slot is drafting); device/cpu are authoritative when stated. This is distinct from --backend-sampling (on by default), which enables/disables the device sampling path itself, also per-request. Samplers the device chain cannot express decline to CPU correctly β€” placement is observable per-slot, not assumed.

6. Quality β€” RYS layer duplication

--repeat-layers re-runs a contiguous block of middle layers, weight-shared: zero extra parameter VRAM, no new GGUF, quant-agnostic. You pay in KV cache and tokens/s proportional to the extra effective layers; you buy quality-per-token.

--repeat-layers 33,34      # +1 layer  (RYS-S)
--repeat-layers 26-34      # +8 layers ([26,34) half-open, RYS-XL)
--repeat-layers 8-12;20-24 # disjoint blocks

Rules that matter:

  • Middle layers only. Duplicating first/last layers reliably produces incoherent output on merge-fragile models β€” this is a model property, not an engine bug; the engine prints a boot advisory when a plan touches the boundary band.
  • Absent flag = identity = byte-identical to stock.
  • Composes with the full stack: quantized KV, DCA, rolling-KV window/spill, sparse-attn (the residency/DCA/sparse sizing paths are effective-plan-aware), and MTP β€” where the draft context deliberately runs the un-duplicated base stack while the target keeps RYS (still lossless: the target verifies every drafted token). Wired across all text archs (dense, MoE, Gemma-4 iSWA dual-cache, Qwen 3.5/3.6 recurrent-hybrid); unsupported archs fail loudly at load rather than silently ignoring the plan.

Finding a good plan: --rys-probe enumerates safe-band blocks, scores each by Ξ”PPL + a task-probe battery, and prints two ready-to-paste templates (most-efficient and max-gain):

sh ./opencoti-llamafile … --rys-probe -m model.gguf -f corpus.txt \
   --rys-probe-widths auto --rys-probe-topk 10

Treat its output as a shortlist to verify with your own eval, not a verdict.


7. Instrumentation β€” monitor & control API

Three planes (full reference: docs/features/introspection.md):

Boot knobs

Everything in Β§Β§3–6 is a boot flag: set at launch, echoed back at runtime. By design, tier/residency/DCA/retention cannot change per request (KV layout would differ).

Per-request control (JSON body fields)

Field Default Effect
session_id "" Session→slot affinity: the same session returns to the slot holding its KV (prevents cross-session eviction at --parallel > 1). Pair with cache_prompt: true.
shared_pool_slot -1 Attach this request to SharedKVPool slot N (read-only prefix share, manual path Β§3.4).
shared_prefix_n_tokens 0 Length of the shared prefix (manual path Β§3.4).
pool_id unset Attach to a control-plane pool (Β§3.5); shared prefix P computed automatically, token-exact. Needs --polykv-max-pools.
overcommit false Skip the enforced admission gate for THIS request (Β§3.5 P7) β€” explicit caller-controlled oversubscription past the pool's floor/target.

Runtime introspection

GET /props β†’ "opencoti" object β€” boot-state echo plus the effective KV state read back from the live cache:

"opencoti": {
  "kv": { "cache_type_k": "q8_0", "cache_type_v": "q4_0",
          "auto_tier": false,
          "effective": { "type_k": "q8_0", "type_v": "q4_0",
                         "n_cells": 524288, "n_cells_resident": 524288,
                         "n_layers_spilling": 0, "fully_resident": true,
                         "is_iswa": true } },
  "residency":  { "kv_residency_mode": 0, "vram_target_mib": 0 },
  "dca":        { "enabled": true, "chunk_size": 0, "yarn_factor": 1.0 },
  "sparse_attn":{ "enabled": false, "block_size": 128, "topk": 0 },
  "speculative":{ "types": ["none","draft-assistant"], "n_max": 3 },
  "kv_reuse":   { "n_parallel": 4, "kv_unified": true, "cache_ram_mib": 8192 },
  "rest_kv":    { "eviction": false, "recent": 256, "layer": -1 },
  "repeat_layers": null
}

kv.effective is the only authoritative record of the auto-tier decision β€” configured != effective is expected when auto-tier engaged. fully_resident / n_layers_spilling tell you whether rolling-KV is streaming.

GET /slots β†’ per-slot "opencoti" object (requires --slots): lifetime draft_n_total / draft_n_accepted / draft_acceptance per slot, plus the slot's current session_id and pool binding. Operational tell: sustained draft_acceptance ≳ 0.95 at turn end usually means the model is looping/ruminating (healthy agentic decode sits ~0.4–0.9) β€” pollable, no log-scraping.

Per-completion timings: cache_n (prefix-reuse hits), draft_n / draft_n_accepted for that response.

Quick recipes

curl -s :8080/props | jq .opencoti                      # what is this server running?
curl -s :8080/props | jq .opencoti.kv.effective         # did auto-tier/spill engage?
curl -s :8080/slots | jq '.[] | {id, acc: .opencoti.draft_acceptance}'

For embedders/tools linking the C API: llama_memory_opencoti_kv_info() (in llama.h) returns the same effective-KV struct.

PolyKV control-plane telemetry (Β§3.5, requires --polykv-max-pools): GET /polykv/tps (SSE or ?once=1) for per-session tps + context budget; GET /polykv/pools/{id}/capacity for admission headroom + compaction_pressure; per-slot n_pool_shared in /slots; PolyKV gauges in /metrics.

GPU-share β€” cooperative peers on one GPU (c5, docs/features/gpu_share.md): multiple opencoti-llamafile processes on the same physical GPU auto-discover each other through a named shared-memory registry (zero-conf, no ports; Linux/macOS/Windows). --gpu-share-weight W (ratio, default 1.0) sets this instance's compute share β€” e.g. weights 2, 1, 1 resolve to 50%/25%/25% β€” and each instance duty-cycle paces its decode loop toward that share only while another peer is actively decoding; a solo or idle-peers instance always runs at full speed. Crashes age out in 3 s. GET /gpu/peers lists the live registry (pid, name, weight, resolved share_pct, measured busy_pct, active, heartbeat_age_ms); the same snapshot rides in /props under opencoti.gpu_share. Caveat: the registry is keyed by device description + ordinal, so peers must see the GPU under the same device view (same CUDA_VISIBLE_DEVICES ordering).

Still log-only

SharedKVPool share/reject events, retention-eviction discards, rolling-KV tactic selection detail, and the auto-tier WARN line currently appear only in the server log.


8. Composition matrix

quant-KV auto-tier rolling-KV PolyKV pool DCA sparse-attn MTP RYS
quant-KV β€” K-only honors βœ… (tiles dequant-on-lift) βœ… βœ… βœ… (the win case) βœ… (turbo+MTP is the top decode combo) βœ…
auto-tier β€” βœ… (it manages spill) βœ… βœ… (probes in DCA state) βœ… βœ… βœ… (sizing is eff-plan-aware)
rolling-KV β€” βœ… βœ… βœ… βœ… βœ… (validated: window spill Γ— RYS on hybrid)
PolyKV pool (SharedKVPool) β€” βœ… βœ… βœ… βœ… (validated: 2-agent share gate Γ— --repeat-layers on A4B; hybrid-GDN omnimerge Γ— NextN MTP full gate, patch 0135)
DCA β€” βœ… βœ… (dual-ctx) βœ… (effβ†’src mapped)
sparse-attn β€” βœ… βœ…
MTP β€” βœ… (draft runs base stack; target keeps RYS)

One guard worth restating: SharedKVPool requires --no-cache-idle-slots (zero-copy with --kv-unified, copy-share on the split cache since c6). Assistant-MTP works on both cache layouts since c6 (the old --kv-unified forcing is retired), so MTP + rolling window + pools compose on the split cache. The rolling window by itself also runs under --kv-unified (Β§3.3, validated 2026-09-19 with two concurrent sessions); the three-way combination has only been gated on the split layout.

Reference "agentic serving" launch (Gemma-4-A4B on a 24 GB card β€” quantized KV + MTP + introspection):

sh ./opencoti-llamafile-<ver>-<tag>-x86_64.llamafile --server --port 8080 \
  -m gemma4-A4B-Q4_K_M.gguf -ngl 99 --flash-attn on \
  -c 262144 --parallel 4 --kv-unified \
  -ctk q8_0 -ctv q4_0 \
  --spec-type draft-assistant --mtp-head gemma4-assistant-A4B-Q8_0.gguf \
  -ngld 99 --spec-draft-n-max 2 \
  --slots

9. Internal / superseded machinery (so you don't chase ghosts)

Present in the patch series but not user-facing knobs anymore:

  • HeadInfer head-split (--headinfer-gpu-heads-frac): retired as a manual knob; it survives as one tactic inside rolling-KV's auto ladder (auto is the only value you should pass, and the adapter does it for you).
  • NEO GPU/CPU FA pipelining (--neo-pipeline): structurally shipped, default off; no measurable win on single-GPU consumer hardware. Leave off.
  • Fused-MoE up-gate (--fused-moe-up-gate): niche (+2.4% decode on OLMoE-class MoE; Gemma-4 already fuses). Default off.
  • Fused-NextN draft graph (OPENCOTI_MTP_FUSED_NEXTN=1): built and shipped (patch 0093) but default off for a measured reason β€” on CUDA it decodes βˆ’7.0 to βˆ’11.8% slower on 4/4 NextN models than the default autoregressive draft loop (which, post-0128, is at upstream parity or better). Leave off. Corrected 2026-08-12: this entry used to blame a non-shape-invariant graph that "rebuilds every cycle". The OPENCOTI_NEXTN_REUSE_TRACE counter measures 1.7–1.8% miss (e.g. hit=1257 miss=23), i.e. the graph is reused ~98% of the time β€” the default-OFF verdict stands on the throughput measurement, but that mechanism claim was wrong and the real cause is still open. On Vulkan the picture splits (35B +11.7%, qwopus27 βˆ’3.3%), so it stays off there too pending a fuller grid. See docs/features/fused_nextn_mtp.md.
  • ScoutAttention, LMCache: design-only / deferred β€” the flags don't exist.

10. Verifying an artifact

Moved to docs/llamafile-artifacts.md (hashes vs MANIFEST.json/SHA256SUMS, embedded-DSO verification without execution, version-string check, the glibc floor contract).


11. Migrating between cuts

From c6 (or earlier) to the current cut β€” behaviour flips

The host binary surface is compatible, but three defaults changed and one path moved. All are visible at boot (--version, --help, the log) and reversible by flag:

  • KV admission is ON by default β€” --admission-poolless ships enforced (Β§3.6): a request that cannot fit its measured KV demand now gets a fast HTTP 429 + Retry-After before prefill instead of failing or evicting after it. Clients that never handled 429s should either handle them (the header tells you when to retry) or launch with --admission-poolless warn / off.
  • Sampler placement defaults to auto and routes (Β§5.3) β€” earlier cuts defaulted to device (and for one cut auto only observed). Pass --sampling-placement device to pin the old behaviour.
  • Auto-selected speculation tapers β€” --auto-mtp-policy ships taper (draft at 1 live slot, off from 2). An explicit --spec-type is unaffected. Pass --auto-mtp-policy always for the old always-draft behaviour.
  • The GPU-payload directory is versioned by the full engine string β€” ~/.llamafile/v/opencoti-<ver>-<tag>/, not ~/.llamafile/v/<ver>/. Scripts that pre-seed or clean the old path must update (see llamafile-artifacts.md).

To a 0.10.5-based cut (upstream base bump)

The first cut on the upstream llamafile 0.10.5 base carries the same opencoti feature set (the whole patch series is ported), with these migration-relevant deltas:

  • TurboQuant GGUF type IDs are renumbered. Upstream 0.10.5 claimed numeric type slot 42, so the turbo family moved (TURBO2_0=43 … TURBO8_0=48). Any GGUF quantized as TQ3_1S/TQ4_1S under the old IDs will not load on a 0.10.5-based build β€” re-quantize it. KV-cache turbo tiers (-ctk/-ctv turbo*) are runtime-only and unaffected; saved session/state files that embed type ids are also invalidated.
  • -fa auto resolution changed internally (a fused-ops registry replaced tensor-name parsing). Same flag surface, same tri-state; the boot log line to look for is now resolve_fused_ops: Flash Attention enabled.
  • Everything else is surface-compatible; per-cut specifics live in that release's RELEASE_NOTES.md.