Download USAGE.md from ManniX-ITA/opencoti-llamafile: direct link, hf CLI and curl.
- Browser
- Download file 62.5 kB
-
https://huggingface.co/ManniX-ITA/opencoti-llamafile/resolve/main/USAGE.md
- Command line
-
hf download hf://ManniX-ITA/opencoti-llamafile/USAGE.md
-
curl -L -o USAGE.md https://huggingface.co/ManniX-ITA/opencoti-llamafile/resolve/main/USAGE.md
opencoti-llamafile β usage guide
How this engine diverges from upstream Mozilla-Ocho llamafile, what the added features are, how each is gated, its knobs and defaults, its limitations, and which features are meant to be used together.
Audience: anyone running the packaged
opencoti-llamafile-<ver>-<tag>-<arch>.llamafile artifact as a local
inference server. This is the narrative guide; the two companion
references are
docs/llamafile-flags.md (every flag / env /
JSON knob, with defaults, reconciled against the engine source) and
docs/llamafile-artifacts.md (downloads,
GPU payloads, side-load paths, verification). Deep-dive design docs
live in docs/features/, measured evaluations in
docs/evaluations/.
Supported / target model families β read this first
opencoti-llamafile loads any GGUF that upstream llama.cpp loads β that part is inherited unchanged. But the opencoti feature set (KV tiers, rolling-KV, DCA, MTP, sparse attention, RYS) is developed, tuned, and correctness-gated on two model families, in a deliberate primary/secondary split:
Gemma-4 β PRIMARY target
| Model | Kind | Notes |
|---|---|---|
| Gemma-4 26B-A4B-128e ("A4B") | MoE | the flagship serving target; MTP-validated with its native gemma4-assistant drafter |
| Gemma-4 12B / 31B | dense | full feature validation incl. RYS + DCA |
| Gemma-4 E2B / E4B | elastic (E-series) | shared-KV elastic layers supported; MTP drafters available |
Gemma-4 is what the engine is for: its unusual head dims (256 and
512), the iSWA sliding/global dual KV cache, and the per-size
gemma4-assistant MTP drafters
all have dedicated kernels and graph paths here that upstream lacks
or handles slowly. --spec-type draft-assistant, D256/D512 FA-VEC +
scalar-MMA decode, iSWA-aware rolling-KV/SharedKVPool/DCA wiring are
all Gemma-4-first features.
Qwen β SECONDARY target / verification family
| Model | Kind | Notes |
|---|---|---|
| Qwen3.5 / Qwen3.6 (e.g. 35B-A3B) | dense / MoE / hybrid (gated-delta-net) | NextN self-spec MTP (--spec-type draft-mtp, no external drafter needed) |
| Qwen2.5-14B-1M | dense, native-1M | the long-context/DCA validation vehicle |
Qwen is the standard-architecture (head_dim 128) counterweight: every feature that ships is verified on it too, and it carries one feature Gemma doesn't β NextN self-speculation (the model's own MTP head drafts; fused multi-step, at/above upstream parity).
Everything else
Other architectures run with upstream behavior and safe fallbacks,
but opencoti features are unvalidated there, and some are
arch-gated: MTP needs NextN tensors (Qwen-style) or a
gemma4-assistant drafter; RYS --repeat-layers supports the
qwen2/qwen3(+MoE)/qwen3.5/qwen3next/gemma-4 forward loops; DCA is
validated on Gemma-4 and Qwen2.5-1M. Quality gates (KLD,
RULER-niah) were run on the two families above β re-gate before
trusting aggressive KV tiers on anything else.
1. Relationship to upstream llamafile
opencoti-llamafile is upstream llamafile plus an additive patch
series (patches/ in the HF repo, vendors/patches/llamafile/ in
the git repo; the upstream base version and the exact patch list for
a given cut are recorded in its RELEASE_NOTES and every artifact's
MANIFEST.json). Three properties are contractual:
- Off means off. Every opencoti feature is opt-in behind a flag, env var, or per-request JSON field. With no opencoti flags set, the engine's compute path is byte-identical to upstream β this is a regression gate on every patch, not an aspiration.
- Lossless by proof, not vibes. Features that touch the forward pass are gated by logit-equivalence / KLD / RULER-retrieval against vanilla, never by "the output looks fine". Speculative decode is verified-lossless (the output is the target model's).
- Single file, zero dependencies. The artifact is a Cosmopolitan
APE: one file runs on Linux/macOS/Windows/BSD, x86_64 and aarch64.
In the full x86_64 artifact the CUDA and Vulkan backends are
embedded and self-extract to
~/.llamafile/v/opencoti-<ver>-<tag>/on first GPU run; the-winvariant ships without them (see llamafile-artifacts.md). TCQ codebooks and quantization tables are compiled in. No installer, no downloads.
What upstream gives you is unchanged: the server API
(/completion, /v1/chat/completions, /props, /slots, β¦), GGUF
loading, sampling, chat templates. opencoti adds serving-efficiency
machinery on top, aimed at multi-session agentic serving on a fixed
VRAM budget: more concurrent sessions per card, longer usable
context, faster decode.
chmod +x opencoti-llamafile-<ver>-<tag>-x86_64.llamafile
sh ./opencoti-llamafile-<ver>-<tag>-x86_64.llamafile --server --port 8080 \
-m model.gguf -ngl 99
# --version β opencoti-<ver>-<tag> ; without --server you get the chat CLI
Note (Linux): launch via
sh ./file.llamafileif your kernel lacks binfmt_misc APE registration.
1.1 Artifact variants, GPU payloads, side-load
Moved to docs/llamafile-artifacts.md: the
full artifact matrix (Linux / Windows / aarch64 / universal, all with
compressed GPU payloads), the published side-load DSO set, the
versioned ~/.llamafile/v/opencoti-<ver>-<tag>/ runtime path, and
byte-level verification. Short version: the host APE is identical in
every variant; pick the file whose embedded payloads match your OS,
or take the small bare -win file and side-load a DSO.
1.2 Backends β CUDA, Vulkan, CPU
One artifact serves three compute backends; selection is
--gpu {auto,nvidia,vulkan,amd,apple,disable} (default auto, which
probes CUDA first, then Vulkan):
- CUDA (
--gpu nvidia) β the primary, fully-validated backend; every opencoti kernel family (turbo/TCQ in-register FA, DCA, sparse attention, MTP verify) has CUDA instances. sm_75β120f on x86_64, sbsa Blackwell-class on aarch64. - Vulkan (
--gpu vulkan, aliasvk; payloads embedded since c7) β AMD, Intel and NVIDIA through one DSO. The opencoti Vulkan port carries the sparse-attention parity gate, streaming flash-attention, LSE emission and coopmat FA shaders; on NVIDIA it is the fallback when CUDA can't load, on AMD/Intel it is the GPU story. Feature parity is asserted knob-for-knob against CUDA where ported (vk==cuda gates), but CUDA remains the performance reference. - CPU β always available, no payload needed; iqk flash-attention
kernels (
--iqk-flash-attn, default auto) accelerate it. Deleting the side-loaded/extracted DSO forces guaranteed CPU serving.
Flash attention is a tri-state (-fa {on,off,auto}, default
auto): auto resolves per device at graph-reserve time and prints
its decision (resolve_fused_ops: Flash Attention enabled on
0.10.5-based builds). Explicit on/off remain authoritative. The
opencoti split-attention paths (head-split, NEO, GPU_STREAM) are
resolver-aware β non-split layers still resolve FA normally.
2. Feature map β what exists and how it's gated
| Feature | Default | Turn on with | Class |
|---|---|---|---|
| Session-keyed KV reuse | off (per request) | session_id JSON field |
latency |
| ReST-KV retention eviction | off | --rest-kv-eviction |
quality-under-overflow |
| KV quantization (scalar) | f16 | -ctk / -ctv |
capacity |
| TurboQuant / TCQ KV tiers | off | -ctk/-ctv turbo* |
capacity |
| Auto KV-tier policy | off | OPENCOTI_KV_AUTO_TIER=1 |
capacity (policy) |
| PolyKV pool (SharedKVPool) | off (per request) | shared_pool_slot JSON field |
multi-agent capacity |
| PolyKV control-plane API | off | --polykv-max-pools N + /polykv/* routes |
multi-agent orchestration |
| Rolling-KV window / spill | auto (engages only under pressure) | --vram-target, --kv-residency-mode |
capacity |
| Mixed-KV spilled tail | off | -ctkt / -ctvt |
capacity |
| DCA long-context | off | --dca on |
context extension |
| Sparse attention (block-selector) | off | --sparse-attn on |
long-ctx decode speed |
| Sparse-V | auto on iSWA+quant-V, else off | TURBO_SPARSE_V_TAU |
decode speed |
| MTP speculative decode | off | --spec-type + drafter |
decode speed |
| RYS layer duplication | off | --repeat-layers |
quality |
| RYS probe | off | --rys-probe |
tooling |
| Lazy slot context | off | --slot-initial-ctx, --slot-shrink-idle-ms |
embedder memory |
| GPU-share TDM pacing | auto (engages only with β₯2 opencoti servers on one GPU) | --gpu-share-weight N, GET /gpu/peers |
multi-server fairness |
| Elastic parallel slots | off | --max-parallel N (+ tps floor / VRAM reserve) |
serving capacity |
| iSWA SWA-pool seq budget | off (= n_seq_max) | --swa-seq-budget B |
serving capacity (VRAM) |
| KV admission gate | on (enforced) |
--admission-poolless {off,warn,enforced} |
serving capacity |
| Sampling placement | auto (routes) |
--sampling-placement {device,cpu,auto} |
decode speed |
| Auto-MTP depth policy | taper |
--auto-mtp-policy {allocator,taper,always,off} |
decode speed |
| Introspection API | always on | GET /props, GET /slots |
observability |
Every boot flag also has an env twin
(OPENCOTI_LLAMAFILE_<SNAKE_CASE> for adapter-typed fields,
LLAMA_ARG_* for llama.cpp-registered ones). The per-flag reference β
every knob, value set, default, and env name β is
docs/llamafile-flags.md; this table is the map,
not the contract.
2.1 Sharing one GPU between servers (gpu-share, c5/c6)
Multiple opencoti-llamafile servers on the SAME physical GPU discover
each other zero-conf (shared-memory peer registry keyed on the PCI bus
id, so CUDA_VISIBLE_DEVICES ordering doesn't matter) and pace
themselves by deterministic weighted time-division: a repeating
period is carved into one contiguous window per active peer,
proportional to --gpu-share-weight (default 1). The split is exact
by construction and work-conserving β an idle peer's time flows to the
busy ones, and a single server pays zero overhead (pacing only
engages with β₯2 active peers). The period adapts for interactive
latency (80β400 ms; worst token stall at 2:1 β 80 ms). Measured on a
3090: 2:1 β 1.97:1 at ~solo aggregate; on Windows (5080): 2.00:1.
Watch peers live via GET /gpu/peers. Full design + measurements:
docs/features/gpu_share.md. Crashed peers age out β€3 s; no daemon,
no MPS, works on Windows.
3. KV capacity stack β PolyKV
PolyKV is the umbrella name for this whole stack: the compressed
shared KV pool. Concretely it is the KV quantization tiers of Β§3.1
plus the multi-agent SharedKVPool of Β§3.4, stacking with the
auto-tier policy (Β§3.2) and the rolling-KV window (Β§3.3), and
orchestrated at runtime through the control-plane API of Β§3.5
(pools/fork/admission/capacity/tps REST surface + the
opencoti-langgraph package). If you arrived here looking for
"PolyKV" from an announcement: Β§3.4 is the shared-prefix pool
itself; Β§3.1 is what the pooled cells are made of; Β§3.5 is how an
agent framework drives it.
These four features share one goal β fit more context / more sessions in fixed VRAM β and are designed to stack. Recommended order of adoption: scalar quant β auto-tier β rolling-KV β turbo tiers β SharedKVPool.
3.0 --kv-unified β one shared cell pool vs per-slot reservations
This flag decides what -c and --parallel actually mean. Get it
wrong and every capacity number in this document is wrong too. The two
layouts are not a tuning preference; they are different allocation
models.
With --kv-unified (recommended for agentic / multi-session)
n_ctx_seq = n_ctx (src/llama-context.cpp:456) and the cache runs as
a single pool of -c cells with one stream
(n_stream = 1, src/llama-kv-cache.cpp:701).
-cis the whole pool, shared. It is NOT divided by--parallel.- Every slot advertises the full
-cas its context. The server setsslot.n_ctx = llama_n_ctx_seq(ctx)with no division (tools/server/server-context.cpp:1894). One long session may legally consume the entire pool. --parallelcosts no base-KV VRAM. Raising it does not enlarge the cache. (The one exception is iSWA β see below.)- There is no isolation. Cells are first-come-first-served, so a greedy session can starve its neighbours. This is exactly why the PolyKV admission gate (Β§3.5) exists: with a shared pool, admission control is not optional, it is the isolation mechanism.
- Pool prefixes are shared zero-copy (Β§3.4).
Provisioning rule of thumb (bug-2287) β size the pool so the expected working set fits:
-c β tokens_per_session Γ (--parallel + --polykv-max-pools)
This is a sizing guideline you apply, not something the engine
enforces. PolyKV pools occupy reserved seq-ids above the slot range
(seq = n_parallel + k), so they need their own share of the pool.
Without --kv-unified (per-slot reservation)
n_ctx_seq = n_ctx / n_seq_max, and the cache splits into
n_stream = --parallel independent per-slot caches.
-cis divided. Each slot gets a hard-reserved-c / --paralleland can never exceed it, no matter how idle the others are.--paralleldirectly divides usable context per session.- Isolation is free β no session can affect another's capacity.
- Idle slots waste their whole reservation.
- Pool prefixes are copy-shared, not zero-copy.
Choosing
| Want | Use |
|---|---|
| Many sessions of varying, unpredictable length | --kv-unified + admission control |
| Long single sessions that should use all VRAM | --kv-unified |
| Hard per-tenant context guarantees | omit --kv-unified |
| Benchmarks where per-slot capacity must be identical and fixed | omit --kv-unified |
The iSWA exception (Gemma-4). On sliding-window models the SWA
cache is the one allocation that scales with the seq-id count, even
under --kv-unified, and it is allocated eagerly in VRAM at boot
(src/llama-kv-cache-iswa.cpp:71):
n_seq_max = --parallel + --polykv-max-pools
size_swa = n_swa Γ n_seq_max + n_ubatch (padded to 256)
--polykv-max-pools counts. A pool occupies a reserved seq-id above
the slot range, so it costs a full SWA window exactly like a slot does.
Budgeting against --parallel alone silently under-counts.
Measured on bs2 GPU0, gemma-4-31B-it-Q6_K, n_swa = 1024,
n_ubatch = 512, -ctk/-ctv q8_0 (boot-only probe, 5/5 exact):
--parallel |
--polykv-max-pools |
n_seq_max |
-c |
base cells | SWA cells |
|---|---|---|---|---|---|
| 64 | 4 | 68 | 417792 | 417792 | 70144 |
| 16 | 4 | 20 | 417792 | 417792 | 20992 |
| 8 | 4 | 12 | 417792 | 417792 | 12800 |
| 8 | 0 | 8 | 417792 | 417792 | 8704 |
| 8 | 4 | 12 | 131072 | 131072 | 12800 |
Read the two invariants off the table: the base cache equals -c
exactly and is identical at n_seq_max 8 and 68, and the SWA cache
is unchanged when -c drops 3.2Γ. They are independent axes.
At 31B's 425 KiB per SWA cell (50 SWA layers Γ 16 kv-heads Γ 256
head_dim Γ K+V at q8_0) the 28.4 GiB of SWA alone** β which is what forces the context budget
down on large-slot Gemma runs. --parallel 64 --polykv-max-pools 4 row is
**--parallel 8 --polykv-max-pools 4 is
~5.2 GiB. Qwen-class models have no SWA split and pay nothing here.
Plan work to make this budget independent of the slot count:
docs/features/elastic_parallel.md.
3.1 KV quantization: scalar types + TurboQuant/TCQ tiers (PolyKV M6)
The KV cache type is set per-tensor-half: -ctk <type> (keys) and
-ctv <type> (values), independently β asymmetric pairs are
first-class (e.g. -ctk q8_0 -ctv q4_0).
Supported types: f16, bf16, q8_0, q6_0, q5_1, q5_0,
q4_0 (scalar), turbo2, turbo3, turbo4, turbo8
(TurboQuant, MSE-optimal with Walsh-Hadamard rotation + InnerQ),
turbo2_tcq, turbo3_tcq (trellis-coded, Viterbi-encoded).
Which to pick (measured):
- 8-bit / 4-bit: use
q8_0/q4_0. The native scalar types dominate turbo8/turbo4 at equal width β turbo earns nothing there. -ctk q8_0 -ctv q4_0is the workhorse asymmetric pair: keys keep 8-bit fidelity (attention logits are K-sensitive), values take the compression.- 3 bits and below is TurboQuant territory:
turbo3Pareto-beats q4_0 (90 vs 129 MiB KV at equal quality, teacher-forced TV 0.0067 vs 0.0094);turbo2is the smallest logit-equivalent KV that exists (~2 bit) β the 256k-context play. TCQ variants trade encode cost for a further fidelity step at the same width. - All shipped tiers pass logit-equivalence gates; decode runs the quantized data in-register in the flash-attention kernel (no f16 materialization) for turbo2/3/4 and TCQ.
Limitations: turbo8 uses a materialize fallback (not fused); at Gemma-4's head_dim 512 only turbo2/turbo3 have fused D=512 instances; prefill on very long prompts uses a hybrid path automatically. Quality validation on Gemma franken-merges must use retrieval (niah), not perplexity.
KVarN (beellama) is NOT a shipped type. It was evaluated against
the tiers above in c8 stage 9 (2026-09-10): at 2β4 bits it is 2β13Γ
better on teacher-forced KLD than the matched turbo/tcq/q4_0 tier and
ties q6_0/q8_0, but it is a 128-token tile format with its own cache
and kernels (not a -ctk type that can be added here), and on an
RTX 3090 it decodes β21 % and prefills β27 % slower than f16 at every
width (our TCQ tiers decode slower still, plain turbo faster). Decision
2026-09-10: KVarN is being integrated (c8 stage 9b) with a prefill
optimization pass to follow; the turbo tiers stay available until
that lands and the retirement decision is taken. Numbers and status:
docs/evaluations/kvarn_vs_turbo.md.
3.2 Auto KV-tier (OPENCOTI_KV_AUTO_TIER=1)
Boot policy: pick the least-compressing scalar pair that keeps the whole KV resident in the VRAM budget; if even that spills, the T* model decides between "small f16 spill" and "quantize one tier down" by predicted tokens/s drop.
Knobs (env): OPENCOTI_KV_AUTO_TIER=1 (master),
OPENCOTI_KV_TSTAR_DROP (target drop, default 20%),
OPENCOTI_KV_TSTAR_MAX_SPILL_MIB (default 800),
OPENCOTI_KV_AUTO_TIER_TAIL=1 (also auto-pick a q4_0 spilled tail).
Explicit -ctv disables auto entirely; explicit -ctk holds K and
walks only V. Dense full-attention models only (iSWA models keep f16).
Read back what it decided: GET /props β .opencoti.kv.effective β
the configured vs effective split exists exactly because auto-tier
may override you.
3.3 Rolling-KV window (residency / spill)
"KV doesn't have to fit." Each layer keeps a device-resident window
of recent positions; the tail [0 β¦ window_start) lives in pinned
host RAM and is streamed through the attention kernel per-tile,
merged exactly via online-softmax (LSE). When everything fits, every
layer is GPU_RESIDENT and the path is byte-identical to vanilla β
the feature only engages under memory pressure.
Flags: --vram-target <MiB> (budget cap; 0 = all free VRAM minus
reserve), --kv-residency-mode {auto,head,window} (default auto;
leave it), -ctkt / -ctvt (distinct, more-compressed types for the
spilled tail β "mixed-KV": f16 recent window β q4_0 tail).
Performance model (RTX 3090, PCIe ~6.5 GB/s): spill decode sits at
the PCIe floor, t(token) β fixed + tail_bytes / link_bw β linear,
no cliff. On a fast-link host (RTX 6000, ~50 GB/s) window-mode spill
is genuinely usable; on consumer PCIe it's a last resort β prefer
quantizing (that's what auto-tier automates).
Limitations: while a window is spilled, context-shift and
prompt-cache-reuse are guarded off (requests bounded at n_ctx);
the compute-buffer reserve for long contexts is measured
automatically at boot (two-pass reserve β no knob).
Multi-session: since c6 the rolling window works with
--parallel N at full speed β each session gets its own windowed
KV stream, and multi-stream windowed decode scales positively
(measured: 2 concurrent windowed sessions aggregate above the solo
rate, per-session β 60% of solo on a PCIe-bound 3090).
Fixed in the c7 re-cut (r2, 2026-09-20) β check which c7 you have. The FIRST c7 byte set (2026-09-03) aborts as soon as a position window engages (
--vram-target/--kv-residency-mode windowwith a context that overflows it), on both cache layouts:fattn-common.cuh:87: GGML_ASSERT(dst->op == GGML_OP_FLASH_ATTN_EXT). It is a 0.10.5-port regression (bug-3369): the streaming-attention op's plain-FA fallbacks hand the new base a node it now asserts on. Patch0253fixes it, and c7 was re-cut with it β same tag, same file names, new bytes. If your x86_64 artifact hashes to3c907bc7β¦you have the broken set: re-download (the r2 x86_64 sha256 starts4f4102d6β¦; full table in the release notes). An r2 binary replaces an already-extracted r1 CUDA DSO by itself β no cache clearing. If you cannot update, keep the KV resident on the old bytes: lower-c, or compress it (-ctk/-ctv, Β§3.1).
The window also works under --kv-unified β so it composes with
the dynamic-slot serving mode (--kv-unified --max-parallel N).
Earlier cuts of this guide called the pair "inherently incompatible";
that is not true of the current engine, which never refuses the
combination. Validated 2026-09-19 on the development build (Qwen3-8B
Q8_0, f16 KV, --parallel 2 --kv-unified --vram-target, two
concurrent ~11k-token sessions = 22 240 cells against a 7 424-cell
resident window, so ~two thirds of the live cells sat in the host
tail): head / middle / tail needles all retrieved in both sessions,
and every answer byte-identical to the fully-resident run and to the
split-cache run. What differs between the layouts is sizing, not
correctness: unified has ONE window over the shared pool (-c cells),
split has one window per stream (-c / --parallel cells each).
PolyKV pools (Β§3.4) work on both layouts too. Plain (non-windowed)
split-KV --parallel serving is also full-speed as of c6 (a 30β60Γ multi-slot slowdown in earlier
cuts was fixed). Quantized spilled tails (-ctk/-ctv q4_0/q8_0 with
an active host tail) keep full-bandwidth bulk staging with multiple
concurrent sessions too (c6 r2): on a ~50 GB/s PCIe host, two
concurrent q4_0-tail sessions aggregate within ~5% of the solo rate
(earlier c6 bytes collapsed ~6Γ in exactly this combination).
3.4 PolyKV SharedKVPool (multi-agent shared prefix)
N agents attending one physical copy of a common prefix (system prompt + tool defs). Per-request JSON, no CLI flag:
{ "shared_pool_slot": 0, "shared_prefix_n_tokens": 4096, β¦ }
Server must run with --no-cache-idle-slots (mandatory β the default
idle-slot save/clear would evict the pooled prefix). Since c6 pools
work on BOTH cache layouts: with --kv-unified the prefix is shared
zero-copy (cell set-membership, one physical copy); on the
split-stream cache (e.g. rolling-KV window mode) the share is a
physical per-stream copy of the prefix β same semantics, higher
footprint (device use stays window-capped in window mode).
What the pool speeds up β measured (Gemma-4-26B-A4B Q4_K_M, RTX 3090, q8/q8 unified KV, Pβ5k-token shared prefix, greedy fixed-length decode, 8 concurrent sessions unless noted):
| axis | naive (N private copies) | shared pool | gain |
|---|---|---|---|
| KV cells (N=8) | ~8Β·P | P + suffixes | 6.9Γ (~306 vs ~9 agents on a fixed buffer) |
| prefill, 8 sessions joining | 22.8 s | 4.6 s | ~5Γ (prefix enters KV once per pool) |
| steady-state batched decode (N=8) | ~190 tok/s | 217 tok/s | +14% (8 queries read one physical prefix β L2 reuse, smaller cell span) |
| multi-turn re-query (N=8) | 99 tok/s | 225 tok/s | 2.3Γ (see note) |
| iso-speed capacity | 8 sessions @ 24.0 tok/s each | β₯12 sessions @ β₯26.9 tok/s each | β₯1.5Γ sessions (crossover not reached at N=12; aggregate 315 tok/s) |
The multi-turn row is iSWA-specific and easy to miss: a private slot that has decoded past its prompt cannot partially rewind (upstream SWA-checkpoint semantics, llama.cpp PR #13194), so re-querying it pays a checkpoint restore or a full re-prefill every turn. The pool slot never decodes, so its prefix never slides β every re-attach is free. Note the capacity row is about per-session speed, not just aggregate: 12 pooled sessions each decode faster than 8 private ones.
Cross-architecture results (RTX 6000 96GB, Pβ5073, GEN=256, A/B/A naive/shared/naive): the pool is validated on all three attention architectures, and the memory axis is architecture-independent (~6.8β6.9Γ at N=8 β it counts cells, not attention math).
| axis | Qwen2.5-14B-1M Q8_0 (pure full attention) | Qwen3.6-27B-Omnimerge-v4 Q4_K_M (hybrid GDN + NextN MTP n=3) |
|---|---|---|
| KV cells (N=8) | 6.78Γ (~406 vs ~9 agents on a fixed 8192-cell buffer) | 6.89Γ (~304 vs ~9 agents) |
| batched decode N=8 | 358 β 403 tok/s (+12.5%) | 106 β 123 tok/s (+15.6%) |
| batched decode N=24 | 412 β 742 tok/s (+80%; 17.2 β 30.9 tok/s per session) | 90 β 125 tok/s (+39.7%; 3.7 β 5.2 per session) |
| shared-only sweep N=32/48/64 | 813 / 871 / 865 tok/s (plateau ~870 near N=48) | 126 / 124 / 124 tok/s (saturates by Nβ24β32) |
The shared-vs-naive decode gain grows with N on both. On hybrid/recurrent models (delta-net, mamba) the absolute aggregate saturates much earlier than on pure attention β the recurrent layers batch worse β so there the pool buys concurrency capacity and memory, not aggregate throughput past Nβ24.
End-to-end agentic A/B β measured (package-courier harness,
opencoti/langgraph/bench/package_courier/drivers/courier-bs2-ab-polykv.sh; bs2 RTX 6000 96 GB GPU 0;
Gemma-4-26B-A4B-it Q4_K_M + assistant-MTP; q8_0/q8_0 unified KV;
100 packages Γ 10 chained steps, seed 1, temperature 0). Everything
that could compensate is switched off: 8 workers / 8 slots frozen
(--tps-floor 0, so scheduler() never starts), and per-slot
context held equal at 6144 tok β -c 73728/(8+4 pools) vs
-c 49152/8, not equal -c, which per bug-2287 would have silently
handed the naive arm 50% more context per slot. The only difference
is the pool.
| axis | naive (--no-polykv) |
shared pool (P=1177) | gain |
|---|---|---|---|
| wall clock, 100 packages | 1125.5 s | 956.6 s | β15.0% |
| delivery latency, mean | 86.6 s | 72.7 s | β16.1% |
| delivery latency, p90 / p99 | 92.7 / 95.1 s | 78.6 / 81.0 s | β15.2% / β14.8% |
| per-turn gen tps, p50 | 11.7 | 14.1 | +21% |
| per-turn gen tps, p90 | 18.3 | 27.5 | +50% |
| generated tok/s over the run | 49.3 | 58.1 | +18% |
| prompt tokens processed | 2,708,083 | 2,444,410 | β9.7% |
| task score | 100/100 | 100/100 | β |
Both arms ran exactly 2200 turns and generated within 0.3% of the same token count (55,469 vs 55,625) β the same work, done two ways.
Why 58 tok/s here and 315 tok/s three tables up β they measure
different things. This is the single most common misreading, so
state it plainly: generated tok/s over the run is a workload
number, not a decode-speed number. It is answer tokens divided by
wall clock, and the wall clock of an agentic run is dominated by
reading, not writing. Decomposed from this run's own turn events:
| fan-out sweep (315) | single-stream MTP (320) | this A/B (58.1) | |
|---|---|---|---|
| GPU | RTX 3090 | RTX 6000 | RTX 6000 |
| sessions decoding | 12, all at once | 1 | 4.14 on average |
| generated per turn | 256 fixed | 256 fixed | 25 tokens |
| prompt : generated | prefill timed separately | n/a | 44 : 1 |
| clock includes prefill? | no β decode-only rounds | no | yes β 40% of engine time |
The agentic engine processed 2,444,410 prompt tokens to produce 55,625 answer tokens, in 956.6 s β i.e. ~2,610 tok/s of total token throughput, of which only the 58.1 are answer tokens. It spent 3,960 GPU-seconds decoding and 2,667 prefilling (86.6% busy across 8 slots). Per session while decoding it ran 14.0 tok/s, not because decode got slower but because 2.8 other sessions were prefilling into the same GPU at the same time β which the decode-only synthetic rounds deliberately exclude.
So all three numbers are true simultaneously: 320 tok/s is what one session sustains with nothing else running, 315 tok/s is what twelve concurrent sessions aggregate to in steady-state decode on a 3090, and 58.1 tok/s is how fast a real agent fleet emits answer tokens while the same GPU also chews through 2.4M tokens of tool output and conversation history. Only compare a number to one measured the same way.
Read this table next to the synthetic ones, not instead of them. It is the first PolyKV measurement on a real agentic workload β variable-length tool-calling turns, MTP speculation live, and a modest 1177-token shared prefix (system + tool schemas) rather than the 5k synthetic one β so the gain is smaller than the +80% the N=24 fan-out sweep shows. It is also the harder number to argue with: with fan-out frozen, no scheduler behaviour can flatter either arm, and the β9.7% prompt-token count is the shared prefix simply not being prefilled eight times. Expect the gain to grow with a longer shared prefix (agent framework + large tool schemas) and with N; expect it to shrink toward zero as the private suffix comes to dominate the prompt.
How to prefill the pool β use the common-prefix token array, not
the document text. Tokenizers merge across the document/suffix
boundary (on the Qwen tokenizer the last prefix token fuses with the
suffix start), so tok(DOC) can be one token longer than the common
prefix the agents actually share β and a pool that is even one token
longer than shared_prefix_n_tokens cannot be shared exactly. The
correct client sequence:
// 1. tokenize the FULL agent prompts and compute
// P = min over agents of commonPrefixLen(tok(DOC), tok(DOC+suffix_i))
// 2. prefill the pool slot with the token array itself (llama.cpp
// /completion accepts token arrays) β pool state == P on ANY tokenizer:
{ "prompt": [/* tok(DOC+suffix_1)[:P] */], "id_slot": 0,
"cache_prompt": true, "n_predict": 1 }
// 3. agents attach with prompts STRICTLY longer than P:
{ "prompt": "<DOC + private suffix>", "id_slot": 1,
"shared_pool_slot": 0, "shared_prefix_n_tokens": P, β¦ }
On attention-only models a text prefill happens to work (the ranged cell copy tolerates the extra token); on hybrid/recurrent targets it silently disables every share β see the gotcha below. The token-array prefill is correct everywhere.
Hybrid/recurrent gotcha (GDN / mamba / qwen35moe-class models).
A recurrent cache has one rolling state per sequence, not per-position
cells, so a pool share is only possible as an exact full-state
share. The server enforces this: the share engages only when
shared_prefix_n_tokens == pool state length and the request prompt
is strictly longer than the shared prefix; anything else logs
poly-kv-pool: hybrid/recurrent target needs exact full-state share β¦ skipping share, full reprocess (bug-2203) and falls back to a full
(correct, slower) reprocess. If you see zero speedup on a hybrid
model β or mass HTTP 500s at high N because N unshared full prompt
copies overflow the unified KV β grep the server log for that WARN:
it almost always means the pool was prefilled with text instead of
the token array. (Older builds crashed outright here β
failed to remove sequence N with p0=β¦ β fixed by patch 0135.)
Sizing note: the pooled prefix pins P cells in both iSWA caches
(global + SWA) for the pool's lifetime. Budget -c for pool prefix
- N session windows + generation headroom, or long-running sessions can exhaust slot allocation mid-round.
Composes with KV quantization (the pool holds quantized cells) and with session KV-reuse. The pool is read-only for consumers; each agent's divergent suffix is private.
Tiering is pinned per-pool, never per-session. The K/V tiers β including the mixed-KV recent-window β compressed-tail pair β are properties of the boot-allocated cache tensors, chosen once at boot (by you or by auto-tier) before any session exists. The window/tail boundary is a per-layer residency budget over the physical cell axis, so a prefix cell is in the resident window or evicted (and quantized exactly once, on eviction) for all sequences simultaneously. Sharing itself is not copy-on-write: a sharer joins the prefix by adding its sequence bit to the existing cells, and a diverging session just appends private suffix cells β there is no per-session copy that could be re-quantized, and no way for two sessions to see the same prefix at different tiers. The flip side: you cannot give one session a higher-precision read of a shared prefix than another; that would require forking the prefix into a private copy, which is exactly the O(N) memory cost the pool exists to avoid.
3.5 PolyKV control-plane API (pools, fork, admission, capacity, tps)
Everything in Β§3.4 manages the pool by hand (you own the slot, the token array, the prefix length). The control-plane API makes pools first-class server objects with a REST surface, built for agentic orchestrators (LangGraph & co.) that spawn/retire sub-agents at runtime. Enable it at boot:
--parallel N --polykv-max-pools M # + --kv-unified for zero-copy shares
With --kv-unified pool prefixes are shared zero-copy; without it
(split cache β required by the rolling-KV window) each attach
physically copies the prefix into the session's stream (c6).
--polykv-max-pools M reserves M pool sequence-ids beyond the slot
range (default 0 = off; every route below 404s when off β the feature
is fully additive). A reserved pool costs no KV cells until it is
materialized.
Pool objects. A pool is a pinned, immutable token prefix living in
the unified KV, addressed by pool_id, arranged in a tree:
children extend (or copy-on-write branch) their parent and share its
cells. The pool json:
{ "pool_id": 1, "parent": 0, "branch_pos": 65, "prefix_len": 82,
"prefix_hash": "β¦", "pinned": false, "children": [],
"orphaned_pin": false,
"admission": { "mode": "advisory", "target_tps_per_session": 0, β¦ } }
Routes (all JSON; also see the Python client below):
| Route | What it does |
|---|---|
POST /polykv/pools |
Create a root pool. Body: one of prompt (text), tokens (array), from_session / from_slot (capture a live session's prefix); plus `pin: true |
GET /polykv/pools |
List all pools + pools_max, seq_ids_used, tree_depth. |
GET /polykv/pools/{id} |
One pool, incl. orphaned_pin (pinned leaf that nothing references β a compaction leak tell). |
POST /polykv/pools/{id}/fork |
Child pool. Body: prompt or tokens = the child's full prefix (see contract below), optional branch_pos: D for a copy-on-write branch sharing only [0, D), pin. |
POST /polykv/pools/{id}/pin / β¦/unpin |
Pin = survive even with zero attached sessions/children. |
POST /polykv/pools/{id}/release |
Release; cells reclaim when the subtree refcount hits zero (ancestor cells survive while descendants live). |
POST /polykv/pools/{id}/sampling |
Per-pool sampling-placement override (see docs/features/polykv_api.md Β§16). |
POST /polykv/pools/{id}/admission |
Set per-pool admission: { "target_tps_per_session": T, "mode": "advisory"|"enforced", "on_saturation": "reject"|"warn", "guarantee_min_sessions": G, "settle_tokens": S, "settle_max_ms": M } (last three = P7, defaults 1 / 48 / 5000). |
GET /polykv/pools/{id}/capacity[?expected_tokens=N] |
Pre-spawn gate: can_admit, reason, headroom_sessions, mean_active_tps, compaction_pressure β [0,1], free cells; P7 adds settling + settle_remaining_ms, n_warming, n_pool_sessions, guaranteed, projected_mean_tps_model / projected_mean_tps_measured / drop_per_admit_ewma, and projected_idle_estimate (all sessions between turns β projection from the last settled mean, < 30 s fresh β "no measurement" is NOT "free capacity"). The expected_tokens arm is checked first: not enough KV headroom β reason: "context headroom exhausted". |
GET /polykv/tps |
Telemetry: aggregate + n_warming + per-session {session_id, tps_ewma, active, ctx_used, ctx_total, ctx_headroom_tokens}. ?once=1 for a snapshot, default is an SSE stream. |
Admission pacing β P7 settled admission (c5). Admitting a new agent
distorts the very measurement the next admission decides on (the tps
EWMA needs ~2 s to absorb it), so /capacity runs a settle window
per pool: after each new-session admit, until that session has decoded
settle_tokens (48) or settle_max_ms (5000) elapsed, the pool
reports settling: true and withholds can_admit. Advisory
orchestrators should treat a settling tick as a no-op (neither spawn
nor shrink); the enforced gate holds the attach briefly instead of
429ing. Sessions still warming up (tps_ewma = 0) are excluded from
mean_active_tps and counted in n_warming. The projection improves
as the pool runs: each admit's observed mean-tps drop folds into
drop_per_admit_ewma, and projected_mean_tps_if_admitted switches
from the n/(n+1) model to mean β drop_ewma once a sample exists β
gate spawns on the projection, not the raw mean. Two escape valves
prevent starvation: guarantee_min_sessions (default 1) always admits
a pool's first N agents even past the floor (only the physical
context arm can refuse), so a fresh or nested pool can never deadlock;
and "overcommit": true on a request explicitly bypasses the enforced
gate. Continuing sessions (an existing sessionβslot affinity) are
never gated β only genuinely new sessions count as admissions.
Attaching sessions. A completion/chat request joins a pool with a single JSON field:
{ "pool_id": 1, "session_id": "agent-7", β¦ }
The server computes the shared prefix P automatically β the
longest token-exact common prefix between the request and the pool
(no shared_prefix_n_tokens bookkeeping; that manual field remains
for the Β§3.4 low-level path). Proof of attach: the response timings.cache_n
β₯ the shared pool prefix on a cold session. Responses carry
X-Session-TPS; under enforced admission a rejected request gets
HTTP 429 + Retry-After, and admitted ones may carry
X-Sessions-Remaining backpressure.
The fork/prefix contract (the #1 integration mistake).
prompt/tokens in fork is the child's FULL prefix [0, L) β
the parent's exact prefix followed by the new suffix β validated
token-exact against the parent (a suffix-only prompt 400s with
child prefix shorter than branch_pos). Text concatenation can shift
tokenizer boundaries, so the robust recipe is the token path:
// child tokens = parent tokens + suffix tokenized WITHOUT specials
POST /tokenize { "content": "<suffix>", "add_special": false }
POST /polykv/pools/{parent}/fork { "tokens": [/* parent β§Ί suffix */] }
And if sessions attach through /v1/chat/completions, the pool
prefix must be the chat-templated system block (e.g.
<|im_start|>system\nβ¦<|im_end|>\n for Qwen) β a raw-text prefix
shares zero tokens with templated requests (cache_n: 0).
Context exhaustion & compaction (see
docs/features/polykv_api.md Β§13 for the full design). Pool-attached
sessions never context-shift: a shift rewrites shared cell
positions and would corrupt every other reader of the prefix. With
--context-shift on, the server WARNs at boot and pooled sessions
stop gracefully at capacity with truncated: true instead of
shifting (non-pooled sessions still shift normally). Compaction is a
prompt rewrite, never an in-place KV operation: watch
compaction_pressure, then re-root β fork the stable ancestor with
ancestor prefix + summary as the new full prefix, migrate sessions
to the new pool_id, unpin + release the old working pool. A pinned
pool left behind shows up as orphaned_pin: true.
Python / LangGraph. The opencoti-langgraph package
(opencoti/langgraph/ in the repo) wraps all of the above:
PolykvClient/AsyncPolykvClient (incl. tokenize() and
compact_by_refork()), OpencotiChatModel (a LangChain
BaseChatModel with pool_id/session_id attach, per-response
session_tps/cache_n metadata and 429 β PoolSaturatedError),
pool-control tools, a capacity-gate node and a CompactionNode. See
its README for graph-level usage.
3.6 Serving capacity β elastic slots, SWA budget, the admission gate
Three c7-era levers turn "how many sessions fit" from a boot-time
guess into something measured and enforced at runtime. They matter
most on iSWA models (Gemma-4) with --kv-unified, where Β§3.0 showed
the SWA cache scaling with the slot count and the base pool being
first-come-first-served.
Elastic slots (--max-parallel N). --parallel becomes the
initial live slot count, not the ceiling: N slots are allocated but
parked, and the elastic tick admits one at a time while two guards
hold β the projected per-slot decode rate stays above
--max-parallel-tps-floor (default 0 = arm off) and free VRAM stays
above --max-parallel-vram-reserve (default 512 MiB; an
external-pressure guard β growing a live slot allocates nothing).
Requires --kv-unified.
SWA sequence budget (--swa-seq-budget B). On iSWA models every
slot and every reserved PolyKV pool pays a full sliding window up
front (n_swa Γ n_seq_max + n_ubatch cells, Β§3.0). This caps the SWA
pool at B sequences instead of n_seq_max β measured tens of GiB on
large-slot Gemma configurations. Sizing and admission use the same
model, so an over-budget configuration is refused coherently rather
than failing later. It can only shrink the pool; inert on non-iSWA
models and with --swa-full.
The KV admission gate (--admission-poolless, default
enforced). Without admission, an over-committed pool makes a
request pay with an eviction or a failure after prefilling. The gate
prices every completion request (the name is historical β it began
with pool-less requests and now covers pooled ones too, after their
PolyKV policy gate) against every KV pool the server has β the base
pool (under --kv-unified) and, on iSWA models, the SWA pool. Demand
is measured, not guessed: the tokenised prompt plus the declared
generation bound (an undeclared bound contributes 0 β this prevents
admission-time over-commitment, it does not predict unbounded growth).
The reservation is taken atomically on the server thread in the
same operation that decides it fits, so simultaneous arrivals cannot
all be admitted against the same free cells; it is released when the
request binds to a slot.
enforced(default): refused requests get HTTP 429 +Retry-Afterbefore any prefill β the client-visible contract is a fast, cheap refusal (measured: ~0.4 s vs ~11 s for the prefill-then-fail path it replaces).warn: admit anyway, name the saturated pool in anX-KV-Pool-Saturatedresponse header.off: pre-admission behaviour, byte-identical.- Per-request
"overcommit": truebypasses the enforced gate explicitly (Β§3.5); pooled sessions additionally get the per-pool settled-admission machinery of Β§3.5.
Pre-spawn capacity checks. An orchestrator that spawns agents
should ask before spawning, not catch 429s after:
GET /polykv/pools/{id}/capacity?expected_tokens=N (Β§3.5) answers
can_admit from live tps + KV headroom. The 429 path is the backstop
for the requests that arrive anyway.
4. Long context
4.1 DCA β Dual Chunk Attention (training-free context extension)
Splits attention into intra-chunk / successive / inter-chunk position
regimes and merges them exactly by LSE, so a model trained at
n_ctx_train serves multiples of it without retraining.
Flags: --dca on (default off), --dca-chunk-size N (default
derives from the model's training context; explicit 8192 is the
validated recipe), --dca-yarn-factor F (default 1.0; measured
neutral for retrieval β leave it). Serve beyond the GGUF's declared
context with
--override-kv <arch>.context_length=int:1048576.
Validated recipe (Gemma-4-A4B, n_ctx_train 256k):
--dca on --dca-chunk-size 8192 -fa on --parallel 1 \
--override-kv gemma4.context_length=int:1048576
Measured retrieval (RULER-VT, n=50): 256k 0.964 Β· 512k 0.996 Β· 768k 0.984 Β· 1M 0.916 β a gentle β7 pp at 4Γ native, no cliff. Counter-proof on Qwen3-8B (native 41k): plain attention collapses at 128k (PPL 19.2) while DCA holds PPL 7.3.
Works on Gemma-4 (its 5 global layers; SWA layers untouched) and Qwen2.5/3/3.5 (all layers). Composes with quantized KV (scalar pairs all pass; q8-DCA decode costs ~2Γ vs f16-DCA), sparse attention, and rolling-KV.
Limitations: DCA caches K un-rope'd β launch-time toggle only (a server booted DCA-on can't switch off per request); expect approximation, not identity, past one chunk. On models that are already native long-context (e.g. Qwen2.5-1M), DCA can only approximate down β don't use it there.
4.2 Sparse attention (Quest block-selector) + sparse-V
Two independent decode-bandwidth levers:
- Block-selector (
--sparse-attn on): per-block min/max key bounds give an upper bound on each block's attention mass; decode visits only the top-K blocks (+ sinks + recent). Flags:--sparse-attn-block-size(128),--sparse-attn-topk(default 0 = visit all blocks, i.e. no skipping; passautofor adaptive max(64, n_blocks/4), or an explicit block count),--sparse-attn-recent,--sparse-attn-sink(1),--sparse-attn-refresh(8 β re-select every N decode steps),--sparse-attn-mode(0). Default off. - Sparse-V: skips V-dequant for negligible-weight positions
inside visited blocks. Self-configuring: on iSWA models with
quantized V it auto-sets Ο=0.05; elsewhere it stays off. Manual
override:
TURBO_SPARSE_V_TAU=<float>.
When to use: long context on quantized KV. The win grows with context (selectivity 0.91@16k β 0.99@40k and climbing) and lives on quantized KV: q8_0 β sparse at 50% coverage measured 1.34Γ decode at niah 100. Both levers stack (1.31Γ combined measured).
When not to use: short contexts or f16 KV on mid-size models β the decode isn't KV-bandwidth-bound there and the selector overhead can make it slower than dense. Ο values don't transfer across models; retune if you override manually.
5. Decode speed β MTP speculative decoding
Lossless speculative decode; the emitted text is the target model's
own (verified). Two flavours, chosen by --spec-type:
5.1 --spec-type draft-assistant (external drafter β Gemma-4)
A small gemma4-assistant drafter GGUF rides the target's
embeddings:
--spec-type draft-assistant --mtp-head gemma4-assistant-A4B-Q8_0.gguf \
-ngld 99 --spec-draft-n-max 2
--mtp-head (alias -md) names the drafter; -ngld 99 matters
(a CPU-resident draft head erases the win). Drafters for
A4B/12B/27B/E2B/E4B are published per-size. Setting mtpHead in the
TS adapter auto-derives the rest.
5.2 --spec-type draft-mtp (NextN self-spec β Qwen)
Qwen 3.5/3.6 GGUFs that embed a NextN/MTP head self-speculate β no second file:
--spec-type draft-mtp --spec-draft-n-max 3
Runs per-slot under --parallel (multi-session capable).
Measured (RTX 3090 + upstream-parity campaign): A4B assistant
decode beats upstream llama.cpp b9859 at every depth (+6.6/+12.9/+7.6%
at n_max 1/2/3); combined with turbo3_tcq KV it reaches ~89 tok/s
vs 52.9 plain (+69%). Qwen-35B NextN sits at parity with upstream.
Recommended depth: --spec-draft-n-max 2β3 (A4B), 3 (Qwen NextN).
Notes/limits: acceptance dips a few pp at depth β₯2 (chained-draft
numerics β expected); since c6 assistant-MTP runs on BOTH cache
layouts under --parallel (split-vs-unified outputs are
token-identical β the old --kv-unified requirement is retired).
Composes with
turbo/TCQ KV tiers (its biggest lever), DCA, and quantized KV. Watch
live acceptance per slot via GET /slots (Β§7).
5.3 Depth policy and sampler placement (auto-selected speculation)
Two policies sit on top of speculation and sampling; both shipped defaults were selected by measurement, so override them only with your own numbers:
--auto-mtp-policy(defaulttaper) governs draft depth when the speculation type was AUTO-selected (an explicit--spec-typealways drafts at full depth):taperdrafts at 1 live slot and switches off from 2 β speculation pays single-stream and costs batched throughput;allocatoris the measured marginal-value allocator;always/offare the escape hatches.--sampling-placement(defaultauto) picks where the sampler chain runs.autoroutes between device and CPU from a measured crossover (and is MTP-policy-aware: it accounts for whether a slot is drafting);device/cpuare authoritative when stated. This is distinct from--backend-sampling(on by default), which enables/disables the device sampling path itself, also per-request. Samplers the device chain cannot express decline to CPU correctly β placement is observable per-slot, not assumed.
6. Quality β RYS layer duplication
--repeat-layers re-runs a contiguous block of middle layers,
weight-shared: zero extra parameter VRAM, no new GGUF, quant-agnostic.
You pay in KV cache and tokens/s proportional to the extra effective
layers; you buy quality-per-token.
--repeat-layers 33,34 # +1 layer (RYS-S)
--repeat-layers 26-34 # +8 layers ([26,34) half-open, RYS-XL)
--repeat-layers 8-12;20-24 # disjoint blocks
Rules that matter:
- Middle layers only. Duplicating first/last layers reliably produces incoherent output on merge-fragile models β this is a model property, not an engine bug; the engine prints a boot advisory when a plan touches the boundary band.
- Absent flag = identity = byte-identical to stock.
- Composes with the full stack: quantized KV, DCA, rolling-KV window/spill, sparse-attn (the residency/DCA/sparse sizing paths are effective-plan-aware), and MTP β where the draft context deliberately runs the un-duplicated base stack while the target keeps RYS (still lossless: the target verifies every drafted token). Wired across all text archs (dense, MoE, Gemma-4 iSWA dual-cache, Qwen 3.5/3.6 recurrent-hybrid); unsupported archs fail loudly at load rather than silently ignoring the plan.
Finding a good plan: --rys-probe enumerates safe-band blocks,
scores each by ΞPPL + a task-probe battery, and prints two
ready-to-paste templates (most-efficient and max-gain):
sh ./opencoti-llamafile β¦ --rys-probe -m model.gguf -f corpus.txt \
--rys-probe-widths auto --rys-probe-topk 10
Treat its output as a shortlist to verify with your own eval, not a verdict.
7. Instrumentation β monitor & control API
Three planes (full reference: docs/features/introspection.md):
Boot knobs
Everything in Β§Β§3β6 is a boot flag: set at launch, echoed back at runtime. By design, tier/residency/DCA/retention cannot change per request (KV layout would differ).
Per-request control (JSON body fields)
| Field | Default | Effect |
|---|---|---|
session_id |
"" |
Sessionβslot affinity: the same session returns to the slot holding its KV (prevents cross-session eviction at --parallel > 1). Pair with cache_prompt: true. |
shared_pool_slot |
-1 |
Attach this request to SharedKVPool slot N (read-only prefix share, manual path Β§3.4). |
shared_prefix_n_tokens |
0 |
Length of the shared prefix (manual path Β§3.4). |
pool_id |
unset | Attach to a control-plane pool (Β§3.5); shared prefix P computed automatically, token-exact. Needs --polykv-max-pools. |
overcommit |
false |
Skip the enforced admission gate for THIS request (Β§3.5 P7) β explicit caller-controlled oversubscription past the pool's floor/target. |
Runtime introspection
GET /props β "opencoti" object β boot-state echo plus the
effective KV state read back from the live cache:
"opencoti": {
"kv": { "cache_type_k": "q8_0", "cache_type_v": "q4_0",
"auto_tier": false,
"effective": { "type_k": "q8_0", "type_v": "q4_0",
"n_cells": 524288, "n_cells_resident": 524288,
"n_layers_spilling": 0, "fully_resident": true,
"is_iswa": true } },
"residency": { "kv_residency_mode": 0, "vram_target_mib": 0 },
"dca": { "enabled": true, "chunk_size": 0, "yarn_factor": 1.0 },
"sparse_attn":{ "enabled": false, "block_size": 128, "topk": 0 },
"speculative":{ "types": ["none","draft-assistant"], "n_max": 3 },
"kv_reuse": { "n_parallel": 4, "kv_unified": true, "cache_ram_mib": 8192 },
"rest_kv": { "eviction": false, "recent": 256, "layer": -1 },
"repeat_layers": null
}
kv.effective is the only authoritative record of the auto-tier
decision β configured != effective is expected when auto-tier
engaged. fully_resident / n_layers_spilling tell you whether
rolling-KV is streaming.
GET /slots β per-slot "opencoti" object (requires --slots):
lifetime draft_n_total / draft_n_accepted / draft_acceptance
per slot, plus the slot's current session_id and pool binding.
Operational tell: sustained draft_acceptance β³ 0.95 at turn end
usually means the model is looping/ruminating (healthy agentic
decode sits ~0.4β0.9) β pollable, no log-scraping.
Per-completion timings: cache_n (prefix-reuse hits),
draft_n / draft_n_accepted for that response.
Quick recipes
curl -s :8080/props | jq .opencoti # what is this server running?
curl -s :8080/props | jq .opencoti.kv.effective # did auto-tier/spill engage?
curl -s :8080/slots | jq '.[] | {id, acc: .opencoti.draft_acceptance}'
For embedders/tools linking the C API:
llama_memory_opencoti_kv_info() (in llama.h) returns the same
effective-KV struct.
PolyKV control-plane telemetry (Β§3.5, requires
--polykv-max-pools): GET /polykv/tps (SSE or ?once=1) for
per-session tps + context budget; GET /polykv/pools/{id}/capacity
for admission headroom + compaction_pressure; per-slot
n_pool_shared in /slots; PolyKV gauges in /metrics.
GPU-share β cooperative peers on one GPU (c5,
docs/features/gpu_share.md): multiple opencoti-llamafile processes
on the same physical GPU auto-discover each other through a named
shared-memory registry (zero-conf, no ports; Linux/macOS/Windows).
--gpu-share-weight W (ratio, default 1.0) sets this instance's
compute share β e.g. weights 2, 1, 1 resolve to 50%/25%/25% β and
each instance duty-cycle paces its decode loop toward that share
only while another peer is actively decoding; a solo or
idle-peers instance always runs at full speed. Crashes age out in
3 s. GET /gpu/peers lists the live registry (pid, name, weight,
resolved share_pct, measured busy_pct, active,
heartbeat_age_ms); the same snapshot rides in /props under
opencoti.gpu_share. Caveat: the registry is keyed by device
description + ordinal, so peers must see the GPU under the same
device view (same CUDA_VISIBLE_DEVICES ordering).
Still log-only
SharedKVPool share/reject events, retention-eviction discards, rolling-KV tactic selection detail, and the auto-tier WARN line currently appear only in the server log.
8. Composition matrix
| quant-KV | auto-tier | rolling-KV | PolyKV pool | DCA | sparse-attn | MTP | RYS | |
|---|---|---|---|---|---|---|---|---|
| quant-KV | β | K-only honors | β (tiles dequant-on-lift) | β | β | β (the win case) | β (turbo+MTP is the top decode combo) | β |
| auto-tier | β | β (it manages spill) | β | β (probes in DCA state) | β | β | β (sizing is eff-plan-aware) | |
| rolling-KV | β | β | β | β | β | β (validated: window spill Γ RYS on hybrid) | ||
| PolyKV pool (SharedKVPool) | β | β | β | β | β
(validated: 2-agent share gate Γ --repeat-layers on A4B; hybrid-GDN omnimerge Γ NextN MTP full gate, patch 0135) |
|||
| DCA | β | β | β (dual-ctx) | β (effβsrc mapped) | ||||
| sparse-attn | β | β | β | |||||
| MTP | β | β (draft runs base stack; target keeps RYS) |
One guard worth restating: SharedKVPool requires
--no-cache-idle-slots (zero-copy with --kv-unified, copy-share
on the split cache since c6). Assistant-MTP works on both cache
layouts since c6 (the old --kv-unified forcing is retired), so
MTP + rolling window + pools compose on the split cache. The rolling
window by itself also runs under --kv-unified (Β§3.3, validated
2026-09-19 with two concurrent sessions); the three-way combination
has only been gated on the split layout.
Reference "agentic serving" launch (Gemma-4-A4B on a 24 GB card β quantized KV + MTP + introspection):
sh ./opencoti-llamafile-<ver>-<tag>-x86_64.llamafile --server --port 8080 \
-m gemma4-A4B-Q4_K_M.gguf -ngl 99 --flash-attn on \
-c 262144 --parallel 4 --kv-unified \
-ctk q8_0 -ctv q4_0 \
--spec-type draft-assistant --mtp-head gemma4-assistant-A4B-Q8_0.gguf \
-ngld 99 --spec-draft-n-max 2 \
--slots
9. Internal / superseded machinery (so you don't chase ghosts)
Present in the patch series but not user-facing knobs anymore:
- HeadInfer head-split (
--headinfer-gpu-heads-frac): retired as a manual knob; it survives as one tactic inside rolling-KV's auto ladder (autois the only value you should pass, and the adapter does it for you). - NEO GPU/CPU FA pipelining (
--neo-pipeline): structurally shipped, default off; no measurable win on single-GPU consumer hardware. Leave off. - Fused-MoE up-gate (
--fused-moe-up-gate): niche (+2.4% decode on OLMoE-class MoE; Gemma-4 already fuses). Default off. - Fused-NextN draft graph (
OPENCOTI_MTP_FUSED_NEXTN=1): built and shipped (patch 0093) but default off for a measured reason β on CUDA it decodes β7.0 to β11.8% slower on 4/4 NextN models than the default autoregressive draft loop (which, post-0128, is at upstream parity or better). Leave off. Corrected 2026-08-12: this entry used to blame a non-shape-invariant graph that "rebuilds every cycle". TheOPENCOTI_NEXTN_REUSE_TRACEcounter measures 1.7β1.8% miss (e.g. hit=1257 miss=23), i.e. the graph is reused ~98% of the time β the default-OFF verdict stands on the throughput measurement, but that mechanism claim was wrong and the real cause is still open. On Vulkan the picture splits (35B +11.7%, qwopus27 β3.3%), so it stays off there too pending a fuller grid. Seedocs/features/fused_nextn_mtp.md. - ScoutAttention, LMCache: design-only / deferred β the flags don't exist.
10. Verifying an artifact
Moved to docs/llamafile-artifacts.md
(hashes vs MANIFEST.json/SHA256SUMS, embedded-DSO verification
without execution, version-string check, the glibc floor contract).
11. Migrating between cuts
From c6 (or earlier) to the current cut β behaviour flips
The host binary surface is compatible, but three defaults changed and
one path moved. All are visible at boot (--version, --help, the
log) and reversible by flag:
- KV admission is ON by default β
--admission-poollessshipsenforced(Β§3.6): a request that cannot fit its measured KV demand now gets a fast HTTP 429 +Retry-Afterbefore prefill instead of failing or evicting after it. Clients that never handled 429s should either handle them (the header tells you when to retry) or launch with--admission-poolless warn/off. - Sampler placement defaults to
autoand routes (Β§5.3) β earlier cuts defaulted todevice(and for one cutautoonly observed). Pass--sampling-placement deviceto pin the old behaviour. - Auto-selected speculation tapers β
--auto-mtp-policyshipstaper(draft at 1 live slot, off from 2). An explicit--spec-typeis unaffected. Pass--auto-mtp-policy alwaysfor the old always-draft behaviour. - The GPU-payload directory is versioned by the full engine string
β
~/.llamafile/v/opencoti-<ver>-<tag>/, not~/.llamafile/v/<ver>/. Scripts that pre-seed or clean the old path must update (see llamafile-artifacts.md).
To a 0.10.5-based cut (upstream base bump)
The first cut on the upstream llamafile 0.10.5 base carries the same opencoti feature set (the whole patch series is ported), with these migration-relevant deltas:
- TurboQuant GGUF type IDs are renumbered. Upstream 0.10.5 claimed
numeric type slot 42, so the turbo family moved (TURBO2_0=43 β¦
TURBO8_0=48). Any GGUF quantized as
TQ3_1S/TQ4_1Sunder the old IDs will not load on a 0.10.5-based build β re-quantize it. KV-cache turbo tiers (-ctk/-ctv turbo*) are runtime-only and unaffected; saved session/state files that embed type ids are also invalidated. -fa autoresolution changed internally (a fused-ops registry replaced tensor-name parsing). Same flag surface, same tri-state; the boot log line to look for is nowresolve_fused_ops: Flash Attention enabled.- Everything else is surface-compatible; per-cut specifics live in
that release's
RELEASE_NOTES.md.