This is a custom inference-server fork built on top of llama.cpp. The custom work in this repository is focused on:
- proving PFlash prompt-token compression can be routed through the llama-server task path;
- implementing KVFlash resident-prefix admission, idle-slot KV eviction, page-directory recall, and hidden-state restore for repeated prompt prefixes;
- providing an OpenAI-compatible smart router for selecting direct, PFlash, and verification paths across llama-server backends.
DFlash speculative decoding is also used by this fork, but it came with the llama.cpp base. The work here does not claim to invent DFlash. It uses the upstream DFlash/speculative path as one serving mode to compare against and combine with the PFlash, KVFlash, and router experiments.
PFlash is the prompt-side compression path added by this fork. It lets the server reduce a long prompt token stream before full prefill, then run the shorter prompt through normal llama-server execution.
Current capabilities:
- Adds request/server controls such as
--pflash-mode,--pflash-keep-ratio,--pflash-drafter,--pflash-score, and--pflash-model. - Runs only when enabled, so default llama-server behavior stays unchanged when PFlash is off.
- Applies at the server task boundary for completion and infill style requests.
- Preserves prompt order and validates compressed token streams before accepting them.
- Supports a built-in deterministic uniform fallback for first/last-preserving token thinning.
- Supports an external helper bridge through
--pflash-drafter, allowing a scorer/helper to choose the kept token subsequence. - Supports the later model-score path where
--pflash-score modelloads a configured--pflash-modelas an in-process scorer context. - Exposes behavior through logs and benchmark scripts so compressed and uncompressed paths can be compared.
What PFlash is for:
- long-context broad analysis where exact retention of every prompt token is less important than reducing prefill cost;
- experiments that compare quality and speed at different keep ratios;
- two-pass routes where a fast compressed first pass can be followed by direct verification when exactness signals are detected.
Limitations:
- PFlash is experimental and must be benchmarked per workload.
- Aggressive compression can drop facts. Exact extraction, code patching, schema output, and all/every/list style prompts should use a direct path or a PFlash-then-direct-verify path.
KVFlash is the KV-cache-side experiment added by this fork. It now moves beyond accounting-only scaffolding: configured prompts are admitted into a resident-prefix budget, idle slot KV state can be evicted when new candidates need space, and evicted prefixes are retained in a page directory for later recall and optional state restoration.
Current capabilities:
- Adds configuration fields such as
--kvflash,--kvflash-policy,--kvflash-tau, and--kvflash-drafter. - Reports KVFlash configuration in
/props. - Maintains a live
resident_poolstatus object with capacity, resident tokens/pages, hit/miss counters, admission counters, eviction counters, mutation counters, hidden-prefix descriptors, and hidden-state restore counters. - Runs an admission pass at the server task boundary for completion and infill prompts.
- Commits successfully-prefilled prompts into the resident pool, bounded by
--kvflash. - Evicts idle slot KV state when resident-token pressure exceeds the configured pool.
- Records evicted prefix pages so repeated future prompts can be detected as page-directory recall hits.
- Captures hidden sequence state for eligible text-only prompts and restores matching hidden prefixes by default on later requests. Set
LLAMA_KVFLASH_HIDDEN_STATE_RESTORE=0to disable this restore path. - Integrates with PFlash hints so compressed prompts can advertise the resident span that should be tracked by KVFlash.
What KVFlash changes in server behavior:
- It can clear real llama.cpp KV state for idle slots when admitting new resident candidates.
- It can restore a matching hidden prefix from saved sequence state and skip reprocessing that restored prefix.
- It does not allocate a separate low-level KV tensor arena; residency is implemented at the server slot and sequence-state layer on top of llama.cpp memory APIs.
- It does not rewrite attention kernels or sparse attention masks. The page directory is used for recall/accounting and state restore, not for a custom attention kernel.
What KVFlash is for:
- reducing repeated-prefix prefill work across requests that share a resident prefix;
- forcing idle-slot KV reclamation under a configured resident-token budget;
- measuring page-level recall, hidden-state restore, and eviction pressure from a live server.
DFlash/speculative decoding is not claimed as custom work from this repository. The DFlash path came from the llama.cpp base that this fork builds on.
How this fork uses it:
- Runs DFlash-compatible drafter/target launches through existing llama.cpp speculative decoding controls, for example
-md,--spec-type draft-dflash, and--spec-draft-n-maxwhere supported by the base. - Uses DFlash as a comparison and composition mode when evaluating PFlash and KVFlash behavior.
- Records DFlash-related launch and benchmark evidence in status and benchmark artifacts.
- Lets the smart router and benchmark scripts compare direct, PFlash, PFlash-plus-verification, and DFlash-enabled server configurations.
What this repo proved around DFlash is integration evidence with the forked server stack, not ownership of the DFlash mechanism itself.
tools/server/qwen36-smart-router.py is custom glue for running multiple server paths behind an OpenAI-style API surface.
Current capabilities:
- Accepts OpenAI-style chat completion requests.
- Estimates prompt size and detects exactness signals such as patches, schemas, exact values, JSON output, all/every/list extraction, and code edits.
- Routes short or exact work to the direct path.
- Routes long broad-analysis work to a PFlash path.
- Routes long exact work to a PFlash-first path followed by direct verification when needed.
- Supports forced routes through request fields such as
route,x_route, orsmart_route. - Supports optional classifier sidecar modes for shadow or active routing experiments.
- Logs routing decisions to JSONL for analysis.
- Can manage or target separate direct and PFlash backend URLs through environment variables.
What the router is for:
- keeping ordinary OpenAI-compatible clients pointed at one endpoint while different llama-server backends run underneath;
- protecting exactness-sensitive tasks from unsafe compression;
- collecting routing evidence for PFlash and direct-verify policies.
The router is not a replacement for llama-server. It is policy glue around one or more llama-server instances.
Custom-work entry points:
tools/server/qwen36-smart-router.py- OpenAI-compatible smart router.tools/server/bench_cp0*.py- benchmark and validation scripts used during PFlash/KVFlash/router experiments.PORT_STATUS.md- checkpoint notes describing what each custom integration slice does and does not do.checkpoint-*.patch- checkpoint patches and implementation history.NOTICE.md- attribution, license, and research-lineage notes.tools/server/- llama-server base code plus the custom integration surface used by this fork.
Build it as a llama.cpp fork, then run the server configurations needed for the experiment being tested.
The EVO-X2 validation stack uses these local GGUF artifacts as the canonical model roles for this fork:
| Role | Local path | Notes |
|---|---|---|
| Target/base model | /home/hawg/models/Qwen3.6-27B-MTP-GGUF-Q4_K_M/Qwen3.6-27B-Q4_K_M.gguf |
Qwen3.6 27B Q4_K_M MTP target. Use this as -m. |
| DFlash drafter | /home/hawg/models/z-lab-Qwen3.6-27B-DFlash/qwen3.6-27b-dflash-zlab-q8_0.gguf |
Z-Lab Qwen3.6 27B DFlash drafter. Use this as -md with --spec-type draft-dflash. Do not use it as the target model. |
| PFlash scorer | /home/hawg/models/draft/Qwen3.5-0.8B-Q4_K_M.gguf |
Qwen3.5 0.8B Q4_K_M scorer model. Use this as --pflash-model with --pflash-score model. |
Use the MTP-directory Qwen3.6 27B Q4_K_M artifact above as the canonical target for this server's DFlash/PFlash validation. The older non-MTP experiment copy at /home/hawg/experiments/lucebox-campaign-20260627/models/target-unsloth/Qwen3.6-27B-Q4_K_M.gguf is not the ideal target for this server.
Typical flow:
- Build the repository for the target hardware.
- Start a direct llama-server backend.
- Start a PFlash-enabled backend when testing prompt compression.
- Configure KVFlash flags when testing resident-prefix admission, idle-slot eviction, and hidden-state restore.
- Use DFlash/speculative flags only as supported by the inherited llama.cpp base.
- Run the smart router when OpenAI-compatible clients need one endpoint that chooses among those paths.
- Compare logs,
/props, benchmark JSON, routing decisions, and output quality.
Exact flags depend on the backend, model, and experiment. Treat PFlash and KVFlash as experimental controls that need measurement on the target hardware. Treat DFlash as inherited llama.cpp speculative decoding that this fork can use in experiments.
This fork is built on top of llama.cpp and ggml. The base runtime, hardware backends, model loading, build system, server foundation, examples, upstream DFlash/speculative decoding support, and upstream documentation are credited to the llama.cpp and ggml authors and remain under the upstream MIT license.
Special thanks and credit go to Lucebox for the PFlash and KVFlash research/prototyping lineage that informed the prompt-compression and KV-cache-residency work in this repository. The custom proof work here is PFlash integration, KVFlash resident-prefix admission and recall, and the OpenAI-compatible router layered on top of llama.cpp.
See NOTICE.md for attribution and license details.
This repository preserves the upstream llama.cpp MIT license. See LICENSE and NOTICE.md.