llama.cpp fork with TurboQuant quantization (turbo2/3/4) and TriAttention GPU-accelerated KV cache pruning. 75 tok/s on Qwen3-8B / RTX 3080.
-
Updated
Jul 2, 2026 - C++
llama.cpp fork with TurboQuant quantization (turbo2/3/4) and TriAttention GPU-accelerated KV cache pruning. 75 tok/s on Qwen3-8B / RTX 3080.
MSVC+CUDA llama.cpp fork for Windows: KDA/GDN, long-context KV placement, TurboQuant/TCQ when measured, speculative draft. Lab defaults — TriAttention opt-in only, not recommended.
Accelerate Kimi Delta Attention computations with high-performance CUTLASS kernels designed for NVIDIA SM90 architectures and beyond.
To associate your repository with the triattention topic, visit your repo's landing page and select "manage topics."