Run Qwen3.6-35B-A3B full native 258K context on 12GB VRAM (llama.cpp): ncmoe cliff rule, q4_0 KV prefill fix, 4 tuned profiles with scripts
-
Updated
Aug 11, 2026 - PowerShell
Run Qwen3.6-35B-A3B full native 258K context on 12GB VRAM (llama.cpp): ncmoe cliff rule, q4_0 KV prefill fix, 4 tuned profiles with scripts
Run Qwen3.6-35B LLM + ComfyUI SDXL/Pony image gen simultaneously on 12GB VRAM — measured config (36-37 tok/s during render), scripts included
Run full 258K native context of Qwen3.6-35B-A3B MoE on 12GB VRAM with llama.cpp; benchmarked profiles, scripts, and tuning guide included.
To associate your repository with the rtx4070-super topic, visit your repo's landing page and select "manage topics."