Validated GLM-5.3 Flash recipe for 2x NVIDIA RTX PRO 6000 Blackwell 96GB: 262K context, EXL3/TR3, adaptive MTP, tools, and vision.
-
Updated
Aug 28, 2026 - Shell
Validated GLM-5.3 Flash recipe for 2x NVIDIA RTX PRO 6000 Blackwell 96GB: 262K context, EXL3/TR3, adaptive MTP, tools, and vision.
GLM-5.3-Flash EXL3 (320B MoE) on 2x NVIDIA DGX Spark — production serving kit, 1M context, 97%+ multi-session prefix caching, DFlash2 spec decode
LLM inference server for ExLlamaV3 / EXL3, with OpenAI- and Anthropic-compatible APIs optimized for Agent workloads.
Containerized private AI lab: TabbyAPI EXL3 + SillyTavern + Open WebUI + Ollama + SearXNG
GLM-5.2 753B MoE at EXL3-TR3 3.0bpw on 4x RTX PRO 6000 (SM120): digest-pinned serving stack, validation suite, results. Upstream: vllm#139 / sparkinfer#49
To associate your repository with the exl3 topic, visit your repo's landing page and select "manage topics."