Capacity-aware single-GPU SGLang benchmarks with MTP A/B, core/extended context matrices, raw telemetry, model/KV memory capture, and generated Markdown/PDF reports.
-
Updated
Aug 1, 2026 - Python
Capacity-aware single-GPU SGLang benchmarks with MTP A/B, core/extended context matrices, raw telemetry, model/KV memory capture, and generated Markdown/PDF reports.
Evidence-led vLLM tuning and qualification notes for NVIDIA CMP 170HX (sm80)
GLM-5.3-Flash (320B MoE) served with vLLM on 8x CMP 170HX (SM80, 64 GB, PCIe Gen2 x4): pipeline-parallel 8, NVFP4 weights, DFlash2 speculative decoding, 1M context. Patches, launch config and benchmarks.
Running large LLMs on pre-Ampere NVIDIA hardware — Tesla V100 (sm_70), RTX 2080 Ti (sm_75), CMP 170HX. Measured benchmarks, vLLM forks, and the hardware side: NVLink on SXM2 carrier boards, driver traps, cooling, used-kit acceptance.
To associate your repository with the cmp-170hx topic, visit your repo's landing page and select "manage topics."