Capacity-aware single-GPU SGLang benchmarks with MTP A/B, core/extended context matrices, raw telemetry, model/KV memory capture, and generated Markdown/PDF reports.
-
Updated
Aug 1, 2026 - Python
Capacity-aware single-GPU SGLang benchmarks with MTP A/B, core/extended context matrices, raw telemetry, model/KV memory capture, and generated Markdown/PDF reports.
Evidence-led vLLM tuning and qualification notes for NVIDIA CMP 170HX (sm80)
GLM-5.3-Flash (320B MoE) served with vLLM on 8x CMP 170HX (SM80, 64 GB, PCIe Gen2 x4): pipeline-parallel 8, NVFP4 weights, DFlash2 speculative decoding, 1M context. Patches, launch config and benchmarks.
Strict-FP32 CUDA SGEMM with shape-dispatched SM80 cp.async kernels and reproducible A100 evidence
Native dual-GPU long-context inference runtime for Qwen3.8-Flash-Next
To associate your repository with the sm80 topic, visit your repo's landing page and select "manage topics."