Qwen3.8 Flash Next FP8 on 4x CMP 170HX: PP3 transformer stages with the 51B PLE table on GPU 3
-
Updated
Aug 30, 2026 - Python
Qwen3.8 Flash Next FP8 on 4x CMP 170HX: PP3 transformer stages with the 51B PLE table on GPU 3
Experimental portable GGML GPU offload for diffuse-cpp across Vulkan, Metal, and CUDA
Proof repo: run a real LLM fully on-device in React Native (Expo + llama.rn). Offline, no cloud, sub-2s cold start.
To associate your repository with the gpu-offload topic, visit your repo's landing page and select "manage topics."