You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
Qwen3.8-27B on one laptop CPU: up to 2.52 token/s, 8 GB tested, no runtime accuracy loss. Native OpenAI-compatible API with function tools; no GPU or Python. | 单颗笔记本 CPU 运行 Qwen3.8-27B:最快 2.52 token/s,最低 8 GB 内存可运行,推理加速不牺牲准确性。原生 OpenAI 兼容接口支持函数工具;无需 GPU 或 Python。
Native macOS control center and local AI agent gateway for Qwen3.8 on Apple Silicon — MLX, DFlash2, OpenAI Responses, Anthropic Messages, Claude Code, Codex, OpenCode and Grok Build.
Qwen 3.8 is LIVE NOW! Can It Survive 3 Brutal Tests? (Qwen 3.8 Max Benchmarks) - Technical guide, 2.4T parameter specifications, token pricing, and Canvas execution test prompts.
Run Qwen3.8-27B (NVFP4 4-bit) locally on an RTX 5090 — one click on Windows 11, one command on Linux. Ships a coding agent, an OpenAI-compatible API, and three serving backends. Private, unmetered, offline.
Run Qwen3.8-27B locally in Claude Code Desktop with vision, tools, native reasoning and Claude Code CLI support. Validated on an NVIDIA RTX 5090 32 GB.