Skip to content
#

m4-max

Here are 5 public repositories matching this topic...

Language: All
Filter by language

Reproducible on-device LLM benchmarks for Apple Silicon (iPhone 17 Pro, M4 Max): Apple Core AI, MLX, llama.cpp, LiteRT-LM and Core ML on the same model and harness, every number with its quantization and capture session; hybrid Mamba-2 models (Nemotron-3 Nano, Granite-4.0-H, Falcon-H1) included.

  • Updated Sep 5, 2026
  • Python

Cursor-Auto / Claude-tier-style serving for local GGUF models on Mac (M4 Max, 64 GB). FastAPI router fronts llama-swap + llama.cpp, classifying each request into a coder, planner, or uncensored-planner tier. OpenAI-compatible API, opencode integration, per-project subshell, one `llmstack` console-script.

  • Updated Aug 17, 2026
  • Python

Add this topic to your repo

To associate your repository with the m4-max topic, visit your repo's landing page and select "manage topics."

Learn more