Reproducible on-device LLM benchmarks for Apple Silicon (iPhone 17 Pro, M4 Max): Apple Core AI, MLX, llama.cpp, LiteRT-LM and Core ML on the same model and harness, every number with its quantization and capture session; hybrid Mamba-2 models (Nemotron-3 Nano, Granite-4.0-H, Falcon-H1) included.
macos ios benchmark granite mlx coreml core-ai on-device-ai apple-silicon llm llama-cpp llm-inference nemotron tokens-per-second coreai m4-max apple-core-ai hybrid-mamba iphone-17-pro ios-27
-
Updated
Sep 5, 2026 - Python