Local inference platform for K/IQ-quant GGUF models on Apple Silicon
-
Updated
Aug 7, 2026 - Python
Local inference platform for K/IQ-quant GGUF models on Apple Silicon
a GGUF inference runner in Rust
GGUF-Runner - Want to run LLMs locally, use this guide, and run with LLAMA.cpp
A complete inference runtime for open-weight large language models, enabling efficient execution through streaming weights, quantization, and memory-aware scheduling
Render is a lightweight, easy-to-use CLI and local server for generating AI images. Run Stable Diffusion (SD 1.5, SDXL) and FLUX models locally with simple commands like pull and run. Powered by Vulkan acceleration and GGUF support for ultra-fast performance. Think Ollama, but for local image generation.
Add a description, image, and links to the gguf-runner topic page so that developers can more easily learn about it.
To associate your repository with the gguf-runner topic, visit your repo's landing page and select "manage topics."