Run GGUF through llama.cpp and SafeTensors through vLLM behind one OpenAI-compatible endpoint. Your coding tools select a model; the switchboard manages the local runtime, process, and resident-model change.
openai llama-cpp vllm local-llm llm-inference llm-serve gguf gguf-models vllm-serve openai-compatible gguf-model-support model-swap llm-service gguf-manager vllm-server model-swapping llm-serving-systems gguf-model gguf-runner models-switcher
-
Updated
Sep 6, 2026 - Rust