Skip to content
#

llm-serving-systems

Here are 2 public repositories matching this topic...

Run GGUF through llama.cpp and SafeTensors through vLLM behind one OpenAI-compatible endpoint. Your coding tools select a model; the switchboard manages the local runtime, process, and resident-model change.

  • Updated Sep 6, 2026
  • Rust

Add this topic to your repo

To associate your repository with the llm-serving-systems topic, visit your repo's landing page and select "manage topics."

Learn more