Move model weights into GPU memory and reusable artifacts into filesystem caches — through P2P RDMA, streaming, GPUDirect Storage, or host-staged POSIX I/O.
Features • Architecture • Benchmarks • Quick Start • Deployment • Docs • Contributing
ModelExpress (MX) starts with a simple question: before loading a model, where does a compatible copy of its weights already live? Rather than treating every replica as an independent cold start, ModelExpress discovers available sources and selects the fastest supported path into GPU memory. Deploy it standalone or alongside vLLM, SGLang, NVIDIA Dynamo, and other inference runtimes.
| LLM serving problem | How ModelExpress helps |
|---|---|
| Models take too long to load | GPU-to-GPU transfer via NIXL/RDMA instead of loading from storage. In P2P mode, weights already serving inference act as the cache—no extra storage. |
| JIT warmup dominates startup | Compatible vLLM and SGLang NIXL TorchInductor, Triton, DeepGEMM, TileLang, CuTe DSL, and FlashInfer JIT caches transfer from a ready replica instead of being rebuilt. |
| Many nodes need the same model | Metadata backends (Redis, K8s CRD) coordinate sharing: one node loads; others receive via P2P or local paths. |
Today, ModelExpress loads DeepSeek-V4-Pro weights from a serving replica in 11 seconds—a 48× speedup over a cold Hugging Face pull. Reusing compatible weights and JIT kernel-cache artifacts reduces process-start-to-API-ready time from 8 minutes 1 second to 1 minute 44 seconds (4.6×).
At startup, ModelExpress probes the capabilities available in the environment and tries loading strategies in priority order:
- Serving peer → GPU — Transfer post-processed weights directly from a compatible replica over NIXL P2P RDMA. Each new replica then joins the source pool, turning scale-out into GPU-to-GPU fan-out.
- Remote or local storage → GPU with ModelStreamer — Fetch safetensor ranges concurrently through a bounded CPU staging buffer while overlapping reads with GPU placement. Tensor-parallel ranks can divide remote reads instead of each downloading the full checkpoint.
- Local storage → GPU with GDS — Use NIXL's multithreaded GPUDirect Storage backend to bypass host-memory staging when the platform supports it.
- Default loader — Fall back to the inference engine's host-staged POSIX I/O path.
The first applicable strategy runs. If a strategy fails before changing model state, ModelExpress continues to the next one. If weights may already have landed, it reinitializes the model before continuing so a partially loaded model is never served.
The ModelExpress control plane discovers compatible sources through Redis, Kubernetes CRDs, or the decentralized k8s-service backend. Weight bytes stay on the data plane and move directly between storage or source and target. For clusters with a shared disk cache, the Model Cache Service also coordinates a single download so concurrent replicas reuse one cached checkpoint instead of multiplying external ingress.
- Cold start reduction — GPU-to-GPU P2P transfer over InfiniBand instead of disk load
- Capability-driven loading — Automatic priority chain: P2P RDMA → ModelStreamer → GDS → native loader, with safe fallback
- HuggingFace caching — PVC-backed cache,
HF_HUB_OFFLINE,ignore_weights,get_model_pathfor Dynamo - P2P GPU transfer — vLLM
modelexpressloader, SGLangremote_instanceloader with themodelexpressbackend, and TRT-LLMPRESHARDEDloader with NVIDIA NIXL over InfiniBand, RoCE, NVLink, EFA, and other supported fabrics - JIT cache transfer — Reuse compatible vLLM and SGLang NIXL compilation caches when replicas scale out
- Metadata backends — Redis, Kubernetes CRD, or decentralized Kubernetes Service routing
- Kubernetes — Helm chart, CRDs/Redis for P2P, no-shared-storage support
- CLI — Health, download, list, validate, clear; init-container support for pre-warming
- ModelStreamer integration — Pipeline concurrent reads from S3, Azure Blob, GCS, or local storage into vLLM and SGLang
- Expanded model pull providers: NGC catalog and Google Cloud Storage in addition to Hugging Face
- GDS (GPUDirect Storage): load model weights directly from NVMe into GPU memory, bypassing the CPU/DRAM copy path
- Lower NIXL registration overhead — Opt in to allocation-level pool registration or a single VMM arena registration
| Runtime | Integration |
|---|---|
| vLLM | Native --load-format modelexpress in 0.23.0+ for P2P weight and JIT cache transfer; older versions use the ModelExpress plugin |
| SGLang | remote_instance + modelexpress backend with transport=nixl or transport=transfer_engine — see docs/SGLANG.md |
| TensorRT-LLM | LoadFormat.PRESHARDED with MxLiveCheckpointLoader for P2P weight transfer (beta) — TRT-LLM examples |
| NVIDIA Dynamo vLLM runtime | --load-format modelexpress for P2P weight and JIT cache transfer — Dynamo P2P example |
| NVIDIA Dynamo SGLang runtime | remote_instance + modelexpress backend for P2P weight transfer; transport=nixl also supports JIT cache transfer — see docs/SGLANG.md |
The ModelExpress server brokers metadata only. Weight bytes move directly from a compatible serving peer, object storage, or file storage into the new inference engine.
Phase 1 — Bootstrap once: The seed pod selects the fastest available storage path, loads and post-processes the weights, registers GPU memory with NIXL, and publishes metadata. Phase 2 — Scale out: Compatible pods discover a serving peer through the control plane and receive weights directly over NIXL GPUDirect RDMA; the ModelExpress server never handles the weight bytes.
- modelexpress_server: gRPC server with configurable metadata backends (Redis, Kubernetes CRD).
- modelexpress_client: Rust CLI for cache management; Python package with inference engine loaders and
MxClientfor gRPC. - modelexpress_common: Protobuf definitions, model-provider traits, and shared configuration.
See Architecture.
The following results use DeepSeek-V4-Pro with vLLM 0.23.0 and TP=8 on an 8×B200 GPU node with NVIDIA ConnectX-7 NICs. Runs used --enable-flashinfer-autotune; timings are end-to-end startup measurements from the benchmark environment and will vary with storage, network, and runtime configuration.
| Loading path | Time | Speedup vs. cold Hugging Face pull |
|---|---|---|
| Cold pull from Hugging Face | 8m 53s | 1× |
| ModelStreamer from S3 | 3m 16s | 2.7× |
| High-throughput local storage, cold page cache | 1m 10s | 7.6× |
| P2P GPU-to-GPU over NIXL/RDMA | 11s | 48× |
| Registration strategy | Time | Speedup |
|---|---|---|
| Per tensor (default) | 8.16s | 1× |
Pool registration (MX_POOL_REG=1) |
1.14s | 7.1× |
VMM arena (MX_VMM_ARENA=1) |
0.79s | 10.3× |
Pool registration and VMM arena registration are alternatives; enable only one.
| Startup path | API ready | Speedup |
|---|---|---|
| Cold start from VAST, no P2P source | 8m 1s | 1× |
| P2P RDMA weights only | 7m | 1.1× |
| P2P RDMA weights and kernel artifacts | 1m 44s | 4.6× |
The artifact-enabled run reused compatible Triton, DeepGEMM, TileLang, CuTe DSL, and FlashInfer caches. ModelExpress transfers these file-backed artifacts between registered host-memory buffers, verifies them, and installs them into the target engine's filesystem cache; they are not loaded into GPU memory.
Requirements: vLLM 0.23.0+, the ModelExpress Python package, NIXL-compatible GPU nodes, and a reachable metadata backend.
git clone https://github.com/ai-dynamo/modelexpress.git
cd modelexpress
python -m pip install ./modelexpress_client/python
export MX_SERVER_ADDRESS=modelexpress-server:8001
export MX_P2P_METADATA=1
export MX_ARTIFACT_TRANSFER=1
vllm serve deepseek-ai/DeepSeek-V4-Pro \
--load-format modelexpress \
--tensor-parallel-size 8 \
--trust-remote-codeStart the first replica with weights available through local storage or MX_MODEL_URI. After it becomes healthy, launch the same command on another compatible node: ModelExpress discovers the serving replica and transfers its post-processed weights directly over NIXL P2P RDMA.
JIT artifact transfer requires MX_P2P_METADATA=1, a central-coordinator backend (redis or kubernetes), and writable target staging and runtime cache directories. With those prerequisites, MX_ARTIFACT_TRANSFER=1 installs compatible artifacts into the new replica's filesystem caches. Omit it for weight-only transfer or when using the decentralized k8s-service backend.
See the Kubernetes P2P example for metadata-server and inference-worker manifests.
kubectl create secret generic hf-token-secret --from-literal=HF_TOKEN=${HF_TOKEN} -n <namespace>
helm install modelexpress ./helm --namespace modelexpress --create-namespaceOverride values-production.yaml for your env. Full config: helm/README.md.
vllm serve deepseek-ai/DeepSeek-V4-Pro \
--load-format modelexpress \
--tensor-parallel-size 8 \
--trust-remote-codevLLM 0.23.0 recognizes the load format natively; the ModelExpress Python package must still be installed in the runtime image. The first instance loads from disk, while subsequent instances receive weights via RDMA. Set MX_ARTIFACT_TRANSFER=1 to transfer compatible JIT caches as well. P2P guide · Server setup.
Set MX_MODEL_URI to an s3://, gs://, or az:// URI or an absolute local path. For tensor-parallel deployments, participating ranks divide remote reads via MX_MS_DISTRIBUTED (on by default; set to 0 to disable); TP=1 ignores the setting. ModelStreamer examples · vLLM recipes.
docker compose -f docker/docker-compose.yml up --buildPrecedence: CLI → env vars (MODEL_EXPRESS_*, MX_*) → YAML → defaults.
| Variable | Default | Description |
|---|---|---|
MODEL_EXPRESS_SERVER_PORT |
8001 |
gRPC port |
MODEL_EXPRESS_CACHE_DIRECTORY |
./cache |
Cache root |
MX_METADATA_BACKEND |
Server: required; client: unset | Server: redis | kubernetes. Client: unset, server, redis, or kubernetes uses a central coordinator; k8s-service enables decentralized Kubernetes Service routing. |
REDIS_URL |
(required for redis) |
Redis connection URL. Alternatively set MX_REDIS_HOST + MX_REDIS_PORT. No localhost fallback. |
MX_SERVER_ADDRESS |
localhost:8001 |
Client-side gRPC server address (P2P). Recommended. |
MODEL_EXPRESS_URL |
localhost:8001 |
Deprecated, pending removal in a future release. Still read by all client paths and takes precedence when both are set; keep setting it during the transition. |
MX_MODEL_URI |
(unset) | Enable ModelStreamer for an object-store URI or absolute local path. |
MX_MS_DISTRIBUTED |
1 |
Divide ModelStreamer reads across tensor-parallel ranks when TP > 1. On by default; set to 0 to disable. |
MX_POOL_REG |
0 |
Register each underlying CUDA allocation once instead of registering every tensor. |
MX_VMM_ARENA |
0 |
Load into a CUDA VMM arena and register the used range once; alternative to MX_POOL_REG. |
MX_P2P_SOURCE_SELECTOR |
random |
Peer ordering policy; set rendezvous_hash for deterministic, minimally disruptive selection. |
cargo run --bin config_gen -- --output model-express.yaml
cargo run --bin modelexpress-server -- --config model-express.yaml --validate-configFull reference: docs/DEPLOYMENT.md.
modelexpress-cli health
modelexpress-cli model download <model-id>
modelexpress-cli model list
modelexpress-cli model validate <model-id>
modelexpress-cli model clear <model-id>cargo test
cargo test --test integration_tests
cargo run --bin test_client -- --test-model "google-t5/t5-small"
./run_integration_tests.sh
cargo bench| Doc | Description |
|---|---|
| Deployment | Server/client config, Docker, K8s, P2P |
| Architecture | Components, gRPC, NIXL, FP8 |
| Benchmarks | Loading paths, NIXL registration, and artifact-transfer results |
| CLI | Full CLI reference |
| Metadata | Redis keys, K8s CRD schema |
| Helm | Kubernetes configuration |
- GDS loader does not scale with TP — Each TP rank reads full checkpoint tensors and vLLM shards them afterward, so GDS/disk reads scale with TP degree. This can reduce or reverse expected GDS speedups versus the default mmap-based disk loader; TP-aware range reads are needed for a full fix. See GDS Reads Full Checkpoint Tensors Under TP.
- DRAM and NVMe-resident shard streaming: Stream shards across workers while keeping weights in DRAM and host local high-speed NVMe.
- RL post-training refit: Make updates receiver-driven—trainer ranks publish the shards they own, rollout workers discover and plan against their target layout, then pull, convert, reshard, and load directly over NIXL.
- Earlier weight availability: Bring weights to prefill earlier; identify prefill workers that can act as strong source nodes.
- Multi-tier cache hierarchy: Promote and demote models across DRAM, NVMe, and PVC tiers based on access patterns.
- Distributed sharded cache: Shard large models across nodes using consistent hashing and parallel shard assembly.
- Training checkpoint management: Cache and reuse CUDA kernel compilations (torch.compile, deepGEMM) and CUDA graphs across restarts.
- Metrics and observability: Cache hit rates, eviction frequency, transfer throughput, and P2P RDMA utilization via Prometheus/OpenTelemetry.
- Predictive prefetching: Pre-warm caches from workload history or scheduling hints.
- Dynamic EPLB (Expert Parallelism Load Balancer): Rebalance MoE expert placement across GPUs at runtime via P2P transfer of expert weights as load shifts.
Contributions welcome. See CONTRIBUTING.md.
pip install pre-commit && pre-commit install
pre-commit run --all-filesIssues: GitHub Issues
Apache 2.0. See LICENSE.





