Convert, quantize, split, validate, and upload ML models for Apple MLX on Apple Silicon.
Tip: if you use Claude Code for MLX ports, the
mlx-portingskill wraps the end-to-end porting workflow (scaffolding, parity testing, attention patterns, pitfalls) and delegates weight conversion tomlx-forge.
- Convert PyTorch checkpoints to MLX format (safetensors, channels-last conv layout)
- Quantize model weights to int4/int8 with selective layer targeting
- Split large unified model files into per-component files for memory-constrained machines
- Validate converted models: file structure, key naming, weight shapes, quantization integrity
- Upload converted models to HuggingFace Hub with auto-derived repo naming, model cards, and collections
| Model | Recipe | Status |
|---|---|---|
| LTX-2.3 (22B video DiT) | ltx-2.3 |
Stable |
| Fish S2 Pro (5B TTS) | fish-s2-pro |
Stable |
| ERNIE-Image (8B text-to-image DiT) | ernie-image |
Stable |
| ERNIE-Image Prompt Enhancer (3B Ministral3 CausalLM) | ernie-image-pe |
Stable |
| V-JEPA 2.1 ViT-L (encoder + predictor, RoPE) | vjepa-2.1-vitl |
Stable |
| V-JEPA 2.0 ViT-L (encoder + predictor + SSv2/Diving48/EK100 probes) | vjepa-2.0-vitl |
Stable |
# From source
git clone https://github.com/dgrauet/mlx-forge.git
cd mlx-forge
pip install -e .
# Or with uv
uv pip install -e .Requires macOS with Apple Silicon and Python 3.11+.
For recipes that load PyTorch .pth checkpoints (e.g. Fish S2 Pro codec):
pip install 'mlx-forge[torch]'# Convert a model (downloads checkpoint from HuggingFace)
mlx-forge convert ltx-2.3
mlx-forge convert fish-s2-pro
mlx-forge convert ernie-image
# Convert with int8 quantization
mlx-forge convert ltx-2.3 --quantize --bits 8
# Preview conversion plan (no download, no writes)
mlx-forge convert fish-s2-pro --dry-run
# Convert from a local checkpoint
mlx-forge convert ltx-2.3 --checkpoint /path/to/checkpoint.safetensorsV-JEPA 2 recipes are different: the weights are Meta torch-hub .pt files
(not on HuggingFace), so they need an explicit local --source and the [torch]
extra. Output dirs are versioned (models/vjepa-2.X-vitl-mlx):
# V-JEPA 2.1 ViT-L (encoder + predictor) — single .pt
mlx-forge convert vjepa-2.1-vitl --source ~/weights/vjepa2_1_vitl_dist_vitG_384.pt
# V-JEPA 2.0 ViT-L (encoder + predictor + all three attentive probes)
mlx-forge convert vjepa-2.0-vitl \
--source ~/weights/vitl.pt \
--ssv2-source ~/weights/ssv2-vitl.pt \
--diving48-source ~/weights/diving48-vitl-256.pt \
--ek100-source ~/weights/ek100-vitl-256.ptConverted MLX weights are published at
dgrauet/vjepa-2.1-vitl-mlx
and dgrauet/vjepa-2.0-vitl-mlx.
See model-specific options in docs/models/.
mlx-forge validate ltx-2.3 models/ltx-2.3-mlx-distilled
mlx-forge validate fish-s2-pro models/fish-s2-pro-mlx
mlx-forge validate ernie-image models/ernie-image-mlxmlx-forge split ltx-2.3 /path/to/unified-model-dir# Upload with auto-derived repo name (reads split_model.json metadata)
mlx-forge upload models/ltx-2.3-mlx-distilled
# Upload to a specific repo or organization
mlx-forge upload models/fish-s2-pro-mlx --repo-id myuser/my-model
mlx-forge upload models/ernie-image-mlx --namespace my-org
# Upload and add to a collection
mlx-forge upload ./my-model --collection "MLX Forge Models"Requires authentication: run huggingface-cli login or set the HF_TOKEN environment variable.
# Quantize any safetensors file
mlx-forge quantize model.safetensors --bits 8
# Only quantize keys with a specific prefix
mlx-forge quantize model.safetensors --key-prefix transformer. --bits 4mlx_forge/
├── cli.py # CLI entry point
├── convert.py # Shared conversion utilities (download, load, classify, process)
├── transpose.py # Conv weight layout transposition (generic)
├── quantize.py # Quantization engine (generic)
├── split.py # Model splitting (generic)
├── validate.py # Validation framework (generic)
├── upload.py # HuggingFace Hub upload + model card (generic)
└── recipes/
├── ltx_23.py # LTX-2.3: key mapping, config, validation
├── fish_s2.py # Fish S2 Pro: Dual-AR TTS + DAC codec
└── ernie_image.py # ERNIE-Image: 8B single-stream text-to-image DiT
Generic tools live at the top level. Model-specific logic lives in recipes. Adding support for a new model means creating a new recipe file.
Create src/mlx_forge/recipes/my_model.py with:
def classify_key(key: str) -> str | None:
"""Map PyTorch key -> component name."""
...
def sanitize_key(key: str) -> str:
"""PyTorch key naming -> MLX key naming."""
...
def convert(args) -> None:
"""Main conversion entry point."""
...
def validate(args) -> None:
"""Model-specific validation."""
...
def add_convert_args(parser) -> None:
"""Register CLI arguments for convert."""
...
def add_validate_args(parser) -> None:
"""Register CLI arguments for validate."""
...Then register it in recipes/__init__.py:
AVAILABLE_RECIPES = {
"ltx-2.3": "mlx_forge.recipes.ltx_23",
"fish-s2-pro": "mlx_forge.recipes.fish_s2",
"ernie-image": "mlx_forge.recipes.ernie_image",
"my-model": "mlx_forge.recipes.my_model",
}PyTorch stores conv weights as (O, I, ...) while MLX expects channels-last (O, ..., I):
| Layer | PyTorch | MLX |
|---|---|---|
| Conv1d | (O, I, K) | (O, K, I) |
| Conv2d | (O, I, H, W) | (O, H, W, I) |
| Conv3d | (O, I, D, H, W) | (O, D, H, W, I) |
| ConvTranspose1d | (I, O, K) | (O, K, I) |
- Only Linear
.weightmatrices are quantized (affine mode with scales + biases) - Conv, norm, embedding, and other layers stay in original precision
- Critical: non-quantizable tensors must be materialized before
mx.quantize()runs, or lazy tensor buffers get evicted
- Source checkpoints are loaded lazily via
mx.load()(memory-mapped, ~0 GB initially) - Components are processed one at a time to stay within 32 GB RAM
- Explicit
gc.collect()+mx.clear_cache()between components - Each weight tensor is individually materialized before quantization to prevent OOM from accumulated lazy computation graphs
Each recipe has its own detailed guide with architecture, key mapping, known gotchas, and validation details:
- LTX-2.3 — 22B video DiT (6 components, Conv3d/Conv1d transposition)
- Fish S2 Pro — 5B TTS (Dual-AR + DAC codec)
- ERNIE-Image — 8B single-stream text-to-image DiT (+ separate 3B
ernie-image-pePrompt Enhancer recipe) - V-JEPA 2 — Meta video world model: ViT-L encoder + predictor, RoPE (2.1
vjepa-2.1-vitl/ 2.0vjepa-2.0-vitl+ attentive probes)
Apache 2.0