Real-time multimodal AI pipelines on Apple Silicon
Download the Latest macOS App Release
Choose the largest file if you are not sure.
VPIPE is a small, embeddable runtime for building local AI applications where video, audio, images, text, tensors, user input, and tool actions move through the same inspectable pipeline. It is designed for Apple Silicon machines and runs its on-device generative stack on a custom metal-compute backend — its own Metal kernels, with no Python and no third-party tensor runtime in the forward pass.
-
Runs MiniMax H3 FL2VA and REF2VA on an Apple Silicon Mac — a 33B model generating video and its soundtrack together, on as little as 16 GB. Now with Turbo LoRA support! See docs/MINIMAX-H3.md
- 3.75s @ 0.5 MP 24p, 8 steps takes 13 minutes on a fanless 15-inch base-model M5 MacBook Air, 16 GB 1
-
Built in C++ for performance and compactness
-
Full modality support packed under 25 MB 2
-
Top-tier inference speed, enabling realtime visual question answering (VQA), automatic speech recognition (ASR), and language-model chat with text-to-speech (TTS)
-
Top-tier diffusion-transformer inference speed with weight streaming, enabling image and video edits on base-model systems with 16 GB of memory — walk through a reference image edit in docs/KLEIN-KV.md
-
Local MCP support: sandboxed file, shell, and Python tools, plus web fetch
-
Mobile-friendly UI enabling remote access from a phone
-
Extra acceleration from the NAX matmul2d and convolution2d units on M5-generation hardware
Install · Quickstart · First example · Overview · Examples
For developers: Requirements · Build from source · Run · Tests · Structure · Acknowledgements · License
Los Angeles Zoo, Nov 2019
SONY ILCE-7RM3, FE 135mm F1.8 GM @ 1/160s, F/1.8
Download the latest release ▸
Open the .dmg, drag Vpipe Manager to Applications, and launch it.
Requires an Apple Silicon Mac running macOS 26 or later.
Two builds are published. They are the same app; they differ only in whether FFmpeg travels with it:
| Download | Size | Pick this if |
|---|---|---|
VpipeManager-<version>-with-ffmpeg.dmg |
~26 MB | You want it to work immediately. Nothing else to install. |
VpipeManager-<version>-slim.dmg |
~15 MB | You already have FFmpeg installed — Homebrew's, say — and would rather use it. |
The bundled copy is a minimal LGPL build: it covers the common formats and keeps hardware H.264/HEVC/ProRes through VideoToolbox, but it has no libx264/x265. If you take the slim build, point the app at your own FFmpeg — Quickstart covers where.
The app is signed and notarized, so it opens without a Gatekeeper warning, and it updates itself: Vpipe Manager ▸ Check for Updates…
What you get. The app is a launcher around the same two binaries the command line uses — it does not reimplement anything. Its real advantage is permissions: Camera, Microphone and Local Network access are granted to the app under its own name, instead of to whichever terminal you happened to run from.
Everything also builds from source. The pipeline core and the non-Apple stages are portable C++20 and build on Linux and Intel macOS; the on-device model stack needs an Apple Silicon Mac.
git clone --recursive https://github.com/tgo-app-dev/vpipe.git
cd vpipe && cmake -S . -B build && cmake --build build -jThat is the whole story when the dependencies are already in place. For prerequisites, build options, the Metal toolchain and the rest, see Requirements and Build from source below.
Getting from a fresh install to a working browser UI. Everything here is in the app; the command-line equivalents are under Run.
1. Choose a work directory. Settings ▸ Work Directory ▸ Path ▸
Choose… — it defaults to ~/vpipe.
This is where everything lives: models/ for anything downloaded or
quantized, sandbox/ for files pipelines are allowed to write, and the
LMDB registry and logs. Pick a volume with room. Prepared models
run to tens of gigabytes — MiniMax H3 downloads ~115 GB and peaks near
~155 GB while it quantizes — and the app shows the free space on the
volume you select. Moving it later means moving all of that.
2. Point at FFmpeg — only if you took the slim build. Settings ▸ FFmpeg ▸ Library Path ▸ Choose…
Homebrew's is normally /opt/homebrew/lib. The row above the
button says whether a usable FFmpeg was found, so you are not guessing.
Use Default returns to the bundled copy. Either way the change
takes effect the next time you start a pipeline or the server — it is
not picked up by something already running.
3. Decide who can reach the web UI. Settings ▸ Web UI ▸ Bind To
- This Mac only (127.0.0.1) — nothing else on the network can connect, and no access key is needed. Start here.
- Automatic (this Mac's LAN address), or a specific interface — so a phone or another computer can connect.
Anything other than This Mac only makes the web UI reachable from your network, and anyone who reaches it can start and stop pipelines, browse the sandbox and drive models on this Mac. An 8-character access key is the only thing in front of that, so use a LAN binding only when you actually need another device, and not on networks you do not trust. The app warns you in place when you select one.
Port defaults to 9876; change it if something else is using it.
HTTPS (Self-Signed) is only needed for the low-latency Preview view
on another device — browsers restrict WebCodecs to secure contexts —
and costs a one-time certificate warning.
4. Start it. Open the Web UI pane and press Start Server, then Open in Browser.
If you chose a LAN address in step 3, the pane also shows a QR
code: point a phone camera at it and the UI opens already
authenticated, with no key to retype. If you kept This Mac only,
there is no QR code — a phone could not reach 127.0.0.1 anyway. The
access key is still shown, but this Mac connects without it.
Open Work Folder and Open Sandbox Folder reveal those
directories in Finder.
What the QR code contains. Only a URL:
http://<this Mac's LAN address>:<port>/<token>. The token is 14 random characters, generated fresh at every start and never written to disk. Nothing about you, your files, your models or your machine is encoded in it — the single identifying detail is the LAN address, and that is a private one like192.168.x.xwhich means nothing outside your own network.Treat it as a password anyway, because the token is not merely information. Scanning it redirects to
/?key=<access key>and hands the access key over, so anyone who can both see the code and reach that address gets exactly the control you have: starting and stopping pipelines, browsing the sandbox, driving models on this Mac. A photo, a screenshot or a screen share is enough for someone on the same network. Restarting the server invalidates it.
The status row along the bottom shows the machine's thermal state. Sustained image or video generation heat-soaks a Mac, a fanless MacBook Air especially. When it reads Throttling, steps are taking longer because of the hardware, not because something has stalled.
Next: the first example, below.
The shortest path from a working install to a model answering you: two pipeline files, ~7.7 GB on disk, and no conversion step to sit through — this checkpoint arrives already quantized.
- Get both pipelines —
prepare-qwen35-9b-optiq-4bit.vpipelineandqwen35-9b-chat.vpipeline(Raw ▸ Save as, or straight fromdocs/pipelines/in a clone). - In the web UI from the Quickstart, open the Pipeline Manager, Load
the
prepare-one and press Start. It downloads the model into your work directory and registers it there. Once, and never again. - Load the chat pipeline, Start it, open the User I/O panel — and type
at the
you>prompt.
you> In one sentence, what is the Pacific Ocean?
The Pacific Ocean is the largest and deepest ocean on Earth, covering more
than 30% of the planet's surface area and separating the continents of Asia
and Australia to the east from North and South America to the west.
Five stages: a text input, the chat stage, a sampler stage carrying the values this checkpoint recommends for itself, and a feedback pair that makes it turn-by-turn. You can attach a picture to a turn — this model reads images too.
The same pipeline from a phone — scan the QR code the server prints and the browser opens already authenticated. Here a photographed receipt is transcribed, then questioned: the follow-up answers from the same context, and the Mac decodes at ~19.7 tok/s throughout.
From a source build the same two files run on the command line, which is where a chat prompt is most at home:
cd ~/vpipe # your work directory
./build/apps/vpipe/vpipe --launch prepare-qwen35-9b-optiq-4bit.vpipeline
./build/apps/vpipe/vpipe --launch qwen35-9b-chat.vpipeline▸ docs/QWEN35-CHAT.md — the walkthrough: what each stage is for, why sampling is a stage rather than a config key, and the knobs worth knowing.
Then: EXAMPLES.md builds the same chat by hand in the web UI, and adds speech transcription. For image editing from a reference photo, docs/KLEIN-KV.md; for text-to-video with sound, docs/MINIMAX-H3.md. Each ships the pipelines it describes.
VPIPE has three main surfaces:
-
Pipeline core — coroutine-based
Jobstages connected by buffered ports, driven by a runtime that launches and drains them concurrently. Stages are composed into a pipeline from a JSON spec; each stage registers under a type name (e.g.rtsp-capture,video-to-rgb,yolo-detection,onvif-discovery,rest-client). This layer is portable C++20. -
On-device generative-model stack (Apple Silicon) — a from-scratch LLM/VLM/ASR/diffusion/video inference stack running on metal-compute, with custom kernels, model loading, quantization support, weight streaming, and resource planning for memory-constrained Macs. It powers stages such as
text-chat,visual-qa,realtime-vqa,audio-transcribe,generate-image,diffusion-conditioner,vae-encode,vae-decode, andgenerate-video. -
Web UI and Composer — a self-contained browser UI for launching, inspecting, profiling, and editing pipelines. The Composer can arrange pipeline editors, previews, image comparison views, text I/O, profiler views, files, logs, and stage-provided panels, then save that layout with the pipeline spec so a workflow can be reopened and reproduced.
Models are loaded from local directories (sharded safetensors / GGUF) and are not bundled with the source.
Platform note. The pipeline core and the non-Apple stages build on Linux and Intel macOS, but the generative-model stack and the CoreML/Metal stages require an Apple Silicon Mac. On arm64 macOS these features are detected and enabled automatically.
Everything above describes the app. The rest of this file is the source tree: what it needs to build, how to drive the same session from the command line, from Python or from a test binary, and where things live.
-
CMake ≥ 3.25 and a C++20 compiler (Apple Clang or a recent Clang/GCC).
-
Git (the build pulls a few dependencies as submodules).
-
FFmpeg development headers —
libavformat,libavcodec,libavutil,libswresample. VPIPE compiles against the headers anddlopens the libraries at runtime, so FFmpeg must also be installed at runtime to decode media.- macOS:
brew install ffmpeg - Debian/Ubuntu:
apt install libavformat-dev libavcodec-dev libavutil-dev libswresample-dev
- macOS:
-
libcurl — used by the
rest-clientstage. Provided by the SDK on macOS; on Linux install e.g.apt install libcurl4-openssl-dev. -
Python 3 + development headers — only for the optional Python extension (built by default). Disable with
-DVPIPE_BUILD_PYTHON=OFFif you don't need it. -
Apple Silicon Mac — for the on-device inference stack and CoreML/Metal stages.
-
Metal shader toolchain (Apple Silicon builds only; recommended, not required) — by default the build compiles
.metalkernel sources into embedded metallibs usingxcrun -sdk macosx metalandxcrun -sdk macosx metallib. These compilers are part of Xcode's Metal Toolchain; the standalone Command Line Tools do not include them, even though the rest of the build never opens Xcode.No toolchain? The build falls back automatically. If
metal/metallibaren't found at configure time, the build switches to runtime-compile mode: it embeds the Metal shader source and compiles each kernel on first use via the OS's built-in runtime compiler (newLibraryWithSource:), which needs no toolchain on the build or run machine. The Metal Toolchain is therefore optional; the tradeoff is a one-time per-kernel compile on first use instead of at build time. Force either mode with-DVPIPE_METAL_RUNTIME_COMPILE=ON|OFF.To get the faster build-time (AOT) path, install the toolchain. Two independent things can leave
metal/metallibunavailable — both surface the same way, aserror: cannot execute tool 'metal'orxcrun: error: unable to find utility "metal":1. The Metal Toolchain isn't installed. On Xcode 26 and later (macOS 26) the Metal Toolchain is no longer bundled with Xcode by default — it's an optional component you download once. Install it from the command line (or via Xcode ▸ Settings ▸ Components ▸ Metal Toolchain ▸ Get):
xcodebuild -downloadComponent metalToolchain # download + installOn air-gapped or CI machines, export once and import where needed:
xcodebuild -downloadComponent metalToolchain -exportPath ~/Downloads xcodebuild -importComponent metalToolchain ~/Downloads/metalToolchain.dmg
2.
xcrunpoints at the Command Line Tools, not Xcode. If you installed the CLT and then Xcode,xcrunoften still resolves to the standalone CLT, which lacksmetal/metallib. Point the toolchain at Xcode:# Direct xcrun at the Xcode app (run once; needs admin) sudo xcode-select --switch /Applications/Xcode.app/Contents/Developer sudo xcodebuild -license accept # accept the license if you haven't
Verify the active developer dir and that both compilers resolve:
xcode-select -p # -> /Applications/Xcode.app/Contents/Developer xcrun -sdk macosx -f metal # -> a path inside Xcode.app (not /Library/Developer/CommandLineTools) xcrun -sdk macosx -f metallib # -> likewise xcrun -sdk macosx metal --version
To switch back to the Command Line Tools later:
sudo xcode-select --switch /Library/Developer/CommandLineTools.
1. Fetch dependencies. LMDB and pugixml are always required; nanobind is needed for the Python bindings, and metal-cpp for the Apple Silicon features:
git submodule update --init extern/lmdb extern/pugixml extern/nanobind extern/metal-cpp2. Configure (out-of-source build directory):
cmake -S . -B buildThis defaults to an optimized Release build, so the binaries you get are performant out of the box. Override the build type explicitly if you want a debug build:
cmake -S . -B build -DCMAKE_BUILD_TYPE=Debug3. Build:
cmake --build build -jcmake --build is generator-agnostic; use it rather than calling make
directly so the build works regardless of which generator CMake selected.
Useful options (pass at configure time with -D):
| Option | Default | Effect |
|---|---|---|
VPIPE_BUILD_PYTHON |
ON |
Build the vpipe Python extension (needs Python dev headers). |
VPIPE_BUILD_APPLE_SILICON |
auto (on for arm64 macOS) | Build the CoreML/metal-compute wrappers and the inference stages. |
CMAKE_BUILD_TYPE |
Release (when unset) |
Set Debug for an unoptimized debug build. |
Install (optional):
cmake --install build --prefix /path/to/installvpipe-web-ui serves a browser-based Pipeline Manager bound to one VPIPE
session. The web assets are embedded in the binary, so no extra files are
needed:
./build/apps/web-ui/vpipe-web-ui # listens on the LAN address, port 9876
./build/apps/web-ui/vpipe-web-ui --bind 127.0.0.1 # this machine onlyThen open the printed URL (e.g. http://localhost:9876). By default it binds
to the machine's LAN address so other devices can connect; remote connections
must supply the 8-character access key printed at startup, while localhost
connects without one. Options: --bind ADDR, --port N (0 = any free port),
--config CFG (inline JSON, a file path, or empty for defaults), --help.
The UI has a phone layout, and typing an 8-character key into a phone is
exactly the friction that stops anyone from using it. --show-qr prints a
QR code to the console next to the usual startup lines:
./build/apps/web-ui/vpipe-web-ui --show-qrPoint a phone camera at it and the UI opens already authenticated — no key
to read off the screen and retype. The phone layout is selected automatically
from the device; ?ui=desktop (or the drawer's Desktop layout) overrides it,
and ?ui=phone is how that layout is developed on a desktop.
How it works, and what it costs:
- Two different secrets. The 8-character access key is short because a human retypes it. The QR link carries its own, longer secret (14 characters of an uppercase alphanumeric alphabet, ~70 bits), because nothing has to read it — it only has to be unguessable.
- One scan, then the key is gone.
GET /<link-secret>redirects to/?key=<access key>; the page adopts the key intosessionStorageand immediately strips it from the URL, so it never lands in the address bar, the history, or a bookmark. The key is tab-scoped and disappears when the tab closes. - The link is never printed. Only the symbol is rendered — writing the URL beside it would put the secret into the scrollback, a screen share, or a terminal log, which is what the QR code exists to avoid. It is a secret in a URL: treat it as only as private as the console displaying it, and restart the server to invalidate it.
- Both secrets are per-run, generated at startup, and neither is written to disk.
Requests from other computers carry the key as an X-Auth-Key header (or a
?key= parameter where the browser cannot set headers, such as a WebSocket
handshake or an <img> source). Only /api/* is gated, and only for
non-loopback peers — static assets stay open so a remote browser can load the
page in order to ask for the key in the first place.
Note. Add
--tlsif you want the low-latency Preview view on a phone or any other LAN client: the browser's WebCodecs API is secure-context only, so it needs HTTPS off localhost. The certificate is self-signed and cached under~/.vpipe/webui-tls, so expect a one-time browser warning — and a QR scan lands on that warning rather than the UI until it is accepted.
vpipe launches pipelines straight from the terminal — a thin command-line
front end over the same session and stages the web UI drives. It dynamically
links libvpipe.
# A full pipeline from a saved spec file, or from inline JSON:
./build/apps/vpipe/vpipe --launch my-pipeline.vpipeline
./build/apps/vpipe/vpipe --launch '{"id":"tick","stages":[{"id":"c","type":"chrono","config":{"count":5}}]}'
# A single stage wrapped in a one-shot pipeline (handy for utility stages):
./build/apps/vpipe/vpipe --launch-stage onvif-discovery
./build/apps/vpipe/vpipe --launch-stage model-fetch \
--stage-cfg model_path=mlx-community/Qwen3.5-4B-MLX-4bit--stage-cfg overrides stage config: key=value after --launch-stage, or
stage-id::key=value to target a stage inside a --launch spec. Repeat
--launch / --launch-stage to run several pipelines concurrently;
Ctrl-C stops them cleanly. vpipe --help lists every option, and
EXAMPLES.md shows fetching a model from the terminal.
The extension lands in build/python/. Importing the package creates a default
session:
PYTHONPATH=build/python python3 -c "import vpipe; print(vpipe.vpipe_version())"Startup configuration is resolved from VPIPE_CONFIG / VPIPE_CONFIG_FILE, an
./init.vpipe file, or built-in defaults. Call vpipe.create_session(config=...)
to make your own session.
The build produces a unit-test executable:
./build/vpipe_test # run everything
./build/vpipe_test --list_tests
./build/vpipe_test --filter '<pattern>' # supports * and ? wildcards
./build/vpipe_test --color off # for captured/non-interactive outputSome tests exercise real models and are gated on environment variables that point at local model directories; when a variable is unset, the corresponding test skips.
| Path | Contents |
|---|---|
pipeline/, common/, interfaces/, include/ |
Pipeline core: jobs, ports, runtime, session, shared services. |
stages/ |
Pipeline stages (capture, decode, detection, REST, the LLM/VLM stages, …). |
generative-models/ |
On-device LLM/VLM/ASR stack (model families, tokenizers, encoders). |
apple-silicon/ |
metal-compute backend and CoreML C++ wrappers. |
gpu-kernels/metal/ |
Metal compute kernels (attention, GEMM, quant, …). |
apps/ |
Executables: vpipe (CLI), web-ui, db-log-reader. |
python/ |
Python bindings (nanobind). |
tests/ |
Unit tests. |
extern/, 3rd-party/ |
Vendored dependencies. |
VPIPE builds on these projects:
- FFmpeg — the multimedia framework vpipe uses to
decode and encode audio and video; compiled against its headers and
dlopened at runtime. LGPL-2.1-or-later (some optional components are GPL). - LMDB — the memory-mapped key-value store behind vpipe's databases: logs, the model registry, camera records. OpenLDAP Public License 2.8.
- pugixml — a light XML parser, used for the SOAP and WS-Discovery exchanges that find ONVIF cameras. MIT.
- nanobind — the C++/Python
binding layer the
vpipePython extension is built with. BSD 3-Clause. - metal-cpp — header-only C++
bindings for Apple's Objective-C runtime; vpipe uses its
Foundationheaders under the CoreML and metal-compute wrappers. Apache-2.0. - pocketfft — a header-only FFT, used by the audio feature extractors to build mel spectrograms. BSD 3-Clause.
MLX (MIT) is Apple's array
framework for machine learning on Apple Silicon. VPIPE does not link MLX and
does not use it in the forward pass, but it does vendor a small set of MLX's
Metal kernel headers — the "steel" GEMM and attention templates and their
supporting helpers — under
gpu-kernels/metal/vendored/mlx/backend/metal/kernels/steel. Those headers
are #included by vpipe's own .metal sources and compiled into the
embedded metallibs.
Full copyright and license texts for everything bundled or vendored are in
THIRD_PARTY_LICENSES.md.
VPIPE is licensed under the Apache License, Version 2.0 — see
LICENSE. Bundled and vendored third-party components and their
licenses are documented in
THIRD_PARTY_LICENSES.md.
Brought to you by T-Go LLC, registered in California.




