From 4cf2cee07d5a5b266df534ab56bea52dbc2460d1 Mon Sep 17 00:00:00 2001 From: Chao Wang <26245345+ChaoWao@users.noreply.github.com> Date: Sun, 26 Jul 2026 08:12:26 -0700 Subject: [PATCH] Add: investigation on the host worker dispatch latency budget MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit After #1499 took a single `Worker.run()` from 150 us to ~51 us, the obvious follow-up was whether another factor of three was available in what remained, and whether a blocking wakeup primitive (#1498) was how to get it. Measured, and the answer to both is no. Recording it so the next person does not re-derive it. A bare fork+shm mailbox round trip with no simpler runtime in it costs 3.5 us against `Worker.run()`'s 53.8 us, which reads as "94% of dispatch latency is our code, not the IPC". That headline is an artifact of benchmarking a single task. Sweeping tasks per `run()` from 1 to 256 splits it into **43.3 us fixed per run() and 8.08 us marginal per task**: the cost is entering and leaving an orchestration, not dispatching work, and a production run submits a graph per `run()` rather than one task. At 256 tasks the fixed part is 2% of the invocation. The same numbers settle #1498 against itself on latency grounds. A pipe wake measures 4.4 us one way on this box, so replacing the poll with a blocking wait adds more than the entire 3.5 us IPC floor to every dispatch, in exchange for CPU, while the latency it was meant to recover is not in the poll at all. What survives is narrower and unrelated: a child spinning while nothing is outstanding holds a core for as long as its parent lives — a CPU question for narrow-core hosts, already smaller since #1495 made orphans exit with their parent. Also records the reconsideration triggers, since both conclusions are conditional: a workload that issues many small `run()` calls stops amortizing the fixed cost, and the CPU argument for a spin-then-block hybrid remains open. Co-Authored-By: Claude Opus 5 --- .../2026-07-host-dispatch-latency-budget.md | 109 ++++++++++++++++++ docs/investigations/README.md | 1 + 2 files changed, 110 insertions(+) create mode 100644 docs/investigations/2026-07-host-dispatch-latency-budget.md diff --git a/docs/investigations/2026-07-host-dispatch-latency-budget.md b/docs/investigations/2026-07-host-dispatch-latency-budget.md new file mode 100644 index 000000000..b03155002 --- /dev/null +++ b/docs/investigations/2026-07-host-dispatch-latency-budget.md @@ -0,0 +1,109 @@ +# Host worker dispatch latency: where the remaining ~50 us goes + +**Date**: 2026-07-26 +**Verdict**: dropped — the cost is fixed per `run()`, not per task, and +amortizes to 2% at realistic graph sizes. Also closes the case for replacing +the mailbox poll with a blocking wakeup primitive: the poll is not where the +latency is. + +## Question + +[#1499](https://github.com/hw-native-sys/simpler/pull/1499) removed the parent's +50 us `TASK_DONE` sleep and took a single `Worker.run()` from 150 us to ~51 us. +The obvious follow-up: is there another factor of three in the remaining 51 us, +and is a blocking wakeup primitive ([#1498](https://github.com/hw-native-sys/simpler/issues/1498)) +the way to get it? + +## What was tried + +Four measurements on aarch64 Linux, 1 sub worker, no-op callables, medians over +200-500 iterations. + +**1. The IPC floor.** `tests/ut/py/test_hostsub_fork_shm.py::_SubWorkerPool` is a +minimal fork + shared-memory mailbox with no simpler runtime in it at all, so a +round trip through it is what the mechanism costs before any of our code runs. + +**2. A full `Worker.run()`** with one sub task, for the difference. + +**3. `cProfile` over 500 dispatches** to see the shape of the runtime's own work. + +**4. Tasks per `run()` swept 1 → 256**, to separate fixed from marginal cost. +This is the measurement that decides the question; the first three only frame it. + +## Result + +```text +bare mailbox round trip (fork + shm, no runtime): 3.5 us +Worker.run(), one sub task: 53.8 us + => runtime overhead per dispatch: 50.3 us (94%) +``` + +94% sitting in our code rather than the IPC looks damning, and taken alone it +argues for optimising the runtime and against #1498. The sweep shows why that +reading is wrong: + +| tasks per `run()` | total | per task | +| ----------------- | ----- | -------- | +| 1 | 51.4 us | 51.38 us | +| 2 | 62.7 us | 31.35 us | +| 4 | 80.2 us | 20.04 us | +| 8 | 114.4 us | 14.30 us | +| 16 | 177.0 us | 11.06 us | +| 64 | 538.7 us | 8.42 us | +| 256 | 2113.0 us | 8.25 us | + +**Fixed: 43.3 us per `run()`. Marginal: 8.08 us per task.** + +The 50 us is almost entirely the cost of *entering and leaving an +orchestration*, not of dispatching work. `cProfile` agrees on the shape: a +no-op dispatch runs ~113 Python calls and ~15 threading lock/notify operations, +spread thin across the submit → wait → drain plumbing with no single hot frame. + +## Why not (now) + +**The fixed cost amortizes to nothing at realistic sizes.** A production run +submits a graph per `run()`, not one task — at 256 tasks the fixed 43 us is 2% +of the invocation. Optimising it would pay off only for a workload that makes +many tiny `run()` calls, and we do not have one. Spending effort on the thin, +call-count-bound Python plumbing to win 43 us that a real graph already hides is +the wrong trade. + +**The marginal cost is already close to the floor.** 8.08 us per task against a +3.5 us bare fork+shm round trip is ~2.3x, and that 2.3x buys arg marshalling, +callable-identity resolution and per-task error reporting that the bare pool +does not do. + +**And it settles #1498 against itself.** A pipe wake measures **4.4 us one way** +on this box (8.85 us round trip, p99 9.62 us). Replacing the poll with a +blocking wait would therefore add more than the entire 3.5 us IPC floor to every +dispatch, in exchange for CPU, while 94% of the latency is somewhere else +entirely. The polling is not the problem, so removing the polling is not the +fix. + +What survives from #1498 is narrower and unrelated to latency: a child spinning +while *nothing* is outstanding holds a core for as long as its parent lives. +That is a CPU-efficiency question for narrow-core hosts, and it shrank further +once [#1495](https://github.com/hw-native-sys/simpler/pull/1495) made orphans +exit with their parent. + +## When to reconsider + +- A workload appears that issues many small `run()` calls — deep L3/L4 + hierarchies, or per-token orchestration — where 43 us fixed stops amortizing. + Measure that workload's tasks-per-`run()` first; if it is above ~16 the fixed + cost is already under 25%. +- The marginal 8.08 us becomes the bottleneck for a SubWorker-heavy graph. Note + this is the Python-callable path; NPU tasks do not go through it. +- Someone proposes a blocking primitive again for **latency** reasons. The + numbers above say it costs latency. Only the CPU argument is open, and only + in a spin-then-block hybrid that never reaches the blocking path while work + is in flight. + +## References + +- [#1499](https://github.com/hw-native-sys/simpler/pull/1499) — removed the + parent's dispatch-path sleep; the measurements here start from its result. +- [#1498](https://github.com/hw-native-sys/simpler/issues/1498) — blocking + wakeup primitive; re-scoped to CPU-only on the strength of these numbers. +- [`.claude/rules/codestyle.md`](../../.claude/rules/codestyle.md) rule 5 — the + dispatch path may spin or block, but never sleep. diff --git a/docs/investigations/README.md b/docs/investigations/README.md index 581adb83c..70d2ad53d 100644 --- a/docs/investigations/README.md +++ b/docs/investigations/README.md @@ -84,6 +84,7 @@ that ...". Newest first. +- [2026-07 — Host worker dispatch latency: where the remaining ~50 µs goes](2026-07-host-dispatch-latency-budget.md) — measured & dropped: after #1499 a `Worker.run()` costs **43.3 µs fixed + 8.08 µs per task**, so the "94% of latency is our runtime, not the IPC" headline is an artifact of benchmarking a *single* task — at 256 tasks/run the fixed cost is 2%. Also settles #1498 against itself for latency: a pipe wake measures 4.4 µs one way against a 3.5 µs bare fork+shm round trip, so blocking *adds* more than the whole IPC floor to every dispatch; only the CPU-while-idle argument survives - [2026-07 — Containing A2/A3 SDMA stream teardown after an AICore fault](2026-07-a2a3-sdma-fault-teardown.md) — CANN exposes no remote-channel retirement fence: even device-confirmed stream frees followed by a soft-reset success that did not complete context teardown can add minutes to a later ordinary fault. Contained by making SDMA an explicit per-Worker opt-in (`enable_sdma`) so only opt-in Workers carry the slow teardown and ordinary Workers keep fast recovery; full closure needs CANN runtime + driver changes (#1425) - [2026-07 — AICore-side arg fill for ALL dispatches (not just the gated path)](2026-07-aicore-fills-all-args-ready-path.md) — measured & rejected for the ready path: making the AICore fill its own `args[]` for every task adds ~1.0 µs to each task's `receive→start` setup (paged_attention_unroll: 349 ns → 1356 ns), because a ready task has no idle doorbell gate to hide the fill. Offload is a win only on `not_ready` (early-dispatch), where the AICore already spins at the gate; shipped design keeps the AICPU filling ready tasks (#1328) - [2026-07 — L2 swimlane AICore: switch-overhead source + FIN-early reorder & ACK-gate](2026-07-aicore-swimlane-switch-overhead-and-ack-gate.md) — measured: the ~0.8 µs inter-task switch is the record write-back `dcci(record,OUT)+dsb` (~0.5 µs) + payload setup (~0.28 µs), inherent and not reducible by moving FIN; the WAIT gap (p99 ~700 µs) dominates decode. Shipped: sample `end_time` after an early FIN, and an AICPU ACK-gate on buffer rotation (release the old buffer only when AICore ACKs the new buffer's first task) to close the FIN-before-record boundary race the reorder introduced