Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions docs/07-rpc-namespace-design.md
Original file line number Diff line number Diff line change
Expand Up @@ -905,7 +905,7 @@ interface cow-api {
/// A non-2xx reply with no typed rejection envelope; `body` is raw text.
record http-failure { status: u16, body: option<string> }

/// A typed orderbook rejection, parsed host-side from `{errorType, description}`.
/// A typed orderbook rejection, parsed host-side from `{errorType, description, data}`.
record order-rejection { status: u16, error-type: string, description: string, data: option<string> }

/// A shared host `fault`, a raw HTTP failure, or a typed order rejection.
Expand Down Expand Up @@ -970,7 +970,7 @@ impl shepherd::cow::cow_api::Host for NexumHostState {
// A non-2xx with no typed rejection envelope surfaces as `http`
// so a caller matches on `status` (e.g. 404) and reads `body`
// only for diagnostics; submit-order parses the orderbook's
// `{errorType, description}` into the `rejected` case instead.
// `{errorType, description, data}` into the `rejected` case instead.
let body = resp.text().await.ok();
return Ok(Err(CowApiError::Http(HttpFailure { status, body })));
}
Expand Down
2 changes: 1 addition & 1 deletion docs/adr/0011-per-interface-typed-errors.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ status: accepted

The host once returned one unified envelope, `host-error`, from every imported function and every module export: a record of `domain` (a stringly subsystem tag), a `host-error-kind` enum, a numeric `code` (a JSON-RPC code or an HTTP status, depending on the caller), a `message`, and an optional opaque `data` blob. Dispatch on the failure cause meant reading `kind`, sometimes cross-checking `domain` and `code`. Modules re-named their own domain in every error they built and prefixed each message with the module name, duplicating context the runtime already had (the module name, the interface).

The envelope conflated two things that want to move independently: the shared cross-domain failure vocabulary (unavailable, timeout, denied, ...) and the per-interface structured detail (a JSON-RPC revert carries a node code and decoded revert bytes; an orderbook rejection carries a typed `{errorType, description}`). Squeezing both through one flat record meant every interface paid for fields it did not use and lost the fields it did.
The envelope conflated two things that want to move independently: the shared cross-domain failure vocabulary (unavailable, timeout, denied, ...) and the per-interface structured detail (a JSON-RPC revert carries a node code and decoded revert bytes; an orderbook rejection carries a typed `{errorType, description, data}`). Squeezing both through one flat record meant every interface paid for fields it did not use and lost the fields it did.

## Decision

Expand Down
14 changes: 14 additions & 0 deletions docs/migration/0.1-to-0.2.md
Original file line number Diff line number Diff line change
Expand Up @@ -167,6 +167,20 @@ interface chain {
}
```

`cow-api-error` adds two cases: a raw `http-failure` for a non-2xx reply with
no typed rejection envelope, and a parsed `order-rejection` carrying the
orderbook's `errorType`/`description` plus its optional structured `data`
payload (re-encoded as a JSON string) so a caller does not have to
re-decode the body itself.

```wit
interface cow-api {
record http-failure { status: u16, body: option<string> }
record order-rejection { status: u16, error-type: string, description: string, data: option<string> }
variant cow-api-error { fault(fault), http(http-failure), rejected(order-rejection) }
}
```

### Author migration

The stringly `domain`/`code` cross-check is gone; dispatch on the typed variant instead.
Expand Down
44 changes: 22 additions & 22 deletions docs/production.md
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
# Production deployment guide

Operator handbook for running `nexum` (Shepherd) in
production. Focused on **concrete artefacts** unit files,
backup recipes, alert rules not the design rationale, which
production. Focused on **concrete artefacts** - unit files,
backup recipes, alert rules - not the design rationale, which
lives in `docs/06-production-hardening.md` (resource enforcement,
restart policy, RPC resilience, logging + metrics design).

Expand All @@ -28,12 +28,12 @@ Before launching:
- [ ] **`engine.toml`** (the production-shape config) exists with:
- `[engine] state_dir = "/var/lib/shepherd"` (or equivalent
persistent path; never `/tmp`).
- `[engine] log_level = "info"` (NOT debug see §5).
- `[engine] log_level = "info"` (NOT debug - see §5).
- `[engine.metrics] enabled = true` and `bind_addr` on
`127.0.0.1:9100` (NOT `0.0.0.0` see §7).
`127.0.0.1:9100` (NOT `0.0.0.0` - see §7).
- One `[chains.<id>]` entry per chain you intend to
subscribe to, with a **paid** WS URL (Alchemy / Infura /
QuickNode public nodes will throttle under sustained
QuickNode - public nodes will throttle under sustained
load, see §6).
- One `[[modules]]` entry per module to load.
- [ ] **`/var/lib/shepherd`** exists, writable by the engine's
Expand All @@ -42,8 +42,8 @@ Before launching:
- [ ] **A Prometheus instance** scraping the engine's `/metrics`
endpoint (§7) and an alert pipeline pointed at the rules in §9.
- [ ] **A log aggregator** ingesting the engine's JSON stdout
(§5) stdout, not a file written by the engine.
- [ ] **An on-call runbook reference** link to this document
(§5) - stdout, not a file written by the engine.
- [ ] **An on-call runbook reference** - link to this document
and to `docs/operations/m3-testnet-runbook.md` (testnet
validation, useful for staging deploys).

Expand All @@ -70,12 +70,12 @@ WorkingDirectory=/opt/shepherd
ExecStart=/opt/shepherd/bin/nexum \
--engine-config /etc/shepherd/engine.toml

# Graceful shutdown engine handles SIGINT/SIGTERM by:
# Graceful shutdown - engine handles SIGINT/SIGTERM by:
# 1. closing chain subscription tasks,
# 2. finishing the in-flight dispatch,
# 3. writing `last_dispatched_block:{chain_id}` to local-store,
# 4. logging `graceful shutdown complete ...` and exiting 0.
# Give it 30 s production runs can have ~5 s of in-flight RPC.
# Give it 30 s - production runs can have ~5 s of in-flight RPC.
KillSignal=SIGINT
TimeoutStopSec=30s

Expand All @@ -91,14 +91,14 @@ RestrictAddressFamilies=AF_INET AF_INET6 AF_UNIX
LockPersonality=true
MemoryDenyWriteExecute=false # wasmtime JIT requires writable+executable memory pages

# Restart policy supervisor handles per-module poison/restart
# Restart policy - supervisor handles per-module poison/restart
# itself, but if the host process exits non-zero (panic, OOM,
# etc.) restart after 5 s. RestartSec=0 would loop fast on
# config errors.
Restart=on-failure
RestartSec=5s

# Resource caps (defence in depth wasmtime is already capping
# Resource caps (defence in depth - wasmtime is already capping
# per-module memory at 64 MiB and fuel at ~1B inst/event).
LimitNOFILE=65536
MemoryMax=2G
Expand Down Expand Up @@ -193,7 +193,7 @@ services:
- shepherd-state:/var/lib/shepherd
- ./engine.toml:/etc/shepherd/engine.toml:ro
ports:
# Bind metrics endpoint to the host loopback only
# Bind metrics endpoint to the host loopback only -
# Prometheus scrapes it via docker network, no public
# exposure.
- "127.0.0.1:9100:9100"
Expand Down Expand Up @@ -286,7 +286,7 @@ The pause-and-copy window is ≤ 1 s on a ~100 MiB local-store
the alloy WS connection survives a brief process stop.

A `redb::Database::backup`-style API (snapshot from within a
read transaction) is on the roadmap track in upstream redb
read transaction) is on the roadmap - track in upstream redb
releases > 2.6.

### 4.3 Restore + integrity check
Expand Down Expand Up @@ -343,16 +343,16 @@ Important fields on every event:
| `target` | Crate + module path. Useful filters: `nexum_runtime`, `nexum_runtime::supervisor`, `nexum_runtime::host::impls::cow_api`. |
| `level` | `TRACE` < `DEBUG` < `INFO` < `WARN` < `ERROR`. **Production should never see `ERROR`** from `nexum_runtime::*` (only from third-party crates the supervisor wraps as warnings). |
| `fields.message` | Human-readable summary. Greppable. |
| `fields.module` | Set on every per-module event supervisor, host calls, guest log emissions. Use this for per-module dashboards. |
| `fields.module` | Set on every per-module event - supervisor, host calls, guest log emissions. Use this for per-module dashboards. |

### 5.2 Retention + aggregation

Two-tier model:

1. **Hot (last 7 days)** full INFO + DEBUG. Lives in your
1. **Hot (last 7 days)** - full INFO + DEBUG. Lives in your
log aggregator (Loki / CloudWatch Logs / Datadog). Used for
incident investigation.
2. **Cold (90 days)** INFO only, drop DEBUG at ingest time.
2. **Cold (90 days)** - INFO only, drop DEBUG at ingest time.
S3 / GCS with lifecycle rule to Glacier at 90 days. Used for
audit + post-mortem.

Expand Down Expand Up @@ -439,7 +439,7 @@ Capacity sizing (per chain):
number of registered orders; budget accordingly.

`shepherd_chain_request_total{outcome="err"}` rate is the
canonical "the RPC is degraded" signal see §9 alerts.
canonical "the RPC is degraded" signal - see §9 alerts.

---

Expand All @@ -459,7 +459,7 @@ network.
| `shepherd_module_restarts_total` | counter | `module` | Increments on every `reinstantiate_one` attempt (per-module restart backoff). |
| `shepherd_module_poisoned` | gauge | `module` | `1` if the module has been quarantined per `POISON_MAX_FAILURES=5` / `POISON_WINDOW=10m`. Stays `1` until process restart. |
| `shepherd_chain_request_total` | counter | `chain_id`, `method`, `outcome` | Every `chain::request` host call. `outcome="err"` rate > 5% = RPC degraded. |
| `shepherd_cow_api_submit_total` | counter | `chain_id`, `outcome` | Every orderbook submit. `outcome="err"` covers both retriable and dropped drill into supervisor logs to discriminate. |
| `shepherd_cow_api_submit_total` | counter | `chain_id`, `outcome` | Every orderbook submit. `outcome="err"` covers both retriable and dropped - drill into supervisor logs to discriminate. |
| `shepherd_stream_reconnects_total` | counter | `kind`, `chain_id`, `module?` | WS reconnect attempts. `kind="block"` is per-chain; `kind="chain-log"` carries the `module` label too. |

### 7.2 Prometheus config snippet
Expand All @@ -483,7 +483,7 @@ series for the gauges + ~30 for the counters.
Resource limits today are compile-time constants. Per-module
overrides via `[engine.limits]` are tracked as a 0.3 follow-up
(referenced from `crates/nexum-runtime/src/runtime/limits.rs`).
The tuning advice below is therefore advisory adjust by
The tuning advice below is therefore advisory - adjust by
changing the constants in `runtime/limits.rs` and rebuilding,
or by ensuring per-module loads fit within the current
defaults.
Expand Down Expand Up @@ -525,7 +525,7 @@ groups:
POISON_WINDOW. Engine has stopped dispatching to it.
Investigate: journalctl -u shepherd | jq 'select(.fields.module=="{{ $labels.module }}")'

# P1: trap rate climbing. Pre-poison signal gives 5 min
# P1: trap rate climbing. Pre-poison signal - gives 5 min
# of warning before ShepherdModulePoisoned fires.
- alert: ShepherdModuleTraps
expr: rate(shepherd_module_errors_total{error_kind="trap"}[5m]) > 0
Expand Down Expand Up @@ -626,7 +626,7 @@ journalctl -u shepherd -f --output=json \
### 10.2 Reset a poisoned module

A poisoned module stays poisoned until process restart (M4
design no live un-poison API yet). The recovery flow:
design - no live un-poison API yet). The recovery flow:

1. Triage the failure: `journalctl -u shepherd | jq 'select(.MESSAGE | fromjson? | .level == "ERROR" or (.fields.message | test("trapped|poisoned")))'`.
2. Fix the underlying bug (in the module's Rust code, or the
Expand Down Expand Up @@ -654,7 +654,7 @@ a 0.3+ follow-up.
There is no `ls-dump` CLI today. Workarounds:

- Boot a one-shot Rust script with `redb::Database::open` (read-
only) against the live file. Safe redb supports concurrent
only) against the live file. Safe - redb supports concurrent
readers + a single writer.
- Stop the engine + use any redb inspector tool against the
copy.
Expand Down
Loading