diff --git a/docs/operations/backtest-reports/backtest-7d-2026-06-22.md b/docs/operations/backtest-reports/backtest-7d-2026-06-22.md deleted file mode 100644 index a1796243..00000000 --- a/docs/operations/backtest-reports/backtest-7d-2026-06-22.md +++ /dev/null @@ -1,289 +0,0 @@ -# Pre-soak backtest - 7d window on Sepolia (2026-07-10T18:55:13Z) - -Replays every collected EthFlow `OrderPlacement` event through the production `ethflow_watcher::strategy::on_chain_logs` code path via `shepherd_sdk_test::MockHost`. The orderbook is **never hit**: the MockHost programs a catch-all 200 for all `cow_api_request` calls so the observe+verify strategy sees every fixture as already indexed. Success is measured by whether the strategy wrote the exact `observed:{uid}` marker to the local store after the 200 confirmation. - -## Run metadata - -| Field | Value | -|---|---| -| Chain | Sepolia (id=11155111) | -| Window | 7d (11066372..11116772) | -| Collected at | 2026-06-22T15:47:06Z | -| RPC | `https://sepolia.drpc.org` | -| Orderbook | `https://api.cow.fi/sepolia/api/v1` | -| EthFlow owner | `0xba3cb449bd2b4adddbc894d8697f5170800eadec` | -| ComposableCoW | `0xfdafc9d1902f4e0b84f65f49f244b32b31013b74` | -| Accept threshold | 95% | - -## EthFlow replay summary - -- Events replayed: **240** -- Observed: **240** (100.0%) - -Accepted (Observed + RejectedExpected): **240/240 = 100.0%** - PASS vs. threshold (95%). - -## Anomalies - -None. Every replayed event landed in `Observed` or `RejectedExpected`. - -## TWAP lane status - -26 `ConditionalOrderCreated` events were collected in this window. **Replay deferred to Phase 2B** because driving `twap_monitor::strategy::on_block` requires walking each watch's `eth_call(getTradeableOrderWithSignature)` per-block - a workload public-tier RPCs refuse (see the baseline-latency finding). The fixtures are committed for the future re-run; the TWAP gap on the sign-off is intentional and tracked separately. - -## Sign-off - -**PASS.** EthFlow replay clears the 95% acceptance bar with no outstanding anomalies. All fixtures were Observed (strategy wrote `observed:{uid}` to local store). Soak is unblocked from the backtest side; remaining blockers are external (paid RPC + VM for the wall-clock run). - -## Reproducing - -```bash -python3 tools/backtest-collect/backtest_collect.py --days 7 -cargo run -p shepherd-backtest -- \ - --fixtures tools/backtest-collect/fixtures-YYYY-MM-DD.json -``` - -## Appendix: per-event classification - -| # | uid | block | timestamp | class | -|---:|---|---:|---:|---| -| 1 | `0x5e43c584..ffffff` | 11066776 | 1781541696 | Observed | -| 2 | `0xf56cba95..ffffff` | 11066777 | 1781541708 | Observed | -| 3 | `0xb47d7c7f..ffffff` | 11066798 | 1781541960 | Observed | -| 4 | `0x4069c497..ffffff` | 11066799 | 1781541972 | Observed | -| 5 | `0xed31c793..ffffff` | 11066818 | 1781542200 | Observed | -| 6 | `0x5da63e4e..ffffff` | 11066819 | 1781542212 | Observed | -| 7 | `0x6c3227cb..ffffff` | 11066844 | 1781542512 | Observed | -| 8 | `0x97fa5d54..ffffff` | 11066844 | 1781542512 | Observed | -| 9 | `0xe5197c47..ffffff` | 11066863 | 1781542740 | Observed | -| 10 | `0xf13dc9a1..ffffff` | 11066863 | 1781542740 | Observed | -| 11 | `0x6118ba38..ffffff` | 11067068 | 1781545200 | Observed | -| 12 | `0xb2efc9f1..ffffff` | 11067190 | 1781546664 | Observed | -| 13 | `0x7f76bfa1..ffffff` | 11067200 | 1781546784 | Observed | -| 14 | `0x2c6cca82..ffffff` | 11067729 | 1781553132 | Observed | -| 15 | `0xe32c8541..ffffff` | 11068198 | 1781558760 | Observed | -| 16 | `0x5af30b3e..ffffff` | 11068620 | 1781563836 | Observed | -| 17 | `0xfa067d01..ffffff` | 11069102 | 1781569632 | Observed | -| 18 | `0xf242a25b..ffffff` | 11069119 | 1781569836 | Observed | -| 19 | `0x8bd36dc7..ffffff` | 11069495 | 1781574348 | Observed | -| 20 | `0x591da4be..ffffff` | 11069501 | 1781574420 | Observed | -| 21 | `0xa9a747d5..ffffff` | 11069948 | 1781579796 | Observed | -| 22 | `0x8b778cc7..ffffff` | 11070077 | 1781581344 | Observed | -| 23 | `0x0a56f14f..ffffff` | 11070107 | 1781581704 | Observed | -| 24 | `0x43445124..ffffff` | 11070306 | 1781584104 | Observed | -| 25 | `0x01026a74..ffffff` | 11070324 | 1781584320 | Observed | -| 26 | `0x71b74cf6..ffffff` | 11070394 | 1781585160 | Observed | -| 27 | `0x8834595a..ffffff` | 11070716 | 1781589024 | Observed | -| 28 | `0x6f4fae10..ffffff` | 11071674 | 1781600520 | Observed | -| 29 | `0x50211d94..ffffff` | 11071675 | 1781600532 | Observed | -| 30 | `0xcd925b4d..ffffff` | 11071676 | 1781600544 | Observed | -| 31 | `0x297cb16d..ffffff` | 11071677 | 1781600556 | Observed | -| 32 | `0x57e24641..ffffff` | 11071679 | 1781600580 | Observed | -| 33 | `0xb30c35c2..ffffff` | 11071681 | 1781600604 | Observed | -| 34 | `0x743f9609..ffffff` | 11071682 | 1781600616 | Observed | -| 35 | `0x713bc286..ffffff` | 11071683 | 1781600628 | Observed | -| 36 | `0x7925f236..ffffff` | 11071684 | 1781600640 | Observed | -| 37 | `0x11d613bf..ffffff` | 11071687 | 1781600676 | Observed | -| 38 | `0xd42de36d..ffffff` | 11072025 | 1781604732 | Observed | -| 39 | `0xe81f3615..ffffff` | 11072165 | 1781606412 | Observed | -| 40 | `0xfe8b70cc..ffffff` | 11072663 | 1781612388 | Observed | -| 41 | `0xfb9ebfe2..ffffff` | 11072665 | 1781612412 | Observed | -| 42 | `0x76c0a5ea..ffffff` | 11072730 | 1781613192 | Observed | -| 43 | `0xb2e7c4b5..ffffff` | 11072738 | 1781613288 | Observed | -| 44 | `0x20484bb9..ffffff` | 11072742 | 1781613336 | Observed | -| 45 | `0x541fb237..ffffff` | 11072774 | 1781613720 | Observed | -| 46 | `0x2404f184..ffffff` | 11072774 | 1781613720 | Observed | -| 47 | `0x30e44c53..ffffff` | 11072791 | 1781613924 | Observed | -| 48 | `0x9a340499..ffffff` | 11072801 | 1781614044 | Observed | -| 49 | `0x7f7b151d..ffffff` | 11072849 | 1781614620 | Observed | -| 50 | `0xb68eeaf4..ffffff` | 11072850 | 1781614632 | Observed | -| 51 | `0x5395405d..ffffff` | 11072881 | 1781615004 | Observed | -| 52 | `0x45d5563b..ffffff` | 11072881 | 1781615004 | Observed | -| 53 | `0x15431ff4..ffffff` | 11072907 | 1781615316 | Observed | -| 54 | `0xab2d5a81..ffffff` | 11072907 | 1781615316 | Observed | -| 55 | `0x3918584c..ffffff` | 11072915 | 1781615412 | Observed | -| 56 | `0x433ded5d..ffffff` | 11072937 | 1781615676 | Observed | -| 57 | `0x38e4a8d5..ffffff` | 11072938 | 1781615688 | Observed | -| 58 | `0x2391bfae..ffffff` | 11073096 | 1781617584 | Observed | -| 59 | `0xf6cd036e..ffffff` | 11073102 | 1781617656 | Observed | -| 60 | `0x157fa6bc..ffffff` | 11073102 | 1781617656 | Observed | -| 61 | `0x70aeba19..ffffff` | 11073107 | 1781617716 | Observed | -| 62 | `0x8d4ca0b6..ffffff` | 11073145 | 1781618172 | Observed | -| 63 | `0x45902e3e..ffffff` | 11073145 | 1781618172 | Observed | -| 64 | `0x19d6b26a..ffffff` | 11073178 | 1781618568 | Observed | -| 65 | `0x8ecd3588..ffffff` | 11073179 | 1781618580 | Observed | -| 66 | `0x2baee5d9..ffffff` | 11073185 | 1781618652 | Observed | -| 67 | `0x0a684eb7..ffffff` | 11073186 | 1781618664 | Observed | -| 68 | `0xccf652d3..ffffff` | 11073220 | 1781619072 | Observed | -| 69 | `0x51bd84d8..ffffff` | 11073220 | 1781619072 | Observed | -| 70 | `0x3f2f050f..ffffff` | 11073251 | 1781619444 | Observed | -| 71 | `0x9b53383b..ffffff` | 11073251 | 1781619444 | Observed | -| 72 | `0xd55d2e7e..ffffff` | 11073259 | 1781619540 | Observed | -| 73 | `0x355e6c2e..ffffff` | 11073304 | 1781620080 | Observed | -| 74 | `0xb75a1e65..ffffff` | 11073309 | 1781620140 | Observed | -| 75 | `0x812f257a..ffffff` | 11073316 | 1781620224 | Observed | -| 76 | `0x25efa78a..ffffff` | 11073319 | 1781620260 | Observed | -| 77 | `0x72aee1ae..ffffff` | 11073340 | 1781620512 | Observed | -| 78 | `0xe10e143e..ffffff` | 11073345 | 1781620572 | Observed | -| 79 | `0x1cb540d1..ffffff` | 11073377 | 1781620956 | Observed | -| 80 | `0x4a27b3b7..ffffff` | 11073446 | 1781621784 | Observed | -| 81 | `0x6f60c9b7..ffffff` | 11073450 | 1781621832 | Observed | -| 82 | `0x2afac9af..ffffff` | 11073618 | 1781623884 | Observed | -| 83 | `0x09eed081..ffffff` | 11073639 | 1781624136 | Observed | -| 84 | `0x0f04118e..ffffff` | 11073639 | 1781624136 | Observed | -| 85 | `0xb1445f46..ffffff` | 11073661 | 1781624400 | Observed | -| 86 | `0x1b9f5ae8..ffffff` | 11073662 | 1781624412 | Observed | -| 87 | `0xbf218ce6..ffffff` | 11073684 | 1781624676 | Observed | -| 88 | `0xcd000d8e..ffffff` | 11073684 | 1781624676 | Observed | -| 89 | `0x41e05a80..ffffff` | 11073706 | 1781624940 | Observed | -| 90 | `0x46728c28..ffffff` | 11073706 | 1781624940 | Observed | -| 91 | `0x7d517f6b..ffffff` | 11073716 | 1781625060 | Observed | -| 92 | `0xcd0993d3..ffffff` | 11073738 | 1781625324 | Observed | -| 93 | `0xd41c9e0b..ffffff` | 11073738 | 1781625324 | Observed | -| 94 | `0x728367f6..ffffff` | 11073789 | 1781625936 | Observed | -| 95 | `0xa78cabeb..ffffff` | 11074351 | 1781632692 | Observed | -| 96 | `0x211dd498..ffffff` | 11074351 | 1781632692 | Observed | -| 97 | `0x5f686fb7..ffffff` | 11074387 | 1781633124 | Observed | -| 98 | `0x3b7b5aaa..ffffff` | 11074387 | 1781633124 | Observed | -| 99 | `0x8cb5b94f..ffffff` | 11074421 | 1781633532 | Observed | -| 100 | `0x7545eda0..ffffff` | 11074421 | 1781633532 | Observed | -| 101 | `0x6c4472c1..ffffff` | 11074463 | 1781634048 | Observed | -| 102 | `0x134a3f32..ffffff` | 11074464 | 1781634060 | Observed | -| 103 | `0x625b15a3..ffffff` | 11074482 | 1781634276 | Observed | -| 104 | `0x8b981bae..ffffff` | 11074491 | 1781634408 | Observed | -| 105 | `0x2316cffb..ffffff` | 11074491 | 1781634408 | Observed | -| 106 | `0x019e8adf..ffffff` | 11076610 | 1781659896 | Observed | -| 107 | `0xe0ccf0ed..ffffff` | 11076616 | 1781659968 | Observed | -| 108 | `0xfcfd0fa6..ffffff` | 11076620 | 1781660016 | Observed | -| 109 | `0x9fabbf08..ffffff` | 11077214 | 1781667204 | Observed | -| 110 | `0x84fd9342..ffffff` | 11077272 | 1781667900 | Observed | -| 111 | `0xa1b4b41b..ffffff` | 11078078 | 1781677608 | Observed | -| 112 | `0x9e4ceb19..ffffff` | 11078078 | 1781677608 | Observed | -| 113 | `0xd6eaa1a9..ffffff` | 11078093 | 1781677788 | Observed | -| 114 | `0x819ff1ec..ffffff` | 11078093 | 1781677788 | Observed | -| 115 | `0xa06aafbb..ffffff` | 11078113 | 1781678028 | Observed | -| 116 | `0xd6f7b83b..ffffff` | 11078114 | 1781678040 | Observed | -| 117 | `0x1ac7b661..ffffff` | 11078214 | 1781679240 | Observed | -| 118 | `0x9fc72d42..ffffff` | 11078215 | 1781679252 | Observed | -| 119 | `0x642b7890..ffffff` | 11078225 | 1781679372 | Observed | -| 120 | `0x7cf68a7b..ffffff` | 11078246 | 1781679624 | Observed | -| 121 | `0xb6e10751..ffffff` | 11078248 | 1781679648 | Observed | -| 122 | `0xd3c38d90..ffffff` | 11078475 | 1781682372 | Observed | -| 123 | `0x22e93a3a..ffffff` | 11078475 | 1781682372 | Observed | -| 124 | `0x5e60eb4e..ffffff` | 11078500 | 1781682672 | Observed | -| 125 | `0xaf138bc2..ffffff` | 11078501 | 1781682684 | Observed | -| 126 | `0xcbb6226f..ffffff` | 11078504 | 1781682720 | Observed | -| 127 | `0x7bc92f46..ffffff` | 11078513 | 1781682828 | Observed | -| 128 | `0x7bf72b31..ffffff` | 11078520 | 1781682912 | Observed | -| 129 | `0x9dd59d5b..ffffff` | 11078520 | 1781682912 | Observed | -| 130 | `0x8f8bed50..ffffff` | 11078539 | 1781683140 | Observed | -| 131 | `0x4e04244d..ffffff` | 11078539 | 1781683140 | Observed | -| 132 | `0xe16973fa..ffffff` | 11078557 | 1781683356 | Observed | -| 133 | `0xfb362d58..ffffff` | 11078558 | 1781683368 | Observed | -| 134 | `0x6068b9d9..ffffff` | 11080029 | 1781701044 | Observed | -| 135 | `0x299fe688..ffffff` | 11080365 | 1781705076 | Observed | -| 136 | `0xbdbb9773..ffffff` | 11082360 | 1781729064 | Observed | -| 137 | `0xbaedfe1c..ffffff` | 11083644 | 1781744496 | Observed | -| 138 | `0x8194545d..ffffff` | 11083764 | 1781745936 | Observed | -| 139 | `0x9c92f5f3..ffffff` | 11083770 | 1781746008 | Observed | -| 140 | `0x98906b97..ffffff` | 11083812 | 1781746512 | Observed | -| 141 | `0x627afacc..ffffff` | 11083859 | 1781747076 | Observed | -| 142 | `0x240bcbde..ffffff` | 11083866 | 1781747160 | Observed | -| 143 | `0x2559867e..ffffff` | 11084049 | 1781749380 | Observed | -| 144 | `0x7ce8d855..ffffff` | 11084079 | 1781749740 | Observed | -| 145 | `0xffc3686e..ffffff` | 11084150 | 1781750592 | Observed | -| 146 | `0x7b51bb4a..ffffff` | 11084744 | 1781757792 | Observed | -| 147 | `0x8b4e73ec..ffffff` | 11085093 | 1781762016 | Observed | -| 148 | `0x8bd2dbf8..ffffff` | 11085093 | 1781762016 | Observed | -| 149 | `0xd3d9da38..ffffff` | 11085123 | 1781762376 | Observed | -| 150 | `0x518e19aa..ffffff` | 11085123 | 1781762376 | Observed | -| 151 | `0x8bb95de2..ffffff` | 11085229 | 1781763648 | Observed | -| 152 | `0x3f80482c..ffffff` | 11085229 | 1781763648 | Observed | -| 153 | `0x6217df15..ffffff` | 11085285 | 1781764320 | Observed | -| 154 | `0xb8b7945e..ffffff` | 11085290 | 1781764380 | Observed | -| 155 | `0xbeaa3ed3..ffffff` | 11085296 | 1781764452 | Observed | -| 156 | `0xac9baa9a..ffffff` | 11085311 | 1781764632 | Observed | -| 157 | `0x5471a5aa..ffffff` | 11085316 | 1781764692 | Observed | -| 158 | `0xb9af72ee..ffffff` | 11085768 | 1781770116 | Observed | -| 159 | `0x577b183c..ffffff` | 11086236 | 1781775732 | Observed | -| 160 | `0xafd52f06..ffffff` | 11086284 | 1781776308 | Observed | -| 161 | `0x4ee00443..ffffff` | 11086290 | 1781776380 | Observed | -| 162 | `0x64427828..ffffff` | 11086493 | 1781778816 | Observed | -| 163 | `0x31cce049..ffffff` | 11087438 | 1781790168 | Observed | -| 164 | `0x927561c1..ffffff` | 11087442 | 1781790216 | Observed | -| 165 | `0x33e40575..ffffff` | 11087462 | 1781790456 | Observed | -| 166 | `0x5fe03b4e..ffffff` | 11087468 | 1781790528 | Observed | -| 167 | `0x5e7f3fe1..ffffff` | 11087473 | 1781790588 | Observed | -| 168 | `0x0cfa5720..ffffff` | 11087748 | 1781793888 | Observed | -| 169 | `0x88e25f26..ffffff` | 11087749 | 1781793900 | Observed | -| 170 | `0x3d47b55b..ffffff` | 11089218 | 1781811528 | Observed | -| 171 | `0x91b7bb98..ffffff` | 11089257 | 1781811996 | Observed | -| 172 | `0x47931578..ffffff` | 11089274 | 1781812200 | Observed | -| 173 | `0x006b940d..ffffff` | 11089296 | 1781812464 | Observed | -| 174 | `0x104f25a0..ffffff` | 11089394 | 1781813640 | Observed | -| 175 | `0x6d296984..ffffff` | 11089725 | 1781817624 | Observed | -| 176 | `0xf5788a8b..ffffff` | 11089920 | 1781819964 | Observed | -| 177 | `0xdd4a43c4..ffffff` | 11090323 | 1781824800 | Observed | -| 178 | `0x3d1098b8..ffffff` | 11091082 | 1781833920 | Observed | -| 179 | `0xe11c2022..ffffff` | 11091089 | 1781834004 | Observed | -| 180 | `0x1b7ecce8..ffffff` | 11091095 | 1781834076 | Observed | -| 181 | `0x17413c36..ffffff` | 11091102 | 1781834160 | Observed | -| 182 | `0xe18203b8..ffffff` | 11091111 | 1781834268 | Observed | -| 183 | `0x362dcbfb..ffffff` | 11091361 | 1781837316 | Observed | -| 184 | `0x7fe88a51..ffffff` | 11093487 | 1781863032 | Observed | -| 185 | `0x0c37aa67..ffffff` | 11094315 | 1781872980 | Observed | -| 186 | `0x2170ca89..ffffff` | 11094609 | 1781876508 | Observed | -| 187 | `0xcc734808..ffffff` | 11094664 | 1781877168 | Observed | -| 188 | `0x30fbd136..ffffff` | 11095162 | 1781883144 | Observed | -| 189 | `0x35e2b54d..ffffff` | 11095647 | 1781888964 | Observed | -| 190 | `0x2300a9f6..ffffff` | 11096098 | 1781894376 | Observed | -| 191 | `0xe2e97306..ffffff` | 11097901 | 1781916048 | Observed | -| 192 | `0x15d6d916..ffffff` | 11098198 | 1781919612 | Observed | -| 193 | `0xd0186dc0..ffffff` | 11098367 | 1781921652 | Observed | -| 194 | `0xe188b861..ffffff` | 11098372 | 1781921712 | Observed | -| 195 | `0xd13cad5b..ffffff` | 11098766 | 1781926452 | Observed | -| 196 | `0xbd8cc161..ffffff` | 11099348 | 1781933460 | Observed | -| 197 | `0x77446407..ffffff` | 11099353 | 1781933520 | Observed | -| 198 | `0xfec45a42..ffffff` | 11100012 | 1781941428 | Observed | -| 199 | `0xd97a9fa0..ffffff` | 11100016 | 1781941476 | Observed | -| 200 | `0xb1c28fdb..ffffff` | 11102012 | 1781965476 | Observed | -| 201 | `0x62ac8489..ffffff` | 11102181 | 1781967528 | Observed | -| 202 | `0x64fc616d..ffffff` | 11103512 | 1781983536 | Observed | -| 203 | `0x809b0200..ffffff` | 11105542 | 1782007956 | Observed | -| 204 | `0x2efe5755..ffffff` | 11105545 | 1782008004 | Observed | -| 205 | `0xa801952a..ffffff` | 11105551 | 1782008076 | Observed | -| 206 | `0xe51fcd2e..ffffff` | 11105566 | 1782008256 | Observed | -| 207 | `0x87e7ed80..ffffff` | 11105602 | 1782008688 | Observed | -| 208 | `0xa249ff22..ffffff` | 11105613 | 1782008820 | Observed | -| 209 | `0xfbe5e123..ffffff` | 11105618 | 1782008880 | Observed | -| 210 | `0x72bf47dc..ffffff` | 11105715 | 1782010044 | Observed | -| 211 | `0x1293c734..ffffff` | 11106911 | 1782024444 | Observed | -| 212 | `0xeeb9580b..ffffff` | 11106932 | 1782024696 | Observed | -| 213 | `0x42e1019e..ffffff` | 11107190 | 1782027792 | Observed | -| 214 | `0x863285b8..ffffff` | 11108038 | 1782037968 | Observed | -| 215 | `0x3545ede8..ffffff` | 11109421 | 1782054600 | Observed | -| 216 | `0x98d90da1..ffffff` | 11112811 | 1782095304 | Observed | -| 217 | `0xf60dc92a..ffffff` | 11112823 | 1782095448 | Observed | -| 218 | `0xbad0af58..ffffff` | 11112886 | 1782096204 | Observed | -| 219 | `0xda12aeee..ffffff` | 11112894 | 1782096300 | Observed | -| 220 | `0x271f933d..ffffff` | 11113205 | 1782100044 | Observed | -| 221 | `0x354f7970..ffffff` | 11114331 | 1782113580 | Observed | -| 222 | `0x52d511d9..ffffff` | 11114353 | 1782113844 | Observed | -| 223 | `0x39ba5342..ffffff` | 11115217 | 1782124248 | Observed | -| 224 | `0x8f627309..ffffff` | 11115224 | 1782124332 | Observed | -| 225 | `0x4d0b40ff..ffffff` | 11115229 | 1782124392 | Observed | -| 226 | `0xdc658bc7..ffffff` | 11115320 | 1782125508 | Observed | -| 227 | `0x3ad5be48..ffffff` | 11115412 | 1782126624 | Observed | -| 228 | `0xa412cc0c..ffffff` | 11115417 | 1782126684 | Observed | -| 229 | `0x6cce9b62..ffffff` | 11115424 | 1782126768 | Observed | -| 230 | `0x45902f9e..ffffff` | 11115429 | 1782126828 | Observed | -| 231 | `0xfac5eea4..ffffff` | 11115499 | 1782127668 | Observed | -| 232 | `0x3f581643..ffffff` | 11115786 | 1782131112 | Observed | -| 233 | `0x303a3415..ffffff` | 11115796 | 1782131232 | Observed | -| 234 | `0xc6bf93cb..ffffff` | 11115816 | 1782131472 | Observed | -| 235 | `0xc70930be..ffffff` | 11116414 | 1782138732 | Observed | -| 236 | `0xbdc6f0ae..ffffff` | 11116631 | 1782141408 | Observed | -| 237 | `0x0bab3b08..ffffff` | 11116635 | 1782141456 | Observed | -| 238 | `0x1916db8f..ffffff` | 11116641 | 1782141528 | Observed | -| 239 | `0xda1c4056..ffffff` | 11116645 | 1782141576 | Observed | -| 240 | `0xa2d2d863..ffffff` | 11116660 | 1782141756 | Observed | - diff --git a/docs/operations/baselines/baseline-latency-2026-06-19.md b/docs/operations/baselines/baseline-latency-2026-06-19.md deleted file mode 100644 index 4faa5ccb..00000000 --- a/docs/operations/baselines/baseline-latency-2026-06-19.md +++ /dev/null @@ -1,60 +0,0 @@ -# CoW orderbook EthFlow indexer baseline (2026-06-22T14:03:22Z) - -Per-chain pairing of every on-chain `EthFlow.OrderPlacement` event in the trailing window with the orderbook's record for the same UID, plus the `(creationDate - block.timestamp)` delta. Each pair is rigorous — the script ABI-decodes the event's GPv2OrderData and derives the OrderUid via EIP-712 before looking it up — so the data is ground-truth, not a temporal-FIFO approximation. - -## Headline finding - -**For EthFlow orders the orderbook indexer sets `creationDate := block.timestamp`** (not the indexer's ingest time), so the historical delta is structurally 0s on every chain. This is the orderbook's intentional behaviour for back-fill-style flows; it is **not** a measurement bug. The implication for the M4 / M5 KPIs is that EthFlow indexer latency cannot be derived from historical orderbook data — the meaningful relayer-latency baseline lives on the TWAP lane (where the orderbook records the indexer's `now()` per child order PUT). TWAP child-latency is tracked as a follow-up since it requires per-part UID derivation from each parent `ConditionalOrderCreated` static input. - -What the run below **is** useful for: confirming the orderbook's `creationDate` semantics across every supported chain, and yielding ground-truth UID ↔ block pairings the M4 e2e harness can cross-check against. - -## Method - -- Window: trailing **7 days** from the run. -- Event source: `eth_getLogs` against the chain's ETH_FLOW_PRODUCTION (ETH_FLOW_SEPOLIA on Sepolia) for the `OrderPlacement` topic. -- Order source: `GET /account/{ETH_FLOW_ADDRESS}/orders` from the chain's cow.fi orderbook, paginated. -- Pairing: per-event EIP-712 UID derivation. For each event the script ABI-decodes the GPv2OrderData payload, computes the order digest against the chain's GPv2Settlement domain, and assembles UID = digest || ethflow_owner || validTo. Each UID is then looked up against the bulk `/account/.../orders` fetch, falling back to `GET /api/v1/orders/{uid}` if the bulk page missed it. No temporal-FIFO approximation. -- Sanity filters: negative deltas dropped (clock skew between block and indexer); deltas > 1 hour dropped (stale/re-indexed order). -- Event cap per chain: **200** (most recent). - -## EthFlow latency, per chain - -| Chain | Events scanned | Orders fetched | Pairs | Median (s) | p95 (s) | -|---|---:|---:|---:|---:|---:| -| Mainnet | 0 | 0 | 0 | n/a | n/a | -| Gnosis | 0 | 0 | 0 | n/a | n/a | -| Arbitrum One | 0 | 0 | 0 | n/a | n/a | -| Base | 0 | 0 | 0 | n/a | n/a | -| Sepolia | 256 | 5000 | 200 | 0.00 | 0.00 | - -## TWAP latency, per chain - -*Not measured in v1 of this baseline.* TWAP requires reconstructing `(t0, n, t)` from each parent `ConditionalOrderCreated` static input and deriving each child order's UID per part, then matching to the orderbook's child orders. Tracked as a follow-up; **EthFlow alone is sufficient anchor for the M4 KPI bar** since both modules share the same dispatch path in shepherd. - -## Notes per chain - -- **Mainnet**: - - RPC-LIMITED: public endpoint (https://eth.drpc.org) refused the log scan even at 50-block chunks (endpoint refused 3 consecutive calls at chunk=31: 408 Client Error: Request Timeout for url: https://eth.drpc.org/). Re-run with a paid endpoint via RPC_URL_* env to get real data; this baseline cell stays blank. Matches the paid-endpoint requirement. -- **Gnosis**: - - RPC-LIMITED: public endpoint (https://gnosis.drpc.org) refused the log scan even at 50-block chunks (endpoint refused 3 consecutive calls at chunk=31: 500 Server Error: Internal Server Error for url: https://gnosis.drpc.org/). Re-run with a paid endpoint via RPC_URL_* env to get real data; this baseline cell stays blank. Matches the paid-endpoint requirement. -- **Arbitrum One**: - - RPC-LIMITED: public endpoint (https://arbitrum.drpc.org) refused the log scan even at 50-block chunks (endpoint refused 3 consecutive calls at chunk=31: 500 Server Error: Internal Server Error for url: https://arbitrum.drpc.org/). Re-run with a paid endpoint via RPC_URL_* env to get real data; this baseline cell stays blank. Matches the paid-endpoint requirement. -- **Base**: - - RPC-LIMITED: public endpoint (https://base.drpc.org) refused the log scan even at 50-block chunks (endpoint refused 3 consecutive calls at chunk=31: 500 Server Error: Internal Server Error for url: https://base.drpc.org/). Re-run with a paid endpoint via RPC_URL_* env to get real data; this baseline cell stays blank. Matches the paid-endpoint requirement. -- **Sepolia**: - - capped to last 200 events of 256 - - match diagnostics: bulk_hit=200 - -## Reproducing - -```bash -python3 tools/baseline-latency/baseline_latency.py \ - --window-days 7 --max-events-per-chain 200 \ - --out docs/operations/baselines/baseline-latency-$(date -u +%Y-%m-%d).md -``` - -Override individual RPCs via env: `RPC_URL_MAINNET`, `RPC_URL_GNOSIS`, `RPC_URL_ARBITRUM`, `RPC_URL_BASE`, `RPC_URL_SEPOLIA_HTTP`. - -## Provenance - -Script: `tools/baseline-latency/baseline_latency.py`. Raw data dump per chain: `tools/baseline-latency/data/`. diff --git a/docs/operations/e2e-prep.md b/docs/operations/e2e-prep.md deleted file mode 100644 index 795dd890..00000000 --- a/docs/operations/e2e-prep.md +++ /dev/null @@ -1,334 +0,0 @@ -# E2E run-prep punch list - -Companion to `docs/operations/e2e-testnet-runbook.md`. This file -captures every **pinned value** for the 2026-06-18 dry run so the operator can copy-paste through the on-chain -actions without re-deriving any UID, address, or calldata. - -If you are running a *later* E2E run (different EOA, different -Safe, different config), do not reuse the UIDs / calldatas — they -are a function of all the pinned config below. Either re-derive -via the Python recipes in this doc, or re-run -`cargo test -p stop-loss --lib cow_1064` to lock the new UID. - ---- - -## 0. Pinned identities (2026-06-18 run) - -| Role | Address | Network | Notes | -|---|---|---|---| -| Test EOA | `0x7bF140727D27ea64b607E042f1225680B40ECa6A` | Sepolia | Bruno-controlled. Funds itself via faucet. | -| Test Safe (single-sig, threshold 1) | `0x14995a1118Caf95833e923faf8Dd155721cd53c2` | Sepolia | EOA is the sole owner. Submits TWAP order. | -| ComposableCoW | `0xfdaFc9d1902f4e0b84f65F49f244b32b31013b74` | Sepolia | Where `create((address,bytes32,bytes),bool)` lands. | -| TWAP handler | `0x6cF1e9cA41f7611dEf408122793c358a3d11E5a5` | Sepolia | `ConditionalOrderParams.handler`. | -| CoWSwapEthFlow | `0xbA3cB449bD2B4ADddBc894D8697F5170800EAdeC` | Sepolia | EthFlow's production deployment; emits `OrderPlacement`. | -| GPv2Settlement | `0x9008D19f58AAbD9eD0D60971565AA8510560ab41` | Sepolia | `setPreSignature(orderUid, signed)` lives here. | -| GPv2VaultRelayer | `0xc92e8bdf79f0507f65a392b0ab4667716bfe0110` | Sepolia | Spender for sell-token ERC-20 approvals. | -| WETH9 | `0xfFf9976782d46CC05630D1f6eBAb18b2324d6B14` | Sepolia | `deposit()` payable wraps ETH; `balanceOf(EOA)` is the sell-side balance. | -| COW Token | `0x0625aFB445C3B6B7B929342a04A22599fd5dBB59` | Sepolia | name="CoW Protocol Token", symbol="COW", decimals=18. | -| GPv2 domain separator | `0xdaee378bd0eb30ddf479272accf91761e697bc00e067a268f95f1d2732ed230b` | Sepolia | EIP-712 domain digest queried from chain. | - -All addresses verified via `eth_getCode > 0` on -`https://ethereum-sepolia-rpc.publicnode.com` as of run prep. - ---- - -## 1. Per-module config pinning - -### stop-loss - -`modules/examples/stop-loss/module.toml` is checked in on the -`feat/e2e-run-config-cow-1064` branch with the production-ready -config for this run. Effective values: - -| Field | Value | Notes | -|---|---|---| -| `oracle_address` | `0x694AA1769357215DE4FAC081bf1f309aDC325306` | Chainlink ETH/USD Sepolia. | -| `decimals` | `8` | Chainlink USD-pair convention. | -| `trigger_price` | `2000.00` | Above the live Sepolia mocked answer (~$1681), `direction=below` → triggers on first block. | -| `owner` | `0x7bF1...Ca6A` | Test EOA. | -| `sell_token` | `0xfFf9...6B14` | WETH9 Sepolia. | -| `buy_token` | `0x0625...BB59` | COW Sepolia. | -| `sell_amount_wei` | `5000000000000000` | 0.005 WETH. | -| `buy_amount_wei` | `20000000000000000000` | 20 COW. Conservative quote at run-prep time. | -| `valid_to_seconds` | `4294967295` | uint32::MAX. | - -### Resulting OrderUid - -The strategy's `build_creation` is pinned by the -`cow_1064_e2e_settings_yield_expected_uid` regression test -(`crates/.../stop-loss/src/strategy.rs`). The canonical UID: - -``` -0xc2b9cb4ea1ee5a86d8049ac09d8f494bf04cca0a68407285f31e2e6379800be87bf140727d27ea64b607e042f1225680b40eca6affffffff -``` - -Decomposition (per `packOrderUidParams`): - -| Offset | Bytes | Field | Value | -|---|---|---|---| -| 0..32 | 32 | `orderDigest` (EIP-712) | `0xc2b9cb4ea1ee5a86d8049ac09d8f494bf04cca0a68407285f31e2e6379800be8` | -| 32..52 | 20 | `owner` | `0x7bf140727d27ea64b607e042f1225680b40eca6a` | -| 52..56 | 4 | `validTo` (uint32) | `0xffffffff` | - -### balance-tracker - -Pinned to the EOA + Safe so the run sees ETH-balance diffs: - -| Field | Value | -|---|---| -| `addresses` | `0x7bF1...Ca6A,0x1499...53c2` | -| `change_threshold` | `1000000000000000` (0.001 ETH) | - ---- - -## 2. On-chain actions for the run window - -> Order: action 1 can be done at any time before/during the run. -> Actions 2-4 should fire **after** the engine prints -> `INFO supervisor ready modules=5 chains=1` so the modules -> observe the events. They are independent; do them in any order. - -### Action 1 (optional, pre-run): wrap 0.01 ETH → 0.01 WETH - -Without WETH, stop-loss will hit `TransferSimulationFailed` -> -`backoff:` write (which is itself a valid terminal-marker per -the acceptance bar). To get the **`submitted:`** path, -wrap first then do action 2. - -- Etherscan: https://sepolia.etherscan.io/address/0xfff9976782d46cc05630d1f6ebab18b2324d6b14#writeContract -- Connect Web3 from the EOA in Metamask -- Function `deposit` → payable value `0.01` ETH → Write - -Verify: `balanceOf(EOA)` returns `10000000000000000` post-tx. - -### Action 2 (optional, only if action 1 done): pre-sign stop-loss order - -- Etherscan: https://sepolia.etherscan.io/address/0x9008d19f58aabd9ed0d60971565aa8510560ab41#writeProxyContract -- Connect Web3 from the EOA -- Function `setPreSignature(bytes orderUid, bool signed)`: - - `orderUid`: - ``` - 0xc2b9cb4ea1ee5a86d8049ac09d8f494bf04cca0a68407285f31e2e6379800be87bf140727d27ea64b607e042f1225680b40eca6affffffff - ``` - - `signed`: `true` -- Write - -Also approve WETH → GPv2VaultRelayer so the settle path is real: - -- Etherscan: https://sepolia.etherscan.io/address/0xfff9976782d46cc05630d1f6ebab18b2324d6b14#writeContract -- Function `approve(address guy, uint256 wad)`: - - `guy`: `0xc92e8bdf79f0507f65a392b0ab4667716bfe0110` - - `wad`: `5000000000000000` (0.005 WETH — matches the order's sell_amount) -- Write - -### Action 3: TWAP conditional order via Safe TX Builder - -Triggers `ConditionalOrderCreated` → twap-monitor writes -`watch:{orderHash}`. The Safe pays the gas (~0.003 ETH); the -order will TRY to settle later but the Safe holds no WETH so -settlement will fail. **That's fine** — only the `create()` -event is required for the acceptance marker. - -- Safe app: https://app.safe.global/transactions/queue?safe=sep:0x14995a1118Caf95833e923faf8Dd155721cd53c2 -- New transaction → Transaction Builder -- Enter contract address: `0xfdaFc9d1902f4e0b84f65F49f244b32b31013b74` -- Toggle "Use custom data (hex encoded)" ON -- Generate the calldata locally (do NOT paste a pinned blob): - -```bash -python3 scripts/_twap_calldata.py -``` - -The helper backdates `t0` by 60 s on every invocation so part 0 is -Ready immediately. The constants (sell/buy tokens, amounts, n, t, -salt) mirror section 4.2; edit there + in the helper in lockstep -if the TWAP shape changes. - -Copy the helper's stdout into the Transaction Builder's custom-data -field. The blob is ~516 bytes - the `create(ConditionalOrderParams, -bool dispatch)` call with a 2-part TWAP from WETH → COW, 0.001 WETH -per part, 600 s between parts, salt pinned to `0x...6670f000`. - -> Historical note: a previously-pinned variant of this calldata -> hardcoded `t0 = 0`, which silently produced an -> `AFTER_TWAP_FINISHED` revert on every poll because -> `calculateValidTo` divided `block.timestamp` by `t` and exceeded -> `n`. Surfaced in the 2026-06-18 dry run. Always derive -> via the helper. - -- ETH value: `0` -- Create batch → Send batch → sign with the EOA - -Expected log within 1-2 Sepolia blocks: - -``` -INFO twap-monitor watch:0x chain_id=11155111 -``` - -### Action 4: EthFlow swap via cow-swap UI - -Triggers `OrderPlacement` → ethflow-watcher writes -`submitted:{uid}` (or `dropped:{uid}` if the orderbook rejects; -both are valid terminal markers). - -Easiest path is the cow-swap UI: - -1. https://swap.cow.fi/#/11155111/swap/ETH/COW (Sepolia) -2. Connect Metamask, EOA selected, network=Sepolia -3. Sell amount: `0.005` ETH -4. Click "Swap" → it builds the EthFlow `createOrder` tx -5. Approve in Metamask - -The UI handles `quoteId` resolution + `appData` IPFS pinning + -EthFlow contract call. Sell amount is small enough to fit in the -~0.05 ETH budget plus gas. - -Expected log within 1-2 Sepolia blocks: - -``` -INFO ethflow-watcher submitted:0x -``` - -If the UI errors out (Sepolia orderbook can be flaky), fallback -to calling EthFlow directly via Etherscan: - -- https://sepolia.etherscan.io/address/0xba3cb449bd2b4adddbc894d8697f5170800eadec#writeContract -- Function `createOrder((address,address,uint256,uint256,bytes32,uint256,uint32,bool,int64))` -- The shape of the tuple needs the orderbook quote endpoint hit - first to get `feeAmount` + `quoteId` — easier to defer to the - UI for the run. - ---- - -## 3. Validation snippets for the operator - -Run these in a separate shell while the engine is up: - -```bash -RPC="wss://eth-sepolia.g.alchemy.com/v2/" # replace -EOA="0x7bF140727D27ea64b607E042f1225680B40ECa6A" -SAFE="0x14995a1118Caf95833e923faf8Dd155721cd53c2" -WETH="0xfFf9976782d46CC05630D1f6eBAb18b2324d6B14" - -# EOA + Safe balances -cast balance $EOA --rpc-url $RPC -cast balance $SAFE --rpc-url $RPC - -# EOA WETH balance + GPv2VaultRelayer allowance -cast call $WETH "balanceOf(address)(uint256)" $EOA --rpc-url $RPC -cast call $WETH "allowance(address,address)(uint256)" \ - $EOA 0xc92e8bdf79f0507f65a392b0ab4667716bfe0110 --rpc-url $RPC - -# Did setPreSignature land? -cast call 0x9008D19f58AAbD9eD0D60971565AA8510560ab41 \ - "preSignature(bytes)(uint256)" \ - 0xc2b9cb4ea1ee5a86d8049ac09d8f494bf04cca0a68407285f31e2e6379800be87bf140727d27ea64b607e042f1225680b40eca6affffffff \ - --rpc-url $RPC -# Returns 1 if pre-signed, 0 otherwise. - -# Mine the supervisor log for terminal markers in real time -journalctl -u shepherd -f --output=json \ - | jq -r '.MESSAGE | fromjson? | select(.fields.message | test("watch:|submitted:|dropped:|backoff:|TRIGGERED")) | "\(.fields.module): \(.fields.message)"' -``` - -(If you don't have `cast` installed: `curl -L https://foundry.paradigm.xyz | bash && foundryup`.) - ---- - -## 4. Recipes for re-deriving the pinned values - -If anything in section 0 drifts, regenerate from these recipes. - -### 4.1 OrderUid - -Either: - -```bash -cargo test -p stop-loss --lib cow_1064 -- --nocapture -``` - -(asserts against the same constants pinned in `module.toml`, -fails loudly if the EIP-712 type-hash or domain separator -shifts). - -Or with raw Python: - -```python -from eth_utils import keccak - -# Replace these 8 values to re-derive -DOMAIN_SEP = bytes.fromhex("daee378bd0eb30ddf479272accf91761e697bc00e067a268f95f1d2732ed230b") -SELL_TOKEN = bytes.fromhex("fFf9976782d46CC05630D1f6eBAb18b2324d6B14") -BUY_TOKEN = bytes.fromhex("0625aFB445C3B6B7B929342a04A22599fd5dBB59") -OWNER = bytes.fromhex("7bF140727D27ea64b607E042f1225680B40ECa6A") -RECEIVER = OWNER -SELL_AMOUNT = 5_000_000_000_000_000 -BUY_AMOUNT = 20_000_000_000_000_000_000 -VALID_TO = 4_294_967_295 - -APP_DATA = bytes.fromhex("b48d38f93eaa084033fc5970bf96e559c33c4cdc07d889ab00b4d63f9590739d") # keccak("{}") -KIND_SELL = keccak(b"sell") -ERC20 = keccak(b"erc20") -TYPE_HASH = keccak(b"Order(address sellToken,address buyToken,address receiver,uint256 sellAmount,uint256 buyAmount,uint32 validTo,bytes32 appData,uint256 feeAmount,string kind,bool partiallyFillable,string sellTokenBalance,string buyTokenBalance)") -pad32 = lambda b: bytes(32-len(b)) + b -uint = lambda v: v.to_bytes(32, "big") -struct_hash = keccak( - TYPE_HASH + pad32(SELL_TOKEN) + pad32(BUY_TOKEN) + pad32(RECEIVER) - + uint(SELL_AMOUNT) + uint(BUY_AMOUNT) + uint(VALID_TO) - + APP_DATA + uint(0) + KIND_SELL - + b"\x00"*32 + ERC20 + ERC20 # partiallyFillable=false -) -order_digest = keccak(b"\x19\x01" + DOMAIN_SEP + struct_hash) -uid = order_digest + OWNER + VALID_TO.to_bytes(4, "big") -print("0x" + uid.hex()) -``` - -### 4.2 ComposableCoW.create() calldata - -```python -import time -from eth_utils import keccak -from eth_abi import encode - -selector = keccak(b"create((address,bytes32,bytes),bool)")[:4] -# Edit these 10 fields to retarget the TWAP -static = encode( - ["(address,address,address,uint256,uint256,uint256,uint256,uint256,uint256,bytes32)"], - [( - "0xfFf9976782d46CC05630D1f6eBAb18b2324d6B14", # sellToken - "0x0625aFB445C3B6B7B929342a04A22599fd5dBB59", # buyToken - "0x14995a1118Caf95833e923faf8Dd155721cd53c2", # receiver - 1_000_000_000_000_000, 500_000_000_000_000_000, # partSellAmount, minPartLimit - int(time.time()) - 60, 2, 600, 0, # t0 (NEVER 0 - see note above), n, t, span - b"\x00" * 32, # appData - )] -) -calldata = selector + encode( - ["(address,bytes32,bytes)", "bool"], - [( - "0x6cF1e9cA41f7611dEf408122793c358a3d11E5a5", # TWAP handler - bytes.fromhex("000000000000000000000000000000000000000000000000000000006670f000"), # salt - static, - ), True] -) -print("0x" + calldata.hex()) -``` - ---- - -## 5. Acceptance checklist for THIS run - -Hand-check at the end of the run (also goes in -`e2e-report-YYYY-MM-DD.md` section 7): - -- [ ] EOA at `0x7bF1...Ca6A` still has ≥ 0.03 ETH remaining -- [ ] twap-monitor logged `watch:0x...` after action 3 -- [ ] ethflow-watcher logged `submitted:0x...` after action 4 -- [ ] stop-loss logged `backoff:` or `TRIGGERED + submitted:` (depending on whether action 1+2 ran) -- [ ] price-alert logged `TRIGGERED` on first block -- [ ] balance-tracker logged a `last:0x7bf1...` write on first block + at least one Warn diff log over the run window -- [ ] `shepherd_module_poisoned{...} == 0` for all 5 modules at end -- [ ] `shepherd_module_errors_total{error_kind="trap"} == 0` for all modules -- [ ] ≥ 1500 Sepolia blocks dispatched (`block delta` in report section 2) - -If all green the run is complete and the 7-day soak can start. diff --git a/docs/operations/e2e-reports/e2e-report-2026-06-18.md b/docs/operations/e2e-reports/e2e-report-2026-06-18.md deleted file mode 100644 index 9cc93774..00000000 --- a/docs/operations/e2e-reports/e2e-report-2026-06-18.md +++ /dev/null @@ -1,242 +0,0 @@ -# E2E testnet integration report — 2026-06-18 - -> Auto-generated by `scripts/e2e-report-gen.sh`. Operator -> review each section + flesh out anomalies + sign off in -> section 8 before committing. - -## 1. Run metadata - -| Field | Value | -|---|---| -| Start (UTC) | 2026-06-18T20:01:58Z | -| End (UTC) | 2026-06-18T21:25:36Z | -| Wall clock | 1h 23m | -| Engine commit | `cd68de0b4764b6836fe06ceb396e771cb7771468` | -| Engine config | `engine.e2e.local.toml` (rendered from `engine.e2e.toml`) | -| RPC provider | drpc.live (Sepolia WS) | -| Engine restarts | 2 (mid-run, to validate PR #47 — see §6.5) | -| Engine commits exercised | `5bcd47b` (pre-PR-47), `acc9654` (PR #47 twap-monitor), `cd68de0` (PR #47 ethflow-watcher) | - -## 2. Chain coverage - -| Chain | First block | Last block | Block delta | -|---|---|---|---| -| Sepolia (11155111) | 11089335 | 11089749 | 415 | - -Acceptance: block delta ≥ 1500 → **FAIL** - -## 3. On-chain actions submitted - -| Action | Tx | -|---|---| -| TWAP ComposableCoW.create() — script (t0=0 bug) | [0xa3d8a36f...4d02d](https://sepolia.etherscan.io/tx/0xa3d8a36f8a7dd8b097635ac59249b908d3f634bf5ede87c9336619e319e4d02d) | -| TWAP ComposableCoW.create() — cow-swap UI | [via UI; observed at block 11089497, indexed at 20:35:49Z, orderHash `0xc4bc4296...`](https://sepolia.etherscan.io/address/0xfdaFc9d1902f4e0b84f65F49f244b32b31013b74) | -| EthFlow.createOrder() — script (empty appData) | [0x622375d8...5731](https://sepolia.etherscan.io/tx/0x622375d89119df6419324ad4e5603688261fb01a4d47d717d686b6dd426b5731) | -| EthFlow.createOrder() — cow-swap UI (rich appData) | [0x82da5ced...b878](https://sepolia.etherscan.io/tx/0x82da5ceda6e28337625a991d4fc7db6b82a1695012b58a6b660ec92b8a88b878) | -| WETH-to-Safe transfer + GPv2VaultRelayer approve | manual via Safe UI (see §6.5) | -| WETH9.deposit() / setPreSignature for stop-loss | _(not run — stop-loss `submitted:` produced via PreSign-orderbook-accept path, see §6.3)_ | - -## 4. Per-module terminal-state markers - -| Module | First marker | Sample line | -|---|---|---| -| twap-monitor | 2026-06-18T20:07:36.495145Z | `indexed watch:0x7bf140727d27ea64b607e042f1225680b40eca6a:0x2ef7e76456176904e518b068744aad0e97a0d6...` | -| ethflow-watcher | 2026-06-18T20:14:00.841145Z | `ethflow backoff 0x104f25a0d633f9f39840723fc7e72a87d327829c9bc541a08ad9c8a62b9ecc9eba3cb449bd2b4ad...` | -| price-alert | 2026-06-18T20:02:10.605669Z | `price-alert: TRIGGERED answer=169974867813 threshold=250000000000 (Below)` | -| balance-tracker | 2026-06-18T20:02:10.772149Z | `balance-tracker 0x7bf140727d27ea64b607e042f1225680b40eca6a changed +50581434977874097 wei (prior=...` | -| stop-loss | 2026-06-18T20:02:12.874405Z | `stop-loss retry on next block (0): orderbook error (DuplicatedOrder): order already exists` | - -## 5. Error counts (Prometheus delta) - -| Metric | Start | End | Delta | -|---|---|---|---| -| `shepherd_event_latency_seconds_count{module="balance-tracker",event_kind="block"}` | 17 | 33 | 16 | -| `shepherd_event_latency_seconds_count{module="ethflow-watcher",event_kind="log"}` | 0 | 1 | 1 | -| `shepherd_event_latency_seconds_count{module="price-alert",event_kind="block"}` | 17 | 33 | 16 | -| `shepherd_event_latency_seconds_count{module="stop-loss",event_kind="block"}` | 17 | 33 | 16 | -| `shepherd_event_latency_seconds_count{module="twap-monitor",event_kind="block"}` | 17 | 33 | 16 | -| `shepherd_event_latency_seconds_sum{module="balance-tracker",event_kind="block"}` | 5.38369 | 9.72033 | 4.33664 | -| `shepherd_event_latency_seconds_sum{module="ethflow-watcher",event_kind="log"}` | 0 | 0.442872 | 0.442872 | -| `shepherd_event_latency_seconds_sum{module="price-alert",event_kind="block"}` | 2.86219 | 5.03446 | 2.17227 | -| `shepherd_event_latency_seconds_sum{module="stop-loss",event_kind="block"}` | 18.835 | 27.4352 | 8.60022 | -| `shepherd_event_latency_seconds_sum{module="twap-monitor",event_kind="block"}` | 0.0018655 | 56.1652 | 56.1633 | -| `shepherd_event_latency_seconds{module="balance-tracker",event_kind="block",quantile="0"}` | 0.310814 | 0.271721 | -0.0390927 | -| `shepherd_event_latency_seconds{module="balance-tracker",event_kind="block",quantile="0.5"}` | 0.334306 | 0.272832 | -0.0614738 | -| `shepherd_event_latency_seconds{module="balance-tracker",event_kind="block",quantile="0.9"}` | 0.334306 | 0.282889 | -0.0514163 | -| `shepherd_event_latency_seconds{module="balance-tracker",event_kind="block",quantile="0.95"}` | 0.334306 | 0.282889 | -0.0514163 | -| `shepherd_event_latency_seconds{module="balance-tracker",event_kind="block",quantile="0.99"}` | 0.334306 | 0.282889 | -0.0514163 | -| `shepherd_event_latency_seconds{module="balance-tracker",event_kind="block",quantile="0.999"}` | 0.334306 | 0.282889 | -0.0514163 | -| `shepherd_event_latency_seconds{module="balance-tracker",event_kind="block",quantile="1"}` | 0.347925 | 0.322888 | -0.0250366 | -| `shepherd_event_latency_seconds{module="price-alert",event_kind="block",quantile="0"}` | 0.141162 | 0.130526 | -0.0106367 | -| `shepherd_event_latency_seconds{module="price-alert",event_kind="block",quantile="0.5"}` | 0.165117 | 0.152575 | -0.0125423 | -| `shepherd_event_latency_seconds{module="price-alert",event_kind="block",quantile="0.9"}` | 0.165117 | 0.152727 | -0.0123897 | -| `shepherd_event_latency_seconds{module="price-alert",event_kind="block",quantile="0.95"}` | 0.165117 | 0.152727 | -0.0123897 | -| `shepherd_event_latency_seconds{module="price-alert",event_kind="block",quantile="0.99"}` | 0.165117 | 0.152727 | -0.0123897 | -| `shepherd_event_latency_seconds{module="price-alert",event_kind="block",quantile="0.999"}` | 0.165117 | 0.152727 | -0.0123897 | -| `shepherd_event_latency_seconds{module="price-alert",event_kind="block",quantile="1"}` | 0.199031 | 0.170941 | -0.0280894 | -| `shepherd_event_latency_seconds{module="stop-loss",event_kind="block",quantile="0"}` | 0.731767 | 0.680018 | -0.051749 | -| `shepherd_event_latency_seconds{module="stop-loss",event_kind="block",quantile="0.5"}` | 0.899515 | 0.719139 | -0.180375 | -| `shepherd_event_latency_seconds{module="stop-loss",event_kind="block",quantile="0.9"}` | 1.3033 | 0.719139 | -0.584161 | -| `shepherd_event_latency_seconds{module="stop-loss",event_kind="block",quantile="0.95"}` | 1.3033 | 0.719139 | -0.584161 | -| `shepherd_event_latency_seconds{module="stop-loss",event_kind="block",quantile="0.99"}` | 1.3033 | 0.719139 | -0.584161 | -| `shepherd_event_latency_seconds{module="stop-loss",event_kind="block",quantile="0.999"}` | 1.3033 | 0.719139 | -0.584161 | -| `shepherd_event_latency_seconds{module="stop-loss",event_kind="block",quantile="1"}` | 1.56857 | 0.740204 | -0.828361 | -| `shepherd_event_latency_seconds{module="twap-monitor",event_kind="block",quantile="0"}` | 8.2e-05 | 0.86952 | 0.869438 | -| `shepherd_event_latency_seconds{module="twap-monitor",event_kind="block",quantile="0.5"}` | 0.000110411 | 1.35921 | 1.35909 | -| `shepherd_event_latency_seconds{module="twap-monitor",event_kind="block",quantile="0.9"}` | 0.000110411 | 1.49466 | 1.49455 | -| `shepherd_event_latency_seconds{module="twap-monitor",event_kind="block",quantile="0.95"}` | 0.000110411 | 1.49466 | 1.49455 | -| `shepherd_event_latency_seconds{module="twap-monitor",event_kind="block",quantile="0.99"}` | 0.000110411 | 1.49466 | 1.49455 | -| `shepherd_event_latency_seconds{module="twap-monitor",event_kind="block",quantile="0.999"}` | 0.000110411 | 1.49466 | 1.49455 | -| `shepherd_event_latency_seconds{module="twap-monitor",event_kind="block",quantile="1"}` | 0.000132833 | 1.94945 | 1.94932 | -| `shepherd_chain_request_total{chain_id="11155111",method="eth_call",outcome="err"}` | 0 | 33 | 33 | -| `shepherd_chain_request_total{chain_id="11155111",method="eth_call",outcome="ok"}` | 34 | 100 | 66 | -| `shepherd_chain_request_total{chain_id="11155111",method="eth_getBalance",outcome="ok"}` | 34 | 66 | 32 | -| `shepherd_cow_api_submit_total{chain_id="11155111",outcome="err"}` | 17 | 67 | 50 | - -## 6. Anomalies + defects - -Four anomalies surfaced by this run. Each was filed as a separate issue. - -### 6.1 SDK + modules: non-empty `appData` hash rejected client-side - -**Status: fixed in this run via PR #47, live-validated in §6.5.** - -`twap-monitor` and `ethflow-watcher` strategies hard-coded -`EMPTY_APP_DATA_JSON` when assembling `OrderCreation`. Any -order with a richer `appData` (cow-swap UI orders carry -partner-id + slippage + quote-id metadata) hit -"app_data JSON digest does not match signed app_data hash" -client-side and was silently skipped. - -Pre-PR-47 evidence (block 11089387, before mid-run restart): -``` -INFO twap-monitor poll watch:0x14995a...:0xc4bc4296... -> Ready -INFO twap-monitor twap submit skipped for 0x14995a1118caf95833e923faf8dd155721cd53c2: - invalid OrderCreation: app_data JSON digest does not match signed app_data hash -``` - -Post-PR-47 (validated in §6.5): the submit body builds with -the matching JSON resolved from `GET /api/v1/app_data/{hash}`, -reaches the orderbook server, and rejects only on -server-side reasons (`DuplicatedOrder` for TWAP, since the UI -already submitted; `ExcessiveValidTo` for EthFlow — see §6.2). - -### 6.2 ethflow-watcher: `ExcessiveValidTo` from Sepolia orderbook - -**Status: open.** - -EthFlow on-chain orders carry `validTo = type(uint32).max` so -cancellation is operator-controlled via the EthFlow contract, -not orderbook-time-bounded. The Sepolia orderbook has a -max-validTo cap that rejects this shape. - -Evidence: -``` -WARN ethflow backoff 0x6d296984...ba3cb449bd2b4adddbc894d8697f5170800eadecffffffff - (0): orderbook error (ExcessiveValidTo): validTo is too far into the future -``` - -Last 4 bytes of UID = `ffffffff` = uint32::MAX. Pending -upstream investigation (Sepolia config drift vs mainnet -behaviour; needs cross-check before filing in -cowprotocol/services). - -### 6.3 stop-loss: `DuplicatedOrder` not classified as `Drop` - -**Status: open.** - -The stop-loss order from the E2E prep smoke (run earlier -on 2026-06-18) is still in the Sepolia orderbook (valid until -2106). The run-1 + run-2 stop-loss strategy re-submits the -same `OrderUid` on every block; orderbook responds -`DuplicatedOrder` (400); `shepherd_sdk::cow::classify_api_error` -maps to `TryNextBlock` and the retry loops forever (76 occurrences -in the first 170 blocks). - -Correct classification: `Drop` (the order is logically already -submitted; nothing to retry). PR sketch: -`crates/shepherd-sdk/src/cow/error.rs` `errorType` arm for -`DuplicatedOrder` → `RetryAction::Drop` + write -`submitted:{uid}` (or new `already-on-server:{uid}` marker). - -This run's stop-loss `submitted:` marker (via the PreSign- -upfront-accept path) was logged during the E2E prep smoke; -the marker persists in the orderbook and was observed -as `DuplicatedOrder` in this run. - -### 6.4 scripts/e2e-onchain.sh: TWAP `t0=0` produces permanently-finished order - -**Status: open.** - -`scripts/e2e-onchain.sh` hardcoded `t0=0` in the TWAP -`create()` calldata. TWAP `validateData` does NOT reject -t0=0 (only checks `t0 >= type(uint32).max`), so the create() -succeeds. But `TWAPOrderMathLib.calculateValidTo` computes -`part = (block.timestamp - 0) / t = ~3M`, which is `>= n=2`, -triggering `AFTER_TWAP_FINISHED` reverts on every -`getTradeableOrderWithSignature` poll. - -Evidence (custom error selector `0xc8fc2725` decoded): -``` -WARN twap-monitor eth_call failed (server returned an error response: - error code 3: execution reverted, data: "0xc8fc272500...616674657220747761702066696e6973686564" - [= ASCII "after twap finished"]) -``` - -Caller-side bug introduced by an AI-drafted helper. Fix is a -2-line edit to the encoder + a new comment; tracked separately. - -### 6.5 Live validation of PR #47 (this run's key methodology note) - -Mid-run, after observing §6.1, three engine binaries were -exercised back-to-back on the same `data/e2e` local-store -(restart preserved watches; no replay of past on-chain events -was needed — the indexed `watch:` keys in the redb survive -process restarts by design): - -| Engine commit | What it validates | -|---|---| -| `5bcd47b` (pre-PR-47) | Surfaces §6.1: twap-monitor + ethflow-watcher both log `submit skipped: digest does not match` for non-empty appData orders | -| `acc9654` (PR #47 twap-monitor) | After restart, the existing `watch:0x14995a...:0xc4bc4296...` (cow-swap UI TWAP) polled to Ready → resolve_app_data succeeded → submit reached orderbook → DuplicatedOrder (the order is already in the orderbook from the UI's original submission). **Client-side digest check was bypassed.** | -| `cd68de0` (PR #47 ethflow-watcher) | New cow-swap UI EthFlow swap submitted (tx `0x82da5ced...`); ethflow-watcher observes the OrderPlacement event with `order.appData = 0xe46e7d0c...` (NON-empty). resolve_app_data calls `GET /api/v1/app_data/0xe46e7d0c...` against the orderbook; orderbook returns `{"fullAppData": "{\"appCode\":\"CoW Swap\",\"environment\":\"production\",\"metadata\":{...,\"quote\":{\"slippageBips\":857,\"smartSlippage\":true}},...}"}`. The SDK extracts `fullAppData`; build_eth_flow_creation produces a body with matching digest; submit reaches orderbook; rejects only on ExcessiveValidTo (§6.2). **Client-side digest check was bypassed for ethflow-watcher too.** | - -The PR #47 fix is therefore live-validated end-to-end against -the real Sepolia orderbook in **both** affected modules. -Section 7's `block delta ≥ 1500` row is the only acceptance -row that does not clear; the engine was restarted twice for -this validation, totalling 415 blocks across the three -generations. A continuous 5h run with PR #47 included from -boot is the natural validation for the 7-day soak -rather than re-running the E2E. - -## 7. Acceptance checklist - -- [ ] block delta ≥ 1500 (got 415) -- [x] all 5 modules emitted ≥ 1 terminal-state marker -- [x] shepherd_module_errors_total{error_kind="trap"} == 0 (offenders: none) -- [x] no module poisoned at end (offenders: none) -- [x] 0 ERROR lines from nexum_engine::* (got 0) -- [x] TWAP + EthFlow on-chain txs submitted - -## 8. Sign-off (operator) - -> Auto-generated report. Operator: in 1-2 sentences confirm whether this run is clean enough to unblock the 7-day soak. If any acceptance row above is `[ ]`, file the defect before signing off. - -**Bruno (operator)** — _pending sign-off_ - -Recommended sign-off text (delete + replace as appropriate): - -> "Run validated the engine + 5-module dispatch path end-to-end against -> live Sepolia. Surfaced 4 anomalies (described in §6); the appData -> issue was fixed in-run via PR #47 and live-validated for both -> twap-monitor and ethflow-watcher (§6.5). Block delta short (415/1500) -> only because the run included two intentional restarts to validate -> the in-flight PR. **The 7-day soak is unblocked** to start on -> PR #47 merged + `feat/e2e-run-config` branch state; the -> other three follow-ups do not block the soak." - -## 9. Attachments - -- Engine log: `engine-combined-20260618.log` -- Metrics start: `metrics-start-20260618T200158Z.txt` -- Metrics end: `metrics-end-20260618T212514Z.txt` diff --git a/docs/operations/e2e-reports/e2e-report.template.md b/docs/operations/e2e-reports/e2e-report.template.md index 632fdac1..3a63cd71 100644 --- a/docs/operations/e2e-reports/e2e-report.template.md +++ b/docs/operations/e2e-reports/e2e-report.template.md @@ -1,9 +1,6 @@ -# E2E testnet integration report — YYYY-MM-DD +# E2E testnet integration report: YYYY-MM-DD -> Copy this file to `e2e-report-YYYY-MM-DD.md` in the same directory -> at the start of the run and fill it in as the run progresses. -> Sections marked **(operator)** must be filled in manually; the rest -> are derived from logs and `/metrics` snapshots. +> Copy to `e2e-report-YYYY-MM-DD.md` in this directory at the start of the run and fill in as it progresses. Sections marked **(operator)** are manual; the rest derive from logs and `/metrics` snapshots. ## 1. Run metadata @@ -15,8 +12,8 @@ | Wall clock | Hh Mm | | Engine commit | (`git rev-parse HEAD`) | | Engine config | `engine.e2e.toml` | -| Run host | (e.g. `bruno@bleu-mbp-m1`, `ec2-...`) | -| RPC provider | (alchemy / infura / publicnode / ...) | +| Run host | | +| RPC provider | | ## 2. Chain coverage @@ -24,8 +21,7 @@ |---|---|---|---|---| | Sepolia (11155111) | | | | | -Target: `block delta >= 1500` to clear the acceptance bar -(>= 1500 Sepolia blocks ≈ 5 h at 12 s block time). +Target: `block delta >= 1500` (>= 5 h at 12 s block time). ## 3. On-chain actions submitted by operator @@ -57,13 +53,11 @@ Target: `block delta >= 1500` to clear the acceptance bar | `sell_token` allowance tx hash | 0x... | | Owner EOA | 0x... | | Expected UID | 0x... | -| Expected detection | stop-loss logs `submitted:{uid}` once oracle trips | +| Expected detection | stop-loss logs `submitted:{uid}` once the oracle trips | ## 4. Per-module terminal-state markers -> Pull from the engine log with the JSON filter -> `jq 'select(.fields.message | test("submitted:|dropped:|backoff:|TRIGGERED|trapped"))'`. -> Each module must show at least ONE marker for the acceptance bar. +> Pull from the log with `jq 'select(.fields.message | test("submitted:|dropped:|backoff:|TRIGGERED|trapped"))'`. Each module must show at least one marker. | Module | First marker timestamp | Marker | Sample line | |---|---|---|---| @@ -75,15 +69,13 @@ Target: `block delta >= 1500` to clear the acceptance bar ## 5. Error counts (from `/metrics` delta) -> Capture two snapshots: at boot (`/metrics > metrics-start.txt`) and -> immediately before shutdown (`/metrics > metrics-end.txt`). Fill in -> the delta column. +> Snapshot at boot and immediately before shutdown; fill the delta column. | Metric | Start | End | Delta | |---|---|---|---| | `shepherd_module_errors_total{module="...",error_kind="trap"}` (per module) | | | | | `shepherd_module_restarts_total{module="..."}` (per module) | | | | -| `shepherd_module_poisoned{module="..."}` (gauge, end-state per module) | n/a | | n/a | +| `shepherd_module_poisoned{module="..."}` (gauge, end-state) | n/a | | n/a | | `shepherd_cow_api_submit_total{outcome="ok"}` | | | | | `shepherd_cow_api_submit_total{outcome="err"}` | | | | | `shepherd_chain_request_total{outcome="ok"}` | | | | @@ -94,9 +86,7 @@ Target: `block delta >= 1500` to clear the acceptance bar ## 6. Anomalies + defects -> Anything outside the expected log shape. Each anomaly that is -> reproducible OR has an unclear root cause must be filed as a -> separate issue and linked here. +> Each reproducible or unexplained anomaly is filed as a separate issue and linked here. | # | Time (UTC) | Module | Summary | |---|---|---|---| @@ -104,29 +94,21 @@ Target: `block delta >= 1500` to clear the acceptance bar ## 7. Acceptance checklist -- [ ] `block delta >= 1500` (≥ 5 h coverage) -- [ ] All 5 modules have ≥ 1 terminal-state marker in section 4 +- [ ] `block delta >= 1500` +- [ ] All 5 modules have >= 1 terminal-state marker in section 4 - [ ] `shepherd_module_errors_total{error_kind="trap"}` for well-behaved modules == 0 - [ ] No `[[modules]]`-listed module is `shepherd_module_poisoned == 1` at end -- [ ] No `ERROR` lines from `nexum_engine` in the supervisor log -- [ ] At least one orderbook submit attempt landed (`ok` or typed - `err` with retry/drop classification) on twap-monitor, - ethflow-watcher, AND stop-loss +- [ ] No `ERROR` lines from `nexum_runtime` in the supervisor log +- [ ] At least one orderbook submit attempt landed on twap-monitor, ethflow-watcher, and stop-loss - [ ] Report committed in this directory - [ ] Defects filed and linked in section 6 ## 8. Sign-off (operator) -> Brief paragraph: ran clean / found N defects / blocking issues for -> soak Y/N. The soak MUST NOT start until this -> section says "no blocking issues". - -… +> Ran clean / found N defects / blocking issues for the soak Y/N. The soak must not start until this says "no blocking issues". ## 9. Attachments -- `engine.log` (full supervisor JSON log; ≥ 4 h) +- `engine.log` (full supervisor JSON log) - `metrics-start.txt` - `metrics-end.txt` -- (optional) `metrics-snapshots/` — every 60 s scrape if a soak-style - Prometheus pull was not running diff --git a/docs/operations/e2e-testnet-runbook.md b/docs/operations/e2e-testnet-runbook.md index 3054b372..047334fc 100644 --- a/docs/operations/e2e-testnet-runbook.md +++ b/docs/operations/e2e-testnet-runbook.md @@ -1,52 +1,24 @@ # E2E testnet runbook -How to exercise **all 5 modules** — twap-monitor, ethflow-watcher, -price-alert, balance-tracker, stop-loss — on a real Sepolia host -**simultaneously for 4-6 hours**. Same shape as the M2 + M3 -runbooks, but this one runs the full production module suite and -captures a structured report (`docs/operations/e2e-reports/`). - -The E2E run is the integration step between unit-test coverage -(MockHost, per-module strategy tests) and the 7-day soak. -The soak validates *stability*; this validates *correctness in a -live dispatch context* and surfaces cross-module bugs the soak -should not be discovering. - -The acceptance bar is: - -- ≥ 1500 Sepolia blocks (≈ 5 h at 12 s block time). -- Each of the 5 modules writes at least one terminal-state marker - (`submitted:` / `dropped:` / `backoff:` / `TRIGGERED` / `last:`). +Runs all 5 production modules (twap-monitor, ethflow-watcher, price-alert, balance-tracker, stop-loss) on a live Sepolia host simultaneously for 4-6 h and captures a structured report under `docs/operations/e2e-reports/`. This is the correctness step between unit-test coverage and the 7-day soak. + +Acceptance bar: + +- >= 1500 Sepolia blocks (~5 h at 12 s block time). +- Each of the 5 modules writes at least one terminal-state marker (`submitted:` / `dropped:` / `backoff:` / `TRIGGERED` / `last:`). - 0 unexpected errors in the supervisor log. - 0 well-behaved modules trapped or poisoned at end of run. -- A committed report + filed defects. - ---- +- A committed report. ## 0. Prerequisites ### Toolchain -Same as the M2 + M3 runbooks (`rustup target add wasm32-wasip2`, -optionally `just`, a Sepolia WS RPC). +Same as the M2 + M3 runbooks (`rustup target add wasm32-wasip2`, `just`, a Sepolia WS RPC). ### RPC -The public Sepolia node (`wss://ethereum-sepolia-rpc.publicnode.com`) -throttles `eth_subscribe` and `eth_call` under sustained load. The -E2E run does at minimum: - -- 1 block subscription (shared across 4 modules — price-alert, - balance-tracker, stop-loss, twap-monitor block-tick). -- 2 log subscriptions (twap-monitor's - `ComposableCoW.ConditionalOrderCreated` + ethflow-watcher's - `CoWSwapEthFlow.OrderPlacement`). -- ≥ 4 `eth_call` per block from price-alert + balance-tracker - (×2 addresses) + stop-loss, + 1 per registered TWAP order - per block. - -Override the `[chains.11155111] rpc_url` in `engine.e2e.toml` -with an Alchemy / Infura WS for the run: +The public Sepolia node throttles `eth_subscribe` and `eth_call` under sustained load. The run holds 1 block subscription (shared across 4 modules), 2 log subscriptions (twap-monitor `ConditionalOrderCreated`, ethflow-watcher `OrderPlacement`), and >= 4 `eth_call` per block. Override `rpc_url` in `engine.e2e.toml` with an Alchemy / Infura WS: ```toml [chains.11155111] @@ -55,277 +27,190 @@ rpc_url = "wss://eth-sepolia.g.alchemy.com/v2/" ### On-chain prep (operator) -The acceptance bar requires real on-chain submissions. Before -launching the run, prepare: - -1. **A funded test EOA on Sepolia** (≥ 0.05 ETH for gas; the same - EOA can satisfy the EthFlow swap + stop-loss `setPreSignature` - sub-tasks). -2. **A Safe (or direct caller) that can call ComposableCoW** on - Sepolia — for the TWAP conditional-order submission. -3. **stop-loss config aligned with that EOA**: update - `modules/examples/stop-loss/module.toml::[config].owner` to the - EOA address you control, and pick a `sell_token` / `buy_token` - pair the EOA holds + has approved to the GPv2VaultRelayer. - See `docs/operations/m3-testnet-runbook.md` section 2 for the - full pre-sign + allowance recipe. +The acceptance bar requires real on-chain submissions. Prepare: -The E2E run will start cleanly without (1)/(2)/(3), but the -acceptance bar requires at least one `submitted:` marker on each -of twap-monitor / ethflow-watcher / stop-loss, and you only get -those by triggering each path on-chain. +1. A funded test EOA on Sepolia (>= 0.05 ETH for gas; also covers the EthFlow swap + stop-loss `setPreSignature`). +2. A Safe (or direct caller) that can call ComposableCoW, for the TWAP conditional-order submission. +3. stop-loss config aligned with that EOA: set `[config].owner` in `modules/examples/stop-loss/module.toml` to the EOA, and pick a `sell_token` / `buy_token` pair the EOA holds and has approved to the GPv2VaultRelayer (M3 runbook section 2 has the pre-sign + allowance recipe). ---- +The run boots without (1)/(2)/(3), but the acceptance bar needs one `submitted:` marker on each of twap-monitor / ethflow-watcher / stop-loss, which only on-chain triggers produce. ## 1. Boot -The engine + all 5 modules + Prometheus `/metrics` endpoint: - ```bash just run-e2e ``` -Equivalent long form: +Long form: ```bash just build-e2e # builds the 5 module .wasm artefacts -cargo build -p nexum-cli -cargo run -p nexum-cli -- --engine-config engine.e2e.toml +cargo run -p shepherd -- --engine-config engine.e2e.toml ``` -### Expected boot sequence (~5 s) +Expected boot (~5 s) ends with: ``` -INFO nexum starting -INFO opening chain RPC provider chain_id=11155111 url="wss://..." INFO metrics exporter listening at /metrics addr=127.0.0.1:9100 -INFO loading module manifest manifest=modules/twap-monitor/module.toml -INFO compiling component component=...twap_monitor.wasm INFO init succeeded module=twap-monitor -INFO loading module manifest manifest=modules/ethflow-watcher/module.toml INFO init succeeded module=ethflow-watcher -INFO loading module manifest manifest=modules/examples/price-alert/module.toml INFO init succeeded module=price-alert -INFO loading module manifest manifest=modules/examples/balance-tracker/module.toml INFO init succeeded module=balance-tracker -INFO loading module manifest manifest=modules/examples/stop-loss/module.toml INFO init succeeded module=stop-loss -INFO supervisor up count=5 INFO supervisor ready modules=5 chains=1 -INFO block subscription open chain_id=11155111 INFO log subscription open chain_id=11155111 module=twap-monitor INFO log subscription open chain_id=11155111 module=ethflow-watcher ``` -If any of `count=5`, `modules=5`, or both log subscriptions are -missing, **stop the run and triage** — running 4-6 h on a -degraded engine wastes time the operator does not get back. - -### Smoke at first block (~12 s after boot) - -Within the first Sepolia block dispatched: +If `modules=5` or either log subscription is missing, stop the run and triage before committing to 4-6 h. -``` -DEBUG dispatch block chain_id=11155111 number=N -DEBUG chain::request method=eth_call # price-alert oracle read -DEBUG chain::request method=eth_getBalance # balance-tracker addr 1 -DEBUG chain::request method=eth_getBalance # balance-tracker addr 2 -DEBUG chain::request method=eth_call # stop-loss oracle read -WARN price-alert: TRIGGERED answer=... threshold=... -``` +## 2. The run -(See `docs/operations/m3-testnet-runbook.md` for the per-module -single-block expectations — the E2E run reproduces those plus -twap-monitor's empty poll loop until a `watch:` is registered.) - ---- - -## 2. The 4-6 h run - -### 2.1 Start the clock - -Pipe the engine output to a JSON log file the operator can mine -with `jq` after the run: +### 2.1 Start the clock and baseline ```bash just run-e2e 2>&1 | tee -a docs/operations/e2e-reports/engine-$(date -u +%Y%m%dT%H%M%SZ).log +curl -s http://127.0.0.1:9100/metrics > docs/operations/e2e-reports/metrics-start.txt ``` -Record `date -u --iso-8601=seconds` and `git rev-parse HEAD` in -section 1 of the report template. +Record `date -u --iso-8601=seconds` and `git rev-parse HEAD` in section 1 of the report. -### 2.2 Capture the metrics baseline +### 2.2 Trigger each on-chain action -```bash -curl -s http://127.0.0.1:9100/metrics > docs/operations/e2e-reports/metrics-start.txt -``` +Run as soon as the supervisor is `ready`: + +1. **TWAP order**: call ComposableCoW from the Safe. Within 1-2 blocks: `INFO twap-monitor watch:{orderHash}`. +2. **EthFlow swap**: execute a small ETH-flow swap from the EOA via the CoW Swap front-end on Sepolia. Within 1-2 blocks: `INFO ethflow-watcher submitted:{uid}` (or a typed `dropped:{uid}`, both terminal markers). +3. **stop-loss trigger**: once the owner EOA has called `setPreSignature` and approved the sell token, lower `trigger_price` in `modules/examples/stop-loss/module.toml` to <= the current Chainlink ETH/USD answer and reload. Within 1 block: `INFO stop-loss TRIGGERED` then `submitted:{uid}`. + +### 2.3 Idle until end of run -### 2.3 Trigger each on-chain action - -Run these as soon as the supervisor is `ready`: - -1. **TWAP order** — call ComposableCoW from your Safe (or directly - if you control the user). Within 1-2 blocks, twap-monitor logs: - ``` - INFO twap-monitor watch:{orderHash} chain_id=11155111 - ``` -2. **EthFlow swap** — execute a small ETH-flow swap from your EOA - via the cow-swap front-end pointed at Sepolia. Within 1-2 blocks - ethflow-watcher logs: - ``` - INFO ethflow-watcher submitted:{uid} - ``` - (or a typed `dropped:{uid}` if the orderbook rejected — both - count as a terminal-state marker for section 4.) -3. **stop-loss trigger** — once your owner EOA has called - `setPreSignature` and approved the sell token, lower - `trigger_price` in `modules/examples/stop-loss/module.toml` to - ≤ the current Sepolia Chainlink ETH/USD answer and reload the - engine (or set it pre-boot if you already know the feed value). - Within 1 block stop-loss logs: - ``` - INFO stop-loss TRIGGERED price=... trigger=... - INFO stop-loss submitted:{uid} - ``` - -### 2.4 Idle until end of run - -Once all three terminal markers are observed and the report's -section 4 has at least one entry per module, leave the engine -running undisturbed for the remainder of the 4-6 h window. - -The operator should watch for these red flags (if any appears, -the run is a defect and section 6 must capture it): +Once all three markers are observed, leave the engine undisturbed for the remainder of the window. Red flags (each is a defect for report section 6): | Red flag | Why it matters | |---|---| -| `ERROR` from `nexum_runtime::*` | Acceptance #5: zero ERROR lines. | -| `module ... trapped:` for a non-fixture module | Trapping production-side modules is a defect. | -| `module ... poisoned` | Quarantine of a real module is a defect. | -| `stream reconnect attempt=N` with N rising | The WS is flapping (RPC issue or bug). One reconnect per chain is fine. | -| `chain::request` `err` rate > 5% | The RPC is degraded. Switch keys / providers. | +| `ERROR` from `nexum_runtime::*` | Acceptance: zero ERROR lines | +| `module ... trapped:` for a non-fixture module | Trapping a production module is a defect | +| `module ... poisoned` | Quarantine of a real module is a defect | +| `stream reconnect attempt=N` with N rising | WS flapping. One reconnect per chain is fine | +| `chain::request` `err` rate > 5% | RPC degraded. Switch keys / providers | -### 2.5 Capture metrics deltas + shutdown - -At the end of the run window: +### 2.4 Capture deltas and shut down ```bash curl -s http://127.0.0.1:9100/metrics > docs/operations/e2e-reports/metrics-end.txt -# Ctrl-C the engine — graceful shutdown writes last_dispatched_block: -# > INFO graceful shutdown complete dispatched_blocks=N dispatched_logs=M uptime_secs=K -``` - -Diff the two snapshots to fill in the report's section 5: - -```bash +# Ctrl-C: graceful shutdown logs `dispatched_blocks=N dispatched_logs=M uptime_secs=K`. diff <(grep '^shepherd_' docs/operations/e2e-reports/metrics-start.txt) \ <(grep '^shepherd_' docs/operations/e2e-reports/metrics-end.txt) ``` ---- +## 3. Report -## 3. Filling in the report - -Copy the template at the start of the run: +Copy the template at the start of the run and fill it as the run progresses: ```bash DATE=$(date -u +%Y-%m-%d) -cp docs/operations/e2e-reports/e2e-report.template.md \ - docs/operations/e2e-reports/e2e-report-${DATE}.md -$EDITOR docs/operations/e2e-reports/e2e-report-${DATE}.md -``` - -Fill sections in this order: - -1. **Section 1 (run metadata)** at boot. -2. **Section 3 (on-chain actions)** as you submit each one. -3. **Section 4 (terminal markers)** as each first marker fires. -4. **Section 5 (metrics)** once `metrics-end.txt` is captured. -5. **Section 6 (anomalies)** continuously — anything unexpected - gets a row + an issue. -6. **Section 7 (acceptance checklist)** at the end — every box - must be `[x]` for the run to pass. -7. **Section 8 (sign-off)** is the gating decision for the - 7-day soak. - -Commit the filled-in report on the same branch as this runbook: - -```bash -git add docs/operations/e2e-reports/e2e-report-${DATE}.md -git commit -m "ops(e2e): report from ${DATE} run" -git push +cp docs/operations/e2e-reports/e2e-report.template.md docs/operations/e2e-reports/e2e-report-${DATE}.md ``` ---- - -## 4. What this does NOT prove - -- **Stability beyond ~5 h** → the 7-day soak (Sepolia + Arb Sepolia). -- **Adversarial resource exhaustion** → a fuel/memory adversarial fixtures run (M4 territory). -- **Security review** → tracked separately. -- **Production deployment story** → `docs/production.md`. -- **Multi-chain isolation under live WS drops** → partially - proven by integration tests; full validation - requires Arb Sepolia + Sepolia simultaneously, which the soak - exercises. - ---- +Every acceptance box in the template's section 7 must be `[x]` for the run to pass. Commit the filled report on this branch. -## 5. Troubleshooting - -Inherits the M2 + M3 runbook tables. E2E-specific: +## 4. Troubleshooting | Symptom | Likely cause | Fix | |---|---|---| -| `supervisor ready modules=4 chains=1` (or less) at boot | One of the 5 module manifests failed to load — likely a missing wasm artefact under `target/wasm32-wasip2/release/` | Re-run `just build-e2e` and verify all 5 `.wasm` files are present. | -| `INFO log subscription open chain_id=11155111` appears only once | One of the two log-subscribing modules failed init | Check the immediately preceding `init failed module=...` line; the failing module's `[capabilities]` or subscription `address` is the usual culprit. | -| RPC drops every ~30 min on `publicnode.com` | Public node rate limits | Switch to Alchemy / Infura per section 0. | -| `stop-loss TRIGGERED` fires immediately on default config | Default `trigger_price = 2500.00` is above Sepolia Chainlink ETH/USD (~$1745) and `direction = "below"`. See M3 runbook §1. | Tune `trigger_price` lower to test the "silent until trigger" path. | -| `twap-monitor` never logs `watch:` | No `ConditionalOrderCreated` event observed on Sepolia during the window | Submit the TWAP order from section 2.3 step 1. | -| `ethflow-watcher` never logs `submitted:` | No `OrderPlacement` event observed on Sepolia during the window | Execute the EthFlow swap from section 2.3 step 2. | +| `supervisor ready modules=4 chains=1` at boot | A module manifest failed to load, likely a missing wasm artefact | Re-run `just build-e2e`; verify all 5 `.wasm` present | +| Only one `log subscription open` | One log-subscribing module failed init | Check the preceding `init failed module=...`; usual culprit is `[capabilities]` or the subscription `address` | +| RPC drops every ~30 min on `publicnode.com` | Public node rate limits | Switch to Alchemy / Infura | +| `stop-loss TRIGGERED` fires immediately | Default `trigger_price` above the feed with `direction = below` | Tune `trigger_price` lower | +| `twap-monitor` never logs `watch:` | No `ConditionalOrderCreated` observed | Submit the TWAP order (2.2 step 1) | +| `ethflow-watcher` never logs `submitted:` | No `OrderPlacement` observed | Execute the EthFlow swap (2.2 step 2) | ---- +## 5. Known Sepolia constraint: EthFlow `validTo = u32::MAX` -## 5.5. Known upstream constraints on Sepolia +EthFlow on-chain orders carry `validTo = type(uint32).max` by design (cancellation is operator-controlled via the EthFlow contract). The Sepolia orderbook's max-validTo cap rejects this shape with `errorType = "ExcessiveValidTo"`, so every EthFlow placement on Sepolia terminates as `Drop`. The strategy recognises this and degrades gracefully: -These are not bugs in shepherd; they are documented gaps between -the on-chain protocol and the Sepolia orderbook's validation -config. The strategy code recognises each and degrades gracefully -(Drop, not retry storm). The soak report should call them out so -the reader does not file them as anomalies. +- `ethflow dropped (400): orderbook error (ExcessiveValidTo)...` at Info level. +- `dropped:{uid}` written once per placement. +- `shepherd_cow_api_submit_total{outcome="err"}` grows by exactly the EthFlow placement count, then stops. -### EthFlow `validTo = u32::MAX` → `ExcessiveValidTo` +This is a testnet orderbook constraint, not a bug; the report should note it so it is not filed as an anomaly. -EthFlow on-chain orders carry `validTo = type(uint32).max` by -design: cancellation is operator-controlled via the EthFlow -contract, not orderbook-time-bounded. `cowprotocol::eth_flow` -documents this as the canonical CoW-side shape on every chain. +## 6. Re-deriving pinned values -The Sepolia orderbook's max-validTo cap rejects this shape with -`errorType = "ExcessiveValidTo"`. Every `POST /api/v1/orders` -ethflow-watcher forwards on Sepolia therefore terminates as -`Drop` (since the host fix; before that fix the same case -manifested as an infinite `backoff:` loop). +If the pinned identities in a run config drift, regenerate. -Operator-visible behaviour after the strategy refinement: +### OrderUid -- `ethflow dropped (400): orderbook error (ExcessiveValidTo)...` -- Log level: **Info** (not Warn). -- `dropped:{uid}` marker written exactly once per placement. -- The soak's Prometheus - `shepherd_cow_api_submit_total{outcome="err"}` curve grows by - exactly the EthFlow placement count, then stops. +```bash +cargo test -p stop-loss --lib cow_1064 -- --nocapture +``` -Upstream confirmation with the cowprotocol/services team is -pending; if mainnet also rejects this shape the design needs -revisiting at the contract level (which is out of scope for -shepherd). +Asserts against the constants in `module.toml`; fails loudly if the EIP-712 type-hash or domain separator shifts. Raw Python equivalent: + +```python +from eth_utils import keccak + +DOMAIN_SEP = bytes.fromhex("daee378bd0eb30ddf479272accf91761e697bc00e067a268f95f1d2732ed230b") +SELL_TOKEN = bytes.fromhex("fFf9976782d46CC05630D1f6eBAb18b2324d6B14") +BUY_TOKEN = bytes.fromhex("0625aFB445C3B6B7B929342a04A22599fd5dBB59") +OWNER = bytes.fromhex("7bF140727D27ea64b607E042f1225680B40ECa6A") +RECEIVER = OWNER +SELL_AMOUNT = 5_000_000_000_000_000 +BUY_AMOUNT = 20_000_000_000_000_000_000 +VALID_TO = 4_294_967_295 + +APP_DATA = bytes.fromhex("b48d38f93eaa084033fc5970bf96e559c33c4cdc07d889ab00b4d63f9590739d") # keccak("{}") +KIND_SELL = keccak(b"sell") +ERC20 = keccak(b"erc20") +TYPE_HASH = keccak(b"Order(address sellToken,address buyToken,address receiver,uint256 sellAmount,uint256 buyAmount,uint32 validTo,bytes32 appData,uint256 feeAmount,string kind,bool partiallyFillable,string sellTokenBalance,string buyTokenBalance)") +pad32 = lambda b: bytes(32-len(b)) + b +uint = lambda v: v.to_bytes(32, "big") +struct_hash = keccak( + TYPE_HASH + pad32(SELL_TOKEN) + pad32(BUY_TOKEN) + pad32(RECEIVER) + + uint(SELL_AMOUNT) + uint(BUY_AMOUNT) + uint(VALID_TO) + + APP_DATA + uint(0) + KIND_SELL + + b"\x00"*32 + ERC20 + ERC20 # partiallyFillable=false +) +order_digest = keccak(b"\x19\x01" + DOMAIN_SEP + struct_hash) +uid = order_digest + OWNER + VALID_TO.to_bytes(4, "big") +print("0x" + uid.hex()) +``` ---- +### ComposableCoW.create() calldata + +Generate locally with `python3 scripts/_twap_calldata.py` (never paste a pinned blob). The helper backdates `t0` by 60 s per invocation so part 0 is Ready immediately; `t0` must never be 0 or every poll reverts `AFTER_TWAP_FINISHED`. Constants: + +```python +import time +from eth_utils import keccak +from eth_abi import encode + +selector = keccak(b"create((address,bytes32,bytes),bool)")[:4] +static = encode( + ["(address,address,address,uint256,uint256,uint256,uint256,uint256,uint256,bytes32)"], + [( + "0xfFf9976782d46CC05630D1f6eBAb18b2324d6B14", # sellToken + "0x0625aFB445C3B6B7B929342a04A22599fd5dBB59", # buyToken + "0x14995a1118Caf95833e923faf8Dd155721cd53c2", # receiver + 1_000_000_000_000_000, 500_000_000_000_000_000, # partSellAmount, minPartLimit + int(time.time()) - 60, 2, 600, 0, # t0 (never 0), n, t, span + b"\x00" * 32, # appData + )] +) +calldata = selector + encode( + ["(address,bytes32,bytes)", "bool"], + [( + "0x6cF1e9cA41f7611dEf408122793c358a3d11E5a5", # TWAP handler + bytes.fromhex("000000000000000000000000000000000000000000000000000000006670f000"), # salt + static, + ), True] +) +print("0x" + calldata.hex()) +``` -## 6. References +## 7. References -- M2 runbook (sister doc): `docs/operations/m2-testnet-runbook.md` -- M3 runbook (sister doc): `docs/operations/m3-testnet-runbook.md` +- M2 + M3 runbooks (sister docs) - Engine config: `engine.e2e.toml` - Report template: `docs/operations/e2e-reports/e2e-report.template.md` diff --git a/docs/operations/load-reports/load-20x20-2026-06-19.md b/docs/operations/load-reports/load-20x20-2026-06-19.md deleted file mode 100644 index 9495fe36..00000000 --- a/docs/operations/load-reports/load-20x20-2026-06-19.md +++ /dev/null @@ -1,84 +0,0 @@ -# Load test report - medium 20x20 - -## 1. Run metadata - -| Field | Value | -|---|---| -| Stamp (UTC) | 2026-06-19T16:03:24Z | -| Wall clock | 120 s (2 min) | -| Engine commit | `feat/load-gen-calibration` head | -| Anvil command | `anvil --fork-url $RPC_URL_SEPOLIA_HTTP --port 8545 --block-time 1` | -| Mock orderbook | `tools/orderbook-mock --port 9999` | -| Modules | `twap-monitor`, `ethflow-watcher` | -| Scenario | medium (20 TWAP + 20 EthFlow per block, 2 min) | - -## 2. Load generator output - -``` -load-gen finished blocks_seen=14 - twap_attempted=280 twap_ok=280 - ethflow_attempted=280 ethflow_ok=280 -``` - -The `blocks_seen=14` is load-gen's perspective - it processes the next block only after finishing the previous burst of 40 submissions. Anvil itself mined **128 blocks** during the run (per `shepherd_event_latency_seconds_count{event_kind="block"}`), so shepherd's supervisor fired 128 block-dispatch cycles. - -## 3. Engine throughput - -| Metric | Delta | Notes | -|---|---|---| -| `shepherd_event_latency_seconds_count{module="twap-monitor",event_kind="block"}` | **128** | One per Anvil block. | -| `shepherd_event_latency_seconds_count{module="twap-monitor",event_kind="log"}` | **280** | 1:1 with load-gen. | -| `shepherd_event_latency_seconds_count{module="ethflow-watcher",event_kind="log"}` | **280** | 1:1 with load-gen. | -| `shepherd_cow_api_submit_total{outcome="ok"}` | **280** | All EthFlow submissions reached the mock orderbook successfully. | -| `shepherd_cow_api_submit_total{outcome="err"}` | **0** | Zero. | -| `shepherd_chain_request_total{method="eth_call",outcome="err"}` | **18 442** | Watch polls reverting (no settle-time allowance); strategy correctly classifies as TryNextBlock. | -| `shepherd_module_errors_total` | **0** | Zero. | - -### Latency - -**twap-monitor block (poll loop over all 280 active watches):** - -| Quantile | Value | -|---|---| -| p50 | 56 ms | -| p95 | 66 ms | -| p99 | 67 ms | -| max | 67 ms | - -**ethflow-watcher log:** p50/p95/p99 = 8 / 9.5 / 12 ms. - -Engine-log-derived dispatch_block max: 471 ms (cold-start outlier, same pattern as the baseline). - -## 4. Mock orderbook stats - -``` -submits_ok = 280 -submits_err = 0 -app_data_lookups = 0 -``` - -## 5. Acceptance vs. medium bar - -| Criterion | Observed | Pass? | -|---|---|---| -| 20 TWAP + 20 EthFlow events delivered per load-gen iteration | 280 + 280 across 14 iterations = exactly 20 per iteration | **PASS** | -| Graceful degradation (`backoff:` markers OK; `shepherd_module_errors_total = 0`) | zero module_errors_total | **PASS** | -| `cow_api_submit{outcome="err"}` stays 0 | zero | **PASS** | -| Zero traps | zero | **PASS** | -| p99 < 2 s (informal carry-over from baseline) | TWAP block p99 = 67 ms | **PASS** (30x margin) | - -**Medium: full PASS.** - -## 6. Scaling observation - -Compared to the baseline (130 watches → 49 ms p99) the medium run holds 280 watches → 67 ms p99 - **sub-linear growth** in dispatch latency, not the strict linear scaling extrapolated earlier. Encouraging signal for the saturation scenario. - -## 7. Attachments - -- Metrics start: `/tmp/shepherd-load/metrics-start-20260619T160324Z.txt` -- Metrics end: `/tmp/shepherd-load/metrics-end-20260619T160324Z.txt` -- Engine + load-gen logs under `/tmp/shepherd-load/`. - -## 8. Sign-off - -**Bruno (operator) - PASS, medium.** Engine handles 20+20 events per load-gen iteration with the same 30x latency margin as baseline, zero errors. Saturation scenario unblocked. diff --git a/docs/operations/load-reports/load-50x50-2026-06-19.md b/docs/operations/load-reports/load-50x50-2026-06-19.md deleted file mode 100644 index 3a2266ff..00000000 --- a/docs/operations/load-reports/load-50x50-2026-06-19.md +++ /dev/null @@ -1,113 +0,0 @@ -# Load test report - saturation 50x50 - -## 1. Run metadata - -| Field | Value | -|---|---| -| Stamp (UTC) | 2026-06-19T16:08:51Z | -| Wall clock | 120 s (2 min) | -| Engine commit | `feat/load-gen-calibration` head | -| Anvil command | `anvil --fork-url $RPC_URL_SEPOLIA_HTTP --port 8545 --block-time 1` | -| Mock orderbook | `tools/orderbook-mock --port 9999` | -| Modules | `twap-monitor`, `ethflow-watcher` | -| Scenario | saturation (50 TWAP + 50 EthFlow per block, 2 min) | - -## 2. Load generator output - -``` -load-gen finished blocks_seen=6 - twap_attempted=300 twap_ok=300 - ethflow_attempted=300 ethflow_ok=300 -``` - -300 + 300 events delivered. Anvil mined **138 blocks** during the -run (per `shepherd_event_latency_seconds_count{event_kind="block"}`) -- the load-gen's `blocks_seen=6` is its own perspective (the burst of -100 sequential tx submissions per iteration takes ~20 s per round, -so it only processes 6 block-tick events from the WS subscription -during the 120 s window). - -## 3. Engine throughput - -| Metric | Delta | Notes | -|---|---|---| -| `shepherd_event_latency_seconds_count{module="twap-monitor",event_kind="block"}` | **138** | One per Anvil block. | -| `shepherd_event_latency_seconds_count{module="twap-monitor",event_kind="log"}` | **300** | 1:1 with load-gen. | -| `shepherd_event_latency_seconds_count{module="ethflow-watcher",event_kind="log"}` | **300** | 1:1 with load-gen. | -| `shepherd_cow_api_submit_total{outcome="ok"}` | **300** | All EthFlow submissions reached the mock orderbook successfully. | -| `shepherd_cow_api_submit_total{outcome="err"}` | **0** | Zero. | -| `shepherd_chain_request_total{method="eth_call",outcome="err"}` | **22 137** | Watch polls (300 watches × ~74 blocks). | -| `shepherd_module_errors_total` | **0** | Zero. | - -### Latency - -**twap-monitor block (poll loop over all 300 active watches):** - -| Quantile | Value | -|---|---| -| p50 | 67 ms | -| p95 | 76 ms | -| p99 | 78 ms | -| max | 88 ms | - -**ethflow-watcher log:** p50/p95/p99 = 8.0 / 8.1 / 8.9 ms - basically flat vs. baseline. - -**twap-monitor log:** p50/p95/p99 = 4.0 / 5.9 / 7.0 ms. - -Engine-log-derived dispatch_block max: 497 ms (cold-start outlier, same pattern). - -## 4. Mock orderbook stats - -``` -submits_ok = 300 -submits_err = 0 -app_data_lookups = 0 -``` - -## 5. Acceptance vs. saturation bar - -| Criterion | Observed | Pass? | -|---|---|---| -| 50 TWAP + 50 EthFlow events delivered per load-gen iteration | 300 + 300 across 6 iterations = exactly 50 per iteration | **PASS** | -| Identify the bottleneck | **Bottleneck is on the load-gen side, not the engine** - see §6 | (informative) | -| `shepherd_module_errors_total = 0` | zero | **PASS** | -| Zero traps | zero | **PASS** | - -**Saturation: PASS - and the test did NOT saturate the engine.** - -## 6. The unexpected finding: engine did not saturate - -The hypothesis going in (informed by lgahdl's PR #9 thread on sequential per-module dispatch) was that 50x50 would push the supervisor past its single-module dispatch budget and surface a per-block latency outlier or a backlog. None of that happened: - -- TWAP block p99 grew from 49 ms (130 watches, baseline) to **78 ms (300 watches, saturation)** - **sub-linear growth.** -- EthFlow log p99 held at **8.9 ms** across all three scenarios - the submit-to-mock round-trip is dominated by the network hop, not engine bookkeeping. -- Zero `shepherd_module_errors_total`, zero traps, zero backoff: markers. -- The cold-start outlier (~500 ms on the first watch-heavy block) is consistent across runs and does not scale with the watch count - it's a one-shot first-block redb / eth_call warmup cost. - -**Actual bottleneck:** load-gen's sequential `eth_sendTransaction` submission. At 100 tx/iteration (50+50) and ~200 ms per submission roundtrip, each iteration takes ~20 s, vs. Anvil's 1 s block time. So the load-gen processes 6 block-events of its own but Anvil mines 138 blocks during the same window. The engine handles those 138 dispatch cycles cleanly. - -### Implications - -1. **lgahdl's sequential-dispatch concern**: not surfaced at this scale (300 watches, 138 dispatch cycles in 2 min). To genuinely test it would require an order of magnitude more watches (3 000 - 10 000) or parallel load generators. -2. **What this proves**: shepherd's M4 supervisor handles **at least 300 concurrent watches and 138 block-dispatch cycles in 2 min** with p99 < 80 ms and zero errors. -3. **What it does NOT prove**: behaviour at 3 000+ watches, behaviour under real-network RPC variability, behaviour over 7 days (the 7-day soak's actual job). - -## 7. Followups - -The bottleneck shifted from engine to load-gen; to actually saturate the engine, future iterations should: - -1. **Run multiple load-gens in parallel**, each impersonating a different EOA (Anvil supports arbitrary impersonation), so the per-EOA nonce serialisation does not gate the throughput. -2. **Use a smaller `--block-time`** (e.g. 100 ms) so blocks emit faster and the engine has to handle more dispatch cycles per second. -3. **Pre-seed thousands of watches via direct redb writes** before starting the dispatch loop, then run a small steady-state load - this isolates dispatch cost from indexing cost. - -These are not blocking for the acceptance sign-off; this is a saturation **target**, not a saturation **failure**. The acceptance bar ("identify the bottleneck") is met: bottleneck identified, on the test-tool side, not the engine side. - -## 8. Attachments - -- Metrics start: `/tmp/shepherd-load/metrics-start-20260619T160851Z.txt` -- Metrics end: `/tmp/shepherd-load/metrics-end-20260619T160851Z.txt` -- Engine + load-gen logs under `/tmp/shepherd-load/`. - -## 9. Sign-off - -**Bruno (operator) - PASS, saturation.** Engine handles 50+50 events per load-gen iteration without flinching. Bottleneck is the test tool, not the engine. The three-scenario acceptance sweep is complete. diff --git a/docs/operations/load-reports/load-50x50-parallel-2026-06-19.md b/docs/operations/load-reports/load-50x50-parallel-2026-06-19.md deleted file mode 100644 index afb3665a..00000000 --- a/docs/operations/load-reports/load-50x50-parallel-2026-06-19.md +++ /dev/null @@ -1,148 +0,0 @@ -# Load test report - aggressive saturation (10 workers, 0.5s blocks) - -> The saturation push the prior `load-50x50` report flagged as "engine -> did not saturate, the bottleneck is on the load-gen side". This run -> removes both load-gen-side limits and finds the engine's actual -> saturation knee. - -## 1. Run metadata - -| Field | Value | -|---|---| -| Stamp (UTC) | 2026-06-19T17:05:40Z | -| Wall clock | 120 s (2 min) | -| Engine commit | `feat/load-gen-calibration` head | -| Anvil command | `anvil --fork-url $RPC_URL_SEPOLIA_HTTP --port 8545 --block-time 0.5` | -| Mock orderbook | `tools/orderbook-mock --port 9999` | -| Modules | `twap-monitor`, `ethflow-watcher` | -| Scenario | saturation-parallel (10 workers × (5 TWAP + 5 EthFlow) per block, `--block-time 0.5`, 2 min) | -| load-gen flags | `--parallel 10 --twap-per-block 5 --ethflow-per-block 5 --block-time 0.5 --duration-min 2` | - -The parallel-mode flag is new: each worker impersonates its own synthetic EOA (`0x57...01` … `0x57...0a`), has its own WS connection + nonce stream, runs its own per-block submission loop. Removes the per-EOA nonce serialisation bottleneck the single-worker saturation report (`load-50x50-2026-06-19.md`) identified. - -## 2. Load generator output - -``` -load-gen finished workers_finished=10 blocks_seen=179 - twap_attempted=895 twap_ok=895 - ethflow_attempted=895 ethflow_ok=895 -``` - -895 TWAP + 895 EthFlow `eth_sendTransaction` acks across 10 workers; zero load-gen-side errors (the first attempt at this run had a sellAmount-overflow bug that blew past the EOA's 1M ETH balance; fixed by namespacing `ethflow_seq` to a 10 000-wide per-worker window). - -## 3. Engine throughput - the saturation signal - -| Metric | Delta | Notes | -|---|---|---| -| `shepherd_event_latency_seconds_count{module="twap-monitor",event_kind="block"}` | **110** | Block events dispatched. With `--block-time 0.5` we expected ~240; the engine saw 110 - **the block stream itself dropped under load**, see §4. | -| `shepherd_event_latency_seconds_count{module="twap-monitor",event_kind="log"}` | **381** | `ConditionalOrderCreated` events delivered. load-gen submitted 895 → only **43%** reached the engine. | -| `shepherd_event_latency_seconds_count{module="ethflow-watcher",event_kind="log"}` | **343** | load-gen submitted 895 → **38%** reached the engine. | -| `shepherd_cow_api_submit_total{outcome="ok"}` | **343** | Matches EthFlow events 1:1 - the engine submitted every event it saw. | -| `shepherd_cow_api_submit_total{outcome="err"}` | **0** | Zero submit errors. | -| `shepherd_chain_request_total{method="eth_call",outcome="err"}` | **31 097** | Watch polls (381 watches × ~80 effective dispatch cycles). | -| `shepherd_module_errors_total` | **0** | Engine never traps. | - -### Latency (Prometheus histogram) - -**twap-monitor block dispatch:** - -| Quantile | Value | -|---|---| -| p50 | **145 ms** | -| p95 | 145 ms | -| p99 | 145 ms | -| **max** | **101 593 ms** ≈ 101 s | - -(The histogram bucketing collapses p50-p99 to the same value because the sample is sparse + bucket-bounded; the `max` is the meaningful upper tail.) - -Engine-log-derived dispatch_block (more granular): -- n = 586 dispatches -- p50 = 4 ms -- p95 = 46 ms -- p99 = 74 ms -- **max = 101 593 ms** (the same 101-second outlier the histogram caught) - -**twap-monitor log + ethflow-watcher log:** histogram-buckets to 0 across all quantiles - per-event indexing + submit completed in < 1 ms even at the peak. The slow path is the watch-polling loop, NOT the indexing or submit. - -## 4. Saturation knee identified - -Two distinct signals - both new vs. the earlier 50×50 run: - -### 4.1 Engine dispatch outlier: 101 s on a single block - -In the prior runs (130 / 280 / 300 watches), the dispatch_block max was bounded between 50 ms and 88 ms steady-state (plus a ~500 ms cold-start outlier on the first watch-heavy block). This run, with 381 active watches and a 0.5 s block time, hit **a 101-second dispatch on at least one block**. That is 200× the prior worst case. - -The likely chain: a 0.5 s block cadence + 381 watches × per-watch `eth_call` against the TWAP handler + 10 parallel WS connections producing log events concurrently → either Anvil's serialised JSON-RPC handling backs up (most likely), the engine's redb writes block, or the per-module dispatch hits a worst-case queue contention. - -Distinguishing among these is the natural follow-up. For the saturation sign-off the headline matters: **the engine has a saturation knee**, it reaches it at ~380 active watches + 10 parallel submitters + 0.5 s block-time on a M-class laptop, and even at that knee it sustains 343 EthFlow round-trips end-to-end + 31 097 `eth_call` polls without producing a single `shepherd_module_errors_total`, `trap`, or `poison`. - -### 4.2 Event-delivery loss: 38-43% of load-gen events never reached the engine - -- 895 TWAP txs → 381 `ConditionalOrderCreated` events delivered. -- 895 EthFlow txs → 343 `OrderPlacement` events delivered. - -That is **a 57-62% drop rate** between the load-gen's `eth_sendTransaction` ack and shepherd's WS subscription. Three plausible causes: - -1. **Anvil's WS subscription buffer overflows** under 10 concurrent connections × 0.5 s block × 10+ log events per block. Anvil is not built for this kind of subscriber load. -2. **Alloy's pubsub client drops events** when its internal channel fills (we DID see "Pubsub service request channel closed" lines in the load-gen output - some workers' WS connections dropped before the 2-min deadline). -3. **Anvil includes only a subset of mempool txs in each block** when the mempool grows faster than the miner can drain (gas-limit-bound or mempool-eviction). - -The block-event drop signal (engine saw 110 of an expected ~240 blocks) is consistent with #1 + #2. - -### 4.3 Engine health under saturation - -Despite the 101 s dispatch outlier and the event-drop ratio: - -- ✓ Zero `shepherd_module_errors_total`. -- ✓ Zero traps. Zero poisoned modules. -- ✓ Every event the engine **did** see was dispatched and submitted: 343 EthFlow → 343 mock orderbook hits, 1:1. -- ✓ One log-side ERROR line, which is the post-teardown WS reset (same as every prior run). - -Shepherd's failure mode under saturation is **graceful degradation, not breakage**. It processes events more slowly when the surrounding system (Anvil + WS transport) cannot keep up; it does not corrupt state, drop events on its own, or kill modules. - -## 5. Comparison across the four saturation runs - -| Scenario | Workers | Block-time | Watches | TWAP block p99 | Engine errors | -|---|---|---|---|---|---| -| baseline 5×5 | 1 | 1 s | 130 | 49 ms | 0 | -| medium 20×20 | 1 | 1 s | 280 | 67 ms | 0 | -| saturation 50×50 | 1 | 1 s | 300 | 78 ms | 0 | -| **saturation-parallel** | **10** | **0.5 s** | **381** | **74 ms (log) / 101 s (max)** | **0** | - -The watch-count grew only modestly (300 → 381), but the surrounding stress (10 connections, 2× block rate) is where the new pressure came from. **The engine itself still scales sub-linearly with watch count - the 101 s outlier is correlated with Anvil + WS, not with watch count.** - -## 6. Bottleneck identified - -In order of severity: - -1. **Anvil + alloy WS subscription** chokes under 10 concurrent subscribers × 0.5 s block cadence. Event-drop ratio 57-62%. -2. **Engine dispatch** has rare worst-case 100-second outliers when polling 380+ watches against a stressed JSON-RPC backend. The dispatch itself is fine; it is waiting on synchronous `eth_call` responses that Anvil cannot serve fast enough. -3. **load-gen** is no longer the bottleneck (was in the prior run). 10 workers in parallel sustain 895 + 895 acks per 2 min. - -For the 7-day soak: this matters because Sepolia's public RPC is closer in shape to Anvil-under-pressure than to a dedicated archive node. The soak should use Alchemy/drpc/QuickNode paid endpoints, not publicnode, OR accept that some event drops will happen and rely on the `eth_getLogs` re-indexing on reconnect. - -## 7. Acceptance - -The saturation scenario's acceptance bar is "identify the bottleneck". Identified: - -1. Engine survives 380+ concurrent watches with zero errors. -2. The dispatch p99 outlier (101 s) at peak load is a **surrounding-system** symptom (Anvil + WS), not an engine bug. -3. 57-62% of upstream events are dropped before they reach the engine under this configuration - **operator must use a faster RPC than publicnode for the 7-day soak**. - -**Saturation-parallel: PASS with caveats** - engine acceptance criteria met; the test surfaces the surrounding infrastructure as the next limiting factor. - -## 8. Followups - -1. **Re-run with a paid Sepolia archive endpoint** (Alchemy / drpc / QuickNode) and confirm the event-drop ratio falls below 5%. This is mostly a one-liner in `scripts/.env`. -2. **Re-run with `anvil --no-mining` + explicit `evm_mine` calls** to remove the timing race entirely. Each block can be packed with N+M txs deterministically. -3. **redb pre-seed** (option 3 from the load-test follow-up list) - bypass `create()` entirely, write 3 000+ watch entries directly to the local-store before engine boot. Isolates "watch-count → dispatch cost" scaling perfectly. Not blocking for this acceptance sign-off. - -## 9. Attachments - -- Metrics start: `/tmp/shepherd-load/metrics-start-20260619T170540Z.txt` -- Metrics end: `/tmp/shepherd-load/metrics-end-20260619T170540Z.txt` -- Engine + load-gen logs under `/tmp/shepherd-load/`. - -## 10. Sign-off - -**Bruno (operator) - PASS, saturation-parallel.** Engine survives the heaviest load we could synthesise without breaking. The saturation knee is real (101 s dispatch outlier, 38-43% event delivery) but the symptoms point at Anvil + WS, not at shepherd. Engine continues to scale sub-linearly with watch count and never produces a `module_errors_total`, trap, or panic. diff --git a/docs/operations/load-reports/load-5x5-2026-06-19.md b/docs/operations/load-reports/load-5x5-2026-06-19.md deleted file mode 100644 index 5e4f9095..00000000 --- a/docs/operations/load-reports/load-5x5-2026-06-19.md +++ /dev/null @@ -1,161 +0,0 @@ -# Load test report — baseline 5×5 - -> Second baseline run on 2026-06-19 after the load-gen calibration landed. -> Supersedes the conditional-pass first run recorded earlier today. - -## 1. Run metadata - -| Field | Value | -|---|---| -| Stamp (UTC) | 2026-06-19T14:48:46Z | -| Wall clock | 60 s | -| Engine commit | `feat/load-gen-calibration` head | -| Engine config | `engine.load.toml` (state_dir=./data/load wiped per run) | -| Anvil command | `anvil --fork-url $RPC_URL_SEPOLIA_HTTP --port 8545 --block-time 1` | -| Sepolia archive | `https://ethereum-sepolia-rpc.publicnode.com` | -| Mock orderbook | `tools/orderbook-mock --port 9999` (no latency, no errors) | -| Modules | `twap-monitor`, `ethflow-watcher` | -| Scenario | baseline (5 TWAP + 5 EthFlow per block, 1 min) | - -## 2. Load generator output - -``` -load-gen finished blocks_seen=26 - twap_attempted=130 twap_ok=130 - ethflow_attempted=130 ethflow_ok=130 -``` - -130 `ComposableCoW.create(...)` + 130 `CoWSwapEthFlow.createOrder(...)` -delivered across 26 Anvil blocks. Counters now reflect *delivered -events* because the load-gen calibration removed the nonce-race + -EthFlow OrderUid dedup that suppressed the first run. - -## 3. Engine throughput (the answer to "how does shepherd do under 5+5/block?") - -### Counts (Prometheus delta) - -| Metric | Delta | Notes | -|---|---|---| -| `shepherd_event_latency_seconds_count{module="twap-monitor",event_kind="block"}` | **64** | One per Anvil block (60 s / ~1 s block; some pre-load-gen + post-load-gen blocks captured). | -| `shepherd_event_latency_seconds_count{module="twap-monitor",event_kind="log"}` | **130** | `ConditionalOrderCreated` indexings; matches load-gen 1:1. | -| `shepherd_event_latency_seconds_count{module="ethflow-watcher",event_kind="log"}` | **130** | `OrderPlacement` dispatches; matches load-gen 1:1. | -| `shepherd_cow_api_submit_total{outcome="ok"}` | **130** | Every EthFlow strategy submit reached the mock orderbook successfully. | -| `shepherd_chain_request_total{method="eth_call",outcome="err"}` | **4 157** | `getTradeableOrderWithSignature` reverts (no settle-time allowance); strategy correctly classifies as `TryNextBlock`. 130 watches × ~32 blocks ≈ 4 160 - matches. | -| `shepherd_module_errors_total` | **0** | Zero traps, zero panics, zero poisoned modules. | - -### Latency - -**twap-monitor block (poll loop over all 130 watches):** - -| Quantile | Value | -|---|---| -| p50 | 34 ms | -| p95 | 45 ms | -| p99 | 49 ms | -| max | 50 ms | - -**twap-monitor log (`ConditionalOrderCreated` decode + persist):** - -| Quantile | Value | -|---|---| -| p50 | 4 ms | -| p95 | 5 ms | -| p99 | 6 ms | -| max | 11 ms | - -**ethflow-watcher log (decode → resolve_app_data → build OrderCreation → mock submit → marker):** - -| Quantile | Value | -|---|---| -| p50 | 8 ms | -| p95 | 10 ms | -| p99 | 11 ms | -| max | 11 ms | - -Engine-log-derived dispatch_block max: 474 ms (one cold-start outlier -on the first block where all 130 watches were freshly indexed and -the eth_call cache was cold). Subsequent blocks 34-50 ms steady. - -## 4. Mock orderbook - -``` -submits_ok = 130 -submits_err = 0 -app_data_lookups = 0 -``` - -The empty appData hash matches the `EMPTY_APP_DATA_HASH` short-circuit -in `shepherd_sdk::cow::resolve_app_data`, so the mock's app-data -endpoint sees zero traffic in this scenario - that's expected -behaviour, not a load-gen miss. - -## 5. Acceptance vs. baseline bar - -| Criterion | Observed | Pass? | -|---|---|---| -| 5 TWAP + 5 EthFlow events delivered per block | 130 TWAP + 130 EthFlow in 26 blocks = exactly 5+5/block | **✓** | -| 100% terminal markers within 3 blocks of event | Every EthFlow dispatch reaches mock + writes marker in 8-11 ms (single Anvil block) | **✓** | -| p99 latency < 2 s | TWAP block p99 = 49 ms; EthFlow log p99 = 11 ms | **✓** (40× margin) | -| Zero fuel exhaust | zero | **✓** | -| Zero traps | zero | **✓** | -| `shepherd_module_errors_total = 0` | zero | **✓** | - -**Baseline: full PASS.** Engine sustains the 5+5/block scenario with -40× margin on the latency bar. - -## 6. Observed bottleneck signal - -The twap-monitor *block* dispatch grows linearly with the watch count: -each block re-polls every watch (`eth_call` of -`getTradeableOrderWithSignature`). At 130 watches the p99 is 49 ms; -extrapolating naively, ~3 000 watches would put us at ~1 s which is -still under the 2 s bar but visible. - -The saturation scenario (50 × 50 = 3 000 events in a 60 s window) is -explicitly designed to test that extrapolation. **Medium 20 × 20 and -saturation 50 × 50 are unblocked - run them in follow-up sessions.** - -## 7. Engine health summary - -- ✓ Zero `shepherd_module_errors_total`. -- ✓ Zero traps, zero `init failed`, zero poisoned modules. -- ✓ All 130 EthFlow submissions reached the mock orderbook (0 errors). -- ✓ All 130 TWAP indexings persisted to the local store. -- ✓ `ConditionalOrderCreated` and `OrderPlacement` event streams - delivered 1:1 from Anvil to the modules with no drops. -- The one `WARN reconnect failed` line late in the engine log is the - expected post-teardown WS reset when `scripts/load-run.sh`'s trap - killed Anvil. Not an anomaly. - -## 8. Followups surfaced by this run - -1. **`scripts/load-bootstrap.sh` PID-file truncation** - on a fresh - run, the bootstrap wipes `/tmp/shepherd-load.pids`, so a previous - run's leaked engine process (port 9100) is invisible to teardown. - We hit this between the calibration smoke and this run; - manual `pkill nexum-engine` was required. Fix: pid-by-port - teardown, or move PID files into per-run timestamps. Not a load - test finding; just operational hygiene. -2. **Cold-start outlier on first watch-heavy block** (474 ms vs. - 34-50 ms steady-state). Probably redb's first-write barrier plus - the cold `eth_call` provider connection. Re-confirm under medium - scenario; if the outlier scales with watch count, worth a - supervisor-side investigation. - -## 9. Attachments - -- Engine log: `/tmp/shepherd-load/engine.log` -- Load-gen log: `/tmp/shepherd-load/load-gen.log` -- Anvil log: `/tmp/shepherd-load/anvil.log` -- Mock log: `/tmp/shepherd-load/orderbook-mock.log` -- Metrics start: `/tmp/shepherd-load/metrics-start-20260619T144846Z.txt` -- Metrics end: `/tmp/shepherd-load/metrics-end-20260619T144846Z.txt` - -(Local-only; auto-archiving into `docs/operations/load-reports/` -remains a follow-up.) - -## 10. Sign-off - -**Bruno (operator) — PASS, baseline.** Engine handles 5+5/block -with 40× margin on latency, zero errors, every event delivered -end-to-end. Medium 20×20 and saturation 50×50 are unblocked. diff --git a/docs/operations/load-testnet-runbook.md b/docs/operations/load-testnet-runbook.md index d4bfb659..09d94c48 100644 --- a/docs/operations/load-testnet-runbook.md +++ b/docs/operations/load-testnet-runbook.md @@ -1,196 +1,85 @@ # Load test runbook -How to stress shepherd's `twap-monitor` + `ethflow-watcher` modules -under synthetic load using a local Anvil fork of Sepolia and a mock -orderbook. +Stresses the `twap-monitor` + `ethflow-watcher` modules under synthetic load using a local Anvil fork of Sepolia and a mock orderbook. Answers one question: how many TWAP+EthFlow events per block the engine dispatches before something breaks. -The acceptance bar is: +Acceptance bar: | Scenario | Per-block load | Expected outcome | |---|---|---| -| Baseline | 5 TWAP + 5 EthFlow | 100% terminal markers within 3 blocks; p99 latency < 2s; zero fuel exhaust; zero traps | -| Medium | 20 TWAP + 20 EthFlow | Graceful degradation - `backoff:` markers OK, `shepherd_module_errors_total` stays 0 | +| Baseline | 5 TWAP + 5 EthFlow | 100% terminal markers within 3 blocks; p99 latency < 2 s; zero fuel exhaust; zero traps | +| Medium | 20 TWAP + 20 EthFlow | Graceful degradation: `backoff:` markers OK, `shepherd_module_errors_total` stays 0 | | Saturation | 50 TWAP + 50 EthFlow | Expected to saturate; report identifies the bottleneck | -This runbook is distinct from -`docs/operations/e2e-testnet-runbook.md` (correctness on live Sepolia) -and the 7-day soak (wall-clock stability). - ---- - ## 0. Prerequisites -### Toolchain - ``` rustup target add wasm32-wasip2 -brew install foundry # for `anvil` + `cast` +brew install foundry # anvil + cast cargo --version >= 1.87 ``` -### Sepolia archive endpoint - -`anvil --fork-url` needs an HTTP archive endpoint to seed the fork. -Add to `scripts/.env`: +`anvil --fork-url` needs an HTTP archive endpoint. Add to `scripts/.env`: ``` RPC_URL_SEPOLIA_HTTP=https://eth-sepolia.g.alchemy.com/v2/ ``` -(Public nodes throttle the initial fork warmup; use Alchemy / drpc / -similar.) - ---- +(Public nodes throttle the fork warmup; use Alchemy / drpc / similar.) ## 1. Boot -The three supporting processes (Anvil, orderbook-mock, engine) live in -the background; `scripts/load-run.sh` is the single entry point. +`scripts/load-run.sh` is the single entry point: ```bash -# baseline (default knobs: 5 TWAP + 5 EthFlow per block, 1 minute) +# baseline (5 TWAP + 5 EthFlow per block, 1 minute) ./scripts/load-run.sh -# medium load -./scripts/load-run.sh --twap-per-block 20 --ethflow-per-block 20 \ - --duration-min 2 --scenario medium +# medium +./scripts/load-run.sh --twap-per-block 20 --ethflow-per-block 20 --duration-min 2 --scenario medium -# saturation probe -./scripts/load-run.sh --twap-per-block 50 --ethflow-per-block 50 \ - --duration-min 2 --scenario saturation +# saturation +./scripts/load-run.sh --twap-per-block 50 --ethflow-per-block 50 --duration-min 2 --scenario saturation ``` The script: -1. Sources `scripts/load-bootstrap.sh` -> starts Anvil (`port 8545`) - and `tools/orderbook-mock` (`port 9999`). -2. Builds `twap-monitor` + `ethflow-watcher` `.wasm`, the - `nexum` binary, and `tools/load-gen`. -3. Starts the engine pointed at `engine.load.toml`. -4. Snapshots `/metrics` from the engine. +1. Sources `scripts/load-bootstrap.sh`: starts Anvil (port 8545) and `tools/orderbook-mock` (port 9999). +2. Builds the two module `.wasm`, the `shepherd` binary, and `tools/load-gen`. +3. Starts the engine on `engine.load.toml`. +4. Snapshots `/metrics`. 5. Runs `tools/load-gen` for the requested duration. 6. Snapshots `/metrics` again. -7. Tears everything down. -8. Drops a report at `docs/operations/load-reports/load-NxM-YYYY-MM-DD.md`. - -If you Ctrl-C, the trap calls `load_teardown` and kills the children -before exit. If something escapes (bash trap missed), run -`./scripts/load-teardown.sh` explicitly. - ---- - -## 2. What each component does - -### Anvil (port 8545) - -``` -anvil --fork-url $RPC_URL_SEPOLIA_HTTP --port 8545 --block-time 1 -``` - -Forks Sepolia at the latest block. Inherits every contract the test -needs (ComposableCoW, CoWSwapEthFlow, TWAP handler, WETH9, COW token) -at their pinned Sepolia addresses, so the test EOA can call -`ComposableCoW.create(...)` and `CoWSwapEthFlow.createOrder(...)` -against real bytecode without any local deployment step. - -`--block-time 1` mines a block per second, matching Sepolia's -~12s cadence... loosely. The point of the load test is to push N+M -transactions into each block, not to mimic mainnet block times. - -### Mock orderbook (port 9999) - -`tools/orderbook-mock` serves the one endpoint shepherd's `cow-api` -host backend hits per submission: +7. Tears everything down and drops a report at `docs/operations/load-reports/load-NxM-YYYY-MM-DD.md`. -- `POST /api/v1/orders` - returns a synthetic 56-byte OrderUid. +Ctrl-C triggers `load_teardown`. If a child escapes, run `./scripts/load-teardown.sh`. -Knobs (set via env in `scripts/load-bootstrap.sh` if needed): +## 2. Components -- `--latency-ms` - inject artificial latency on every response. -- `--error-rate` - fraction of POST /orders responses that return a - recognised `ApiError` envelope. Alternates between - `InsufficientFee` (`TryNextBlock`) and `InvalidSignature` (`Drop`). - -For the saturation probe, leaving `latency_ms=0` and `error_rate=0` -isolates the engine-side bottleneck from orderbook-side variability. - -### Engine (engine.load.toml) - -- `[chains.11155111] rpc_url = "ws://localhost:8545"` -- `[extensions.cow.orderbook_urls] 11155111 = "http://localhost:9999"` -- Prometheus enabled on `127.0.0.1:9100` -- `state_dir = ./data/load` (wiped at the start of every run) -- Module list: `twap-monitor` + `ethflow-watcher` only - -### Load generator (tools/load-gen) - -Connects to the Anvil WebSocket, calls `anvil_impersonateAccount` + -`anvil_setBalance` on the pinned EOA -(`0x7bF140727D27ea64b607E042f1225680B40ECa6A`), then in a loop, every -new block, fires N `ComposableCoW.create(...)` calls plus M -`CoWSwapEthFlow.createOrder(...)` calls. Each create uses a fresh -salt (counter-derived) so the txs do not collide on the -ComposableCoW dedup check. - -`anvil_impersonateAccount` skips signing entirely - one fewer -overhead under load. - ---- +- **Anvil (port 8545)**: `anvil --fork-url $RPC_URL_SEPOLIA_HTTP --port 8545 --block-time 1`. Forks Sepolia at the latest block, inheriting ComposableCoW, CoWSwapEthFlow, the TWAP handler, WETH9, and COW at their pinned addresses, so the test EOA calls real bytecode with no local deployment. +- **Mock orderbook (port 9999)**: `tools/orderbook-mock` serves `POST /api/v1/orders`, returning a synthetic 56-byte OrderUid. Knobs (env in `scripts/load-bootstrap.sh`): `--latency-ms` injects response latency; `--error-rate` returns a fraction as an `ApiError` envelope, alternating `InsufficientFee` (`TryNextBlock`) and `InvalidSignature` (`Drop`). Leave both 0 for the saturation probe to isolate the engine-side bottleneck. +- **Engine (`engine.load.toml`)**: RPC `ws://localhost:8545`; cow orderbook URL `http://localhost:9999`; Prometheus on `127.0.0.1:9100`; `state_dir = ./data/load` (wiped each run); modules `twap-monitor` + `ethflow-watcher`. +- **Load generator (`tools/load-gen`)**: connects to the Anvil WS, calls `anvil_impersonateAccount` + `anvil_setBalance` on the pinned EOA, then each new block fires N `ComposableCoW.create(...)` + M `CoWSwapEthFlow.createOrder(...)` calls, each with a fresh counter-derived salt. ## 3. Acceptance reading -After a run, the report at -`docs/operations/load-reports/load-NxM-YYYY-MM-DD.md` carries: - -- mock-orderbook stats (success vs. error count) - matches load-gen's - reported submit-attempt count, modulo `error_rate`. -- load-gen tail - submit success/failure breakdown per block. -- engine log tail - watch for `module trap`, `poisoned`, - `init failed`, `WS reconnect`. -- metrics delta filename pair (auto-delta lands in a follow-up). +The report at `docs/operations/load-reports/load-NxM-YYYY-MM-DD.md` carries mock-orderbook stats, the load-gen submit breakdown, the engine log tail, and the metrics snapshot pair. Look at: -Look at: +- `shepherd_event_latency_seconds{module="twap-monitor"}` quantiles: p99 < 2 s for baseline. +- `shepherd_cow_api_submit_total{outcome="ok"}`: tracks the load-gen success count. +- `shepherd_module_errors_total`: must stay 0 for baseline/medium; any non-zero count on saturation is the headline. +- `shepherd_chain_request_total{method="eth_call"}`: twap-monitor polls via `eth_call`; the count shows how hard the poll races the next block. -- `shepherd_event_latency_seconds{module="twap-monitor"}` quantiles - - p99 < 2s for the baseline scenario. -- `shepherd_cow_api_submit_total{outcome="ok"}` - should track the - load-gen success count. -- `shepherd_module_errors_total` - must stay 0 for baseline/medium; - any non-zero count on saturation is the headline. -- `shepherd_chain_request_total{method="eth_call"}` - twap-monitor - polls via `eth_call`; the count tells you how aggressively the - poll is racing the next block. - ---- - -## 4. What this does NOT prove - -- WS reconnect resilience (7-day soak). -- Diverse appData / order-shape correctness (the watcher replays real - placements from contract genesis at startup). -- Multi-day memory drift (7-day soak). -- Real-orderbook 4xx variety (only a live-orderbook run exercises this). -- Provider rate-limit handling on the live network. - -This test answers exactly one question: "How many TWAP+EthFlow events -per block can shepherd dispatch before something breaks?" Use it -alongside the soak, not instead of it. - ---- - -## 5. Troubleshooting +## 4. Troubleshooting | Symptom | Cause | Fix | |---|---|---| -| Anvil exits within 5s | Forking endpoint rejected | Check `RPC_URL_SEPOLIA_HTTP` is an archive endpoint, not a pruned node. Alchemy free tier works. | -| `cargo build --target wasm32-wasip2` fails on `wit-bindgen` | Toolchain stale | `rustup target add wasm32-wasip2` (re-run; may have rolled). | -| Engine never reaches `supervisor ready` | wasm artefacts not built | The script builds them, but a stale `target/wasm32-wasip2/release/*` from another branch can collide. `rm -rf target/wasm32-wasip2` and rerun. | -| `/metrics` never comes up | Port 9100 in use | Edit `engine.load.toml` `bind_addr` (and the curl URL in `scripts/load-run.sh`). | -| `load-gen` errors with "EOA not impersonated" | Anvil restarted mid-run | `scripts/load-teardown.sh && scripts/load-run.sh` from scratch. | - ---- +| Anvil exits within 5 s | Forking endpoint rejected | Ensure `RPC_URL_SEPOLIA_HTTP` is an archive endpoint | +| `wasm32-wasip2` build fails on `wit-bindgen` | Toolchain stale | `rustup target add wasm32-wasip2` | +| Engine never reaches `supervisor ready` | Stale wasm artefacts | `rm -rf target/wasm32-wasip2` and rerun | +| `/metrics` never comes up | Port 9100 in use | Edit `engine.load.toml` `bind_addr` and the curl URL in `scripts/load-run.sh` | +| `load-gen` errors with "EOA not impersonated" | Anvil restarted mid-run | `scripts/load-teardown.sh && scripts/load-run.sh` | -## 6. References +## 5. References - Sister doc (live Sepolia E2E): `docs/operations/e2e-testnet-runbook.md` - Engine config: `engine.load.toml` diff --git a/docs/operations/m2-testnet-runbook.md b/docs/operations/m2-testnet-runbook.md index a8cc9782..8891d17d 100644 --- a/docs/operations/m2-testnet-runbook.md +++ b/docs/operations/m2-testnet-runbook.md @@ -1,37 +1,18 @@ # M2 testnet runbook (Sepolia) -How to actually run the M2 modules - twap-monitor and ethflow-watcher - -on Sepolia and exercise the full path the unit tests cannot: real -`eth_subscribe` streams, real `eth_call` reverts, real orderbook -submissions. +Runs twap-monitor and ethflow-watcher on Sepolia against real `eth_subscribe` streams, `eth_call` reverts, and orderbook submissions. -Two flavours: +Two flavours share the same boot: -1. **Smoke run**: boot the engine, watch the supervisor pick up every - `ConditionalOrderCreated` / `OrderPlacement` log that lands on - Sepolia. Passive; you do not produce traffic. 15-30 min wall clock. -2. **Round-trip run**: smoke run plus you author a TWAP order via a - Sepolia Safe and an EthFlow swap via the public CoW Swap UI. The - engine indexes / decodes / submits. 1-2 h. - -Both share the same boot. The round-trip is the smoke run with a hand -on the wheel. - ---- +1. Smoke run: boot the engine and watch it pick up every `ConditionalOrderCreated` / `OrderPlacement` log that lands on Sepolia. Passive, 15-30 min. +2. Round-trip run: the smoke run plus you author a TWAP order via a Sepolia Safe and an EthFlow swap via the CoW Swap UI, 1-2 h. ## 0. Prerequisites -- Rust toolchain matching `rust-toolchain.toml` (nightly with - `wasm32-wasip2` target). `rustup target add wasm32-wasip2` once. -- `just` (`cargo install just` or `brew install just`). -- Sepolia RPC. Public endpoint in `engine.m2.toml` works for short - runs; switch to Alchemy/Infura with a key for anything past ~20 min. -- For the round-trip: - - A Sepolia EOA with some test ETH ([Alchemy faucet](https://sepoliafaucet.com)). - - A [Sepolia Safe](https://app.safe.global/?chain=sep) (only for the - TWAP half). - ---- +- Rust toolchain matching `rust-toolchain.toml` (nightly with `wasm32-wasip2`). `rustup target add wasm32-wasip2` once. +- `just`. +- Sepolia RPC. The public endpoint in `engine.m2.toml` works for short runs; switch to Alchemy/Infura for anything past ~20 min. +- Round-trip only: a Sepolia EOA with test ETH, and a Sepolia Safe for the TWAP half. ## 1. Smoke run @@ -39,28 +20,20 @@ on the wheel. just run-m2 ``` -Equivalent long form: +Long form: ```bash cargo build -p twap-monitor --target wasm32-wasip2 --release cargo build -p ethflow-watcher --target wasm32-wasip2 --release -cargo run -p nexum-cli -- --engine-config engine.m2.toml +cargo run -p shepherd -- --engine-config engine.m2.toml --pretty-logs ``` -### What you should see in the first ~5 seconds (observed) +Expected boot (~5 s): ``` INFO nexum_runtime nexum starting INFO nexum_runtime::host::provider_pool opening chain RPC provider chain_id=11155111 url="wss://..." -INFO nexum_runtime::supervisor loading module manifest manifest=modules/twap-monitor/module.toml -[manifest] required capabilities: logging, local-store, chain, cow-api -INFO nexum_runtime::supervisor compiling component component=target/wasm32-wasip2/release/twap_monitor.wasm -INFO nexum_runtime::host::impls::logging twap-monitor init module="twap-monitor" INFO nexum_runtime::supervisor init succeeded module=twap-monitor -INFO nexum_runtime::supervisor loading module manifest manifest=modules/ethflow-watcher/module.toml -[manifest] required capabilities: logging, local-store, chain, cow-api -INFO nexum_runtime::supervisor compiling component component=target/wasm32-wasip2/release/ethflow_watcher.wasm -INFO nexum_runtime::host::impls::logging ethflow-watcher init module="ethflow-watcher" INFO nexum_runtime::supervisor init succeeded module=ethflow-watcher INFO nexum_runtime::supervisor supervisor up count=2 INFO nexum_runtime supervisor ready modules=2 chains=1 @@ -69,172 +42,72 @@ INFO nexum_runtime::runtime::event_loop log subscription open module=twap-monit INFO nexum_runtime::runtime::event_loop log subscription open module=ethflow-watcher chain_id=11155111 ``` -Then every ~12s (Sepolia block time): - -``` -INFO nexum_runtime::runtime::event_loop dispatch block chain_id=11155111 number=N -``` +Then a `dispatch block` line every ~12 s (Sepolia block time). -### What to verify +Verify: | Check | How | |---|---| -| Both modules booted | `module_count: 2` + 2 `loaded module` lines | +| Both modules booted | `count=2` + 2 `init succeeded` lines | | Subscriptions wired | 2 log subs + 1 block sub | -| No traps in the first 10 blocks | `alive: 2` stays at 2; no `module ... trapped` lines | +| No traps in the first 10 blocks | no `module ... trapped` lines | | State persistence works | `ls data/m2/` shows `ls.redb` growing | -### Stopping cleanly +Ctrl-C to stop. Remove `./data/m2/` between runs for a fresh slate. -Ctrl-C. Tear down `./data/m2/` between runs if you want a fresh slate. - -### Common surprises +## 2. Round-trip run -- **Public RPC throttles after a few minutes.** Symptom: `eth_subscribe` - reconnects in a loop. Fix: switch to Alchemy/Infura. Edit the - `[chains.11155111]` block in `engine.m2.toml` (env-substitution is - not wired yet). -- **You see `eth_call failed (...); defaulting to TryNextBlock`.** This - is twap-monitor polling watches that are still empty (no - `ConditionalOrderCreated` indexed yet). Expected on a fresh `./data/m2`. -- **You see NO log dispatches for hours.** Sepolia has low ComposableCoW - / EthFlow traffic. The smoke run is mostly a "stay alive" test until - you produce events yourself (see round-trip below). +Same boot; you produce the events. ---- +### 2a. TWAP half (Safe + Compose) -## 2. Round-trip run +ComposableCoW expects the conditional-order owner to be an EIP-1271 verifier, so the TWAP flow runs behind a Safe, not an EOA. -Same boot as #1; you produce the events. - -### 2a. TWAP half (via Safe + Compose) - -The TWAP flow lives behind a Safe, not an EOA, because ComposableCoW -expects the conditional-order owner to be an EIP-1271 verifier. - -1. **Create a Sepolia Safe** at . - Single signer with your EOA is fine. Fund it with ~0.05 Sepolia - ETH (gas) and ~10 of a Sepolia ERC-20 you want to sell. -2. **Install the Compose app** in the Safe. CoW Protocol publishes the - ComposableCoW Watch Tower as a Safe app on Sepolia. - - In Safe -> Apps -> Add custom app: use the URL from - README ("Add to - Safe"). -3. **Author a TWAP order**. Compose UI -> "TWAP". Recommended for the - first run: - - Sell: 1 of your test ERC-20. - - Buy: any Sepolia stable. - - Split into 2 parts, 5-minute interval, validity 30 min. - - Confirm + sign the Safe tx. -4. **Watch the engine logs.** Within ~12s of the Safe tx confirming, - you should see: +1. Create a Sepolia Safe at (single signer with your EOA). Fund it with ~0.05 Sepolia ETH and ~10 of a Sepolia ERC-20 to sell. +2. Add the ComposableCoW Compose app (Safe -> Apps -> Add custom app, URL from the composable-cow README). +3. Author a TWAP order in the Compose UI: sell 1 test ERC-20, buy any Sepolia stable, 2 parts, 5-minute interval, 30-minute validity. Sign the Safe tx. +4. Within ~12 s of the tx confirming: ``` INFO twap-monitor indexed watch:0x:0x - ``` - Then on the next blocks where the tranche is ready: - ``` INFO twap-monitor poll watch:... -> Ready INFO twap-monitor submitted submitted:0x ``` - Sometimes you see `TryAtEpoch(t)` instead of `Ready` - that means - the tranche is gated until time `t`. Wait the configured interval. -5. **Confirm on the orderbook.** Get the UID from the log, then: +`TryAtEpoch(t)` instead of `Ready` means the tranche is gated until time `t`; wait the configured interval. +5. Confirm on the orderbook (settlement on Sepolia is spotty; reaching the orderbook is the goal): ```bash curl https://api.cow.fi/sepolia/api/v1/orders/0x ``` - You should see the order JSON back. Trade settlement on Sepolia is - spotty (solvers do not always pick up); the goal of this test is - that the order reached the orderbook, not that it filled. -### 2b. EthFlow half (via swap.cow.fi) +### 2b. EthFlow half (swap.cow.fi) -EthFlow does not need a Safe - any EOA works. +Any EOA works. -1. Go to (Sepolia native - ETH selector). -2. Connect your EOA, select a small swap (e.g. 0.001 SETH -> any - token), confirm. -3. The CoWSwapEthFlow contract on Sepolia - (`0xbA3cB4...EadeC`) emits `OrderPlacement`. -4. **Watch the engine logs:** +1. Go to . +2. Connect the EOA, select a small swap (e.g. 0.001 SETH -> any token), confirm. +3. CoWSwapEthFlow (`0xbA3cB4...EadeC`) emits `OrderPlacement`. Expected log: ``` INFO ethflow-watcher ethflow submitted 0x ``` - If you see `ethflow backoff 0x ...` instead: orderbook - classified the submit as retriable. Wait one block, the watcher - does not retry on its own today (planned for M4 supervisor - restart wiring). - - If you see `ethflow dropped 0x ...`: orderbook rejected - permanently (most likely `DuplicateOrder` - CoW Swap submits the - order itself first, ethflow-watcher races and loses). Expected; the - `dropped:{uid}` row is the regression guard, not the - failure signal here. +`ethflow backoff 0x` means the orderbook classified the submit as retriable; wait one block. `ethflow dropped 0x` means a permanent rejection (commonly `DuplicateOrder`, since CoW Swap submits the order first and the watcher races it); the `dropped:{uid}` row is the expected marker. -### What "passing M2 round-trip" looks like - -- At least one `submitted:{uid}` row in `data/m2/ls.redb` written by - each module. -- Both modules still alive (`alive: 2`) at the end of the run. -- Zero `module ... trapped` lines in the engine log. -- `curl api.cow.fi/sepolia/api/v1/orders/` returns the order JSON - for at least one submitted UID (`null` means the orderbook never - accepted; non-null means we round-tripped). - ---- +Passing round-trip: at least one `submitted:{uid}` row per module in `data/m2/ls.redb`, both modules alive at the end, zero `trapped` lines, and `curl api.cow.fi/sepolia/api/v1/orders/` returns the order JSON for at least one UID. ## 3. Inspecting state after a run -The local-store is a redb file. Quick inspection without writing a -tool: - -```bash -# Build the example mini-CLI the engine ships -cargo run -p nexum-cli --bin ls-dump -- data/m2/ls.redb 2>/dev/null \ - || echo "no ls-dump bin in 0.2 - read via the engine on next boot" -``` - -Today the canonical way to read the store is to boot the engine again -on the same `state_dir`: the supervisor logs every `watch:` / -`submitted:` / `dropped:` row it loads. A proper inspector is -production-hardening scope (M4). +The local store is a redb file with no standalone dump tool. Reboot the engine on the same `state_dir`: the supervisor logs every `watch:` / `submitted:` / `dropped:` row it loads. ---- - -## 4. What this run does NOT prove - -- **Throughput / soak stability**. That is the 7-day soak. -- **Cross-module isolation under load**. That is the 4-6h - multi-module E2E run. The local-store namespace test guarantees the - invariant in unit; the runbook above is a single-Safe / single-EOA - setup. -- **Resource-limit enforcement under adversarial guests**. Fuel + memory tests (M4 territory). -- **Security review**. Tracked separately (M4 territory). - -The M2 runbook covers: "does the engine actually boot the two M2 -modules end-to-end against Sepolia, route real subscription events -through the wit-bindgen + WitBindgenHost path, and round-trip orders -to the CoW orderbook". That is the deliverable M2 is responsible for. - ---- - -## 5. Troubleshooting +## 4. Troubleshooting | Symptom | Likely cause | Fix | |---|---|---| -| `connection refused` / WS retries | Public node throttled | Switch RPC to Alchemy / Infura | -| `module twap-monitor trapped: OutOfFuel` | Dispatch path exceeded fuel budget | Almost certainly an upstream issue, file as a separate issue; raise `[engine.limits]` fuel temporarily | +| `connection refused` / WS retries | Public node throttled | Switch RPC to Alchemy / Infura in `engine.m2.toml` | +| `module twap-monitor trapped: OutOfFuel` | Dispatch exceeded fuel budget | File an issue; raise `[engine.limits]` fuel temporarily | | `eth_call failed (rate limited)` repeatedly | Public node | Same as above | -| `ParseManifestError: missing capability cow-api` | Engine version mismatch with module.toml | `cargo build -p nexum-cli --release` and use the fresh binary | -| `data/m2/ls.redb` not created | `state_dir` not writable | Check permissions, or change `state_dir` in `engine.m2.toml` | - ---- +| `ParseManifestError: missing capability cow-api` | Engine/module.toml version mismatch | `cargo build -p shepherd --release` and use the fresh binary | +| `data/m2/ls.redb` not created | `state_dir` not writable | Check permissions or change `state_dir` in `engine.m2.toml` | -## 6. References +## 5. References - Engine config schema: `crates/nexum-runtime/src/engine_config.rs` - M2 modules: `modules/twap-monitor/`, `modules/ethflow-watcher/` -- ADR-0005 (cow-api routing): `docs/adr/0005-cow-api-via-cached-orderbookapi.md` -- ADR-0006 (twap + ethflow helpers): `docs/adr/0006-cow-twap-ethflow-host-helpers.md` -- ADR-0009 (host trait surface): `docs/adr/0009-host-trait-surface.md` -- M2 PRs in `bleu/nullis-shepherd`: #2-#11 +- ADR-0005 (cow-api routing), ADR-0006 (twap + ethflow helpers), ADR-0009 (host trait surface) diff --git a/docs/operations/m3-edge-case-validation.md b/docs/operations/m3-edge-case-validation.md deleted file mode 100644 index 62b681ab..00000000 --- a/docs/operations/m3-edge-case-validation.md +++ /dev/null @@ -1,230 +0,0 @@ -# M3 testnet edge-case validation (2026-06-18) - -Five edge cases run against the live `engine.m3.toml` boot on Sepolia. -Each takes ~10-15 s of wall clock; together they exercise the error -paths the runbook section 1 cannot cover passively. **All five -passed with one minor observation** (init-failed module stays -`alive=true`; safe in practice, worth a follow-up issue). - -Run on commit `feat/m3-edge-case-validation` tip; engine debug log -level. - ---- - -## 1.1 Bad RPC URL -> structured connect error, clean exit - -**Mutation**: `engine.m3.toml` `rpc_url = "wss://nonexistent.example.com"`. - -**Observed**: - -``` -INFO nexum starting -INFO opening chain RPC provider chain_id=11155111 url="wss://nonexistent.example.com" -Error: connect chain 11155111: IO error: failed to lookup address information: - nodename nor servname provided, or not known -``` - -**Verdict**: ✅ engine exits with structured `connect chain N: ...` -error chain. No panic, no retry loop, no silent hang. Operator -gets a clear cue to fix the URL. - -**Implication**: an operator misconfiguring an RPC URL fails fast and -loud. Combined with the supervisor restart loop (planned for M4), -this gives "kill engine, fix config, restart, no orphaned state". - ---- - -## 1.2 Bad oracle address -> module Warn + stays alive - -**Mutation**: `modules/examples/price-alert/module.toml::[config]` -`oracle_address = "0x0000000000000000000000000000000000000001"` (an -EOA with no code; `eth_call` returns empty bytes). - -**Observed**: boot clean; on the first block: - -``` -WARN price-alert: latestRoundData decode failed: - ABI decoding failed: buffer overrun while deserializing -``` - -Engine stays at `supervisor up count=3`; balance-tracker and -stop-loss continue to operate normally. - -**Verdict**: ✅ module gracefully handles upstream giving the wrong -shape. The decode error names the failing call (`latestRoundData`) -and the failure mode (buffer overrun), so an operator can correlate -to a misconfigured `oracle_address` without reading source. - -**Implication**: validates the SDK error model end-to-end: -`chain::request` returns Ok with empty bytes, `parse_eth_call_result` -returns `Some(vec![])`, `latestRoundDataCall::abi_decode_returns` -fails with `alloy_sol_types::Error::Buffer overrun`, the strategy's -`map_err` surfaces it as a `Warn` log via `LoggingHost::log`. All -four host traits + the `cow` helper path exercised. - ---- - -## 1.3 Capability mismatch -> boot rejects module - -**Mutation**: `modules/examples/stop-loss/module.toml::[capabilities]` -`required = ["logging"]` (dropped `chain`, `local-store`, `cow-api`). - -**Observed**: - -``` -INFO loading module manifest manifest=modules/examples/stop-loss/module.toml -[manifest] required capabilities: logging -INFO compiling component component=...stop_loss.wasm -Error: load module target/wasm32-wasip2/release/stop_loss.wasm - -Caused by: - 0: capability violation in target/wasm32-wasip2/release/stop_loss.wasm - 1: component imports `cow-api` (shepherd:cow/cow-api@0.1.0) but it - is not listed in [capabilities].required or [capabilities].optional -``` - -Engine exits with non-zero. The whole boot fails because the -supervisor cannot honour the (intentionally under-declared) manifest. - -**Verdict**: ✅ the capability security boundary is enforced at module -load, not deferred to first host call. Error chain identifies the -specific `cow-api` import that the manifest does not authorise. This -is the `enforce capability declarations at module -instantiation` invariant working in production. - -**Implication**: a malicious or buggy module cannot import a host -capability without explicitly declaring it. This is the M3 SDK -contract's core security guarantee. - ---- - -## 1.4 Malformed `[config]` -> init returns typed `InvalidInput` - -**Mutation**: `modules/examples/price-alert/module.toml::[config]` -`threshold = "not-a-number"`. - -**Observed**: - -``` -INFO loading module manifest manifest=modules/examples/price-alert/module.toml -WARN init failed - module=price-alert - kind=invalid_input - "invalid [config]: threshold: non-digit character in - \"not-a-number\"" -INFO balance-tracker init: 2 addresses, ... -INFO init succeeded module=balance-tracker -INFO stop-loss init: owner=..., trigger=..., ... -INFO init succeeded module=stop-loss -INFO supervisor up count=3 -``` - -**Verdict**: ✅ init failure isolated to the offending module. -Balance-tracker and stop-loss boot normally. The export returns -`fault.invalid-input`; the supervisor supplies the module name and -derives the log `kind` from the fault label, with a clear message -identifying the field + the invalid character. - -**Update (landed in this PR series)**: the supervisor now -flips `alive = false` when `init` returns `Err`, and the boot log -shows `supervisor up loaded=3 alive=2` so the discrepancy is -visible. Re-running scenario 1.4 against live Sepolia after the fix: - -``` -WARN init failed - module loaded but marked dead; dispatcher will skip it - module=price-alert kind=invalid_input - "invalid [config]: threshold: non-digit character in 'not-a-number'" -INFO supervisor up loaded=3 alive=2 -``` - -Subsequent block dispatches reach only the 2 alive modules; the -init-failed price-alert is now skipped by the dispatch fast-path -without surfacing the no-op fuel cost. Regression test: -`supervisor::tests::init_failure_marks_module_dead_and_excludes_from_dispatch`. - ---- - -## 1.5 Persistence cross-restart -> redb file preserved - -**Mutation**: boot 1 with `rm -rf data/m3` (fresh state), then boot 2 -without rm. - -**Observed**: - -``` -=== Boot 1 (fresh) === -INFO balance-tracker init: 2 addresses, ... -INFO init succeeded module=balance-tracker -(stopped after 14s) - -=== State after boot 1 === -total 7200 --rw-r--r-- brunotavaresdosanjos 3686400 data/m3/local-store.redb - -=== Boot 2 (state preserved) === -INFO balance-tracker init: 2 addresses, ... -INFO init succeeded module=balance-tracker -(stopped after 14s) -``` - -Both boots clean; `local-store.redb` file size stable (3.6 MB - redb -pre-allocates pages; actual key/value content is bytes, not MB). - -**Verdict**: ✅ the redb file survives `kill -TERM` cleanly, can be -re-opened on the next boot, and the supervisor reads from it -without corruption. This validates the 32-byte hash prefix -namespace in production: modules wrote -keys, the engine shut down, modules re-attached on restart, no -panic. - -**Implication**: the local-store invariant -(`namespaces_isolate_modules` unit test + cross-restart durability) -is now confirmed against a real Sepolia run. Combined with the -supervisor integration tests, this is sufficient -evidence that local-store persistence works at the production -boundary, not only in mocks. - -**Caveat**: there is no built-in CLI to dump the redb contents, so -visual confirmation of specific keys (`last:0x...`, etc.) requires -either re-booting the engine on the same state_dir or writing an -ad-hoc inspector. Filed as a future M4-territory nice-to-have. - ---- - -## Summary - -| # | Scenario | Verdict | New issue? | -|---|---|---|---| -| 1.1 | Bad RPC URL | ✅ structured error + clean exit | no | -| 1.2 | Bad oracle address | ✅ Warn + module alive + clear decode error | no | -| 1.3 | Capability mismatch | ✅ boot rejects with structured error chain | no | -| 1.4 | Malformed `[config]` | ✅ typed `InvalidInput`; init-failed module marked dead + excluded from dispatch | resolved in this PR series | -| 1.5 | Cross-restart persistence | ✅ redb file preserved + re-attaches cleanly | no (a state-dump CLI would help; M4 nice-to-have) | - -**One follow-up issue**: in `Supervisor::load`, when `init` returns -`Err(fault)`, set `alive=false` (or skip pushing the module into -`self.modules`). Subsequent dispatch wastes fuel on a no-op -short-circuit otherwise. Safe today; cleanup before M4. - -**Not in scope here** (M4 territory, already filed): -- Fuel exhaustion (M4 territory) -- Memory exhaustion (M4 territory) -- Module trap during `on_event` + restart with backoff (M4 territory) -- WS reconnect logic instead of bail → not filed (current behaviour - is documented in `runtime/event_loop.rs` as "0.3 fix") - ---- - -## How to reproduce - -Each scenario is a one-line config mutation + `just run-m3` (or the -equivalent `cargo run`). Mutations are listed inline above. Restore -config between runs: - -```bash -git checkout modules/examples/price-alert/module.toml \ - modules/examples/stop-loss/module.toml \ - engine.m3.toml -``` - -Tested on commit `` at 2026-06-18, Sepolia public WS. diff --git a/docs/operations/m3-testnet-runbook.md b/docs/operations/m3-testnet-runbook.md index 8f2f9335..6e0b4a17 100644 --- a/docs/operations/m3-testnet-runbook.md +++ b/docs/operations/m3-testnet-runbook.md @@ -1,208 +1,99 @@ # M3 testnet runbook (Sepolia) -How to exercise the M3 example modules - price-alert, balance-tracker, -stop-loss - on Sepolia. Same shape as the M2 runbook but the modules -are different: - -- **price-alert** validates SDK `chain` helpers + Chainlink ABI decode. - Read-only; no on-chain or orderbook action. -- **balance-tracker** validates SDK `chain::request` (raw RPC) + - `local-store` per-key diff persistence. Read-only. -- **stop-loss** validates the full M3 surface: `chain::request` + - `local-store` dedup + `cow-api::submit-order` with - `Signature::PreSign`. Will attempt to submit a real CoW order to the - Sepolia orderbook when the oracle price crosses the trigger. - -In other words: M3 exercises the *strategy*-side SDK surface that M2 -modules eventually consume. The runbook below validates everything in -~8 seconds of wall clock against the real Sepolia ETH/USD Chainlink -feed. - ---- +Exercises the example modules price-alert, balance-tracker, and stop-loss on Sepolia: -## 0. Prerequisites +- price-alert: SDK `chain` helpers + Chainlink ABI decode. Read-only. +- balance-tracker: SDK `chain::request` (raw RPC) + `local-store` per-key diff persistence. Read-only. +- stop-loss: `chain::request` + `local-store` dedup + `cow-api::submit-order` with `Signature::PreSign`. Submits a real CoW order to the Sepolia orderbook when the oracle price crosses the trigger. -- Same as the M2 runbook (Rust nightly + `wasm32-wasip2`, `just` - optional, Sepolia RPC). -- For stop-loss to actually settle an order (not just submit and get - rejected) you also need: - - An EOA matching `[config] owner = ...` in - `modules/examples/stop-loss/module.toml` that has called - `setPreSignature(orderUid, true)` on the GPv2Settlement Sepolia - contract for the computed UID. - - That EOA holds + has approved enough of `sell_token` to settle. +All three subscribe to blocks only and start working immediately; a single Sepolia block (~12 s) drives each through its full strategy. - Without those, stop-loss will hit `TransferSimulationFailed` (or - `InvalidSignature` / `InsufficientAllowance`) and log it as a - retriable error or drop. **That outcome alone validates the - orderbook round-trip** - same shape as the M2 EthFlow validation. +## 0. Prerequisites ---- +- Same as the M2 runbook (Rust nightly + `wasm32-wasip2`, `just`, Sepolia RPC). +- For stop-loss to settle (not just submit and get rejected): + - An EOA matching `[config] owner` in `modules/examples/stop-loss/module.toml` that has called `setPreSignature(orderUid, true)` on the GPv2Settlement Sepolia contract for the computed UID. + - That EOA holds and has approved enough `sell_token` to settle. -## 1. Smoke + active run +Without those, stop-loss hits `TransferSimulationFailed` (or `InvalidSignature` / `InsufficientAllowance`) and logs it as a retriable error or drop. That outcome still validates the orderbook round-trip. -The M3 modules all subscribe to blocks only and start working -immediately - there is no `[[subscription]] kind = "chain-log"` to wait for. -A single Sepolia block (~12 s) drives all three through their full -strategy. +## 1. Smoke + active run ```bash just run-m3 ``` -Equivalent long form: +Long form: ```bash cargo build -p price-alert --target wasm32-wasip2 --release cargo build -p balance-tracker --target wasm32-wasip2 --release cargo build -p stop-loss --target wasm32-wasip2 --release -cargo run -p nexum-cli -- --engine-config engine.m3.toml +cargo run -p shepherd -- --engine-config engine.m3.toml --pretty-logs ``` -### What you should see in the first ~10 seconds (observed) +Expected boot (~10 s): ``` INFO nexum starting -INFO opening chain RPC provider chain_id=11155111 url="wss://..." -INFO loading module manifest manifest=modules/examples/price-alert/module.toml -[manifest] required capabilities: logging, chain -INFO compiling component component=...price_alert.wasm -INFO price-alert init: oracle=0x694aa1769357215de4fac081bf1f309adc325306 - threshold=250000000000 direction=Below every_n_blocks=1 INFO init succeeded module=price-alert -INFO loading module manifest manifest=modules/examples/balance-tracker/module.toml -[manifest] required capabilities: logging, chain, local-store -INFO compiling component component=...balance_tracker.wasm -INFO balance-tracker init: 2 addresses, threshold=100000000000000000 wei INFO init succeeded module=balance-tracker -INFO loading module manifest manifest=modules/examples/stop-loss/module.toml -[manifest] required capabilities: logging, chain, local-store, cow-api -INFO compiling component component=...stop_loss.wasm -INFO stop-loss init: owner=0x70997970c51812dc3a010c7d01b50e0d17dc79c8 - trigger=250000000000 sell=0x6810e776880c02933d47db1b9fc05908e5386b96 - buy=0xfff9976782d46cc05630d1f6ebab18b2324d6b14 INFO init succeeded module=stop-loss -INFO supervisor up count=3 INFO supervisor ready modules=3 chains=1 INFO block subscription open chain_id=11155111 ``` -Then on the FIRST Sepolia block dispatch (~5-15s after boot): +On the first Sepolia block dispatch (~5-15 s after boot): ``` -DEBUG chain::request chain_id=11155111 method=eth_call # price-alert reads oracle +DEBUG chain::request method=eth_call # price-alert reads oracle WARN price-alert: TRIGGERED answer=174553978080 threshold=250000000000 (Below) -DEBUG chain::request chain_id=11155111 method=eth_getBalance # balance-tracker addr 1 -DEBUG chain::request chain_id=11155111 method=eth_getBalance # balance-tracker addr 2 -DEBUG chain::request chain_id=11155111 method=eth_call # stop-loss reads oracle -DEBUG cow-api::submit-order chain_id=11155111 bytes=561 -WARN stop-loss retry on next block (0): orderbook error (TransferSimulationFailed): - sell token cannot be transferred +DEBUG chain::request method=eth_getBalance # balance-tracker addr 1 +DEBUG chain::request method=eth_getBalance # balance-tracker addr 2 +DEBUG chain::request method=eth_call # stop-loss reads oracle +DEBUG cow-api::submit-order bytes=561 +WARN stop-loss retry on next block (0): orderbook error (TransferSimulationFailed): sell token cannot be transferred ``` -That single block proves the entire M3 strategy surface end-to-end: -oracle read + ABI decode + multi-key local-store + cow-api submit + -typed retry classification, all routed through real wit-bindgen + -WitBindgenHost + supervisor dispatch on a live testnet. - -### Why TRIGGERED fires immediately +That block proves the M3 strategy surface end-to-end: oracle read + ABI decode + multi-key local-store + cow-api submit + typed retry classification. -The default `threshold = "2500.00"` in `module.toml::[config]` is -above the Sepolia Chainlink ETH/USD feed (which tracks a stale or -mocked value, often around $1745). Direction is `below`, so the very -first poll trips the alert. Tune `threshold` if you want to test the -"silent" path. +Why TRIGGERED fires immediately: the default `trigger_price` in `module.toml` is above the Sepolia Chainlink ETH/USD feed (a stale or mocked value), and `direction = below`, so the first poll trips. Raise `trigger_price` to test the silent path. -### Why stop-loss logs TransferSimulationFailed - -The default `owner = 0x70997970...` in stop-loss's config is the -canonical hardhat test EOA (`anvil` account index 1). It does not own -or approve the `sell_token` on Sepolia, so the orderbook simulates -the would-be settle and rejects with -`TransferSimulationFailed`. **This is the orderbook returning a typed -error - the full submit path worked.** The module's -`classify_api_error` SDK helper correctly tagged it as retriable -(`TryNextBlock`), so the watch is left in place for the next block. - -For the silent ("idle until trigger") run path, set `owner` to a real -EOA with the right allowances + pre-signature - see section 2 below. - ---- +Why stop-loss logs TransferSimulationFailed: the default `owner` does not own or approve `sell_token` on Sepolia, so the orderbook simulates the settle and rejects with a typed error. The `classify_api_error` SDK helper tags it retriable (`TryNextBlock`) and leaves the watch for the next block. ## 2. Active validation (optional) -To see stop-loss actually submit + persist `submitted:{uid}` you need -to set up a real signed order: - -1. Pick a Sepolia EOA you control. -2. In `modules/examples/stop-loss/module.toml`, set `owner = "0x..."` - to that EOA. -3. Choose a `sell_token` / `buy_token` pair the EOA holds. -4. Compute the OrderUid the module will submit (the `build_creation` - helper in `strategy.rs` shows the construction; you can also boot - the engine once with a high trigger so it stays idle, then - simulate-decode the would-be submit by reading the supervisor's - debug log). -5. Call `GPv2Settlement.setPreSignature(uid, true)` from that EOA on - Sepolia. -6. Approve `sell_token` to the GPv2VaultRelayer for the sell amount. -7. Lower the `trigger_price` in `module.toml` so the next poll fires. +To see stop-loss submit and persist `submitted:{uid}`: + +1. Set `owner` in `modules/examples/stop-loss/module.toml` to a Sepolia EOA you control. +2. Choose a `sell_token` / `buy_token` pair the EOA holds. +3. Compute the OrderUid (see `build_creation` in `strategy.rs`). +4. Call `GPv2Settlement.setPreSignature(uid, true)` from that EOA. +5. Approve `sell_token` to the GPv2VaultRelayer for the sell amount. +6. Lower `trigger_price` so the next poll fires. On the next block: ``` INFO stop-loss TRIGGERED price=... trigger=... -DEBUG cow-api::submit-order ... INFO stop-loss submitted submitted:0x ``` -This is the M3 equivalent of the M2 EthFlow validation: same -end-to-end surface, different module. - ---- - ## 3. State inspection -`./data/m3/ls.redb` accumulates the `last:{addr}` keys -(balance-tracker), `submitted:{uid}` / `dropped:{uid}` (stop-loss). -Same caveat as M2 - no `ls-dump` CLI today; reboot the engine on the -same `state_dir` and the supervisor logs every key it loads. - -`rm -rf ./data/m3` between runs for a fresh slate. - ---- - -## 4. What this does NOT prove - -Same boundary as M2's section 4: +`./data/m3/ls.redb` accumulates `last:{addr}` (balance-tracker) and `submitted:{uid}` / `dropped:{uid}` (stop-loss). No standalone dump tool: reboot the engine on the same `state_dir` and the supervisor logs every key it loads. `rm -rf ./data/m3` for a fresh slate. -- Throughput / 7-day soak. -- Cross-module isolation under load (the 4-6 h E2E run). -- Adversarial resource exhaustion (M4 territory). -- Security review (M4 territory). -- `app_data` resolution for stop-loss orders with non-empty metadata - -> M5 (typed `Cow` client with `raw_request`). - ---- - -## 5. Troubleshooting - -Most of the M2 runbook's section 5 applies verbatim. M3-specific: +## 4. Troubleshooting | Symptom | Likely cause | Fix | |---|---|---| -| `module stop-loss trapped: TransferSimulationFailed` | Trap vs warn confusion | The "sell token cannot be transferred" line is a Warn, not a trap. Module stays alive. Read again carefully. | -| Engine bails immediately with `log stream ended (WebSocket dropped?)` | Pre-fix M1 bug | Should not happen on this commit. The fix lands in `runtime/event_loop.rs`: `select_all` over empty `Vec` is replaced with `stream::pending()`. Regression test at `supervisor::tests::run_does_not_bail_when_both_stream_kinds_are_empty`. | -| `price-alert: TRIGGERED` does not fire | Oracle returned shape we cannot decode, or Sepolia public node throttled the `eth_call` | Check for `eth_call failed` warnings; switch to Alchemy. | -| `balance-tracker` only logs 1 of 2 addresses | RPC dropped a request mid-block | Same RPC throttle path; switch RPC. | - ---- +| `module stop-loss trapped: TransferSimulationFailed` | Trap vs warn confusion | `sell token cannot be transferred` is a Warn, not a trap; the module stays alive. | +| `price-alert: TRIGGERED` does not fire | Undecodable oracle shape, or throttled `eth_call` | Check for `eth_call failed`; switch to Alchemy. | +| `balance-tracker` logs only 1 of 2 addresses | RPC dropped a request mid-block | Switch RPC. | -## 6. References +## 5. References - M3 modules: `modules/examples/{price-alert,balance-tracker,stop-loss}/` -- SDK helpers exercised: `crates/shepherd-sdk/src/{chain,cow}/` -- ADR-0009 (host trait surface): `docs/adr/0009-host-trait-surface.md` -- M3 PRs in `bleu/nullis-shepherd`: #12-#26 (SDK + examples + tutorial + QA cleanup) -- M3 fix tail PRs: #27-#31 (CI matrix, rustdoc gate, doctests, supervisor integration, M2 runbook) -- M2 runbook (sister doc, same shape): `docs/operations/m2-testnet-runbook.md` +- SDK chain helpers: `crates/nexum-sdk/src/chain/` +- ADR-0009 (host trait surface) +- M2 runbook (sister doc): `docs/operations/m2-testnet-runbook.md` diff --git a/docs/operations/soak-runbook.md b/docs/operations/soak-runbook.md index 76e3feea..90acb573 100644 --- a/docs/operations/soak-runbook.md +++ b/docs/operations/soak-runbook.md @@ -1,146 +1,66 @@ -# 7-Day Soak Runbook +# 7-day soak runbook -How to run the **7-day unattended stability soak** — all 5 modules on -Sepolia, continuously, with hourly metrics snapshots. - -## Purpose - -Grant evidence. The milestones require: - -- **M1 (24h):** snapshots from hour 0-24 proving sustained operation. -- **M2 (48h):** snapshots from hour 0-48 proving 48-hour stability. -- **M4 (7-day):** full run artifact set (logs + snapshots + start/end metrics). - -The soak validates *stability* — it is the step after the E2E run -(`docs/operations/e2e-testnet-runbook.md`) which validates correctness. -Do not start the soak unless the E2E run has passed its acceptance bar. - ---- +Runs all 5 modules on Sepolia continuously and unattended, with hourly metrics snapshots, to validate stability. The soak follows the E2E run (`docs/operations/e2e-testnet-runbook.md`), which validates correctness; do not start it until the E2E run has passed its acceptance bar. ## How it works -Two Docker containers managed by `docker-compose.soak.yml`: - -- **engine** — the nexum binary with `restart: unless-stopped`. Docker - handles log rotation (json-file driver, 500 MB × 14 files ≈ 7 GB cap) - and automatic crash recovery. -- **snapshotter** — an Alpine container that runs `scripts/soak-snapshot.sh`: - captures a baseline on start, then scrapes `/metrics` every hour and - writes `metrics-snap-.txt` to `docs/operations/soak-reports/` - (bind-mounted from the host so files are immediately accessible). +Two containers managed by `docker-compose.soak.yml`: ---- +- **engine**: the shepherd binary with `restart: unless-stopped`. Docker handles log rotation (json-file driver, 500 MB x 14 files) and crash recovery. +- **snapshotter**: an Alpine container running `scripts/soak-snapshot.sh`: captures a baseline on start, then scrapes `/metrics` every hour to `metrics-snap-.txt` under `docs/operations/soak-reports/` (bind-mounted from the host). ## Pre-flight checklist - [ ] Docker installed and running. -- [ ] Paid RPC endpoint with WebSocket support. Public nodes (e.g. - `wss://ethereum-sepolia-rpc.publicnode.com`) throttle `eth_subscribe` - under sustained load. Alchemy or Infura growth tier recommended. +- [ ] Paid RPC endpoint with WebSocket support (public nodes throttle `eth_subscribe` under sustained load). - [ ] `SEPOLIA_RPC_URL` set in the repo-root `.env`: ```bash echo "SEPOLIA_RPC_URL=wss://eth-sepolia.g.alchemy.com/v2/YOUR_KEY" >> .env ``` - Docker Compose reads `.env` automatically and forwards the variable - into the engine container. -- [ ] ≥ 20 GB free disk (engine logs capped at ~7 GB by Docker log - rotation; snapshots are negligible). -- [ ] Machine will not sleep: - - **macOS:** System Settings → Battery → Prevent automatic sleep when power - adapter is connected. - - **Linux:** `systemd-inhibit --what=sleep` or configure Docker to start - on boot (`sudo systemctl enable docker`). -- [ ] E2E run has passed its acceptance bar (see `e2e-testnet-runbook.md`). -- [ ] **Set `SHEPHERD_IMAGE` to an image that exists.** The compose - default (`ghcr.io/nullislabs/shepherd:latest`) is only published on - pushes to `main`; until a release lands there, `pull` fails with - `denied`. Two working options: - - **(a) Build locally from the soak commit** (recommended — the image - digest in the evidence then provably matches the reviewed tree): +- [ ] >= 20 GB free disk (engine logs cap at ~7 GB via rotation). +- [ ] Machine will not sleep (macOS: prevent sleep on power adapter; Linux: `systemd-inhibit --what=sleep` or `sudo systemctl enable docker`). +- [ ] E2E run has passed its acceptance bar. +- [ ] `SHEPHERD_IMAGE` set to an image that exists. The compose default (`ghcr.io/nullislabs/shepherd:latest`) is only published on pushes to `main`; until then, `pull` fails with `denied`. Either build locally from the soak commit: ```bash docker build -t shepherd-soak:$(git rev-parse --short HEAD) . echo "SHEPHERD_IMAGE=shepherd-soak:$(git rev-parse --short HEAD)" >> .env ``` - - **(b) Publish via CI:** trigger the `docker` workflow with - `gh workflow run docker.yml --ref develop`, wait for it to push, then - pin the tag it printed: - ```bash - echo "SHEPHERD_IMAGE=ghcr.io/nullislabs/shepherd:sha-abc1234" >> .env - docker compose -f docker-compose.soak.yml pull - ``` - - Record the resolved image in the evidence set either way: +or publish via CI (`gh workflow run docker.yml --ref develop`), then pin the printed tag and `docker compose -f docker-compose.soak.yml pull`. Record the resolved image: ```bash docker inspect --format '{{.Config.Image}} {{.Image}}' soak-engine \ > docs/operations/soak-reports/image-pin.txt # after `up` ``` ---- - ## Starting the soak ```bash docker compose -f docker-compose.soak.yml up -d -``` - -Docker starts the engine, waits for it to become healthy (up to 90 s), then -starts the snapshotter which immediately captures `metrics-start-.txt`. - -Check it started cleanly: - -```bash docker compose -f docker-compose.soak.yml ps -# Both services should show "Up" and engine shows "(healthy)" +# Both services show "Up"; engine shows "(healthy)". ``` ---- - ## Monitoring -**Follow engine logs live:** +Follow engine logs: ```bash docker compose -f docker-compose.soak.yml logs -f engine ``` -**Filter per-module activity markers:** +Filter per-module markers (`docker logs`, not `docker compose logs`, so jq is not fed the service-name prefix; the JSON formatter flattens fields, message at `.message`): ```bash -# `docker logs` (not `docker compose logs`): compose prefixes every -# line with the service name, which breaks jq. The engine's JSON -# formatter flattens event fields to the top level, so the message -# lives at `.message`. docker logs -f soak-engine \ | jq -r 'select(.message // "" | test("watch:|submitted:|dropped:|backoff:|TRIGGERED")) | .message' ``` -**Check snapshot count:** +Snapshot count (expect one file per completed hour; >= 23 at 24 h): ```bash ls docs/operations/soak-reports/metrics-snap-*.txt | wc -l ``` -Expect one file per completed hour. At 24 h you should see ≥ 23 files. - -**Scrape live metrics:** - -```bash -curl http://127.0.0.1:9100/metrics -``` - -**Check both containers are alive:** - -```bash -docker compose -f docker-compose.soak.yml ps -``` - -**Memory evidence (start once, right after `up`):** - -The Prometheus exporter has no process collector, so `/metrics` carries -no RSS — without this loop there is no data to prove memory stayed flat -across 7 days. Run it on the host: +Memory evidence (start once, right after `up`). The exporter has no process collector, so `/metrics` carries no RSS; this host-side loop is the only memory record: ```bash nohup sh -c 'while true; do @@ -151,18 +71,14 @@ done' >/dev/null 2>&1 & echo $! > docs/operations/soak-reports/.memory-loop.pid ``` ---- - ## Stopping cleanly ```bash -# Capture final metrics before stopping. +# Final metrics before stopping. curl -sf http://127.0.0.1:9100/metrics \ > "docs/operations/soak-reports/metrics-end-$(date -u +%Y%m%dT%H%M%SZ).txt" -# Capture the restart record: RestartCount > 0 means Docker recovered a -# crash during the run — that must be visible in the evidence, not -# discovered by a reviewer noticing counters reset between snapshots. +# Restart record: RestartCount > 0 means Docker recovered a crash; that must be in the evidence. docker inspect soak-engine --format \ 'restarts={{.RestartCount}} started={{.State.StartedAt}} oom={{.State.OOMKilled}}' \ > "docs/operations/soak-reports/engine-state-$(date -u +%Y%m%dT%H%M%SZ).txt" @@ -170,97 +86,54 @@ docker inspect soak-engine --format \ # Stop the host-side memory loop. kill "$(cat docs/operations/soak-reports/.memory-loop.pid)" 2>/dev/null || true -# Bring everything down (SIGINT → graceful shutdown → SIGKILL after 30 s). +# Bring everything down (SIGINT -> graceful shutdown -> SIGKILL after 30 s). docker compose -f docker-compose.soak.yml down ``` ---- +## Evidence artefacts -## Evidence artifacts +All files in `docs/operations/soak-reports/`: -All files in `docs/operations/soak-reports/` constitute the grant evidence: - -| Artifact | Pattern | Purpose | +| Artefact | Pattern | Purpose | |---|---|---| -| Engine log | `engine.log.gz` (see below) | Full operation history | +| Engine log | `engine.log.gz` | Full operation history | | Image pin | `image-pin.txt` | Image + digest the run executed | | Baseline metrics | `metrics-start-.txt` | Counter values at t=0 | -| Hourly snapshots | `metrics-snap-.txt` × N | Hourly Prometheus scrapes | +| Hourly snapshots | `metrics-snap-.txt` x N | Hourly Prometheus scrapes | | Final metrics | `metrics-end-.txt` | Counter values at shutdown | -| Memory samples | `memory.log` | Hourly RSS/CPU — proves no leak | +| Memory samples | `memory.log` | Hourly RSS/CPU | | Restart record | `engine-state-.txt` | RestartCount / OOMKilled | +Save the engine log compressed (uncompressed it approaches the ~7 GB cap; attach the `.gz`, do not commit it): + ```bash -# Save the engine log compressed. Uncompressed it can approach the -# ~7 GB rotation cap — do NOT commit it to the repo; attach the .gz to -# the evidence bundle. Note rotation drops anything past the cap, so -# archive before the run ends if log volume runs high. -docker logs soak-engine 2>&1 | gzip \ - > docs/operations/soak-reports/engine.log.gz +docker logs soak-engine 2>&1 | gzip > docs/operations/soak-reports/engine.log.gz ``` -These satisfy: - -- **M1 (24h):** `metrics-snap-*.txt` files timestamped within 0-24 h of `metrics-start-*.txt`. -- **M2 (48h):** `metrics-snap-*.txt` files timestamped within 0-48 h. -- **M4 (7-day):** The full artifact set above covering ≥ 7 days of uptime. - -To extract the shutdown summary from the log: +Extract the shutdown summary: ```bash -docker compose -f docker-compose.soak.yml logs engine \ - | grep "graceful shutdown complete" | tail -1 +docker compose -f docker-compose.soak.yml logs engine | grep "graceful shutdown complete" | tail -1 ``` ---- - ## Troubleshooting -**Engine container exited early:** +Engine container exited early: ```bash docker compose -f docker-compose.soak.yml logs --tail=100 engine ``` -Common causes: +- OOM kill: `docker inspect soak-engine | jq '.[0].State'`, look for `OOMKilled: true`. Raise the `memory` limit in `docker-compose.soak.yml` or reduce module count. +- RPC errors: look for `connection refused` / `rate limit`. Switch to a paid endpoint. +- WASM trap: look for `module trapped` / `module poisoned`. File a bug. -- OOM kill: `docker inspect soak-engine | jq '.[0].State'` — look for - `OOMKilled: true`. Increase the `memory` limit in `docker-compose.soak.yml` - or reduce module count. -- RPC errors: look for `connection refused` or `rate limit` in the log. - Switch to a paid endpoint with higher rate limits. -- WASM trap: look for `module trapped` or `module poisoned`. File a bug. - -**Snapshotter not producing files:** - -Verify the engine is healthy and the metrics port is up: - -```bash -curl -v http://127.0.0.1:9100/metrics -docker compose -f docker-compose.soak.yml ps -``` - -**Restarting an interrupted run:** - -Both services carry `restart: unless-stopped`, so an engine crash is -recovered automatically — Docker restarts it, the snapshotter keeps -looping through failed scrapes until the engine is healthy again, and -no operator action is needed. Check whether this happened with: - -```bash -docker inspect soak-engine --format 'restarts={{.RestartCount}} oom={{.State.OOMKilled}}' -``` +Snapshotter not producing files: verify the engine is healthy and the metrics port is up (`curl -v http://127.0.0.1:9100/metrics`). -Manual intervention is only needed when the operator stopped the stack -(`down`) or the host rebooted without Docker auto-start: +Restarting an interrupted run: both services carry `restart: unless-stopped`, so an engine crash recovers automatically. Manual intervention is only needed after an operator `down` or a host reboot without Docker auto-start: ```bash docker compose -f docker-compose.soak.yml up -d ``` -The engine's local store is in the `soak-state` Docker volume and -persists across restarts — no state is lost. A new -`metrics-start-.txt` appears only when the snapshotter *container* -restarts (crash or manual `up`), not on engine-only recoveries. -Preserve all files from every run — reviewers can see the combined -coverage, and the `engine-state` snapshot explains any counter resets. +The local store lives in the `soak-state` Docker volume and persists across restarts. A new `metrics-start-.txt` appears only when the snapshotter container restarts, not on engine-only recoveries. Preserve files from every run; the `engine-state` snapshot explains any counter resets.