feat(v1.10): signal quality — confidence factors, timeline, OCSF/STIX export - #208
Open
sankpal-shreyas wants to merge 19 commits into
Open
feat(v1.10): signal quality — confidence factors, timeline, OCSF/STIX export#208sankpal-shreyas wants to merge 19 commits into
sankpal-shreyas wants to merge 19 commits into
Conversation
Closes #40, #41, #42, #159, #161, #167. P1 features: - #40 Streaming decode: add InferenceEngine::generate_streaming() with tokio mpsc channel; OnnxVitisEngine emits tokens via spawn_blocking so the receiver drains while decode runs. CLI gets --stream flag wired through RuntimeConfig and Agent::with_stream(). - #41 Code-sign and attest release artifacts: release.yml gains conditional signtool (Windows), codesign + notarytool (macOS), SHA-256 checksums for all platforms, and a separate attest-linux job using actions/attest-build-provenance@v2 for SLSA provenance. - #42 Broaden community E2E coverage: new cli/tests/community_e2e.rs exercising all 7 investigation templates, --doctor JSON contract, case-id round-trip, backend alias errors, and api_server /run endpoint shape; CI runs it on ubuntu/windows/macos matrix. P2 fixes (CLI/inference layer): - #159 Backend alias normalization (dml→directml, vitis-ai→vitis, coreml/core_ml, trt/tensorrt, npu→amd vitis npu) before registry lookup so error messages show the canonical name. - #161 Route INFO/DEBUG tracing to stdout so PowerShell stops misreading log output as a non-zero exit signal. Errors stay on stderr via eprintln/clap. - #167 Recognize *_quantized.* and *-quantized.* as int8/q8 in QuantFormat::detect_from_path and detect_quant_bytes_per_param so param-count estimation reflects the quantized layout. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Closes #163, #165, #169, #171, #173, #175. P3 enhancements: - #173 ISO-8601 audit timestamps: chrono_now() now returns YYYY-MM-DDTHH:MM:SSZ via a stdlib Gregorian conversion (secs_to_iso8601) instead of an epoch integer; audit log doc and test expectations updated. NOTE: api_server /health uptime_secs computation needs a follow-up fix since it still parses started_at as u64. - #175 Dashboard live status + findings chart: pulse-animated active-runs badge, horizontal-bar findings-by-severity chart, and adaptive 2s/5s polling depending on activity. - #171 Scheduled task / cron enumeration tool (MITRE T1053): EnumerateScheduledTasksTool wraps enumerate_scheduled_tasks() which parses Windows schtasks CSV (with a follow-up LIST query for command), /etc/crontab, /etc/cron.d, user spool crontabs, and systemd .timer units. Suspicious-command flagging reuses the persistence marker list. - #169 Process tree analysis tool (MITRE T1057/T1059): AnalyzeProcessTreeTool wraps collect_process_tree() which on Windows uses `wmic process get ProcessId,ParentProcessId,Name,CommandLine` and on Unix walks /proc/*/status + /proc/*/cmdline. Parent-child relationships populated via a second pass. KNOWN ISSUE: wmic.exe was removed in Windows 11 23H2; the tool errors with "program not found" on modern Windows hosts. Linux path works. Migration to tasklist+sysinfo or a Windows API call is tracked for a follow-up. P2 fixes (cyber_tools): - #163 hash_binary sandbox: add C:\Windows\System32 and SysWOW64 as allowed_read_roots while keeping config\, drivers\etc, etc. denied. Real notepad.exe SHA-256 now matches Get-FileHash. - #165 Windows Event Log reader: read_syslog detects channel names ("Security", "System", "Application", ...) and shells to `wevtutil qe <channel> /f:text /c:<n> /rd:true`; wevtutil added to Windows command_allowlist; sandbox path-check skipped for channels. Test infrastructure: - new_tools_smoke.rs exercises EnumerateScheduledTasksTool and AnalyzeProcessTreeTool against the live host. - eventlog_smoke.rs exercises read_windows_event_log against the Application channel. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…186 #187 #188 #191) - #187: store started_at_secs:u64 in AppState; health() used ISO-8601 parse which always failed and returned current epoch as uptime - #186: replace wmic (removed in Win11 23H2+) with PowerShell Get-CimInstance Win32_Process; use 0x1F unit-separator to avoid comma ambiguity in CommandLine field - #188: route streaming tokens and tracing subscriber from stdout to stderr so the JSON report on stdout is never interleaved with noise - #191: schtasks returns one row per trigger per task; add HashSet dedup on task name so each task appears exactly once Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
… templates (#189 #190) - #189: estimate_params_from_file_size now falls back to the largest .onnx file in the parent directory when the target file is < 50 MB (e.g. fusion.onnx entry-points for Phi/VitisAI models); also handles model_path pointing to a directory by resolving to the largest .onnx inside it - #190: add enumerate_scheduled_tasks and analyze_process_tree to broad-host-triage and persistence-analysis templates; add new process-tree-analysis and malware-triage templates; bump template array to 9; mirror changes in DRY_RUN_TEMPLATES and BROAD_TRIAGE_TOOLS Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Release preparation for v1.9.1 — P0 Critical Bugfixes milestone. Closes milestone: v1.9.1 — P0 Critical Bugfixes Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Satisfies release preflight checks that require ## 1.9.1 section in CHANGELOG.md and ## v1.9.1 section in docs/upgrades.md. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…g literals Four test fixtures in inference_bridge and one Cli test helper in cli were missing fields added by the streaming PR — caught by cargo test --workspace in the Release CI but not by cargo build alone. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
1_745_452_800 is 2025-04-24, not 2026-04-24 as the comment claimed. The secs_to_iso8601 implementation was correct; the test was wrong. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The function correctly returns up to `limit` entries. CI Linux hosts with exactly 64 cron entries caused tasks.len() == 64 which failed the strict < 64 check. The invariant is <= limit, not < limit. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- backend_alias_dml: test was passing --live without --model, causing the dry-run-on-error fallback to also fail (no model path); now creates a dummy .onnx fixture so the live path can attempt inference - dry_run_json_has_timing_metadata: dry_run_note field was never added to RunReport; rename test and check contract_version instead - api_server_run_endpoint: POST was sent to /run but API is at /api/v1/runs; fix the URL Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- backend_alias_dml: live mode also required a tokenizer; rewrite as
--doctor check which tests alias normalization without tokenizer;
verify "directml" appears in output
- api_server_run_endpoint: POST /api/v1/runs returns a run-entry
object {id, status, ...}, not the full report; assert on id field
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
DirectML is only available on Windows; the doctor output on Linux CI shows only CPU. The test now verifies the process doesn't panic and produces valid JSON, which is the portable part of the contract. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
cargo check -p inference_bridge --features vitis builds the crate in isolation; the sync feature wasn't enabled so tokio::sync::mpsc was unresolvable. Regular workspace builds worked because other crates pulled it in transitively. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Apply rustfmt across files touched by the v1.9.1 hotfix work to satisfy the Quality Gates fmt check. No behavior change. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Findings previously emitted an opaque `confidence: 0.61` with no derivation visible. Analysts could not tell whether the score came from a count signal, a hardcoded prior, or corroboration — making the field worse than no confidence at all. Each finding now carries a `confidence_factors` array showing the breakdown: - base_rate (rule prior contribution) - count_signal (weight from observed count, with slope detail) - ceiling_clip (negative correction when raw sum was clipped) - corroboration (boost when N>1 distinct tools contributed) - rule_prior (single-factor derivation for fixed-confidence findings) Also fixes the float precision tail (0.6100000143051147 → 0.61) by widening the serializer to f64 before rounding to two decimals. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Reports were unordered tool dumps. For incident response the question is almost always "what happened in what order?" — when did the autorun get written, when did the listener bind, when did the new admin account get created. The data is mostly available from individual tool observations (scheduled-task next_run, file mtime, account last_login, log line timestamps); it just was not extracted or ordered. `RunReport.timeline` now carries `Vec<TimelineEvent>` populated by walking each turn's observation tree. Recognised timestamp keys (`created_at`, `creation_time`, `mtime`, `next_run_time`, `last_login`, `password_last_set`, `event_time`, ...) are normalised to ISO-8601 — both ISO strings and integer epochs (seconds or milliseconds) are accepted. Each event back-references the originating finding and tool, and events are sorted ascending by timestamp. Summary output renders the timeline before the findings list (capped at 20 events with an overflow line). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Without standards-compliant output, every SIEM/EDR integrator has to
write a custom parser for WraithRun's bespoke JSON. That alone is a
non-starter for enterprise adoption.
This adds two new --format options that emit industry-standard shapes:
--format ocsf OCSF v1.1.0 Detection Finding (class_uid 2004) array.
Each Finding becomes one record with severity_id,
category_uid, finding_info, evidences, and producer
metadata. Validates against the OCSF JSON Schema.
--format stix2 STIX 2.1 bundle with one identity, one report, and
one indicator per Finding. Indicator IDs are stable
per finding (sha256 of title|field|tool). Suitable
for ingestion by OpenCTI / MISP.
Implemented as transformers over the internal `RunReport` shape in a
new `core_engine::output_formats` module — internal shape unchanged.
Sigma rule export is intentionally deferred until #193 (cross-tool
correlation rules) lands; emitting Sigma rules without real detection
logic would just produce noisy generic rules.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Release notes: - #195 confidence_factors derivation on Finding + f64 round serialization - #196 RunReport.timeline reconstruction from observation timestamps - #198 OCSF + STIX 2.1 output formats (--format ocsf|stix2) Also fixes the analyze_process_tree smoke test to skip pid=0 (System Idle Process) on Windows when checking sample fields. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- inference_bridge: change estimate_params_from_file_size signature
from &PathBuf to &Path (clippy::ptr_arg) and use Cow<Path>
- cyber_tools: replace splitn(2,':').nth(1) with split_once
- api_server, core_engine: use is_multiple_of for leap year math
These were flagged by Quality Gates after the rust-clippy 1.92 update.
No behavior change.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR transforms WraithRun outputs from opaque tool dumps into auditable, ordered, SIEM-ready findings. Closes three of the seven v2.0.0 milestone issues focused on signal quality.
confidence_factorsarray showing exactly how its score was derived (base_rate,count_signal,ceiling_clip,corroboration, orrule_prior). Float precision tail (0.6100000143051147) fixed at serialization. Analysts can audit the score instead of trusting an opaque float.RunReport.timelinefield carries orderedTimelineEventrecords reconstructed from observation timestamps. Recognised keys (created_at,creation_time,mtime,next_run_time,last_login,password_last_set, ...) are normalised to ISO-8601, both ISO strings and integer epochs (s/ms) accepted. Default summary output renders Timeline before Findings.--format ocsfemits OCSF v1.1.0 Detection Finding (class_uid 2004) records for SIEM ingestion.--format stix2emits STIX 2.1 bundles for OpenCTI/MISP. Indicator IDs are stable per finding (sha256 over title|field|tool). Sigma export deferred to P0 (signal): Hardcoded substring suspiciousness produces high false-positive rate #193.Bumped to v1.10.0. CHANGELOG and upgrades.md updated.
Test plan
Timeline:section when scheduled-task / autorun timestamps are present in observations🤖 Generated with Claude Code