Skip to content

launch-kit evals: frozen scenarios with a deterministic count, S6 and S7, the 2026-09-02 run - #16

Merged
hblee12294 merged 1 commit into
mainfrom
evals-s6-s7
Sep 2, 2026
Merged

launch-kit evals: frozen scenarios with a deterministic count, S6 and S7, the 2026-09-02 run#16
hblee12294 merged 1 commit into
mainfrom
evals-s6-s7

Conversation

@hblee12294

Copy link
Copy Markdown
Member

The scenarios are declared frozen so runs compare, and every run now records one deterministic number beside the judged pass/fail: the problems vos validate <kit>/kit.json reports across the run's kits (plugin ≥0.20.0 reads each asset's bytes against the channel specs). Two scenarios join from two real runs (Harper v2.9, Karakeep v0.33): S6, store screenshots are the real page, full bleed, no zoom, no chrome; S7, the poster's shot is a named zoom apex, never the cold open, and every card is a real PNG. The 2026-09-02 results file scores both runs and records the two findings that became plugin fixes.

…re screenshots, S7 the poster's shot; the 2026-09-02 run
@hblee12294
hblee12294 marked this pull request as ready for review September 2, 2026 15:22
@hblee12294
hblee12294 merged commit e62aa35 into main Sep 2, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant