Skip to content

[skills-eval] dotnet-experimental: 3 skills, 47% pass — 1 to strengthen, 2 keep #894

Description

@AbhitejJohn

Context — cross-family skill evaluation: dotnet-experimental

This issue is self-contained: it captures everything a skill author needs to act on the dotnet-experimental plugin without opening the full report.

What this measures. Every runnable skill in dotnet/skills was run through Vally (0.7) on a cross-family matrix: 5 executor model families — opus-4.8, gpt-5.5, sonnet-4.6, haiku-4.5, mai-flash — each judged by a different family (judge ≠ executor; default judge = latest Opus, or GPT when Opus is the executor). For every skill, a skilled run is compared against a baseline (no-skill) run and scored per executor. This removes single-model and self-judging bias, so a skill that only helps one model family — or only its own family's judge — is visible.

  • Data source: cross-family CI grid (run 29228914412 + backfills) — 5 executors × 85 runnable skills, 419 scored cells over 84 skills.
  • Row grain: one row per skill, aggregated across its (up to 5) executor cells. avgN is the mean trial count behind the cells (trials 1–17; higher = more statistically trustworthy). thin-N flags directional-only rows.

dotnet-experimental at a glance (portfolio scorecard)

Plugin Skills Cells Pass Impact Tie-trials Err avg ΔTok avg ΔTurns avg ΔTools Headline
dotnet-experimental 3 15 47% 0.474 7 0 +17,562 +0.56 +0.53 Healthy; keep-polish

Insights. 3 skill(s); mean impact 0.47; 3/3 help ≥1 frontier model. Exemplars (win incl. frontier): exp-test-maintainability.

Address first: nothing critical — polish the STRENGTHEN rows below.

How to read the table

Each skill is scored skilled vs. baseline on these axes:

Signal Column What it means Good
Breadth Families ✓ (n/5) How many of the 5 model families the skill helps (per-family pass count, max 5/5) 4–5 / 5
Where Passed on Which families passed. Frontier = latest Opus + latest GPT are bold; Sonnet 4.6 / Haiku 4.5 / MAI Flash are mid/low-weight frontier ✓
Magnitude Impact (−1…+1) How strongly the judge prefers skilled over baseline ≥ 0.4
Decisiveness Ties▵ Trials where the judge saw no difference → skill is inert low
Safety Loss▵ Trials where skilled was WORSE than baseline → skill misfires ~0
Reliability Err Trials that errored/crashed in setup or judging 0
Efficiency ΔTok / ΔTurns / ΔTools Extra tokens / agent turns / tool calls vs baseline ≤ 0
Confidence avgN Mean trials behind the verdict; low N = directional only ≥ 3
Invocation Call% Share of skilled trials where the model actually invoked the skill ~100%

Families ✓ is cell-level (max 5/5); Ties▵/Loss▵ are trial-level tallies summed across all families (including the ones where the skill failed). A high Families ✓ next to non-zero Loss▵ is not a contradiction — see Passed on and the Action text for where losses landed.

Action buckets (each skill has one primary action; [flags] note secondary concerns):

Bucket Priority Meaning
FIX-RELIABILITY 🔴 P0 Errored trials / no verdict — stabilize the harness before trusting the score
FIX-DISCOVERY 🔴 P0 Model doesn't invoke it (Call% < 50%) — a triggering/description problem
FIX-REGRESSION 🔴 P0 Skilled is worse than baseline on many trials — the skill misfires
ADD-DECISIVENESS 🟠 P1 Called ~100% but ties dominate, ~0 impact — inert; needs sharper behavioral steps
TRIM-COST 🟠 P1 Passes but with heavy token/turn overhead — trim verbosity
EXEMPLAR 🟢 keep Broad, strong, reliable win — use as a template
EFFICIENT-WIN 🟢 protect Wins and cuts turns/tools — the ideal shape
KEEP-POLISH 🟢 Solid majority win; minor polish + more trials
STRENGTHEN 🟡 P2 Marginal/mixed lift — sharpen triggers & success criteria

Per-skill actions

Skill Families ✓ Passed on (frontier bold) Impact Ties▵ Loss▵ Err avgN Call% ΔTok ΔTurns ΔTools Action
exp-simd-vectorization 3/5 GPT, Haiku, MAI 0.57 1 2 0 3.8 100% +36,894 +1.27 +1.33 KEEP-POLISH · 🟢 Solid majority win on GPT, Haiku, MAI. Losses/ties are concentrated in the non-passing cells (Opus, Sonnet), not the 3 passing ones — the passes are clean. To firm up: (1) raise trials on Opus, Sonnet; (2) investigate why a frontier model (Opus) didn't pass — that's the highest-value gap; (3) minor wording polish only. ([frontier-miss], [frontier-regressed])
exp-test-maintainability 2/5 Opus, Haiku 0.51 1 2 0 4 100% +7,893 +0 -0.25 EFFICIENT-WIN · ✅ Better and cheaper. Wins on Opus, Haiku while saving 0.25 tools. What's good: it adds value without bloat — the ideal shape. Protect the brevity; don't let it grow. ⚠️ Frontier miss on GPT — worth a look. ([frontier-miss], [efficient])
exp-mock-usage-analysis 2/5 GPT, Haiku 0.34 5 2 0 6 100% +7,899 +0.4 +0.5 STRENGTHEN · 🟡 Marginal/mixed (passes 2/5 — GPT, Haiku). Diagnosis: mostly ties — too generic/non-prescriptive. Try: (1) sharpen the trigger so it fires only where it wins; (2) add 1–2 opinionated, concrete steps that change behaviour; (3) add trials to separate signal from noise. ([frontier-miss])

Generated from the cross-family Call-to-Action report (CALL-TO-ACTION.md §4–§5; companion IMPACT-ANALYSIS.md). Regenerate the underlying tables with node deep-metrics.mjs "$env:TEMP\cf-ci" agg-cinode gen-cta-tables.mjs agg-ci. Numbers are directional where avgN is low; treat single-trial cells as hypotheses to confirm with more runs.

Metadata

Metadata

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions