Improve .NET performance skill review workflow#912
Conversation
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: c7960b02-a0d3-4c71-8b65-3b238a9b7b42
Skill Coverage Report
Uncovered:
|
Five-family no-skill vs skill A/B evidenceRun: https://github.com/dotnet/skills/actions/runs/29625558673 This run started from current Pre-dispatch validation passed: pinned actionlint 1.7.7 across all 16 hand-authored workflows; all 4 focused Vally adapter reliability tests; static Quality resultAll five executor jobs completed successfully and produced the expected 110 raw trajectories (5 families × 11 stimuli × 2 variants). Vally produced 47 matched comparisons and 8 explicit unmatched baseline-error cases; there were 0 comparison-judge errors. Across matched comparisons the descriptive W/T/L count was 37/7/3, but I do not pool preference scores across families because the executor/judge pair differs.
The causal evidence is the within-run pairing above. It supports a clear quality gain for Haiku and strong positive but incomplete evidence for GPT and MAI. It does not establish a credible gain for Opus, and Sonnet is too incomplete to interpret beyond the six matched trials. Per-scenario evidenceEach cell is descriptive only: there is one trial per family/scenario, so scenario-level results are not statistically significant.
Key judge evidence:
Explicit missing/error evidence: GPT baseline failed to execute the truncation scenario; MAI baselines failed the Ordinal/FrozenDictionary and truncation scenarios; Sonnet baselines failed Aggregate/Replace, CurrentCulture/regex budget, Ordinal/FrozenDictionary, branched Replace, and locale hierarchy. These eight pairs are unmatched and were not silently converted into wins. Efficiency (separate from quality)These are Vally's standard treatment-minus-baseline deltas over matched trajectories.
Raw means below use every trajectory with available metrics; baseline sample counts are GPT 10, Haiku 11, MAI 9, Opus 11, Sonnet 6, while every treatment has 11. Token columns are input/output/total/cache-read/cache-write.
Aggregate tool-call breakdown
Reproducibility against run 29537550152This is secondary evidence only; the within-run A/B above is causal.
The broad pattern reproduced: strong credible positive preference for Haiku; strong positive GPT preference (current formal result limited only by one unmatched baseline); positive but incomplete MAI evidence; Opus positive point estimate with CI crossing zero; Sonnet positive point estimate but unreliable because of missing baselines. With one trial per scenario, all scenario-level findings remain descriptive rather than statistically significant. |
Corrected old-vs-new skill A/BThis supersedes my earlier no-skill-vs-skill comment. That run established that the skill is useful, but did not measure PR #912's incremental change. This run compares the existing Run: https://github.com/dotnet/skills/actions/runs/29628141153 Exact inputs
All five artifacts recorded the same hashes. All 10 run summaries confirmed one trial per stimulus and the expected old/new skill paths. Pre-dispatch validation passed: pinned actionlint 1.7.7, all four focused Vally reliability tests, static Quality resultThe run produced all expected 110 raw trajectories, 54 matched comparisons, one explicit unmatched MAI baseline executor error, and zero comparison errors. Treatment had 14 wins, 16 ties, and 24 losses descriptively across matched family/scenario pairs; scores are not pooled because judge families differ.
No family showed a credible quality improvement. Opus showed a statistically credible regression. Haiku and Sonnet also leaned toward the old skill, although their intervals cross zero. GPT leaned toward treatment but was inconclusive; MAI was exactly neutral on matched preference and had one unmatched baseline. Scenario evidenceOne trial per family/scenario makes these descriptive, not independently significant.
The systematic losses align with the workflow change: the old skill's more exhaustive scans produced broader coverage, exact counts/locations, and positive findings that treatment sometimes omitted. The clearest regression was branched
For Opus specifically, treatment had no wins: 0W/5T/6L. Its losses were generally small ( Treatment did have real strengths: it improved focused coverage for per-call Dictionary, Ordinal/FrozenDictionary, and some span scenarios. These gains did not offset the broader losses. EfficiencyThese are treatment-minus-baseline deltas over matched trajectories. Negative means PR #912 used fewer resources.
Raw means confirm the intended mechanism: for example, GPT tool calls fell from 29.45 to 13.82; Haiku from 11.73 to 6.45; Sonnet from 10.27 to 5.18. All recorded metric This establishes an efficiency/quality trade-off, not an overall improvement: PR #912 makes GPT, Haiku, Opus, and Sonnet materially less resource-intensive, but provides no credible quality gain and causes a credible Opus regression. RecommendationDo not accept PR #912 as written. The old-vs-new evidence does not support the change as an overall improvement. A narrower revision may be worthwhile:
|
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: c7960b02-a0d3-4c71-8b65-3b238a9b7b42
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: c7960b02-a0d3-4c71-8b65-3b238a9b7b42
Follow-up to #886.
This stacked PR applies only the evidence-backed skill guidance changes after the eval criteria introduced by #905:
Baseline coverage support
The automated skill coverage report is unchanged at 11/21 units (52.4%) on the frozen evaluator and this stacked follow-up. The ten uncovered units are five final-output validation checks (including the checklist requirement) and five separate pitfall guardrails; this PR therefore does not claim increased declarative coverage or broaden the evaluator. It preserves the existing baseline while changing how the agent gathers evidence during covered focused reviews.
Scenarios 8 and 9 provide the targeted behavioral probe for this workflow: they review hot-path string/collection code with outcome-focused criteria and no prescribed search vocabulary. In the independent five-family combined rerun used as supporting—not causal—evidence, both scenarios activated in all 10 skilled cells, all 10 graders passed, 53/55 rubric criteria were satisfied, average skilled tool calls fell 41.3% versus the historical run, and skilled-versus-baseline overhead fell from +5.93 to +1.36 calls (-77%). Those results support source-first review and batched confirmation while the unchanged 11/21 coverage boundary makes clear what remains untested.
#905 must merge first so the evaluation criteria are frozen before this skill-only treatment is considered. This PR intentionally does not modify eval YAML, fixtures, workflows, the vally adapter, or reliability behavior.
Validation:
skill-validator check --skills plugins/dotnet-diag/skills/analyzing-dotnet-performancemarkdownlint-cli2 plugins/dotnet-diag/skills/analyzing-dotnet-performance/SKILL.mdplugins/dotnet-diag/skills/analyzing-dotnet-performance/SKILL.md