Skip to content

Add llm-judges rule: judge-run methodology distilled from SPEC-B - #67

Open
yulonglin wants to merge 1 commit into
mainfrom
worktree-llm-judges-rule
Open

Add llm-judges rule: judge-run methodology distilled from SPEC-B#67
yulonglin wants to merge 1 commit into
mainfrom
worktree-llm-judges-rule

Conversation

@yulonglin

Copy link
Copy Markdown
Owner

Adds claude/rules/llm-judges.md — judge-run methodology distilled from the SPEC-B transcript-review build: LLM judge over regex for meaning classification, one call per sample, rationale-before-verdict, quote-grounded positives, absence-fields caveat, surface statement, blinding, sha-pinned versioned prompts, append-only resumable JSONL persistence, parse failures reported as numbers.

One sentence to review closely (my addition resolving an ambiguity Codex flagged): "Naming the construct being measured is not a blinding violation — the judge must know what to look for; what it must not see is which verdict the hypothesis favours, or which experimental condition the sample came from."

https://claude.ai/code/session_01EMvJA9K5BsQFSgbuFcqCWc

Codifies the LLM-judge conventions validated by the transcript-judge e2e run:
one call per sample, rationale-before-verdict with verbatim quotes, absence
fields cannot be quote-grounded, state the surface, blind the judge (with the
construct-naming clarification), sha-versioned prompts, append-only JSONL
persistence, parse failures reported as counts.

Claude-Session: https://claude.ai/code/session_01EMvJA9K5BsQFSgbuFcqCWc
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant