Skip to content

SLM-418: build a demo replay-preference pairs corpus + run it (DSH5-10) - #1125

Merged
Tyler-R-Kendrick merged 2 commits into
claude/great-dirac-ni43ohfrom
claude/great-dirac-ni43oh-replay-pairs-run
Jul 27, 2026
Merged

SLM-418: build a demo replay-preference pairs corpus + run it (DSH5-10)#1125
Tyler-R-Kendrick merged 2 commits into
claude/great-dirac-ni43ohfrom
claude/great-dirac-ni43oh-replay-pairs-run

Conversation

@Tyler-R-Kendrick

Copy link
Copy Markdown
Owner

Summary

Stacked on #1124 (the sixth slice's render_replay_preference_pair/render_replay_preference_pairs converter). This is the seventh slice of SLM-418's replay-grounded preference work (DSH5-10).

  • Adds scripts/build_replay_preference_pairs.py: builds one small, deterministic, honestly-labeled fixture_or_scratch conversation exercising three of the seven named patterns (edit-then-undo, undo-then-redo, checkout-another-state) with a toy zero-argument cycling operator, extracts rows via extract_replay_preference_rows, renders them via SLM-418: render replay preference rows into trainable pairs (DSH5-10) #1124's render_replay_preference_pairs, and writes a real pairs.jsonl. This is the first real, on-disk pairs corpus this feature line has ever produced — everything before this was an in-memory row/pair inside a unit test.
  • Confirmed (again, independently) while scoping this: no real captured ConversationTraceV1 corpus exists anywhere in this repo, and nothing ingests one — build_symbolic_operator_corpus synthesizes traces combinatorially from existing gold DSL records, it does not read captured usage. The demo trace here is explicitly scratch, not a stand-in for that missing corpus.
  • 2 of 3 rows render, honestly reported. undo_then_redo never renders — not a bug, a structural fact: for any deterministic, zero-argument operator, redo and "reapply the same operator at the same input state" are, by construction, identical text, so the renderer's own dedup guard (from SLM-418: render replay preference rows into trainable pairs (DSH5-10) #1124) correctly declines rather than emitting a self-contradictory pair. The script's printed report says pairs_dropped: 1 rather than hiding it.
  • First real end-to-end training run using DSH5-10 rows, one command chain: a scratch SFT checkpoint (wf_smoke_v2, seed 0, 8 steps — last_loss=32.610084533691406, identical to the 16+ prior independently-verified rows in the smoke-loop ledger, confirming this is the same deterministic artifact those rows already verified) → the new demo pairs corpus → scripts/train_preference.py train (6 steps, 2 pairs, last_loss=1.077). Both commands well under MAX_RUN_MINUTES=3.
  • Still not a training or held-out-benefit claim. n_pairs=2 on a scratch fixture with no held-out split says the pipeline runs end to end for real — nothing about whether the signal helps the model or should train the DSH3-selected policy head the issue actually asks about.
  • harness.preference.replay_pairs bumped v1 -> v2 (adds the new script + test to its watched paths).

Full disposition in docs/design/dsh5-10-replay-preference-rows.md's new "Seventh slice" section.

Test plan

  • env -u NODE_OPTIONS pytest -q tests/test_dsl/test_replay_preference.py tests/test_dsl/test_operator_merge.py tests/test_dsl/test_operator_conversation.py tests/test_harnesses/preference/test_replay_pairs.py tests/test_scripts/test_build_replay_preference_pairs.py tests/test_evals/test_advanced_operator_disposition.py tests/test_scripts/test_validate_advanced_operator_disposition.py71 passed (69 from SLM-418: render replay preference rows into trainable pairs (DSH5-10) #1124 + 2 new)
  • ruff check clean on both new files
  • python -m scripts.verify_version_stamps --check --base origin/mainok (1 component(s) touched)
  • python -m scripts.repo_policyok
  • python -m scripts.verify_decode_invariants — clean
  • Real training run executed: scripts.train_model (SFT) → scripts.build_replay_preference_pairsscripts.train_preference train, all three commands run for real this session, outputs inspected (not committed — outputs/ is gitignored)

Generated by Claude Code

Seventh slice, stacked on #1124 (the sixth slice's row->PreferencePair
renderer). Adds scripts/build_replay_preference_pairs.py: builds one small,
deterministic, honestly-labeled fixture_or_scratch conversation exercising
three of the seven named patterns (edit-then-undo, undo-then-redo,
checkout-another-state), extracts rows, renders them, and writes a real
pairs.jsonl -- the first real, on-disk pairs corpus this feature line has
ever produced.

2 of 3 rows render; undo_then_redo never does, for a structural reason, not
a bug: for any deterministic zero-argument operator, redo and reapplying
the same operator at the same input state are, by construction, identical
text, so the renderer's own dedup guard correctly declines rather than
emitting a self-contradictory pair. Confirmed (again) while scoping this
slice: no real captured ConversationTraceV1 corpus exists anywhere in this
repo, and no harness ingests one -- build_symbolic_operator_corpus
synthesizes traces combinatorially from gold DSL records, it doesn't read
captured usage.

Also ran the first real end-to-end chain using DSH5-10 rows: a scratch SFT
checkpoint (wf_smoke_v2, seed 0, 8 steps -- last_loss identical to the
16+ prior independently-verified rows in the smoke-loop ledger, confirming
this is the same deterministic artifact those rows already verified), the
new demo pairs corpus, and one bounded scripts/train_preference.py train
call against it (6 steps, 2 pairs, last_loss=1.077). Both well under
MAX_RUN_MINUTES=3. Still not a training or held-out-benefit claim: n_pairs=2
on a scratch fixture says the pipeline runs end to end for real, nothing
about whether the signal helps the model or should train the DSH3-selected
policy head the issue actually asks about.

harness.preference.replay_pairs bumped v1 -> v2 (adds the new script+test
to its watched paths).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EDctCUvBtGrdMvgNd34oHR
@coderabbitai

coderabbitai Bot commented Jul 27, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 6f2e6928-ff12-4e8f-b909-bf96331850de

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch claude/great-dirac-ni43oh-replay-pairs-run

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@vercel

vercel Bot commented Jul 27, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
slm-training Ready Ready Preview, Comment Jul 27, 2026 2:46pm

Request Review

…#1126)

Eighth slice, stacked on #1125 (the seventh slice's demo trace + first real
training run). Adds build_demo_merge_scenario to
scripts/build_replay_preference_pairs.py: a second scratch fixture (two
branches forked from a shared base, editing disjoint node refs), mirroring
the exact disjoint-target shape tests/test_dsl/test_operator_merge.py
already verifies merges cleanly, replayably, and order-invariantly. Reaches
merge_success, the one named pattern the seventh slice's single-trace
corpus structurally cannot (extract_merge_preference_row never operates on
a shared ConversationTraceV1).

main() now combines both sources into one report and one pairs.jsonl: 4
rows (3 trace-scan + 1 merge_success), 3 render -- undo_then_redo is still
the only drop, for the same structural reason the seventh slice documented.

Reran the full training chain against the richer corpus: same deterministic
SFT checkpoint (last_loss=32.610084533691406, matching every prior
wf_smoke_v2/seed-0/8-step row), preference training now over 3 pairs
instead of 2 (9 steps, last_loss=0.531). Still not a training or
held-out-benefit claim -- a larger, more structurally diverse scratch
corpus, not evidence the signal helps the model.

harness.preference.replay_pairs bumped v2 -> v3.


Claude-Session: https://claude.ai/code/session_01EDctCUvBtGrdMvgNd34oHR

Co-authored-by: Claude <noreply@anthropic.com>
@Tyler-R-Kendrick
Tyler-R-Kendrick merged commit a6b3ed4 into claude/great-dirac-ni43oh Jul 27, 2026
5 of 6 checks passed
@Tyler-R-Kendrick
Tyler-R-Kendrick deleted the claude/great-dirac-ni43oh-replay-pairs-run branch July 27, 2026 14:46
Tyler-R-Kendrick added a commit that referenced this pull request Jul 27, 2026
* SLM-418: build a demo replay-preference pairs corpus + run it (DSH5-10)

Seventh slice, stacked on #1124 (the sixth slice's row->PreferencePair
renderer). Adds scripts/build_replay_preference_pairs.py: builds one small,
deterministic, honestly-labeled fixture_or_scratch conversation exercising
three of the seven named patterns (edit-then-undo, undo-then-redo,
checkout-another-state), extracts rows, renders them, and writes a real
pairs.jsonl -- the first real, on-disk pairs corpus this feature line has
ever produced.

2 of 3 rows render; undo_then_redo never does, for a structural reason, not
a bug: for any deterministic zero-argument operator, redo and reapplying
the same operator at the same input state are, by construction, identical
text, so the renderer's own dedup guard correctly declines rather than
emitting a self-contradictory pair. Confirmed (again) while scoping this
slice: no real captured ConversationTraceV1 corpus exists anywhere in this
repo, and no harness ingests one -- build_symbolic_operator_corpus
synthesizes traces combinatorially from gold DSL records, it doesn't read
captured usage.

Also ran the first real end-to-end chain using DSH5-10 rows: a scratch SFT
checkpoint (wf_smoke_v2, seed 0, 8 steps -- last_loss identical to the
16+ prior independently-verified rows in the smoke-loop ledger, confirming
this is the same deterministic artifact those rows already verified), the
new demo pairs corpus, and one bounded scripts/train_preference.py train
call against it (6 steps, 2 pairs, last_loss=1.077). Both well under
MAX_RUN_MINUTES=3. Still not a training or held-out-benefit claim: n_pairs=2
on a scratch fixture says the pipeline runs end to end for real, nothing
about whether the signal helps the model or should train the DSH3-selected
policy head the issue actually asks about.

harness.preference.replay_pairs bumped v1 -> v2 (adds the new script+test
to its watched paths).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EDctCUvBtGrdMvgNd34oHR

* SLM-418: add merge_success to the demo pairs corpus + rerun (DSH5-10)

Eighth slice, stacked on #1125 (the seventh slice's demo trace + first real
training run). Adds build_demo_merge_scenario to
scripts/build_replay_preference_pairs.py: a second scratch fixture (two
branches forked from a shared base, editing disjoint node refs), mirroring
the exact disjoint-target shape tests/test_dsl/test_operator_merge.py
already verifies merges cleanly, replayably, and order-invariantly. Reaches
merge_success, the one named pattern the seventh slice's single-trace
corpus structurally cannot (extract_merge_preference_row never operates on
a shared ConversationTraceV1).

main() now combines both sources into one report and one pairs.jsonl: 4
rows (3 trace-scan + 1 merge_success), 3 render -- undo_then_redo is still
the only drop, for the same structural reason the seventh slice documented.

Reran the full training chain against the richer corpus: same deterministic
SFT checkpoint (last_loss=32.610084533691406, matching every prior
wf_smoke_v2/seed-0/8-step row), preference training now over 3 pairs
instead of 2 (9 steps, last_loss=0.531). Still not a training or
held-out-benefit claim -- a larger, more structurally diverse scratch
corpus, not evidence the signal helps the model.

harness.preference.replay_pairs bumped v2 -> v3.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EDctCUvBtGrdMvgNd34oHR

* SLM-418: first step onto the DSH3 typed policy target (DSH5-10)

Ninth slice, stacked on #1126. Every prior slice since the sixth has
flagged the same gap: the TwoTower PreferencePair path (#1124-#1126) exists
only for tooling compatibility -- the disposition doc's own remaining-scope
note names typed_operator_policy.py's TypedOperatorPolicyScorer, not
TwoTower, as the issue's actual training target. This slice takes the
first real step onto it, and honestly narrows what's reachable there.

Adds src/slm_training/harnesses/experiments/argument_preference.py:
build_argument_preference_example renders only pronoun_focus_followup rows
into a TypedOperatorArgumentPreferenceExampleV1 -- the one named pattern
whose chosen/rejected actions are the same operator with a differing
argument for the same slot. The other six patterns are history controls
(undo/redo/checkout) or merge, none of which has a row in
OperatorPolicyInputV1.action_rows (built only from legal_set.entries);
wiring them needs a real scope decision this slice leaves open rather than
guesses at. typed_operator_argument_preference_loss is a Bradley-Terry
pairwise margin over CandidateScoringHead logits (surrogate, not textbook
DPO, same honesty class as train_preference.py's own dpo_loss).

Real, structural finding, not a bug: training on the actual
pronoun-focus-followup fixture provably cannot move the loss. Both sibling
refs share every field ReferenceModelViewV1 exposes -- they differ only by
semantic_fingerprint, which is deliberately stripped from every model input
for anti-identity-leakage. Two feature-identical candidates get
byte-identical embeddings through the shared-weight CandidateScoringHead
regardless of any parameter update, so the loss sits at
-log_sigmoid(0)=ln(2) structurally. Proven exact (pytest.approx) and
distinct from a mechanism failure via a second test with a synthetic,
feature-distinguishable pair, where training does reduce the loss.
Consequence for future corpus-building: only pairs whose candidates differ
in an allowed feature carry learnable signal for this scorer -- worth
knowing before anyone builds a real corpus and is puzzled why it plateaus.

harness.experiments.argument_preference registered fresh (v1). No change
to typed_operator_policy.py itself (only consumed).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EDctCUvBtGrdMvgNd34oHR

* SLM-418: add partial-rollback and fork-choose demo scenarios (DSH5-10) (#1128)

Stacked on #1127. Tenth slice of SLM-418's replay-grounded preference
work: extends scripts/build_replay_preference_pairs.py with two more
standalone scratch traces (build_demo_partial_rollback_scenario,
build_demo_fork_choose_scenario) covering the two named patterns
(partial_rollback, fork_then_choose_one_branch) extraction has supported
since the fifth slice but the demo TwoTower corpus never exercised.
6 of 7 named patterns are now in this corpus; only pronoun_focus_followup
remains absent, deliberately, since the ninth slice already covers it on
the separate typed_operator_policy path. Reran the full SFT -> build-pairs
-> preference-training chain against the richer 6-pair corpus.


Claude-Session: https://claude.ai/code/session_01Hzz1DzercsjW7yBQaddo4Z

Co-authored-by: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants