Skip to content

Qualify the First Real METR-LA Candidate Run #41

Description

@RMKruse

Parent

#35

What to build

Type: HITL.

Execute the single-seed primary comparison with real METR-LA artifacts and review the resulting training and evaluation story before scaling out. Confirm data availability and compute budget, inspect split and preprocessing provenance, review the selected checkpoint and interpretability artifacts, and freeze or reject the run recipe for the multi-seed Merit Probe.

Acceptance criteria

  • The required METR-LA artifacts and available compute budget are recorded before execution.
  • One declared seed completes the deterministic NAMLSS and optimizer-trained GraphNAMLSS primary comparison or records a concrete blocker.
  • The review confirms temporal split discipline, train-only scaling, validation-based selection, and post-selection test evaluation.
  • Training history, selected checkpoint, predictions, Local Shape Contributions, Graph Contributions, and Graph Gate Diagnostics are inspected for plausibility.
  • The outcome either freezes a reproducible run recipe for multi-seed execution or records bounded corrective actions.
  • No Further-Investment Gate decision is made from this single qualifying run.

Blocked by

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions