The alpha release pipeline for DUNE charge-light matching, covering two detectors (ND-LAr and the 2x2 demonstrator) and simulation + data.
- ND-LAr — end-to-end charge-light matching: front-stage track/shower placement, Phase 2 large-cluster scan, V2 light rescue (the
phase25_trial2_v_alpha_testmodule), Phase 3 small-cluster matrix association. Perceiver weights (~490 MB) ship as a GitHub Release asset; variance prediction is optional (constant-std fallback). This is the version integrated into flow. - 2x2 — a port of the same algorithm to the 2x2 geometry/light system, living under
TwoByTwo/, with its own perceiver (sim and data) and matcher. All 2x2 models are small (~3.5–6 MB) and committed in-repo, so the 2x2 workflows are grab-and-run with no separate download.
Both detectors emit the same per-file .pt schema (calib_hit_t0_reco etc., documented in config.yaml).
There are four scripts/run_<det>_<kind>.sh entry points. Three are runnable today; ND-data has no algorithm yet and is an intentionally-empty placeholder.
| workflow | script | status | models |
|---|---|---|---|
| ND simulation | scripts/run_nd_sim.sh |
✅ runnable | perceiver = Release asset (download once) |
| ND data | scripts/run_nd_data.sh |
⬜ empty placeholder (no ND-data pipeline yet) | — |
| 2x2 simulation | scripts/run_2x2_sim.sh |
✅ runnable | bundled in-repo |
| 2x2 data | scripts/run_2x2_data.sh |
✅ runnable | bundled in-repo |
The default is the error-matrix formulation — the validated, conservative baseline. A newer region-grow method also ships, selectable per-run, but is not the default:
VERSION |
name | what it does |
|---|---|---|
v1.0 (default) |
error-matrix | greedy per-TPC brightest-first small-cluster association, unit-variance χ² (the ND vAlpha formulation). |
v0.1 |
region-grow + tiebreaker | the development line toward the first serious release: cluster-guided spatial region-growing (confident light-matched clusters propagate t0 to neighbours; tuned conf_cos=0.55, light_margin=0.04) plus a learned-variance tiebreaker for ambiguous t0 candidates. Higher efficiency on the sim validation set (~+0.8 pp overall, +1.3 pp low-E); the variance tiebreaker itself is roughly neutral. (v2.0 = legacy alias.) |
v0.1-fx |
χ² family-expand (experimental) | cosine-free agglomerative family expansion: error-matrix-arbitrated spatial families grown from decisive seeds (TwoByTwo/matcher/family_expand_2x2.py). Beats v0.1 on the 2x2 sim aggregate (96.3% vs 95.7% vs 90.4% baseline on the hard-sample benchmark) but this 2x2 prototype scores at base=0 — known multi-flash caveat; the ND port (M5p1/postpass_v01.py) fixes it with remove-and-rescore. |
bash scripts/run_2x2_sim.sh # v1.0 error-matrix (default)
VERSION=v0.1 bash scripts/run_2x2_sim.sh # opt into region-grow + tiebreakergit clone https://github.com/MadivB/CLMatching_AlphaRelease.git
cd CLMatching_AlphaRelease
python scripts/check_install.py # 2x2 assets are bundled -> all OK, no download
# On a 4-GPU interactive node (8 workers, auto-aggregates per-event -> per-file .pt):
salloc -A dune -q interactive -C gpu --gpus-per-node=4 -N 1 -t 30 \
srun -N1 -n1 --gpus-per-node=4 bash scripts/run_2x2_sim.sh # or run_2x2_data.sh
# Result (per input FLOW file):
# output/2x2_sim_v1.0/pt_outputs/<basename>.qlmatch2x2.ptProcess a specific file (sim or data) by passing it positionally or via FILE=:
bash scripts/run_2x2_sim.sh /path/to/MiniRun6.4_1E19_RHC.flow.0000123.FLOW.hdf5
FILE=/path/to/packet-XXXX.FLOW.hdf5 bash scripts/run_2x2_data.shThe 2x2 layout (matcher + perceiver + bundled models + driver) lives under TwoByTwo/; see TwoByTwo/README_2x2.md for the package internals.
Everything below documents the ND-LAr workflow. Its perceiver weights (~490 MB) are too big to commit to git and ship instead as a GitHub Release asset. The variance-prediction model is optional (the pipeline runs with a constant-std fallback when absent).
The v0.1 development line (spatial-guided family assignment + tiebreaking) is
available for ND as a post-Phase-3 post-pass in M5p1/postpass_v01.py,
opt-in via the production CLI — the default remains the bit-identical vAlpha baseline:
python -m M5p1.phase25_trial2_v_alpha_test --files ... --out-dir ... \
--postpass v0.1 # chi2 family-expand (or: v0.1-rg = cosine region-grow)It runs before the final snapshot, so hit_timestamps_post_phase3 and the whole
NPZ → .pt chain transparently reflect it; the backbone guard is re-verified
after the pass. Only blob clusters (labels ≥ split_index) with a uniform t0 can
move; the Phase-1 track/shower backbone never does.
| variant | what it is | measured (63 MiniProdN5p1 events, energy-weighted ±160 ns) |
|---|---|---|
| (none, default) | vAlpha baseline | 92.63% |
v0.1 / v0.1-fx |
χ² family-expand: agglomerative spatial families arbitrated by the error matrix, every score against the residual base (remove-and-rescore) | 92.83% (+0.20 pp; 56/63 events improved, worst single-event −0.10 pp) |
v0.1-rg |
cosine region-grow (2x2-tuned gates) | 92.88% (+0.25 pp; 58/63 improved, but worst single-event −0.62 pp) |
The two variants trade a sliver of aggregate for robustness: the χ² family-expand is magnitude-aware (a faint displaced fragment cannot latch onto a bright flash on spatial pattern alone), which is why its worst-case regression is 6× smaller. Assigned-hit fraction is unchanged (98.85%) — v0.1 only moves assigned blobs. Post-pass cost is negligible (~7–11 s on a ~9 min event).
For iterating on the stages before the Phase 3 small-cluster error-matrix association, use the backbone-only launcher:
# on a GPU node (salloc as usual):
bash scripts/run_backbone_only.sh [/path/to/file.FLOW.hdf5]It wraps run_v_alpha_test_pt_one_file.sh with two new opt-in flags on the
M5p1.phase25_trial2_v_alpha_test CLI:
| flag | effect |
|---|---|
--skip-phase3 |
stop after V2 light rescue; Phase 3 (and any --postpass) never runs. The NPZ keeps its full schema with hit_timestamps_post_phase3 == hit_timestamps_post_v2. |
--predict-min-energy-mev 50 |
skip perceiver image prediction for clusters that only Phase 3 could match: single-TPC, non-backbone (label ≥ split_index) clusters with E ≤ 50 MeV. Backbone track/shower labels and multi-TPC clusters are always predicted, so the front stage + Phase 2 + V2 see exactly the images they would in a full run. Only valid together with --skip-phase3 (enforced — if Phase 3 ran, it would silently skip the filtered clusters instead of matching them), and must be ≤ 50 (enforced — Phase 2 treats single-TPC clusters with E > 50 MeV as primaries). |
Launcher knobs: PREDICT_MIN_E_MEV (default 50; empty string = predict all
clusters), OUT_DIR (default output/backbone_only_e<threshold>/, one
directory per threshold — completed events are skipped via their ok-JSONs, so
never re-run a different configuration into the same directory; force
recomputation with EXTRA_ARGS="--no-skip-existing"), plus everything
run_v_alpha_test_pt_one_file.sh already accepts (FILE, N_GPUS,
N_WORKERS_PER_GPU, PY, ...). Aggregation defaults to SKIP_AGGREGATE=1
because the aggregator writes Mode-A flow files back in-place — don't feed
partial quick-test results into a real file (opt in with SKIP_AGGREGATE=0).
Both flags default to off; without them the module is bit-identical to the released v0.1 behavior.
--refine-intersections (launcher: REFINE_INTERSECTIONS=1) enables a
post-clustering pass (M5p1/first_stage_matching/intersection_refine.py)
that improves hit assignment where structures cross, on the principle that a
track should keep an approximately uniform MeV/cm linear density along its
axis:
- track × cluster/shower: inside the geometric contact window the track
may run up to
ir_soft_factor(1.5×) above its own baseline density; only the excess above that soft cap is shed to the cloud, halo-first, with a per-bin hard floor at 1.0× baseline (a track can never be severed or drained). Endpoint windows are skipped by default (Bragg-peak guard;ir_endpoint_soft_factor=3.0opts into bounded shedding there). - track × track (different labels): crossing-window hits are first geometrically pinned (stage-2.5 scoring against context-fit models, large 0.35 margin), then the truly-ambiguous remainder is apportioned so both tracks continue through the crossing at the same level relative to their own baselines; the donor never drops below 1.0× its baseline continuation, and a new-gap veto rolls back any pair that would sever a track.
Hits move only between existing labels, only inside detected intersection
windows, at most once per hit; noise (-1) is untouched; no label is
created, renumbered, or retyped; the pass is deterministic and its worst
case is bounded near vanilla (energy caps, pair caps, commit-time invariant
checks with automatic fallback). Track-ends-on-track contacts are vetoed by
the same Bragg guard (ir_endpoint_guard_cm) — apportioning there would
strip a stopping track's dE/dx rise. Knobs live on ClusteringConfig
(intersection_refine_* / ir_*); stats land in
ClusteringResult.debug["intersection_refinement"]. Measured on a 176k-hit
MiniProdN5p1 event: ~0.35 s (vs ~70 s clustering), ~470 hits (0.3%)
re-assigned across ~30 active crossings.
# 1. Pick where you want the install to live, then clone.
INSTALL_DIR=/path/to/where/you/want/it
mkdir -p "$INSTALL_DIR" && cd "$INSTALL_DIR"
git clone https://github.com/MadivB/CLMatching_AlphaRelease.git
cd v_alpha_test
# 2. Preflight check (no GPU needed): tells you exactly what's missing
# and how to fix it.
python scripts/check_install.py
# Expected on a fresh clone: perceiver MISSING (required), pulse OK,
# variance optional. Exit code 1.
# 3. Download the perceiver weights (~490 MB; ~30 s on a fast network)
# into the path paths.yaml expects:
mkdir -p NewMLSection/runs/ndfull_run_distributed
curl -L -o NewMLSection/runs/ndfull_run_distributed/checkpoint.pt \
https://github.com/MadivB/CLMatching_AlphaRelease/releases/download/v0.1.0/checkpoint.pt
# 4. Verify the SHA matches release.yaml (paranoia check; recommended).
sha256sum NewMLSection/runs/ndfull_run_distributed/checkpoint.pt
# Expected:
# 38655cca2b50f2caa643ef572fb80c77332611eafd3a831215cbe0f117473ac5 ...
# 5. Re-verify the install (should now be all green).
python scripts/check_install.py
# Expected: "All required assets are present.", exit code 0.
# 6. Optional: edit paths.yaml if any of your assets/data live somewhere
# other than the defaults. paths.yaml is the single source of truth.
# 7. Run the single-file smoke test on a GPU node (assuming that you are on nersc) (~6-10 min wall clock,
# 8 workers across 4 GPUs, auto-aggregates per-event NPZ -> per-file .pt).
salloc -A dune -q interactive -C gpu --gpus-per-node=4 -N 1 -t 30 \
srun -N1 -n1 --gpus-per-node=4 \
bash scripts/run_v_alpha_test_pt_one_file.sh
# 8. Inspect the result.
python scripts/inspect_pt.py output/test_one_file/pt_outputs/*.v_alpha_test.ptThe launcher writes outputs to output/test_one_file/ inside your clone (override with OUT_DIR=...).
Expected coverage on the default test file: ~98.7% of prompt hits get a finite t0 in calib_hit_t0_reco.
If check_install.py reports a missing required asset, it prints the exact path it tried, the download URL, the copy-pasteable download command, and the paths.yaml key to edit.
After the salloc lands, in another login shell:
cd "$INSTALL_DIR"/v_alpha_test # same install dir as above
tail -f output/test_one_file/parallel8_logs/worker*.logsalloc -A dune -q interactive -C gpu --gpus-per-node=4 -N 1 -t 90 \
srun -N1 -n1 --gpus-per-node=4 \
bash scripts/run_v_alpha_test_pt_parallel8.shThe pipeline can be driven three ways, all sharing the same engine and output schema:
| # | mode | script | when to use |
|---|---|---|---|
| 1 | batch submission | scripts/submit_production_robust.sh [N] |
mass production; launches N preemption-robust SLURM chains that self-resubmit and cooperate via atomic file claims |
| 2 | interactive folder | scripts/run_interactive_forward_0000000.sh |
run inside an existing salloc GPU node; processes a whole folder forward, cooperating with any batch chains |
| 3 | single file | scripts/process_one_flow_file.sh <flow.hdf5> [out_dir] |
run inside an existing salloc GPU node; process exactly one FLOW file |
All three use 8 workers (2 per GPU × 4 GPUs) and auto-aggregate per-event NPZ shards into one per-file .pt.
Example for mode 3 (already on an interactive GPU node):
bash scripts/process_one_flow_file.sh \
/global/cfs/cdirs/dunepro/people/abooth/nd-production/output/MiniProdN5/run-ndlar-flow/MiniProdN5p1_NDComplex_FHC.flow.full.sanddrift/FLOW/0000000/MiniProdN5p1_NDComplex_FHC.flow.full.sanddrift.0000123.FLOW.hdf5
# -> output/single/<basename>/pt_outputs/<basename>.v_alpha_test.ptEvery external file the pipeline loads is listed in paths.yaml:
| asset | required? | default location |
|---|---|---|
perceiver_charge_light_relation |
yes | NewMLSection/runs/ndfull_run_distributed/checkpoint.pt (download from GitHub Release) |
pulse_template |
yes | assets/avg_pulse.npy (bundled, ~4 KB) |
variance_prediction |
optional | NewMLSection/var_prediction/runs/.../best_model.pt (constant-std fallback if missing) |
input_data.default_data_dir |
optional | NERSC default; override via CLI or paths.yaml |
Each path: can be absolute or repo-relative. path_candidates: lets you list multiple fallbacks.
You can also point the resolver at a different YAML via V_ALPHA_TEST_PATHS_YAML=/path/to/your.yaml.
v_alpha_test/
├── README.md # this file
├── paths.yaml # USER-EDITABLE asset paths (perceiver, pulse, variance)
├── config.yaml # per-file .pt output schema + field provenance
├── release.yaml # release manifest (sha256s, asset URLs, distribution)
├── assets/
│ └── avg_pulse.npy # bundled pulse template (4 KB)
├── M5p1/ # M5p1 python package (front stage + V2 + Phase 3 + resolver)
│ └── first_stage_matching/
│ └── asset_resolver.py # reads paths.yaml, validates, friendly errors
├── NewMLSection/ # perceiver model code (weights downloaded separately)
└── scripts/
├── check_install.py # validates paths.yaml; exits 1 on missing required assets
├── aggregate_to_pt.py # per-event NPZ shards -> per-file .pt
├── inspect_pt.py # peek at a per-file .pt
├── run_v_alpha_test_pt_one_file.sh # 8-worker single-file launcher (auto-aggregates)
└── run_v_alpha_test_pt_parallel8.sh # 8-worker 10-file launcher (auto-aggregates)
The launcher scripts auto-detect the repo location from their own path — they work from any clone, no editing needed.
See config.yaml for the full schema. Highlights:
Per-prompt-hit fields (size n_calib_hits):
| field | dtype | filled by |
|---|---|---|
calib_hit_t0_reco |
float32 | full pipeline (Front + Phase 2 + V2 + Phase 3); hit_timestamps_post_phase3 scattered via event.hit_refs |
prompt_hit_t_cluster_id |
int16 | front-stage labels_global re-labeled by every V2 spatial+light move (each move yields a brand-new id past the original cluster count) |
Per-merged-hit fields (size n_calib_final_hits, vBeta3-compatible):
| field | dtype | filled by |
|---|---|---|
calib_final_hit_t0_reco |
float32 | aggregator: calib_hit_t0_reco[prompt_idx[i]] where prompt_idx = charge/calib_prompt_hits/ref/charge/calib_final_hits/ref[:, 0] |
calib_final_hit_cluster_id |
int16 | aggregator: same prompt-index lookup against prompt_hit_t_cluster_id |
calib_final_hit_prompt_index |
int64 | aggregator: the column-0 ref above |
Counts + metadata:
| field | type | filled by |
|---|---|---|
n_calib_hits, n_assigned, n_unassigned |
int | aggregator |
n_calib_final_hits, n_calib_final_assigned, n_calib_final_unassigned |
int | aggregator |
processed_event_ids, all_event_ids |
int64 | aggregator |
event_summaries, failed_events |
list[dict] | aggregator |
version, algorithm, input_file, calib_final_hit_source |
str | aggregator |
Sentinels: unassigned prompt and merged hits have *_t0_reco = -1.0 and *_cluster_id = -1.
python scripts/inspect_pt.py output/test_one_file/pt_outputs/*.v_alpha_test.ptIf you ran the batch but didn't auto-aggregate, run the aggregator separately:
python scripts/aggregate_to_pt.py \
--shard-dir output/test_one_file \
--output-dir output/test_one_file/pt_outputsThe default paths in paths.yaml and the launchers are NERSC-friendly out of the box. Run on a 4-GPU GPU-node interactive allocation:
salloc -A dune -q interactive -C gpu --gpus-per-node=4 -N 1 -t 30 \
srun -N1 -n1 --gpus-per-node=4 \
bash scripts/run_v_alpha_test_pt_one_file.shFor 10 files / ~130 events in ~30 min:
salloc -A dune -q interactive -C gpu --gpus-per-node=4 -N 1 -t 90 \
srun -N1 -n1 --gpus-per-node=4 \
bash scripts/run_v_alpha_test_pt_parallel8.sh- The perceiver charge-light relation weights (
checkpoint.pt, ~490 MB) — GitHub Release asset - The variance-prediction
.pt(when produced) — GitHub Release asset, optional
Both are loaded by M5p1.first_stage_matching.load_first_stage_models using the paths from paths.yaml. Missing required assets trigger a friendly error with download instructions.