One subdir benchmark for the search_artifacts MCP tool via rac find <query> corpus/ --json [--type T]. Retrieval quality at more realistic scale: ≥ 30 artifacts across all five types (SAB- ids), six query categories (decision_lookup, feature_lookup, disambiguation, supersession, cross_type, type_filter), four cases each.
MRR is gated to catch within-top-5 ordering regressions the P@1 / R@5 floors are blind to; hard negatives are judged against the full returned list (the top-5 window under-enforced the "never surface a superseded decision" claim — the tool serves the full match list).
Roadmap: rac/roadmaps/benchmarks/tool-benchmarks.md in rac-core [roadmap:tool-benchmarks].
One subdir benchmark for the
search_artifactsMCP tool viarac find <query> corpus/ --json [--type T]. Retrieval quality at more realistic scale: ≥ 30 artifacts across all five types (SAB-ids), six query categories (decision_lookup,feature_lookup,disambiguation,supersession,cross_type,type_filter), four cases each.MRR is gated to catch within-top-5 ordering regressions the P@1 / R@5 floors are blind to; hard negatives are judged against the full returned list (the top-5 window under-enforced the "never surface a superseded decision" claim — the tool serves the full match list).
Roadmap:
rac/roadmaps/benchmarks/tool-benchmarks.mdin rac-core [roadmap:tool-benchmarks].