Skip to content

search-artifacts benchmark: ranked retrieval with MRR and full-list negatives #4

Description

@tcballard

One subdir benchmark for the search_artifacts MCP tool via rac find <query> corpus/ --json [--type T]. Retrieval quality at more realistic scale: ≥ 30 artifacts across all five types (SAB- ids), six query categories (decision_lookup, feature_lookup, disambiguation, supersession, cross_type, type_filter), four cases each.

MRR is gated to catch within-top-5 ordering regressions the P@1 / R@5 floors are blind to; hard negatives are judged against the full returned list (the top-5 window under-enforced the "never surface a superseded decision" claim — the tool serves the full match list).

Roadmap: rac/roadmaps/benchmarks/tool-benchmarks.md in rac-core [roadmap:tool-benchmarks].

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions