I've been using FATE-H/X datasets to evaluate lemma-decomposition-based provers (Seed-Prover, Leanstral, and some experimental pipelines of my own), and I noticed there's currently no central place to compare results across systems the way PutnamBench (https://trishullab.github.io/PutnamBench/leaderboard.html) or MiniF2F leaderboards do. A few systems (e.g. Leanstral 1.5) already report FATE-H/X numbers in their own papers/blog posts, but there's no canonical place to see them side by side.
It would be cool to see a leaderboard system similar to PutnamBench
I don't mind contributing and creating a page that same way Putnam bench did however I wanted your approval
I've been using FATE-H/X datasets to evaluate lemma-decomposition-based provers (Seed-Prover, Leanstral, and some experimental pipelines of my own), and I noticed there's currently no central place to compare results across systems the way PutnamBench (https://trishullab.github.io/PutnamBench/leaderboard.html) or MiniF2F leaderboards do. A few systems (e.g. Leanstral 1.5) already report FATE-H/X numbers in their own papers/blog posts, but there's no canonical place to see them side by side.
It would be cool to see a leaderboard system similar to PutnamBench
I don't mind contributing and creating a page that same way Putnam bench did however I wanted your approval