Company-neutral workflow kit for creating, reviewing, calibrating, and packaging Harbor / Terminal-Bench tasks with LLM agents.
-
Updated
May 18, 2026 - Python
Company-neutral workflow kit for creating, reviewing, calibrating, and packaging Harbor / Terminal-Bench tasks with LLM agents.
An automated media verification & reconciliation engine resolving conflicting SQLite attestations, Ed25519 signed release manifests, and lost Git verifier policies via Git reflog recovery.
An offline benchmark task using Terraform, Python, and Graphviz to reconstruct corrupted machine learning model lineage metadata, query Hugging Face dataset split counts, and render color-coded lineage graphs.
An AI coding benchmark task evaluating automated rebase conflict resolution, MLflow tracking security, secret redaction, and Hugging Face model evaluation safety.
Add a description, image, and links to the benchmark-tasks topic page so that developers can more easily learn about it.
To associate your repository with the benchmark-tasks topic, visit your repo's landing page and select "manage topics."