Benchmark pipeline for evaluating file-level issue localization in repo-level LLM repair.
-
Updated
Jun 16, 2026 - Python
Benchmark pipeline for evaluating file-level issue localization in repo-level LLM repair.
Companion code for "Building AI Agents from First Principles": build a working coding agent from scratch in plain Python, one mechanism per episode, from a bare while loop to multi-agent orchestration. No frameworks, any OpenAI-compatible model.
Add a description, image, and links to the swe-bench-verified topic page so that developers can more easily learn about it.
To associate your repository with the swe-bench-verified topic, visit your repo's landing page and select "manage topics."