Version 1.0.0 Β· CHANGELOG Β· License: MIT
CLI for automatic Python generation and review using two Ollama models: a developer (DEV) and a QA reviewer. Code is formatted with Black, optionally checked with flake8, and executed in a Docker sandbox or locally. The loop continues until QA and the linter agree (with safeguards against infinite churn) or the iteration limit is reached. The interface uses Rich (banner, spinners, panels, bugs table, progress).
At a glance: presets (--preset) and presentation mode (--demo), optional HTML report, Ollama preflight and clearer runtime errors, OLLAMA_HOST for remote servers, pytest plus GitHub Actions CI.
- Startup β After config is loaded and DEV/QA model names are resolved (including CLI overrides), BugHunter queries Ollama for the local model list. If the server is unreachable or those models are not installed, the program exits with a short hint (use
--skip-model-checkto skip this, e.g. for debugging). - DEV (model from config) β generates Python code from the task description.
- Black β code is formatted (
line_lengthfromconfig.yml) before linting. - flake8 β static check (PEP 8, line length). Errors go into the bugs table and into DEV feedback.
- Execution β generated code is saved to
solution.pyand run:- Docker (default):
python:3.11-slim, read-only mount ofsolution.py,--network none,--memory=128m,--cpus=0.5,--rm. A temporary env check and optional test snippet are injected only for the run; the file is restored afterward sosolution.pystays clean. - Local fallback β if Docker is not installed or the daemon is unreachable, code runs with the local Python (with the same temporary injection and restoration).
- Docker (default):
- QA β second model receives task, code, linter and runtime output; should start the verdict with
VERDICT: PASSorVERDICT: ISSUES(seeconfig.yml). The app parses that line when present. - Iterations β QA and linter feedback are sent back to DEV for fixes; cycle repeats (limit set in config or via
-i). - Anti-loop β if the model returns identical code while linter or QA still report issues, the loop stops ("Agent Stuck in Loop").
- Final cleanup β when the target is achieved,
solution.pyis overwritten with only the approved code (Black-formatted, no debug or test snippets).
LLM output is normalized: the first ```python ... ``` block is extracted with a regex so extra text before/after does not end up in the file. Iteration logs are written to bughunter_log.txt in the background; the Rich UI is shown in the console.
QA verdict β if the QA response contains VERDICT: PASS or VERDICT: ISSUES (case-insensitive), that line decides success; otherwise the tool falls back to detecting PASS in the text (legacy).
- Docker sandbox: subprocess timeout 15 seconds (infinite loops or slow I/O in the container).
- Local fallback (when Docker is missing or unreachable): 5 seconds.
These differ because container startup adds overhead; local runs are meant to fail fast.
- Python 3.11+ (aligned with CI; requires a current stdlib and typing style used in the repo)
- Ollama with a running server and models (names in
config.yml; defaultqwen2.5-coder:7bfor DEV and QA). To use a remote Ollama instance, set theOLLAMA_HOSTenvironment variable (URL including scheme, e.g.http://192.168.1.10:11434) before starting BugHunter; the official Python client reads it automatically. - Docker (optional) β for sandboxed execution. If Docker is not installed or not running, execution falls back to the local Python interpreter.
- Optional: flake8 on your PATH for linting generated code (also listed in
requirements.txtfor dev/CI).
pip install -r requirements.txtDependencies: ollama, black, PyYAML, rich, pytest (for the test suite in tests/).
- Start Ollama and pull the models named in
config.yml(defaults expectollama pull qwen2.5-coder:7bunless you pass--model/--qa-model). - Install:
python -m venv .venv && .venv/bin/python -m pip install -r requirements.txt - Verify:
.venv/bin/python main.py --version - Run:
.venv/bin/python main.pyor e.g..venv/bin/python main.py --preset fibo --demo
Docker is optional; if it is missing or the daemon is down, execution uses the same interpreter as in step 2.
From the project root, with dependencies installed (recommended: virtual environment and pip install -r requirements.txt):
python -m pytest tests/If your venv uses a different Python than the pip you invoked, use the same interpreter for both install and test, for example .venv/bin/python -m pip install -r requirements.txt then .venv/bin/python -m pytest tests/.
On GitHub, CI (.github/workflows/ci.yml) runs pytest and flake8 on push and pull requests to main / master for Python 3.11β3.13. Flake8 rules live in .flake8 (line length 88; long lines are ignored as E501 so the tool matches Black-focused workflows).
With a virtual environment:
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txtEnsure Ollama is running and required models are pulled (names from config.yml):
ollama pull qwen2.5-coder:7bOptional: run flake8 locally the same way as CI:
python -m flake8 main.py ui_utils.py tests/For sandboxed runs, have Docker installed and the daemon running. If not, BugHunter will warn and use local execution.
Settings are in config.yml in the project root (next to main.py). Example:
models:
dev: "qwen2.5-coder:7b"
qa: "qwen2.5-coder:7b"
settings:
max_iterations: 5
line_length: 88
prompts:
developer: |
You are a senior Python developer. Write clean, self-contained code.
IMPORTANT:
1. Never exceed {line_length} characters per line.
...
qa: |
Analyze the code for logic, PEP 8 compliance, and task completion.
...- models.dev / models.qa β Ollama model names for developer and QA.
- settings.max_iterations β maximum number of loop iterations.
- settings.line_length β line length limit for Black and flake8.
- prompts.developer / prompts.qa β system prompts;
{line_length}is available in the developer prompt.
If the config file is missing or empty, the script exits with an error. By default the file config.yml in the project root (next to main.py) is used; override with -c / --config PATH.
Exit codes: 0 β finished the hunt loop (success or max iterations). 1 β config / startup checks (missing config, empty YAML, Ollama unreachable or missing models during preflight). 2 β Ollama request failed during DEV or QA chat (e.g. connection dropped mid-run).
python main.pyWithout arguments, the default task is a short one (implement add(a, b)). Use --preset for built-in demos: sum (same as default), fibo (Fibonacci list), path (nested dict path resolver). If you pass both a positional task and --preset, the preset wins and the positional task is ignored (with an info message).
--demo turns on presentation mode: Rich section rules between iterations and a short pause so output is easier to follow when presenting or recording.
Custom task as positional argument:
python main.py "Write a function that returns the factorial of n."Flags override config:
| Argument | Description |
|---|---|
--version |
Print BugHunter version and exit (no run). |
task |
Task description (optional positional; default is a short add(a, b) task). |
-c, --config PATH |
YAML config file (default: config.yml next to main.py). |
--preset NAME |
Built-in task: sum, fibo, or path. Overrides positional task when set. |
--demo |
Presentation mode: section dividers between iterations and short pauses. |
--skip-model-check |
Do not verify Ollama reachability or that DEV/QA models are installed locally. |
-i, --iters N |
Max iterations (overrides settings.max_iterations). |
--model NAME |
Ollama model for DEV (overrides models.dev). |
--qa-model NAME |
Ollama model for QA (overrides models.qa). |
--html-report PATH |
After the run, write a static HTML summary of each iteration (task, code, linter, runtime, QA). |
-t, --test PYCODE |
Python code to append at the end of the script for testing (e.g. print(my_func(10))). Injected only during execution; not saved to solution.py. |
Examples:
python main.py
python main.py --version
python main.py --preset fibo --demo
python main.py --preset path -i 8
python main.py --preset fibo --html-report report.html
python main.py -c ./config.yml --qa-model mistral
python main.py "Sum of a list of numbers"
python main.py "Factorial" -i 10 --model llama3
python main.py "Write a sum function" --test "print(sum(5, 5))"
python main.py "Parse CSV into dict" --iters 3- Banner β BugHunter title panel at start.
- LLM status β Spinner with "LLM generating code..." / "LLM analyzing code..." during model calls.
- [AI Assistant] panels β DEV output (code with syntax highlight). QA output is rendered as Markdown inside the panel when the model uses headings, lists, or emphasis.
- Bugs table β After QA: error type, line, severity, recommendation (including flake8 output and QA verdict).
- Progress β Iteration bar with percentage and time remaining.
- Summary β Panel with paths to
solution.pyand log file; then panel with final code (Syntax, monokai theme).
- solution.py β final generated code only (no debug or test snippets). Overwritten with clean, Black-formatted code when the target is achieved. In
.gitignore. - bughunter_log.txt β step-by-step iteration log, including Docker command and raw stdout/stderr when using the sandbox. Written in the background; in
.gitignore. - Optional HTML report β with
--html-report PATH, a single static HTML file is written after the hunt: task plus one section per completed iteration (code, linter, runtime, QA; long fields are truncated for file size). Safe for sharing (content is HTML-escaped in<pre>blocks).
- main.py β config load, argparse,
BugHunterclass (Docker/local execution, regex extraction, final clean write), entry point underif __name__ == "__main__". - ui_utils.py β Rich console, status spinners, panels (including Markdown for QA), tables, syntax highlight, background file logging, optional HTML report writer.
- pyproject.toml β pytest and Black tool defaults.
- .flake8 β flake8 defaults for local runs and CI.
- .github/workflows/ci.yml β GitHub Actions: pytest + flake8 on supported Python versions.
- CHANGELOG.md β release notes.
- LICENSE β MIT.
- config.yml β models, limits, prompts (required to run).
- tests/ β
pytestunit tests for verdict parsing, markdown stripping, config load, CLI parser, linter row parsing, HTML report, Ollama helpers.