Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
159 commits
Select commit Hold shift + click to select a range
064973e
feat: evidence-grade NIM discovery + all-modality cost-quality benchm…
seonghobae Aug 4, 2026
b48f396
merge: stack NIM benchmark on provider-egress security base
seonghobae Aug 4, 2026
0e19b3a
docs(changelog): record NIM benchmark evidence slice
seonghobae Aug 4, 2026
66c8d55
ci: apply reviewed NIM transport hardening
seonghobae Aug 4, 2026
9d546a7
fix(nim): harden egress and equalize policy budgets
seonghobae Aug 4, 2026
785fb15
fix(nim): install evidence and transport hardening
seonghobae Aug 4, 2026
e07a504
test(nim): prove secure transport and equal budgets
seonghobae Aug 4, 2026
d0312ea
docs(nim): document pinned egress and equal total budgets
seonghobae Aug 4, 2026
c7bcad2
chore(ci): remove temporary NIM repair workflow
seonghobae Aug 4, 2026
e1ae745
chore(ci): stage one-time NIM source repair
seonghobae Aug 4, 2026
1015608
chore(ci): remove privileged one-shot NIM repair workflow
seonghobae Aug 4, 2026
07128ac
chore(ci): stage exact-head NIM security repair
seonghobae Aug 4, 2026
a8393d7
chore(ci): harden bounded NIM repair job
seonghobae Aug 4, 2026
4cc0f55
chore(ci): stage bounded NIM source integration
seonghobae Aug 4, 2026
9ad742a
chore(ci): remove overlapping one-shot NIM workflow
seonghobae Aug 4, 2026
4fd96f1
fix(ci): replace write-capable PR repair with isolated publisher
seonghobae Aug 4, 2026
4bb7c85
docs(changelog): record reviewed NIM access evidence
seonghobae Aug 4, 2026
d9bedd7
chore(ci): stage equal-budget regression repair
seonghobae Aug 4, 2026
c4d47d2
ci: complete bounded NIM source integration
seonghobae Aug 4, 2026
db6ebb9
chore(ci): remove redundant NIM test-only repair workflow
seonghobae Aug 4, 2026
dba1d7e
ci: run bounded NIM source integration from PR checks
seonghobae Aug 4, 2026
a3741b5
fix(ci): execute bounded NIM source transformation deterministically
seonghobae Aug 4, 2026
18d3fe7
fix(ci): preserve embedded transformation literals
seonghobae Aug 4, 2026
27eb1b8
fix(security): remove PR-head OIDC source mutation
seonghobae Aug 4, 2026
ba6b34b
fix(security): delete privileged branch-trigger repair workflow
seonghobae Aug 4, 2026
e4578da
fix(security): delete branch-controlled OIDC repair source
seonghobae Aug 4, 2026
f80e503
test(security): pin PR workflow OIDC boundary
seonghobae Aug 4, 2026
bc77bfc
ci(review): export read-only exact-head workspace
seonghobae Aug 4, 2026
a71ebbc
ci(review): include inert transformation evidence
seonghobae Aug 4, 2026
6df72ee
test(ci): require secretless NIM dry-run job
seonghobae Aug 5, 2026
09fd519
fix(ci): isolate NIM dry runs from provider secrets
seonghobae Aug 5, 2026
2b622bd
docs(security): record NIM workflow secret isolation
seonghobae Aug 5, 2026
40b48ba
docs(changelog): record secretless NIM dry runs
seonghobae Aug 5, 2026
348e5e6
docs(benchmark): explain workflow credential isolation
seonghobae Aug 5, 2026
667b311
chore(review): stage verified NIM repair patch
seonghobae Aug 5, 2026
2498648
ci(nim): stage reviewed patch part 1 of 7
seonghobae Aug 5, 2026
4552084
ci(nim): stage reviewed patch part 2 of 7
seonghobae Aug 5, 2026
45c151e
ci(nim): stage reviewed patch part 3 of 7
seonghobae Aug 5, 2026
9dac6b2
ci(nim): stage reviewed patch part 4 of 7
seonghobae Aug 5, 2026
fecd25f
ci(nim): stage reviewed patch part 5 of 7
seonghobae Aug 5, 2026
19be572
ci(nim): stage reviewed patch part 6 of 7
seonghobae Aug 5, 2026
736893c
ci(nim): stage reviewed patch part 7 of 7
seonghobae Aug 5, 2026
07f28a9
ci(nim): apply exact reviewed benchmark integration
seonghobae Aug 5, 2026
7d5d6aa
ci: retrigger reviewed NIM integration with scoped permissions
seonghobae Aug 5, 2026
fd6b074
ci: verify and export reviewed NIM integration read-only
seonghobae Aug 5, 2026
cb7e9b3
ci: bind NIM verification to security ancestor
seonghobae Aug 5, 2026
3878e5b
fix(ci): bind NIM verification to pull-request head
seonghobae Aug 5, 2026
5bbedc2
fix(ci): reconcile reviewed NIM patch per file
seonghobae Aug 5, 2026
d98fe6c
fix(ci): reconstruct reviewed NIM tree from patch base
seonghobae Aug 5, 2026
428736c
fix(ci): bind reviewed NIM reconstruction to exact head
seonghobae Aug 5, 2026
e4e03c4
fix(ci): run reviewed NIM verification read-only on push
seonghobae Aug 5, 2026
6cd2012
fix(ci): use runner temp variable at execution time
seonghobae Aug 5, 2026
7595919
fix(ci): discover the exact reviewed patch base
seonghobae Aug 5, 2026
a93ec20
chore(ci): inspect the staged NIM integration payload
seonghobae Aug 5, 2026
f11193d
fix(ci): replay the reviewed NIM patch stack
seonghobae Aug 5, 2026
83b7906
fix(ci): rehydrate the reviewed patch index
seonghobae Aug 5, 2026
e9bfb85
fix(ci): execute the patch rehydration helper
seonghobae Aug 5, 2026
88274c0
fix(ci): apply the reviewed patch against the rehydrated index
seonghobae Aug 5, 2026
03b30e5
fix(ci): verify the reconstructed reviewed tree
seonghobae Aug 5, 2026
e256a9a
fix(ci): locate the exact reviewed patch context
seonghobae Aug 5, 2026
012f7c7
fix(ci): build the wheel in an isolated backend environment
seonghobae Aug 5, 2026
dde856f
ci(nim): publish the authenticated verified source tree
seonghobae Aug 5, 2026
8049ca8
feat(nim): integrate verified provider-neutral benchmark
seonghobae Aug 5, 2026
d1516a5
docs(nim): record immutable integration provenance
seonghobae Aug 5, 2026
f090ffc
test(ci): require stacked pull-request coverage
seonghobae Aug 5, 2026
d091b69
ci: validate stacked pull requests
seonghobae Aug 5, 2026
0d66e0a
ci: fuzz stacked pull requests
seonghobae Aug 5, 2026
928214c
ci: secure stacked pull requests
seonghobae Aug 5, 2026
c759cd1
test(ci): require isolated wheel builds
seonghobae Aug 5, 2026
8e19d64
fix(ci): isolate benchmark wheel builds
seonghobae Aug 5, 2026
e7f7238
ci(review): export read-only PR90 workspace
seonghobae Aug 5, 2026
c72f1cb
test(benchmark): stage complete-budget preflight red tests part 1
seonghobae Aug 5, 2026
cb20657
test(benchmark): stage complete-budget preflight red tests part 2
seonghobae Aug 5, 2026
636c3e0
fix(benchmark): stage complete request planner part 3
seonghobae Aug 5, 2026
1053f19
docs(benchmark): stage complete-budget evidence and verification part 4
seonghobae Aug 5, 2026
638a126
ci(benchmark): apply complete-budget preflight test-first
seonghobae Aug 5, 2026
1f6da53
ci(review): stage exact PR90 request-plan patch
seonghobae Aug 5, 2026
8a31de7
ci(review): refresh exact PR90 request-plan patch
seonghobae Aug 5, 2026
87496c8
ci(review): verify and publish exact PR90 request-plan fix
seonghobae Aug 5, 2026
b3c0471
ci(review): inspect exact PR90 patch hash
seonghobae Aug 5, 2026
2efcf08
ci(review): export decoded PR90 patch evidence
seonghobae Aug 5, 2026
a06f9c5
ci(review): verify exact staged PR90 patch
seonghobae Aug 5, 2026
77cc8cc
ci(review): publish verified PR90 request-plan fix
seonghobae Aug 5, 2026
796e6e9
test(benchmark): align scheduled complete-plan contract
seonghobae Aug 5, 2026
5956214
test(benchmark): require request-plan provenance schema
seonghobae Aug 5, 2026
9b72816
fix(nim): require a complete catalog benchmark plan
seonghobae Aug 5, 2026
a2af363
merge: integrate provider-egress security base into NIM benchmark
github-actions[bot] Aug 5, 2026
c5b3876
docs(adr): record exact NIM security integration evidence
seonghobae Aug 5, 2026
d43adaa
test(benchmark): expose undersized locked evaluation manifest
seonghobae Aug 5, 2026
16b5016
feat(benchmark): expand locked manifest to evidence floor
seonghobae Aug 5, 2026
ac13a56
test(benchmark): bind evidence floor to existing hard ceiling
seonghobae Aug 5, 2026
62e0e59
docs(benchmark): record thirty-task evidence floor
seonghobae Aug 5, 2026
43c629f
docs(benchmark): align request plan and evidence floor
seonghobae Aug 5, 2026
ff51584
docs(benchmark): doctor thirty-task evidence floor
seonghobae Aug 5, 2026
bec1129
test(benchmark): expose undersized default request caps
seonghobae Aug 5, 2026
cdfe08e
ci(benchmark): test and repair default request caps
seonghobae Aug 5, 2026
37c78b4
ci(benchmark): rearm default request-cap repair
seonghobae Aug 5, 2026
545cd64
ci(benchmark): add exact-head reopen trigger
seonghobae Aug 5, 2026
e56d1a3
ci(benchmark): use resolvable checkout action
seonghobae Aug 5, 2026
5216171
ci(benchmark): repair complete-plan default finalizer
seonghobae Aug 5, 2026
8cae8d4
test(benchmark): bind evidence floor to canonical request planner
seonghobae Aug 5, 2026
e31f901
ci(benchmark): cover all complete-plan regression anchors
seonghobae Aug 5, 2026
da8ec5f
ci(benchmark): repair invalid one-shot YAML
seonghobae Aug 5, 2026
1b97e48
ci(benchmark): split workflow-authorized publication
seonghobae Aug 5, 2026
0ebb1d5
fix(benchmark): align complete-plan source evidence
github-actions[bot] Aug 5, 2026
9e4b50b
fix(benchmark): align manual request cap with complete plan
seonghobae Aug 5, 2026
821b58e
ci(benchmark): remove completed one-shot workflow
seonghobae Aug 5, 2026
7b23bac
test(nim): cover buyer-facing complete request plan
seonghobae Aug 5, 2026
5c1a3f6
ci(nim): include request-plan regression in quality gate
seonghobae Aug 5, 2026
f5a494f
test: require NIM parser Atheris instrumentation
seonghobae Aug 5, 2026
48f007d
fix: instrument NIM parser imports for Atheris
seonghobae Aug 5, 2026
2736113
docs: distinguish historical NIM integration evidence
seonghobae Aug 5, 2026
b60ec5d
test: cover NIM review regressions
seonghobae Aug 5, 2026
bdc9673
chore: defer broader NIM regression batch
seonghobae Aug 5, 2026
646c417
docs: make NIM receipt head-stable
seonghobae Aug 5, 2026
318feb9
docs: record NIM fuzz instrumentation repair
seonghobae Aug 5, 2026
63438d0
test(nim): require completion-dependent README evidence status
seonghobae Aug 5, 2026
6fc53f1
docs(nim): clarify completion-dependent evidence status
seonghobae Aug 5, 2026
4a79964
test(nim): normalize README contract whitespace
seonghobae Aug 5, 2026
1217fc9
docs(changelog): record NIM evidence-status clarification
seonghobae Aug 5, 2026
11353d4
test(security): inspect every pull-request workflow structurally
seonghobae Aug 5, 2026
67c5c84
test(nim): enforce ordered non-empty secret-isolation steps
seonghobae Aug 5, 2026
989d7fe
test(security): parse pull-request branch filters structurally
seonghobae Aug 5, 2026
4329081
test(nim): lock review regressions before repair
seonghobae Aug 5, 2026
715ee76
fix(fuzz): require normalized NIM catalog depth failures
seonghobae Aug 5, 2026
8c4ee86
test: skip fixture-bound standalone NIM cases
seonghobae Aug 5, 2026
4ca0448
fix: close NIM evidence review gaps
seonghobae Aug 5, 2026
aed3e7c
docs: record NIM evidence hardening
seonghobae Aug 5, 2026
a5d529b
test: require NIM review regressions in coverage gate
seonghobae Aug 5, 2026
8603bcf
ci: cover current NIM review regressions
seonghobae Aug 5, 2026
e9f2072
docs: record exact NIM regression gate
seonghobae Aug 5, 2026
ffd03ba
test(nim): require complete CSV model-assignment evidence
seonghobae Aug 5, 2026
a9bf8bc
feat(nim): preserve model assignment evidence in CSV artifacts
seonghobae Aug 5, 2026
d031f53
feat(nim): fail closed until CSV evidence is complete
seonghobae Aug 5, 2026
554a220
ci(nim): gate CSV evidence adapter at 100% coverage
seonghobae Aug 5, 2026
2db0737
docs(nim): record CSV assignment evidence parity
seonghobae Aug 5, 2026
dfc300a
docs(nim): doctor CSV assignment evidence boundary
seonghobae Aug 5, 2026
5c2c695
refactor(nim): preserve CSV mode without cleanup branches
seonghobae Aug 5, 2026
7dd5245
test(nim): cover CSV evidence fail-closed edge cases
seonghobae Aug 5, 2026
816707f
ci(nim): execute CSV evidence edge regressions
seonghobae Aug 5, 2026
00483f3
test(nim): make routing disclaimer contract case-insensitive
seonghobae Aug 5, 2026
88b0e96
ci(nim): preserve benchmark docstring contract and gate adapter
seonghobae Aug 5, 2026
c9fa27c
test(ci): require pull request workflows to checkout exact heads
seonghobae Aug 5, 2026
5e46772
ci: checkout exact pull request heads in test jobs
seonghobae Aug 5, 2026
f7320aa
ci: checkout exact pull request heads in fuzz jobs
seonghobae Aug 5, 2026
16f1409
ci: checkout exact pull request heads in security jobs
seonghobae Aug 5, 2026
692e49a
docs(ci): record exact-head pull request verification boundary
seonghobae Aug 5, 2026
1ef2c97
docs(ci): doctor exact-head versus merge-test evidence
seonghobae Aug 5, 2026
f042953
merge: synchronize NIM benchmark stack with PR #96
seonghobae Aug 5, 2026
2b84e1e
test(nim): define transactional artifact publication contract
seonghobae Aug 6, 2026
ef28784
fix(nim): publish complete evidence sets transactionally
seonghobae Aug 6, 2026
ed52bc8
test(nim): route CSV wrapper fakes through private staging
seonghobae Aug 6, 2026
5771957
ci(nim): include transactional publication regressions
seonghobae Aug 6, 2026
e1f19ff
test(nim): write fake artifacts into precreated staging
seonghobae Aug 6, 2026
c446251
test(nim): reuse wrapper-created staging directory
seonghobae Aug 6, 2026
30c046f
test(nim): cover transactional publication failure edges
seonghobae Aug 6, 2026
6ed1fbf
ci(nim): include publication edge coverage
seonghobae Aug 6, 2026
5e46ba5
test(nim): cover duplicate split output argument
seonghobae Aug 6, 2026
1cea0c0
docs(nim): record transactional evidence-set publication
seonghobae Aug 6, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 5 additions & 1 deletion .github/workflows/fuzz.yml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,6 @@ on:
push:
branches: [main]
pull_request:
branches: [main]
schedule:
# Weekly deeper run (longer per-target budget via FUZZ_SECONDS).
- cron: "41 4 * * 2"
Expand All @@ -26,6 +25,7 @@ jobs:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.sha || github.sha }}
persist-credentials: false

- name: Set up Python
Expand All @@ -50,6 +50,7 @@ jobs:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.sha || github.sha }}
persist-credentials: false

- name: Set up Python
Expand Down Expand Up @@ -83,6 +84,9 @@ jobs:
- name: Fuzz orchestration engine
run: python fuzz/fuzz_orchestration.py -max_total_time=${FUZZ_SECONDS} -artifact_prefix=crash- fuzz/corpus/orchestration

- name: Fuzz NIM model-catalog parser
run: python fuzz/fuzz_nim_catalog.py -max_total_time=${FUZZ_SECONDS} -artifact_prefix=crash- fuzz/corpus/nim_catalog

- name: Upload crash artifacts
if: failure()
uses: actions/upload-artifact@330a01c490aca151604b8cf639adc76d48f6c5d4 # actions/upload-artifact@v5
Expand Down
149 changes: 149 additions & 0 deletions .github/workflows/nim-benchmark.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,149 @@
name: NIM benchmark

# Evidence-grade NVIDIA NIM model discovery and cost-quality benchmark.
# Manual dispatch defaults to a deterministic dry run. Monthly scheduled runs
# are live but use a conservative hard request cap. Dry execution receives no
# provider credential; only the live step can read NVIDIA_NIM_API_KEY.

on:
workflow_dispatch:
inputs:
dry_run:
description: "Dry run without contacting NVIDIA"
type: boolean
default: true
max_total_requests:
description: "Hard cap on provider requests for this run"
type: number
default: 2000
pricing_scenario:
description: "Optional reviewed pricing-scenario JSON path"
type: string
default: ""
schedule:
- cron: "23 3 5 * *"

permissions:
contents: read

concurrency:
group: nim-benchmark
cancel-in-progress: false

jobs:
dry_run_benchmark:
name: Deterministic NIM benchmark dry run
if: github.event_name == 'workflow_dispatch' && inputs.dry_run == true
runs-on: ubuntu-latest
timeout-minutes: 30
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
persist-credentials: false

- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # actions/setup-python@v6
with:
python-version: "3.12"

- name: Install pinned runtime
run: |
python -m pip install --require-hashes -r requirements.lock
python -m pip install --no-deps -e .

- name: Run dry benchmark
env:
MAX_REQUESTS: ${{ inputs.max_total_requests }}
PRICING_SCENARIO: ${{ inputs.pricing_scenario }}
PROVENANCE_GIT_SHA: ${{ github.sha }}
PROVENANCE_RUN_ID: ${{ github.run_id }}
run: |
set -euo pipefail
extra_args=()
if [ -n "$PRICING_SCENARIO" ]; then
extra_args+=(--pricing-scenario "$PRICING_SCENARIO")
fi
python -m contextual_orchestrator nim-benchmark \
--dry-run \
"${extra_args[@]}" \
--task-manifest examples/nim_task_manifest.json \
--output-dir benchmark_artifacts \
--max-total-requests "$MAX_REQUESTS" \
--git-sha "$PROVENANCE_GIT_SHA" \
--workflow-run-id "$PROVENANCE_RUN_ID"

- name: Upload dry-run artifacts
uses: actions/upload-artifact@330a01c490aca151604b8cf639adc76d48f6c5d4 # actions/upload-artifact@v5
with:
name: nim-benchmark-dry-${{ github.run_id }}
path: benchmark_artifacts/
retention-days: 30
if-no-files-found: error

live_benchmark:
name: Live NIM catalog benchmark
if: github.event_name == 'schedule' || (github.event_name == 'workflow_dispatch' && inputs.dry_run != true)
runs-on: ubuntu-latest
timeout-minutes: 60
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
persist-credentials: false

- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # actions/setup-python@v6
with:
python-version: "3.12"

- name: Install pinned runtime
run: |
python -m pip install --require-hashes -r requirements.lock
python -m pip install --no-deps -e .

- name: Resolve live parameters
id: live_params
env:
EVENT_NAME: ${{ github.event_name }}
INPUT_MAX_REQUESTS: ${{ inputs.max_total_requests }}
INPUT_PRICING: ${{ inputs.pricing_scenario }}
run: |
set -euo pipefail
if [ "$EVENT_NAME" = "schedule" ]; then
echo "max_requests=2000" >> "$GITHUB_OUTPUT"
echo "pricing_scenario=" >> "$GITHUB_OUTPUT"
else
echo "max_requests=${INPUT_MAX_REQUESTS}" >> "$GITHUB_OUTPUT"
echo "pricing_scenario=${INPUT_PRICING}" >> "$GITHUB_OUTPUT"
fi

- name: Run live benchmark
env:
NVIDIA_NIM_API_KEY: ${{ secrets.NVIDIA_NIM_API_KEY }}
MAX_REQUESTS: ${{ steps.live_params.outputs.max_requests }}
PRICING_SCENARIO: ${{ steps.live_params.outputs.pricing_scenario }}
PROVENANCE_GIT_SHA: ${{ github.sha }}
PROVENANCE_RUN_ID: ${{ github.run_id }}
run: |
set -euo pipefail
extra_args=()
if [ -n "$PRICING_SCENARIO" ]; then
extra_args+=(--pricing-scenario "$PRICING_SCENARIO")
fi
python -m contextual_orchestrator nim-benchmark \
"${extra_args[@]}" \
--task-manifest examples/nim_task_manifest.json \
--output-dir benchmark_artifacts \
--max-total-requests "$MAX_REQUESTS" \
--git-sha "$PROVENANCE_GIT_SHA" \
--workflow-run-id "$PROVENANCE_RUN_ID"

- name: Upload live benchmark artifacts
if: always()
uses: actions/upload-artifact@330a01c490aca151604b8cf639adc76d48f6c5d4 # actions/upload-artifact@v5
with:
name: nim-benchmark-live-${{ github.run_id }}
path: benchmark_artifacts/
retention-days: 90
if-no-files-found: error
3 changes: 2 additions & 1 deletion .github/workflows/security.yml
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,6 @@ on:
push:
branches: [main]
pull_request:
branches: [main]
schedule:
- cron: "17 3 * * 1"
workflow_dispatch:
Expand All @@ -34,6 +33,7 @@ jobs:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.sha || github.sha }}
persist-credentials: false

- name: Initialize CodeQL
Expand All @@ -54,6 +54,7 @@ jobs:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.sha || github.sha }}
persist-credentials: false

- name: Set up Python
Expand Down
56 changes: 55 additions & 1 deletion .github/workflows/tests.yml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,6 @@ on:
push:
branches: [main]
pull_request:
branches: [main]

permissions:
contents: read
Expand All @@ -21,6 +20,7 @@ jobs:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.sha || github.sha }}
persist-credentials: false

- name: Set up Python
Expand All @@ -36,3 +36,57 @@ jobs:

- name: Run full test suite
run: python -m pytest -q

nim_benchmark_quality:
name: NIM benchmark coverage, docstrings, and package smoke
runs-on: ubuntu-latest
steps:
- name: Checkout repository
uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # actions/checkout@v7
with:
ref: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.sha || github.sha }}
persist-credentials: false

- name: Set up Python
uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # actions/setup-python@v6
with:
python-version: "3.12"

- name: Install hash-locked quality tools
run: python -m pip install --require-hashes -r requirements-opencode-review-ci.txt

- name: Prove complete benchmark coverage and public docstrings
run: |
set -euo pipefail
python -m coverage erase
python -m coverage run --branch \
--source=contextual_orchestrator.nim_benchmark,contextual_orchestrator.nim_csv_evidence \
-m pytest \
tests/test_nim_benchmark.py \
tests/test_nim_benchmark_budget_view.py \
tests/test_nim_benchmark_release_acceptance.py \
tests/test_nim_benchmark_review_regressions.py \
tests/test_nim_benchmark_workflow_contract.py \
tests/test_nim_artifact_publication.py \
tests/test_nim_artifact_publication_edges.py \
tests/test_nim_csv_evidence.py \
tests/test_nim_csv_evidence_edges.py \
-q
python -m coverage report \
--include=contextual_orchestrator/nim_benchmark.py,contextual_orchestrator/nim_csv_evidence.py \
--show-missing \
--fail-under=100
python -m interrogate -f 100 contextual_orchestrator/nim_benchmark.py
python -m interrogate -f 100 contextual_orchestrator/nim_csv_evidence.py

- name: Build, install, and import the wheel
run: |
set -euo pipefail
rm -rf dist "$RUNNER_TEMP/nim-wheel-site"
python -m pip wheel --no-deps . --wheel-dir dist
python -m pip install --no-deps \
--target "$RUNNER_TEMP/nim-wheel-site" \
dist/contextual_orchestrator-*.whl
cd "$RUNNER_TEMP"
PYTHONPATH="$RUNNER_TEMP/nim-wheel-site" \
python -c "import contextual_orchestrator; import contextual_orchestrator.nim_benchmark; import contextual_orchestrator.nim_csv_evidence"
6 changes: 6 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -12,3 +12,9 @@ tempcred.txt

# hypothesis fuzzing DB
.hypothesis/

# coverage measurement data
.coverage

# local benchmark artifacts (uploaded by CI, not committed)
benchmark_artifacts/
28 changes: 26 additions & 2 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,19 +6,43 @@ The format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and

## [Unreleased]

### Added

- Add an optional provider-neutral NVIDIA NIM benchmark harness that dynamically discovers the live `/v1/models` catalog, probes every discovered model under bounded concurrency and a hard request cap, records machine-readable capability outcomes, and compares direct, route-once, bounded-conduct, and explicit pricing-scenario policies over a locked task manifest.
- Add deterministic no-egress benchmark dry runs, secret-redacted JSON/CSV/Markdown evidence artifacts, paired bootstrap uncertainty, quality-latency and quality-hypothetical-cost Pareto frontiers, all-modality catalog fuzzing, and a manually gated benchmark workflow.
- Add a validated deterministic one-frame H.264 MP4 probe fixture, complete preflight reservation for every discovered model-capability cell plus the full evaluation envelope, and a thirty-task locked manifest that reaches the declared paired-evidence floor without creating an automatic routing recommendation.
- Add direct benchmark quality gates for 100% production statement/branch coverage, 100% public docstrings, wheel build/install/import smoke testing, and optional-import isolation.

### Security

- Restrict the private plain-HTTP provider seam to `localhost` or literal loopback IP addresses, reject URL userinfo before connection, dial directly without ambient proxy lookup, reject all redirect responses, and close failed resources deterministically.
- Pin each HTTPS provider connection to the exact public addresses approved during validation, preserve the original hostname for TLS verification, bypass environment proxy resolution, and reject redirects to close DNS-rebinding and credential-forwarding SSRF paths.
- Integrate DNS-pinned provider dispatch directly into `ModelClient` so package import performs no optional-adapter monkey-patching or order-dependent class mutation.
- Reject provider hosts that resolve to any non-globally-routable address, including RFC 6598 shared address space, while retaining explicit multicast, private, loopback, link-local, and reserved-address protections.
- Document narrowly scoped Semgrep suppressions for parameter-bound database queries, the explicit development-only TLS verification opt-out, and provider URLs that pass the egress guard.
- Document narrowly scoped Semgrep suppressions for parameter-bound database queries and the explicit development-only TLS verification opt-out.
- Remove the NIM benchmark's dynamic `urllib.request.urlopen` sink and compatibility monkeypatch path; live discovery, probes, and evaluation now use direct validation-time-address-pinned TLS with original-host SNI/certificate verification, no proxy lookup, no redirect following, and deterministic cleanup.
- Bound every NIM provider response to 8 MiB and fail closed before an oversized catalog, probe, or evaluation body can exhaust benchmark-runner memory.
- Split scheduled dry and live benchmark jobs so zero-egress dry runs never receive `NVIDIA_NIM_API_KEY`; only the bounded live benchmark step receives the GitHub Secret.
- Remove temporary branch-writing/source-export repair mechanisms and the optional benchmark monkeypatch module from the mergeable tree.
- Import the NIM catalog parser while Atheris import instrumentation is active, with an AST regression contract that prevents parser branches from silently losing coverage guidance.

### Changed

- Make every repository-local pull-request checkout in Tests, Fuzz, and Security select `github.event.pull_request.head.sha` rather than GitHub's synthetic merge ref, while preserving `github.sha` for push, schedule, and manual runs and retaining non-persisted checkout credentials.
- Preserve every benchmark cell's step, role, agent, and model assignment in CLI and scheduled-workflow CSV artifacts as deterministic JSON; fail closed before publishing success when JSON/CSV identities are duplicated, malformed, or incomplete, and replace the enriched CSV atomically.
- Exclude zero-success benchmark policies from quality Pareto frontiers with explicit evidence labels, validate every Markdown-consumed report field before artifact writes, normalize excessive catalog JSON depth to the catalog domain error, preserve positive sub-second provider timeouts, make the standalone NIM test runner fixture-safe, and execute these review regressions inside the 100% branch-coverage gate.
- Expose a stable complete-run request planning view and align API, CLI, manual-workflow, and deterministic test caps with the locked thirty-task evidence floor, preserving fail-before-probe behavior.
- Fail closed after catalog discovery but before capability egress when the complete all-model probe and equal-budget evaluation plan cannot fit the configured hard request cap; the monthly 2,000-request ceiling covers the representative 127-model, thirty-task, seven-worker plan requiring 1,564 requests and still rejects larger plans before partial probing.
- Pin Atheris by Python interpreter so the Python 3.11 fuzz job and the newer central coverage-evidence image both install a published, hash-locked wheel.
- Record the reviewed current NVIDIA NIM General FAQ as expiring evidence for free Developer Program hosted-endpoint prototyping access, while keeping NVIDIA AI Enterprise production licensing and every hypothetical model rate explicitly separate.
- Require live hypothetical pricing scenarios to carry reviewed source, reviewer, review date, validity horizon, rate basis, uncertainty, and explicit rates; reject unreviewed, incomplete, future-dated, or expired price evidence before provider egress.
- Give direct, route-once, conduct, and reviewed cheapest-worker cells one equal total prompt-plus-completion token budget and one common five-call envelope, with configured-versus-observed evidence in every cell.
- Keep the optional NIM adapter lazy: importing the runtime package no longer imports the benchmark or mutates benchmark globals.
- Record immutable source-artifact digests and exact Git tree identity in the integration evidence so buyers and reviewers can reproduce the accepted benchmark source independently of transient workflow state.

### Documentation

- Add APA 7 doctoring for Python environment-marker semantics, Atheris artifact availability and hashes, and the supported-platform uncertainty boundary.
- Add APA 7 doctoring for Python environment-marker semantics, Atheris artifact availability and hashes, the NIM benchmark validity boundary, the thirty-task evidence floor, and supported-platform uncertainty.
- Record the CI trust boundary between generic coverage and native fuzz execution, including the evidence-preserving retry rule for branch-referenced reusable workflows.
- Make the NIM security-integration receipt head-stable: GitHub pull-request metadata and exact-head Checks are authoritative, while older commit and workflow identifiers remain explicitly historical only.
- Clarify that the bundled thirty-task NIM manifest may reach `evidence_review_required` only when the paired-task and 90% completion floors are met; otherwise it remains `insufficient_evidence`, and no artifact changes production routing automatically.
Loading
Loading