Skip to content

docs: Qwen3.8-27B local-model evaluation (EN/TR/DE) - #32

Merged
pretyflaco merged 1 commit into
mainfrom
docs/qwen3.8-local-eval
Aug 15, 2026
Merged

docs: Qwen3.8-27B local-model evaluation (EN/TR/DE)#32
pretyflaco merged 1 commit into
mainfrom
docs/qwen3.8-local-eval

Conversation

@pretyflaco

Copy link
Copy Markdown
Owner

Refreshes docs/local-model-evaluation.md with the Qwen3.8-27B evaluation, run through millet's real two-pass code path (not a harness).

Headline

With the v0.15.1 Pass-1 truncation fix (#31), qwen3.8:27b matches or exceeds the cloud baseline on topic/action/question coverage across English, Turkish, and German meetings — 5/5 format, localized non-English headers, no hallucinated speakers, ~55-100s/meeting on an RTX 3090. It supersedes the earlier 'gpt-oss:20b is best local' conclusion.

Results (via millet, v0.15.1)

Meeting Lang Qwen3.8 (T/A/D/Q) Baseline (T/A/D) Format
E EN 20/8/4/7 11/6/3 5/5
F EN 10/24/5/10 18/20/4 5/5
G EN 20/16/6/10 10/12/4 5/5
H TR 10/1/3/7 6/1/0 5/5
I DE 15/4/3/8 15/3/0 5/5

What changed in the doc

  • New 'Qwen3.8-27B Evaluation (2026-08-15)' section + appendix entry.
  • Reconciled stale claims: default model (qwen3.5:9b + MILLET_SUMMARY_MODEL, not 'changed to gpt-oss'); gpt-oss demoted to fallback; Sonnet 'only model' line softened; Phase-1 marked shipped.
  • Documented the v0.15.1 truncation fix as a prerequisite, and the known 'decisions over-inclusion' weakness.

Caveats (stated in-doc)

  • ~2× Sonnet latency; mild over-count of 'decisions'; non-English sample still small (1 TR + 1 DE).

Docs-only; no release. Metrics only — no transcript/meeting content (meetings anonymized E-I).

…dations

Records the Qwen3.8-27B evaluation run through millet's real two-pass code
path (not a harness): matches/exceeds the cloud baseline on topic/action/
question coverage across English, Turkish, and German meetings, 5/5 format,
localized non-English headers, no hallucinated speakers, ~55-100s/meeting.

Requires the v0.15.1 Pass-1 output-reserve fix or it silently truncates.
Supersedes the earlier 'gpt-oss:20b is best local' conclusion; reconciles
the stale default-model and 'Sonnet is the only model' claims. Metrics only
- no transcript content.
@pretyflaco
pretyflaco merged commit f4a4d8a into main Aug 15, 2026
3 checks passed
@pretyflaco
pretyflaco deleted the docs/qwen3.8-local-eval branch August 15, 2026 10:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant