Skip to content

release: 0.2.2 — corrected grain-ladder measurements - #54

Merged
lusoris merged 1 commit into
masterfrom
release/v0.2.2
Aug 30, 2026
Merged

release: 0.2.2 — corrected grain-ladder measurements#54
lusoris merged 1 commit into
masterfrom
release/v0.2.2

Conversation

@lusoris

@lusoris lusoris commented Aug 30, 2026

Copy link
Copy Markdown
Contributor

v0.2.1 shipped the grain ladder built on an 86-title, single-scene probe. That sample was too small to be trustworthy and several of its published figures were wrong, so the released documentation was actively misleading. This cuts 0.2.2 so the released docs carry the corrected measurement.

What 0.2.2 corrects

Re-measured at 752 titles across 2230 scenes, three scenes per title:

bucket v0.2.1 shipped 0.2.2 error
2010s TV 0.00822 0.00462 +78%
1960s film 0.00823 0.00486 +69%
2010s film 0.00649 0.00455 +43%
2020s film 0.00215 0.00328 -34%
overall median 0.00666 0.00484 +38%

Two conclusions v0.2.1 stated were false: the 1960s are not the grainiest decade (mid-pack; the 1970s-80s peak), and TV does not stay grainier than film into the 2010s (they converge).

Root cause, measured: within-title scene-to-scene spread has a median of 0.71x the title's own median and p90 of 2.02x, so a one-scene probe misestimates by ~36% typically and 100%+ for a tenth of titles.

New finding, invisible at the old sample

1080p masters are 2.2x grainier than 4K masters (median 0.00724 vs 0.00323, like-for-like native crops). 4K releases are denoised hard enough to compress, so the impairment has already been removed. Bench material should be 1080p.

Scope

Documentation and bench methodology only — no library or filter code change. The ladder table and the netflix-bar / --grain <= 16 recommendation are unchanged, since those were measured directly rather than sampled.

Gates

21/21 fast suite, clang-format clean, changelog render in sync, meson version-skew assertion passes.

🤖 Generated with Claude Code

v0.2.1 shipped the grain ladder built on an 86-title, single-scene probe. That
sample was too small to be trustworthy and several of its published figures were
wrong, so the released documentation was misleading.

0.2.2 carries the re-measurement: 752 titles across 2230 scenes, three scenes per
title. Individual decade buckets moved by up to 78% and the overall median by 38%;
two conclusions were false (the 1960s are not the grainiest decade, and TV does not
stay grainier than film into the 2010s). It also surfaces the largest effect in the
data, invisible at the smaller sample: 1080p masters are 2.2x grainier than 4K
masters, which is where bench material should come from.

No library or filter code change -- documentation and bench methodology only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@lusoris
lusoris merged commit 93bef12 into master Aug 30, 2026
2 of 4 checks passed
@lusoris
lusoris deleted the release/v0.2.2 branch August 30, 2026 20:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant