fix: safely decode ARM64 StreamVByte tails - #2
Merged
ajroetker merged 2 commits intoJul 24, 2026
Conversation
… groups The NEON batch decoder requires a full 16-byte load window per group. Groups without one take unreliable tail paths: on linux/arm64 the kernel zero-fills their output while still reporting full decoded/dataConsumed success, so the corruption is undetectable from the return values. Route all decodes through decodeStreamVByteFull, which hands the kernel only the provably safe group prefix, scalar-decodes the remainder, and redoes everything scalar if the kernel deviates from the plan. This removes the need for HWY_NO_SIMD=1 as a correctness mitigation.
Red on the parent commit's raw kernel calls (silent zeroed tails, the TestOpen/TestStreamVByte* arm64 failures); green with the safe-prefix helper under both NEON and HWY_NO_SIMD=1. Includes a truncation case that cuts into real value bytes past the final group's zero padding.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Root cause
The ARM64 batch decoder requires a full 16-byte load window for each four-value group. For short final groups, it could zero-fill output while still reporting complete decoded and consumed counts. Zapx therefore had no signal that location columns were corrupted.
This helper computes the safe SIMD prefix from the control bytes and completes the tail with the scalar decoder. It retains SIMD for the valid prefix and removes the need for HWY_NO_SIMD=1 as a correctness mitigation.
Validation
Related incident: antflydb/antfly#381