Skip to content

feat: add hwy.Tile[T] 2D matrix accumulator with multi-target codegen - #53

Open
ajroetker wants to merge 1 commit into
mainfrom
feature/2d-tile-primitives.md
Open

feat: add hwy.Tile[T] 2D matrix accumulator with multi-target codegen#53
ajroetker wants to merge 1 commit into
mainfrom
feature/2d-tile-primitives.md

Conversation

@ajroetker

Copy link
Copy Markdown
Owner

Add a tile-aware abstraction layer for 2D matrix operations that lowers to target-specific code across all supported backends:

  • Scalar fallback: hwy.Tile[T] with nested loop implementations
  • NEON: TileFloat32x4/Float64x2 using vfmaq_laneq lane-broadcast FMA
  • AVX2: TileFloat32x8/Float64x4 using broadcast+FMA on archsimd vectors
  • AVX512: TileFloat32x16/Float64x8 using broadcast+FMA on archsimd vectors
  • C/:asm targets: AST translator emits NEON/SVE intrinsics with tile structs

hwygen transforms tile operations (TileZero, OuterProductAdd/Sub, TileReadRow, TileStoreRow, TileLoadCol, NewTile, TileDim) using the existing IsMethod pattern for SIMD targets and direct hwy.* calls for fallback.

Add a tile-aware abstraction layer for 2D matrix operations that lowers
to target-specific code across all supported backends:

- Scalar fallback: hwy.Tile[T] with nested loop implementations
- NEON: TileFloat32x4/Float64x2 using vfmaq_laneq lane-broadcast FMA
- AVX2: TileFloat32x8/Float64x4 using broadcast+FMA on archsimd vectors
- AVX512: TileFloat32x16/Float64x8 using broadcast+FMA on archsimd vectors
- C/:asm targets: AST translator emits NEON/SVE intrinsics with tile structs

hwygen transforms tile operations (TileZero, OuterProductAdd/Sub,
TileReadRow, TileStoreRow, TileLoadCol, NewTile, TileDim) using the
existing IsMethod pattern for SIMD targets and direct hwy.* calls for
fallback.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant