Skip to content

Recurrent GDN/KDA for decoding phase #171

Description

@learning-chip

https://github.com/huawei-csl/megagdn-pto only does chunk GDN for prefill. sgl-kernel-npu/fused_sigmoid_gating_recurrent.py has Triton baseline for both GDN and KDA decoding.

This is an easy pure-vector memory-bound kernel, so not expecting much room for performance improvements. Just for feature completeness.

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions