Skip to content

Flash attention makes inference dramatically slower (v1.4.0) #118

Description

@Onkitova

Is it just my setup or "flash attention" option is not working properly in version 1.4.0?

Image Image

CPU inference is also slower, but not that x-times worse. Tried both: quantized and non-quantized models. I have heard that Vulkan's support of flash attention is not mature enough -- might it be related to this?

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions