Is it just my setup or "flash attention" option is not working properly in version 1.4.0?
CPU inference is also slower, but not that x-times worse. Tried both: quantized and non-quantized models. I have heard that Vulkan's support of flash attention is not mature enough -- might it be related to this?
Is it just my setup or "flash attention" option is not working properly in version 1.4.0?
CPU inference is also slower, but not that x-times worse. Tried both: quantized and non-quantized models. I have heard that Vulkan's support of flash attention is not mature enough -- might it be related to this?