A light llama-like llm inference framework based on the triton kernel.
python3 attention continues llm llm-inference llama3 continuous-batching flash-attention-3 triton-kernels qwen3-moe qwen3-vl tensor-parallel
-
Updated
Aug 31, 2026 - Python