High-Performance Rendering Framework on Stream Architectures
-
Updated
Aug 19, 2026 - C++
High-Performance Rendering Framework on Stream Architectures
Ascend NPU fork of nanochat for LLM training with torch_npu/HCCL (experimental)
NPU MFU Analyzer - 大模型训练性能分析工具 | LLM Training Performance Analyzer for Ascend NPU
End-to-end edge AI scene recognition on Huawei Ascend 310: PyTorch, ONNX, C++ inference, V4L2, and MJPEG streaming.
Run nanochat training efficiently on Huawei Ascend NPUs with minimal code changes, supporting tokenizer, pretraining, and evaluation workflows.
Native AscendC Mamba2 selective scan / SSD forward-backward custom operator for Huawei Ascend 910B3 and 950PR, with CANN, torch_npu, A100 benchmarks and msprof profiling.
Industrial defect detection on Huawei Ascend NPU with YOLO, ONNX-to-OM conversion, and C/C++/Python inference.
驱动只有网页终端、在WAF后的竞赛/云pod:反向SSH隧道+命令通道,从自己机器全自动发命令拿结果(含踩坑与安全加固)
Deploy GLM-4.6V-Flash (9B dense VLM) on Huawei Ascend 910B NPU with vLLM - multimodal, OpenAI API, single/dual-card serving, reproducible benchmarks.
Add a description, image, and links to the huawei-ascend topic page so that developers can more easily learn about it.
To associate your repository with the huawei-ascend topic, visit your repo's landing page and select "manage topics."