MUG-U is a Multimodal Large Language Model (MLLM) that supports text, image, and video inputs, enabling powerful understanding, reasoning, and generation capabilities. Developed by the Shopee MUG team, MUG-U leverages the latest advancements and cutting-edge technologies in multimodal modeling.
2024.02.15 We released the MUG-U API.
| Model | Date | API | Note |
|---|---|---|---|
| MUG-U-7B | 2025.02.06 | infer | Qwen2.5-7B |
python infer_api.py| Benchmark | MUG-U-7B |
|---|---|
| MMBench-V1.1test | 81.8 |
| MMStar | 66.6 |
| MMMUval | 54.3 |
| MathVistatestmini | 74.8 |
| HallBenchavg | 51.3 |
| AI2Dtest | 88.9 |
| OCRBench | 91.1 |
| MMVet | 63 |
| Average | 71.5 |