Hi, thanks for sharing the detailed evaluation. I am interested in whether larger-EP colocated configurations were considered for the MiniMax-M2.5 comparison.
My understanding from the evaluation setup is that the colocated vLLM baseline uses one 4-GPU node with DP=4 and EP=4, while FastAFD's best MiniMax-M2.5 result uses 18 nodes at a 17:1 attention-to-FFN ratio (68 attention GPUs and 4 FFN GPUs, with FFN EP=4).
Did you consider comparing FastAFD with a colocated baseline using a larger EP degree at the same deployment scale—for example, EP=72 on the full NVL72 rack, or the largest EP degree supported by the model and runtime?
I am curious whether FastAFD would still retain its per-GPU throughput advantage over a rack-scale colocated layout with wider EP.
I may be missing a constraint or trade-off in the larger-EP setup, and I would appreciate any clarification or measurements you can share.
Hi, thanks for sharing the detailed evaluation. I am interested in whether larger-EP colocated configurations were considered for the MiniMax-M2.5 comparison.
My understanding from the evaluation setup is that the colocated vLLM baseline uses one 4-GPU node with DP=4 and EP=4, while FastAFD's best MiniMax-M2.5 result uses 18 nodes at a 17:1 attention-to-FFN ratio (68 attention GPUs and 4 FFN GPUs, with FFN EP=4).
Did you consider comparing FastAFD with a colocated baseline using a larger EP degree at the same deployment scale—for example, EP=72 on the full NVL72 rack, or the largest EP degree supported by the model and runtime?
I am curious whether FastAFD would still retain its per-GPU throughput advantage over a rack-scale colocated layout with wider EP.
I may be missing a constraint or trade-off in the larger-EP setup, and I would appreciate any clarification or measurements you can share.