Skip to content

[WIP] [feat] A5 deep fused moe supported - #602

Open
666syh wants to merge 4 commits into
sgl-project:mainfrom
666syh:a5_fused_moe
Open

[WIP] [feat] A5 deep fused moe supported#602
666syh wants to merge 4 commits into
sgl-project:mainfrom
666syh:a5_fused_moe

Conversation

@666syh

@666syh 666syh commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

Motivation

  • Deep Fused MoE is now supported on Ascend 950. This operator fuses the Dispatch, expert computation (GMM), and Combine stages into a single large operator so that communication and computation can be overlapped, improving end-to-end MoE execution efficiency.
    This PR introduces the corresponding implementation updates in csrc/ and synchronizes the related examples.

Modifications

  • Host

    • Updated the host-side operator registration and launch flow for Deep Fused MoE.
    • Added the Ascend 950 fused execution path for Dispatch + GMM + Combine.
    • Aligned parameter preparation and invocation flow with the new fused operator behavior.
  • Kernel

    • Updated the Ascend 950 Deep Fused MoE kernel implementation.
    • Integrated Dispatch, GMM, and Combine into a unified execution flow.
    • Improved communication/computation overlap to reduce stage-by-stage synchronization overhead.
  • Tiling

    • Updated tiling calculation and configuration logic for the fused path.
    • Adjusted tile partitioning, buffer layout, and execution granularity for Ascend 950.
    • Ensured consistency between tiling outputs and kernel-side execution requirements.

Testing

  • Original Fused MoE operator based on Ascend 910C:
    [Rank 0] Difference between base and fused recv_count -> max: 0, mean: 0
    [Rank 1] Difference between base and fused recv_count -> max: 0, mean: 0
    [Rank 5] Difference between base and fused recv_count -> max: 0, mean: 0
    [Rank 2] Difference between base and fused recv_count -> max: 0, mean: 0
    [Rank 3] Difference between base and fused recv_count -> max: 0, mean: 0
    [Rank 7] Difference between base and fused recv_count -> max: 0, mean: 0
    [Rank 4] Difference between base and fused recv_count -> max: 0, mean: 0
    [Rank 6] Difference between base and fused recv_count -> max: 0, mean: 0
    [Rank 4] baseline_time= 267.56 us
    [Rank 4] fused_moe_time= 218.31 us
    [Rank 5] baseline_time= 268.56 us
    [Rank 5] fused_moe_time= 218.16 us
    [Rank 1] baseline_time= 268.03 us
    [Rank 1] fused_moe_time= 218.09 us
    [Rank 3] baseline_time= 268.43 us
    [Rank 3] fused_moe_time= 217.98 us
    [Rank 0] baseline_time= 268.18 us
    [Rank 0] fused_moe_time= 218.01 us
    [Rank 6] baseline_time= 268.16 us
    [Rank 6] fused_moe_time= 218.42 us
    [Rank 2] baseline_time= 267.80 us
    [Rank 2] fused_moe_time= 218.01 us
    [Rank 7] baseline_time= 268.22 us
    [Rank 7] fused_moe_time= 218.20 us
  • Fused MoE operator based on Ascend 950:
    | Rank | Local Experts | Fused Counts         | Sum |
    |-----:|--------------:|----------------------|----:|
    |    0 |             8 | [24, 24, 24, 24, 24, 24, 24, 24] | 192 |
    |    1 |             8 | [24, 24, 24, 24, 24, 24, 24, 24] | 192 |
    |    2 |             8 | [24, 24, 24, 24, 24, 24, 24, 24] | 192 |
    |    3 |             8 | [24, 24, 24, 24, 24, 24, 24, 24] | 192 |
    |    4 |             8 | [24, 24, 24, 24, 24, 24, 24, 24] | 192 |
    |    5 |             8 | [24, 24, 24, 24, 24, 24, 24, 24] | 192 |
    |    6 |             8 | [24, 24, 24, 24, 24, 24, 24, 24] | 192 |
    |    7 |             8 | [24, 24, 24, 24, 24, 24, 24, 24] | 192 |
    Accuracy check passed. avg_diff=0.000000, max_diff=0.000000, calc_diff=0.000000
    Profiled NPU op time:
    small-op breakdown (mean over ranks):
    | Stage        |    Avg (us) |    Min (us) |    Max (us) |  Share % |
    |--------------|-------------:|-------------:|-------------:|---------:|
    | dispatch     |        18.48 |        17.29 |        19.96 |    4.43% |
    | gmm1         |       236.29 |       233.34 |       238.96 |   56.63% |
    | swiglu       |        10.96 |        10.46 |        11.67 |    2.63% |
    | requant      |         6.26 |         5.65 |         7.04 |    1.50% |
    | gmm2         |       121.90 |       120.52 |       123.68 |   29.22% |
    | combine      |        23.35 |        20.39 |        26.99 |    5.60% |
    | total        |       417.24 |       414.34 |       421.63 |  100.00% |
    fused buffer path (mean over ranks):
    | Stage        |    Avg (us) |    Min (us) |    Max (us) |  Share % |
    |--------------|-------------:|-------------:|-------------:|---------:|
    | fused        |       401.67 |       389.58 |       573.27 |  100.00% |
    speedup=1.0388x, delta_pct=-3.73%
    rank skew summary:
    | Rank | Dispatch (us) | GMM1 (us) | SwiGLU (us) | Requant (us) | GMM2 (us) | Combine (us) | Small Total (us) | Fused (us) |
    |-----:|--------------:|----------:|------------:|-------------:|----------:|-------------:|-----------------:|-----------:|
    |    0 |         18.68 |    235.81 |       10.92 |         6.29 |    121.91 |        23.61 |           417.22 |     401.65 |
    |    1 |         18.20 |    236.05 |       10.92 |         6.35 |    121.89 |        23.88 |           417.28 |     401.68 |
    |    2 |         18.52 |    236.92 |       10.83 |         6.18 |    122.14 |        22.61 |           417.20 |     401.69 |
    |    3 |         18.50 |    236.09 |       10.99 |         6.34 |    121.81 |        23.64 |           417.37 |     401.70 |
    |    4 |         18.23 |    236.32 |       10.77 |         6.17 |    121.68 |        24.01 |           417.19 |     401.70 |
    |    5 |         18.69 |    236.44 |       11.01 |         6.37 |    122.12 |        22.51 |           417.14 |     401.56 |
    |    6 |         18.03 |    236.45 |       11.20 |         6.16 |    121.85 |        23.51 |           417.20 |     401.68 |
    |    7 |         18.97 |    236.25 |       11.04 |         6.21 |    121.83 |        23.02 |           417.32 |     401.68 |
    | mean |         18.48 |    236.29 |       10.96 |         6.26 |    121.90 |        23.35 |           417.24 |     401.67 |

Benchmarking and Profiling

  • Use single operators to simulate framework-side model invocation.
    • premise:
      • FP8 quant
      • no share experts
      • Topk distribution is uniform.
    • case:
      Dimension value
      EP 8
      hidden 7168
      moe-intermediate-size 3072
      topk 6
    • performance:
      num-tokens-per-device num-experts-per-device Dispatch + GMM + Combine (us) Fused Moe (us) improvement
      32 6 323.59 302.68 6.4%
      64 6 366.38 320.99 12.4%
      96 6 431.81 341.94 20.8%
      128 6 493.37 360.92 26.8%

Checklist

  • Format your code.
  • Add unit tests. (N/A for docs)
  • Update documentation.
  • Provide accuracy and speed benchmark results. (N/A for docs)

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@666syh 666syh changed the title [feat] A5 deep fused moe supported [WIP] [feat] A5 deep fused moe supported Jul 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants