跳转至

GEMM, MoE & Quantization

本页暂无中文版。以下为英文原文。

21 ops, 152 workloads — Linear Algebra (GEMM) 10 · Mixture of Experts 9 · Quantization 2.

One table per op, one row per workload. Ratio is the fastest other implementation's device time divided by ours, so green is faster than it, plain is level with it, red is slower. Times are in ms. How these numbers are taken.

Linear Algebra (GEMM)

BmmFp8NKFwd (5 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
mha-decode-b32-pv-per-tensor-float8_e4m3fn1.05×0.00902flashinfer-bmm-fp8
torch-fp32-ref
0.0095
0.2365
238
mha-decode-b64-qk-per-tensor-float8_e4m3fn1.37×0.0157flashinfer-bmm-fp8
torch-fp32-ref
0.0216
0.2880
273
moe-prefill-b128-per-tensor-float8_e4m3fn1.05×0.1319flashinfer-bmm-fp8
torch-fp32-ref
0.1384
5.4095
1,042
square-b4-1k-per-tensor-float8_e4m3fn1.10×0.0119flashinfer-bmm-fp8
torch-fp32-ref
0.0131
0.2935
722
square-b8-2k-per-tensor-float8_e4m3fn1.06×0.1196flashinfer-bmm-fp8
torch-fp32-ref
0.1262
4.0650
1,150

grouped_gemm_nt (2 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
nt-batch16-m4096-n4096-k4096-bfloat160.99×0.2266torch
torch-compile
torch-ref
0.2250
2.2494
2.2824
606
nt-batch16-m4096-n4096-k4096-float161.15×0.2327torch
torch-compile
torch-ref
0.2682
2.2945
2.3274
591

grouped_gemm_3wg_baselines (12 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
GLM-5-744B-down-T=1310721.15×38.4690triton-tma
deepgemm
torch
triton
44.1777
44.2571
45.2484
57.7138
686
GLM-5-744B-down-T=2621441.00×77.8555deepgemm
torch
triton-tma
triton
77.9715
82.1943
84.3373
115.0890
678
GLM-5-744B-down-T=327681.00×9.6454deepgemm
torch
triton-tma
triton
9.6108
9.9411
11.4355
14.6883
684
GLM-5-744B-down-T=655360.98×19.1609deepgemm
triton-tma
torch
triton
18.7690
20.9368
24.6188
28.9118
689
GLM-5-744B-up-T=1310721.00×74.8445deepgemm
torch
triton-tma
triton
75.0057
80.4625
84.9312
106.2630
705
GLM-5-744B-up-T=2621441.00×151.3490deepgemm
triton-tma
torch
triton
151.8770
163.4910
170.0550
216.2940
697
GLM-5-744B-up-T=327681.02×18.1814deepgemm
torch
triton-tma
triton
18.4979
23.0794
23.1377
27.6289
726
GLM-5-744B-up-T=655361.06×37.5122torch
deepgemm
triton-tma
triton
39.6492
41.1302
41.2708
54.6895
703
Llama4-128E-down-T=1310721.01×15.2954deepgemm
torch
triton-tma
triton
15.5135
15.8305
18.8759
23.0366
719
Llama4-128E-up-T=1310720.99×31.0676deepgemm
torch
triton-tma
triton
30.8841
32.1667
40.8244
52.2075
708
qwen3.5-397B-down-T524290.95×6.9651deepgemm
torch
triton-tma
triton
6.6147
7.3609
9.0427
10.2177
631
qwen3.5-397B-up-T524291.00×12.6930deepgemm
torch
triton-tma
triton
12.6909
13.5139
18.4964
19.6010
693

BmmFwd (16 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
mha-decode-b64-pv-bfloat161.02×0.0239torch-cublas
flaggems
0.0244
0.0407
179
mha-decode-b64-pv-float161.02×0.0239torch-cublas
flaggems
0.0244
0.0407
179
mha-decode-b64-qk-bfloat160.94×0.0225torch-cublas
flaggems
0.0212
0.0260
191
mha-decode-b64-qk-float160.94×0.0225torch-cublas
flaggems
0.0212
0.0260
191
moe-prefill-b128-bfloat160.74×0.2911torch-cublas
flaggems
0.2152
0.2952
472
small-b8-128-bfloat161.16×0.00272flaggems
torch-cublas
0.00317
0.00323
12.3
small-b8-128-float161.16×0.00272flaggems
torch-cublas
0.00317
0.00323
12.3
square-b16-512-bfloat160.90×0.0133torch-cublas
flaggems
0.0120
0.0151
323
square-b16-512-float160.91×0.0132torch-cublas
flaggems
0.0121
0.0151
324
square-b32-256-bfloat161.07×0.00656torch-cublas
flaggems
0.00704
0.0079
164
square-b32-256-float161.07×0.00659torch-cublas
flaggems
0.00707
0.0079
163
square-b4-4k-bfloat160.74×1.0435torch-cublas
flaggems
0.7705
0.9716
527
square-b8-1k-bfloat160.76×0.0407torch-cublas
flaggems
0.0310
0.0448
422
square-b8-1k-float160.77×0.0406torch-cublas
flaggems
0.0311
0.0449
423
square-b8-2k-bfloat160.73×0.2802torch-cublas
flaggems
0.2045
0.2735
491
square-b8-2k-float160.74×0.2832torch-cublas
flaggems
0.2089
0.2760
485

grouped_gemm_nn (1 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
nn-batch16-m4096-n4096-k4096-float160.79×0.3407torch
torch-compile
torch-ref
0.2696
0.2750
0.3074
403

BmmFp8KNFwd (5 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
mha-decode-b32-pv-per-tensor-float8_e4m3fn0.38×0.0647flashinfer-bmm-fp8
torch-fp32-ref
0.0249
0.2360
33.2
mha-decode-b64-qk-per-tensor-float8_e4m3fn0.43×0.1154flashinfer-bmm-fp8
torch-fp32-ref
0.0497
0.2888
37.2
moe-prefill-b128-per-tensor-float8_e4m3fn0.69×0.9006flashinfer-bmm-fp8
torch-fp32-ref
0.6245
5.3998
153
square-b4-1k-per-tensor-float8_e4m3fn0.91×0.0390flashinfer-bmm-fp8
torch-fp32-ref
0.0355
0.2940
220
square-b8-2k-per-tensor-float8_e4m3fn2.03×0.3060flashinfer-bmm-fp8
torch-fp32-ref
0.6227
4.0584
449

GemmW4A16Fwd (6 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
compile-smoke-rect-128x256x256-float160.52×0.00586torch-dequantized-matmul0.003072.87
compile-smoke-square-64x64x128-float160.63×0.00429torch-dequantized-matmul0.002690.245
decode-hbm-streaming-threshold-float160.62×0.0606marlin-fp32
marlin-fp16
torch-dequantized-matmul
0.0379
0.0380
0.0748
4.43
decode-l2-resident-ish-float160.66×0.0330marlin-fp16
marlin-fp32
torch-dequantized-matmul
0.0217
0.0219
0.0468
4.06
decode-long-k-pressure-float160.50×0.2834marlin-fp16
marlin-fp32
torch-dequantized-matmul
0.1408
0.1414
0.3227
4.74
decode-non-power2-low-cta-float160.55×0.0745marlin-fp16
marlin-fp32
torch-dequantized-matmul
0.0406
0.0407
0.0876
3.94

grouped_gemm_tn (1 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
tn-batch16-m4096-n4096-k4096-float160.45×0.7820torch
torch-compile
torch-ref
0.3525
0.5222
0.5251
176

GemmFwd (14 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
ds-v3-decode-down-bfloat160.54×0.0246torch-cublas
deepgemm
0.0132
0.0138
153
ds-v3-decode-gate-up-bfloat160.25×0.0678torch-cublas
deepgemm
0.0172
0.0213
57.2
ds-v3-prefill-attn-proj-bfloat160.61×0.5398deepgemm
torch-cublas
0.3301
0.3314
446
ds-v3-prefill-attn-proj-float160.62×0.5448torch-cublas0.3368441
ds-v3-prefill-down-bfloat160.56×0.3214deepgemm
torch-cublas
0.1798
0.1804
374
ds-v3-prefill-gate-up-bfloat160.52×0.3370torch-cublas
deepgemm
0.1766
0.1810
368
k-dominant-7168x16384-bfloat160.61×2.0598deepgemm
torch-cublas
1.2622
1.2683
467
mid-m16-attn-bfloat160.37×0.0658torch-cublas
deepgemm
0.0245
0.0339
14.3
mid-m32-attn-bfloat160.37×0.0662torch-cublas
deepgemm
0.0243
0.0305
28.4
mid-m64-down-bfloat160.64×0.0207torch-cublas
deepgemm
0.0132
0.0135
90.9
mid-m96-gate-up-bfloat160.25×0.0687torch-cublas
deepgemm
0.0169
0.0220
42.3
square-1k-nn-bfloat160.50×0.0145torch-cublas
flaggems
0.0072
0.0115
148
square-1k-nn-float160.50×0.0145torch-cublas
flaggems
0.00726
0.0118
148
wide-n-24576-bfloat160.49×0.9017deepgemm
torch-cublas
0.4419
0.4524
343

GemmFp8Fwd (15 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
ds-v3-decode-down-block128-float8_e4m3fn0.25×0.0377flashinfer-fp8-blockscale-sm90
torch-scaled-mm
0.00925
0.2793
99.6
ds-v3-decode-down-per-tensor-float8_e4m3fn0.41×0.0253deepgemm
torch-scaled-mm
0.0104
0.2448
148
ds-v3-decode-gate-up-block128-float8_e4m3fn0.09×0.1481flashinfer-fp8-blockscale-sm90
torch-scaled-mm
0.0129
0.2794
26.2
ds-v3-decode-gate-up-per-tensor-bias-float8_e4m3fn2.13×0.1166torch-scaled-mm0.248133.2
ds-v3-decode-gate-up-per-tensor-float8_e4m3fn2.09×0.1162torch-scaled-mm0.242333.4
ds-v3-prefill-attn-proj-block128-float8_e4m3fn0.28×0.7704flashinfer-fp8-blockscale-sm90
torch-scaled-mm
0.2147
6.6789
312
ds-v3-prefill-down-block128-float8_e4m3fn0.32×0.4457flashinfer-fp8-blockscale-sm90
torch-scaled-mm
0.1431
3.3979
270
ds-v3-prefill-down-per-tensor-float8_e4m3fn0.50×0.2109deepgemm
torch-scaled-mm
0.1056
3.3435
570
ds-v3-prefill-gate-up-block128-float8_e4m3fn0.36×0.3870flashinfer-fp8-blockscale-sm90
torch-scaled-mm
0.1408
3.5322
320
ds-v3-prefill-gate-up-per-tensor-float8_e4m3fn6.72×0.5108torch-scaled-mm3.4349243
gemv-down-m1-block128-float8_e4m3fn0.17×0.0446flashinfer-fp8-blockscale-sm90
torch-scaled-mm
0.00778
0.1644
0.658
gemv-down-m1-per-tensor-float8_e4m3fn0.39×0.0258deepgemm
torch-scaled-mm
0.0101
0.1304
1.14
k-dominant-7168x16384-block128-float8_e4m3fn0.22×3.5879flashinfer-fp8-blockscale-sm90
torch-scaled-mm
0.7726
26.4188
268
small-batch-down-m8-per-tensor-float8_e4m3fn0.31×0.0267deepgemm
torch-scaled-mm
0.00829
0.1666
8.81
wide-n-24576-block128-float8_e4m3fn0.37×1.0266flashinfer-fp8-blockscale-sm90
torch-scaled-mm
0.3822
8.5216
301

Mixture of Experts

MoeGateUpFwd (2 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
deepseek-v3-decode-gate-up-bfloat162.27×3.4594torch-ref
torch-compile
6.6366
7.8601
69.5
deepseek-v3-prefill-gate-up-bfloat166.14×4.4089torch-ref
torch-compile
6.9703
27.0804
436

MoeGroupedGemmNopadFwd (4 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
deepseek-v3-decode-down-bfloat162.92×1.9105torch-ref
torch-compile
2.6916
5.5871
62.9
deepseek-v3-decode-gate-up-bfloat161.56×3.7451torch-ref
torch-compile
5.1691
5.8544
64.2
deepseek-v3-prefill-down-bfloat1612.00×2.1523torch-ref
torch-compile
2.8422
25.8232
447
deepseek-v3-prefill-gate-up-bfloat162.51×4.2967torch-ref
torch-compile
5.4109
10.8057
448

MoePermuteAlignFwd (12 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
deepseek-v3-decode-int322.29×0.0148triton0.0337·
deepseek-v3-medium-int322.36×0.0177triton0.0419·
deepseek-v3-prefill-int321.97×0.0378triton0.0743·
deepseek-v3-small-int322.20×0.0153triton0.0337·
kimi-k2-decode-int322.87×0.0169triton0.0484·
kimi-k2-medium-int322.58×0.0217triton0.0559·
kimi-k2-prefill-int322.09×0.0410triton0.0856·
kimi-k2-small-int322.47×0.0195triton0.0481·
qwen3-decode-int321.57×0.0108triton0.0169·
qwen3-medium-int322.12×0.0141triton0.0298·
qwen3-prefill-int322.51×0.0318triton0.0800·
qwen3-small-int321.50×0.0121triton0.0181·

MoeUnpermuteFwd (8 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
large-hidden-decode-bfloat162.39×0.00698vllm0.01660.0164
large-hidden-medium-bfloat161.37×0.0214vllm0.02932.75
large-hidden-prefill-bfloat161.05×0.1329vllm0.13913.53
large-hidden-small-bfloat162.27×0.00787vllm0.01790.466
small-hidden-decode-bfloat161.57×0.0057vllm0.008930.00863
small-hidden-medium-bfloat161.28×0.0116vllm0.01482.18
small-hidden-prefill-bfloat161.10×0.0615vllm0.06753.27
small-hidden-small-bfloat161.52×0.00653vllm0.009940.241

MoePermuteNopadFwd (20 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
deepseek-v3-decode-bfloat161.25×0.00925vllm0.0116·
deepseek-v3-ep2-decode-bfloat160.0087···
deepseek-v3-ep2-medium-bfloat160.0280···
deepseek-v3-ep2-prefill-bfloat160.2101···
deepseek-v3-medium-bfloat161.37×0.0337vllm0.0460·
deepseek-v3-prefill-bfloat160.97×0.2789vllm0.2700·
deepseek-v3-small-bfloat161.32×0.0104vllm0.0138·
kimi-k2-decode-bfloat161.10×0.0106vllm0.0117·
kimi-k2-medium-bfloat161.29×0.0356vllm0.0460·
kimi-k2-prefill-bfloat160.95×0.2855vllm0.2709·
kimi-k2-small-bfloat161.17×0.0118vllm0.0139·
qwen3-235b-decode-bfloat161.43×0.00803vllm0.0115·
qwen3-235b-ep2-medium-bfloat160.0264···
qwen3-235b-medium-bfloat161.47×0.0313vllm0.0461·
qwen3-235b-prefill-bfloat160.97×0.2686vllm0.2612·
qwen3-235b-small-bfloat161.53×0.00902vllm0.0138·
qwen3-30b-decode-bfloat161.67×0.0063vllm0.0105·
qwen3-30b-medium-bfloat161.40×0.0207vllm0.0290·
qwen3-30b-prefill-bfloat160.91×0.1419vllm0.1295·
qwen3-30b-small-bfloat161.74×0.0072vllm0.0125·

FusedTopK (8 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
1-128-8-softmax-norenormalize1.42×0.00426vllm0.006050.000541
1-384-8-sigmoid-renormalize1.00×0.00829vllm0.008290.000834
32-128-8-softmax-norenormalize1.12×0.00739vllm0.008290.00997
32-384-8-sigmoid-renormalize0.81×0.0119vllm0.00970.0186
4096-128-8-softmax-norenormalize1.47×0.0110vllm0.01610.86
4096-384-8-sigmoid-renormalize1.17×0.0203vllm0.02381.4
512-128-8-softmax-norenormalize1.16×0.00781vllm0.009020.151
512-384-8-sigmoid-renormalize0.83×0.0126vllm0.01050.281

FusedMoeFwd (6 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
deepseek-v3-decode-bfloat161.02×5.4253vllm5.519166.5
deepseek-v3-prefill-bfloat161.06×8.3550vllm8.8652345
kimi-k2-decode-bfloat1614.56×3.8889torch-ref56.635792.8
kimi-k2-prefill-bfloat1617.88×7.8753torch-ref140.8010366
qwen3-235b-decode-bfloat161.03×2.7727vllm2.8575130
qwen3-235b-prefill-bfloat161.20×6.0667vllm7.2789476

FusedMoEExpertsNopadPersistent3WGFwd (6 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
deepseek-v3-decode-bfloat161.02×5.4159vllm-triton5.512066.6
deepseek-v3-ep2-decode-bfloat162.7228··133
deepseek-v3-ep2-prefill-bfloat164.1410··697
deepseek-v3-prefill-bfloat161.04×8.4890vllm-triton8.8313340
qwen3-235b-decode-bfloat161.03×2.7726vllm-triton2.8545130
qwen3-235b-prefill-bfloat161.20×6.0374vllm-triton7.2463478

SharedFusedMoE (5 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
1-384-8-7168-2048-18432-sigmoid-renormalize-correctionbias-2.827-bfloat160.17×2.5280vllm0.42620.592
2048-384-8-7168-2048-18432-sigmoid-renormalize-correctionbias-2.827-bfloat160.59×19.5743vllm11.6008157
32-384-8-7168-2048-18432-sigmoid-renormalize-correctionbias-2.827-bfloat160.84×4.7455vllm3.964710.1
4096-384-8-7168-2048-18432-sigmoid-renormalize-correctionbias-2.827-bfloat160.45×32.5463vllm14.6270188
512-384-8-7168-2048-18432-sigmoid-renormalize-correctionbias-2.827-bfloat161.09×8.0566vllm8.769795.2

Quantization

FP8LightningIndexerFwd (1 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
lightning-indexer-s8k-h32-d64-bfloat1680.40×0.6174torch-compile
torch-ref
49.6420
111.9970
55.6

FP8QuantFwd (3 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
kv-index-4k-d128-float321.06×0.00397torch-compile
torch-ref
0.00422
0.0155
0.794
kv-index-8k-d64-bfloat160.99×0.00275torch-compile
torch-ref
0.00272
0.0169
1.15
kv-index-8k-d64-float162.43×0.00275torch-compile
torch-ref
0.00669
0.0167
1.15