Skip to content

Attention

15 ops, 119 workloads.

One table per op, one row per workload. Ratio is the fastest other implementation's device time divided by ours, so green is faster than it, plain is level with it, red is slower. Times are in ms. How these numbers are taken.

Attention

DeepSeekSparseAttentionDecodeWithKVCacheFwd (2 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
longer-kv-lower-topk-float163.62×0.5005torch-gather
torch-sdpa
torch-compile
torch-ref
1.8136
5.2718
16.2783
19.3187
292
single-batch-mainstream-float162.83×1.8607torch-sdpa
torch-compile
torch-ref
5.2580
16.6112
18.9702
314

MultiHeadLatentAttentionDecodeWithKVCacheFwd (6 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
deepseek-v2-32k-bfloat162.15×0.1188torch-compile
torch-ref
0.2558
0.2779
181
deepseek-v2-32k-float162.12×0.1188torch-compile
torch-ref
0.2520
0.2742
181
deepseek-v2-4k-bfloat163.50×0.0375torch-compile
torch-ref
0.1313
0.1639
287
deepseek-v2-4k-float163.42×0.0374torch-compile
torch-ref
0.1280
0.1646
287
deepseek-v3-32k-bfloat161.39×0.1179torch-compile
torch-ref
0.1635
0.1711
91.1
deepseek-v3-4k-bfloat163.24×0.0216torch-compile
torch-ref
0.0700
0.0849
248

MultiHeadAttentionDecodePagedWithKVCacheFwd (4 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
batch2-page256-float161.72×0.00568flashinfer
fa3
0.00976
0.0185
0.738
longer-cache-float161.81×0.00533flashinfer
fa3
0.00966
0.0182
0.394
shorter-cache-float162.01×0.00464flashinfer
fa3
0.00931
0.0179
0.226
single-token-page128-float161.54×0.006flashinfer0.009250.699

GroupedQueryAttentionDecodeWithKVCacheFwd (17 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
llama-70b-32k-bfloat161.08×0.1376fa3
flashinfer
0.1493
0.4417
31.2
llama-70b-32k-float161.09×0.1382fa3
flashinfer
0.1504
0.3868
31.1
llama-70b-4k-bfloat161.08×0.0791fa3
flashinfer
0.0852
0.2274
27.2
llama-70b-4k-float161.08×0.0793fa3
flashinfer
0.0855
0.1995
27.1
llama-8b-32k-bfloat161.04×0.2567fa3
flashinfer
0.2678
0.4963
16.7
llama-8b-32k-float161.04×0.2579fa3
flashinfer
0.2689
0.4288
16.7
llama-8b-4k-bfloat161.02×0.1503fa3
flashinfer
0.1530
0.2566
14.3
llama-8b-4k-float161.02×0.1513fa3
flashinfer
0.1540
0.2239
14.2
llama-8b-4k-softcap50-float1682.46×0.1616torch-sdpa13.326113.3
qwen3-30b-a3b-bs1-128k-float161.09×0.0764flashinfer
fa3
0.0832
0.0929
28.1
qwen3-30b-a3b-bs1-16k-float161.21×0.0182flashinfer
fa3
0.0219
0.0281
14.8
qwen3-30b-a3b-bs1-1k-float161.39×0.00698flashinfer
fa3
0.0097
0.0174
2.4
qwen3-30b-a3b-bs1-256k-float161.04×0.1362flashinfer
fa3
0.1420
0.1618
31.5
qwen3-30b-a3b-bs1-32k-float161.22×0.0283flashinfer
fa3
0.0347
0.0376
18.9
qwen3-30b-a3b-bs1-4k-float161.21×0.0096flashinfer
fa3
0.0116
0.0212
6.99
qwen3-30b-a3b-bs1-64k-float161.16×0.0456flashinfer
fa3
0.0531
0.0576
23.6
qwen3-30b-a3b-bs1-8k-float161.07×0.0132flashinfer
fa3
0.0141
0.0232
10.2

GroupedQueryAttentionPrefillFwd (16 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
llama-70b-prefill-dense-q-lt-kv-bfloat161.00×0.1238flashinfer
torch-ref
0.1235
3.7621
521
llama-70b-prefill-dense-q-lt-kv-float161.00×0.1252flashinfer
torch-ref
0.1251
3.7608
515
llama-70b-prefill-fp8tc-bn224-s3584-bfloat160.71×0.7495fa3
torch-sdpa-dequant
0.5289
1.0509
562
llama-70b-prefill-fp8tc-bn224-s3584-float160.71×0.7495fa3
torch-sdpa-dequant
0.5302
1.0531
562
llama-70b-prefill-fp8tc-bn224-s7168-bfloat160.71×2.8453fa3
torch-sdpa-dequant
2.0226
3.4318
592
llama-70b-prefill-fp8tc-bn224-s7168-float160.71×2.8449fa3
torch-sdpa-dequant
2.0283
3.4318
592
llama-8b-prefill-dense-bfloat161.06×0.0369flashinfer
torch-ref
0.0393
1.1011
233
llama-8b-prefill-dense-float161.05×0.0373flashinfer
torch-ref
0.0393
1.1007
231
llama-8b-prefill-dense-q-lt-kv-bfloat161.01×0.1243flashinfer
torch-ref
0.1252
4.1005
518
llama-8b-prefill-dense-q-lt-kv-float161.00×0.1263flashinfer
torch-ref
0.1265
4.0955
510
llama-8b-prefill-dense-sm-scale-0.125-float161.06×0.0371flashinfer
torch-ref
0.0393
1.1009
232
llama-8b-prefill-dense-softcap50-float161.09×0.0421flashinfer
torch-ref
0.0457
1.2966
205
llama-8b-prefill-fp8tc-bn224-s1792-bfloat160.67×0.1290fa3
torch-sdpa-dequant
0.0863
0.2255
408
llama-8b-prefill-fp8tc-bn224-s1792-float160.67×0.1288fa3
torch-sdpa-dequant
0.0860
0.2270
408
llama-8b-prefill-fp8tc-bn224-s896-bfloat160.62×0.0454fa3
torch-sdpa-dequant
0.0283
0.0926
290
llama-8b-prefill-fp8tc-bn224-s896-float160.63×0.0453fa3
torch-sdpa-dequant
0.0286
0.0918
290

MultiHeadAttentionDecodeWithKVCacheFwd (8 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
llama-70b-32k-bfloat161.01×0.9798fa3
flashinfer
0.9864
0.9966
4.38
llama-70b-32k-float161.01×0.9811fa3
flashinfer
0.9864
0.9962
4.38
llama-70b-4k-bfloat161.00×0.5139fa3
flashinfer
0.5143
0.5298
4.18
llama-70b-4k-float161.00×0.5139fa3
flashinfer
0.5150
0.5308
4.18
llama-8b-32k-bfloat161.00×0.9820fa3
flashinfer
0.9862
0.9989
4.37
llama-8b-32k-float161.01×0.9809fa3
flashinfer
0.9866
0.9989
4.38
llama-8b-4k-bfloat161.00×0.5104fa3
flashinfer
0.5116
0.5303
4.21
llama-8b-4k-float161.00×0.5111fa3
flashinfer
0.5127
0.5294
4.2

GroupedQueryAttentionPrefillPagedWithKVCacheFwd (7 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
gqa-prefill-paged-fp8-cache-softcap50-b4-prefix4k-chunk512-p64-float160.2071··88.1
gqa-prefill-paged-softcap50-b4-prefix4k-chunk512-p64-float160.91×0.1499fa30.1367122
llama-8b-prefill-paged-b8-prefix4k-chunk512-p64-full-rope-float161.9500··150
llama-8b-prefill-paged-fp8-cache-b8-prefix4k-chunk512-p64-float162.0061··146
qwen35-9b-prefill-paged-fp8-cache-b8-prefix32k-chunk1k-p64-float1655.9919··160
qwen35-9b-prefill-paged-fullattn-b8-prefix32k-chunk1k-p64-partial-rope64-float1660.6283··147
qwen35-9b-prefill-paged-fullattn-mixed-b8-p64-partial-rope64-float1630.7355··108

GroupedQueryAttentionFwd (8 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
llama-70b-long-bfloat160.82×0.1617fa3
flashinfer
0.1322
0.1598
425
llama-70b-long-float160.82×0.1627fa3
flashinfer
0.1339
0.1619
422
llama-70b-short-bfloat160.84×0.0380fa3
flashinfer
0.0317
0.0392
226
llama-70b-short-float160.84×0.0382fa3
flashinfer
0.0319
0.0391
225
llama-8b-long-bfloat160.83×0.1613fa3
flashinfer
0.1331
0.1612
426
llama-8b-long-float160.83×0.1627fa3
flashinfer
0.1348
0.1615
422
llama-8b-short-bfloat160.86×0.0369fa3
flashinfer
0.0319
0.0393
233
llama-8b-short-float160.86×0.0371fa3
flashinfer
0.0319
0.0395
232

GroupedQueryAttentionDecodePagedWithKVCacheFwd (8 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
serving-405b-p256-float160.47×0.0563fa30.026619.1
serving-70b-p256-float160.54×0.0685fa3
flashinfer
0.0368
0.0572
15.7
serving-70b-p64-float160.90×0.0496flashinfer0.044521.6
serving-8b-long-p64-float161.36×0.2206flashinfer0.299619.5
serving-8b-p256-float160.48×0.1683fa3
flashinfer
0.0812
0.1253
12.8
serving-8b-p64-float160.75×0.1669flashinfer0.125212.9
serving-8b-p64-softcap50-float160.71×0.1765flashinfer0.125212.2
throughput-8b-p64-float160.60×0.2518flashinfer0.15078.53

MultiHeadAttentionFwd (8 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
llama-70b-long-bfloat160.82×0.1672fa3
flashinfer
0.1368
0.1628
411
llama-70b-long-float160.83×0.1679fa3
flashinfer
0.1395
0.1644
409
llama-70b-short-bfloat160.83×0.0425fa3
flashinfer
0.0353
0.0411
202
llama-70b-short-float160.83×0.0428fa3
flashinfer
0.0354
0.0412
201
llama-8b-long-bfloat160.82×0.1669fa3
flashinfer
0.1365
0.1620
412
llama-8b-long-float160.82×0.1682fa3
flashinfer
0.1383
0.1652
408
llama-8b-short-bfloat160.83×0.0425fa3
flashinfer
0.0355
0.0409
202
llama-8b-short-float160.82×0.0425fa3
flashinfer
0.0348
0.0412
202

GroupedQueryAttentionSlidingWindowFwd (8 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
llama-70b-long-w1024-bfloat160.78×0.1516fa3
flashinfer
0.1184
0.1514
340
llama-70b-long-w1024-float160.78×0.1528fa3
flashinfer
0.1199
0.1531
338
llama-70b-short-w256-bfloat160.86×0.0395fa3
flashinfer
0.0340
0.0408
164
llama-70b-short-w256-float160.86×0.0396fa3
flashinfer
0.0342
0.0411
164
llama-8b-long-w1024-bfloat160.78×0.1517fa3
flashinfer
0.1183
0.1529
340
llama-8b-long-w1024-float160.79×0.1529fa3
flashinfer
0.1209
0.1553
337
llama-8b-short-w256-bfloat160.85×0.0397fa3
flashinfer
0.0340
0.0411
163
llama-8b-short-w256-float160.86×0.0398fa3
flashinfer
0.0342
0.0412
162

GroupedQueryAttentionSlidingWindowVarlenFwd (8 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
llama-70b-long-w1024-bfloat160.78×0.6661flashinfer
fa3
0.5171
0.5271
310
llama-70b-long-w1024-float160.78×0.6677flashinfer
fa3
0.5185
0.5281
309
llama-70b-short-w256-bfloat160.75×0.0930flashinfer
fa3
0.0693
0.0834
139
llama-70b-short-w256-float160.74×0.0931flashinfer
fa3
0.0691
0.0834
139
llama-8b-long-w1024-bfloat160.77×0.3492fa3
flashinfer
0.2700
0.2734
295
llama-8b-long-w1024-float160.77×0.3518fa3
flashinfer
0.2714
0.2767
293
llama-8b-short-w256-bfloat160.73×0.0567flashinfer
fa3
0.0411
0.0471
114
llama-8b-short-w256-float160.73×0.0569flashinfer
fa3
0.0414
0.0472
114

GroupedQueryAttentionBwd (8 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
llama-70b-long-bfloat160.57×1.0192fa30.5770169
llama-70b-long-float160.72×0.8080fa30.5802213
llama-70b-short-bfloat160.39×0.4092fa30.158852.5
llama-70b-short-float160.81×0.1961fa30.1592109
llama-8b-long-bfloat160.47×1.2427fa30.5891138
llama-8b-long-float160.71×0.8322fa30.5925206
llama-8b-short-bfloat160.40×0.4154fa30.165451.7
llama-8b-short-float160.82×0.2029fa30.1661106

GroupedQueryAttentionPrefillVarlenFwd (3 workloads)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
llama-70b-prefill-varlen-q-lt-kv-bf160.50×0.1963fa3
torch-ref
0.0985
2.7647
219
llama-8b-prefill-varlen-mixed-fp160.44×0.1405fa3
torch-ref
0.0614
1.6792
143
llama-8b-prefill-varlen-uniform-fp160.57×0.1251fa3
torch-ref
0.0715
2.0383
206

MultiHeadAttentionBwd (8 workloads · ✅)

Workload Ratio Device time Alternatives Throughput
alt / ours ms name ms TFLOP/s
llama-70b-long-bfloat160.50×1.1015fa30.5495156
llama-70b-long-float160.62×0.8926fa30.5515192
llama-70b-short-bfloat160.31×0.4565fa30.143147
llama-70b-short-float160.59×0.2444fa30.143387.9
llama-8b-long-bfloat160.42×1.3118fa30.5471131
llama-8b-long-float160.61×0.9026fa30.5513190
llama-8b-short-bfloat160.31×0.4562fa30.143347.1
llama-8b-short-float160.59×0.2435fa30.143588.2