Linear Attention & SSM¶
本页暂无中文版。以下为英文原文。
22 ops, 158 workloads.
One table per op, one row per workload. Ratio is the fastest other implementation's device time divided by ours, so green is faster than it, plain is level with it, red is slower. Times are in ms. How these numbers are taken.
Linear Attention / SSM¶
CBProducerFwd (2 workloads)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
mamba2-2p7b-b4-s2k-bfloat16 | 4.48× | 0.0118 | torch | 0.0531 | 22.7 |
mamba2-780m-b1-s4k-float16 | 5.30× | 0.00714 | torch | 0.0378 | 18.8 |
DeltaNetDecodeFwd (8 workloads · ✅)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
delta-decode-serving-b1-h16-d128-bfloat16 | 4.57× | 0.0031 | torch-compiletorch | 0.0142 0.0353 | 0.508 |
delta-decode-serving-b1-h32-d128-bfloat16 | 4.81× | 0.00336 | torch-compiletorch | 0.0162 0.0391 | 0.938 |
delta-decode-serving-b1-h48-d128-bfloat16 | 5.42× | 0.00355 | torch-compiletorch | 0.0192 0.0434 | 1.33 |
delta-decode-serving-b1-h64-d128-bfloat16 | 4.51× | 0.00384 | torch-compiletorch | 0.0173 0.0438 | 1.64 |
delta-decode-serving-b1-h8-d128-bfloat16 | 4.60× | 0.00285 | torch-compiletorch | 0.0131 0.0335 | 0.277 |
delta-decode-serving-b8-h32-d128-bfloat16 | 4.38× | 0.0087 | torch-compiletorch | 0.0381 0.0896 | 2.9 |
delta-decode-serving-b8-h48-d128-bfloat16 | 3.22× | 0.0123 | torch-compiletorch | 0.0396 0.1111 | 3.07 |
delta-decode-serving-b8-h64-d128-bfloat16 | 3.15× | 0.0163 | torch-compiletorch | 0.0515 0.1420 | 3.09 |
EngramGateConvBwd (3 workloads · ✅)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
bwd-b1-s128-d256-bfloat16 | 3.24× | 0.0168 | torch-compiletorch | 0.0544 0.1838 | 0.118 |
bwd-b1-s32-d256-float16 | 4.48× | 0.0111 | torch-compiletorch | 0.0498 0.1677 | 0.0444 |
bwd-b2-s64-d512-float16 | 3.01× | 0.0198 | torch-compiletorch | 0.0595 0.2003 | 0.199 |
GatedDeltaNetPrefillBTHDFwd (26 workloads)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
fallback-gdn-prefill-b1-s4k-h16-d128-bthd-bfloat16 | 2.49× | 0.0794 | fla | 0.1974 | 108 |
fallback-gdn-prefill-b1-s4k-h16-d128-bthd-float16 | 2.48× | 0.0792 | fla | 0.1960 | 109 |
qwen35-gdn-prefill-b1-s128k-h16-d128-bthd-bfloat16 | 4.50× | 1.2814 | fla | 5.7677 | 215 |
qwen35-gdn-prefill-b1-s128k-h16-d128-bthd-float16 | 4.59× | 1.2572 | fla | 5.7713 | 219 |
qwen35-gdn-prefill-b1-s128k-h32-d128-bthd-bfloat16 | 3.50× | 2.5076 | fla | 8.7773 | 219 |
qwen35-gdn-prefill-b1-s128k-h32-d128-bthd-float16 | 3.58× | 2.4458 | fla | 8.7477 | 225 |
qwen35-gdn-prefill-b1-s128k-h48-d128-bthd-bfloat16 | 3.30× | 3.8074 | fla | 12.5781 | 217 |
qwen35-gdn-prefill-b1-s128k-h48-d128-bthd-float16 | 3.34× | 3.7766 | fla | 12.5997 | 218 |
qwen35-gdn-prefill-b1-s128k-h64-d128-bthd-bfloat16 | 3.26× | 4.7856 | fla | 15.5906 | 230 |
qwen35-gdn-prefill-b1-s128k-h64-d128-bthd-float16 | 3.35× | 4.6645 | fla | 15.6272 | 236 |
qwen35-gdn-prefill-b1-s32k-h16-d128-bthd-bfloat16 | 3.96× | 0.3711 | fla | 1.4701 | 185 |
qwen35-gdn-prefill-b1-s32k-h16-d128-bthd-float16 | 4.01× | 0.3653 | fla | 1.4649 | 188 |
qwen35-gdn-prefill-b1-s32k-h32-d128-bthd-bfloat16 | 3.17× | 0.7005 | fla | 2.2183 | 196 |
qwen35-gdn-prefill-b1-s32k-h32-d128-bthd-float16 | 3.23× | 0.6862 | fla | 2.2146 | 200 |
qwen35-gdn-prefill-b1-s32k-h48-d128-bthd-bfloat16 | 2.97× | 1.0652 | fla | 3.1612 | 194 |
qwen35-gdn-prefill-b1-s32k-h48-d128-bthd-float16 | 3.02× | 1.0493 | fla | 3.1685 | 196 |
qwen35-gdn-prefill-b1-s32k-h64-d128-bthd-bfloat16 | 3.11× | 1.2532 | fla | 3.8987 | 219 |
qwen35-gdn-prefill-b1-s32k-h64-d128-bthd-float16 | 3.19× | 1.2242 | fla | 3.9055 | 225 |
qwen35-gdn-prefill-b1-s64k-h16-d128-bthd-bfloat16 | 4.13× | 0.7069 | fla | 2.9162 | 194 |
qwen35-gdn-prefill-b1-s64k-h16-d128-bthd-float16 | 4.17× | 0.6972 | fla | 2.9043 | 197 |
qwen35-gdn-prefill-b1-s64k-h32-d128-bthd-bfloat16 | 3.45× | 1.2767 | fla | 4.4014 | 215 |
qwen35-gdn-prefill-b1-s64k-h32-d128-bthd-float16 | 3.53× | 1.2458 | fla | 4.3951 | 221 |
qwen35-gdn-prefill-b1-s64k-h48-d128-bthd-bfloat16 | 3.25× | 1.9437 | fla | 6.3077 | 212 |
qwen35-gdn-prefill-b1-s64k-h48-d128-bthd-float16 | 3.29× | 1.9168 | fla | 6.3109 | 215 |
qwen35-gdn-prefill-b1-s64k-h64-d128-bthd-bfloat16 | 3.21× | 2.4273 | fla | 7.7806 | 226 |
qwen35-gdn-prefill-b1-s64k-h64-d128-bthd-float16 | 3.29× | 2.3772 | fla | 7.8149 | 231 |
EngramGateConvFwd (3 workloads · ✅)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
fwd-b1-s128-d256-bfloat16 | 2.60× | 0.00448 | torch-compiletorch-ref | 0.0116 0.0802 | 0.176 |
fwd-b1-s32-d256-float16 | 2.72× | 0.004 | torch-compiletorch-ref | 0.0109 0.0740 | 0.0493 |
fwd-b2-s64-d512-float16 | 2.46× | 0.00506 | torch-compiletorch-ref | 0.0124 0.0856 | 0.312 |
SSDDecodeFwd (4 workloads)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
mamba2-1p3b-decode-b1-bfloat16 | 2.30× | 0.00397 | torch-compiletorch-ref | 0.00912 0.0304 | 1.06 |
mamba2-1p3b-decode-b1-float16 | 2.28× | 0.00397 | torch-compiletorch-ref | 0.00906 0.0304 | 1.06 |
mamba2-2p7b-decode-b8-float16 | 1.84× | 0.0163 | torch-compiletorch-ref | 0.0301 0.1119 | 2.57 |
mamba2-780m-decode-b32-float16 | 1.92× | 0.0361 | torch-compiletorch-ref | 0.0695 0.2409 | 2.79 |
SSDStatePassingFwd (6 workloads · ✅)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
mamba2-1p3b-b1-s4k-dstate-bfloat16 | 2.10× | 0.00195 | torch-compilemambatorch-ref | 0.0041 0.00835 0.1229 | 0.135 |
mamba2-1p3b-b1-s4k-dstate-float16 | 2.10× | 0.00195 | torch-compilemambatorch-ref | 0.0041 0.00842 0.1215 | 0.135 |
mamba2-1p3b-b1-s4k-dstate-init-states-bfloat16 | 1.11× | 0.00197 | torch-compilemambatorch-ref | 0.00218 0.0087 0.1220 | 0.134 |
mamba2-1p3b-b1-s4k-dstate-init-states-float16 | 1.71× | 0.002 | torch-compilemambatorch-ref | 0.00342 0.0087 0.1206 | 0.132 |
mamba2-1p3b-b1-s4k-flat-init-states-float32 | 0.93× | 0.0219 | torch-compilemambatorch-ref | 0.0204 0.0216 0.1270 | 0.765 |
mamba2-2p7b-b2-s32k-dstate-float16 | 5.65× | 0.0106 | mambatorch-compiletorch-ref | 0.0599 0.0912 1.1537 | 0.497 |
GatedDeltaNetDecodeFwd (8 workloads · ✅)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
gdn-decode-serving-b1-h16-d128-bfloat16 | 1.27× | 0.0033 | fla | 0.00419 | 0.478 |
gdn-decode-serving-b1-h32-d128-bfloat16 | 1.30× | 0.00362 | fla | 0.00469 | 0.872 |
gdn-decode-serving-b1-h48-d128-bfloat16 | 1.37× | 0.00384 | fla | 0.00525 | 1.23 |
gdn-decode-serving-b1-h64-d128-bfloat16 | 1.38× | 0.00419 | fla | 0.00579 | 1.5 |
gdn-decode-serving-b1-h8-d128-bfloat16 | 1.29× | 0.0031 | fla | 0.004 | 0.254 |
gdn-decode-serving-b8-h32-d128-bfloat16 | 1.67× | 0.00874 | fla | 0.0146 | 2.89 |
gdn-decode-serving-b8-h48-d128-bfloat16 | 1.56× | 0.0124 | fla | 0.0193 | 3.05 |
gdn-decode-serving-b8-h64-d128-bfloat16 | 1.56× | 0.0161 | fla | 0.0252 | 3.14 |
SSDChunkScanFwd (4 workloads · ✅)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
mamba2-1p3b-b2-s32k-float16 | 1.39× | 1.4677 | mambatorch-compiletorch-ref | 2.0346 9.9818 40.0654 | 93.6 |
mamba2-2p7b-b4-s2k-bfloat16 | 1.30× | 0.2372 | mambatorch-compiletorch-ref | 0.3088 1.6394 6.5129 | 90.5 |
mamba2-780m-b1-s4k-bfloat16 | 1.34× | 0.0759 | mambatorch-compiletorch-ref | 0.1018 0.5090 1.9619 | 84.9 |
mamba2-780m-b1-s4k-float16 | 1.38× | 0.0732 | mambatorch-compiletorch-ref | 0.1007 0.5073 1.9596 | 88.1 |
GatedDeltaNetBTHDFwd (12 workloads)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
gdn-bthd-b1-s4k-h16-d128-bfloat16 | 1.12× | 0.1742 | fla | 0.1948 | 58.6 |
gdn-bthd-b1-s4k-h16-d128-float16 | 1.11× | 0.1750 | fla | 0.1935 | 58.3 |
gdn-bthd-b2-s16k-h4-d64-bfloat16 | 0.91× | 0.4327 | fla | 0.3932 | 19.9 |
gdn-bthd-b2-s16k-h4-d64-float16 | 0.91× | 0.4290 | fla | 0.3907 | 20 |
gdn-bthd-b2-s2k-h4-d64-bfloat16 | 1.02× | 0.0664 | fla | 0.0675 | 16.2 |
gdn-bthd-b2-s2k-h4-d64-float16 | 1.00× | 0.0669 | fla | 0.0672 | 16 |
gdn-bthd-b2-s32k-h4-d64-bfloat16 | 3.95× | 0.1950 | fla | 0.7693 | 88.1 |
gdn-bthd-b2-s32k-h4-d64-float16 | 3.94× | 0.1949 | fla | 0.7679 | 88.2 |
gdn-bthd-b2-s4k-h4-d64-bfloat16 | 0.95× | 0.1145 | fla | 0.1083 | 18.8 |
gdn-bthd-b2-s4k-h4-d64-float16 | 0.94× | 0.1151 | fla | 0.1077 | 18.7 |
gdn-bthd-b2-s8k-h4-d64-bfloat16 | 0.93× | 0.2204 | fla | 0.2056 | 19.5 |
gdn-bthd-b2-s8k-h4-d64-float16 | 0.94× | 0.2192 | fla | 0.2051 | 19.6 |
SSDChunkStateFwd (6 workloads · ✅)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
mamba2-1p3b-b2-s32k-seq-idx-float16 | 1.41× | 0.4492 | mambatorch-compiletorch-ref | 0.6315 16.6703 173.1990 | 154 |
mamba2-2p7b-b4-s2k-bfloat16 | 1.22× | 0.0656 | mambatorch-compiletorch-ref | 0.0802 2.4482 27.1195 | 164 |
mamba2-780m-b1-s4k-bfloat16 | 1.10× | 0.0240 | mambatorch-compiletorch-ref | 0.0265 0.6755 8.1408 | 135 |
mamba2-780m-b1-s4k-float16 | 1.05× | 0.0238 | mambatorch-compiletorch-ref | 0.0250 0.6331 8.1358 | 136 |
mamba2-780m-b1-s4k-seq-idx-bfloat16 | 1.01× | 0.0290 | mambatorch-compiletorch-ref | 0.0292 0.7870 8.1521 | 112 |
mamba2-780m-b1-s4k-seq-idx-float16 | 1.21× | 0.0287 | mambatorch-compiletorch-ref | 0.0347 0.7499 8.1473 | 113 |
Mamba2Fwd (8 workloads)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
mamba2-1p3b-b1-s8k-dt-bias-float16 | 1.08× | 0.2904 | mambatorch-compiletorch-ref | 0.3123 2.0138 6.9145 | 89.9 |
mamba2-1p3b-b1-s8k-dt-bias-init-states-float16 | 1.08× | 0.2899 | mambatorch-compiletorch-ref | 0.3125 2.0123 6.9059 | 90.1 |
mamba2-1p3b-b1-s8k-float16 | 1.08× | 0.2909 | mambatorch-compiletorch-ref | 0.3127 2.0180 6.9113 | 89.8 |
mamba2-1p3b-b1-s8k-init-states-float16 | 1.07× | 0.2915 | mambatorch-compiletorch-ref | 0.3126 2.0172 6.9060 | 89.6 |
mamba2-2p7b-b1-s2k-bfloat16 | 0.99× | 0.1100 | mambatorch-compiletorch-ref | 0.1093 0.6874 2.1548 | 74 |
mamba2-2p7b-b1-s2k-dt-bias-bfloat16 | 1.00× | 0.1091 | mambatorch-compiletorch-ref | 0.1096 0.6885 2.1567 | 74.6 |
mamba2-2p7b-b1-s2k-dt-bias-init-states-bfloat16 | 1.01× | 0.1099 | mambatorch-compiletorch-ref | 0.1106 0.6791 2.1555 | 74.1 |
mamba2-2p7b-b1-s2k-init-states-bfloat16 | 1.00× | 0.1106 | mambatorch-compiletorch-ref | 0.1108 0.6784 2.1521 | 73.6 |
GLADecodeFwd (8 workloads · ✅)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
gla-decode-serving-b1-h16-d128-bfloat16 | 0.90× | 0.00742 | torch-compileflatorch | 0.00669 0.00704 0.0318 | 0.142 |
gla-decode-serving-b1-h32-d128-bfloat16 | 0.94× | 0.00784 | flatorch-compiletorch | 0.00736 0.00835 0.0358 | 0.269 |
gla-decode-serving-b1-h48-d128-bfloat16 | 1.13× | 0.00794 | flatorch-compiletorch | 0.00896 0.0103 0.0404 | 0.398 |
gla-decode-serving-b1-h64-d128-bfloat16 | 1.09× | 0.00816 | flatorch-compiletorch | 0.00886 0.00973 0.0419 | 0.516 |
gla-decode-serving-b1-h8-d128-bfloat16 | 0.81× | 0.00742 | torch-compileflatorch | 0.00598 0.00672 0.0304 | 0.0709 |
gla-decode-serving-b8-h32-d128-bfloat16 | 1.10× | 0.0159 | flatorch-compiletorch | 0.0175 0.0211 0.0895 | 1.06 |
gla-decode-serving-b8-h48-d128-bfloat16 | 0.96× | 0.0231 | flatorch-compiletorch | 0.0223 0.0246 0.1208 | 1.09 |
gla-decode-serving-b8-h64-d128-bfloat16 | 0.89× | 0.0304 | flatorch-compiletorch | 0.0269 0.0320 0.1579 | 1.11 |
DeltaNetFwd (8 workloads)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
dn-b2-s16k-h4-d64-bfloat16 | 0.78× | 0.4729 | fla | 0.3700 | 2.27 |
dn-b2-s16k-h4-d64-float16 | 0.77× | 0.4726 | fla | 0.3661 | 2.27 |
dn-b2-s2k-h4-d64-bfloat16 | 0.99× | 0.0629 | fla | 0.0624 | 2.13 |
dn-b2-s2k-h4-d64-float16 | 0.99× | 0.0628 | fla | 0.0620 | 2.14 |
dn-b2-s4k-h4-d64-bfloat16 | 0.91× | 0.1097 | fla | 0.0993 | 2.45 |
dn-b2-s4k-h4-d64-float16 | 0.90× | 0.1095 | fla | 0.0985 | 2.45 |
dn-b2-s8k-h4-d64-bfloat16 | 0.81× | 0.2346 | fla | 0.1910 | 2.29 |
dn-b2-s8k-h4-d64-float16 | 0.81× | 0.2337 | fla | 0.1894 | 2.3 |
DeltaNetBwd (8 workloads)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
dn-bwd-b2-s16k-h4-d64-bfloat16 | 0.87× | 1.0041 | fla | 0.8701 | 2.14 |
dn-bwd-b2-s16k-h4-d64-float16 | 0.87× | 0.9924 | fla | 0.8642 | 2.16 |
dn-bwd-b2-s2k-h4-d64-bfloat16 | 0.87× | 0.1316 | fla | 0.1145 | 2.04 |
dn-bwd-b2-s2k-h4-d64-float16 | 0.87× | 0.1307 | fla | 0.1133 | 2.05 |
dn-bwd-b2-s4k-h4-d64-bfloat16 | 0.83× | 0.2619 | fla | 0.2167 | 2.05 |
dn-bwd-b2-s4k-h4-d64-float16 | 0.83× | 0.2595 | fla | 0.2144 | 2.07 |
dn-bwd-b2-s8k-h4-d64-bfloat16 | 0.85× | 0.5107 | fla | 0.4362 | 2.1 |
dn-bwd-b2-s8k-h4-d64-float16 | 0.86× | 0.5052 | fla | 0.4323 | 2.13 |
GatedDeltaNetPrefillBHTDFwd (4 workloads · ✅)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
bhtd-fallback-gdn-prefill-b1-s4k-h16-d128-bfloat16 | 0.78× | 0.2526 | fla | 0.1979 | 34 |
bhtd-fallback-gdn-prefill-b1-s4k-h16-d128-float16 | 0.78× | 0.2508 | fla | 0.1960 | 34.2 |
bhtd-qwen35-gdn-prefill-b1-s128k-h64-d128-bfloat16 | 0.89× | 17.5479 | fla | 15.5794 | 62.7 |
bhtd-qwen35-gdn-prefill-b1-s128k-h64-d128-float16 | 0.90× | 17.4198 | fla | 15.6294 | 63.1 |
GLABwd (8 workloads)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
gla-bwd-b2-s16k-h4-d64-bfloat16 | 0.75× | 1.4487 | fla | 1.0809 | 1.48 |
gla-bwd-b2-s16k-h4-d64-float16 | 0.71× | 1.5163 | fla | 1.0812 | 1.42 |
gla-bwd-b2-s2k-h4-d64-bfloat16 | 0.80× | 0.1844 | fla | 0.1481 | 1.46 |
gla-bwd-b2-s2k-h4-d64-float16 | 0.81× | 0.1829 | fla | 0.1482 | 1.47 |
gla-bwd-b2-s4k-h4-d64-bfloat16 | 0.79× | 0.3648 | fla | 0.2879 | 1.47 |
gla-bwd-b2-s4k-h4-d64-float16 | 0.78× | 0.3686 | fla | 0.2882 | 1.46 |
gla-bwd-b2-s8k-h4-d64-bfloat16 | 0.77× | 0.7267 | fla | 0.5569 | 1.48 |
gla-bwd-b2-s8k-h4-d64-float16 | 0.75× | 0.7446 | fla | 0.5576 | 1.44 |
GLAFwd (8 workloads)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
gla-init-b2-s16k-h4-d64-bfloat16 | 0.74× | 0.6112 | fla | 0.4522 | 1.76 |
gla-init-b2-s16k-h4-d64-float16 | 0.75× | 0.6180 | fla | 0.4652 | 1.74 |
gla-init-b2-s4k-h4-d64-bfloat16 | 0.76× | 0.1564 | fla | 0.1191 | 1.72 |
gla-init-b2-s4k-h4-d64-float16 | 0.80× | 0.1569 | fla | 0.1254 | 1.71 |
gla-noinit-b2-s2k-h4-d64-bfloat16 | 0.68× | 0.0969 | fla | 0.0659 | 1.39 |
gla-noinit-b2-s2k-h4-d64-float16 | 0.71× | 0.0986 | fla | 0.0702 | 1.36 |
gla-noinit-b2-s8k-h4-d64-bfloat16 | 0.71× | 0.3114 | fla | 0.2200 | 1.72 |
gla-noinit-b2-s8k-h4-d64-float16 | 0.79× | 0.3132 | fla | 0.2476 | 1.71 |
GatedDeltaNetBHTDFwd (8 workloads)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
gdn-bhtd-b2-s16k-h4-d64-bfloat16 | 0.62× | 0.6370 | fla | 0.3936 | 13.5 |
gdn-bhtd-b2-s16k-h4-d64-float16 | 0.61× | 0.6363 | fla | 0.3908 | 13.5 |
gdn-bhtd-b2-s2k-h4-d64-bfloat16 | 0.78× | 0.0866 | fla | 0.0675 | 12.4 |
gdn-bhtd-b2-s2k-h4-d64-float16 | 0.78× | 0.0866 | fla | 0.0672 | 12.4 |
gdn-bhtd-b2-s4k-h4-d64-bfloat16 | 0.75× | 0.1445 | fla | 0.1082 | 14.9 |
gdn-bhtd-b2-s4k-h4-d64-float16 | 0.72× | 0.1488 | fla | 0.1077 | 14.4 |
gdn-bhtd-b2-s8k-h4-d64-bfloat16 | 0.65× | 0.3166 | fla | 0.2056 | 13.6 |
gdn-bhtd-b2-s8k-h4-d64-float16 | 0.65× | 0.3138 | fla | 0.2051 | 13.7 |
GatedDeltaNetBwd (8 workloads)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
gdn-bwd-b2-s16k-h4-d64-bfloat16 | 0.65× | 1.4680 | fla | 0.9536 | 1.46 |
gdn-bwd-b2-s16k-h4-d64-float16 | 0.65× | 1.4237 | fla | 0.9204 | 1.51 |
gdn-bwd-b2-s2k-h4-d64-bfloat16 | 0.68× | 0.2049 | fla | 0.1401 | 1.31 |
gdn-bwd-b2-s2k-h4-d64-float16 | 0.66× | 0.2017 | fla | 0.1337 | 1.33 |
gdn-bwd-b2-s4k-h4-d64-bfloat16 | 0.67× | 0.3875 | fla | 0.2584 | 1.39 |
gdn-bwd-b2-s4k-h4-d64-float16 | 0.66× | 0.3809 | fla | 0.2495 | 1.41 |
gdn-bwd-b2-s8k-h4-d64-bfloat16 | 0.67× | 0.7504 | fla | 0.5018 | 1.43 |
gdn-bwd-b2-s8k-h4-d64-float16 | 0.67× | 0.7228 | fla | 0.4867 | 1.49 |
DaCumsumFwd (5 workloads · ✅)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
mamba2-1p3b-b8-s2k-bfloat16 | 0.42× | 0.0154 | mambatorch-compiletorch-ref | 0.00643 0.0114 0.0912 | 0.477 |
mamba2-1p3b-b8-s2k-dt-bias-bfloat16 | 0.43× | 0.0148 | mambatorch-compiletorch-ref | 0.00637 0.0115 0.0964 | 0.565 |
mamba2-2p7b-b2-s32k-dt-bias-float16 | 0.37× | 0.0599 | mambatorch-compiletorch-ref | 0.0223 0.0346 0.2369 | 0.7 |
mamba2-780m-b1-s4k-dt-bias-float16 | 0.81× | 0.00429 | mambatorch-compiletorch-ref | 0.00346 0.0048 0.0742 | 0.367 |
mamba2-780m-b1-s4k-float16 | 0.65× | 0.00509 | mambatorch-compiletorch-ref | 0.00333 0.00477 0.0713 | 0.27 |
EngramDecodeFwd (3 workloads · ✅)¶
| Workload | Ratio | Device time | Alternatives | Throughput | |
|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | |
decode-b1-dmem512-d256-float16 | 0.40× | 0.0330 | torch-compiletorch-ref | 0.0133 0.0946 | 0.0161 |
decode-b4-dmem1024-d512-float16 | 0.31× | 0.0827 | torch-compiletorch-ref | 0.0256 0.1212 | 0.102 |
decode-b8-dmem512-d256-bfloat16 | 0.63× | 0.0334 | torch-compiletorch-ref | 0.0212 0.1114 | 0.127 |