Initial matched 4K / 128 speed matrix
| Vanilla format | Vulkan / RADV | TheRock ROCm 7.15 | TG winner | ||||||
|---|---|---|---|---|---|---|---|---|---|
| PP tok/s | TG tok/s | TTFP | MTP accepted | PP tok/s | TG tok/s | TTFP | MTP accepted | ||
| BF16 | 80.84 | 13.69 | 43.60 s | 106 / 124 | 261.90 | 9.47 | 13.46 s | 99 / 160 | Vulkan |
| UD-Q8_K_XL | 136.49 | 14.86 | 25.83 s | 98 / 169 | 200.47 | 14.67 | 17.58 s | 98 / 169 | Vulkan +1.3% |
| UD-Q6_K_XL | 129.31 | 14.75 | 27.26 s | 97 / 173 | 223.33 | 19.37 | 15.78 s | 99 / 160 | TheRock |
| UD-Q5_K_XL | 156.21 | 29.00 | 22.57 s | 103 / 139 | 224.44 | 18.14 | 15.71 s | 100 / 156 | Vulkan |
| UD-Q4_K_XL | 217.26 | 28.26 | 16.23 s | 100 / 156 | 258.08 | 19.08 | 13.66 s | 99 / 160 | Vulkan |
| ROCmFP4 | 282.78 | 34.21 | 12.47 s | 93 / 133 | 320.00 | 38.24 | 11.02 s | 98 / 116 | TheRock |
| ROCmFPX UltraQuality 7.61 BPW, cap 3 | 147.07 | 18.67 | 23.97 s | 90 / 110 | 231.32 | 17.57 | 15.24 s | 90 / 110 | Vulkan +6.2% |
PP is prompt processing. TG is generated-token throughput. TTFP includes uncached prompt ingestion. MTP accepted reports accepted draft tokens over proposed draft tokens. Higher PP/TG and lower TTFP are better. The UltraQuality row uses its latest matched 3,525-prompt / 128-generation cap-3 gate; the other rows retain their original matrix settings. The validated gates and cap selection below are the current tuning reference.
Validated new-TheRock and specialized-runner speed gates
These optimized-build gates are used alongside the Hard Mode audit. The stock lanes run stock llama.cpp; ROCmFP4 and UltraQuality run the required ROCmFPX or dedicated Vulkan implementations.
| Lane | Backend / runner | PP tok/s | Generation tok/s | Draft acceptance |
|---|---|---|---|---|
| BF16 | Stock / new TheRock | 336.69 | 8.83 | 90 / 110 (81.82%) |
| UD-Q8_K_XL | Stock / new TheRock | 294.38 | 14.20 | 90 / 110 (81.82%) |
| UD-Q6_K_XL | Stock / new TheRock | 254.04 | 18.08 | 90 / 108 (83.33%) |
| UD-Q5_K_XL | Stock / new TheRock | 280.39 | 20.08 | 90 / 111 (81.08%) |
| UD-Q4_K_XL | Stock / new TheRock | 295.80 | 21.48 | 91 / 108 (84.26%) |
| ROCmFP4 | ROCmFPX / new TheRock | 277.55 | 33.65 | 90 / 108 (83.33%) |
| ROCmFP4 | Dedicated Vulkan, cap 4 | 263.12 | 34.49 | Suite counters |
| ROCmFPX UltraQuality 7.61 BPW | Matched Vulkan | 147.07 | 18.67 | 90 / 110 (81.82%) |
| ROCmFPX UltraQuality 7.61 BPW | ROCmFPX / new TheRock | 231.32 | 17.57 | 90 / 110 (81.82%) |
The dedicated Vulkan ROCmFP4 values are active counters from its earlier valid suite run and retain that lane's cap-4 profile. All other rows are standardized 3.5K-prompt / 128-generation gates.
Tool-Eval Hard Mode quality
The optional Hard Mode suite covers TC-70 through TC-84: 15 scenarios scored on tool choice, argument construction, recovery, state handling, and verification. Quality is intentionally reported separately from the short speed gate.
| Lane | Backend / runner | Score | Pass / partial / fail | Runtime |
|---|---|---|---|---|
| ROCmFP4 | Dedicated Vulkan, MTP cap 4 | 70 | 9 / 3 / 3 | 261 s |
| UD-Q8_K_XL | Stock llama.cpp, new TheRock, cap 3 | 67 | 9 / 2 / 4 | 501 s |
| BF16 | Stock llama.cpp, new TheRock, cap 3 | 63 | 8 / 3 / 4 | 659 s |
| UD-Q4_K_XL | Stock llama.cpp, new TheRock, cap 3 | 63 | 8 / 3 / 4 | 454 s |
| ROCmFP4 | ROCmFPX, new TheRock, cap 3 | 63 | 8 / 3 / 4 | 330 s |
| UD-Q6_K_XL | Stock llama.cpp, new TheRock, cap 3 | 60 | 8 / 2 / 5 | 465 s |
| UD-Q5_K_XL | Stock llama.cpp, new TheRock, cap 3 | 57 | 7 / 3 / 5 | 443 s |
| ROCmFPX UltraQuality 7.61 BPW | Matched Vulkan, cap 3 | 60 | 8 / 2 / 5 | 614 s |
| ROCmFPX UltraQuality 7.61 BPW | ROCmFPX, new TheRock, cap 3 | 60 | 8 / 2 / 5 | 479 s |
Comparable stored Hard Mode evidence
| Model / profile | Score | Comparison note |
|---|---|---|
| Laguna S 2.1 Q4 | 87 | Target-only Vulkan; highest stored 2.0.7 Hard Mode result found |
| DeepSeek V4 Flash API | 83 | Clean API rerun |
| Step 3.7 QualityPlus | 83 | Prior best, temperature 1, seed 43 |
| Hy3 FPX-iFP2 | 80 | Native MTP, intended profile |
| Step 3.7 AGENTKV | 80 | Prior best, temperature 1, seed 43 |
| Qwen3.6 35B MoEQuality 7.07 BPW | 73 | Prior comparable local result |
| Qwen3.6 35B vanilla Q8 | 70 | Prior comparable local result |
| Atlas PR #353 NVFP4 MTP | 63 | Same Qwen3.6 27B family |
| Qwen3.6 35B ROCmFP4 | 60 | Prior comparable local result |
The intended-profile rows do not all share one sampler. They are useful cross-model context, not a claim that quantization alone caused every score difference.
Why MTP cap 3 replaced cap 6
The initial speed matrix used a six-token Q4 draft cap. A sweep on the same stock Q4/new-TheRock long request showed that six was over-drafting: acceptance fell to 59.04% and actual generation slowed.
| Draft cap | Generation tok/s | Accepted / drafted | Acceptance | Decision |
|---|---|---|---|---|
| 6 | 18.69 | 98 / 166 | 59.04% | Rejected: excessive drafting |
| 4 | 20.90 | 95 / 126 | 75.40% | Better, still below target |
| 3 | 22.07 | 91 / 108 | 84.26% | Selected: fastest |
| 2 | 20.02 | 81 / 91 | 89.01% | Highest acceptance, lower speed |
Acceptance is a diagnostic, not the objective by itself. Cap 3 won because it delivered the highest measured generation throughput while clearing 80% acceptance.
Artifact footprint and effective bits per weight
| Vanilla format | GGUF size | Effective BPW | Basis |
|---|---|---|---|
| BF16 | 54.658 GB | 16.00 | Two GGUF shards |
| UD-Q8_K_XL | 35.776 GB | 10.48 | Full artifact |
| UD-Q6_K_XL | 26.015 GB | 7.62 | Full artifact |
| UD-Q5_K_XL | 20.351 GB | 5.96 | Full artifact |
| UD-Q4_K_XL | 17.909 GB | 5.24 | Full artifact |
| ROCmFP4 | 14.817 GB | 4.34 | Full artifact |
| ROCmFPX UltraQuality | 26.005 GB | 7.61 | Full vanilla artifact |
Effective BPW is total GGUF bytes × 8 divided by 27,320,697,856 model parameters. It includes GGUF metadata and therefore measures the downloadable artifact, not merely the nominal tensor label. The run logs captured whole-host GTT on this unified-memory APU, not a clean per-process loaded allocation, so no misleading “VRAM loaded” column is reported. Runtime memory would also include KV cache, MTP context, and compute buffers and would vary with context and backend.
What the matrix shows
The backend winner depends on the quantization path. Vulkan leads generation for BF16, Q8, Q5, Q4, and UltraQuality; TheRock leads Q6 and ROCmFP4. TheRock consistently processes this prompt faster, but that PP advantage does not automatically translate into faster token generation.
UltraQuality is backend-stable in this harness. Matched Vulkan and new-TheRock runs both scored 60 with the same 8/2/5 scenario outcome and identical inference controls. TheRock processed the speed-gate prompt 57.3% faster, while Vulkan generated 6.2% faster; both accepted 90 of 110 MTP draft tokens.
At 60, UltraQuality matches the Q6 Hard Mode result, trails Q8 by seven points, and exceeds Q5 by three. The identical per-scenario backend outcome indicates that Vulkan versus new TheRock did not determine the measured tool-quality score.
The UltraQuality artifact is built from the same vanilla Unsloth BF16 source as the rest of the matrix using an importance matrix and the ranked leave-32 ROCmFPX/Q6_K-splice tensor policy. No fine-tuned derivative weights are included.
Method and scope
Initial speed matrix
- Same runtime source commit:
5d04ce30c. - Vulkan0 via RADV versus ROCm0 via TheRock 7.15.
- One slot, full GPU offload, flash attention, 8K server context.
- The matrix prompt generated exactly 128 tokens with cache off and temperature 0.
- Original rows use MTP cap 6 except ROCmFP4 cap 4; UltraQuality uses the matched 3,525-token cap-3 validation gate.
Hard Mode matrix
- Tool-Eval Bench 2.0.7 optional Hard Mode, TC-70 through TC-84.
- Stock formats use stock llama.cpp; only ROCmFPX formats use their specialized runners.
- TheRock means the new TheRock 7.15 runtime.
- Matched Qwen controls: embedded Jinja template, reasoning off, temperature 0, top-p 0.95, top-k 20, min-p 0, seed 123, 64K context, and MTP cap 3.
- Results characterize this Strix Halo host and these exact builds.
