Crown Citadel Group Ciru Inference Labllm.ciru.ai / research Research Index

Crown Citadel Research Report · July 22, 2026

Qwen3.6 27B Vanilla MTP: Vulkan vs TheRock

Matched served-API speed measurements, Tool-Eval Hard Mode quality, MTP cap selection, and effective bits-per-weight across seven vanilla Qwen3.6 27B formats on AMD Strix Halo.

336.69Fastest validated prompt gate, BF16 new TheRock.
34.49Fastest validated generation, ROCmFP4 Vulkan.
70 / 100Best Qwen Hard Mode score, ROCmFP4 Vulkan.
Cap 3Best measured Q4 MTP generation throughput.

Initial matched 4K / 128 speed matrix

Vanilla formatVulkan / RADVTheRock ROCm 7.15TG winner
PP tok/sTG tok/sTTFPMTP acceptedPP tok/sTG tok/sTTFPMTP accepted
BF1680.8413.6943.60 s106 / 124261.909.4713.46 s99 / 160Vulkan
UD-Q8_K_XL136.4914.8625.83 s98 / 169200.4714.6717.58 s98 / 169Vulkan +1.3%
UD-Q6_K_XL129.3114.7527.26 s97 / 173223.3319.3715.78 s99 / 160TheRock
UD-Q5_K_XL156.2129.0022.57 s103 / 139224.4418.1415.71 s100 / 156Vulkan
UD-Q4_K_XL217.2628.2616.23 s100 / 156258.0819.0813.66 s99 / 160Vulkan
ROCmFP4282.7834.2112.47 s93 / 133320.0038.2411.02 s98 / 116TheRock
ROCmFPX UltraQuality 7.61 BPW, cap 3147.0718.6723.97 s90 / 110231.3217.5715.24 s90 / 110Vulkan +6.2%

PP is prompt processing. TG is generated-token throughput. TTFP includes uncached prompt ingestion. MTP accepted reports accepted draft tokens over proposed draft tokens. Higher PP/TG and lower TTFP are better. The UltraQuality row uses its latest matched 3,525-prompt / 128-generation cap-3 gate; the other rows retain their original matrix settings. The validated gates and cap selection below are the current tuning reference.

Validated new-TheRock and specialized-runner speed gates

These optimized-build gates are used alongside the Hard Mode audit. The stock lanes run stock llama.cpp; ROCmFP4 and UltraQuality run the required ROCmFPX or dedicated Vulkan implementations.

LaneBackend / runnerPP tok/sGeneration tok/sDraft acceptance
BF16Stock / new TheRock336.698.8390 / 110 (81.82%)
UD-Q8_K_XLStock / new TheRock294.3814.2090 / 110 (81.82%)
UD-Q6_K_XLStock / new TheRock254.0418.0890 / 108 (83.33%)
UD-Q5_K_XLStock / new TheRock280.3920.0890 / 111 (81.08%)
UD-Q4_K_XLStock / new TheRock295.8021.4891 / 108 (84.26%)
ROCmFP4ROCmFPX / new TheRock277.5533.6590 / 108 (83.33%)
ROCmFP4Dedicated Vulkan, cap 4263.1234.49Suite counters
ROCmFPX UltraQuality 7.61 BPWMatched Vulkan147.0718.6790 / 110 (81.82%)
ROCmFPX UltraQuality 7.61 BPWROCmFPX / new TheRock231.3217.5790 / 110 (81.82%)

The dedicated Vulkan ROCmFP4 values are active counters from its earlier valid suite run and retain that lane's cap-4 profile. All other rows are standardized 3.5K-prompt / 128-generation gates.

Tool-Eval Hard Mode quality

The optional Hard Mode suite covers TC-70 through TC-84: 15 scenarios scored on tool choice, argument construction, recovery, state handling, and verification. Quality is intentionally reported separately from the short speed gate.

LaneBackend / runnerScorePass / partial / failRuntime
ROCmFP4Dedicated Vulkan, MTP cap 4709 / 3 / 3261 s
UD-Q8_K_XLStock llama.cpp, new TheRock, cap 3679 / 2 / 4501 s
BF16Stock llama.cpp, new TheRock, cap 3638 / 3 / 4659 s
UD-Q4_K_XLStock llama.cpp, new TheRock, cap 3638 / 3 / 4454 s
ROCmFP4ROCmFPX, new TheRock, cap 3638 / 3 / 4330 s
UD-Q6_K_XLStock llama.cpp, new TheRock, cap 3608 / 2 / 5465 s
UD-Q5_K_XLStock llama.cpp, new TheRock, cap 3577 / 3 / 5443 s
ROCmFPX UltraQuality 7.61 BPWMatched Vulkan, cap 3608 / 2 / 5614 s
ROCmFPX UltraQuality 7.61 BPWROCmFPX, new TheRock, cap 3608 / 2 / 5479 s

Comparable stored Hard Mode evidence

Model / profileScoreComparison note
Laguna S 2.1 Q487Target-only Vulkan; highest stored 2.0.7 Hard Mode result found
DeepSeek V4 Flash API83Clean API rerun
Step 3.7 QualityPlus83Prior best, temperature 1, seed 43
Hy3 FPX-iFP280Native MTP, intended profile
Step 3.7 AGENTKV80Prior best, temperature 1, seed 43
Qwen3.6 35B MoEQuality 7.07 BPW73Prior comparable local result
Qwen3.6 35B vanilla Q870Prior comparable local result
Atlas PR #353 NVFP4 MTP63Same Qwen3.6 27B family
Qwen3.6 35B ROCmFP460Prior comparable local result

The intended-profile rows do not all share one sampler. They are useful cross-model context, not a claim that quantization alone caused every score difference.

Why MTP cap 3 replaced cap 6

The initial speed matrix used a six-token Q4 draft cap. A sweep on the same stock Q4/new-TheRock long request showed that six was over-drafting: acceptance fell to 59.04% and actual generation slowed.

Draft capGeneration tok/sAccepted / draftedAcceptanceDecision
618.6998 / 16659.04%Rejected: excessive drafting
420.9095 / 12675.40%Better, still below target
322.0791 / 10884.26%Selected: fastest
220.0281 / 9189.01%Highest acceptance, lower speed

Acceptance is a diagnostic, not the objective by itself. Cap 3 won because it delivered the highest measured generation throughput while clearing 80% acceptance.

Artifact footprint and effective bits per weight

Vanilla formatGGUF sizeEffective BPWBasis
BF1654.658 GB16.00Two GGUF shards
UD-Q8_K_XL35.776 GB10.48Full artifact
UD-Q6_K_XL26.015 GB7.62Full artifact
UD-Q5_K_XL20.351 GB5.96Full artifact
UD-Q4_K_XL17.909 GB5.24Full artifact
ROCmFP414.817 GB4.34Full artifact
ROCmFPX UltraQuality26.005 GB7.61Full vanilla artifact

Effective BPW is total GGUF bytes × 8 divided by 27,320,697,856 model parameters. It includes GGUF metadata and therefore measures the downloadable artifact, not merely the nominal tensor label. The run logs captured whole-host GTT on this unified-memory APU, not a clean per-process loaded allocation, so no misleading “VRAM loaded” column is reported. Runtime memory would also include KV cache, MTP context, and compute buffers and would vary with context and backend.

What the matrix shows

The initial 4K/128 matrix favored ROCmFP4 on TheRock. It led that matrix's prompt processing at 320.00 tok/s and generation at 38.24 tok/s, while posting the lowest measured TTFP at 11.02 seconds. In the later validated gates, BF16/new-TheRock led PP at 336.69 tok/s and the dedicated Vulkan ROCmFP4 lane led generation at 34.49 tok/s.

The backend winner depends on the quantization path. Vulkan leads generation for BF16, Q8, Q5, Q4, and UltraQuality; TheRock leads Q6 and ROCmFP4. TheRock consistently processes this prompt faster, but that PP advantage does not automatically translate into faster token generation.

UltraQuality is backend-stable in this harness. Matched Vulkan and new-TheRock runs both scored 60 with the same 8/2/5 scenario outcome and identical inference controls. TheRock processed the speed-gate prompt 57.3% faster, while Vulkan generated 6.2% faster; both accepted 90 of 110 MTP draft tokens.

At 60, UltraQuality matches the Q6 Hard Mode result, trails Q8 by seven points, and exceeds Q5 by three. The identical per-scenario backend outcome indicates that Vulkan versus new TheRock did not determine the measured tool-quality score.

The UltraQuality artifact is built from the same vanilla Unsloth BF16 source as the rest of the matrix using an importance matrix and the ranked leave-32 ROCmFPX/Q6_K-splice tensor policy. No fine-tuned derivative weights are included.

Method and scope

Initial speed matrix

  • Same runtime source commit: 5d04ce30c.
  • Vulkan0 via RADV versus ROCm0 via TheRock 7.15.
  • One slot, full GPU offload, flash attention, 8K server context.
  • The matrix prompt generated exactly 128 tokens with cache off and temperature 0.
  • Original rows use MTP cap 6 except ROCmFP4 cap 4; UltraQuality uses the matched 3,525-token cap-3 validation gate.

Hard Mode matrix

  • Tool-Eval Bench 2.0.7 optional Hard Mode, TC-70 through TC-84.
  • Stock formats use stock llama.cpp; only ROCmFPX formats use their specialized runners.
  • TheRock means the new TheRock 7.15 runtime.
  • Matched Qwen controls: embedded Jinja template, reasoning off, temperature 0, top-p 0.95, top-k 20, min-p 0, seed 123, 64K context, and MTP cap 3.
  • Results characterize this Strix Halo host and these exact builds.
UltraQuality artifact identity. The tested vanilla GGUF is 26.005 GB and measures 7.6146 effective BPW against 27,320,697,856 parameters. Its tensor assignment is the ranked leave-32 UltraQuality policy, calibrated with the vanilla model's importance matrix.