Crown Citadel GroupCiru Inference Labllm.ciru.ai / research

Crown Citadel Research Report

2-bit Escha, high-quant quality

A two-bit-class Qwen3.6-35B-A3B model leads the released 35B HermesAgent field and remains close to the best higher-quant coding results.

2-bit-class EschaMoE W212.3 GB weightsQwen3.6 35B-A3BReleased-model fieldBenchmarked resultUpdated 2026-08-03

Executive readout

The result is quality density. Escha’s 2-bit-class W2 format leads the released 35B HermesAgent field and stays within a few percentage points of the best higher-quant coding runs.

2-bit W212.3 GB model weights2b gate/up · 3b down · INT8 dense
90 / 100HermesAgent-20rank #1 of 12 model/quant results
90.9%HumanEval+ plus2.4 pp from released best
75.7%MBPP+ plus286/378 · 1.9 pp from released best
29.73%BigCodeBench Hard44/148 · 2.0 pp from released best
87 / 100Tool Eval · 6980/100 on the 15-task hard set

Bottom line. At 12.3 GB of model weights, Escha is 35% smaller than the next-smallest 19.0 GB entry and about one-third the size of a 36.9 GB Q8 model. It posts the best released-model 35B HermesAgent-20 score while landing only 2.4 points behind the best HumanEval+ result, 1.9 behind the best MBPP+ result, and 2.0 behind the best comparable BigCodeBench Hard result.

HermesAgent-20

The headline ranks 12 model-and-quant combinations using each combination’s highest complete 20-scenario score. Bar labels pair quality score with model-weight footprint.

Awesome-pyecharts

What stands out: Escha’s 90/100 exceeds the next-best released-model result at 88 despite using the lowest-bit weight format in the comparison. Bars are labeled with the quant wherever it is identifiable.

Coding and tool use

HumanEval+ and MBPP+ use EvalPlus plus pass@1, BigCodeBench uses pass@1, and Tool Eval reports its native 100-point score. Each chart keeps only the highest score for a model and quant; bar labels include model-weight footprint.

87 / 100Tool Eval · 69
80 / 100Tool Eval hard · 15
95.1%HumanEval base
90.9%HumanEval plus
75.7%MBPP plus
29.73%BigCode Hard
Awesome-pyecharts
Awesome-pyecharts
Awesome-pyecharts

MBPP scoring is complete. MBPP+: 90.7% base / 75.7% plus. Mbpp/84 remains a preserved failure. The score uses the original 378 first samples.

Released 35B quality ledger

Highest scored result for each public 35B model, quant, and suite, plus the Escha results. Weight footprint is the model-weight artifact size; runtime memory additionally includes context/KV cache and backend workspace.

DateFamilySuitePublic releaseQuantWeightsTasksScore
20260702T2evalplushumanevalQwen3.6-35B-A3B · Chadrock v2 seriesROCmFPX MoEQuality 7.07 BPW31.4 GB16490.24
20260702T1evalplusmbppQwen3.6-35B-A3B · Unsloth GGUFQ6_K_XL32.6 GB37877.51
20260702T1evalplushumanevalQwen3.6-35B-A3B · Unsloth GGUFQ6_K_XL32.6 GB16491.46
20260626T0bigcodebenchbigcodebench-hard-instructOrnith 1.0 35BQ4_K_M21.2 GB14827.70
20260607T0bigcodebenchbigcodebench-hard-instructQwen3.6-35B-A3B · Chadrock v2 seriesROCmFP419.0 GB14831.76
20260604T0evalplushumanevalQwen3.6-35B-A3B · Chadrock seriesStrix Lean mixed quant19.0 GB16491.46
20260602T0evalplusmbppQwen3.6-35B-A3B · Chadrock v2 seriesROCmFP419.0 GB37876.98
20260602T0evalplushumanevalQwen3.6-35B-A3B · Chadrock v2 seriesROCmFP419.0 GB16490.85
20260527T0bigcodebenchbigcodebench-hard-instructQwen3.6-35B-A3B · Crown Halo DynamicDynamic mixed quant22.6 GB14829.05
20260523T1evalplushumanevalQwen3.6-35B-A3B · Crown Halo DynamicDynamic mixed quant22.6 GB16489.02
20260523T1evalplusmbppQwen3.6-35B-A3B · Crown Halo DynamicDynamic mixed quant22.6 GB37874.87
2026-08-03hermesagent-20official-20Qwen3.6-35B-A3B · Chadrock v2 seriesROCmFP419.0 GB2083.00
2026-08-03hermesagent-20official-20Escha W2Escha W2 / hybrid 2–3b + INT812.3 GB2090.00
2026-08-03evalplushumanevalEscha W2Escha W2 / hybrid 2–3b + INT812.3 GB16490.85
2026-08-03evalplusmbppEscha W2Escha W2 / hybrid 2–3b + INT812.3 GB37875.66
2026-08-03bigcodebenchbigcodebench-hard-instructEscha W2Escha W2 / hybrid 2–3b + INT812.3 GB14829.73
2026-08-03tool-eval-benchstandard-69Escha W2Escha W2 / hybrid 2–3b + INT812.3 GB6987.00
2026-08-03tool-eval-benchhard-15Escha W2Escha W2 / hybrid 2–3b + INT812.3 GB1580.00
2026-07-27tool-eval-benchstandard-69Ornith 1.0 35BDualView FPX7 + Q8 MTP33.5 GB6989.13
2026-07-27bigcodebenchbigcodebench-hard-instructOrnith 1.0 35BDualView FPX7 + Q8 MTP33.5 GB14827.70
2026-07-27evalplushumanevalOrnith 1.0 35BDualView FPX7 + Q8 MTP33.5 GB16493.29
2026-07-26hermesagent-20official-20Ornith 1.0 35BDualView FPX7 + Q8 MTP33.5 GB2087.00
2026-07-26hermesagent-20official-20Ornith 1.0 35BQ7S8 hybrid32.6 GB2088.00
2026-07-25hermesagent-20official-20Ornith 1.0 35BQ836.9 GB2088.00
2026-07-21hermesagent-20official-20Qwen3.6-35B-A3BQ836.9 GB2069.00
2026-07-02hermesagent-20official-20Qwen3.6-35B-A3B · Chadrock v2 seriesROCmFPX MoEQuality 7.07 BPW31.4 GB2082.00
2026-07-02hermesagent-20official-20Qwen3.6-35B-A3B · Unsloth GGUFQ6_K_XL32.6 GB2071.00
2026-06-26hermesagent-20official-20Ornith 1.0 35BQ4_K_M21.2 GB2082.00
2026-06-05hermesagent-20official-20Qwen3.6-35B-A3B · Chadrock seriesROCmFP419.0 GB2079.00
2026-06-05hermesagent-20official-20Qwen3.6-35B-A3B · Crown Halo DynamicDynamic mixed quant22.6 GB2070.00
2026-06-05hermesagent-20official-20Qwen3.6-35B-A3B · Chadrock seriesStrix Lean mixed quant19.0 GB2073.00
2026-05-28bfclbfcl-v4-all_scoringQwen3.6-35B-A3B · Crown Halo DynamicDynamic mixed quant22.6 GB521716.86
2026-05-28bfclbfcl-v4-non_liveQwen3.6-35B-A3B · Crown Halo DynamicDynamic mixed quant22.6 GB139083.00
2026-05-28livecodebenchrelease_latest-codegeneration-2025-01-01Qwen3.6-35B-A3B · Crown Halo DynamicDynamic mixed quant22.6 GB18236.81