Escha-W2 (2-bit, 10.2 GB) vs Bonsai PQ2 (ternary, 7.2 GB) on official Google IFEval — 541 prompts, thinking off — with a 4.00bpw EXL3 build as the reference, not a contestant.
PQ2 is 30% smaller than Escha and about half the 4bpw reference — 27B plus 32k context on one 16 GB 4080 SUPER. PTQ1 is Bonsai's smaller sibling (speed-only entry, no IFEval run).
PQ2 decodes 1.35× faster than Escha. The 4bpw reference isn't in this fight: 8-way parallel on a 5090 it cleared all 541 prompts in 6.8 min vs ~59–64 min sequential.
Escha: SGLang 1.2.2, fp8 KV, greedy temp 0, sequential. PQ2 / PTQ1: Prism llama.cpp, greedy, sequential — all on one RTX 4080 SUPER 16 GB. Reference: r0b0tlab EXL3 4.00bpw + DFlash2, Aider sampler temp 0.7 / top_p 0.8 / top_k 20, 8-way on RTX 5090. Official Google IFEval, 541 prompts, max 2048 tokens, thinking off. Decode = tg128 after 128-token prefill; 4bpw ≈13.5 GB = 27B × 4 bits.
| Metric | r0b0t EXL3 · reference | Escha-W2 | Bonsai PQ2 | Escha − PQ2 |
|---|---|---|---|---|
| Strict prompt-level | 84.1% | 81.9% | 79.5% | +2.4 pts |
| Loose prompt-level | 87.2% | 84.5% | 82.6% | +1.9 pts |
| Strict instruction-level | 89.0% | 87.3% | 86.5% | +0.8 pts |
| Loose instruction-level | 91.5% | 89.3% | 88.5% | +0.8 pts |
| Instruction | n | Escha | PQ2 | 4bpw ref |
|---|
Bonsai: llama-bench. Escha: HTTP. The 4bpw reference was not ladder-benched on the 4080; its Aider panel on the 5090 measured 276 aggregate output tok/s at 8 workers.
| Prefill | Escha tg | PQ2 tg | PTQ1 tg |
|---|---|---|---|
| 128 | 41.9 | 56.5 | 62.2 |
| 2k | 42.2 | 55.7 | 59.9 |
| 4k | 41.8 | 55.0 | 59.7 |
| 8k | 41.3 | 52.8 | 56.4 |
| 16k | 31.7 | 50.6 | 54.8 |
| 32k | 29.5 | 46.0 | 49.0 |
| Prefill | Escha pp | PQ2 pp | PTQ1 pp |
|---|---|---|---|
| 128 | 762 | 1393 | 902 |
| 2k | 1612 | 1529 | 925 |
| 4k | 1621 | 1573 | 938 |
| 8k | 1609 | 1636 | 922 |
| 16k | 1268 | 1655 | 903 |
| 32k | 1326 | 1545 | 839 |
| ub | PTQ1 pp/tg | PQ2 pp/tg |
|---|---|---|
| 256 | 921 / 67.6 | 1945 / 62.1 |
| 512 | 916 / 67.7 | 1933 / 63.1 |
| 1024 | 903 / 68.5 | 1857 / 62.2 |
| 2048 | 950 / 68.7 | 1758 / 62.2 |
The competition is between the two hyper-compressed builds — Escha-W2 (2-bit, 10.2 GB) and Bonsai PQ2 (ternary, 7.2 GB) — both running solo on a 16 GB 4080 SUPER. The r0b0t EXL3 4.00bpw build is the reference: a normal, non-hyper-compressed quant showing where an uncrushed model lands (dashed marker). Escha takes quality on every aggregate metric (+2.4 strict prompt); PQ2 takes decode speed (×1.35) and is 30% smaller. Headline is strict prompt-level IFEval — every instruction in a prompt must pass. Speed ladders are 4080-only; the reference wall-clock (6.8 min) came from 8-way concurrency on a 5090.