ESCHA VS BONSAI — IFEVAL — 27B
Escha-W2vsBonsai PQ2 IFEval · n=541 18 Sep 2026

Two ways to crush 27B.One 16 GB card.

Escha-W2 (2-bit, 10.2 GB) vs Bonsai PQ2 (ternary, 7.2 GB) on official Google IFEval — 541 prompts, thinking off — with a 4.00bpw EXL3 build as the reference, not a contestant.

Quality lead · +2.4 pts
Escha-W2
81.9%
Strict prompt-level IFEval
10.2GB
Weights
42t/s
Decode
~64min
541 prompts
VS
Speed lead · ×1.35 decode
Bonsai PQ2
79.5%
Strict prompt-level IFEval
7.2GB
Weights
57t/s
Decode
~59min
541 prompts
The reference — r0b0t EXL3 4.00bpw · RTX 5090 84.1 strict · 87.2 loose · 89.0 inst.
Weights on disk GB
PTQ1
5.95 GB
PQ2
7.21 GB
Escha
10.2 GB
4bpw ref
≈13.5 GB

PQ2 is 30% smaller than Escha and about half the 4bpw reference — 27B plus 32k context on one 16 GB 4080 SUPER. PTQ1 is Bonsai's smaller sibling (speed-only entry, no IFEval run).

Decode speed tok/s · short ctx
PTQ1
62.2 t/s
PQ2
56.5 t/s
Escha
41.9 t/s
4bpw ref
276 t/s agg.

PQ2 decodes 1.35× faster than Escha. The 4bpw reference isn't in this fight: 8-way parallel on a 5090 it cleared all 541 prompts in 6.8 min vs ~59–64 min sequential.

Escha: SGLang 1.2.2, fp8 KV, greedy temp 0, sequential. PQ2 / PTQ1: Prism llama.cpp, greedy, sequential — all on one RTX 4080 SUPER 16 GB. Reference: r0b0tlab EXL3 4.00bpw + DFlash2, Aider sampler temp 0.7 / top_p 0.8 / top_k 20, 8-way on RTX 5090. Official Google IFEval, 541 prompts, max 2048 tokens, thinking off. Decode = tg128 after 128-token prefill; 4bpw ≈13.5 GB = 27B × 4 bits.

01

IFEval numbers

◂ marks best of the two compressed quants
Metricr0b0t EXL3 · referenceEscha-W2Bonsai PQ2Escha − PQ2
Strict prompt-level84.1%81.9%79.5%+2.4 pts
Loose prompt-level87.2%84.5%82.6%+1.9 pts
Strict instruction-level89.0%87.3%86.5%+0.8 pts
Loose instruction-level91.5%89.3%88.5%+0.8 pts
Escha-W2
10.2 GB · 2.0-bit
SGLang 1.2.2 · fp8 KV · Triton · graphs on · 32k ctx · RTX 4080 SUPER · sequential
Bonsai PQ2_0
7.21 GB · ternary
Prism llama.cpp · -b 2048 -ub 512 · 32k ctx · RTX 4080 SUPER · sequential
r0b0tlab EXL3 · ref
≈13.5 GB · 4.00bpw
+ DFlash2 · ExLlamaV3 cq3 · 32 slots / 262k pool · Aider 8 workers · RTX 5090 (Dunamis)
02

By category

strict · instruction-level
Escha PQ2 4bpw reference
03

By subcategory

strict · sorted by Escha
InstructionnEschaPQ24bpw ref
04

Decode speed

tg128 after prefill · tok/s · 4080 SUPER

Bonsai: llama-bench. Escha: HTTP. The 4bpw reference was not ladder-benched on the 4080; its Aider panel on the 5090 measured 276 aggregate output tok/s at 8 workers.

Escha PQ2 PTQ1
PrefillEscha tgPQ2 tgPTQ1 tg
12841.956.562.2
2k42.255.759.9
4k41.855.059.7
8k41.352.856.4
16k31.750.654.8
32k29.546.049.0
05

Prefill speed

pp · tok/s · 4080 SUPER
Escha PQ2 PTQ1
PrefillEscha ppPQ2 ppPTQ1 pp
1287621393902
2k16121529925
4k16211573938
8k16091636922
16k12681655903
32k13261545839
06

Fine print

Bonsai ubatch @ 4k prefill
ubPTQ1 pp/tgPQ2 pp/tg
256921 / 67.61945 / 62.1
512916 / 67.71933 / 63.1
1024903 / 68.51857 / 62.2
2048950 / 68.71758 / 62.2
How to read the poster

The competition is between the two hyper-compressed builds — Escha-W2 (2-bit, 10.2 GB) and Bonsai PQ2 (ternary, 7.2 GB) — both running solo on a 16 GB 4080 SUPER. The r0b0t EXL3 4.00bpw build is the reference: a normal, non-hyper-compressed quant showing where an uncrushed model lands (dashed marker). Escha takes quality on every aggregate metric (+2.4 strict prompt); PQ2 takes decode speed (×1.35) and is 30% smaller. Headline is strict prompt-level IFEval — every instruction in a prompt must pass. Speed ladders are 4080-only; the reference wall-clock (6.8 min) came from 8-way concurrency on a 5090.