Crown Citadel Group Ciru Inference Lab llm.ciru.ai / research Benchmarks

Crown Citadel Research Index

LLM Research Pages

An index of Ciru Inference Lab research notes, benchmark reports, quantization comparisons, tool-loop investigations, and serving experiments published on llm.ciru.ai.

22LLM research pages in /research.
4Related LLM reports in /reports.
1Broader infrastructure and energy research page.
1Former index preserved as its own page.

Research Pages

These are the current LLM-related research pages. The previous /research/ document has been preserved as the Hermes tool-loop findings page.

Qwen3.8 Flash5 passes × 84 casesAnimated race

Tool Eval Suite 3.8 Flash

Seven entrants and 2,940 scored attempts: recorded tool-call races, core and hard scores, wall time, native PP/TG, and complete runtime configurations.

Open Suite
Strata v0.1.40Strix HaloFull source

Strata Flash

Recorded Flash race and comparison: decode, zero-replay prefill, recommended tool and Hermes scores, numerical fidelity, memory, and audited benchmark source on GitHub.

Open Comparison
Flash comparisonQwen APIAnimated race

Flash Race

Recorded tool and Hermes-20 races across eight local stacks and Qwen3.8 Flash API.

Open Race
Qwen3.8 FlashAMD Strix Halo8 stacks · Open evidence

The Ultimate Qwen Flash 3.8 Strix Showdown

Eight deployments compared across tool use, Hermes agent tasks, decode, prefill, numerical fidelity, and memory. Interactive races replay saved first attempts, with public test source, sanitized run logs, protocol locks, and checksums on GitHub.

Open the Showdown
Qwen3.8 27BIFEvalRTX 4080 SUPER

Qwen3.8-27B: Escha-W2 vs Bonsai PQ2

Head-to-head comparison of two compressed 27B builds on a 16 GB GPU, covering 541 IFEval prompts, decode and prefill speed, and model size, with a 4.00 BPW reference.

Open Report
Ornith 1.5Strix Halo256K context

Ornith 1.5 CIRU HALO AGENT

Interactive comparisons against Q4_K_XL and ROCmFP4: concurrency, prefill, cached long-context returns, prose, BF16 agreement, EvalScope, and repeated Hermes agent runs. Includes measured limitations and unavailable cached C8 baseline results.

Open Benchmarks
Qwen3.8 v3ROCm10 / VulkanInteractive charts

Qwen3.8 Flash v3: serving, quality & wall time

V3 serving at 4K, 64K and full capacity; mixed-task wall times, Hermes agent outcomes and coding regression checks. Includes historical v2 comparisons and EvalScope timings.

Explore Comparison
Qwen3.8 Flash NextAMD Ryzen AI HaloIndependent Research

How Ciru Turned AMD Ryzen AI Halo Into a 128 GB AI Lab

Independent, community-driven Ciru report on a route-aware Q5-class target, exact SSD-backed FP8 learned memory, device-resident MTP runtime, and measured performance and quality evidence.

Open Report
Ling 3.0 FlashConcurrencyAnimated Race

vLLM Dominates llama.cpp Under Concurrency

At five simultaneous Ling 3.0 Flash requests, native vLLM delivers 2.85× aggregate decode throughput and clears the batch 46.60 seconds before AtomicChat llama.cpp.

Run the Race
Ling 3.0 FlashStrix HaloROCm

One APU, One Week: A 97.5× Serving Turnaround

Interactive research edition tracing the Ling 3.0 Flash ROCm campaign from its initial serving baseline through kernel, concurrency, MTP, and promotion-gate evidence.

Open Report
2-bit quality35BEscha

Escha W2 vs Released 35B Models

Interactive Python/ECharts comparison of Escha's 2-bit-class quality across HermesAgent-20, EvalPlus, BigCodeBench, and Tool Eval against released 35B model lines, with quants identified from the artifacts.

Open Report
Qwen3.6 27BVulkan vs TheRockMTP

Vanilla Qwen3.6 27B MTP Backend Matrix

One-table 4K/128 served comparison across BF16, Q8, Q6, Q5, Q4, ROCmFP4, and vanilla ROCmFPX UltraQuality on AMD Strix Halo.

Open Matrix
RTX 4080 SUPERBonsai 27BDSpark

Ternary Bonsai 27B on a 16 GB GPU

Successful target-only and DSpark benchmarks on the RTX 4080 SUPER, with context scaling, throughput, VRAM, and reproducible launch settings.

Open Benchmarks
ROCmFP4InteractiveQuantization

Why ROCmFP4 Beats Traditional Q4

Interactive implementation guide to Codebook10, dual-scale blocks, MSE-optimal UE4M3 scale search, and the tensor-aware STRIX recipe.

Open Guide
FPX-IFP2ValidatedSSD Cache

FPX-IFP2: Four Numbers, Ten Bytes

ROCmFP2 codebook and AMD dot-path report with a validated HY3 imatrix build, full Tool‑Eval 88, and a completed disk-only target+draft cache proof.

Open Report
Step 3.73 AuditsCandidate Review

Step 3.7 Candidate Audit Comparison

Side-by-side synthesis of the GROK 4.5, GLM 5.2, and SOL Ultra audits, with preserved source reports, shared findings, material disagreements, and a combined retest protocol.

Open Comparison
AgentsWorld ModelHermesAgent

AgentWorld as a Public-Contract World Model

Research experiment on agent-world design, public contracts, trace evidence, and benchmark behavior across HermesAgent and BigCodeBench workflows.

Open Page
ServingQwen3.6 27BMTP

Atlas vs BF16 Qwen3.6 27B

Comparison of Atlas NVFP4 MTP-K2 and BF16 GGUF MTP inference, including throughput, memory, wall time, and HermesAgent-20 evidence.

Open Page
ROCmFP4UnslothMTP

Chadrock ROCmFP4 vs Unsloth Upstream Q4/Q6

Comparison of Chadrock ROCmFP4 against upstream Unsloth Q4/Q6 paths, including MTP and non-MTP serving behavior.

Open Page
Tool LoopFailure AnalysisHermes

How a Read-Only Tool Call Became a Self-Reinforcing Loop

The original Hermes tool-loop findings report, moved out of index.html so the main research route can serve this index.

Open Page
QuantizationQwable7.61 BPW

Ultra Quality 7.61 BPW

Research report on what changed from the 7.37 BPW quality model, with local Strix Halo evidence and quality/speed implications.

Open Page
ROCmFPXSpeedUnsloth

Why ROCmFP4 / ROCmFPX Runs Faster Than Stock Unsloth GGUF

Visual speed report explaining the measured serving advantages of ROCmFP4 and ROCmFPX paths over stock Unsloth GGUF lanes.

Open Page
StepFunTool CallingTemplate Fix

StepFun Tool Calling: Why Eval Loops Happen and How To Fix Them

Implementation note explaining the StepFun tool-call message loop, why eval harnesses can misformat it, and how to patch templates and clients.

Open Page
AMD Strix Halo64 BuildsInteractive Gallery

64 Builds, Two AMD Desktops

Explore every playable app and game from a dual Strix Halo run: 32 simultaneous builds on each machine, live speed telemetry, and a complete thumbnail wall.

Open Showcase
Gemma4 26B32 SlotsArtifacts

Gemma4 26B QAT MTP 32-Slot Artifact Run

Gallery and run audit for 32 simultaneous local artifact workers, including final outputs, cycle snapshots, timing, completion state, and system-memory evidence.

Open Gallery
AtlasROCmFP4Feasibility

Feasibility of Porting ROCmFP4 to Atlas

Now includes measured native-HIP Qwen3.6 27B NVFP4 serving on AMD Strix Halo: 256.5 tok/s prefill, 16.68 tok/s decode, and a 63/100 hard tool-use score.

Open Report

Broader Research

Infrastructure and energy research connected to local compute, data-center systems, and practical deployment economics.

Waste HeatData CentersEnergy Systems

Recovering Heat from Bitcoin Miners and Data-Center Chips

Engineering and economic evaluation of Wankel-style engines, organic Rankine cycles, and direct heat reuse for recovering value from compute waste heat.

Open Report

Related Reports

These LLM research reports currently live under /reports/, so the links below point outside the /research/ folder while keeping the main research index complete.

Phone GGUFQVACTool-Eval

Phone Tool-Eval Scoreboard

Standalone comparison of phone-targeted QVAC and llama.cpp GGUF results with score, VRAM, latency, token, throughput, category, and scenario charts.

Open Report
Speculative Decodingllama.cppPolicy

Dynamic Draft Proposal

Research note proposing a context-conditioned speculative decoding policy for llama.cpp serving.

Open Report
HermesAgent-20ROCmFP6Q6

HermesAgent-20: ROCmFP6 vs Unsloth Q6

Benchmark comparison of Qwable ROCmFP6 and Unsloth Q6 on the Strix Halo local benchmark host.

Open Report
ROCmFP6QualityStrix Halo

ROCmFP6 Strix Quality

Research report on Q6-level quality with served-speed gains for the ROCmFP6 Strix quality lane.

Open Report