Crown Citadel Group Ciru Inference Lab llm.ciru.ai / research Benchmarks

Crown Citadel Research Index

LLM Research Pages

An index of Ciru Inference Lab research notes, benchmark reports, quantization comparisons, tool-loop investigations, and serving experiments published on llm.ciru.ai.

17LLM research pages in /research.
4Related LLM reports in /reports.
1Broader infrastructure and energy research page.
1Former index preserved as its own page.

Research Pages

These are the current LLM-related pages served directly under /research/. The previous /research/ document has been preserved as the Hermes tool-loop findings page.

Ling 3.0 FlashConcurrencyAnimated Race

vLLM Dominates llama.cpp Under Concurrency

At five simultaneous Ling 3.0 Flash requests, native vLLM delivers 2.85× aggregate decode throughput and clears the batch 46.60 seconds before AtomicChat llama.cpp.

Run the Race
Ling 3.0 FlashStrix HaloROCm

One APU, One Week: A 97.5× Serving Turnaround

Interactive research edition tracing the Ling 3.0 Flash ROCm campaign from its initial serving baseline through kernel, concurrency, MTP, and promotion-gate evidence.

Open Report
2-bit quality35BEscha

Escha W2 vs Released 35B Models

Interactive Python/ECharts comparison of Escha's 2-bit-class quality across HermesAgent-20, EvalPlus, BigCodeBench, and Tool Eval against released 35B model lines, with quants identified from the artifacts.

Open Report
Qwen3.6 27BVulkan vs TheRockMTP

Vanilla Qwen3.6 27B MTP Backend Matrix

One-table 4K/128 served comparison across BF16, Q8, Q6, Q5, Q4, ROCmFP4, and vanilla ROCmFPX UltraQuality on AMD Strix Halo.

Open Matrix
RTX 4080 SUPERBonsai 27BDSpark

Ternary Bonsai 27B on a 16 GB GPU

Successful target-only and DSpark benchmarks on the RTX 4080 SUPER, with context scaling, throughput, VRAM, and reproducible launch settings.

Open Benchmarks
ROCmFP4InteractiveQuantization

Why ROCmFP4 Beats Traditional Q4

Interactive implementation guide to Codebook10, dual-scale blocks, MSE-optimal UE4M3 scale search, and the tensor-aware STRIX recipe.

Open Guide
FPX-IFP2ValidatedSSD Cache

FPX-IFP2: Four Numbers, Ten Bytes

ROCmFP2 codebook and AMD dot-path report with a validated HY3 imatrix build, full Tool‑Eval 88, and a completed disk-only target+draft cache proof.

Open Report
Step 3.73 AuditsCandidate Review

Step 3.7 Candidate Audit Comparison

Side-by-side synthesis of the GROK 4.5, GLM 5.2, and SOL Ultra audits, with preserved source reports, shared findings, material disagreements, and a combined retest protocol.

Open Comparison
AgentsWorld ModelHermesAgent

AgentWorld as a Public-Contract World Model

Research experiment on agent-world design, public contracts, trace evidence, and benchmark behavior across HermesAgent and BigCodeBench workflows.

Open Page
ServingQwen3.6 27BMTP

Atlas vs BF16 Qwen3.6 27B

Comparison of Atlas NVFP4 MTP-K2 and BF16 GGUF MTP inference, including throughput, memory, wall time, and HermesAgent-20 evidence.

Open Page
ROCmFP4UnslothMTP

Chadrock ROCmFP4 vs Unsloth Upstream Q4/Q6

Comparison of Chadrock ROCmFP4 against upstream Unsloth Q4/Q6 paths, including MTP and non-MTP serving behavior.

Open Page
Tool LoopFailure AnalysisHermes

How a Read-Only Tool Call Became a Self-Reinforcing Loop

The original Hermes tool-loop findings report, moved out of index.html so the main research route can serve this index.

Open Page
QuantizationQwable7.61 BPW

Ultra Quality 7.61 BPW

Research report on what changed from the 7.37 BPW quality model, with local Strix Halo evidence and quality/speed implications.

Open Page
ROCmFPXSpeedUnsloth

Why ROCmFP4 / ROCmFPX Runs Faster Than Stock Unsloth GGUF

Visual speed report explaining the measured serving advantages of ROCmFP4 and ROCmFPX paths over stock Unsloth GGUF lanes.

Open Page
StepFunTool CallingTemplate Fix

StepFun Tool Calling: Why Eval Loops Happen and How To Fix Them

Implementation note explaining the StepFun tool-call message loop, why eval harnesses can misformat it, and how to patch templates and clients.

Open Page
Gemma4 26B32 SlotsArtifacts

Gemma4 26B QAT MTP 32-Slot Artifact Run

Gallery and run audit for 32 simultaneous local artifact workers, including final outputs, cycle snapshots, timing, completion state, and system-memory evidence.

Open Gallery
AtlasROCmFP4Feasibility

Feasibility of Porting ROCmFP4 to Atlas

Technical and commercial assessment of an Atlas ROCmFP4 path on AMD Strix Halo, including likely integration points, engineering scope, risks, and validation gates.

Open Report

Broader Research

Infrastructure and energy research connected to local compute, data-center systems, and practical deployment economics.

Waste HeatData CentersEnergy Systems

Recovering Heat from Bitcoin Miners and Data-Center Chips

Engineering and economic evaluation of Wankel-style engines, organic Rankine cycles, and direct heat reuse for recovering value from compute waste heat.

Open Report

Related Reports

These LLM research reports currently live under /reports/, so the links below point outside the /research/ folder while keeping the main research index complete.

Phone GGUFQVACTool-Eval

Phone Tool-Eval Scoreboard

Standalone comparison of phone-targeted QVAC and llama.cpp GGUF results with score, VRAM, latency, token, throughput, category, and scenario charts.

Open Report
Speculative Decodingllama.cppPolicy

Dynamic Draft Proposal

Research note proposing a context-conditioned speculative decoding policy for llama.cpp serving.

Open Report
HermesAgent-20ROCmFP6Q6

HermesAgent-20: ROCmFP6 vs Unsloth Q6

Benchmark comparison of Qwable ROCmFP6 and Unsloth Q6 on the Strix Halo local benchmark host.

Open Report
ROCmFP6QualityStrix Halo

ROCmFP6 Strix Quality

Research report on Q6-level quality with served-speed gains for the ROCmFP6 Strix quality lane.

Open Report