vLLM Dominates llama.cpp Under Concurrency
At five simultaneous Ling 3.0 Flash requests, native vLLM delivers 2.85× aggregate decode throughput and clears the batch 46.60 seconds before AtomicChat llama.cpp.
Run the RaceCrown Citadel Research Index
An index of Ciru Inference Lab research notes, benchmark reports, quantization comparisons, tool-loop investigations, and serving experiments published on llm.ciru.ai.
/research./reports.These are the current LLM-related pages served directly under /research/. The previous /research/ document has been preserved as the Hermes tool-loop findings page.
At five simultaneous Ling 3.0 Flash requests, native vLLM delivers 2.85× aggregate decode throughput and clears the batch 46.60 seconds before AtomicChat llama.cpp.
Run the RaceInteractive research edition tracing the Ling 3.0 Flash ROCm campaign from its initial serving baseline through kernel, concurrency, MTP, and promotion-gate evidence.
Open ReportInteractive Python/ECharts comparison of Escha's 2-bit-class quality across HermesAgent-20, EvalPlus, BigCodeBench, and Tool Eval against released 35B model lines, with quants identified from the artifacts.
Open ReportOne-table 4K/128 served comparison across BF16, Q8, Q6, Q5, Q4, ROCmFP4, and vanilla ROCmFPX UltraQuality on AMD Strix Halo.
Open MatrixSuccessful target-only and DSpark benchmarks on the RTX 4080 SUPER, with context scaling, throughput, VRAM, and reproducible launch settings.
Open BenchmarksInteractive implementation guide to Codebook10, dual-scale blocks, MSE-optimal UE4M3 scale search, and the tensor-aware STRIX recipe.
Open GuideROCmFP2 codebook and AMD dot-path report with a validated HY3 imatrix build, full Tool‑Eval 88, and a completed disk-only target+draft cache proof.
Open ReportSide-by-side synthesis of the GROK 4.5, GLM 5.2, and SOL Ultra audits, with preserved source reports, shared findings, material disagreements, and a combined retest protocol.
Open ComparisonResearch experiment on agent-world design, public contracts, trace evidence, and benchmark behavior across HermesAgent and BigCodeBench workflows.
Open PageComparison of Atlas NVFP4 MTP-K2 and BF16 GGUF MTP inference, including throughput, memory, wall time, and HermesAgent-20 evidence.
Open PageComparison of Chadrock ROCmFP4 against upstream Unsloth Q4/Q6 paths, including MTP and non-MTP serving behavior.
Open PageThe original Hermes tool-loop findings report, moved out of index.html so the main research route can serve this index.
Research report on what changed from the 7.37 BPW quality model, with local Strix Halo evidence and quality/speed implications.
Open PageVisual speed report explaining the measured serving advantages of ROCmFP4 and ROCmFPX paths over stock Unsloth GGUF lanes.
Open PageImplementation note explaining the StepFun tool-call message loop, why eval harnesses can misformat it, and how to patch templates and clients.
Open PageGallery and run audit for 32 simultaneous local artifact workers, including final outputs, cycle snapshots, timing, completion state, and system-memory evidence.
Open GalleryTechnical and commercial assessment of an Atlas ROCmFP4 path on AMD Strix Halo, including likely integration points, engineering scope, risks, and validation gates.
Open ReportInfrastructure and energy research connected to local compute, data-center systems, and practical deployment economics.
Engineering and economic evaluation of Wankel-style engines, organic Rankine cycles, and direct heat reuse for recovering value from compute waste heat.
Open ReportThese LLM research reports currently live under /reports/, so the links below point outside the /research/ folder while keeping the main research index complete.
Standalone comparison of phone-targeted QVAC and llama.cpp GGUF results with score, VRAM, latency, token, throughput, category, and scenario charts.
Open ReportResearch note proposing a context-conditioned speculative decoding policy for llama.cpp serving.
Open ReportBenchmark comparison of Qwable ROCmFP6 and Unsloth Q6 on the Strix Halo local benchmark host.
Open ReportResearch report on Q6-level quality with served-speed gains for the ROCmFP6 Strix quality lane.
Open Report