Ciru
CiruINFERENCE LAB
QWEN3.8 FLASH · TOOL EVAL · OCTOBER 2026

Tool Eval Suite 3.8 Flash

Benchmark clock 00:00.00 replay speed ×10
Speed
start0:00

passpartial creditfailvaries across passesin progress

Race feed

    Recorded points · all entrants

      7ENTRANTS
      5PASSES EACH
      69 + 15CASES PER PASS
      2,940SCORED ATTEMPTS
      OFFTHINKING
      FIVE-PASS AVERAGES

      Results at a glance

      Seeds 123–127 · concurrency 1
      EntrantCore 69Hard 15CombinedSuite wallNative TGNative PP

      Scores use earned points ÷ available points. Wall time covers the scored suite. Native PP/TG use each runtime’s processed/accepted token counters; profile caching is enabled. Hosted native timings are unavailable.

      QUALITY · TIME · THROUGHPUT

      The comparison

      All seven entrants

      Core, hard & combined

      points % · higher is better

      Full suite wall time

      minutes · five pass range + mean

      Scores across five passes

      Quality & completion time

      combined points % × suite minutes

      Native token generation

      accepted tok/s · weighted across five passes

      Native prompt processing

      processed tok/s · tool workload with cache

      Category scores

      A–O core · P hard · five-pass mean

      Charts preserve each entrant’s native accounting and serving settings. PP here describes the tool-call workload, rather than a cold-prefill sweep.

      EVERY CHECKPOINT

      84 cases, five passes

      Each tile shows mean points out of 2. Select a checkpoint for its five recorded outcomes and times.

      MODEL CARDS · RUNTIMES · EFFECTIVE SETTINGS

      Configuration ledger

      Download complete settings ↓
      HardwareAMD Ryzen AI MAX+ 395 · Radeon 8060S · gfx1151 · 128 GiB unified memory · NixOS
      ProtocolTool Eval 2.0.7 · 69 core + 15 hard · 5 passes · serial requests · thinking off
      Context & output262,144 evaluation context · output follows each recorded serving configuration
      Budget & initialization32 turns · 1,800 s elapsed timeout · no scored retries · one unscored one-token transport probe per pass
      SamplingEach pinned model card’s non-thinking recommendation; base-card inheritance recorded per entrant
      Scoring & replayOriginal scorer · pass 2 / partial 1 / fail 0 · replay sums scored-case durations, excluding the probe and gaps between cases

      Hardware and serving details are recorded per entrant. Dataset fixture date: 20 March 2026. Harness revision: 6da446e695e8b9119b9fd7fcca8b3fd8d0092f47.