Crown Citadel Group Ciru Inference Lab llm.ciru.ai / tool-eval
Tool-Eval Bench 2.0.7 / Published model index

Tool Eval Index

Pick any two published Tool-Eval runs to compare the full-suite score, pass profile, speed, weak sections, and per-section deltas side by side.

Score delta 0

Score And Runtime

Left Right

Outcome Mix

Pass Partial Fail

Section Scores

Left Right

Head To Head

Metric Left Right Delta

Standalone Scoreboards

Linked artifact
Phone class / generated 2026-07-07

Phone Tool-Eval Scoreboard

QVAC direct runs and llama.cpp proxy runs compared across the full 69-scenario Tool-Eval suite plus short QVAC smoke runs.

Best full score: 89 123 / 138 points QVAC + proxy
Open Scoreboard

All Tool Eval Results

Select any two above