Hollow markers & dashed spokes: axis not yet scored
G
42
A
93
U
—
N
—
T
n/d
L
—
E
—
T
35.5
Specification
- Parameters
- 8.19B
- Architecture
- qwen3
- Size on disk
- 4.3 GB
- Quantization
- 4bit
- Format
- MLX
- Class
- Standard
- Reasoning (CoT)
- Yes — emits reasoning tokens
- Internal ID
- M31
- Mean speed
- 84.5 tok/s across suites
- Model card
- huggingface.co
Suite results
Suite Score Avg / 20 Tests tok/s Run
General capability (13-task real-workload suite) 1a 109.2 / 260 8.4 13/13 73.7 run 2026-09-06
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| A1 | Dev/Ops Scripting | 7/20 | 97 | |
| A2 | Dev/Ops Scripting | 14/20 | 101.7 | |
| A3 | Dev/Ops Scripting | 7/20 | 60.7 | |
| A4 | Dev/Ops Scripting | 16/20 | 68.3 | |
| A5 | Dev/Ops Scripting | 9/20 | 64.5 | |
| A6 | Dev/Ops Scripting | 4/20 | 68.5 | |
| B1 | Document Processing | 17/20 | 74.3 | |
| B2 | Document Processing | 0/20 | — | |
| B3 | Document Processing | 0/20 | — | |
| C1 | Content Production | 13/20 | 71.4 | |
| C2 | Content Production | 1/20 | 58.2 | |
| C3 | Content Production | 13/20 | 69.8 | |
| C4 | Content Production | 8/20 | 76 |
Category mean Dev/Ops Scripting 9.5Document Processing 5.7Content Production 8.8
Agentic tool-calling & protocol adherence 1b 148.8 / 160 18.6 8/8 95.2 run 2026-09-06
| Test | Category | Score | tok/s | |
|---|---|---|---|---|
| WA1 | Tool Calling & Protocol Adherence | 20/20 | 95.4 | |
| WA2 | Tool Calling & Protocol Adherence | 20/20 | 99.4 | |
| WA3 | Tool Calling & Protocol Adherence | 20/20 | 96.9 | |
| WA4 | Tool Calling & Protocol Adherence | 20/20 | 97.3 | |
| WA5 | Tool Calling & Protocol Adherence | 20/20 | 98.3 | |
| WB1 | Agent Robustness & State Management | 18/20 | 94.2 | |
| WB2 | Agent Robustness & State Management | 17/20 | 98.9 | |
| WB3 | Agent Robustness & State Management | 14/20 | 81 |
Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 16.3
Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Rows with a ▶ expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.