← Leaderboard

Ornith-1.0-35B 4bit MLX

Candidate Rank #19 of 37 · 3/8 GAUNTLET progress
Generalist — 85.5 / 100 G Agentic — 98 / 100 A Understanding — pending U Needle — pending N Thinking — pending T Live — pending L Engineering — pending E Throughput — 48.3 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
85.5
A
98
U
N
T
n/d
L
E
T
48.3

Specification

Parameters
35B total (MoE; active-per-token not confirmed in model card)
Architecture
qwen3_5_moe
Size on disk
20.4 GB
Quantization
4bit
Format
MLX
Class
Standard
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M28
Mean speed
115.6 tok/s across suites
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests tok/s Run
General capability (13-task real-workload suite) 1a 222.3 / 260 17.1 13/13 118.9 run 2026-08-22
Test Category Score tok/s
A1 Dev/Ops Scripting 14/20 117.5
A2 Dev/Ops Scripting 14/20 124
A3 Dev/Ops Scripting 19/20 121.6
A4 Dev/Ops Scripting 19/20 121.4
A5 Dev/Ops Scripting 19/20 128.3
A6 Dev/Ops Scripting 20/20 121.1
B1 Document Processing 19/20 123.8
B2 Document Processing 20/20 121.5
B3 Document Processing 0/20
C1 Content Production 20/20 112.2
C2 Content Production 19/20 109.7
C3 Content Production 19/20 111.5
C4 Content Production 20/20 113.7

Category mean Dev/Ops Scripting 17.5Document Processing 13Content Production 19.5

Agentic tool-calling & protocol adherence 1b 156.8 / 160 19.6 8/8 112.3 run 2026-08-22
Test Category Score tok/s
WA1 Tool Calling & Protocol Adherence 20/20 109.3
WA2 Tool Calling & Protocol Adherence 20/20 99.3
WA3 Tool Calling & Protocol Adherence 20/20 114.1
WA4 Tool Calling & Protocol Adherence 20/20 83
WA5 Tool Calling & Protocol Adherence 20/20 115.3
WB1 Agent Robustness & State Management 20/20 127.3
WB2 Agent Robustness & State Management 20/20 126.8
WB3 Agent Robustness & State Management 17/20 123.4

Category mean Tool Calling & Protocol Adherence 20Agent Robustness & State Management 19

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Rows with a expand to the per-test breakdown. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.