← Leaderboard

Qwen3.6-27B Dense 8bit MLX

Keeper RAN THE GAUNTLET Rank #3 of 33 · 8/8 GAUNTLET progress
Generalist — 91 / 100 G Agentic — 100 / 100 A Understanding — 93.2 / 100 U Needle — 99 / 100 N Thinking — 97.2 / 100 T Live — 68.3 / 100 L Engineering — 91.5 / 100 E Throughput — 8.2 / 100 T
G
91
A
100
U
93.2
N
99
T
97.2
L
68.3
E
91.5
T
8.2

Specification

Parameters
27B dense
Architecture
qwen3_5
Size on disk
29.53 GB
Quantization
8bit
Format
MLX
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
M7
Mean speed
13.4 tok/s across suites
Stall census
2 stalls in 71 observed tests (2.8%)
Reasoning appetite
2,881 tokens mean · 16,310 max
Model card
lmstudio.ai

Suite results

Suite Score Avg / 20 Tests tok/s
General capability (13-task real-workload suite) 1a 236.6 / 260 18.2 13/13 15.2
Agentic tool-calling & protocol adherence 1b 160 / 160 20 8/8 16.6
Coding depth 1c 109.8 / 120 18.3 6/6 13.9
Doc/OCR vision 1d 153.6 / 160 19.2 8/8 14.1
Doc/OCR — real-degraded tier 1d2 70 / 80 17.5 4/4 13.2
Content-production depth 1e 115.8 / 120 19.3 6/6 16.4
Long-context retrieval & synthesis 1g 118.2 / 120 19.7 6/6 10.5
Long-context multi-needle (MRCR) 1g2 60 / 60 20 3/3 9.2
Live one-shot builds (runtime-verified) 1h 77 / 80 19.3 4/4
Production replay (real agent workload) 1i 141.6 / 240 11.8 12/12 11.5

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.