Local-LLM proving ground

Proving local agents before they act.

GAUNTLET Bench runs local models through real workloads — agentic, live, verified — and publishes the evidence.

All scores measured on Apple M5 Max, 128 GB unified memory, LM Studio

Current podium

#1 Generalist — 99 / 100 G Agentic — 100 / 100 A Understanding — 90.1 / 100 U Needle — 89.3 / 100 N Thinking — 96.5 / 100 T Live — 75.7 / 100 L Engineering — 98.5 / 100 E Throughput — 5.5 / 100 T

Qwen3.8-27B Q6_K GGUF

27B dense · Q6_K GGUF

8/8 13.2 tok/s
#2 Generalist — 95.5 / 100 G Agentic — 98 / 100 A Understanding — 79.7 / 100 U Needle — 77.8 / 100 N Thinking — 88 / 100 T Live — 55.5 / 100 L Engineering — 91 / 100 E Throughput — 27.7 / 100 T

Gemma-4 26B-A4B 8bit MLX

26B total / 4B active per token · 8bit MLX

8/8 66.2 tok/s
#3 Generalist — 91 / 100 G Agentic — 100 / 100 A Understanding — 93.2 / 100 U Needle — 99 / 100 N Thinking — 97.2 / 100 T Live — 68.3 / 100 L Engineering — 91.5 / 100 E Throughput — 5.6 / 100 T

Qwen3.6-27B Dense 8bit MLX

27B dense · 8bit MLX

8/8 13.4 tok/s

Leaderboard — 35 models · click a column to sort

# Model Params Quant / Format G A U N T L E T Progress tok/s
1 Qwen3.8-27B Q6_K GGUF 27B dense Q6_K GGUF 9910090.189.396.575.798.55.5 8/8 13.2
2 Gemma-4 26B-A4B 8bit MLX 26B total / 4B active per token 8bit MLX 95.59879.777.88855.59127.7 8/8 66.2
3 Qwen3.6-27B Dense 8bit MLX 27B dense 8bit MLX 9110093.29997.268.391.55.6 8/8 13.4
4 Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX MTPLX 8bit MLX 27B 8bit MLX 95.597.58891.545.51006.4 7/8 15.3
5 Mistral Small 3.2 24B 8bit MLX 24B 8bit MLX 7992.578.780100705.4 7/8 12.9
6 Qwen3.6-35B-A3B 4bit MLX 35B total / 3B active per token 4bit MLX 9387.591.297.193.837.4 6/8 89.4
7 Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX Q8_0 GGUF 27B Q8_0 GGUF 8999.590.597.7894.6 6/8 11
8 Qwen3.6-35B-A3B Unsloth Dynamic UD-Q8_K_XL MLX (Brooooooklyn) 35B total / 3B active per token UD-Q8_K_XL mixed-precision MLX 93.59591.210028.6 5/8 68.4
9 Qwen3.5-122B-A10B 4bit MLX 122B total / 10B active per token 4bit MLX 92.59590.210016.5 5/8 39.4
10 Qwen3-Coder-Next 80B 4bit MLX 80B 4bit MLX 899510071.528.2 5/8 67.3
11 Laguna S 2.1 (Q4_K_M GGUF) 118B total / 8B active (MoE, 10-of-256 experts + 1 shared) Q4_K_M GGUF 869610083.522.3 5/8 53.3
12 Bonsai 27B Ternary 2bit MLX 27B 2bit (ternary, PrismML extreme compression) MLX 82.597.593.19917.2 5/8 41.1
13 Nemotron 3 Nano 30B-A3B 4bit MLX 30B total / ~3B active per token 4bit MLX 79.59272.896.736.2 5/8 86.7
14 Qwen3.8-27B Q4_K_M GGUF 27B dense Q4_K_M GGUF 94.599.595.79.3 4/8 22.2
15 Qwen3.8-27B Q8_0 GGUF 27B dense Q8_0 GGUF 94.599.51009.6 4/8 22.9
16 Gemma-4 E4B 4B 8bit MLX 4B 8bit MLX 82.59310031.8 4/8 76
17 Devstral Small 2 24B 6bit MLX 24B 6bit MLX 78.594.51008.5 4/8 20.3
18 Muse-Glimmer 30B (GGUF, kquant 17GB) 30B (unconfirmed -- filename-derived, see note) kquant (unconfirmed exact scheme -- filename says "kquant", not a standard llama.cpp quant label) GGUF 75771007.6 4/8 18.2
19 GPT-OSS 120B 120B mxfp4 MLX 65.810019.2 3/8 45.9
20 Qwen3.8-27B MLX 6bit (lmstudio-community) 27B dense 6bit MLX 51.975n/d6.7 3/8 15.9
21 Gemma 3 1B QAT 4bit MLX 1B QAT 4bit MLX 27.539n/d100 3/8 239.2
22 Qwen3-VL 30B-A3B Instruct 4bit MLX 30B total / 3B active per token 4bit MLX 81.710020.6 3/8 49.2
23 Qwen3.6-35B-A3B MLX 8bit Uniform (lmstudio-community) 35B total / 3B active per token 8bit uniform MLX 85.5n/d36.6 2/8 87.5
24 Gemma-4 E4B 7.5B Q4_K_M GGUF 7.5B Q4_K_M GGUF 78n/d37.4 2/8 89.4
25 Devstral Small 2 24B 4bit MLX 24B 4bit MLX 74.5n/d14.5 2/8 34.7
26 Hermes-4 70B 4bit MLX 70B 4bit MLX 73n/d4.3 2/8 10.4
27 Qwen3-4B-2507 (non-thinking) 4bit MLX 4B 4bit MLX 68.5n/d57.5 2/8 137.6
28 GPT-OSS 120B Fable-5 Distilled 120B mxfp4 MLX 62.7n/d14.6 2/8 34.8
29 GPT-OSS 20B (OpenAI mxfp4) 20B mxfp4 MLX 62.3n/d36.4 2/8 87.1
30 Gemma 3 4B QAT 4bit MLX 4B QAT 4bit MLX 61.5n/d68.3 2/8 163.3
31 DeepSeek-R1-Distill-Qwen-32B 8bit MLX 32B 8bit MLX 60.5n/d5.5 2/8 13.1
32 DeepSeek-R1-Distill-Llama 70B 8bit MLX 70B 8bit MLX 58n/d2.4 2/8 5.7
33 GPT-OSS 20B MLX 20B MLX 52.3n/d14.1 2/8 33.8
34 GPT-OSS Safeguard 20B MLX 20B MLX 50.9n/d24.3 2/8 58
35 Phi-3.5-mini-instruct 4bit MLX 3.8B 4bit MLX 31.5n/d46.1 2/8 110.2

Axes are 0–100. = suite not yet run for that model. n/d = Thinking axis needs ≥15 observed tests to score (see methodology). Default order: GAUNTLET progress, then Generalist score.