Local-LLM proving ground

Proving local agents before they act.

GAUNTLET Bench runs local models through real workloads — agentic, live, verified — and publishes the evidence.

39models 10suites 907per-test results

All scores measured on Apple M5 Max, 128 GB unified memory, LM Studio

Current podium

1 Generalist — 99.5 / 100 G Agentic — 96 / 100 A Understanding — 93.3 / 100 U Needle — 91.2 / 100 N Thinking — 96.5 / 100 T Live — 80.5 / 100 L Engineering — 96.5 / 100 E Throughput — 5.8 / 100 T

Qwen3.8-27B Q6_K GGUF

27B dense · Q6_K GGUF

8/8 13.8 tok/s
2 Generalist — 95.5 / 100 G Agentic — 99.5 / 100 A Understanding — 81 / 100 U Needle — 77.8 / 100 N Thinking — 88 / 100 T Live — 65.5 / 100 L Engineering — 98.5 / 100 E Throughput — 27.9 / 100 T

Gemma-4 26B-A4B 8bit MLX

26B total / 4B active per token · 8bit MLX

8/8 66.4 tok/s
3 Generalist — 90.5 / 100 G Agentic — 100 / 100 A Understanding — 96.3 / 100 U Needle — 100 / 100 N Thinking — 97.2 / 100 T Live — 79 / 100 L Engineering — 100 / 100 E Throughput — 5.7 / 100 T

Qwen3.6-27B Dense 8bit MLX

27B dense · 8bit MLX

8/8 13.4 tok/s

Leaderboard — 39 models · click a column to sort

# Model Params Quant / Format G A U N T L E T Progress tok/s
1 Qwen3.8-27B Q6_K GGUF 27B dense Q6_K GGUF 99.59693.391.296.580.596.55.8 8/8 13.8
2 Gemma-4 26B-A4B 8bit MLX 26B total / 4B active per token 8bit MLX 95.599.58177.88865.598.527.9 8/8 66.4
3 Qwen3.6-27B Dense 8bit MLX 27B dense 8bit MLX 90.510096.310097.2791005.7 8/8 13.4
4 Mistral Small 3.2 24B 8bit MLX 24B 8bit MLX 80.5917980.210062.586.55.6 8/8 13.2
5 Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX MTPLX 8bit MLX 27B 8bit MLX 95.59990.891.5531006.4 7/8 15.3
6 Qwen3.6-35B-A3B 4bit MLX 35B total / 3B active per token 4bit MLX 9487.591.297.19438.9 6/8 92.5
7 Qwen3.6-27B Fable-Fusion-711 Uncensored Heretic NEO-MAX Q8_0 GGUF 27B Q8_0 GGUF 89.510091.297.798.54.6 6/8 11
8 Qwen3-Coder-Next 80B 4bit MLX 80B 4bit MLX 88.594.526.510091.524 6/8 57.1
9 Qwen3.6-35B-A3B Unsloth Dynamic UD-Q8_K_XL MLX (Brooooooklyn) 35B total / 3B active per token UD-Q8_K_XL mixed-precision MLX 93.59691.210028.5 5/8 67.8
10 Qwen3.5-122B-A10B 4bit MLX 122B total / 10B active per token 4bit MLX 9094.591.210016.4 5/8 38.9
11 Laguna S 2.1 (Q4_K_M GGUF) 118B total / 8B active (MoE, 10-of-256 experts + 1 shared) Q4_K_M GGUF 879910087.522.6 5/8 53.6
12 Bonsai 27B Ternary 2bit MLX 27B 2bit (ternary, PrismML extreme compression) MLX 819293.19917.4 5/8 41.4
13 Nemotron 3 Nano 30B-A3B 4bit MLX 30B total / ~3B active per token 4bit MLX 78.592.57596.736.5 5/8 86.7
14 Devstral Small 2 24B 6bit MLX 24B 6bit MLX 7897100908.8 5/8 20.8
15 GPT-OSS 120B 120B mxfp4 MLX 65.4601008515.6 5/8 37.2
16 Qwen3.8-27B Q4_K_M GGUF 27B dense Q4_K_M GGUF 95.510095.79.4 4/8 22.2
17 Qwen3.8-27B Q8_0 GGUF 27B dense Q8_0 GGUF 9495.51009.7 4/8 22.9
18 Gemma-4 E4B 4B 8bit MLX SMALL 4B 8bit MLX 84.594.510032 4/8 76
19 Muse-Glimmer 30B (GGUF, kquant 17GB) 30B (unconfirmed -- filename-derived, see note) kquant (unconfirmed exact scheme -- filename says "kquant", not a standard llama.cpp quant label) GGUF 74761007.7 4/8 18.2
20 Ornith-1.0-35B 4bit MLX 35B total (MoE; active-per-token not confirmed in model card) 4bit MLX 86.398n/d48.5 3/8 115.2
21 Gemma-4 E4B 7.5B Q4_K_M GGUF SMALL 7.5B Q4_K_M GGUF 7826.5n/d23.7 3/8 56.3
22 GPT-OSS 20B (OpenAI mxfp4) 20B mxfp4 MLX 63.7n/d7143.4 3/8 103.2
23 Gemma 3 1B QAT 4bit MLX SMALL 1B QAT 4bit MLX 28.533n/d100 3/8 237.7
24 Qwen3-VL 30B-A3B Instruct 4bit MLX 30B total / 3B active per token 4bit MLX 83.310019.9 3/8 47.2
25 Qwen3.6-35B-A3B MLX 8bit Uniform (lmstudio-community) 35B total / 3B active per token 8bit uniform MLX 88n/d36.8 2/8 87.5
26 Hermes-4 70B 4bit MLX 70B 4bit MLX 74n/d4.4 2/8 10.4
27 Devstral Small 2 24B 4bit MLX 24B 4bit MLX 73.5n/d14.6 2/8 34.7
28 Qwen3-4B-2507 (non-thinking) 4bit MLX SMALL 4B 4bit MLX 68n/d57.1 2/8 135.6
29 GPT-OSS 120B Fable-5 Distilled 120B mxfp4 MLX 63.3n/d14.6 2/8 34.8
30 DeepSeek-R1-Distill-Llama 70B 8bit MLX 70B 8bit MLX 62.5n/d2.4 2/8 5.7
31 DeepSeek-R1-Distill-Qwen-32B 8bit MLX 32B 8bit MLX 60.5n/d5.5 2/8 13.1
32 Gemma 3 4B QAT 4bit MLX SMALL 4B QAT 4bit MLX 57.5n/d68.7 2/8 163.3
33 GPT-OSS 20B MLX 20B MLX 54.3n/d14.2 2/8 33.8
34 GPT-OSS Safeguard 20B MLX 20B MLX 53n/d24.4 2/8 58
35 Phi-3.5-mini-instruct 4bit MLX SMALL 3.8B 4bit MLX 36n/d46.4 2/8 110.2
36 Qwen3.8-27B MLX 8bit (lmstudio-community) 27B dense 8bit MLX 0n/d6.5 2/8 15.5
37 Qwen3.8-27B MTPLX Optimized-Speed 4bit MLX (Youssofal) 27B dense 4bit MLX 0n/d8.4 2/8 20
38 Qwen3.8-27B MLX 6bit (lmstudio-community) 27B dense 6bit MLX 75n/d7.5 2/8 17.8
39 Qwen3.8-27B MLX 8bit (lmstudio-community) 27B dense 8bit MLX 87.5n/d6.6 2/8 15.8

SMALL = under 8B total parameters (phone/edge class) — scored on the same suites, ranked in the same list; compare within class.
Axes are 0–100. = suite not yet run for that model. n/d = Thinking axis needs ≥15 observed tests to score (see methodology). Default order: GAUNTLET progress, then Generalist score.