← Leaderboard

GPT-OSS 120B

Keeper Rank #19 of 33 · 3/8 GAUNTLET progress
Generalist — 65.8 / 100 G Agentic — pending A Understanding — pending U Needle — pending N Thinking — 100 / 100 T Live — pending L Engineering — pending E Throughput — 28.1 / 100 T

Hollow markers & dashed spokes: axis not yet scored

G
65.8
A
U
N
T
100
L
E
T
28.1

Specification

Parameters
120B
Architecture
gpt_oss
Size on disk
63.41 GB
Quantization
mxfp4
Format
MLX
Reasoning (CoT)
Yes — emits reasoning tokens
Internal ID
G4
Mean speed
45.9 tok/s across suites
Stall census
0 stalls in 15 observed tests (0.0%)
Model card
huggingface.co

Suite results

Suite Score Avg / 20 Tests tok/s
General capability (13-task real-workload suite) 1a 171 / 260 19 9/13 40
Content-production depth 1e 112.8 / 120 18.8 6/6 51.7

Each test is LLM-judged 0–20 against a fixed rubric; suite max = tests × 20. Where fewer tests ran than the suite total, unrun tests count as zero toward the axis score. See methodology.