Grok 4.5 and GPT-5.6 land on APEX

SpaceXAI's Grok 4.5 debuts at #2 on APEX-SWE and takes the overall Integration lead, and OpenAI's Sol reaches frontier-level performance on APEX-Agents while delivering impressive token efficiency.

Grok 4.5's coding debut ranks #2 on APEX-SWE at 51.2% Pass@1, with the overall Integration lead at 65.0%. Sol's efficiency also stands out: 49.8 Pass@1 points per million tokens on APEX-Agents, 1.5x the field, on 757k tokens per trajectory while peers spend 0.9M to 1.4M.

Grok 4.5 debuts at #2 on APEX-SWE

Grok 4.5 debuts at #2 on APEX-SWE with 51.2% Pass@1 (±6.0), behind only Fable 5. On APEX-Agents it lands at #7 with 34.2% Pass@1 and a 50.9% mean score, more than double Grok 4 on both metrics in a year.

APEX-SWE measures economically valuable software engineering work like multi-step build tasks and diagnosis and debugging, while APEX-Agents measures long-horizon professional work across investment banking, corporate law, and management consulting.

Grok 4.5 results on APEX-SWE and APEX-Agents

01 · New #1 leader for Integration

On APEX-SWE's Integration tasks, Grok 4.5 posts 65.0% Pass@1, the top score recorded for any model, with Observability at 37.3% Pass@1, ranking #3. Overall it climbs from Grok 4's 21.0% to 51.2%, a 30.2-point gain in a year.

02 · Top 5 in Investment Banking

On APEX-Agents, Grok 4.5 posts 37.3% Pass@1 with a 44.2% mean in investment banking (#5 in the domain), 25.4% with a 52.2% mean in corporate law, and 39.9% with a 56.2% mean in management consulting, its strongest domain and its largest jump from Grok 4's 12.0%.

03 · Room to climb in corporate law

Corporate law is Grok 4.5's lowest Pass@1 domain at 25.4%, but its 52.2% mean shows the model already completes most of each legal task. The model gets close and doesn't finish. Fix that, and corporate law is the biggest source of remaining gains.

Grok 4.5 scores by APEX-Agents domain

GPT-5.6 Sol reaches near-frontier accuracy on half the tokens

01 · Fewer tokens with many light steps

Sol averages 28k tokens per tool call compared with Opus 4.8 at 63k. It also takes more steps than most of the frontier, but each step is less than half the weight. Sol operates in many light moves rather than a few heavy ones.

02 · Strong in Law and Consulting

Sol's overall APEX-Agents Pass@1 is 37.7%, against Fable 5 at 43.3%. The gap is not spread evenly: Sol matches Opus 4.8 and GPT-5.5 in corporate law and management consulting, scoring 60.2% and 59.5%. In investment banking it scores 42.3%, giving up 8 to 11 points to Opus 4.8's 50.6% and Fable 5's 53.8%. Its headroom is concentrated in heavy quantitative modeling work.

03 · Token efficient, less consistent

Given 4 attempts, Sol solves 48.5% of tasks by Pass@4 but only 28.1% on all four attempts by Pass^4. That is a 20.4-point capability-to-reliability gap, wider than GPT-5.5's 15.0 points between 45.8% and 30.8%. Sol can do more than its Pass@1 suggests; it just does not do it the same way twice. The distribution is bimodal: 28% of tasks pass on all four runs, 51% never pass, and only 99 of 480 are flaky.

04 · xHigh captures most of Sol's ceiling

Higher reasoning effort buys a little more accuracy. On APEX-Agents, max reasoning reaches 39.8%, and with pro mode enabled, 40.0%, both 2 to 2.5 points above xHigh's 37.7% at 2.4 to 5x the token spend. On APEX-SWE, max reasoning lifts Pass@1 from 39.7% to 41.2%. xHigh already captures most of Sol's ceiling at the most efficient setting.

GPT-5.6 Sol token efficiency on APEX-Agents