xAI released Grok 4.7 on September 21, 2026, and published a set of lab-reported benchmark results with the launch, including 71.0% on DeepSWE v1.1 at high effort. In its Grok 4.7 announcement, signed as SpaceXAI, the company calls the model “our most capable model for coding and knowledge work.”

On this page

Availability

xAI says that “Grok 4.7 is available today in Cursor and Grok Build. It is also available through the Grok API, third-party coding harnesses, and model routers and cloud platforms.”

What xAI says changed

According to the announcement, Grok 4.7 “works longer on difficult tasks, checks its own work more carefully, and comes with our best-calibrated safeguards to date.” xAI says the model “uses a new, larger base model compared to Grok 4.6” and was “trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete.” The company also says it trained Grok 4.7 “to natively understand the Grok Bot harness,” aimed at conversational tasks and general knowledge work.

Benchmark results reported by xAI

All scores below are xAI’s own, as published in the announcement. xAI marks the DeepSWE v1.1 result as run at high effort.

Area (as labelled by xAI) Benchmark Grok 4.7
Software engineering DeepSWE v1.1 (high effort) 71.0%
Software engineering CursorBench 4.0 46.3%
Multi-hour terminal work Terminal-Bench 4.0 37.6%
Multi-hour office work AA-Briefcase v1.1 1,657
Legal work Harvey Legal Agent Benchmark 19.6%
Clinical reasoning HealthBench Professional 56.7%

xAI describes CursorBench 4.0 as a benchmark “which stresses longer-running coding tasks.” The table spans coding, terminal work, office tasks, legal work and clinical reasoning, in line with the company’s stated training emphasis on problems that take many hours to complete.

How to read the numbers

These are lab-reported figures, measured by xAI on its own setup, and they cover the benchmarks xAI chose to publish. Scores on the same benchmark from a different lab, a benchmark maintainer or a third-party evaluator can differ with the harness, effort setting and benchmark version, so comparisons hold best when those details match.

What it changes

For developers, Grok 4.7 arrives in Cursor and xAI’s own Grok Build on day one, with API and cloud access alongside. By xAI’s description, it is now the company’s most capable model for coding and knowledge work, built on a larger base model than Grok 4.6.

Grok 4.7’s scores are recorded with their source and date on our benchmark leaderboard, and the live state of xAI’s service is on the Grok status page.

Sources