xAI released Grok 4.7 on September 21, 2026, and published a set of lab-reported benchmark results with the launch, including 71.0% on DeepSWE v1.1 at high effort. In its Grok 4.7 announcement, signed as SpaceXAI, the company calls the model “our most capable model for coding and knowledge work.”
On this page
Availability
xAI says that “Grok 4.7 is available today in Cursor and Grok Build. It is also available through the Grok API, third-party coding harnesses, and model routers and cloud platforms.”
What xAI says changed
According to the announcement, Grok 4.7 “works longer on difficult tasks, checks its own work more carefully, and comes with our best-calibrated safeguards to date.” xAI says the model “uses a new, larger base model compared to Grok 4.6” and was “trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete.” The company also says it trained Grok 4.7 “to natively understand the Grok Bot harness,” aimed at conversational tasks and general knowledge work.
Benchmark results reported by xAI
All scores below are xAI’s own, as published in the announcement. xAI marks the DeepSWE v1.1 result as run at high effort.
| Area (as labelled by xAI) | Benchmark | Grok 4.7 |
|---|---|---|
| Software engineering | DeepSWE v1.1 (high effort) | 71.0% |
| Software engineering | CursorBench 4.0 | 46.3% |
| Multi-hour terminal work | Terminal-Bench 4.0 | 37.6% |
| Multi-hour office work | AA-Briefcase v1.1 | 1,657 |
| Legal work | Harvey Legal Agent Benchmark | 19.6% |
| Clinical reasoning | HealthBench Professional | 56.7% |
xAI describes CursorBench 4.0 as a benchmark “which stresses longer-running coding tasks.” The table spans coding, terminal work, office tasks, legal work and clinical reasoning, in line with the company’s stated training emphasis on problems that take many hours to complete.
How to read the numbers
These are lab-reported figures, measured by xAI on its own setup, and they cover the benchmarks xAI chose to publish. Scores on the same benchmark from a different lab, a benchmark maintainer or a third-party evaluator can differ with the harness, effort setting and benchmark version, so comparisons hold best when those details match.
What it changes
For developers, Grok 4.7 arrives in Cursor and xAI’s own Grok Build on day one, with API and cloud access alongside. By xAI’s description, it is now the company’s most capable model for coding and knowledge work, built on a larger base model than Grok 4.6.
Grok 4.7’s scores are recorded with their source and date on our benchmark leaderboard, and the live state of xAI’s service is on the Grok status page.




