xAI's Grok 4.7 Lifts Terminal-Bench 4.0 From 20.3% to 38.0% on a Larger Base Model and a Longer RL Run
xAI released Grok 4.7 on September 21, 2026, reporting a Terminal-Bench 4.0 score of 38.0% against the 20.3% of its predecessor Grok 4.6. Pricing is $2 per million input tokens and $6 per million output tokens, with a fast variant that delivers twice the output speed at twice the price. xAI attributes the gains to a new, larger base model and a longer reinforcement learning run on a harder mix of tasks. ASAP reproduces the published benchmark table and then separates what these numbers establish from what they leave open.
Grok 4.7 takes first place in three of the seven benchmark rows
The comparison table xAI published sets Grok 4.7 against Grok 4.6, GPT-5.6 Sol, and Fable 5.1 across seven rows. On CursorBench 4.0, Grok 4.7 scores 46.3%, ahead of Grok 4.6 at 40.4% and GPT-5.6 Sol at 41.7% but behind Fable 5.1 at 51.8%. Its DeepSWE v1.1 score is 71.0%, carrying an asterisk that xAI marks as a "high effort" result; on the same row Grok 4.6 is 65.2%, GPT-5.6 Sol is 72.7%, and Fable 5.1 is 70.0%.
Three rows go to Grok 4.7. EEBench is 64.0%, ahead of Grok 4.6 at 53.0%, GPT-5.6 Sol at 39.4%, and Fable 5.1 at 56.4%. AA Briefcase v1.1 is 1,657, above Grok 4.6 at 1,546 and GPT-5.6 Sol at 1,487 but below Fable 5.1 at 1,678. The Harvey Legal Agent Benchmark is where the gap is widest: 19.6% against Grok 4.6 at 15.8%, GPT-5.6 Sol at 2.5%, and Fable 5.1 at 6.7%. HealthBench Professional runs the other way at 56.7%, behind GPT-5.6 Sol at 60.5% and Fable 5.1 at 62.1%. Outside the table, xAI also reports a GDPval Elo score of 1,695.
The 17.7-point Terminal-Bench jump is the largest single move in the release
Terminal-Bench 4.0 is the row with the biggest generational gain, rising 17.7 percentage points from 20.3% to 38.0%. EEBench follows at 11.0 points, then HealthBench Professional at 8.2, CursorBench 4.0 at 5.9, DeepSWE v1.1 at 5.8, and the Harvey Legal Agent Benchmark at 3.8. That ordering is the clearest available picture of which axis this generation pushed.
The training changes xAI describes line up with that distribution. The company says the model "was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete," and that as a result "the model is better at verifying its own work and managing longer context." Terminal-Bench and EEBench are less about single-response quality than about carrying a task through many steps while checking intermediate results, and that is exactly where the gains cluster. HealthBench Professional, which probes knowledge and judgment, gained 8.2 points and still trails both competitors.
What this outlines is a division of labor rather than a ranking. Grok 4.7 cedes first place in the code-editor setting of CursorBench and in software repair on DeepSWE, and claims ground on long-running terminal work and legal agent tasks. No single model wins every column, so the practical question at a fixed price point becomes which column a team is actually buying.
A leading 19.6% on Harvey Legal is still a low absolute number
Grok 4.7's 19.6% on the Harvey Legal Agent Benchmark leads the table and also means the model fails four attempts out of five. GPT-5.6 Sol scores 2.5% and Fable 5.1 scores 6.7% on the same row. Both facts hold at once: the spread between models is large, and all three models mostly fail the task.
Mixing relative and absolute readings produces bad decisions here. Describing 19.6% against 2.5% as "roughly eight times the competition" is an accurate transcription, but a team weighing deployment needs the other sentence, which is that the task succeeds about one time in five. CursorBench at 46.3% and Terminal-Bench at 38.0% are likewise below half. A benchmark table ranks models against each other; it does not establish that a task can be automated, and this release stretches that distinction furthest on the legal row.
The DeepSWE asterisk calls for the same caution. If 71.0% is a "high effort" measurement, then the 72.7% and 70.0% printed beside it belong on the same line only if those were produced under comparable compute conditions, which the announcement does not state. Attaching the footnote rather than hiding it is to xAI's credit, and a footnoted value is still a different kind of value from its unfootnoted neighbors.
Reporting safety numbers inside the capability release is the structural change
Two safety figures appear alongside the performance table: a 3.3% allowance rate for risky dual-use prompts on HackerBench v0.3, and 62.4% on the LatchBio biosafety benchmark. xAI states that "Grok 4.7 was built with an entirely new safeguard stack," that "it is the strongest model we've tested on refusals and jailbreak resistance," and that "in dual-use domains like cybersecurity and biological work, it leads on both utility for benign tasks and safe refusal on dangerous ones."
The design detail worth noticing is that the metric runs in two directions rather than one. The 3.3% figure measures how much dangerous material got through, and xAI pairs it with the phrase "while rarely blocking legitimate security work." Raising refusal rates alone makes a safety number look excellent and makes the model useless to security practitioners; optimizing purely for utility does the reverse. Binding both into a single reported metric closes off the path of buying safety through over-refusal.
These figures are nonetheless self-reported, and the announcement does not publish the item composition or scoring rubric for either the HackerBench or LatchBio benchmark. Competitors have an incentive to reproduce and contest capability scores; for safety scores, neither the contesting party nor the reproduction path is obvious. Including safety numbers in a capability release is progress, and the verification structure behind those numbers is still missing.
What a team should check before switching
Separating what is priced from what is undisclosed is the safer first step. Three things are fixed: $2 per million input tokens, $6 per million output tokens, and a fast variant at twice the speed and twice the price. Absent from the announcement are the context window size and the API model identifier strings, along with any statement about regional availability. The context window matters most here, because xAI explicitly claims improved long-context management without attaching a number to it, so teams working on long documents have reason to wait for that figure.
Deployment surface is broad. xAI says Grok 4.7 is available the same day in Cursor and Grok Build, and through the Grok API, third-party coding harnesses, model routers, and cloud platforms. For a team already routing across several models, the switching cost is close to a configuration change, which turns the decision from whether to adopt into which class of work to move. The table points toward multi-hour terminal work and document-centric agent tasks, and away from single-shot code fixes and medical knowledge queries, where other models lead in these same rows.
One further dependency deserves attention. xAI says Grok 4.7 was trained "to natively understand the Grok Bot harness," which means optimization for a specific execution environment is baked into the model. That is precisely the condition under which published benchmark scores fail to reproduce inside a different in-house harness, so running an internal task set before committing is not a step to skip.
Open questions
The first question this release does not answer is scale: xAI describes "a new, larger base model compared to Grok 4.6" and publishes neither a parameter count nor a training token budget. Without that, the 17.7-point Terminal-Bench gain cannot be split between the larger base model and the longer RL run, and xAI itself reports the two changes together without apportioning credit.
The second question is cost per result. Pricing is public at $2 input and $6 output, but as the "high effort" footnote on DeepSWE implies, the token volume required to reach a given score differs by model. Multiplying a table score by a table price requires average output tokens per task, which is not disclosed. The third question is reproduction: rows such as CursorBench and Terminal-Bench can be run externally, independent results will land within days, and whether the self-reported table survives that is the real verification point for this launch.
Source: xAI's official announcement "Grok 4.7" (x.ai/news/grok-4-7, September 21, 2026). Every score, price, and quotation above was checked directly against that document by ASAP.

AI & tech,
read in depth
Beyond the headlines — into the context and the structure
AGI Soon As Possible · asapai.co.kr