AGI Soon As Possible · Deep reads on AI & tech
Article

Reading the Grok 4.6 eval table row by row: coding gains, and 26% on terminal work

2026-08-13 · 9 min read

SpaceXAI released Grok 4.6 on August 12, 2026, reporting a score of 61 on the Artificial Analysis Intelligence Index, a composite of nine benchmarks, which ties GPT-5.6 Sol. Pricing is $2 per million input tokens and $6 per million output tokens, with a fast variant at twice that price, and the model is available from launch day in Cursor, Grok Build, and the API. Going down the same announcement's table row by row tells a more divided story: CursorBench v3.2 comes in at 69.9% against GPT-5.6 Sol Max's 67.2%, while Terminal-Bench v3.0 lands at 26% against 34.6%. ASAP works from SpaceXAI's own announcement to separate the categories where this model wins from the ones where it does not.

The table splits differently on every row

SpaceXAI's published table places Grok 4.6 High alongside Grok 4.5 High, GPT-5.6 Sol Max, and Fable 5 Max across ten evaluations, and the winner changes from row to row. Generational improvement and competitive position have to be read separately for the table to make sense.

Generationally, all ten rows moved up. The Artificial Analysis Intelligence Index went from 56 to 61, GDPVal-AA v2 from 1526 to 1753, and CursorBench v3.2 from 66.7% to 69.9%. The largest single jump is DeepSWE v1.1, which rose from 54% to 65.9%, a gain of 11.9 percentage points, followed by Terminal-Bench v3.0 moving from 15.7% to 26%, a gain of 10.3 points. AA-Briefcase climbed from 1313 to 1577 and APEX-Agents from 47.1% to 57.5%.

Against competitors the picture changes. Grok 4.6 beats both GPT-5.6 Sol Max and Fable 5 Max on exactly three rows: GDPVal-AA v2 (1753 to 1728 to 1741), AA-Briefcase (1577 to 1502 to 1574), and Harvey LAB (15.8% to 2.5% to 11.3%). CursorBench v3.2 at 69.9% clears GPT-5.6 Sol Max's 67.2% but falls 0.6 points short of Fable 5 Max's 70.5%, and the Artificial Analysis Intelligence Index at 61 ties GPT-5.6 Sol Max while trailing Fable 5 Max's 62 by a point.

The losses are equally clear. DeepSWE v1.1 at 65.9% trails both GPT-5.6 Sol Max's 73% and Fable 5 Max's 70%, and Terminal-Bench v3.0 at 26% trails 34.6% and 34.1% by a wide margin. FrontierCode v1.1 Extended at 61.3% edges GPT-5.6 Sol Max's 60.6% but sits below Fable 5 Max's 63.6%, and APEX-Agents at 57.5% falls short of Fable 5 Max's 59.2%. APEX-SWE reads 56.4% against Fable 5 Max's 58.8%, with the GPT-5.6 Sol Max cell left empty.

Grok 4.5 served as the teacher for Grok 4.6

SpaceXAI states in the announcement that it used Grok 4.5 to regenerate the SFT trajectories that fed Grok 4.6's supervised fine-tuning stage. One generation manufactured the training data for the next.

The announcement describes three stages. First came a longer supplemental training run than Grok 4.5 received, using curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe. Then Grok 4.5 regenerated SFT trajectories across reasoning efforts, agent harnesses, and domains including STEM, software engineering, and knowledge work, with problematic traces filtered out by model-based checks. Finally the model was trained on a wide range of agentic RL tasks spanning knowledge work, general coding, and domain-specific environments for kernel optimization, web development, and computer-aided design.

The announcement also records a behavioral change observed on longer runs. Self-testing and verification appeared more often, with the model checking its own work before moving on. On visual and interactive projects, SpaceXAI reports that given a concrete product idea the model could establish an application's structure and visual language in a single pass.

How to read 26% on terminal work from a model built for long-running agents

Terminal-Bench v3.0 at 26% is the number in this release that most needs explaining, because it appears to contradict the long-running-agent framing the announcement leads with. The contradiction softens once the benchmarks are sorted by what they actually measure.

The strength the announcement emphasizes is "turning a broad product idea into a working first version": researching an unfamiliar domain, structuring the application, implementing core interactions, and refining through rounds of feedback. Success there is judged by whether the artifact looks and behaves credibly to a person, and the rows where Grok 4.6 leads, including CursorBench, AA-Briefcase, and Harvey LAB, sit closer to that kind of judgment.

Terminal-Bench and DeepSWE are a different species. They ask whether an environment's stated conditions were met, judged mechanically, and a single broken step drops the whole attempt to a failure. That both of Grok 4.6's largest deficits fall in this family is unlikely to be coincidence. Producing a strong first pass and driving a fixed verification to completion are not the same axis of capability.

Read alongside the generational delta, the direction is right. Terminal-Bench rose 10.3 points from 15.7% to 26%, one of the steepest relative improvements in the table. But 26% in absolute terms means three attempts in four do not finish. In interactive use, where a person reviews output and issues the next instruction, that gap is easy to miss; in an unattended automation pipeline it becomes the failure rate directly.

More cells in the table are not bolded than are

SpaceXAI states that the best score per evaluation is shown in bold, and Grok 4.6 holds that position in three of ten rows. The announcement's headline sentence and the full table do not say the same thing.

The summary claims are that Grok 4.6 "achieves frontier intelligence across several agentic coding and knowledge work benchmarks" and that it "matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is a composite score of nine benchmarks." Both are accurate. The comparison partner in the tie is GPT-5.6 Sol, and on that same index Fable 5 Max scores 62, one point higher. The model chosen for the headline comparison is not the table's leader.

This structure points at something worth checking in every model announcement. A composite index compresses nine benchmarks into one figure and erases per-row variance. A profile like Grok 4.6's, strong in one family and well behind in another, is exactly the profile a single composite number describes least well. Some secondary coverage framed this release as leading GPT on both coding and terminal performance, but the announcement's own table puts Terminal-Bench v3.0 at 26% against 34.6%.

The evaluation conditions carry their own caveat. SpaceXAI states that third-party model scores are the best of self-reported or publicly available results, which is a condition favorable to competitors rather than one that inflates them. What the announcement does not define is what its own "High" setting denotes.

Where $2 per million input tokens places this model

Grok 4.6 costs $2 per million input tokens and $6 per million output tokens, with a fast variant at double, and those figures only mean something read against the eval table. When performance varies by task family, price is the variable that decides the choice.

The notable feature of the structure is the 3x ratio between input and output. That shape favors work that reads a lot and writes a little, such as ingesting a large codebase or document set and returning a short judgment. Work that continuously generates long code is dominated by the output rate instead, and the "establish an application's structure in one pass" use the announcement highlights sits on the output-heavy side.

The fast variant at twice the price fits the stated use as well. Long-running agents accumulate latency across many steps, so speed carries real value. Set against the completion rate implied by 26% on Terminal-Bench, however, failed runs still consume tokens. Automation budgets are set by the total cost of running until success, not by the unit cost of a successful run.

The offer of 2x included usage in Grok Build and Cursor for the first week bears directly on evaluation timing. It is an invitation to test against your own workload while the extra allowance lasts, and real running cost after that window has to be estimated separately.

What a team in Korea should check before adopting Grok 4.6

The first decision for an engineering team evaluating this model is whether the use is interactive or unattended, because the eval table answers those two cases differently. Rather than importing the table wholesale, pick the rows that correspond to your own task types.

For workflows where a person reviews output and issues the next instruction, CursorBench at 69.9% and AA-Briefcase at 1577 are the relevant rows. Draft quality and the polish of visual artifacts are what users actually feel. For workflows wired into CI and run without human intervention, Terminal-Bench at 26% and DeepSWE at 65.9% are closer to the baseline, and at those levels retry policy and failure handling matter more than model choice.

Legal and document-heavy organizations should read Harvey LAB at 15.8% differently. The absolute value is low, but against GPT-5.6 Sol Max's 2.5% and Fable 5 Max's 11.3% it is the widest relative gap in the table. That benchmark targets English-language legal work, and the announcement offers no basis for transferring the result to Korean legal documents. Any local adoption decision still needs a re-measurement on an in-house task set.

Access paths are broad. Beyond Cursor, Grok Build, and the API, the announcement lists partners including OpenRouter, Vercel, and Cloudflare. Teams already on those paths can run a comparison without a separate contract.

Open questions: context length and evaluation settings

The largest omission is context length, and its absence from a release built around long-running agents is itself informative. The more steps a task takes, the more the ceiling on carried context governs performance, and the published announcement does not state that number.

Evaluation settings are open too. Every first-party score is labeled Grok 4.6 High or Grok 4.5 High, but the announcement never defines what "High" means in terms of reasoning effort or time budget. The training description mentions regenerating SFT trajectories across reasoning efforts, which suggests multiple effort tiers exist, yet where the table's "High" sits among them, and which pricing variant it corresponds to, is not confirmed anywhere in the text.

A third gap is the content of the safety evaluation. SpaceXAI states that pre-deployment testing was its widest-ever suite and that extensive post-deployment and third-party testing was performed, but provides no figures or named evaluations and no pass criteria. What appears instead is a description of calibration, saying safeguards were tuned in line with the model's capabilities to keep it helpful and safe in areas such as vulnerability patching, accelerating the engineering design cycle, and augmenting AI research. That tells you the direction without providing anything verifiable.

Source: ASAP analysis based on SpaceXAI's official announcement "Introducing Grok 4.6" (August 12, 2026)

ASAP — AGI Soon As Possible

AI & tech,
read in depth

Beyond the headlines — into the context and the structure

AGI Soon As Possible · asapai.co.kr

← All posts