NVIDIA Blackwell Takes No. 1 in the AI-Agent Hardware Benchmark
NVIDIA's Blackwell GB300 NVL72 platform took the top spot in "AA-AgentPerf," a new benchmark that measures hardware performance for AI agents. Introduced by Artificial Analysis in 2026, the benchmark reproduces real coding-agent workloads to evaluate throughput per unit of power, and Blackwell was found to handle roughly 20 times more AI agents per megawatt than the previous-generation Hopper. It shows that in the age of AI agents, "power efficiency" has emerged as a key competitive metric.
Watch the Yardstick, Not the Ranking
The most telling thing about this result is not the ranking itself but the measure used to produce it. In the inaugural AA-AgentPerf results released by Artificial Analysis, the Blackwell GB300 NVL72 posted the top performance, beating AMD's MI355X in the same evaluation. The model used in the evaluation was DeepSeek V4 Pro, released around April 2026, and NVIDIA stated that Blackwell handled about 20 times more agents per megawatt than the previous-generation Hopper. What decided the winner was not tokens per second or raw FLOPs, the metrics that long ruled hardware comparisons, but "Agents per Megawatt." In other words, the target of measurement has moved from a chip's peak speed to a whole system's sustainable throughput.
How to Read the 20x Figure
A roughly 20-times gap per megawatt is unusually large for a single generational step. But it would be a mistake to read it as a per-chip performance multiplier. AA-AgentPerf replays real coding-agent workflows spanning up to 200 turns and more than 100,000 tokens, and it measures how many agents a single system can handle simultaneously while meeting a service-level agreement (SLA). So the 20x should be read as a figure produced not by a single chip but by the entire rack-scale NVL72 system under the long-context, repeated-inference load that is characteristic of agents. The more natural interpretation is that the generational gap widens precisely in the regime where accumulating conversation piles pressure on memory and bandwidth.
Why Power, Not Speed, Is the Real Constraint for Agents
An agent performs not a single inference but a chain of dozens to hundreds of them. Even to reach the same outcome, it burns far more compute and power than a one-off chatbot. In that structure, how many agents you can sustain concurrently on the same power matters more than how fast the chip is, because that ratio effectively becomes unit cost. When power is a largely fixed cost and a resource that is itself hard to secure, throughput per megawatt is closer to a business metric than a performance one. The fact that the benchmark's core metric has shifted here shows the grammar of competition moving from "who is faster" to "who lasts longer on the same electricity."
What It Means for Korea and for Practitioners
In Korea, where grid headroom is thin and competition over data-center sites and power is fierce, this metric lands with particular weight. The center of gravity in purchasing decisions is likely to move from catalog peak speed to actual agent capacity per contracted megawatt. Still, it is worth remembering that no single benchmark speaks for everything. Because the model and workload used here are tuned to a DeepSeek V4 Pro coding-agent flow, whether the same gap reproduces under other models or inference patterns needs separate verification. Practitioners would do well to re-measure throughput per megawatt under conditions close to their own workloads.
Sources: Artificial Analysis · NVIDIA Blog · Crypto Briefing

AI & tech,
read in depth
Beyond the headlines — into the context and the structure
AGI Soon As Possible · asapai.co.kr