ASAPAGI Soon As Possible · Deep reads on AI & tech
Article

K-EXAONE 2.0 released: LG AI Research's 750B open-weight model leads on 3 of 24 benchmarks and trails all three rivals on 17

2026-07-31 · 8 min read

LG AI Research released K-EXAONE 2.0, a Mixture-of-Experts language model with 750 billion total parameters and 37 billion active per token, on Hugging Face under the Apache 2.0 license on July 31, 2026. Read directly, the 24-benchmark table published in its own model card shows the model ahead of all three comparison models, Qwen3.5, GLM-5.1 and DeepSeek-V4 Pro max, on exactly three rows: OpenAI-MRCR at 94.4, KGC-Safety at 99.8 and ROK-Fortress at 89.5. It sits below all three on 17 rows. ASAP works from the table LG AI Research published itself, rather than the "overtakes foreign models" summary in Korean press coverage, to place this model accurately.

Two of the three headline claims reproduce from the table, one does not

LG AI Research led its announcement with three numbers: an average of 70.1 across 24 benchmarks, a 30% gain on three coding and agentic tasks, and a lead of more than 10% over GLM-5.1 on long context. The second and third reproduce exactly from the model card. The coding and agentic category, made up of SciCode, SWE Bench Verified and Terminal-Bench 2.1, averages 38.4 for the previous K-EXAONE and 49.8 for 2.0, a gain of 29.6% that matches the stated 30%. The long-context category of OpenAI-MRCR, AA-LCR and Ko-LongBench averages 80.1 for 2.0 against 72.5 for GLM-5.1, a 10.5% lead that matches the stated figure.

The first number does not reproduce. A flat mean of the 24 scores in the card gives 73.3 for K-EXAONE 2.0 and 66.9 for the previous K-EXAONE (calculated by ASAP directly from the model card table). Averaging the nine category averages instead gives 73.4, essentially the same. Neither method yields the announced 70.1 and 63.3. The gap between announced and computed values sits at roughly 3.2 points in both cases, a consistent offset that points to a different evaluation configuration or benchmark set behind the announcement rather than an arithmetic slip.

On relative improvement the two methods agree. The announced pair of 70.1 against 63.3 is a 10.7% rise; the card-derived pair of 73.3 against 66.9 is a 9.5% rise. The claim of more than 10% improvement over the previous generation holds in direction. Anyone citing the absolute level, however, should note that 70.1 cannot be verified against the published table. That is why this article works from individual benchmark scores rather than the average.

What the three leading scores have in common

K-EXAONE 2.0 beats all three comparison models on three rows: OpenAI-MRCR at 94.4, KGC-Safety at 99.8 and ROK-Fortress at 89.5. OpenAI-MRCR measures retrieval of specific information from long context, where the margin over Qwen3.5 at 93.0 and DeepSeek-V4 Pro max at 92.9 is narrow but the gap to GLM-5.1 at 71.5 reaches 22.9 points. KGC-Safety and ROK-Fortress evaluate Korean-language safety and norm compliance, and there the margins over GLM-5.1 at 69.3 and DeepSeek-V4 Pro max at 47.6 exceed 30 points.

These three share a property: each is either designed around domestic criteria or was targeted explicitly during training. ROK-Fortress moved from 60.9 in the previous generation to 89.5, a 28.6-point gain that reads as deliberate investment rather than a byproduct of scale. KGC-Safety at 99.8 is effectively saturated, leaving no headroom to claim in a future release.

For the same reason, this strength does not transfer. A lead on Korean safety benchmarks is decisive in domestic public-sector and financial procurement that adopts those benchmarks as criteria, and gives an overseas developer no reason to choose the model. OpenAI-MRCR at 94.4 is the exception, because long-context retrieval is language-independent and therefore the one advantage that travels. Compressed to a single line, this model's external competitive claim rests on one benchmark.

What a 3.2x parameter increase bought, and what it did not

K-EXAONE 2.0 scales the previous model's 236 billion total parameters to 750 billion, a factor of 3.2, and its 23 billion active parameters to 37 billion, a factor of 1.6. Only 4.9% of parameters fire per token, roughly one in twenty. The stack runs 78 main layers, 2 dense and 76 sparse, plus 1 multi-token-prediction layer, and routes to 8 of 256 experts alongside 1 shared expert. Context length is 262,144 tokens, vocabulary is 153,600, and the knowledge cutoff is Q2 2025.

Summing the per-benchmark changes shows where the increase went. Across all 24 rows the total gain is 152.3 points, and the top five rows account for 116.9 of them, or 76.8% (calculated by ASAP). OpenAI-MRCR leads at 52.3 to 94.4, a gain of 42.1, followed by ROK-Fortress from 60.9 to 89.5, SWE Bench Verified from 49.4 to 68.2, PolyMath from 57.4 to 71.3 and Terminal-Bench 2.1 from 30.3 to 43.8.

Four rows moved backwards. HMMT Feb 2026 fell from 80.7 to 78.4, HRM8K-KSM from 91.9 to 91.1, MMLU-Pro from 83.8 to 83.5 and GlobalMMLU-Lite from 86.9 to 86.6. The tool-use task τ³-Banking is unchanged at 14.2 across both generations.

The distribution says that 3.2x the parameters bought a specific bundle of abilities rather than general capability. Long-context retrieval, code repair, terminal operation and Korean safety rose; world knowledge and competition mathematics stalled or slipped. MMLU-Pro landing 0.3 points below a model one-third the size deserves attention on its own. Parameter count alone does not move knowledge benchmarks, and this generation's gains came from post-training design and data composition instead. The practical implication is that parameter count is a poor predictor of what the next release will do.

A sovereign model finishes last on Korean knowledge evaluation

K-EXAONE 2.0 scores 69.1 on KMMLU-Pro, below Qwen3.5 at 77.4, GLM-5.1 at 75.8 and DeepSeek-V4 Pro max at 80.5 in the same table. The other Korean row, CLIcK, places it at 84.2, outside the 88.7 to 91.6 band the three comparison models occupy. Of the three Korean rows, it beats a comparison model on only two: HRM8K-KSM at 91.1 against GLM-5.1's 89.4, and Ko-LongBench at 89.6 against GLM-5.1's 83.6.

This result invites a second look at the premise of Korea's National AI Foundation Model project. That premise holds that a domestic model performs better in Korean and in Korean context, while the card's own table reports three Chinese open-weight models ahead on Korean knowledge evaluation. The Korean domain where the domestic model does lead clearly is safety and norms, not knowledge. What domestic development delivered is closer to "better aligned to Korean standards" than to "knows Korea better."

That distinction bears directly on adoption decisions. For a service where Korean question-answering accuracy is the core requirement, this table does not put K-EXAONE 2.0 first by default. For public-sector and financial deployments where domestic norm compliance and safety filtering are procurement requirements, KGC-Safety at 99.8 and ROK-Fortress at 89.5 open gaps in the 30-point range over the alternatives. The same model earns very different grades by use case, which means the sentence "a Korean model beat the foreign ones" supports no decision at all.

One further caveat applies to the whole table: every figure is self-reported. SK Telecom's A.X K2, released two days earlier on July 29, 2026, self-reports 80.5 on a benchmark of the same name, KMMLU-Pro. Placing the two scores side by side would require matched thinking budgets, prompts and grading procedures, and the two model cards do not share those conditions. The absence of a common evaluation regime for Korean models submitted to the same national program is itself a problem this release exposes.

What the switch to Apache 2.0 actually opens up

K-EXAONE 2.0 ships under Apache 2.0, a change from the previous model, which was distributed under the bespoke K-EXAONE AI Model License Agreement. The earlier license required separate agreement, while 2.0 permits commercial use and derivative distribution with no further negotiation. Language support also widened from six, Korean, English, Spanish, German, Japanese and Vietnamese, to ten with the addition of French, Italian, Portuguese and Polish.

The practical effect of the license change is the entry condition for domestic development teams. Few Korean organizations can serve a 750-billion-parameter model directly, but Apache 2.0 removes the legal obstacle to distilling from it as a teacher model or building smaller domain-specific derivatives. FP8 and NVFP4 quantized builds plus a DSpark variant landing alongside the original on day one point the same direction. The real reach of this release is decided by the smaller derivatives it enables, not by the 750-billion-parameter model itself.

Three things left to verify

Judgment on K-EXAONE 2.0 will settle on three indicators over the second half of 2026. The first is third-party reproduction. Apache 2.0 open weights let outside researchers rebuild the 24-row table, and whether OpenAI-MRCR at 94.4 and ROK-Fortress at 89.5 survive independent evaluation determines the credibility of the whole table. The configuration behind the announced 70.1 average belongs on the same list.

The second is the outcome of the National AI Foundation Model project. K-EXAONE 2.0 was submitted to that program, which also fields competitors including SK Telecom's A.X K2. Whether the selection criterion is overall performance, Korean specialization, or safety and norm compliance changes which row of this table decides the result. On the evidence available, the third criterion is where K-EXAONE 2.0 stands strongest.

The third is how quickly derivatives appear. A 750-billion-parameter footprint is too heavy for most Korean organizations to run as released, and Apache 2.0 makes compression and specialization straightforward. Whether those two conditions produce deployable smaller models within months is what settles this release. That single question matters more to the domestic ecosystem than any of the 24 rows in the benchmark table.

Source: LG AI Research official model card LGAI-EXAONE/K-EXAONE-2.0-750B-A37B (Hugging Face, Apache 2.0, July 31, 2026), with reporting from The Korea Herald and The Korea Times, compiled by ASAP

ASAP — AGI Soon As Possible

AI & tech,
read in depth

Beyond the headlines — into the context and the structure

AGI Soon As Possible · asapai.co.kr

← All posts