AGI Soon As Possible · Deep reads on AI & tech
Article

Anthropic handed Claude an entire protein binder design campaign through one prompt, and 354 binders came back across 14 of 15 targets

2026-08-21 · 10 min read

Anthropic's technical report of August 18, 2026 states that Claude Opus 4.8 and Mythos Preview are capable of running 24- to 48-hour de novo binder design campaigns against 16 targets with no human input into any design decision, and that 354 of 1,320 designs bound across the 15 targets that yielded interpretable measurements. The overall hit rate is 27%, and among the designs each campaign ranked first for its target, 49% bound. On the E3 ligase subunit RBX1, where a recent open competition saw 9 of 245 de novo designs bind, 28 of Claude's 90 designs bound, and the tightest reached a KD of 3.9 nM against 45 nM for the competition's winning entry re-synthesized on the same plate. ASAP walks through the experimental design and the numbers, then argues that the most consequential detail is not the hit rate but the fact that two-thirds of the 16,000-word prompt was not about science.

What the humans did was name the targets and fund the compute

Human involvement in these campaigns is limited to three points: naming the 16 targets by name, UniProt accession, organism, and oligomeric state; funding a cloud GPU account with USD $50,000 or $10,000 per campaign; and placing synthesis orders for the designs that came back. Which region of the target to model, which surface to bind, which tools to use in what proportion, and which designs made the final ranked 30 were all decided by Claude.

Campaigns ran in two formats. Multi-target campaigns lasted 48 hours and covered 14 targets at once, once per model with Claude Opus 4.8 and Mythos Preview. Single-target campaigns lasted 24 hours, with Mythos Preview running against all 16 targets and Opus 4.8 against TNFα, latent GDF-8, and mature GDF-8. Opus 4.8 ran at maximum reasoning effort and Mythos Preview at the high setting.

No design tool was pre-installed. The protocol instructed Claude to build each one from its public repository into a container image and validate it within the first hour. When a campaign closed, it returned 30 ranked designs per target, and two contract research organizations, Adaptyv Bio and Twist Bioscience, synthesized every design exactly as delivered and measured binding. Orders went out with sequences shuffled and with no indication of model, campaign, or rank.

One target dropped out of the analysis entirely. The mature GDF-8 dimer aggregated under assay conditions and gave no interpretable measurement at either CRO, so its 120 designs were excluded, leaving 15 targets and 1,320 designs.

Two-thirds of the 16,000-word prompt is operations, not protein science

The protocol prompt Claude received is roughly 16,000 words, and two-thirds of it covers not protein science but how to pace a fixed GPU budget against the wall clock, how to delegate work to a two-layer team of sub-agents and supervise it, how to verify results before reporting them, and when to report. Anthropic iterated these sections over test campaigns until Claude reliably sustained 24- to 48-hour runs, then froze them before the reported experiments.

That split is the most practically useful fact in the report. The remaining third, the scientific knowledge, is already public in papers and documentation; what made an agent survive 24 to 48 hours alone was the operational layer that governs burn rate, delegation, and self-verification. The only human messages that entered any campaign were short, non-technical instructions to resume after a session died for infrastructure reasons, and Claude detected and worked around the remaining infrastructure incidents on its own.

How the prompt names tools is equally hands-off. It lists license-cleared, installable options by category and requires seven designated structure-generation methods to contribute at least 50 backbones each to every target's pool. Beyond that, Claude picked. Per-design provenance in the release shows PXDesign produced 358 of the ordered designs, RFdiffusion3 267, Genie 3 185, and FreeBindCraft 135.

On RBX1 the open competition bound 9 of 245 and Claude bound 28 of 90

RBX1 results are the sharpest head-to-head in the report: an open design competition bound 9 of 245 de novo designs, while 28 of the 90 designs Claude delivered across three campaigns bound. Claude's tightest binder measured a KD of 3.9 nM, against 45 nM for the competition's winning entry re-synthesized and measured on the same plate.

Per-target performance varied enormously. In Mythos Preview's multi-target campaign, TREM2 yielded 72 binders out of 90 designs, VEGF-A 54 of 90, and IL-7Rα 49 of 90. At the other end sit BBF-14 with 3 of 90, 15-PGDH with 1 of 30, and MBP with 0 of 90.

Format mattered too. In the 48-hour multi-target format, Mythos Preview hit 26.7% and Opus 4.8 hit 22.6%. In the 24-hour single-target format, Mythos Preview reached 35.1%, a figure the report pairs with the note that this format spends roughly 2.8 times the compute budget per target.

Six targets had prior open competitions: TREM2, RBX1, latent GDF-8, Nipah G, EGFR, and 15-PGDH. Claude's designs bound at higher rates than the field on four of them. On TREM2 the comparison is 72 of 90 for Claude against 36 of 94 for the hackathon entries, though counting exactly as the organizers did, by Adaptyv Bio's measurement alone, lowers Claude's TREM2 count from 72 to 65.

Cross-species reactivity was only a secondary objective of the prompt and still produced numbers. Of 233 binders tested against the mouse ortholog of their target, 130 also bound it, including 68 of 69 on TREM2 and 32 of 53 on VEGF-A.

The three failed targets say more than the successful ones

MBP, the maltose-binding protein of Escherichia coli, is the single target where Claude produced no binder at all, and all 90 designs from three campaigns failed. A known MBP binder run as a plate control bound normally at Adaptyv Bio, which places the failure in the design rather than the assay.

The second hard case is BBF-14, a 110-residue β-barrel that does not occur in nature and was itself designed de novo. Claude bound only 3 of 90 designs against it, and 15-PGDH yielded just 1 binder from 30 designs.

What makes these three targets important is that the failures were not predictable in advance. After the campaigns closed, Anthropic re-scored every ordered design with ten public co-folding predictors, and designs against MBP, BBF-14, and 15-PGDH scored nearly as high as designs against VEGF-A and Nipah G, which produced binders in bulk. Co-folding confidence separated binders from non-binders within a target's ranking but did not flag which campaigns would fail.

That reshapes the cost model. If computation alone cannot tell a good campaign from a doomed one, the true cost of an autonomous campaign includes synthesis and measurement, not just the GPU budget. It also makes the multi-target format look like the more practical default: when failures cannot be screened out beforehand, spreading a budget across targets loses less than concentrating it on one.

Why the 27% headline requires the caveats printed alongside it

Anthropic's stated comparison point is the 10 to 15% hit rate typical of protein design campaigns today, which means this campaign's 27% is roughly double the top of that range. Four caveats, all of them in the report itself, govern how far that comparison travels.

The first is the counting unit. The 354 figure counts sequences, and sequence variants from one backbone are not independent. Counting only the best-ranked sequence from each of the 809 generated backbones gives 200 binders, or 24.7%.

The second is replication. Each combination of model, format, and target ran once. Model, format, and run-to-run variation are therefore confounded, and the gap between Mythos Preview at 26.7% and Opus 4.8 at 22.6% is not evidence of a model difference. Anthropic states plainly that it has described campaigns rather than models.

The third is target selection. Most targets tested are extensively characterized in the literature, so extending this hit rate to novel targets with thin structural records is a claim the data does not support.

The fourth concerns what was measured. Five oligomeric targets, TNFα, VEGF-A, Nipah G, latent GDF-8, and 15-PGDH, present a multivalent analyte, their KD values are apparent rather than true, and they account for 100 of the 354 binders.

The bottleneck moved from the model to human time

Every design and structure-prediction model Claude used in these campaigns is open-source, and none was pre-installed. Anthropic's concluding emphasis is accessibility rather than performance: one frozen protocol served all 16 targets, which means laboratories that have targets of interest but no computational protein design expertise are within reach of running campaigns like these.

Attaching costs sharpens the claim. A multi-target campaign consumed USD $50,000 of cloud GPU over 48 hours, and a single-target campaign $10,000 over 24 hours. For a campaign covering 14 targets at once, per-target compute lands under $4,000. Set against the days to weeks a human operator spends orchestrating a single campaign, that figure reads as labor converted into compute.

What changes is the kind of expertise a team needs. Installing tools, choosing epitopes, and setting filter thresholds moved into the prompt, which leaves target selection, interpretation of binding data, and synthesis budget allocation on the human side. For a biotech startup, the practical implication is to invest in target hypotheses and a CRO measurement pipeline before building a computational design team.

The accessibility is not symmetrical, however you read it. The expertise required did not disappear; it moved into the prompt, and Anthropic wrote and tuned those 16,000 words over test campaigns. Because the prompts are released, reproducing the protocol is cheap, but adapting it to a new target class or a new design objective returns the work to experts.

Binders are not drugs, and the report says so in its first limitation

The chief limitation Anthropic records is that the evidence here is binding, not structure or function. No design was tested for activity or solved structurally, every pose in the report is a prediction, and affinities on the five oligomeric targets are apparent values.

The release is unusually complete. Anthropic published the prompts, kickoff messages, and document corpus, computational models of all 1,440 designs with Claude's ranks and per-design provenance, and both CROs' raw binding data for the 1,320 designs with reliable measurements, on Hugging Face under CC BY 4.0 for data and MIT for scripts. Every binding statistic in the report is therefore recomputable by a third party.

Three questions are left open. One is whether the same hit rate holds on novel targets with thin literature. Another is whether binding translates into function, meaning whether these binders actually block or modulate their targets. The last is variance: until the same protocol runs repeatedly against the same target, a single campaign's hit rate is one observation rather than an expected value.

What the report does establish is unambiguous. An agent given only target names installed its own tools, chose its own epitopes, filtered and ranked its own candidates, and delivered 1,320 designs of which 354 bound. An entire workflow that consumed days to weeks of expert time was replaced by one prompt and two days of GPU budget, and the fact that this format worked at all defines the next set of questions.

Source: ASAP analysis based on Anthropic's technical report "Autonomous de novo protein binder design with Claude" (Amir Shanehsazzadeh, August 18, 2026) and Anthropic's official research post

ASAP — AGI Soon As Possible

AI & tech,
read in depth

Beyond the headlines — into the context and the structure

AGI Soon As Possible · asapai.co.kr

← All posts