Claude Protein Design: What 14 of 15 Actually Means
Fourteen of fifteen targets is a coverage figure, not a success rate. The report's real result is 354 confirmed binders from 1,320 designs, and a campaign structure worth copying.
On 20 August 2026 Anthropic published results from a protein design campaign, reporting that Claude produced binders against 14 of 15 targets, with 354 confirmed binders out of 1,320 total designs. Adaptyv Bio and Twist Bioscience synthesised and tested the designs in wet labs without modifying them. The write-up is on Anthropic's research site.
The 14 of 15 figure is the one that travelled, and it is the least informative number in the report. It says the model got at least one working design against almost every target given enough attempts. With 1,320 designs across 15 targets, that is roughly 88 shots per target. Breadth of coverage under that budget is a much weaker claim than it sounds.
Three numbers, doing three different jobs
Number | What it measures | How much weight it carries |
|---|---|---|
14 of 15 targets | Coverage: at least one hit somewhere | Low, given ~88 designs per target |
22.6% to 35.1% hit rate | Efficiency per design attempt | High, this is the comparable figure |
354 confirmed binders | Absolute output of the campaign | Medium, scales with budget |
6 targets with high affinity | Whether hits are strong enough to matter | High, affinity is the gate |
The hit rate is where the argument actually is. Anthropic reports 26.7 percent for Mythos Preview and 22.6 percent for Opus 4.8 in the multi-target setting, rising to 35.1 percent for Mythos Preview when given a full 24 hours per single target, against a stated industry baseline of roughly 10 to 15 percent. Doubling a hit rate is a real result if the baseline is fair. That is the load-bearing assumption, and it is one a reader outside the field cannot check.
Affinity is the second gate. A binder that attaches weakly is a laboratory observation; a binder that attaches tightly enough to do something is a starting point for a drug. The report claims high-affinity binders against at least six targets and designs matching or exceeding best published affinity against at least four. Those are the sentences to hold onto.
What the Claude protein design campaign did not show
Maltose-binding protein produced nothing. Ninety designs, zero confirmed binders. The report says so plainly, and it is the most useful paragraph in the document, because it shows the capability is uneven rather than general. A system that succeeds broadly and fails completely on one target is behaving like a system with structure to its competence, not a system that has solved the problem.
Anthropic also says it intends to follow up with more extensive characterisation. Read that as: these are binding results, not efficacy results, and the distance between the two is where most drug candidates die.
The part worth copying, if you build agents
Strip out the biology and the campaign is a template for running an agent against an expensive external grader. The setup, as described in the report:
A prompt of roughly 30,000 tokens, which is a specification document rather than an instruction.
Internet access plus a curated corpus of papers, so the agent could ground itself in the literature.
Connectors to Google Drive, Slack, Gmail and BioRxiv, giving it the same information surfaces a human researcher would use.
GPU access for specialised models, so the agent could call domain tools rather than reason its way to a structure in prose.
A hard compute budget, up to 12,500 H100 hours for multi-target runs or 2,500 per single target.
A physical laboratory as the grader, which cannot be gamed, flattered or reward-hacked.
That last item is the one most agent projects lack. The reason evaluation is hard for agents is that the grader is usually another model or a human skimming output, and both are cheap to fool. A wet lab costs money and time, and returns a verdict that owes nothing to how persuasive the agent sounded. If you are building anything agentic, the question the campaign raises is not whether you can afford a lab. It is what your equivalent of one is, and our notes on what an AI eval actually is are the practical entry point to that.
The compute budget matters too. Twelve thousand GPU hours to design proteins is a plausible research spend and an implausible product spend. Anyone reading this as a template for their own agentic AI system should note that the interesting behaviour appeared under a budget most teams do not have.
How to read vendor-published results generally
This report is stronger than most because the validation was physical and external. Adaptyv Bio and Twist Bioscience made and tested the designs; the model did not grade its own homework. That is a real structural advantage over a benchmark score, and it is worth saying so.
It still arrives with the usual caveats attached to anything a lab publishes about itself: the baseline comparison is chosen by the party being compared against it, the targets were selected, and one inconclusive target was dropped from 16 to 15. None of that is misconduct. It is the ordinary shape of a vendor result, and the same reading habits apply as with inflated AI benchmark claims and with model system cards: find the denominator, find the baseline, and find the thing that failed.
If you follow model releases at all closely, our guide to keeping up with AI news covers the triage version of this, which is deciding in ninety seconds whether a result deserves the twenty minutes.
Frequently asked
Did Claude design a drug?
No. It designed binders, meaning proteins that attach to a target. Binding is the first step of a very long process. Nothing in the report claims a therapeutic candidate.
Which targets were involved?
The report names clinically significant targets including PD-L1, TREM2, TNF alpha and EGFR, alongside maltose-binding protein, where all 90 designs failed.
Who verified the results?
Adaptyv Bio and Twist Bioscience, two contract laboratories, synthesised and tested the designs physically. Anthropic ran the design campaign and published the write-up.
Is a 22 to 35 percent hit rate good?
Against the 10 to 15 percent baseline the report cites, yes, roughly double. Whether that baseline is the right comparison for these specific targets is the question a specialist would push on, and the report is the only source for it.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


