Skip to main content

Claude Found a CRISPR-Like DNA System—What Do We Actually Know?

Claude Found a CRISPR-Like DNA System—What Do We Actually Know?

I open a lot of AI announcements and close most of them inside a minute. This one I sat with — not because of the headline ("Claude discovered a CRISPR-like enzyme system"), but because of one line in a session log that Anthropic published along with it, and because a company that spends a lot of energy warning about AI risk just shipped the first result out of an actual molecular biology lab it built in the Bay Area.

So I downloaded the 40-page preprint, read the methods, and went through the reviews. Here is the line from the agent's transcript first, because it may be my favorite thing I've read all year:

“[The DNA next to the RT] is spectacular: I can see by eye a tandem repeat array … that's a CRISPR-like … repeat array?!”— Claude (Mythos 5), during the search, as quoted in Anthropic's announcement

Now the caveat every headline skipped: Anthropic does not know what this system does. The preprint says it plainly — they have not shown the enzyme is active, and they have not shown that the RNAs this array produces are its working material. There is no gene editor here yet. The scientists who reviewed it are split between "genuinely intriguing" and "nothing indicates this is a rival to CRISPR," and both of those are true at the same time, which is what makes the whole thing interesting.

So here is where things actually stand:

Item What's actually true right now
What was found ART — "array-associated reverse transcriptases": a reverse transcriptase gene, a partner gene of unknown function, and an array of 3–21 short repeating sequences. Found mostly in bacteriophage DNA.
Who found it Claude agents ran a 21.5-hour campaign — 949 sessions, ~215.6M tokens. The brief was written by humans; all wet-lab work was done by humans.
What the lab showed The array is transcribed into discrete short RNAs — up to 8% of all phage RNA 15 minutes into an infection. Reproduced in E. coli from plasmids.
What nobody has shown That the enzyme is active at all; that those RNAs are its substrates; that ART cuts, copies, or pastes DNA; what it does for the phage.
What was already known The reverse transcriptase itself — earlier genome studies had described it. The repeats and the partner gene went unnoticed until this report.
Status Preprint + technical report, not peer reviewed. Anthropic says experiments are ongoing.
The catch Anthropic reran the whole hunt ten more times. All ten missed the array.

So what is ART, exactly?

A reverse transcriptase (RT) is an enzyme that copies RNA back into DNA. It's one of the workhorse enzymes of molecular biology — cDNA synthesis, a lot of genome-engineering tooling, and half the reason retroviruses work. In bacteria, RTs rarely work alone: they typically sit next to a partner protein and a non-coding RNA in the same little cassette. Retrons, for example, use an RT to write a dedicated small RNA into DNA, which then teams up with an effector protein to fight off phages.

The system Claude found — ART — has three parts, and it's the combination that's unusual. There's the RT itself, which carries a long stretch of about 180 extra residues before its polymerase domain, where most RTs have fifty or fewer. There's a partner gene of unknown function right beside it. And then there's the part that made the agent stop mid-screen: a long array of short DNA repeats, evenly spaced, sitting upstream of the enzyme.

The numbers, from the preprint: 95 distinct ART-family RT clusters found in jumbo phages and predicted viral contigs, 28 of which carry a detectable repeat array. The arrays run from 0.3 to 4.1 kb and hold 3 to 21 copies of a repeat that's 15–49 nucleotides long, each with a palindromic core of about 15 nucleotides. The spacers between copies run 120–220 nucleotides, and importantly, they're unrelated to each other within one array — which is how CRISPR holds its bank of different guides.

And the phrase everyone ran with — "CRISPR-like" — is about that layout, not the mechanism. There are no cas genes anywhere near these loci. No Cas9, no cutting machinery, nothing that says "editor." It looks like a CRISPR array the way a filing cabinet looks like a filing cabinet, which is not the same as knowing what's in the drawers.

How the hunt actually ran

The team gave Claude one prompt: search a database of 1.9 billion protein clusters for interesting new reverse transcriptase systems. That was the human contribution, besides the lab work. Agents running on Claude Mythos 5 did the rest — one agent planning and running each task, a second reviewing it, and the agents opening new follow-up tasks as they went.

The campaign log: 949 agent sessions, 21.5 hours of wall-clock time, ~215.6 million tokens. Along the way the agents recovered roughly 200,000 RT clusters, classified them into nine RT classes, scored 3,564 candidate protein families, and filed 19 reports for human review. Anthropic's own framing of what that replaces: "For an expert scientist, this type of analysis can take weeks to months of work."

Bar chart on a log scale showing the search narrowing from 1.9 billion protein clusters to about 200,000 RT clusters, 3,564 candidate families, 19 reports, and one system described
The campaign ledger, plotted: 1.9 billion clusters in, one system described. My chart, from the numbers in the technical report — note these are different kinds of things (clusters, families, reports), not one nested funnel.

ART came out of a side path. It wasn't one of the top candidates with a polished report — a worker agent, reading raw DNA next to an unusual enzyme, noticed the repeats "by eye" and pulled on that thread. Which is a good moment to note that the same lab published the sessions, so you can read the hunt yourself instead of taking a summary's word for it.

Figure 2 from the preprint: diagram of the path agents took from a census candidate to the ART locus, plus upstream regions of related systems
The path from census candidate to ART locus, from the preprint itself — including the step where the worker reads the upstream gap and the repeats jump out. image: Anthropic technical report (preprint)

The moment it noticed

The flank the agent was reading is called L0050 — 2,900 base pairs of ordinary-looking As, Cs, Gs and Ts lifted out of a metagenomic dataset. This is what it was staring at when the exclamation happened:

Animated graphic of the raw DNA sequence the agent was reading, a FASTA-style block of 2,900 nucleotides labeled L0050 upstream region
The actual screen: the raw DNA the agent was reading when it spotted the repeat pattern. No one had flagged this region before. image: Anthropic

What's interesting is what happened after. It didn't just name the pattern and move on. It counted the repeats, measured their spacing, compared the layout against known systems, searched the literature for prior descriptions, and filed a structured finding — a claim, its evidence, and a confidence score between 0 and 1. That's the workflow of a careful grad student, executed without anyone watching.

I keep coming back to how mundane the input was. This wasn't a model spotting a pattern in a vector database or a pre-computed alignment. It was an agent reading a text file of A/C/G/T and noticing that part of it repeats on a rhythm — and then having the taste to chase it. (Turns out that last part is the fragile bit. More below.)

What the wet lab actually showed — and the sentence that matters

The computational find is the fun part; the wet-lab verification is the part that keeps it honest. And there's a real experiment here. Using published RNA-seq data from a phage infection — jumbo phage SA1 infecting Staphylococcus lentus, sampled at 5, 15 and 55 minutes — Anthropic's scientists found that the ART array is highly transcribed. At 15 minutes in, the array-derived RNAs account for up to 8% of all phage RNA, making them among the most abundant transcripts in the infection. Within the array, the transcripts resolve into discrete, short species whose boundaries are reproducible across replicates.

Figure 6 from the preprint: RNA-seq coverage over the SA1 ART locus at 5, 15 and 55 minutes after infection, transcript abundance in TPM, and predicted RNA secondary structures
Fig. 6: RNA-seq across the infection. The array RNA (blue) peaks near 10⁴ TPM at 15 minutes — an order of magnitude above the enzyme's own mRNA. image: Anthropic technical report (preprint)

They also expressed the SA1 ART system on plasmids in E. coli — both the native locus and a heterologous-promoter construct — and saw the same discrete short RNAs, which rules out "this is just infection noise." So: transcription, confirmed. The array produces a set of distinct RNA pieces, in large excess over the enzyme itself.

A gloved hand in a molecular biology lab holding a small tube of orange liquid
Where the verification half happened: Anthropic's Bay Area lab — BSL-1/BSL-2 only, no human pathogens, and every experiment run by human scientists. image: Anthropic

Now, the sentence from the preprint's discussion that I'd put in every headline:

“Beyond these observations, we have not shown that the RT is active or that the unit RNAs are its substrates. Whether the RT and its partner interact, and what the system does for the phage, are currently unknown.”— Anthropic technical report (preprint, not peer reviewed)

Read that carefully, because it draws the line exactly where the hype doesn't. What's verified: the array is transcribed, the RNAs are real and abundant, the arrangement is new. What's not: that the enzyme does anything. They haven't shown catalysis. They haven't shown the RNAs are what the enzyme works on. Nobody knows what this system does for the phage it lives in — and Anthropic says its "work to understand the primary function of ARTs is ongoing."

Is it CRISPR? No — but that's not the interesting question

The "next CRISPR" framing is what the market reacted to and what most headlines implied, so let's do it honestly. Here's the layout of an ART locus, drawn the way the preprint draws it:

Dark schematic of the ART locus: an array of 14 short repeats with spacers, a partner gene, and a reverse transcriptase gene with a long N-terminal domain
The ART locus: repeat array upstream, then the partner gene, then the RT. My schematic, redrawn from the preprint — not to scale. The repeat array is the CRISPR-looking part; everything else is not.

Kevin Blake, a microbiologist at Washington University School of Medicine, gave the bluntest version of the skeptical case to Al Jazeera: "Because the identified array is 'CRISPR-like', some have leaped to conclude Anthropic discovered the 'next CRISPR' — ie, a Nobel Prize-winning gene editing technology." His point is that CRISPR the technology — the editing tool — is a very different thing from CRISPR in nature, which is basically a bacterium's immune system. And there's a deeper point underneath: there are countless CRISPR-like sequences out there that nobody has ever catalogued, because we've studied a tiny fraction of the microbial world. Finding one more odd arrangement is not, by itself, a breakthrough.

"There's nothing to indicate this is a rival to CRISPR-the-technology, or could be developed into any kind of therapeutic or practical application," Blake said. And strictly speaking, that's correct. No one has shown ART edits anything, or is even enzymatically active.

But — and I think this is where the interesting question lives — look at how CRISPR itself actually happened. The repeats were first noticed in 1987 in an E. coli gene nobody was curious about. It took 20 years to figure out that they're an immune system, and 25 to turn them into an editing tool. Restriction enzymes: 17 years from the first odd observation to the first purified enzyme. Taq polymerase and PCR: 19 years from Yellowstone bacterium to a standard lab technique. Every one of these is "odd repeats somebody almost ignored" at the start.

Timeline chart showing restriction enzymes 1953 to 1970, Taq/PCR 1969 to 1988, CRISPR 1987 to 2012, and ART found 2026 with function still unknown
From odd sequence to lab tool: 17, 19, and 25 years for the three famous systems. ART is at year zero, on a preprint. My chart — dates from the standard histories of each system.

So the right question isn't "is ART the next CRISPR." It's "is ART at 1987?" — a strange, unexplained arrangement in a genome, found before anyone knows what to do with it. If it is, we won't know for years. And the reviewers who are positive seem to be saying exactly that. Stanley Qi, a bioengineering professor at Stanford, called it "incredibly exciting" — and pointed at why: "What stands out is its ability to recognize an unusual biological pattern that was difficult to detect before, and to pursue it comprehensively as a research question."

The catch: ten reruns, zero repeats

Here's the part I respect the most, because it would have been so easy to leave out. The authors asked themselves whether the discovery was reproducible inside their own harness, and ran the same campaign ten more times. Nearly every rerun sampled ART loci during the census. In two of them, workers even investigated the lineage as follow-up. But not one of the ten read the DNA upstream of the enzyme — and all ten missed the array entirely.

The authors blame the size of the search and the agents' own unpredictability. And they went further: in fixed benchmark tests, their four most capable models (Opus 5.5, Mythos 5.1, Mythos 5, and Opus 5) described the array accurately in at least 90% of attempts when the DNA was placed directly in context — but with ordinary files and tools, the rate fell as low as 32%. Digging into the transcripts: 39% of attempts never read a contiguous stretch of 200 nucleotides or more, meaning they never saw more than about one repeat unit. When a model did read at least 200 nucleotides, its recognition rate jumped 16 to 32 percentage points.

I find this the most useful result in the whole package. It says: the model can do this, but only if it actually reads the DNA — and left to its own devices, a third of the time it doesn't. That's a humility footnote on the "AI discovered it autonomously" banner, and it's also the most fixable problem in the pipeline.

One more layer, and this one is genuinely unusual. Anthropic dug into Mythos 5's internals while it read the sequence, decomposing the model's activity into individual signals, and found two that fire on tandem repeats — the paper calls them "repeat-signal 1" and "repeat-signal 2." Both were silent before the array. Repeat-signal 2 strengthened by the third repeat copy and stayed on afterward. When the researchers shuffled the nucleotides of each copy in place, the signals went quiet — repeat-signal 1 silenced on 12 of the 14 copies, repeat-signal 2 on all 14. In other words, they found the machinery behind the noticing: a repeat detector, and it was firing right before the agent wrote its little exclamation.

The paper frames noticing an anomaly as the start of scientific discovery. Going inside the model and identifying the specific internal signals behind that anomaly-detection is the closest thing I've seen to showing *how* the machine does research, rather than just that it did.

What the scientists actually said

The strongest endorsement doesn't come from Anthropic. It comes from Feng Zhang — one of the pioneers of CRISPR genome editing at MIT and the Broad Institute — who reviewed the preprint before publication:

Portrait photo of Feng Zhang in a dark jacket and glasses
Feng Zhang reviewed the preprint. His verdict: "genuinely intriguing and merits further investigation." image: Wikimedia Commons / CC BY-SA 4.0
“This is an exciting example of how AI agents can contribute to biological discovery. The identification of RNA-repeat arrays associated with reverse transcriptases is genuinely intriguing and merits further investigation. I hope this work encourages more scientists to explore how AI can support their research.”— Feng Zhang, MIT / Broad Institute

That's a careful, deliberately limited endorsement, and I mean that as a compliment. He's praising the identification and the method, not the biology — "merits further investigation" is scientist for "call me when there's data."

Dario Amodei's own post on X is unusually well-calibrated too, for a CEO: "Its precise function, biotechnological utility (if any), or level of significance is not yet clear, but at minimum it is work I would have been proud to do as a PhD student." He also confirmed the division of labor — the work was done "mostly, though not entirely, by Claude" — and said the crew is not letting the model run experiments: "Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment, with appropriate safeguards in place, but we aren't doing that today." The lab operates only at BSL-1 and BSL-2 and handles no human pathogens, per Anthropic.

Portrait photo of Dario Amodei, CEO of Anthropic, in a dark shirt
Dario Amodei: "at minimum it is work I would have been proud to do as a PhD student." image: TechCrunch / Wikimedia Commons (CC BY 2.0)

He also gave the tell on prior work: the enzyme itself was identified in previous studies, and a Stanford team had recently described a similar — but distinct — RT system with a non-coding array, evolved independently. What's new is the specific combination, and who noticed it. Amodei: "mostly, though not entirely, by Claude."

Wall Street read the headline, not the preprint

On Wednesday, gene-editing stocks slid as the news crossed: CRISPR Therapeutics down about 6%, Beam down 6%, Editas down 8%, Intellia down 3%, and Prime Medicine — after a bounce the day before — down roughly 12%. Retail sentiment on the tickers, per Stocktwits, flipped bearish on some names. That's the market pricing in a competitor for a tool that, on the evidence, doesn't exist yet. It's also, honestly, a pretty good real-time gauge of how far the framing ran ahead of the paper — the preprint doesn't even claim the enzyme is active.

What I'm watching next

If ART is real, here's the sequence of things that should happen over the next few months. First: peer review, or at least independent labs taking the reported sequences and checking the family. Second: somebody purifies the enzyme and demonstrates catalysis — that's the single most important missing piece, and it's a standard experiment. Third: the substrate question — whether those short RNAs are actually what the enzyme works on. And either way, I want to see the follow-up on the reproducibility problem, because "one lucky agent in twelve runs" is the real bottleneck between a fun story and a discovery pipeline.

My take

Not a gene editor. A lead with an asterisk — and in this case the asterisk is where the actual story is.

What Anthropic shipped this week is two things at once. The biology is a maybe: an odd, unexplained arrangement around a known enzyme, with clean RNA evidence and zero functional evidence, sitting in a preprint. Could be a future tool, could be a footnote, and nobody — including the people who found it — can tell you which yet.

The process is the more interesting half, and it's where I'd spend my attention even if ART turns out to be nothing. The hunt was auditable: session logs published, a technical report with the exact token and session ledgers, a ten-run reproducibility study that reports its own failure, and an interpretability pass that found the repeat-detector inside the model. I've read a lot of "AI is doing science now" posts this year. This is the first one where I could go and check the receipts — including the receipts for the parts that didn't work.

Amodei draws the curve from math to biology and thinks the same exponential is coming for the lab. That's a forecast, not evidence; he'd be the first to tell you the wet lab doesn't scale like a GPU cluster. But the boring parts of this package — an agent reading raw DNA character by character, filing confidence-scored reports, the rerun study — are what a real research assistant looks like, and that part already works well enough to find something twelve straight runs couldn't.

I'll post an update when something changes: when the paper gets reviewed, when someone shows the enzyme doing anything at all (or shows it doesn't), or when another lab digs up more of these systems. Any of the three might happen first.

Sources and notes. Everything here is quoted from primary sources and reporting I read directly — I didn't run any experiments, and the only numbers I derived myself are the three charts, which are labeled as mine. Primary: Anthropic's announcement (Sep 23, 2026) · the technical report / preprint (not peer reviewed) · Anthropic's X post · Amodei's X post. Reporting: Al Jazeera (Qi, Blake quotes) · TechCrunch · The Next Web (rerun study, fixed-test numbers) · METAL · Stocktwits coverage of the gene-editing tape.

Related on this blog. Claude Opus 5.5: 40% Cheaper — Except at Max Effort — one of the four models that passed the preprint's own array-recognition benchmark. · Pace the Frontier: AI's Rivals Agree They Should Slow Down — the safety debate this story keeps colliding with. · Claude Code Weekly Limits Fell 17% — the tool the agent campaign actually ran on.

Comments