Five years before Anthropic announced that Claude had found a strange new biological system hiding in viral DNA, scientists had already looked at part of it.
In 2021, researchers studying three unusually large viruses that infect Staphylococcus aureus described a predicted reverse transcriptase—an enzyme that copies RNA into DNA—and a long stretch of nearby DNA they suspected might contain a noncoding RNA.
They recorded the enzyme and the neighboring region. What they did not identify was the larger repeat-array architecture Anthropic’s researchers now call ART.
Then an AI agent came along.
While searching DNA around similar reverse transcriptases, one Claude agent noticed something odd: a long array of repeating sequences sitting next to the enzyme and another nearby gene. Anthropic’s researchers eventually named the arrangement array-associated reverse transcriptases, or ART.
And because those repeats vaguely resemble part of the architecture behind CRISPR, one of the most consequential gene-editing technologies ever developed, the obvious headline practically wrote itself.
There is just one fairly important complication.
Nobody knows what ART does.
That makes Anthropic’s experiment interesting for a reason that has less to do with discovering “the next CRISPR” and more to do with what AI might change about science itself.
Claude did not pull an entirely unknown enzyme out of the biological ether. It looked at existing genomic data, followed an unusual clue around a previously described enzyme and recognized a relationship humans had not characterized.
The interesting capability here may be scientific attention: giving machines enough autonomy to decide which weird things in enormous datasets deserve a second look.
Now scientists have to figure out whether that attention is reliable.
What Claude actually found
Anthropic’s experiment was large.
According to the company, roughly 950 Claude agents spent 21 hours investigating reverse transcriptases across enormous genomic databases, consuming about 210 million tokens in the process.
The agents collected more than 200,000 reverse transcriptases, identified roughly 3,500 candidate systems and narrowed those down to 20 for deeper analysis.
One agent eventually wandered into the finding that became ART.
Rather than stopping at the protein it had been investigating, it examined the surrounding DNA, spotted the unusual repeat pattern and pursued the lead. Human researchers then reviewed the finding and ran follow-up experiments in Anthropic’s wet lab.
Those experiments supplied another piece of the puzzle: the repeat array produces distinct short RNA molecules.
That makes the system worth studying. But it leaves the most important questions unanswered.
Researchers have not established what the reverse transcriptase actually does in the ART system, whether those RNAs guide or interact with it, or whether the system can target, copy, cut or edit DNA in any controlled way.
That is why the CRISPR comparison needs an asterisk the size of a lab coat.
ART is CRISPR-like in one narrow sense: both contain arrangements involving repeated DNA sequences. CRISPR’s enormous usefulness comes from what its components actually do together.
ART’s function is still a blank space.
The more interesting test may be attention
There is another way to look at Anthropic’s result.
Forget gene editing for a moment.
The striking part is that Claude was not simply asked a narrowly defined question and sent off to calculate the answer. Anthropic gave its agents a broad research objective, and one agent decided that an unexpected feature in the surrounding DNA deserved investigation.
That resembles one of the messier parts of real science: deciding which anomaly is interesting enough to chase.
Biology has no shortage of data. Genomic databases contain a staggering number of sequences, proteins and relationships that no scientist has time to inspect individually.
Traditional bioinformatics tools already help researchers search that universe. A scientist can design software to find repeats, compare related genes, cluster similar proteins and examine which genes tend to appear together.
In fact, independent computer scientist Dimitri Perrin noted that a purpose-built conventional bioinformatics pipeline could probably have found the ART pattern too. Anthropic’s study does not include a matched experiment showing that its agent approach beats conventional genome mining on the same data, personnel, time and cost.
So the useful question is not whether AI can perform a search humans simply cannot.
It is whether software that can run tools, inspect results and choose which unexpected clues to pursue lets researchers explore more promising paths than existing workflows do.
That question is still open.
Now comes the awkward part: can the process be repeated?
Scientific methods become useful when researchers can depend on them.
And Anthropic’s own technical report contains an important warning about how dependent this kind of discovery can be on the workflow around the model.
When Anthropic later tested Claude models on the relevant ART sequences under controlled conditions, the strongest models identified the repeat array in at least 90% of attempts when the DNA was placed directly in their context.
Give the models tools and files to navigate, however, and performance became much less reliable. In one model-and-tool configuration, recognition fell to 32%.
The paper offers a revealing clue about why.
In many failed file-based attempts, the model never read enough contiguous DNA to encounter the relevant signal properly. The problem was not necessarily that the biology had suddenly become too sophisticated. The research workflow sometimes failed to put the model in front of the evidence it needed.
That is exactly why the result matters.
An autonomous research system is more than its underlying model. Reliability depends on whether the agent searches the right data, invokes the right tools, follows useful leads and actually inspects the evidence it retrieves.
A brilliant model wrapped in a brittle research workflow can still miss the clue.
This is also why reproducibility in AI-assisted science may require more than rerunning the same prompt. Stanford researchers have explored a similar idea with Paper2Agent, which turns research papers into executable agents so that scientific methods become easier to inspect and reuse.
For systems like Anthropic’s, scientists need to know how frequently important signals are found, how many false leads the process produces, what gets missed and how those results compare with other methods.
One successful hunt demonstrates possibility.
A reproducible process demonstrates a method.
The word “discovery” is doing a lot of work here
There is another distinction worth keeping straight.
Claude did not discover the reverse transcriptase itself.
The 2021 study of the MarsHill bacteriophage had already described a predicted reverse transcriptase and noted a roughly 1,200-base-pair noncoding region immediately upstream.
Anthropic’s reported contribution is different: its agent recognized that reverse transcriptases like this one appeared beside a distinctive repeat array and partner gene, then connected those features into what the researchers propose is a previously uncharacterized biological system.
That may sound like scientific hair-splitting. It isn't.
As AI becomes a more routine part of research, “the AI discovered X” risks collapsing several very different contributions into one phrase.
An AI system can retrieve an overlooked observation. It can recognize a pattern across existing data. It can generate a hypothesis. It can propose an experiment. Human researchers can then test that hypothesis, establish a mechanism and determine whether the finding survives independent replication.
Each step matters.
They are not interchangeable.
ART is a particularly useful example because the chain is unusually visible: earlier researchers generated and interpreted the genomic data; an AI agent noticed a different pattern inside it; Anthropic’s scientists decided the lead was worth pursuing; humans ran the physical experiments; and the system’s actual function remains unresolved.
“AI discovered it” fits nicely into a headline.
Science is going to need a longer vocabulary.
Anthropic now sits on both sides of the microscope
There is also a structural wrinkle to this new model of AI-assisted science.
Anthropic is not merely supplying tools scientists use to do research. It now operates its own life-sciences research program and wet lab while also controlling access to some of its most capable biology models.
Days before announcing ART, the company introduced a Life Sciences Verification Program designed to give vetted researchers access to models with more permissive biology safeguards.
Anthropic says those restrictions address the dual-use risks of powerful biological AI systems.
That does not imply wrongdoing or an inherent conflict. It does create a practical reproducibility question.
If a model company uses its own advanced systems, internal agent infrastructure and laboratory to produce a discovery, what exactly does an outside researcher need to reproduce the work?
The underlying sequence data may be public, while reproducing the exact workflow can still require the same model access, orchestration, compute, candidate-selection rules and experimental setup.
That distinction matters because the future of AI-enabled biology will depend on more than whether one company can produce promising leads. Outside researchers also need enough access and methodological detail to evaluate how those leads were produced.
What would actually settle this
ART now has two separate tests ahead of it.
The biological test is straightforward in principle, even if the experiments themselves are hard: researchers need to determine what the system does. Does the reverse transcriptase interact with the repeat-derived RNAs? Are those RNAs necessary? What biological reaction occurs? Can the system be controlled?
Until those experiments exist, ART is an intriguing biological lead rather than a demonstrated new gene-editing technology.
The second test is about the AI.
An independent team needs to show whether an agent-led search can reliably surface discoveries like this, ideally against conventional genome-mining approaches working from the same starting point.
That comparison would measure something much more useful than whether Claude happened to find one interesting pattern.
It could show whether this kind of agent-led search changes the cost, time, researcher effort or candidate yield involved in exploring genomic data.
Because that may be where the real advantage lies.
Scientists already know how to search genomes. Researcher attention for deciding which computational leads deserve deeper analysis and scarce laboratory time is much harder to scale.
AI agents can investigate thousands of candidates in parallel, chase odd side paths and generate more hypotheses than a research team could reasonably inspect one by one.
That creates its own bottleneck.
When machines can produce hundreds or thousands of plausible leads, the scarce resource shifts toward deciding which ones are worth believing, testing and funding.
That connects ART to a much broader shift already underway in AI-driven biology: models are getting better at generating possible answers faster than laboratories can validate them.
ART may turn out to be biologically important. It may turn out to be a curiosity. Its resemblance to CRISPR may eventually look prescient, or merely cosmetic.
The more immediate test is much less glamorous.
Can another laboratory verify what it does? Can another research team reproduce the search? Can scientists measure when this agent workflow succeeds, when it fails and whether it actually improves on the tools they already have?
That is the boring part of science.
It is also the part that turns an interesting result into something other scientists can rely on.