Anthropic Claude multi-agent campaign discovers ART (array-associated reverse transcriptases)
Source: <https://www.anthropic.com/news/claude-discovers-novel-enzyme-system>
Preprint PDF: <https://www-cdn.anthropic.com/22573675ada52a8ca8a97a1a4b4326b2f208a071.pdf>
Title: *Autonomous AI agents discover reverse transcriptases with tandem repeat arrays*
Published: 2026-09-23 (Anthropic news + preprint)
What ART is, and where the public materials stop
ART (array-associated reverse transcriptases) is a sequence-and-neighbourhood label that bundles three things, mostly found in jumbo phage genomes:
1. A reverse transcriptase (RT) — an enzyme family that copies RNA into DNA; 2. A partner gene sitting next to it; 3. A uniformly spaced DNA repeat array upstream of the RT, visually reminiscent of a CRISPR array but lacking the adjacent classical cas-gene cluster.
The news post and the preprint both state: the underlying RT is not new — RTs have been described in prior literature. The novelty is that an agent bundled repeat array + partner gene + RT into a nameable system. On the wet-lab side, public RNA-seq and E. coli heterologous expression show the array producing a set of discrete short RNAs, and the array-derived RNA can be very abundant early in infection. The preprint's Discussion is direct: there is currently no proof that the RT is active, that the unit RNA is its substrate, how the RT and partner interact, or what this system does for the phage. Feng Zhang's preprint comment is "exciting" and worth continuing — it is a signal of expert interest, not product approval.
The campaign, not a slogan
The preprint frames the harness as an auditable campaign, not as "AI discovered something".
Orchestration. A research brief was split into analysis phases, then into small tasks. Each small task put one Claude Code worker in charge of proposing and executing the plan and a supervisor in charge of reviewing both. The supervisor could open follow-ups from observation. Plan, result, and review go into a shared record; what the campaign hands back to humans is a written report, not a chat-window screenshot.
Scale (per preprint):
| Item | Number | |------|--------| | Protein-cluster search space | ~1.9×10⁹ | | Wall-clock | 21.5 h | | Agent sessions | 949 | | Tokens | 2.156×10⁸ | | Agent-hours | 77 | | Tasks | 119 (98 follow-ups triggered by observation) | | RT clusters recovered | ~2×10⁵ | | Partner families scored in neighbourhood | 3,564 (sampling ~1.1×10⁴ RT loci) | | Candidate families in deep dive | 16+1 | | End reports | 19 |
The later part of the funnel is more "scientific" than the early part. Of 17 candidate partner families, only 3 were confirmed as RT-associated but previously unreported; another 14 were dropped or set aside — annotation artefacts, parts of known systems, or just common neighbourhood residents. The worker also flagged individual RTs with anomalous flanking DNA, of which 3 were classified as new RT lineages. ART landed on the branch where the brief asked for partner proteins but a non-coding-DNA observation escalated it. The news post summarises this as "around 20 of the most interesting reports"; the preprint says 19. Cite the preprint for campaign numbers.
The discovery moment: the key action is reading raw DNA
The brief was "find novel RT systems by new partner-gene associations" — it did not name repeat arrays. The defining feature of ART is the non-coding repeat upstream of the RT. The session record shows the worker, while reading contig flanking sequence, naming the tandem repeat array out loud, counting repeats, measuring spacing, comparing layouts to known RT/CRISPR/retron families, checking literature for the same pattern, then handing the report to humans for review. The preprint later uses an internal signal-analysis pass to argue that Mythos 5 has an interpretable response when reading repeat DNA, then naturally names it — the authors describe this as one side of "genomic vision".
A follow-up with interactive Claude Science expanded the family: at 90% identity, 95 ART-related RT clusters were found, and 28 had a detectable repeat array upstream of the RT. The repeat unit is on the order of ~200 nt; conservation of repeats between related phages is not identical to the "highly variable" CRISPR-spacer model — these are sequence-level descriptions, still not functional conclusions.
For coding-agent teams, this section is more useful than the CRISPR analogy: the brief must allow observation to open a new path, the shared record has to absorb 98 follow-ups, and the actual key raw string has to be read into context. Remove any one of those three and the "discovery" narrative collapses.
Rerunning the same campaign ten times: every rerun misses the array
The preprint does not sell ART as "press the button and you get ART". They reran the same campaign ten times.
The layered result matters:
- Almost every rerun that completed the census did sample ART-related loci;
- Two reruns had a worker that treated that lineage as a follow-up;
- But: no rerun read the DNA upstream of the RT, so every rerun missed the array.
The authors attribute this to two factors: the broad protein search space around RTs, and harness non-determinism. So "ten reruns did not reproduce" should not be translated as "the first one was a hallucination". The more accurate engineering sentence: ART-family loci keep showing up in the protein funnel, but array identification depends on a "read enough upstream DNA" path that none of the reruns walked again.
Fixed-input benchmark: more tools, lower array hit rate
To separate "does the model know ART" from "did the campaign get lucky", the team ran a fixed-input benchmark: hand the model ART-related sequence directly — either inline in the context, or in the file system with analysis tools enabled — and ask it to write the features it considers important. A Mythos 5 judge scored against 10 author-defined ART features.
Seven Claude models show a clear split: the stronger cohort is Opus 5.5, Mythos 5.1, Mythos 5, Opus 5; the weaker cohort includes Opus 4.6, Opus 4.8, Sonnet 5. The original campaign ran on Mythos 5, so the authors credit capability tier.
When the three pieces are split out, the difficulty is asymmetric:
- RT protein: even with only two protein sequences, strong models describe it fairly accurately;
- Partner gene: more tools and information improve the description;
- Array: the preprint is emphatic — more information and tooling suppress the accurate description of the array.
Specific percentages:
- With loci only in context, the strongest models describe the array accurately ≥90% of attempts;
- With the tooling-enabled tiers, it can drop as low as 32% (e.g. Opus 5 level 4 in the paper);
- With file use, 39% of attempts never read in a contiguous ≥200 nt of DNA — i.e. not even one full repeat unit;
- Once the model reads ≥200 nt, all models gain roughly 16–32 percentage points;
- The longer the read-in, the four strong models' aggregate identification rate climbs from about 29% to 76%, with Mythos 5 reaching 96%.
For local multi-agent research pipelines this set is a direct warning: the default "summarise first, then call tools, read little" pattern saves tokens in code repos and systematically loses signal on data where the repeat pattern has to be seen. The problem is not necessarily that the model "isn't smart enough" — it's that the key window never entered the context.
Six engineering rules that fit local agent work
1. Briefs reserve observation budget. 98 of 119 tasks are follow-ups. If a supervisor cannot open a new task because of unusual flanking sequence, ART-style "out-of-brief features" do not surface. 2. Shared record is the audit surface, not decoration. 949 sessions living in private session context cannot be replayed. The preprint can discuss rerun and transcript because plan/result/review is on disk. 3. Number buckets are written separately, not collapsed into a slogan. Track separately: did a full-campaign rerun read the array again; what is the fixed-input description accuracy; where is the wet-lab RNA evidence; is function known? 4. For tasks that must see raw data, override the default read-in. Force a contiguous window, a minimum nt/byte threshold, or inject the key segment into context; do not assume "file tool available = model will read enough". 5. Rerun failure on a weaker model does not falsify a strong-model campaign. Same harness, Sonnet tier vs Mythos 5 tier shows a clear array-identification gap. Comparison experiments must pin the model snapshot, not just the "agent framework" name. 6. Humans still own the non-outsourceable squares. Bulk candidate elimination, wet-lab work, functional interpretation, and the decision of when to publish a preprint remain human. The agent contribution is large-scale sweep + anomaly nomination.
Gaps between headlines and the materials
The Verge and similar outlets emphasise the CRISPR analogy and "nearly 1,000 agents / 21 hours / 210M tokens", and note the company still does not know the actual function, plus the overlap between life-sciences narrative and listings / research-team expansion timing. These are legitimate news framings. Engineering readers who stop at the headline will miss the two hardest pieces in the preprint: ten reruns all miss the array, and the array identification rate can drop to about 30% under the tool path.
The news post's motivation sentence is also explicit: the early share is to show Claude capability and to let outsiders see what they are working on; they are also recruiting research collaborators. Separating "capability demo + early scientific share" from "system function clarified" keeps the headline from dragging the reader away.
Bottom line
What is pinned down by the public materials:
- A genome-mining campaign, mostly Mythos 5 + Claude Code worker/supervisor, finished in roughly 21.5 hours, 949 sessions, 2.156e8 tokens, and handed back candidate reports including ART;
- ART, as a sequence-and-neighbourhood label, has early experimental support on the RNA-expression side;
- Function unknown — this is not a validated gene-editing product;
- Full-campaign reruns show the array-identification path is fragile; fixed-input benchmarks show whether the model reads enough long DNA matters more than whether tools are available.
For anyone running Hermes Agent, Claude Code, Codex or a self-built local agent, the directly usable engineering judgement is: the quality of a multi-agent scientific harness is, in large part, "did the key raw data enter the context" + "can observation turn into a follow-up" + "can humans audit reports bucket by bucket". The preprint writes those three buckets as auditable numbers; the CRISPR headline is only a visual analogy.