7,000 rare diseases have no approved therapy, and for most of them nobody has ever written down a serious mechanistic hypothesis. MEDICLAUDE writes one, continuously, one disease after another. The hard part is not generating ideas — it is throwing away the wrong ones. So every idea has to survive a second AI built to destroy it.
There are more than 7,000 recognised rare diseases and roughly 300 million people living with one. The overwhelming majority have no approved therapy. Not because the biology is uniformly intractable — because a few thousand patients cannot repay a decade of development. For most of these conditions, nobody has ever sat down and written a serious mechanistic hypothesis.
Language models can write that hypothesis in minutes. The problem is that they can also write a beautiful, confident, completely wrong one, and a plausible falsehood in medicine is worse than silence. Fluency is not correctness, and a single model reviewing its own work grades itself generously.
So MEDICLAUDE never lets one model decide. Candidate therapies from FABLE are put through four rounds of elimination by OPUS 5 — a separate model, in a separate context, given the opposite incentive: it is rewarded for destroying, not for producing. Surviving four rounds of that is weak evidence, and weaker than an outside expert would be. It is still far stronger than one model marking its own work.
The output is not a cure and does not pretend to be. It is a testable increment: a hypothesis specific enough that a bench scientist can say what experiment would kill it, and an explicit record of what already died.
Each cycle pulls one disease from the Monarch Initiative knowledge graph — causal genes, phenotype spectrum with frequencies, gene pleiotropy, and phenotype-matched model organisms. The challenger enters six to eight candidate therapies. Then the killing starts.
Reads the genetic evidence and enters a slate of candidate therapies — the conservative repurposing plays and the aggressive long shots together. Then defends each survivor, round after round, against attacks it cannot see coming. Abandoning a candidate it cannot defend is a legitimate move.
A separate model given the same evidence and one instruction: kill what does not deserve to survive. It hunts invented drugs, misremembered trial outcomes, missing causal steps, impossible therapeutic windows, and claims no experiment could ever disprove. Its rulings are final.
Does the agent exist, and is it even aimed at this disease? The obviously broken die here.
The full causal chain, defended. Skip a step or contradict the data and the candidate dies.
Can it be built, can it reach the tissue, and has it already failed in the clinic?
Name the cheapest experiment that would prove it wrong. No such experiment, no survival.
A kill is permanent — there is no appeal and no revision. What reaches the archive has been attacked four separate ways by a model that was rewarded for destroying it. When nothing survives, that is published as a wipeout, because a disease where every plausible idea dies is a real result and worth knowing.
Every claim is tagged at the source: [KG] from the knowledge graph,
[KNOWN] from literature, [INFERRED] reasoned, [SPECULATIVE]
flagged as unsupported. The gauntlet checks the tags too — inference dressed up as established fact
is grounds for a kill.
A model that has read the literature can produce a confident, correct-sounding paragraph about a disease without deriving anything. That failure mode is invisible when every disease in the set is unsolved, because there is no answer to check against.
Spinal muscular atrophy sits in the set as a control. SMA is solved — nusinersen, risdiplam and onasemnogene abeparvovec are approved, and the SMN1/SMN2 mechanism is textbook. When the gauntlet reaches it, the trace can be read against a known answer. Did it reconstruct the SMN2 copy-number logic from the graph, or restate what it already knew?
The control validates nothing on its own. It calibrates how much weight to put on the diseases where there is nothing to check against — which is the entire rest of the set.
Both models stream token by token as they work — the researcher's draft, then the critic tearing into it, then the verdict. Nothing is edited, curated, or retried for a better answer. Refusals, truncations and failures print exactly as they occur.
Every disease resolves to a MONDO identifier against Monarch before entering the queue. The set grows on its own — the researcher proposes related disorders, and they join the frontier.
Newest first. Click any entry to expand its surviving hypotheses and the ones the critic killed.
Monarch Initiative v3 knowledge graph — a public, keyless API unifying curated rare-disease resources including OMIM, Orphanet and HPO annotations. Per disease we pull the entity record, causal and correlated gene associations, the phenotype spectrum with frequency qualifiers, gene pleiotropy, and phenotype-nearest diseases and mouse model genes by semantic similarity.
Two models in opposition, both reached through OpenRouter. Both see identical evidence — the critic must be able to check every claim against the same source the researcher used.
Nothing here is reviewed by a human. Findings are marked UNREVIEWED because that is
what they are — there is no expert review queue behind this site. Contested stages are flagged.
Failures, refusals and empty results are printed rather than hidden, and the archive publishes
killed hypotheses alongside surviving ones.
/events — SSE stream of the live gauntlet