Ask while the protocol can still change.
Will this design survive Phase III — and what would have to change for it to?
Intelligence is the surface you address before you commit. It takes a trial as drafted and returns a calibrated interval for the outcome, together with the specific design levers that move it. It exists for the narrow window in which the answer is still actionable: once the protocol is registered and enrolling, a probability is a forecast; while it is a draft, the same probability is a decision.
validation scores · the design-time prediction head on held-out trials
Every faint line is a trial this design could still become. The levers decide which branch you land on; once the protocol locks, the cloud collapses to one. A hand-authored schematic, not a measurement.
What Intelligence actually does for you.
A late-stage trial costs years and hundreds of millions of dollars. Most trials that fail were not the wrong drug — they were the wrong trial: an endpoint too blunt to see the effect, a population that diluted it, a cut-off set a notch too wide. Intelligence reads your trial while it is still a draft and tells you two things a spreadsheet never can — how likely this exact design is to succeed, and the handful of changes that would most improve its odds.
The difference is timing. Once the protocol is registered and enrolling, a prediction is just a forecast you watch come true. While it is still a draft, the same prediction is a decision you can still act on — and every answer arrives with the specific change that would move it.
Point Intelligence at your draft protocol. It returns a calibrated range for the outcome and ranks the design decisions by how much each one moves that range — so the redesign conversation starts from evidence, not opinion, while the protocol can still be edited.
A calibrated head over one shared trial representation anchored to a causal map. The counterfactual panel reports intervention effects on five fixed design axes with an epistemic / aleatoric uncertainty split, subgroup (CATE) conditioning, a two-path disagreement monitor, and an Ed25519-signed audit row per prediction.
One lever, moved. Same drug, same trial — only the enrichment cut-off changes. As drafted, the outcome interval sits below the bar and the trial misses; revised, it clears. A hand-authored synthetic fixture, not a measurement.
A score tells you the odds. Only Intelligence hands you the lever.
The industry standard is a trained model that returns a probability. That answers “how likely?” and stops. The commercial question is the next one — “so what do I change?” — and it is structurally out of reach for a model that was never built to hold a trial's design as something you can move.
Same molecule. One selection decision. Opposite result.
The design lever was the patient, not the drug.
Gefitinib in a broad, unselected population of previously-treated lung-cancer patients (the ISEL trial, versus placebo) showed no clear benefit. On that evidence the molecule looks like a failure.
Then the same drug was tested first-line in patients selected for an EGFR mutation (IPASS). In mutation-positive patients it beat chemotherapy on the trial's progression-free endpoint; in mutation-negative patients it did worse. The drug never changed. The selection decision decided whether its effect was visible.
That enrolment lever is exactly what Intelligence prices before a protocol locks — while who-you-enrol is still a choice, not a post-mortem.
Public record, drawn from the published literature (Thatcher et al., Lancet 2005; Mok et al., NEJM 2009). Talaria was not run on this trial — it is shown to make the mechanism concrete. No Talaria prediction is attached to any real molecule.
Four things travel with every answer.
A calibrated interval, never a point
The output is a range with an explicit uncertainty envelope, split into the part that shrinks with more evidence and the part that does not. A single number would imply a precision the evidence does not support.
A counterfactual panel
Five design levers — enrichment cutoff, primary endpoint, population, powering and enrollment, line of therapy — each carrying its own causal interval. This is the part a leaderboard model cannot produce.
Subgroup-conditional estimates
The pooled interval fans into strata that frequently disagree. A design that clears on the pooled estimate can still fail on the population you are actually able to enroll.
A calibration state and a signed audit row
Whether the model considers itself trustworthy for this trial today, and a cryptographically signed record of the model, features and substrate versions behind the answer.
The full shape of an Intelligence answer.
One head ships today — the subgroup-conditional design-time predictor, with everything below travelling with each answer. It sits at the front of a roster of prediction surfaces (57 across the trial lifecycle), each a calibrated head over the same trial representation, built to customer pull.
Calibrated outcome interval
A range with an explicit epistemic / aleatoric split — never a false-precision point.
Ships todayCounterfactual lever panel
Five design levers, each with its own causal interval — the part a leaderboard cannot produce.
Ships todaySubgroup (CATE) fan
The pooled interval resolved into the strata you can actually enrol.
Ships todayCalibration / trust state
Whether the model stands behind this trial today, or declines to.
Ships todaySigned audit row
An Ed25519 record of model, feature and substrate versions behind every answer.
Ships todayThe prediction roster
Dozens of further heads — endpoint, enrolment, safety and more — designed and queued against the same substrate.
Designed · queuedHonest split: item 06 is designed and queued, not deployed. We name what ships and what does not.
One substrate, addressed at a different moment.
Every prediction is a calibrated head over one shared representation of a trial, anchored to a causal map. New predictive capacity is added as a head plus a counterfactual panel — never as an eighth base learner. That constraint is what keeps the levers interpretable: the panel reports changes in a fixed causal structure rather than the shifting attributions of an ever-growing ensemble.
One substrate, many heads. New predictive capacity is added as a head plus a counterfactual panel over the same representation — never an eighth base learner — which is what keeps the levers interpretable. A hand-authored schematic, not a measurement.
The levers are reported unranked — an order we cannot yet defend would be the most useful-looking, least honest thing on this page. One head ships today; the rest of the roster is designed and queued. The validation scores are held-out figures, not a warranty — live performance depends on your trial. No FDA authorization exists; a SaMD De Novo submission is a planned milestone, not an achieved one.
An interactive demo on a synthetic composite.
Every value in the demo is a hand-authored fixture. It is there to show the shape of the output and how it responds — not to report a measurement.