M01 · Talaria Intelligence

Home  ·  One engine, three surfaces  ·  Intelligence

M01 · Prospective · Shipped core

Ask while the protocol can still change.

Will this design survive Phase III — and what would have to change for it to?

Intelligence is the surface you address before you commit. It takes a trial as drafted and returns a calibrated interval for the outcome, together with the specific design levers that move it. It exists for the narrow window in which the answer is still actionable: once the protocol is registered and enrolling, a probability is a forecast; while it is a draft, the same probability is a decision.

Takes
A trial draft
Asked
Before it runs
Returns
An interval & its levers
Status
Shipped core

validation scores · the design-time prediction head on held-out trials

The window this surface lives in
While it's a draft, the outcome is still plural — and movable.
clears ↑ your draft, today

Every faint line is a trial this design could still become. The levers decide which branch you land on; once the protocol locks, the cloud collapses to one. A hand-authored schematic, not a measurement.

In plain terms

What Intelligence actually does for you.

Plain Englishno jargon — the one-paragraph version

A late-stage trial costs years and hundreds of millions of dollars. Most trials that fail were not the wrong drug — they were the wrong trial: an endpoint too blunt to see the effect, a population that diluted it, a cut-off set a notch too wide. Intelligence reads your trial while it is still a draft and tells you two things a spreadsheet never can — how likely this exact design is to succeed, and the handful of changes that would most improve its odds.

The difference is timing. Once the protocol is registered and enrolling, a prediction is just a forecast you watch come true. While it is still a draft, the same prediction is a decision you can still act on — and every answer arrives with the specific change that would move it.

For the clinical lead

Point Intelligence at your draft protocol. It returns a calibrated range for the outcome and ranks the design decisions by how much each one moves that range — so the redesign conversation starts from evidence, not opinion, while the protocol can still be edited.

For the data-science reviewer

A calibrated head over one shared trial representation anchored to a causal map. The counterfactual panel reports intervention effects on five fixed design axes with an epistemic / aleatoric uncertainty split, subgroup (CATE) conditioning, a two-path disagreement monitor, and an Ed25519-signed audit row per prediction.

DESIGN LEVERS Enrichment cutoff Primary endpoint Population Powering Line of therapy one lever moved · four held fixed PREDICTED OUTCOME clears → as drafted · misses the endpoint as revised · one lever moved, it clears 0 0.5 1 probability of clearing (schematic)

One lever, moved. Same drug, same trial — only the enrichment cut-off changes. As drafted, the outcome interval sits below the bar and the trial misses; revised, it clears. A hand-authored synthetic fixture, not a measurement.

Why we're ahead

A score tells you the odds. Only Intelligence hands you the lever.

The industry standard is a trained model that returns a probability. That answers “how likely?” and stops. The commercial question is the next one — “so what do I change?” — and it is structurally out of reach for a model that was never built to hold a trial's design as something you can move.

The conventional path · a leaderboard model
A number, and a shrug
Returns one probability, with no account of what would move it.
Trained to rank trials, not to price a design change; the levers are invisible to it.
A pooled score — the subgroup you can actually enrol is averaged away.
No statement of whether it should be trusted on your trial today.
No verifiable record of what produced the number.
Talaria Intelligence
A decision you can act on
A causal interval for each of five design levers — the change worth making, not just the odds.
Built over a fixed causal map, so a lever's effect stays interpretable, not an ever-shifting attribution.
A subgroup (CATE) fan — the estimate for the population you can actually enrol.
A calibration / trust state — whether the model stands behind this trial today.
An Ed25519-signed audit row per prediction — model, features and substrate versions, verifiable.
A real trial, on the public record

Same molecule. One selection decision. Opposite result.

Gefitinib in non-small-cell lung cancer · ISEL (2005) → IPASS (2009)

The design lever was the patient, not the drug.

Gefitinib in a broad, unselected population of previously-treated lung-cancer patients (the ISEL trial, versus placebo) showed no clear benefit. On that evidence the molecule looks like a failure.

Then the same drug was tested first-line in patients selected for an EGFR mutation (IPASS). In mutation-positive patients it beat chemotherapy on the trial's progression-free endpoint; in mutation-negative patients it did worse. The drug never changed. The selection decision decided whether its effect was visible.

That enrolment lever is exactly what Intelligence prices before a protocol locks — while who-you-enrol is still a choice, not a post-mortem.

SAME MOLECULE · gefitinib effect visible → unselected population EGFR-selected no clear benefit ISEL · vs placebo effect visible IPASS · first-line

Public record, drawn from the published literature (Thatcher et al., Lancet 2005; Mok et al., NEJM 2009). Talaria was not run on this trial — it is shown to make the mechanism concrete. No Talaria prediction is attached to any real molecule.

What it returns

Four things travel with every answer.

01

A calibrated interval, never a point

The output is a range with an explicit uncertainty envelope, split into the part that shrinks with more evidence and the part that does not. A single number would imply a precision the evidence does not support.

02

A counterfactual panel

Five design levers — enrichment cutoff, primary endpoint, population, powering and enrollment, line of therapy — each carrying its own causal interval. This is the part a leaderboard model cannot produce.

03

Subgroup-conditional estimates

The pooled interval fans into strata that frequently disagree. A design that clears on the pooled estimate can still fail on the population you are actually able to enroll.

04

A calibration state and a signed audit row

Whether the model considers itself trustworthy for this trial today, and a cryptographically signed record of the model, features and substrate versions behind the answer.

What we can provide

The full shape of an Intelligence answer.

One head ships today — the subgroup-conditional design-time predictor, with everything below travelling with each answer. It sits at the front of a roster of prediction surfaces (57 across the trial lifecycle), each a calibrated head over the same trial representation, built to customer pull.

01

Calibrated outcome interval

A range with an explicit epistemic / aleatoric split — never a false-precision point.

Ships today
02

Counterfactual lever panel

Five design levers, each with its own causal interval — the part a leaderboard cannot produce.

Ships today
03

Subgroup (CATE) fan

The pooled interval resolved into the strata you can actually enrol.

Ships today
04

Calibration / trust state

Whether the model stands behind this trial today, or declines to.

Ships today
05

Signed audit row

An Ed25519 record of model, feature and substrate versions behind every answer.

Ships today
06

The prediction roster

Dozens of further heads — endpoint, enrolment, safety and more — designed and queued against the same substrate.

Designed · queued

Honest split: item 06 is designed and queued, not deployed. We name what ships and what does not.

How it works

One substrate, addressed at a different moment.

Every prediction is a calibrated head over one shared representation of a trial, anchored to a causal map. New predictive capacity is added as a head plus a counterfactual panel — never as an eighth base learner. That constraint is what keeps the levers interpretable: the panel reports changes in a fixed causal structure rather than the shifting attributions of an ever-growing ensemble.

Trial draft as drafted One shared trial representation + causal map one substrate Calibrated head how likely Counterfactual panel what to change Signed, audited output Ed25519-signed

One substrate, many heads. New predictive capacity is added as a head plus a counterfactual panel over the same representation — never an eighth base learner — which is what keeps the levers interpretable. A hand-authored schematic, not a measurement.

Scope, stated plainly

The levers are reported unranked — an order we cannot yet defend would be the most useful-looking, least honest thing on this page. One head ships today; the rest of the roster is designed and queued. The validation scores are held-out figures, not a warranty — live performance depends on your trial. No FDA authorization exists; a SaMD De Novo submission is a planned milestone, not an achieved one.

See it run

An interactive demo on a synthetic composite.

Every value in the demo is a hand-authored fixture. It is there to show the shape of the output and how it responds — not to report a measurement.

The other two surfaces

Same engine, different moment.