M03 · Talaria Sentinel

Home  ·  One engine, three surfaces  ·  Sentinel

M03 · Real-time · Built

The protocol is frozen. The world is not.

Is the ground it was designed on still the ground — and is the model still trustworthy for it?

A Phase III trial can run for years. Over that span the standard of care moves, competitors read out, and the population that actually enrolls drifts away from the one that was powered for. None of this is visible from inside the protocol, because the protocol cannot change. Sentinel watches the ground underneath a running trial and reports when it has moved far enough to matter.

Takes
A live trial
Asked
During conduct
Returns
An interval & its levers
Status
Built
Why a frozen protocol goes stale
The protocol holds still. The world does not.
PROTOCOL frozen at design divergence → withdraw standard of care competitive field enrolling population

Three surfaces drift while the protocol cannot move. A design-time prediction is anchored to where these lines started; Sentinel measures how far they have travelled — and withdraws before the number misleads. A hand-authored schematic, not a measurement.

In plain terms

What Sentinel actually does for you.

Plain Englishno jargon — the one-paragraph version

A Phase III trial can take years, and the world does not hold still while it runs. The standard of care moves, the population that actually enrolls drifts, competitors read out. A prediction made when the protocol was drafted can quietly go stale — still confident, no longer earned — and nothing inside the frozen protocol will tell you.

Sentinel watches the live trial against the ground it was designed on. While that ground still holds, it tracks quietly. When the ground has moved too far, it does the unusual thing — it says so and withdraws its own calibration, rather than reporting a confident but stale number.

For the trial sponsor

Sentinel monitors your running trial and tells you when the assumptions it was designed on no longer hold — before the readout, not after. What to do about it stays your decision; the job here is to make sure the drift is never a surprise.

For the data-science reviewer

Continuous drift and effective-coverage monitoring against the design-time reference distribution; fail-closed calibration withdrawal when population or standard-of-care drift exceeds tolerance; Mondrian-conditioned coverage tracking, each state change carrying a signed audit row.

THE GROUND IT WAS DESIGNED ON design-time reference band actual ground · standard of care · enrolling population ground drifts out of band drift exceeds tolerance CALIBRATION COVERAGE calibration ACTIVE calibration WITHDRAWN reports no stale number protocol frozen drift threshold readout trial running over time (schematic)

It withdraws before it misleads. As the trial runs, the ground drifts off the distribution the model was calibrated on; past the tolerance threshold, Sentinel reports no number rather than a stale one. A hand-authored synthetic fixture, not a measurement.

Why we're ahead

Every other model scores once. Only Sentinel reports its own expiry.

A model trained at design time will happily return the same confident number three years later, long after the world it was fitted to has moved on. Knowing when to stop trusting a number is a capability, not an afterthought — and it is the one thing a score-and-forget model structurally cannot do.

The conventional path · a static, one-time model
Confident to the end
Scores once at design time and reports that number for years, as its assumptions rot.
No sense of whether the enrolling population still matches the one it was powered for.
Confident even when the standard of care has moved underneath it.
No account of which surface moved.
No verifiable record of any state change.
Talaria Sentinel
It knows when to stop
Continuous drift monitoring against the design-time reference distribution.
Effective coverage, not nominal — what the intervals actually deliver now.
Fail-closed withdrawal — it says “I can no longer stand behind this” rather than emit a stale number.
The specific mover — standard of care, competitive field, or enrolling population.
A signed audit row per state change.
A real trial, on the public record

The trial didn't change. The standard of care did.

Standard-of-care drift · dexamethasone / RECOVERY (2020)

A design-time assumption can expire mid-trial.

A Phase III can run for years, and the ground it was designed on can move mid-conduct. In 2020 the RECOVERY trial showed dexamethasone cut mortality in hospitalised COVID-19 patients who needed oxygen or ventilation — and it became standard of care almost overnight.

Every other COVID-19 trial then running had its control-arm baseline reset underneath it: comparator-arm patients were now receiving a treatment that had not existed when the protocol was written. (The same pattern recurs in oncology, where a chemotherapy control arm can be overtaken by immunotherapy becoming standard mid-trial.)

A prediction fixed at design time can quietly go stale while the trial runs. Sentinel is what notices — and says so.

STANDARD OF CARE, OVER TIME staleness design-time standard of care (assumed) actual standard of care new SOC adopted mid-trial protocol frozen readout

Public record (RECOVERY Collaborative Group, NEJM 2021; the immunotherapy example is a general, well-recognised pattern). Talaria was not run on these trials — they are shown to make the mechanism concrete. No Talaria prediction is attached to any real trial.

What it returns

Four things travel with every answer.

01

A drift severity, per stratum

Movement is rarely uniform. A trial can be stable overall while one enrolling subgroup has drifted past the point where the original powering assumption holds.

02

Effective coverage, not nominal coverage

What the intervals are actually delivering now, rather than what they were built to deliver under conditions that may no longer obtain.

03

An explicit trust state

Active, degraded, or withdrawn. The system is designed to fail closed: when it can no longer stand behind a calibration it says so, rather than continuing to emit a confident-looking number.

04

The specific mover

Which of the watched surfaces — standard of care, competitive field, enrolling population — accounts for the change, so the finding is actionable rather than merely alarming.

What we can provide

Everything Sentinel returns during conduct.

Sentinel is Built. Every answer travels with what moved, how far, and whether the calibration still stands — up to and including the decision to withdraw it.

01

Drift severity, per stratum

Movement is rarely uniform.

Built
02

Effective coverage

What the intervals actually deliver now, not nominal.

Built
03

Trust state

Active, degraded, or withdrawn (fail-closed).

Built
04

The specific mover

Standard of care, competitive field, or enrolling population.

Built
05

Signed audit row

A verifiable record per state change.

Built
06

Documented thresholds

Where degraded becomes withdrawn is a governance choice, stated as such.

Built

Sentinel is Built. Its most important output is its own withdrawal — a monitor that never declines to answer is a number generator with a schedule.

How it works

One substrate, addressed at a different moment.

Sentinel monitors the substrate a prediction was conditioned on and re-evaluates whether that conditioning still holds. The engineering commitment is the unusual part: the system's most important output is its own withdrawal. A monitor that never declines to answer is not a monitor, it is a number generator with a schedule.

Running trial during conduct Design-time reference distribution the ground it was designed on compared against Live substrate SOC · population · competitors Drift + effective- coverage monitor Trust state active / degraded / withdrawn fail-closed

The output is a verdict on itself. Sentinel re-evaluates whether the conditioning a prediction relied on still holds, and fails closed when it cannot. A hand-authored schematic, not a measurement.

Scope, stated plainly

Sentinel cannot change a frozen protocol — it tells you what you are now running, which is a different trial from the one you designed. Drift detection is not outcome prediction: a stable reading is not a favourable one, and a drifting one is not a prediction of failure. Thresholds are governance choices, documented as such. Every figure shown publicly is a hand-authored synthetic composite. No FDA authorization exists; pre-deployment, not for clinical use.

See it run

An interactive demo on a synthetic composite.

Every value in the demo is a hand-authored fixture. It is there to show the shape of the output and how it responds — not to report a measurement.

The other two surfaces

Same engine, different moment.