The protocol is frozen. The world is not.
Is the ground it was designed on still the ground — and is the model still trustworthy for it?
A Phase III trial can run for years. Over that span the standard of care moves, competitors read out, and the population that actually enrolls drifts away from the one that was powered for. None of this is visible from inside the protocol, because the protocol cannot change. Sentinel watches the ground underneath a running trial and reports when it has moved far enough to matter.
Three surfaces drift while the protocol cannot move. A design-time prediction is anchored to where these lines started; Sentinel measures how far they have travelled — and withdraws before the number misleads. A hand-authored schematic, not a measurement.
What Sentinel actually does for you.
A Phase III trial can take years, and the world does not hold still while it runs. The standard of care moves, the population that actually enrolls drifts, competitors read out. A prediction made when the protocol was drafted can quietly go stale — still confident, no longer earned — and nothing inside the frozen protocol will tell you.
Sentinel watches the live trial against the ground it was designed on. While that ground still holds, it tracks quietly. When the ground has moved too far, it does the unusual thing — it says so and withdraws its own calibration, rather than reporting a confident but stale number.
Sentinel monitors your running trial and tells you when the assumptions it was designed on no longer hold — before the readout, not after. What to do about it stays your decision; the job here is to make sure the drift is never a surprise.
Continuous drift and effective-coverage monitoring against the design-time reference distribution; fail-closed calibration withdrawal when population or standard-of-care drift exceeds tolerance; Mondrian-conditioned coverage tracking, each state change carrying a signed audit row.
It withdraws before it misleads. As the trial runs, the ground drifts off the distribution the model was calibrated on; past the tolerance threshold, Sentinel reports no number rather than a stale one. A hand-authored synthetic fixture, not a measurement.
Every other model scores once. Only Sentinel reports its own expiry.
A model trained at design time will happily return the same confident number three years later, long after the world it was fitted to has moved on. Knowing when to stop trusting a number is a capability, not an afterthought — and it is the one thing a score-and-forget model structurally cannot do.
The trial didn't change. The standard of care did.
A design-time assumption can expire mid-trial.
A Phase III can run for years, and the ground it was designed on can move mid-conduct. In 2020 the RECOVERY trial showed dexamethasone cut mortality in hospitalised COVID-19 patients who needed oxygen or ventilation — and it became standard of care almost overnight.
Every other COVID-19 trial then running had its control-arm baseline reset underneath it: comparator-arm patients were now receiving a treatment that had not existed when the protocol was written. (The same pattern recurs in oncology, where a chemotherapy control arm can be overtaken by immunotherapy becoming standard mid-trial.)
A prediction fixed at design time can quietly go stale while the trial runs. Sentinel is what notices — and says so.
Public record (RECOVERY Collaborative Group, NEJM 2021; the immunotherapy example is a general, well-recognised pattern). Talaria was not run on these trials — they are shown to make the mechanism concrete. No Talaria prediction is attached to any real trial.
Four things travel with every answer.
A drift severity, per stratum
Movement is rarely uniform. A trial can be stable overall while one enrolling subgroup has drifted past the point where the original powering assumption holds.
Effective coverage, not nominal coverage
What the intervals are actually delivering now, rather than what they were built to deliver under conditions that may no longer obtain.
An explicit trust state
Active, degraded, or withdrawn. The system is designed to fail closed: when it can no longer stand behind a calibration it says so, rather than continuing to emit a confident-looking number.
The specific mover
Which of the watched surfaces — standard of care, competitive field, enrolling population — accounts for the change, so the finding is actionable rather than merely alarming.
Everything Sentinel returns during conduct.
Sentinel is Built. Every answer travels with what moved, how far, and whether the calibration still stands — up to and including the decision to withdraw it.
Drift severity, per stratum
Movement is rarely uniform.
BuiltEffective coverage
What the intervals actually deliver now, not nominal.
BuiltTrust state
Active, degraded, or withdrawn (fail-closed).
BuiltThe specific mover
Standard of care, competitive field, or enrolling population.
BuiltSigned audit row
A verifiable record per state change.
BuiltDocumented thresholds
Where degraded becomes withdrawn is a governance choice, stated as such.
BuiltSentinel is Built. Its most important output is its own withdrawal — a monitor that never declines to answer is a number generator with a schedule.
One substrate, addressed at a different moment.
Sentinel monitors the substrate a prediction was conditioned on and re-evaluates whether that conditioning still holds. The engineering commitment is the unusual part: the system's most important output is its own withdrawal. A monitor that never declines to answer is not a monitor, it is a number generator with a schedule.
The output is a verdict on itself. Sentinel re-evaluates whether the conditioning a prediction relied on still holds, and fails closed when it cannot. A hand-authored schematic, not a measurement.
Sentinel cannot change a frozen protocol — it tells you what you are now running, which is a different trial from the one you designed. Drift detection is not outcome prediction: a stable reading is not a favourable one, and a drifting one is not a prediction of failure. Thresholds are governance choices, documented as such. Every figure shown publicly is a hand-authored synthetic composite. No FDA authorization exists; pre-deployment, not for clinical use.
An interactive demo on a synthetic composite.
Every value in the demo is a hand-authored fixture. It is there to show the shape of the output and how it responds — not to report a measurement.