Clinical trial dynamics
Scripted actions on 20,000 active trials
Close recruitment and keep treating: six-month completion probability +59% on average. Trials differ.
Compare actions and outcomes
| Action taken at day 30 | Recruitment shortfall | Termination | Completion | Timeline postponed > 6 months |
|---|---|---|---|---|
| Trial left alone: six-month probability | 0.83% | 0.39% | 3.24% | 4.59% |
| Change the design | +61% | +31% | +23% | +20% |
| Change the sponsor | +34% | +12% | +13% | +18% |
| Close recruitment, keep treating | −45% | +8% | +59% | −35% |
| Suspend the trial | +24% | +14% | −7% | −37% |
| Pull the completion date in | +14% | +28% | +59% | +22% |
| Add sites | +6% | +5% | −3% | +2% |
| Verify the record, nothing else | 0% | +4% | −5% | −1% |
Change is relative to the trial left alone, averaged over 20,000 active trials on 1 March 2025, six months ahead, and sixteen real examples of each action. In absolute terms the moves are tenths of a percentage point on rare outcomes and one to two points on completion and timeline. These are the model's responses to an action, not measured effects of taking it; the only way to grade them is against a sponsor's own past decisions, which the public registry does not record.
Closing recruitment
Closing recruitment, by trial
Completion probability +5 points for the top tenth of trials. Almost no change for the bottom half.
Half the trials account for the change in the portfolio forecast.
Responses across trials and actions
Half the trials gain more than a point of completion probability when recruitment closes. One in six gains more than a point of shortfall risk from a design change. One in 25 moves more than a point in either direction when the record is merely verified.
Per-trial changes are the 10th, 50th and 90th percentiles across the 20,000 trials: closing recruitment moves completion by +0.0, +1.0 and +5.0 points; a design change moves shortfall by −0.0, +0.2 and +1.4 points; verification moves completion by −0.8, −0.1 and +0.2 points.
The model learned from actions sponsors chose to take. These are model responses to scenarios, rather than measured effects of interventions.
Following real events
Tracking after 1 to 8 real events
59,769 sequences, eight recorded events per trial. Tracking scores 0.83 to 0.88.
model, updated with each real eventno update: the trial as it stood at the start
How tracking was measured
Skill is one minus the model's error divided by the error of the never-updated picture, measured in the model's own representation against the state it computes from the full record after each event. The first event arrives a median 31 days in; the eighth, 304.
Most sequences span about a year. The model updates its state after each supplied event, the same operation used for scripted scenarios. For current forecasts, we read the latest full record: following eight events ranks outcomes about halfway between a stale reading and a fresh one.
Using the model
Ranking scores and 13 scripted actions
Ranking 0.80–0.86 on recruitment, termination, completion and timeline. Actions examined one at a time.
The same action can be compared across trials, or different actions on the same trial.
Forecasts, scenarios and consistency checks
Ranking scores use 0.5 as the chance baseline. Shortfall and termination probabilities are calibrated within a tenth of a percentage point.
The forecasts do not identify which event will happen next. Across 64 imagined futures per trial, outcome probabilities end within about a third of a point of one another and today’s forecast. For quiet trials, the model reads the record fresh rather than simulating six months of silence.
Scenario results come from one training run evaluated on a development period. Sequences of planned actions have not yet been evaluated.
A trial left alone must forecast about the same as the trial read today, and the model's own imagined futures must land where today's forecast points. Both checks pass on recruitment, termination, completion and timeline. An earlier version of the what-if read failed the first check because the quiet months were run with an input the model had never seen in training; the failure was traced, the read corrected, and every what-if number here is from the corrected version.
Sampled futures: next event chosen by the model, real example payloads, six 30-day steps, twelve monthly anchors of 20,000 trials. Ranking accuracy of the imagined-future ensemble against today's read: shortfall 0.82 vs 0.84, termination 0.815 vs 0.817, completion 0.81 vs 0.83, timeline 0.84 vs 0.86.
Related decks
This deck is the scripted-action half. The risk-ranking half is Trial Risk. Both are combined in Clinical Trial Oversight.