Strange Lab · Research
Clinical trial research · Part 2 of 2, scripted actions · September 2026

Clinical trial dynamics

Scripted actions on 20,000 active trials

Close recruitment and keep treating: six-month completion probability +59% on average. Trials differ.

Change is relative to the trial left alone. Sixteen real examples of each action.
Compare actions and outcomes
Action taken at day 30Recruitment shortfallTerminationCompletionTimeline postponed > 6 months
Trial left alone: six-month probability0.83%0.39%3.24%4.59%
Change the design+61%+31%+23%+20%
Change the sponsor+34%+12%+13%+18%
Close recruitment, keep treating−45%+8%+59%−35%
Suspend the trial+24%+14%−7%−37%
Pull the completion date in+14%+28%+59%+22%
Add sites+6%+5%−3%+2%
Verify the record, nothing else0%+4%−5%−1%

Change is relative to the trial left alone, averaged over 20,000 active trials on 1 March 2025, six months ahead, and sixteen real examples of each action. In absolute terms the moves are tenths of a percentage point on rare outcomes and one to two points on completion and timeline. These are the model's responses to an action, not measured effects of taking it; the only way to grade them is against a sponsor's own past decisions, which the public registry does not record.

Closing recruitment

Closing recruitment, by trial

Completion probability +5 points for the top tenth of trials. Almost no change for the bottom half.

+5 pointscompletion probability for the top tenth of trials
Almost no changefor the bottom half

Half the trials account for the change in the portfolio forecast.

Responses across trials and actions

Half the trials gain more than a point of completion probability when recruitment closes. One in six gains more than a point of shortfall risk from a design change. One in 25 moves more than a point in either direction when the record is merely verified.

Per-trial changes are the 10th, 50th and 90th percentiles across the 20,000 trials: closing recruitment moves completion by +0.0, +1.0 and +5.0 points; a design change moves shortfall by −0.0, +0.2 and +1.4 points; verification moves completion by −0.8, −0.1 and +0.2 points.

The model learned from actions sponsors chose to take. These are model responses to scenarios, rather than measured effects of interventions.

Following real events

Tracking after 1 to 8 real events

59,769 sequences, eight recorded events per trial. Tracking scores 0.83 to 0.88.

0 0.5 1.0 1 2 3 4 5 6 7 8 real events fed in, one per step how much of the trial's movement it captures 0.84 after one event 0.83 after eight a picture that never updates

model, updated with each real eventno update: the trial as it stood at the start

How tracking was measured

Skill is one minus the model's error divided by the error of the never-updated picture, measured in the model's own representation against the state it computes from the full record after each event. The first event arrives a median 31 days in; the eighth, 304.

Most sequences span about a year. The model updates its state after each supplied event, the same operation used for scripted scenarios. For current forecasts, we read the latest full record: following eight events ranks outcomes about halfway between a stale reading and a fresh one.

Using the model

Ranking scores and 13 scripted actions

Ranking 0.80–0.86 on recruitment, termination, completion and timeline. Actions examined one at a time.

0.80–0.86ranking scores across recruitment, termination, completion and timeline outcomes
13 actionsscripted and examined one at a time

The same action can be compared across trials, or different actions on the same trial.

Forecasts, scenarios and consistency checks

Ranking scores use 0.5 as the chance baseline. Shortfall and termination probabilities are calibrated within a tenth of a percentage point.

The forecasts do not identify which event will happen next. Across 64 imagined futures per trial, outcome probabilities end within about a third of a point of one another and today’s forecast. For quiet trials, the model reads the record fresh rather than simulating six months of silence.

Scenario results come from one training run evaluated on a development period. Sequences of planned actions have not yet been evaluated.

A trial left alone must forecast about the same as the trial read today, and the model's own imagined futures must land where today's forecast points. Both checks pass on recruitment, termination, completion and timeline. An earlier version of the what-if read failed the first check because the quiet months were run with an input the model had never seen in training; the failure was traced, the read corrected, and every what-if number here is from the corrected version.

Sampled futures: next event chosen by the model, real example payloads, six 30-day steps, twelve monthly anchors of 20,000 trials. Ranking accuracy of the imagined-future ensemble against today's read: shortfall 0.82 vs 0.84, termination 0.815 vs 0.817, completion 0.81 vs 0.83, timeline 0.84 vs 0.86.

Related decks

This deck is the scripted-action half. The risk-ranking half is Trial Risk. Both are combined in Clinical Trial Oversight.