The question
Which trials fall short on recruitment in the next six months
We need to know which trials will fall short, and notice when a later record is unlike the past.
- That takes how a trial is expected to move, and a note when something is off.
- A rule on phase, sponsor type, size and geography does not do this. An LLM cannot read hundreds of thousands of trials, or millions of registry events.
- We use a world model of the trial registry to rank later trials and flag the ones unlike earlier activity.
With the model
60%of the trials that went on to fall short were in the model's riskiest fifthWithout the model
34%with a rule built from phase, sponsor type, size and geographyYale · NCT04168710
Top-flagged trials recorded shortfalls at twelve times the overall rate
Yale, 1 March 2025. Within six months, enrollment was below half the target.
Across 400 sponsors: 1,044 recruitment shortfalls expected, 982 recorded.
Where the flags concentrate
Many top-flagged recruitment risks already have a “reason stopped” event. The shortfall is written into the public record later.
Recruitment shortfalls were concentrated in US academic medical centers, including Montefiore, Miami, Maryland, Northwell and Yale: three to four percent of active trials, five times the average. Termination risk was concentrated in large pharma, including Roche, AbbVie, Sanofi, Amgen, Pfizer and Novartis.
Model of record: aact11-r0-readout. 5,037,391 registry events, 31,613 sponsors. Ranking accuracy six months ahead, AUROC: recruitment shortfall 0.80, termination 0.81, timeline slip 0.84, completion 0.79, against chance 0.50. Shortfall and termination probabilities are calibrated to within a tenth of a percentage point.
Review priorities
Watching the riskiest fifth catches 60% of shortfalls
A screen on phase, sponsor type, size and geography catches 34%.
modelcohort screen
All outcomes and comparison details
| If you watch the riskiest | Recruitment shortfalls caught | Terminations caught | Timeline slips caught | Completions caught |
|---|---|---|---|---|
| 10% | 39% vs 20% | 46% vs 32% | 44% vs 18% | 39% vs 16% |
| 20% | 60% vs 34% | 65% vs 47% | 66% vs 31% | 58% vs 29% |
| 30% | 74% vs 45% | 76% vs 59% | 80% vs 42% | 71% vs 40% |
| 50% | 90% vs 63% | 90% vs 74% | 95% vs 61% | 88% vs 62% |
Model, with the cohort screen after "vs". The bottom 30% of the model's list holds 2 to 3% of the recruitment, termination and slip problems.
Six months, every active trial; each column uses the better of the two heads fitted for that question. The cohort screen uses 510 cells and scores each trial with the historical rate of its phase × sponsor type × size × country-count cell. Watching at random, 10% of attention catches 10% of problems. Chart: recruitment shortfall.
Comparing possible actions
The same action does not move every trial the same way
Design change, add sites, close recruitment. Compared with the trial left alone, across 20,000 trials.
That average hides large differences between trials.
Compare actions and outcomes
| Action taken at day 30 | Recruitment shortfall | Termination | Completion | Timeline postponed > 6 months |
|---|---|---|---|---|
| Trial left alone: six-month probability | 0.83% | 0.39% | 3.24% | 4.59% |
| Change the design | +61% | +31% | +23% | +20% |
| Change the sponsor | +34% | +12% | +13% | +18% |
| Close recruitment, keep treating | −45% | +8% | +59% | −35% |
| Suspend the trial | +24% | +14% | −7% | −37% |
| Pull the completion date in | +14% | +28% | +59% | +22% |
| Add sites | +6% | +5% | −3% | +2% |
| Verify the record, nothing else | 0% | +4% | −5% | −1% |
Change is relative to the trial left alone, averaged over 20,000 active trials and sixteen real examples of each action. In absolute terms the moves are tenths of a percentage point on rare outcomes and one to two points on completion and timeline. These are the model's responses to an action, not measured effects of taking it; the only way to grade them is against a sponsor's own past decisions, which the public registry does not record.
Closing recruitment
Closing recruitment moves the top tenth, not the bottom half
Completion probability +5 points for the top tenth of trials. Almost no change for the bottom half.
Half the trials account for the change in the portfolio forecast.
Responses across trials and actions
Half the trials gain more than a point of completion probability when recruitment closes. One in six gains more than a point of shortfall risk from a design change. One in 25 moves more than a point in either direction when the record is merely verified.
Per-trial changes are the 10th, 50th and 90th percentiles across the 20,000 trials: closing recruitment moves completion by +0.0, +1.0 and +5.0 points; a design change moves shortfall by −0.0, +0.2 and +1.4 points; verification moves completion by −0.8, −0.1 and +0.2 points.
The model learned from actions sponsors chose to take. These are model responses to scenarios, rather than measured effects of interventions.
Following real events
The model follows a trial as events arrive
59,769 sequences, eight recorded events per trial. Tracking skill 0.83 to 0.88.
model, updated with each real eventno update: the trial as it stood at the start
How tracking was measured
Skill is one minus the model's error divided by the error of the never-updated picture, measured in the model's own representation against the state it computes from the full record after each event. The first event arrives a median 31 days in; the eighth, 304.
Most sequences span about a year. The model updates its state after each supplied event, the same operation used for scripted scenarios. For current forecasts, we read the latest full record: following eight events ranks outcomes about halfway between a stale reading and a fresh one.
Full results: Trial Risk and Trial Dynamics.