Strange Lab · Research
Clinical trial research · Full deck · September 2026

The question

Which trials fall short on recruitment in the next six months

Explore clinical-trial event histories

We need to know which trials will fall short, and notice when a later record is unlike the past.

  • That takes how a trial is expected to move, and a note when something is off.
  • A rule on phase, sponsor type, size and geography does not do this. An LLM cannot read hundreds of thousands of trials, or millions of registry events.
  • We use a world model of the trial registry to rank later trials and flag the ones unlike earlier activity.

With the model

60%of the trials that went on to fall short were in the model's riskiest fifth

Without the model

34%with a rule built from phase, sponsor type, size and geography

Yale · NCT04168710

Top-flagged trials recorded shortfalls at twelve times the overall rate

Yale, 1 March 2025. Within six months, enrollment was below half the target.

1 in 12top-flagged trials recorded shortfalls within six months
1 in 140active trials recorded shortfalls overall

Across 400 sponsors: 1,044 recruitment shortfalls expected, 982 recorded.

Where the flags concentrate

Many top-flagged recruitment risks already have a “reason stopped” event. The shortfall is written into the public record later.

Recruitment shortfalls were concentrated in US academic medical centers, including Montefiore, Miami, Maryland, Northwell and Yale: three to four percent of active trials, five times the average. Termination risk was concentrated in large pharma, including Roche, AbbVie, Sanofi, Amgen, Pfizer and Novartis.

Model of record: aact11-r0-readout. 5,037,391 registry events, 31,613 sponsors. Ranking accuracy six months ahead, AUROC: recruitment shortfall 0.80, termination 0.81, timeline slip 0.84, completion 0.79, against chance 0.50. Shortfall and termination probabilities are calibrated to within a tenth of a percentage point.

Review priorities

Watching the riskiest fifth catches 60% of shortfalls

A screen on phase, sponsor type, size and geography catches 34%.

0% 50% 100% 0% 20% 40% 60% 80% 100% share of trials watched, riskiest first shortfalls caught model 60% cohort screen 34% random

modelcohort screen

All outcomes and comparison details
If you watch the riskiestRecruitment shortfalls caughtTerminations caughtTimeline slips caughtCompletions caught
10%39% vs 20%46% vs 32%44% vs 18%39% vs 16%
20%60% vs 34%65% vs 47%66% vs 31%58% vs 29%
30%74% vs 45%76% vs 59%80% vs 42%71% vs 40%
50%90% vs 63%90% vs 74%95% vs 61%88% vs 62%

Model, with the cohort screen after "vs". The bottom 30% of the model's list holds 2 to 3% of the recruitment, termination and slip problems.

Six months, every active trial; each column uses the better of the two heads fitted for that question. The cohort screen uses 510 cells and scores each trial with the historical rate of its phase × sponsor type × size × country-count cell. Watching at random, 10% of attention catches 10% of problems. Chart: recruitment shortfall.

Comparing possible actions

The same action does not move every trial the same way

Design change, add sites, close recruitment. Compared with the trial left alone, across 20,000 trials.

Close recruitment and keep treating the patients already enrolled: completion probability up one to two points on average.

That average hides large differences between trials.

Compare actions and outcomes
Action taken at day 30Recruitment shortfallTerminationCompletionTimeline postponed > 6 months
Trial left alone: six-month probability0.83%0.39%3.24%4.59%
Change the design+61%+31%+23%+20%
Change the sponsor+34%+12%+13%+18%
Close recruitment, keep treating−45%+8%+59%−35%
Suspend the trial+24%+14%−7%−37%
Pull the completion date in+14%+28%+59%+22%
Add sites+6%+5%−3%+2%
Verify the record, nothing else0%+4%−5%−1%

Change is relative to the trial left alone, averaged over 20,000 active trials and sixteen real examples of each action. In absolute terms the moves are tenths of a percentage point on rare outcomes and one to two points on completion and timeline. These are the model's responses to an action, not measured effects of taking it; the only way to grade them is against a sponsor's own past decisions, which the public registry does not record.

Closing recruitment

Closing recruitment moves the top tenth, not the bottom half

Completion probability +5 points for the top tenth of trials. Almost no change for the bottom half.

+5 pointscompletion probability for the top tenth of trials
Almost no changefor the bottom half

Half the trials account for the change in the portfolio forecast.

Responses across trials and actions

Half the trials gain more than a point of completion probability when recruitment closes. One in six gains more than a point of shortfall risk from a design change. One in 25 moves more than a point in either direction when the record is merely verified.

Per-trial changes are the 10th, 50th and 90th percentiles across the 20,000 trials: closing recruitment moves completion by +0.0, +1.0 and +5.0 points; a design change moves shortfall by −0.0, +0.2 and +1.4 points; verification moves completion by −0.8, −0.1 and +0.2 points.

The model learned from actions sponsors chose to take. These are model responses to scenarios, rather than measured effects of interventions.

Following real events

The model follows a trial as events arrive

59,769 sequences, eight recorded events per trial. Tracking skill 0.83 to 0.88.

0 0.5 1.0 1 2 3 4 5 6 7 8 real events fed in, one per step how much of the trial's movement it captures 0.84 after one event 0.83 after eight a picture that never updates

model, updated with each real eventno update: the trial as it stood at the start

How tracking was measured

Skill is one minus the model's error divided by the error of the never-updated picture, measured in the model's own representation against the state it computes from the full record after each event. The first event arrives a median 31 days in; the eighth, 304.

Most sequences span about a year. The model updates its state after each supplied event, the same operation used for scripted scenarios. For current forecasts, we read the latest full record: following eight events ranks outcomes about halfway between a stale reading and a fresh one.

Full results: Trial Risk and Trial Dynamics.