FAQ

Strange Lab builds world models for work. They learn how a company or domain changes, forecast what happens next, and compare those forecasts under different possible actions.

What does Strange Lab build?

Every organisation already generates a continuous record of its world: mail, chat, documents, tickets, code, customer activity, decisions, and outcomes. Strange Lab turns that record into a source-linked timeline, reconstructs the state of the organisation at any point, and trains a world model on how that state changes.

The result is a living model of the organisation: memory of what happened, a view of what is happening now, forecasts of what comes next, and a way to compare modelled scenarios before choosing an action.

How do the pieces build on each other?

Each layer makes the next one possible:

  1. Record: activity from existing systems becomes a common stream of dated, source-linked events.
  2. State: those events become an as-of snapshot of the world — what was known, what was moving, and what was stuck.
  3. World model: the model learns how one state tends to become another, including how its forecasts differ across possible actions.
  4. Predictions: that learned structure becomes forecasts people can use: outcome, timing, risk, next actor, surprise, or likely consequence.
  5. Decisions: people and agents query those forecasts when deciding what to do next.
  6. Learning loop: later events score the forecast and can be retained as training evidence for a subsequent retraining cycle.

The layers reinforce one another. The same event spine powers replay, training, evaluation, and live forecasts. Later outcomes provide evidence for a governed retraining cycle.

What is a world model?

A world model learns the dynamics of one particular world rather than language from the internet at large. It carries the world's state forward by asking what is true now, what is likely to happen next, and how the forecast changes across possible actions. When there is no ground truth for the action not taken, the comparison shows how the model's forecast responds to a proposal rather than what the proposal would cause in the real world.

For a company, the world is its people, work, systems, customers, and decisions. For World Router, the world is an AI task moving through prompts, tools, tests, workers, and results. For our other public experiments, it can be a software project, the FDA oncology pathway, the U.S. economy, or Bismarck's political world. The machinery is the same: build the record, learn its dynamics, predict the next state.

How is this different from asking an LLM?

An LLM is a powerful interface to language. It works best when the problem is bounded: you know what to ask, which context to provide, and what a good answer looks like.

Open-ended work is unbounded. You can't know in advance which PRs are useful, or which workflows matter, or which projects are at risk, or where an agent should focus. A world model learns those dynamics from the full event history, and the LLM can use it as a core tool and interpret the result. The world model predicts; the LLM interprets.

How is this different from company memory, enterprise search, or process mining?

Company memory and enterprise search retrieve what the record says. Process mining reconstructs recurring flows and bottlenecks. Strange Lab uses the same source-linked events to learn how the state changes, test forecasts against what happened later, and compare the model's forecast under possible actions. Retrieval and process views can sit inside the system, but its defining job is prediction.

How do you test a forecast?

We stop the record at a historical cutoff and hide what happened later. The model reconstructs the state at that moment, makes its forecast, and the subsequent record grades it.

The holdout matches the real question. Project Confirm holds out entire drug assets. Project OSS predicts outcomes and timing from repository history. Bismarck gives the world model and GPT the same pre-cutoff evidence, then scores both against the hidden future. The result is an inspectable answer key, not a retrospective story.

What has the testing shown?

  • World Router applies a personal world model to live Codex work. Across 386 recorded worker runs from Rohit's own Codex history, it cut estimated model spend by 37% against assigning every run to Sol High.
  • Project OSS applies this to an unbounded repository problem: which PRs are useful, what is at risk, and where should an agent focus? Its single pooled model is trained on 7.04 million events from eight repositories. Its 28-day merge ranking beats structured-feature baselines in all eight, and it transfers to OpenClaw without retraining.
  • Project Confirm Oncology turns 17,703 FDA, trial, submission, and label events into 3,275 held-out test windows from entirely unseen drug assets. It reaches 0.857 AUROC on the next regulatory gate and finds a withdrawal signal with a median lead of about 24 months.
  • Bismarck beats GPT on 36 of 40 hidden-history forecasts.
  • Macro History records 33% lower error than pure momentum across the 159 largest-change windows. Its interactive public snapshot contains 5,339 records through June 9, 2026; it is a dated laboratory, not a current feed.

One record-to-world-model loop now works across software operations, clinical regulation, politics, and the economy. The Evidence page carries the scored results, comparators, and underlying receipts.

What does this provide inside a company?

The company gets an operational model it can query. It can reconstruct what was known when a decision was made, find repeated workflows and broken handoffs, forecast completion and delay, identify unusual work, measure process drift and key-person risk, and compare the model's forecasts for acting now versus later.

That same model gives agents a map of the organisation. Instead of automating whichever task is easiest to describe, an agent can target the workflows that matter, use the state of the company when choosing an action, and learn from the outcome.

How does it fit with existing systems?

Mail, chat, Drive, Jira, Salesforce, code hosts, and internal tools remain the systems where work happens. Strange Lab connects their permitted activity into one dated event layer and sits above them as the organisation's memory, simulation, and prediction system.

Because every event retains its source, a forecast can be traced back through the state that produced it. The organisation gets one model of how its work fits together while its people keep working in the tools they already use.

What data is needed, where does it run, and who controls it?

There is no universal event count. The minimum is a dated stream with repeated actions and observable outcomes, plus enough history to reserve an honest holdout. Bismarck was built from one dense PDF. Our current company work begins with a narrow, permitted slice and runs locally.

By default, company capture uses metadata, snippets, source links, and permissions rather than full document text. The customer controls its sources, retention, deletion, redaction, model training, and any sharing outside its environment. Customer-specific models and evaluations remain separate from public research.

Can an agent act on the model automatically?

Yes, at the level of autonomy the organisation chooses. A deployment can begin with read-only analysis, move into shadow recommendations, add approval-gated actions, and automate narrow workflows once their outcomes can be checked reliably.

The world model makes that progression measurable. Every proposed action starts from an explicit state, every forecast can be recorded, and every result can be graded and retained as evidence for a later governed retraining cycle.

What can I see today?

Start with the worlds we have built so far: Life Sciences, Software, Cybersecurity, Games, and Fiction. Each page uses current work as a concrete example, and this set will grow as we model more worlds. You can also try Bismarck or the dated Macro History research world. The Evidence page carries the scored results.

Where does this go?

As the event record grows, the same world model becomes a deeper operational memory, a better forecasting system, and a more capable decision layer for both people and agents. A company can ask what is happening, what happens next, what the model forecasts under different possible actions, and which option looks most promising within those forecasts.

Can we work together?

We partner with people who have a genuinely open-ended problem, or work that resists being reduced to a fixed workflow. Talk to us.