All articles Research

Foundation Models for Causal Inference (CausalFM)

CausalFM is a transformer-based foundation model trained on simulated causal worlds so it can reason about cause and effect, estimate treatment effects, and answer counterfactual what-if questions rather than just detect correlations.

CausalFM is a new type of AI model designed to learn cause–effect relationships, not just statistical patterns. Where most models answer “what usually goes together?”, CausalFM aims to answer “what causes what?” — a key step toward AI that reasons like a scientist. It sits at the intersection of two threads that matured over 2024–2026: foundation models (large pre-trained transformers) and the Prior-Data Fitted Network (PFN) line of work that turned Bayesian inference into a single forward pass.

CausalFM = a foundation model that can answer “what causes what?” and “what happens if we change something?” — in context, without retraining.

The problem it solves

Traditional AI — including large language models and most deep-learning systems — learns correlations: it notices that things occur together. For example, it learns that smoking and cancer appear together in data, but it doesn't truly understand whether smoking causes cancer.

CausalFM instead learns cause → effect relationships, so it can distinguish "these happen together" from "this one produces that one."

The gap matters because of confounding. A variable that influences both a treatment and an outcome creates a correlation that is not causal. The classic illustration: ice-cream sales correlate with drowning deaths, but neither causes the other — summer heat drives both. A purely correlational model will happily "predict" drownings from ice-cream sales and be badly wrong the moment you intervene (e.g., ban ice cream to save swimmers).

flowchart TD H["Hidden cause: hot weather"] --> I["Ice-cream sales"] H --> D["Drownings"] I -. "spurious correlation" .-> D

CausalFM's job is to recover the arrows in diagrams like this — the structure — and then answer questions that require knowing that structure, not just the joint distribution over the variables.

The ladder of causation

Judea Pearl's Ladder of Causation is the cleanest way to frame what CausalFM adds. It describes three rungs of reasoning, each strictly more powerful than the one below.

RungQuestion typeFormal operatorExamplePlain LLM?
1. AssociationSeeing / observingP(y | x)“Patients on drug A tend to recover.”Yes
2. InterventionDoing / actingP(y | do(x))“If we give drug A, recovery rises.”Not reliably
3. CounterfactualImagining / retrospectionP(y_x | x', y')“This patient recovered — would they have without the drug?”No

The key distinction is between P(y | x) (observing that x happened) and P(y | do(x)) (forcing x to a value, severing whatever normally determined it). Standard prediction lives entirely on Rung 1. CausalFM targets Rungs 2 and 3, which is exactly where high-stakes decisions live: policy, medicine, pricing, and treatment allocation.

How it works

CausalFM builds on the transformer architecture but adds a few key ideas:

  • Trained on simulated worlds. It uses synthetic data generated from causal models — typically structural causal models (SCMs) with randomly sampled graphs and mechanisms — learning scenarios of the form “if A changes, what happens to B?”
  • Prior-Data Fitted Networks (PFNs). The model is trained on defined cause-effect rules (priors) and then learns to approximate Bayesian posterior inference in a single forward pass, generalizing beyond any single prior instance.
  • In-context causal reasoning. Given a dataset in its context window, it can answer “what caused this?” and “what if we change X?” — a form of counterfactual reasoning — without gradient updates at test time.

Structural causal models as the data-generating prior

A structural causal model represents each variable as a deterministic function of its direct causes plus an independent noise term:

X := f_X(U_X)
Z := f_Z(X, U_Z)
Y := f_Y(X, Z, U_Y)

Here the assignment operator := is directional: changing X propagates to Y, but not vice versa. An intervention do(X = x) replaces the first line with the constant X := x, cutting the arrows coming into X while leaving everything downstream intact. This surgery on the graph is precisely what a correlational model cannot represent, and it is what CausalFM learns to emulate by seeing many such worlds.

Prior-Data Fitted Networks in depth

PFNs (Müller et al., popularized through TabPFN for tabular prediction and extended with a stronger TabPFN v2 in 2025) reframe learning as in-context Bayesian inference. Instead of fitting parameters to one dataset, you:

  1. Define a prior — a distribution over datasets. For CausalFM this prior is a distribution over SCMs: sample a graph, sample mechanisms, sample noise, then draw observational and interventional samples.
  2. Pre-train a transformer to predict held-out targets given the rest of a sampled dataset, minimizing cross-entropy. Averaged over the prior, this objective drives the network toward the true Bayesian posterior predictive distribution.
  3. At inference, feed a real dataset as context. The network outputs a calibrated predictive (or interventional) distribution in one forward pass — no per-dataset training loop.

The elegant consequence: all the expensive computation happens once, offline, during pre-training. The prior encodes your causal assumptions, and the network amortizes inference over every dataset the prior can produce. This is why PFN-style models are fast at test time — inference is a forward pass, not an optimization.

AspectClassic per-dataset modelPFN / CausalFM
Where fitting happensAt test time, per datasetOnce, during pre-training
Test-time costFull training loopSingle forward pass
Assumptions live inModel class + hyperparametersThe prior over SCMs
OutputPoint estimate or one posteriorAmortized posterior predictive

The training and inference pipeline

The full loop, from prior to answer, looks like this:

flowchart TD P["Prior over SCMs
graphs + mechanisms + noise"] --> S["Sample synthetic datasets
observational + interventional"] S --> T["Pre-train transformer
predict held-out targets"] T --> M["CausalFM weights (frozen)"] R["Real dataset in context"] --> M M --> O["Posterior predictive:
P(y given do(x)), effects, counterfactuals"]

Two properties are worth stressing. First, the prior is a modeling choice: broaden it (more graph shapes, nonlinear mechanisms, heavier-tailed noise) and the model becomes more robust but harder to train; narrow it and the model is sharper but more brittle to prior misspecification. Second, because interventional and counterfactual queries are simulated during pre-training, the network can answer Rung-2 and Rung-3 questions at test time even though the real dataset it is handed may be purely observational — provided the query is identifiable from that data under the assumed structure.

A worked example

Input: a patient took drug A and recovered.

Unlike a standard LLM, CausalFM can reason about questions such as:

  • Did the drug actually cause the recovery?
  • What would have happened if the patient hadn't taken it?
  • What if we changed the dosage?

Concretely, suppose severity Z influences both whether a patient receives the drug X and whether they recover Y — a textbook confounder. The naive association P(Y=recover | X=drug) is contaminated by Z, because sicker patients may be preferentially treated. The causal quantity we actually want is the interventional average:

ATE = E[Y | do(X=1)] - E[Y | do(X=0)]

Under the back-door criterion, adjusting for Z identifies it:

P(Y | do(X=x)) = sum_z  P(Y | X=x, Z=z) * P(Z=z)
flowchart TD Z["Severity Z (confounder)"] --> X["Drug X"] Z --> Y["Recovery Y"] X --> Y

CausalFM effectively learns to perform this adjustment implicitly: because its pre-training distribution contained many confounded worlds where the correct answer required back-door adjustment, the network internalizes the pattern and applies it in context. The counterfactual question — "would this recovered patient have recovered without the drug?" — is Rung 3 and requires the full SCM, including the noise term realized for that individual; this is where CausalFM's SCM-based prior is essential and a plain regression model cannot follow.

Causal estimands it targets

A general causal foundation model aims to cover a family of standard estimands within one interface:

EstimandMeaningTypical use
ATEAverage treatment effect across a populationDoes the policy help on average?
CATEConditional (per-subgroup) treatment effectWho benefits most? Personalization
ATTEffect on the treatedWas treating these units worth it?
CounterfactualOutcome for a specific unit under an alternative actionIndividual what-if, attribution

Classical identification strategies map onto CausalFM's capabilities: back-door adjustment when confounders are observed, the front-door criterion when a mediator is available but confounders are not, and instrumental variables when an exogenous nudge exists. A causal foundation model is attractive because it can, in principle, select and apply the right strategy from the structure implied by its context rather than requiring the analyst to hand-code each one.

Why it's a breakthrough

  • Moves AI beyond correlation — it can reason about why things happen, not just what co-occurs.
  • Works in critical domains — medicine, economics, and policy decisions, where acting on a spurious correlation is costly or dangerous.
  • A general framework — it can address back-door and front-door causal problems and instrumental-variable analysis within one model, rather than a bespoke estimator per problem.
  • Amortized and fast — like TabPFN, inference is a forward pass, so a practitioner gets an answer in seconds instead of standing up a custom causal pipeline.
  • Rides the 2025–2026 tabular FM wave — the same in-context, prior-fitted recipe that made TabPFN v2 competitive with tuned gradient-boosted trees on small tabular data now extends naturally to causal queries.

How it compares

Model typeWhat it primarily learnsHighest rung reached
GPT / LLMPatterns (correlation) over languageAssociation
Reasoning modelsStep-by-step problem solvingAssociation (+ verbal what-if)
TabPFN / tabular PFNAmortized predictive posteriorAssociation
Hand-built causal estimatorOne estimand under fixed assumptionsIntervention / Counterfactual
CausalFMCause-and-effect (causality), amortizedIntervention + Counterfactual

Limitations and open problems

A causal foundation model is not magic, and it inherits the fundamental constraints of causal inference:

  • Identifiability is a hard wall. No model can recover an effect that the data cannot identify. If a confounder is unobserved and no valid instrument or front-door path exists, the answer is genuinely undetermined — CausalFM can at best report a range, not a point.
  • Prior misspecification. If the real data-generating process falls outside the model's prior over SCMs (unusual graph size, mechanism shape, or noise), estimates degrade — the causal analogue of out-of-distribution failure.
  • Context-window limits. Like other PFNs, the amount of data it can condition on in one pass is bounded, which constrains very large datasets or very high-dimensional problems.
  • Assumptions are still assumptions. Back-door and front-door validity, no interference between units (SUTVA), and correct variable definitions remain the analyst's responsibility. The model automates estimation, not causal thinking.
  • Validation is intrinsically hard. Counterfactuals are never observed, so evaluating Rung-3 answers relies on synthetic benchmarks or rare randomized experiments.

Key takeaways

  • CausalFM combines the foundation-model and PFN recipes: pre-train once on simulated causal worlds, then answer causal queries in context.
  • It targets Rungs 2 and 3 of Pearl's ladder — interventions and counterfactuals — where plain LLMs and standard predictors fail.
  • Its power and its fragility both come from the same place: the prior over structural causal models.
  • It automates causal estimation, but not the identification assumptions that make estimation valid.

One-line intuition: CausalFM tries to make AI think like a scientist, not just a pattern recognizer — a transformer-based foundation model trained on causal data so it can perform cause-effect reasoning and answer “what happens if we change something?”

← Back to all articles