← All posts

Imputation in Clinical Research: Prediction vs Causal Thinking Explained

Clinical Epidemiology ResearchUniqcret doctor knowledgesData Analytics or Statistics
On this page

A Practical Guide to Prediction, Causal Inference, and Longitudinal Data

1. Why Imputation Needs Careful Thinking

Missing data occur in almost all clinical datasets:

Imputation is used to replace missing values so that:

However, imputation is not a neutral technical step.It directly affects:

The same imputation strategy can be correct in one study and wrong in another.


2. First Rule: Define Your Research Goal

Before choosing any imputation method, you must clearly define:

What question am I trying to answer?

There are two fundamentally different goals:

  1. Prediction → “Can we predict an outcome for a future patient?”
  2. Causal inference → “What is the true effect of an exposure or treatment?”

These goals follow different logic and require different imputation rules.


3. Prediction Models (CPMs): Why Y Must Not Be Imputed

3.1 What Is a Prediction Model?

Clinical Prediction Models (CPMs) aim to:

Examples:

In real clinical use:

3.2 Do Not Impute the Outcome (Y)

In prediction studies:

Never impute YNever use Y to help impute X

This rule is absolute.

3.3 Data Leakage Explained Simply

Data leakage happens when information that would not be available in real life is used during model development.

If Y is used during imputation:

This creates a false sense of accuracy.

Think of it like this:

It is like giving students the exam answersand then claiming they are excellent test-takers.

3.4 Why Performance Becomes Overestimated

Using Y in imputation causes:

But when applied to new patients:

This is one of the most common reasons prediction models fail in practice.

3.5 Conceptual Error: Prediction vs Explanation

Prediction asks:

“Given what we know now, what will happen?”

Using Y during imputation answers a different question:

“Given the outcome, can we reconstruct the predictors?”

That is a retrospective explanation, not prediction.


4. Causal Inference: Why the Rules Are Different

4.1 What Is Causal Analysis?

Causal (etiologic or therapeutic) research asks:

“What would have happened if the exposure or treatment were different?”

The goal is to estimate:

Here, the focus is:

Prediction performance is irrelevant.

4.2 Using Y in Imputation Can Be Acceptable

In causal analysis, using Y in multiple imputation can be appropriate if done carefully.

Why?

Because:

When missingness is Missing At Random (MAR):

This aligns with:

4.3 Key Difference in Mindset

AspectPredictionCausal Inference
GoalFuture accuracyTruth of effect
Use of YForbiddenOften acceptable
Performance metricsCentralSecondary
Bias controlLimitedPrimary focus

5. Longitudinal Data: Why Y Can Be Imputed

5.1 What Makes Longitudinal Data Special?

Longitudinal data include:

Examples:

Missingness often occurs:

5.2 Imputing Y Is Often Reasonable

In longitudinal settings:

Therefore: ✅ Imputing missing Y values is often appropriate

5.3 The Critical Rule: Respect Time Order

You may use:

You must NOT use:

Violating time order creates temporal leakage, which is as harmful as data leakage.


6. Practical Summary Table

Study TypeImpute Y?Use Y to Impute X?Key Reason
Clinical prediction models❌ No❌ NoPrevent data leakage
Machine learning prediction❌ No❌ NoReal-world realism
Causal/etiologic studies✅ Sometimes✅ SometimesReduce bias
Therapeutic effect estimation✅ Sometimes✅ SometimesCausal validity
Longitudinal outcomes✅ Yes (with rules)✅ YesTemporal structure

7. Common Mistakes to Avoid

Key Insight

Imputation is part of study design, not just data cleaning.

Good imputation cannot fix a poorly defined research question.


Final Takeaways

0
Message for International and Thai ReadersUnderstanding My Medical Context in ThailandRead more →Message for International and Thai ReadersUnderstanding My Broader Content Beyond MedicineRead more →

Comments

No comments yet. Be the first to share your thoughts.

Sign in to comment