Why 1 Minus Kaplan-Meier Overstates Risk When Patients Can Die First

Clinical Epidemiology ResearchMethodology and Research DesignUniqcret doctor knowledges
Why 1 Minus Kaplan-Meier Overstates Risk When Patients Can Die First
On this page

อ่านฉบับภาษาไทย (Thai version)

Abstract

When patients can die before the event of interest, death is not censoring: it makes the event impossible rather than unobserved. One minus Kaplan-Meier (the usual survival curve flipped into a risk, with deaths treated as lost follow-up) estimates a hypothetical risk without death, under an untestable assumption, and is never lower than the actual risk. The actual risk is the cumulative incidence function, estimated by the Aalen-Johansen estimator, which weights each event fraction by the share still event-free. In a hand example, 20 of 100 patients die before any recurrence: 1 minus Kaplan-Meier gives 12.5% and the cumulative incidence 10%. A second hand example follows twelve cancer patients for up to 12 months: the cumulative incidence of recurrence is 0.380, against 0.489 from 1 minus Kaplan-Meier. In a simulated heart-failure trial of 1,500 patients, 1 minus Kaplan-Meier overstates the first-admission risk in both arms. This article shows how to compute and check the cumulative incidence, compare groups with Gray's test and report both causes.


Visual summary. Simulated data.

A recurrence risk that looks too high

An oncology team sees older patients a few weeks after cancer surgery. At that visit it quotes each of them a risk that the cancer returns within five years. The figure comes from the team's own follow-up records, analysed as one minus a Kaplan-Meier curve for recurrence. Patients who died of other causes were treated as censored, as if their records had simply stopped.

A patient's daughter asks why the quoted risk is higher than the share of past patients who had a recurrence within five years. Many of them, she points out, died of heart disease or pneumonia first. She has a point: a patient who has died cannot have a recurrence later, so a death is not a lost record.

Censoring and competing events are different

Censoring means that follow-up ended without the event while the event could still happen later, for example because the patient moved away or the study closed. A competing event makes the event of interest impossible afterwards, such as death before recurrence. Data with competing events carry competing risks, and Part 1 of this series maps where they sit among time-to-event questions.

After censoring, recurrence is possible but unseen. After death, it will never happen. Treating a death as censoring therefore pictures a future the patient cannot have [1, 2].

Let $T$ be the time from surgery to the first event of either kind. The cumulative incidence function of cause $k$, written $\mathrm{CIF}_k(t)$, is the probability that the first event is of cause $k$ and happens by time $t$. Here the causes are recurrence (R) and death before recurrence (D), giving $\mathrm{CIF}_R(t)$ and $\mathrm{CIF}_D(t)$. All-cause event-free survival, $S(t) = P(T > t)$, is the probability of being alive and recurrence-free at time $t$.

Hand example: one hundred patients, deaths first

Hand example. One hundred patients are followed after surgery, and nobody is lost to follow-up. Early on, 20 of them die of other causes before anyone has had a recurrence. Later, 10 of the 80 survivors have a recurrence, and the other 70 reach the end of follow-up alive and recurrence-free.

  1. Cumulative incidence

    \[ \mathrm{CIF}_R = \frac{10}{100} = 10\% \]

    With nobody censored before the last event, this is the share of the starting group who had a recurrence.

  2. 1 minus Kaplan-Meier, deaths censored

    \[ 1 - \mathrm{KM} = \frac{10}{80} = 12.5\% \]

    KM is the Kaplan-Meier estimate of staying recurrence-free. It drops the 20 deaths from the risk set, the patients still followed and event-free, as if they were lost records, leaving 80 at risk.

  3. The gap

    \[ 12.5\% - 10\% = 2.5 \text{ percentage points} \]

    The 20 deaths have been treated as patients who could still recur.

Result: 1 minus Kaplan-Meier gives 12.5% and the cumulative incidence 10%, a contrast that rests on all 20 deaths coming before the first recurrence.

The order of events matters for 1 minus Kaplan-Meier only. If some recurrences came before some deaths, it would fall between 10% and 12.5%, and if all 10 came first it would equal 10%. The cumulative incidence stays at 10% in every order, because with nobody censored before the last event it is the plain proportion who recurred.

The Aalen-Johansen estimator

Real cohorts have censoring, so the cumulative incidence needs an estimator rather than a simple proportion. The Aalen-Johansen estimator builds it one event time at a time [1]:

$$\widehat{\mathrm{CIF}}_k(t) = \sum_{t_j \le t} \hat S(t_{j-1})\,\frac{d_{kj}}{n_j}$$

A hat marks an estimate from data. The sum runs over the distinct times $t_j$ up to $t$ at which any first event occurs. At each, $n_j$ is the number still event-free and under follow-up just before $t_j$ (the risk set), and $d_{kj}$ is the number of first events of cause $k$ at $t_j$.

$\hat S(t_{j-1})$ is the all-cause event-free survival just after the previous event time, with $\hat S(t_0) = 1$. It is itself a Kaplan-Meier estimate in which every first event counts, whatever its cause. Each term is therefore the share of the cohort still event-free times the fraction of the risk set with an event of cause $k$. Censored patients simply leave the risk set, on the assumption that their future risks resemble those of patients still followed.

Why the risk of death matters

In continuous time the same idea shows why a competing event changes the risk of recurrence. A hazard is the event rate at a given moment among patients still event-free. The cause-specific hazard of recurrence, $\lambda_R(t)$, counts only recurrences. Then

$$\mathrm{CIF}_R(t) = \int_0^t S(u)\,\lambda_R(u)\,du$$

The cumulative incidence adds up the recurrence rate over time, weighted by the share still alive and recurrence-free. Death lowers $S(u)$, so the risk of recurrence depends on the hazard of death as well as on the hazard of recurrence [1, 3]. A treatment that raises the death rate can lower the cumulative incidence of recurrence without changing the hazard of recurrence.

Hand example: twelve patients, step by step

Hand example. Twelve patients are followed for up to 12 months after cancer surgery. In order: death at month 2, recurrence at 3, censored at 4, death at 5, recurrences at 6 and 7, censored at 8, death at 9, recurrence at 10, censored at 11 and two censored at 12. That is 4 recurrences, 3 deaths before recurrence and 5 censored patients.

  1. Month 2: death

    \[ \widehat{\mathrm{CIF}}_D = 1 \times \frac{1}{12} = 0.083 \]

    All 12 are at risk, and $\hat S$ drops to $1 - \tfrac{1}{12} = 0.917$.

  2. Month 3: recurrence

    \[ \widehat{\mathrm{CIF}}_R = \frac{11}{12} \times \frac{1}{11} = \frac{1}{12} = 0.083 \]

    Eleven are at risk, and $\hat S$ drops to 0.833.

  3. Month 4: censoring

    One patient leaves the risk set, and no estimate changes.

  4. Month 5: death

    \[ \widehat{\mathrm{CIF}}_D = 0.083 + 0.833 \times \frac{1}{9} = 0.176 \]

    Nine are at risk, and $\hat S$ drops to 0.741.

  5. Month 6: recurrence

    \[ \widehat{\mathrm{CIF}}_R = 0.083 + 0.741 \times \frac{1}{8} = 0.176 \]

    Eight are at risk, and $\hat S$ drops to 0.648.

  6. Month 7: recurrence

    \[ \widehat{\mathrm{CIF}}_R = 0.176 + 0.648 \times \frac{1}{7} = 0.269 \]

    Seven are at risk, and $\hat S$ drops to 0.556.

  7. Months 8 and 9: censoring, then death

    \[ \widehat{\mathrm{CIF}}_D = 0.176 + 0.556 \times \frac{1}{5} = 0.287 \]

    Five are at risk at month 9, and $\hat S$ drops to 0.444.

  8. Month 10: recurrence

    \[ \widehat{\mathrm{CIF}}_R = 0.269 + 0.444 \times \frac{1}{4} = 0.380 \]

    Four are at risk, and $\hat S$ drops to 0.333. The censorings at months 11 and 12 change nothing.

Result: At month 12, $\hat S = 0.333$, $\widehat{\mathrm{CIF}}_R = 0.380$ and $\widehat{\mathrm{CIF}}_D = 0.287$.

The Aalen-Johansen table for the twelve patients

Hand example. Each increment is $\hat S$ just before the event multiplied by $1/n_j$ (one event at each time), and it goes to the column of the cause that occurred. The censorings at months 4, 8, 11 and 12 only shrink the risk set.
MonthFirst eventAt risk, $n_j$$\hat S$ just beforeIncrement$\widehat{\mathrm{CIF}}_R$$\widehat{\mathrm{CIF}}_D$$\hat S$ just after
2Death121.0000.0830.0000.0830.917
3Recurrence110.9170.0830.0830.0830.833
5Death90.8330.0930.0830.1760.741
6Recurrence80.7410.0930.1760.1760.648
7Recurrence70.6480.0930.2690.1760.556
9Death50.5560.1110.2690.2870.444
10Recurrence40.4440.1110.3800.2870.333

A built-in check: everyone is somewhere

At any time each patient is in exactly one of three states: event-free, recurred first or died first. So the three probabilities add to one:

$$1 = S(t) + \mathrm{CIF}_R(t) + \mathrm{CIF}_D(t)$$

Event-free survival and the cumulative incidences of all causes share out the whole cohort. At month 12 in the hand example, $0.333 + 0.380 + 0.287 = 1.000$, and every row of the table passes the same check up to rounding. A table that fails it contains a slip, or a 1 minus Kaplan-Meier value where a cumulative incidence belongs.

1 minus Kaplan-Meier for the twelve patients, deaths censored

Hand example. The same twelve patients, with recurrence as the event and each death censored. $\widehat{\mathrm{KM}}_R(t)$ is the Kaplan-Meier estimate of staying recurrence-free by time $t$. At each recurrence it is multiplied by the factor shown, the share of the risk set that stays recurrence-free; the risk sets are those of the Aalen-Johansen table. The deaths at months 2, 5 and 9 shrink the risk set but never lower the curve.
MonthAt riskFactorKaplan-Meier, recurrence-free1 minus Kaplan-Meier
31110/110.9090.091
687/80.7950.205
776/70.6820.318
1043/40.5110.489

Same risk sets, a different multiplier

At month 12, 1 minus Kaplan-Meier is 0.489, against a cumulative incidence of recurrence of 0.380. It is higher by 0.109, or 1.29 times as high. Written as a sum, 1 minus Kaplan-Meier has the same shape as the Aalen-Johansen estimator:

$$1 - \widehat{\mathrm{KM}}_R(t) = \sum_{t_j \le t} \widehat{\mathrm{KM}}_R(t_{j-1})\,\frac{d_{Rj}}{n_j}$$

The only change is the multiplier. Aalen-Johansen multiplies each recurrence fraction $d_{Rj}/n_j$ by $\hat S$, which drops at deaths. Kaplan-Meier multiplies it by its own recurrence-free curve, which does not. So 1 minus Kaplan-Meier is never lower than the cumulative incidence, and it is higher once a death comes before a recurrence [1, 2].

At month 10, for example, Aalen-Johansen adds $0.444 \times \tfrac{1}{4} = 0.111$, while Kaplan-Meier adds $0.682 \times \tfrac{1}{4}$. The three deaths have lowered $\hat S$ but left the Kaplan-Meier curve untouched, so every later recurrence weighs more in 1 minus Kaplan-Meier.

A hypothetical risk, not an arithmetic error

What does 0.489 estimate, then? It is the probability of recurrence by month 12 in a hypothetical world in which death before recurrence could not happen. It also assumes that removing death would not change the hazard of recurrence: the patients who died would have recurred at the same rate as those still followed. The data cannot test that assumption, because nobody is observed after death [1, 3].

So 1 minus Kaplan-Meier is not an arithmetic error. It answers a different, hypothetical question, which may be of interest when that world is the explicit target and its assumption is stated. The risk the daughter asked about is the cumulative incidence, the probability that a patient like her parent actually has a recurrence [2, 4].

Hand example. The slider sets how many of the 100 patients die before any recurrence (a higher hazard of death means more die first), while the share of survivors who recur stays fixed. 1 minus Kaplan-Meier, which treats deaths as censored, therefore stays at 12.5%, and the cumulative incidence of recurrence falls further below it. A strip below repeats the comparison in the simulated heart-failure trial (simulated data).

Comparing groups: Gray's test

Gray's test asks whether the cumulative incidence curves of one cause are the same across groups, such as two treatment arms, over follow-up [5]. It plays the part that the log-rank test, a test of whether survival curves differ, plays for Kaplan-Meier curves. Unlike a log-rank test with deaths censored, it keeps patients with the competing event in the comparison.

A small P value is evidence that the curves differ somewhere over follow-up; it does not say by how much at a stated time. Nor does it say why: a lower cumulative incidence of recurrence can come from a lower hazard of recurrence or from a higher hazard of death. When curves cross, differences of opposite sign can cancel and the test can miss them.

Regression for competing risks is the subject of Part 5 of this series: the cause-specific Cox model, a regression for the hazard of one cause among patients still event-free, and the Fine-Gray model for the cumulative incidence.

The same contrast in a simulated heart-failure trial

The rest of the article analyses one simulated dataset. A randomised trial enrols 1,500 adults after a heart-failure admission, 750 per arm, and follows them for 24 to 36 months. The outcome is the first heart-failure admission after randomisation, and death before any admission is the competing event. The drug is fictional, and the simulation lets it lower both the admission rate and the rate of cardiovascular death.

The opening scene was about cancer and this trial is about heart failure, but the logic is the same. In both, the event of interest can happen only while the patient is alive. Read 'admission' for 'recurrence', and every formula above applies.

In the data, 683 patients had a first admission (364 control, 319 treated), and 514 died before any admission (258 and 256). Another 303 (128 control, 175 treated) were censored without either event, at the close of the study or on loss to follow-up. Losses in the simulation are unrelated to risk, as both estimators assume. The simulation was built from known rates, so each estimate can be checked against the simulated truth.

First admission by 12 and 24 months: cumulative incidence versus 1 minus Kaplan-Meier

Simulated data. Point estimates of probabilities, shown to four decimals as computed. Each estimate is set beside its simulated truth, so confidence intervals are not shown, although the R function cuminc also returns the variances they are built from. The simulated trial data and the true values behind the simulated-truth rows are not published, so the table can be read but not rerun.
Quantity (probability)Control, 12 monthsTreated, 12 monthsControl, 24 monthsTreated, 24 months
Cumulative incidence of admission, Aalen-Johansen estimate0.35340.26970.46060.4116
Cumulative incidence of admission, simulated truth0.34170.27340.45360.3899
1 minus Kaplan-Meier for admission, deaths censored, estimate0.41320.30310.58990.5188
Risk of admission if death were removed, simulated truth0.40100.31240.59500.4935
Cumulative incidence of death before admission, Aalen-Johansen estimate0.22870.20940.33300.3105

Reading the table: 1 minus Kaplan-Meier is higher in both arms

By 24 months the cumulative incidence of admission is 46.1% in the control arm and 41.2% in the treated arm. Gray's test compares the two curves over the whole of follow-up [5]. It gives a chi-square statistic of 5.99 on 1 degree of freedom, P = 0.014. It cannot say whether the difference comes from the drug's action on admission or on death.

One minus Kaplan-Meier gives 59.0% and 51.9% at 24 months, higher in both arms. The gap is already there at 12 months, 35.3% against 41.3% in the control arm, and it widens as deaths accumulate. Each estimator lands within a few percentage points of the simulated truth for its own quantity.

Kaplan-Meier manages that here only because the simulation makes death independent of a patient's tendency to be admitted. No real trial can check that assumption. Aalen-Johansen needs no assumption about death, only that censoring is unrelated to prognosis.

The last row is easy to misread. Death before admission is not all deaths, and its cumulative incidence depends on both hazards. The drug lowers admissions, so more treated patients stay event-free and at risk of dying first, which offsets their lower death rate. That is why the two arms end up close, although the simulation lets the drug lower cardiovascular death.

The code: a hand-built Aalen-Johansen sum checked in Stata, Gray's test in R

In both scripts, t_rec is the time to the first admission, death or censoring, and rec_status is 0 for censored, 1 for admission and 2 for death before admission. The Stata excerpt checks the hand-built Aalen-Johansen sum against the user-written command stcompet [6], in which compet1(2) names code 2 as the competing event. Neither official Stata nor stcompet offers Gray's test, so the P value comes from cuminc in the R package cmprsk.

Stata: the Aalen-Johansen sum, checked against stcompet

Stata code w3_sim.do (lines 160-167 of 637)
* cross-check the hand calculation against stcompet when it is installed
capture which stcompet
if _rc == 0 {
    stcompet ci_pkg = ci, compet1(2) by(arm)
    generate double ci_gap = abs(ci_pkg - aj_cif) if rec_status == 1
    summarize ci_gap, meanonly
    display "stcompet versus hand Aalen-Johansen, largest absolute gap: " %12.10f r(max)
}
Output of the run w3_sim.log
. * cross-check the hand calculation against stcompet when it is installed
. capture which stcompet

. if _rc == 0 {
.     stcompet ci_pkg = ci, compet1(2) by(arm)
.     generate double ci_gap = abs(ci_pkg - aj_cif) if rec_status == 1
(817 missing values generated)
.     summarize ci_gap, meanonly
.     display "stcompet versus hand Aalen-Johansen, largest absolute gap: " %12.10f r(max)
stcompet versus hand Aalen-Johansen, largest absolute gap: 0.0000000000
. }
Simulated data. An excerpt of the simulation script, with the run's output beside it; the simulated trial file itself is not published. Earlier lines of the script, not shown, set admission as the event with stset and build the Aalen-Johansen sum by hand in the variable aj_cif, as in the twelve-patient table. The excerpt checks that sum against stcompet, and the output shows the largest gap between the two: zero to ten decimals.

R: Gray's test and the cumulative incidences with cmprsk

R code w3_sim_r.R (lines 95-99 of 384)
  # Aalen-Johansen cumulative incidence and Gray's test (cmprsk::cuminc)
  ci <- cuminc(d$t_rec, d$rec_status, group = d$arm, cencode = 0)
  print(ci$Tests)
  tp <- timepoints(ci, c(12, 24))$est
  print(tp)
Output of the run w3_sim_r.log
        stat        pv df
1 5.98748833 0.0144077  1
2 0.06172269 0.8037936  1
           12        24
0 1 0.3534148 0.4606164
1 1 0.2697090 0.4116034
0 2 0.2287071 0.3330379
1 2 0.2093512 0.3104714
Simulated data. An excerpt of the simulation script, with the run's output beside it; the simulated trial file itself is not published. The lines sit inside a function that the script runs on the simulated trial, passed in as d. The first table is Gray's test, row 1 for admission and row 2 for death before admission: stat is the chi-square statistic, pv the P value and df the degrees of freedom. In the second, each row label gives the arm (0 control, 1 treated) and the cause (1 admission, 2 death before admission), and the columns are 12 and 24 months.

Common misreadings and their fixes

  • "Kaplan-Meier is wrong under competing risks."

    Its arithmetic is sound; it answers a different question.

    Fix: One minus Kaplan-Meier estimates the risk in a hypothetical world without the competing event, assuming its removal would not change the hazard of recurrence. That is a legitimate but hypothetical quantity; for the risk patients actually face, use the cumulative incidence.

  • "A death is just another censored patient."

    After censoring the event could still happen; after death it cannot.

    Fix: Code death before the event as a competing event whenever the question is the risk patients face.

  • "1 minus recurrence-free survival is the risk of recurrence."

    If the curve's event is recurrence or death, whichever comes first, 1 minus it is $\mathrm{CIF}_R(t) + \mathrm{CIF}_D(t)$; if deaths were censored, it is the hypothetical quantity above.

    Fix: Name the event definition before reading the curve.

  • "A lower cumulative incidence of recurrence shows that the treatment acts on recurrence."

    A treatment that raises the death rate, with no effect on the hazard of recurrence, still lowers this cumulative incidence, because fewer patients stay alive to recur.

    Fix: Report the cumulative incidence of every cause, and use the cause-specific hazards to ask how the treatment acts.

  • "A subdistribution hazard ratio of 0.70 means a 30% lower risk."

    The Fine-Gray model compares subdistribution hazards, in which patients who had the competing event stay in the risk set; a ratio of hazards is not a ratio of risks.

    Fix: A subdistribution hazard ratio of 0.70 means a lower subdistribution hazard in the treated arm, which, when the subdistribution hazards are proportional, implies a lower cumulative incidence of the event at every time. It does not mean a 30% lower risk. Read the absolute difference in risk at a stated time from the cumulative incidence curves.

What to do in your own analysis

Glossary

competing event (เหตุการณ์แข่งขัน)
An event, such as death before recurrence, that makes the event of interest impossible afterwards.
all-cause survival
Here all-cause event-free survival, $S(t)$: the probability of no first event of any cause by time $t$.
cumulative incidence function (อุบัติการณ์สะสม)
$\mathrm{CIF}_k(t)$: the probability that the first event is of cause $k$ and occurs by time $t$.
Aalen-Johansen estimator (ตัวประมาณ Aalen-Johansen)
The estimator of the cumulative incidence that adds, at each event time, $\hat S$ just before that time multiplied by the fraction of the risk set with an event of that cause.
Gray's test (การทดสอบของ Gray)
A test of whether the cumulative incidence curves of one cause are equal across groups.

References

  1. Putter H, Fiocco M, Geskus RB. Tutorial in biostatistics: competing risks and multi-state models. Stat Med. 2007;26(11):2389-2430. doi:10.1002/sim.2712 https://doi.org/10.1002/sim.2712
  2. Austin PC, Lee DS, Fine JP. Introduction to the analysis of survival data in the presence of competing risks. Circulation. 2016;133(6):601-609. doi:10.1161/CIRCULATIONAHA.115.017719 https://doi.org/10.1161/CIRCULATIONAHA.115.017719
  3. Andersen PK, Geskus RB, de Witte T, Putter H. Competing risks in epidemiology: possibilities and pitfalls. Int J Epidemiol. 2012;41(3):861-870. doi:10.1093/ije/dyr213 https://doi.org/10.1093/ije/dyr213
  4. Wolbers M, Koller MT, Stel VS, Schaer B, Jager KJ, Leffondré K, Heinze G. Competing risks analyses: objectives and approaches. Eur Heart J. 2014;35(42):2936-2941. doi:10.1093/eurheartj/ehu131 https://doi.org/10.1093/eurheartj/ehu131
  5. Gray RJ. A class of K-sample tests for comparing the cumulative incidence of a competing risk. Ann Stat. 1988;16(3):1141-1154. doi:10.1214/aos/1176350951 https://doi.org/10.1214/aos/1176350951
  6. Coviello V, Boggess M. Cumulative incidence estimation in the presence of competing risks. Stata J. 2004;4(2):103-112. doi:10.1177/1536867X0400400201 https://doi.org/10.1177/1536867X0400400201

Key takeaways

  • A death before recurrence is a competing event, not censoring, because it makes recurrence impossible rather than unobserved.
  • One minus Kaplan-Meier with deaths censored estimates a hypothetical risk in a world without death, under an assumption the data cannot test, and it is never lower than the cumulative incidence.
  • The Aalen-Johansen estimator adds, at each event time, the share still event-free times the fraction of the risk set with an event of that cause.
  • Event-free survival and the cumulative incidences of all causes add to one, which gives every table a built-in check.
  • Gray's test weighs the evidence that cumulative incidence curves differ, not by how much or why, so report the cumulative incidence of every cause at named times.

Related in the wiki: [[time-to-event-survival-analysis]]

0
Message for International and Thai ReadersUnderstanding My Medical Context in ThailandRead more →Message for International and Thai ReadersUnderstanding My Broader Content Beyond MedicineRead more →

Comments

No comments yet. Be the first to share your thoughts.

Sign in to comment