Why 1 Minus Kaplan-Meier Overstates Risk When Patients Can Die First

On this page
อ่านฉบับภาษาไทย (Thai version)
Abstract
When patients can die before the event of interest, death is not censoring: it makes the event impossible rather than unobserved. One minus Kaplan-Meier (the usual survival curve flipped into a risk, with deaths treated as lost follow-up) estimates a hypothetical risk without death, under an untestable assumption, and is never lower than the actual risk. The actual risk is the cumulative incidence function, estimated by the Aalen-Johansen estimator, which weights each event fraction by the share still event-free. In a hand example, 20 of 100 patients die before any recurrence: 1 minus Kaplan-Meier gives 12.5% and the cumulative incidence 10%. A second hand example follows twelve cancer patients for up to 12 months: the cumulative incidence of recurrence is 0.380, against 0.489 from 1 minus Kaplan-Meier. In a simulated heart-failure trial of 1,500 patients, 1 minus Kaplan-Meier overstates the first-admission risk in both arms. This article shows how to compute and check the cumulative incidence, compare groups with Gray's test and report both causes.
A recurrence risk that looks too high
An oncology team sees older patients a few weeks after cancer surgery. At that visit it quotes each of them a risk that the cancer returns within five years. The figure comes from the team's own follow-up records, analysed as one minus a Kaplan-Meier curve for recurrence. Patients who died of other causes were treated as censored, as if their records had simply stopped.
A patient's daughter asks why the quoted risk is higher than the share of past patients who had a recurrence within five years. Many of them, she points out, died of heart disease or pneumonia first. She has a point: a patient who has died cannot have a recurrence later, so a death is not a lost record.
Censoring and competing events are different
Censoring means that follow-up ended without the event while the event could still happen later, for example because the patient moved away or the study closed. A competing event makes the event of interest impossible afterwards, such as death before recurrence. Data with competing events carry competing risks, and Part 1 of this series maps where they sit among time-to-event questions.
After censoring, recurrence is possible but unseen. After death, it will never happen. Treating a death as censoring therefore pictures a future the patient cannot have [1, 2].
Let $T$ be the time from surgery to the first event of either kind. The cumulative incidence function of cause $k$, written $\mathrm{CIF}_k(t)$, is the probability that the first event is of cause $k$ and happens by time $t$. Here the causes are recurrence (R) and death before recurrence (D), giving $\mathrm{CIF}_R(t)$ and $\mathrm{CIF}_D(t)$. All-cause event-free survival, $S(t) = P(T > t)$, is the probability of being alive and recurrence-free at time $t$.
Hand example: one hundred patients, deaths first
Hand example. One hundred patients are followed after surgery, and nobody is lost to follow-up. Early on, 20 of them die of other causes before anyone has had a recurrence. Later, 10 of the 80 survivors have a recurrence, and the other 70 reach the end of follow-up alive and recurrence-free.
-
Cumulative incidence
\[ \mathrm{CIF}_R = \frac{10}{100} = 10\% \]
With nobody censored before the last event, this is the share of the starting group who had a recurrence.
-
1 minus Kaplan-Meier, deaths censored
\[ 1 - \mathrm{KM} = \frac{10}{80} = 12.5\% \]
KM is the Kaplan-Meier estimate of staying recurrence-free. It drops the 20 deaths from the risk set, the patients still followed and event-free, as if they were lost records, leaving 80 at risk.
-
The gap
\[ 12.5\% - 10\% = 2.5 \text{ percentage points} \]
The 20 deaths have been treated as patients who could still recur.
Result: 1 minus Kaplan-Meier gives 12.5% and the cumulative incidence 10%, a contrast that rests on all 20 deaths coming before the first recurrence.
The order of events matters for 1 minus Kaplan-Meier only. If some recurrences came before some deaths, it would fall between 10% and 12.5%, and if all 10 came first it would equal 10%. The cumulative incidence stays at 10% in every order, because with nobody censored before the last event it is the plain proportion who recurred.
The Aalen-Johansen estimator
Real cohorts have censoring, so the cumulative incidence needs an estimator rather than a simple proportion. The Aalen-Johansen estimator builds it one event time at a time [1]:
$$\widehat{\mathrm{CIF}}_k(t) = \sum_{t_j \le t} \hat S(t_{j-1})\,\frac{d_{kj}}{n_j}$$A hat marks an estimate from data. The sum runs over the distinct times $t_j$ up to $t$ at which any first event occurs. At each, $n_j$ is the number still event-free and under follow-up just before $t_j$ (the risk set), and $d_{kj}$ is the number of first events of cause $k$ at $t_j$.
$\hat S(t_{j-1})$ is the all-cause event-free survival just after the previous event time, with $\hat S(t_0) = 1$. It is itself a Kaplan-Meier estimate in which every first event counts, whatever its cause. Each term is therefore the share of the cohort still event-free times the fraction of the risk set with an event of cause $k$. Censored patients simply leave the risk set, on the assumption that their future risks resemble those of patients still followed.
Why the risk of death matters
In continuous time the same idea shows why a competing event changes the risk of recurrence. A hazard is the event rate at a given moment among patients still event-free. The cause-specific hazard of recurrence, $\lambda_R(t)$, counts only recurrences. Then
$$\mathrm{CIF}_R(t) = \int_0^t S(u)\,\lambda_R(u)\,du$$The cumulative incidence adds up the recurrence rate over time, weighted by the share still alive and recurrence-free. Death lowers $S(u)$, so the risk of recurrence depends on the hazard of death as well as on the hazard of recurrence [1, 3]. A treatment that raises the death rate can lower the cumulative incidence of recurrence without changing the hazard of recurrence.
Hand example: twelve patients, step by step
Hand example. Twelve patients are followed for up to 12 months after cancer surgery. In order: death at month 2, recurrence at 3, censored at 4, death at 5, recurrences at 6 and 7, censored at 8, death at 9, recurrence at 10, censored at 11 and two censored at 12. That is 4 recurrences, 3 deaths before recurrence and 5 censored patients.
-
Month 2: death
\[ \widehat{\mathrm{CIF}}_D = 1 \times \frac{1}{12} = 0.083 \]
All 12 are at risk, and $\hat S$ drops to $1 - \tfrac{1}{12} = 0.917$.
-
Month 3: recurrence
\[ \widehat{\mathrm{CIF}}_R = \frac{11}{12} \times \frac{1}{11} = \frac{1}{12} = 0.083 \]
Eleven are at risk, and $\hat S$ drops to 0.833.
-
Month 4: censoring
One patient leaves the risk set, and no estimate changes.
-
Month 5: death
\[ \widehat{\mathrm{CIF}}_D = 0.083 + 0.833 \times \frac{1}{9} = 0.176 \]
Nine are at risk, and $\hat S$ drops to 0.741.
-
Month 6: recurrence
\[ \widehat{\mathrm{CIF}}_R = 0.083 + 0.741 \times \frac{1}{8} = 0.176 \]
Eight are at risk, and $\hat S$ drops to 0.648.
-
Month 7: recurrence
\[ \widehat{\mathrm{CIF}}_R = 0.176 + 0.648 \times \frac{1}{7} = 0.269 \]
Seven are at risk, and $\hat S$ drops to 0.556.
-
Months 8 and 9: censoring, then death
\[ \widehat{\mathrm{CIF}}_D = 0.176 + 0.556 \times \frac{1}{5} = 0.287 \]
Five are at risk at month 9, and $\hat S$ drops to 0.444.
-
Month 10: recurrence
\[ \widehat{\mathrm{CIF}}_R = 0.269 + 0.444 \times \frac{1}{4} = 0.380 \]
Four are at risk, and $\hat S$ drops to 0.333. The censorings at months 11 and 12 change nothing.
Result: At month 12, $\hat S = 0.333$, $\widehat{\mathrm{CIF}}_R = 0.380$ and $\widehat{\mathrm{CIF}}_D = 0.287$.
The Aalen-Johansen table for the twelve patients
| Month | First event | At risk, $n_j$ | $\hat S$ just before | Increment | $\widehat{\mathrm{CIF}}_R$ | $\widehat{\mathrm{CIF}}_D$ | $\hat S$ just after |
|---|---|---|---|---|---|---|---|
| 2 | Death | 12 | 1.000 | 0.083 | 0.000 | 0.083 | 0.917 |
| 3 | Recurrence | 11 | 0.917 | 0.083 | 0.083 | 0.083 | 0.833 |
| 5 | Death | 9 | 0.833 | 0.093 | 0.083 | 0.176 | 0.741 |
| 6 | Recurrence | 8 | 0.741 | 0.093 | 0.176 | 0.176 | 0.648 |
| 7 | Recurrence | 7 | 0.648 | 0.093 | 0.269 | 0.176 | 0.556 |
| 9 | Death | 5 | 0.556 | 0.111 | 0.269 | 0.287 | 0.444 |
| 10 | Recurrence | 4 | 0.444 | 0.111 | 0.380 | 0.287 | 0.333 |
A built-in check: everyone is somewhere
At any time each patient is in exactly one of three states: event-free, recurred first or died first. So the three probabilities add to one:
$$1 = S(t) + \mathrm{CIF}_R(t) + \mathrm{CIF}_D(t)$$Event-free survival and the cumulative incidences of all causes share out the whole cohort. At month 12 in the hand example, $0.333 + 0.380 + 0.287 = 1.000$, and every row of the table passes the same check up to rounding. A table that fails it contains a slip, or a 1 minus Kaplan-Meier value where a cumulative incidence belongs.
1 minus Kaplan-Meier for the twelve patients, deaths censored
| Month | At risk | Factor | Kaplan-Meier, recurrence-free | 1 minus Kaplan-Meier |
|---|---|---|---|---|
| 3 | 11 | 10/11 | 0.909 | 0.091 |
| 6 | 8 | 7/8 | 0.795 | 0.205 |
| 7 | 7 | 6/7 | 0.682 | 0.318 |
| 10 | 4 | 3/4 | 0.511 | 0.489 |
Same risk sets, a different multiplier
At month 12, 1 minus Kaplan-Meier is 0.489, against a cumulative incidence of recurrence of 0.380. It is higher by 0.109, or 1.29 times as high. Written as a sum, 1 minus Kaplan-Meier has the same shape as the Aalen-Johansen estimator:
$$1 - \widehat{\mathrm{KM}}_R(t) = \sum_{t_j \le t} \widehat{\mathrm{KM}}_R(t_{j-1})\,\frac{d_{Rj}}{n_j}$$The only change is the multiplier. Aalen-Johansen multiplies each recurrence fraction $d_{Rj}/n_j$ by $\hat S$, which drops at deaths. Kaplan-Meier multiplies it by its own recurrence-free curve, which does not. So 1 minus Kaplan-Meier is never lower than the cumulative incidence, and it is higher once a death comes before a recurrence [1, 2].
At month 10, for example, Aalen-Johansen adds $0.444 \times \tfrac{1}{4} = 0.111$, while Kaplan-Meier adds $0.682 \times \tfrac{1}{4}$. The three deaths have lowered $\hat S$ but left the Kaplan-Meier curve untouched, so every later recurrence weighs more in 1 minus Kaplan-Meier.
A hypothetical risk, not an arithmetic error
What does 0.489 estimate, then? It is the probability of recurrence by month 12 in a hypothetical world in which death before recurrence could not happen. It also assumes that removing death would not change the hazard of recurrence: the patients who died would have recurred at the same rate as those still followed. The data cannot test that assumption, because nobody is observed after death [1, 3].
So 1 minus Kaplan-Meier is not an arithmetic error. It answers a different, hypothetical question, which may be of interest when that world is the explicit target and its assumption is stated. The risk the daughter asked about is the cumulative incidence, the probability that a patient like her parent actually has a recurrence [2, 4].
Comparing groups: Gray's test
Gray's test asks whether the cumulative incidence curves of one cause are the same across groups, such as two treatment arms, over follow-up [5]. It plays the part that the log-rank test, a test of whether survival curves differ, plays for Kaplan-Meier curves. Unlike a log-rank test with deaths censored, it keeps patients with the competing event in the comparison.
A small P value is evidence that the curves differ somewhere over follow-up; it does not say by how much at a stated time. Nor does it say why: a lower cumulative incidence of recurrence can come from a lower hazard of recurrence or from a higher hazard of death. When curves cross, differences of opposite sign can cancel and the test can miss them.
Regression for competing risks is the subject of Part 5 of this series: the cause-specific Cox model, a regression for the hazard of one cause among patients still event-free, and the Fine-Gray model for the cumulative incidence.
The same contrast in a simulated heart-failure trial
The rest of the article analyses one simulated dataset. A randomised trial enrols 1,500 adults after a heart-failure admission, 750 per arm, and follows them for 24 to 36 months. The outcome is the first heart-failure admission after randomisation, and death before any admission is the competing event. The drug is fictional, and the simulation lets it lower both the admission rate and the rate of cardiovascular death.
The opening scene was about cancer and this trial is about heart failure, but the logic is the same. In both, the event of interest can happen only while the patient is alive. Read 'admission' for 'recurrence', and every formula above applies.
In the data, 683 patients had a first admission (364 control, 319 treated), and 514 died before any admission (258 and 256). Another 303 (128 control, 175 treated) were censored without either event, at the close of the study or on loss to follow-up. Losses in the simulation are unrelated to risk, as both estimators assume. The simulation was built from known rates, so each estimate can be checked against the simulated truth.
First admission by 12 and 24 months: cumulative incidence versus 1 minus Kaplan-Meier
| Quantity (probability) | Control, 12 months | Treated, 12 months | Control, 24 months | Treated, 24 months |
|---|---|---|---|---|
| Cumulative incidence of admission, Aalen-Johansen estimate | 0.3534 | 0.2697 | 0.4606 | 0.4116 |
| Cumulative incidence of admission, simulated truth | 0.3417 | 0.2734 | 0.4536 | 0.3899 |
| 1 minus Kaplan-Meier for admission, deaths censored, estimate | 0.4132 | 0.3031 | 0.5899 | 0.5188 |
| Risk of admission if death were removed, simulated truth | 0.4010 | 0.3124 | 0.5950 | 0.4935 |
| Cumulative incidence of death before admission, Aalen-Johansen estimate | 0.2287 | 0.2094 | 0.3330 | 0.3105 |
Reading the table: 1 minus Kaplan-Meier is higher in both arms
By 24 months the cumulative incidence of admission is 46.1% in the control arm and 41.2% in the treated arm. Gray's test compares the two curves over the whole of follow-up [5]. It gives a chi-square statistic of 5.99 on 1 degree of freedom, P = 0.014. It cannot say whether the difference comes from the drug's action on admission or on death.
One minus Kaplan-Meier gives 59.0% and 51.9% at 24 months, higher in both arms. The gap is already there at 12 months, 35.3% against 41.3% in the control arm, and it widens as deaths accumulate. Each estimator lands within a few percentage points of the simulated truth for its own quantity.
Kaplan-Meier manages that here only because the simulation makes death independent of a patient's tendency to be admitted. No real trial can check that assumption. Aalen-Johansen needs no assumption about death, only that censoring is unrelated to prognosis.
The last row is easy to misread. Death before admission is not all deaths, and its cumulative incidence depends on both hazards. The drug lowers admissions, so more treated patients stay event-free and at risk of dying first, which offsets their lower death rate. That is why the two arms end up close, although the simulation lets the drug lower cardiovascular death.
The code: a hand-built Aalen-Johansen sum checked in Stata, Gray's test in R
In both scripts, t_rec is the time to the first admission, death or censoring, and rec_status is 0 for censored, 1 for admission and 2 for death before admission. The Stata excerpt checks the hand-built Aalen-Johansen sum against the user-written command stcompet [6], in which compet1(2) names code 2 as the competing event. Neither official Stata nor stcompet offers Gray's test, so the P value comes from cuminc in the R package cmprsk.
Stata: the Aalen-Johansen sum, checked against stcompet
* cross-check the hand calculation against stcompet when it is installed
capture which stcompet
if _rc == 0 {
stcompet ci_pkg = ci, compet1(2) by(arm)
generate double ci_gap = abs(ci_pkg - aj_cif) if rec_status == 1
summarize ci_gap, meanonly
display "stcompet versus hand Aalen-Johansen, largest absolute gap: " %12.10f r(max)
}
. * cross-check the hand calculation against stcompet when it is installed
. capture which stcompet
. if _rc == 0 {
. stcompet ci_pkg = ci, compet1(2) by(arm)
. generate double ci_gap = abs(ci_pkg - aj_cif) if rec_status == 1
(817 missing values generated)
. summarize ci_gap, meanonly
. display "stcompet versus hand Aalen-Johansen, largest absolute gap: " %12.10f r(max)
stcompet versus hand Aalen-Johansen, largest absolute gap: 0.0000000000
. }
R: Gray's test and the cumulative incidences with cmprsk
# Aalen-Johansen cumulative incidence and Gray's test (cmprsk::cuminc)
ci <- cuminc(d$t_rec, d$rec_status, group = d$arm, cencode = 0)
print(ci$Tests)
tp <- timepoints(ci, c(12, 24))$est
print(tp)
stat pv df
1 5.98748833 0.0144077 1
2 0.06172269 0.8037936 1
12 24
0 1 0.3534148 0.4606164
1 1 0.2697090 0.4116034
0 2 0.2287071 0.3330379
1 2 0.2093512 0.3104714
Common misreadings and their fixes
-
"Kaplan-Meier is wrong under competing risks."
Its arithmetic is sound; it answers a different question.
Fix: One minus Kaplan-Meier estimates the risk in a hypothetical world without the competing event, assuming its removal would not change the hazard of recurrence. That is a legitimate but hypothetical quantity; for the risk patients actually face, use the cumulative incidence.
-
"A death is just another censored patient."
After censoring the event could still happen; after death it cannot.
Fix: Code death before the event as a competing event whenever the question is the risk patients face.
-
"1 minus recurrence-free survival is the risk of recurrence."
If the curve's event is recurrence or death, whichever comes first, 1 minus it is $\mathrm{CIF}_R(t) + \mathrm{CIF}_D(t)$; if deaths were censored, it is the hypothetical quantity above.
Fix: Name the event definition before reading the curve.
-
"A lower cumulative incidence of recurrence shows that the treatment acts on recurrence."
A treatment that raises the death rate, with no effect on the hazard of recurrence, still lowers this cumulative incidence, because fewer patients stay alive to recur.
Fix: Report the cumulative incidence of every cause, and use the cause-specific hazards to ask how the treatment acts.
-
"A subdistribution hazard ratio of 0.70 means a 30% lower risk."
The Fine-Gray model compares subdistribution hazards, in which patients who had the competing event stay in the risk set; a ratio of hazards is not a ratio of risks.
Fix: A subdistribution hazard ratio of 0.70 means a lower subdistribution hazard in the treated arm, which, when the subdistribution hazards are proportional, implies a lower cumulative incidence of the event at every time. It does not mean a 30% lower risk. Read the absolute difference in risk at a stated time from the cumulative incidence curves.
What to do in your own analysis
- List every way follow-up can end, and sort each into censoring (the event could still happen) or a competing event (it cannot).
- When the question is the risk patients face, report the cumulative incidence at prespecified times with confidence intervals, not 1 minus Kaplan-Meier.
- Plot the cumulative incidence of every cause, so that a lower curve for one cause is read against the others.
- If a risk without the competing event is genuinely the target, say so and state its untestable assumption.
- Check that event-free survival and the cumulative incidences add to one, and consider giving Gray's test only beside the absolute difference at a named time.
Glossary
- competing event (เหตุการณ์แข่งขัน)
- An event, such as death before recurrence, that makes the event of interest impossible afterwards.
- all-cause survival
- Here all-cause event-free survival, $S(t)$: the probability of no first event of any cause by time $t$.
- cumulative incidence function (อุบัติการณ์สะสม)
- $\mathrm{CIF}_k(t)$: the probability that the first event is of cause $k$ and occurs by time $t$.
- Aalen-Johansen estimator (ตัวประมาณ Aalen-Johansen)
- The estimator of the cumulative incidence that adds, at each event time, $\hat S$ just before that time multiplied by the fraction of the risk set with an event of that cause.
- Gray's test (การทดสอบของ Gray)
- A test of whether the cumulative incidence curves of one cause are equal across groups.
References
- Putter H, Fiocco M, Geskus RB. Tutorial in biostatistics: competing risks and multi-state models. Stat Med. 2007;26(11):2389-2430. doi:10.1002/sim.2712 https://doi.org/10.1002/sim.2712
- Austin PC, Lee DS, Fine JP. Introduction to the analysis of survival data in the presence of competing risks. Circulation. 2016;133(6):601-609. doi:10.1161/CIRCULATIONAHA.115.017719 https://doi.org/10.1161/CIRCULATIONAHA.115.017719
- Andersen PK, Geskus RB, de Witte T, Putter H. Competing risks in epidemiology: possibilities and pitfalls. Int J Epidemiol. 2012;41(3):861-870. doi:10.1093/ije/dyr213 https://doi.org/10.1093/ije/dyr213
- Wolbers M, Koller MT, Stel VS, Schaer B, Jager KJ, Leffondré K, Heinze G. Competing risks analyses: objectives and approaches. Eur Heart J. 2014;35(42):2936-2941. doi:10.1093/eurheartj/ehu131 https://doi.org/10.1093/eurheartj/ehu131
- Gray RJ. A class of K-sample tests for comparing the cumulative incidence of a competing risk. Ann Stat. 1988;16(3):1141-1154. doi:10.1214/aos/1176350951 https://doi.org/10.1214/aos/1176350951
- Coviello V, Boggess M. Cumulative incidence estimation in the presence of competing risks. Stata J. 2004;4(2):103-112. doi:10.1177/1536867X0400400201 https://doi.org/10.1177/1536867X0400400201
Key takeaways
- A death before recurrence is a competing event, not censoring, because it makes recurrence impossible rather than unobserved.
- One minus Kaplan-Meier with deaths censored estimates a hypothetical risk in a world without death, under an assumption the data cannot test, and it is never lower than the cumulative incidence.
- The Aalen-Johansen estimator adds, at each event time, the share still event-free times the fraction of the risk set with an event of that cause.
- Event-free survival and the cumulative incidences of all causes add to one, which gives every table a built-in check.
- Gray's test weighs the evidence that cumulative incidence curves differ, not by how much or why, so report the cumulative incidence of every cause at named times.
Related in the wiki: [[time-to-event-survival-analysis]]