Recurrent Events: Counting Every Admission Without Pretending They Are Independent

On this page
อ่านฉบับภาษาไทย (Thai version)
Abstract
Patients with heart failure are admitted again and again, yet many trials analyse only the time to the first admission. Recurrent-event models keep every admission, using a start-stop layout with one row per interval at risk. The Andersen-Gill model fits one Cox model to all admissions on a single clock from randomisation, with a standard error clustered on patient because one patient's admissions are correlated. The Prentice-Williams-Peterson model stratifies by admission number, with time measured from randomisation (total time) or from the previous admission (gap time). In a simulated trial of 1,500 patients with 1,313 admissions, the Andersen-Gill hazard ratio, comparing admission rates between arms, was 0.73 (95% CI 0.63 to 0.85). Clustering raised the standard error of its logarithm from 0.0555 to 0.0738 without moving the estimate. Death ends the count, so treating it as ordinary censoring (follow-up ending without the event) changes the question. This article shows how to build the rows, choose the clock and report deaths alongside admissions.
A trial that counts only the first admission
A steering committee is reviewing the analysis plan of a randomised heart-failure trial. The trial enrols 1,500 adults after a heart-failure admission and follows them for 24 to 36 months. The drug is fictional and the data are simulated.
The primary outcome is the time to the first heart-failure admission. Over follow-up, some patients are admitted again and again, a few of them four or more times, and most are never admitted. A cardiologist on the committee asks how to use every admission without treating one patient's repeated admissions as if they came from different people.
Admissions are recurrent events: events that can happen more than once in the same patient. The answer has three parts. Lay the data out with one row per interval at risk, and count every admission with a standard error that respects the patient. Then report deaths alongside, because a patient who has died can no longer be admitted.
What a first-event analysis leaves out
Take four patients followed for up to 24 months, a hand example rather than trial data. R1 was admitted at months 3 and 8 and was alive and in follow-up at month 24. R2 was admitted at month 5 and died at month 12. R3 was never admitted and was followed to month 24.
R4 was admitted at months 2, 4 and 10 and was lost to follow-up at month 18. The four patients had 6 admissions in all. An analysis of the time to the first event keeps one admission per patient: 3 of the 6.
Censoring means that follow-up ends without the event, so the patient is known to be event-free only up to that time. In a first-event analysis, R1, R2 and R4 leave at their first admission, and R3 is censored at month 24. R1's second admission and R4's second and third never enter the analysis. A patient admitted three times looks the same as one admitted once.
One row per interval at risk
Recurrent-event models take a counting process view: each patient carries a running count of admissions that steps up by one at every admission. The data go into a start-stop layout, with one row for each interval in which the patient is at risk of the next admission.
Each row has five entries. start and stop are the months since randomisation at which the interval opens and closes. event is 1 if the interval ends in an admission and 0 if it ends in censoring or death. enum is the number of the admission the patient is at risk for, and gap is stop minus start.
A patient with $k$ admissions therefore has $k + 1$ rows, the last ending at death or censoring. In the simulated trial analysed below, 1,500 patients with 1,313 admissions give 2,813 rows.
Hand example: four patients, ten rows
Hand example, times in months since randomisation. Each interval is written (start, stop]: it excludes its start and includes its stop. An admission closes a row with event = 1, and the end of follow-up closes the last row with event = 0.
-
R1: admitted at 3 and 8, censored at 24
\[ (0, 3],\ (3, 8],\ (8, 24] \]
Three rows: two end in admissions, the third at censoring. Gap times 3, 5 and 16.
-
R2: admitted at 5, died at 12
\[ (0, 5],\ (5, 12] \]
Two rows. The second ends at death and is coded event = 0, like censoring. Gap times 5 and 7.
-
R3: never admitted, censored at 24
\[ (0, 24] \]
One row, event = 0, gap time 24.
-
R4: admitted at 2, 4 and 10, lost at 18
\[ (0, 2],\ (2, 4],\ (4, 10],\ (10, 18] \]
Four rows. Gap times 2, 2, 6 and 8.
-
Rows
\[ 3 + 2 + 1 + 4 = 10 \]
One row per interval at risk.
-
Admissions
\[ 2 + 1 + 0 + 3 = 6 \]
A first-event analysis keeps 3 of them.
-
Crude admission rate
\[ \frac{6}{24 + 12 + 24 + 18} = \frac{6}{78} = 0.077 \]
Admissions per patient-month of follow-up, counting every admission.
Result: Four patients give 10 rows, 6 admissions and 78 patient-months of follow-up, a crude rate of 0.077 admissions per patient-month.
The ten start-stop rows of the hand example
| Patient | start (months) | stop (months) | event | enum | gap (months) |
|---|---|---|---|---|---|
| R1 | 0 | 3 | 1 | 1 | 3 |
| R1 | 3 | 8 | 1 | 2 | 5 |
| R1 | 8 | 24 | 0 | 3 | 16 |
| R2 | 0 | 5 | 1 | 1 | 5 |
| R2 | 5 | 12 | 0 | 2 | 7 |
| R3 | 0 | 24 | 0 | 1 | 24 |
| R4 | 0 | 2 | 1 | 1 | 2 |
| R4 | 2 | 4 | 1 | 2 | 2 |
| R4 | 4 | 10 | 1 | 3 | 6 |
| R4 | 10 | 18 | 0 | 4 | 8 |
Andersen-Gill: every admission on one clock
The Andersen-Gill model fits a Cox model to the start-stop rows, so every admission counts toward one hazard ratio (HR) [1].
$$h_i(t) = h_0(t)\,\exp(\beta A_i)$$Here $h_i(t)$ is the admission rate at month $t$ since randomisation for a patient $i$ who is still at risk, and $A_i$ is 1 for treatment and 0 for control. $h_0(t)$ is the control-arm rate, left unspecified as in any Cox model, $\beta$ is the log hazard ratio and $\exp(\beta)$ is the hazard ratio. One $h_0(t)$ serves the first, second and later admissions alike, so admissions are treated as interchangeable.
Time runs on one clock from randomisation, called total time. At each moment the risk set, the patients who could have the event then, holds every patient alive and in follow-up. When R4 is admitted for the second time at month 4, all four hand-example patients are in it.
Why the standard error is clustered on patient
Admissions from one patient are not independent. Some patients are more prone to admission than others and are admitted more often throughout, so one admission predicts more. A model-based, or naive, standard error treats the 1,313 admissions as if they came from 1,313 unrelated people. With positive dependence, as here, it comes out too small.
The robust (sandwich) standard error first adds up everything each patient contributes to the estimate. It then measures how much these patient totals vary, which keeps it valid when admissions cluster within patients. Clustering leaves the hazard ratio itself unchanged, and that estimate can be read as a ratio of average admission rates between the arms among patients still alive and in follow-up. The robust standard error keeps the interval valid under that reading, without assuming that one patient's admissions are independent [2].
A sandwich standard error repairs the standard error when the variance assumption is wrong. It does not repair a wrong mean model, and a biased coefficient keeps its bias.
In the simulated trial the Andersen-Gill hazard ratio is 0.73. The naive standard error of its logarithm, 0.0555, would give a 95% CI of 0.66 to 0.82. Clustering on patient raises the standard error to 0.0738 and widens the interval to 0.63 to 0.85. The estimate does not move; only its uncertainty does.
- Stata:
stset stop, id(id) failure(event == 1) time0(start) exit(time .), thenstcox arm, vce(cluster id) - R:
coxph(Surv(start, stop, event) ~ arm + cluster(id), data = rl, ties = "breslow"), whererlholds the start-stop rows
Prentice-Williams-Peterson: admissions in order
The Prentice-Williams-Peterson (PWP) model respects the order of admissions [3]. It stratifies the Cox model by admission number, so each admission number gets its own baseline rate. A patient enters the risk set for admission $k$ only after admission $k - 1$.
$$h_{ik}(t) = h_{0k}(t)\,\exp(\beta A_i)$$Here $k$ is the admission number, which defines the stratum, and $h_{0k}(t)$ is the control-arm rate for admission $k$ at time $t$ on the chosen clock. The hazard ratio $\exp(\beta)$ is shared across strata unless one is fitted per stratum. It compares the arms among patients with the same number of previous admissions.
PWP runs on either of two clocks. In total time (PWP-TT), $t$ counts months since randomisation and each row keeps its start and stop. In gap time (PWP-GT), $t$ counts months since the previous admission: the clock restarts at 0 after each admission, so each row runs from 0 to its gap.
In the hand example the four strata hold 4, 3, 2 and 1 rows. At month 4 in total time, only R1 and R4 are at risk of a second admission: R2 has not yet had its first, and R3 never has one. In gap time, R4's second admission comes 2 months after its first. The risk set is then every patient's interval at risk for a second admission that lasted at least 2 months: R1 (5 months), R2 (7) and R4 (2).
Total time suits a question about time since randomisation, for example when a treatment effect may build up or fade during the trial. Gap time suits a process that starts afresh after each admission, where the question is how long until the next one. The clock is best chosen in the analysis plan, before the data are seen.
Later admissions, selected patients
Later strata are small and selected. Few patients reach them, so the simulated analysis pools the fourth and later admissions into one stratum. Patients who reach a later admission also tend to be more prone to admission, and within that stratum the arms are no longer the randomised groups.
PWP-TT gives HR 0.78 (95% CI 0.69 to 0.87) and PWP-GT gives 0.77 (0.69 to 0.87), both with the clustered standard error. In a simulated trial of 100,000 patients both settle near 0.76, not 0.70. Treated patients who reach a second admission are, on average, more prone to admission than control patients who do, and that pulls the comparison toward 1. Clustering changes the PWP standard error far less than the Andersen-Gill one, probably because conditioning on the number of previous admissions already absorbs much of the dependence.
- Stata:
stcox arm, strata(enum_c) vce(cluster id), after the start-stopstsetfor total time and afterstset gap, failure(event == 1)for gap time - R:
coxph(Surv(start, stop, event) ~ arm + strata(enum_c) + cluster(id), data = rl, ties = "breslow"), withSurv(gap, event)for gap time;enum_cis the admission number with four and above pooled
Two other models: Wei-Lin-Weissfeld and shared frailty
The Wei-Lin-Weissfeld (WLW) model treats the first, second and third admissions as separate outcomes, each timed from randomisation [4]. Every patient counts as at risk for every admission number from the start, even for a third admission before a second. Because the risk set for a later admission does not wait for the earlier ones, its hazard ratio also reflects the effect on those earlier admissions, so WLW suits several distinct event types better than ordered admissions of one kind.
A shared frailty model gives each patient a random multiplier, the frailty, that scales all of that patient's admission rates. Its hazard ratio compares a treated and an untreated patient with the same frailty. The simulated trial was generated this way, with a frailty of mean one and variance 0.6 drawn from a gamma distribution (a distribution of positive values) and a rate ratio of 0.70.
Because the frailty is independent of arm, death and censoring, patients still in follow-up have the same average frailty in both arms at every month. The patient-level and population-level rate ratios are therefore both 0.70 here. They can differ when death depends on frailty, or when the risk set is limited to patients with a given history, as in a first-event or PWP analysis.
Death stops the count
Death is a terminal event: once a patient dies, no further admission can happen. Andersen-Gill and PWP code the last row of a patient who died as event = 0, like censoring, and the count models below simply end that patient's follow-up at death. Read as censoring, that assumes the patients who died would have been admitted at the same rate as similar patients still alive. In heart failure that is rarely credible, because the patients who die are often the sickest, and the sickest are admitted most.
Death therefore cannot be treated as independent censoring. Coding it as event = 0 is a convention of the layout that fixes the question: the hazard ratio describes the admission rate among patients still alive. A drug that left that rate unchanged but shortened life would still lower the admissions per randomised patient, because its patients would spend less time alive.
In the simulated trial the drug lowers the hazard of cardiovascular death (hazard ratio 0.70), so treated patients live longer and have more time to be admitted. With 100,000 simulated patients, the mean number of admissions per patient is 0.93 with control and 0.71 with treatment. That is a smaller relative difference than the rate ratio of 0.70. Death was generated independently of the frailty, which is why the Andersen-Gill value with 100,000 simulated patients is 0.70; real data give no such guarantee.
Three approaches keep death in view. A composite outcome counts cardiovascular deaths as events alongside admissions, so each such death counts against the drug, although one death can still remove several later admissions. The mean frequency function estimates the expected number of admissions per patient by each month, with patients who die adding no more admissions, so it is read beside the survival curves for death [5]. Joint models, which link admissions and death through a shared frailty, go further and are beyond this article.
For the first admission, death before admission is a competing event, one that stops the event of interest from ever happening. One minus the Kaplan-Meier curve then estimates the proportion who would be admitted if no one could die first, assuming that removing death would not change the admission rate: a hypothetical quantity. The cumulative incidence function gives the proportion actually admitted by a given month when death can come first; Part 4 of this series covers both.
Counting instead of timing: negative binomial regression
A simpler route ignores timing and counts each patient's admissions. Poisson regression models the count with follow-up time as the exposure, so it compares admissions per month of follow-up. The logarithm of follow-up enters as an offset, a term whose coefficient is fixed at one. Poisson regression assumes that the variance of a count equals its mean.
Frailty breaks that assumption, because a few patients have many admissions and most have none; this extra variation is called overdispersion. Negative binomial regression adds a patient-level multiplier with variance $\alpha$ that absorbs it. In the simulated trial $\hat\alpha$ is 0.67, where the hat marks an estimate, close to the generator's frailty variance of 0.6. The negative binomial rate ratio is 0.72 (95% CI 0.62 to 0.83).
The Poisson rate ratio is 0.73, and its naive and robust standard errors, 0.0555 and 0.0738, equal those of Andersen-Gill to four decimals. With one row per patient, the robust standard error of a Poisson model already respects the patient. Count models give a rate ratio and a mean count, not the timing of admissions, and they handle death no better than the Cox-type models. A separate article on Poisson regression covers rates in more depth.
- Stata:
nbreg n_adm arm, exposure(fu_months) irr, wheren_admis the count andfu_monthsthe follow-up in months - R:
glm.nb(n_adm ~ arm + offset(log(fu_months)), data = w3)from the MASS package
The simulated trial: every model side by side
The simulated trial has 1,500 patients, 750 per arm, and 1,313 admissions: 727 with control and 586 with treatment. At least one admission occurred in 364 control and 319 treated patients. The mean number of admissions per patient is 0.97 with control and 0.78 with treatment. The generator's admission rate ratio is 0.70.
A Cox model for the first admission only, with deaths before admission censored, uses 683 of the 1,313 admissions. It gives HR 0.77 (95% CI 0.66 to 0.89), and its value in a simulated trial of 100,000 patients is 0.75, not 0.70. This is the cause-specific Cox model of Part 5 of this series.
The population-level first-admission hazard ratio drifts from 0.70 at randomisation to 0.76 after one year and 0.80 after two years, while the patient-level ratio stays at 0.70. A single first-admission hazard ratio therefore averages a ratio that is not constant over follow-up. The frailest control patients are admitted early and leave the risk set faster than the frailest treated patients. So the control patients who remain are less prone to admission than the treated patients who remain.
Every model on the same simulated admissions
| Model | Clock | Ratio (95% CI) | Model-based standard error, log scale | Robust standard error, log scale | Value with 100,000 simulated patients |
|---|---|---|---|---|---|
| Cox, first admission only | Total time | 0.77 (0.66 to 0.89) | not shown | not needed: one row per patient | 0.75 |
| Andersen-Gill | Total time | 0.73 (0.63 to 0.85) | 0.0555 | 0.0738 | 0.70 |
| PWP total time | Total time, by admission number | 0.78 (0.69 to 0.87) | 0.0559 | 0.0581 | 0.76 |
| PWP gap time | Gap time, by admission number | 0.77 (0.69 to 0.87) | 0.0559 | 0.0587 | 0.76 |
| Poisson | Counts per month of follow-up | 0.73 (0.63 to 0.85) | 0.0555 | 0.0738 | not computed |
| Negative binomial | Counts per month of follow-up | 0.72 (0.62 to 0.83) | 0.0729 | not used | 0.71 |
Stata: Andersen-Gill with a standard error clustered on patient
* Andersen-Gill: every admission counts, total time, robust (sandwich) SE clustered on patient
stset stop, id(id) failure(event == 1) time0(start) exit(time .)
stcox arm, vce(cluster id)
. * Andersen-Gill: every admission counts, total time, robust (sandwich) SE clustered on patient
. stset stop, id(id) failure(event == 1) time0(start) exit(time .)
Survival-time data settings
ID variable: id
Failure event: event==1
Observed time interval: (start, stop]
Exit on or before: time .
--------------------------------------------------------------------------
2,813 total observations
0 exclusions
--------------------------------------------------------------------------
2,813 observations remaining, representing
1,500 subjects
1,313 failures in multiple-failure-per-subject data
29,562.808 total analysis time at risk and under observation
At risk from t = 0
Earliest observed entry t = 0
Last observed exit t = 35.94684
. stcox arm, vce(cluster id)
Failure _d: event==1
Analysis time _t: stop
Exit on or before: time .
ID variable: id
Iteration 0: Log pseudolikelihood = -9007.5576
Iteration 1: Log pseudolikelihood = -8991.7416
Iteration 2: Log pseudolikelihood = -8991.7416
Refining estimates:
Iteration 0: Log pseudolikelihood = -8991.7416
Cox regression with no ties
No. of subjects = 1,500 Number of obs = 2,813
No. of failures = 1,313
Time at risk = 29,562.8078
Wald chi2(1) = 17.80
Log pseudolikelihood = -8991.7416 Prob > chi2 = 0.0000
(Std. err. adjusted for 1,500 clusters in id)
------------------------------------------------------------------------------
| Robust
_t | Haz. ratio std. err. z P>|z| [95% conf. interval]
-------------+----------------------------------------------------------------
arm | .7325 .054048 -4.22 0.000 .6338715 .8464749
------------------------------------------------------------------------------
R: the same model, naive and robust standard errors in one table
# Andersen-Gill: every admission counts, total time, robust (sandwich) SE clustered on patient
ag <- coxph(Surv(start, stop, event) ~ arm + cluster(id), data = rl, ties = "breslow")
print(summary(ag))
> ag <- coxph(Surv(start, stop, event) ~ arm + cluster(id),
+ data = rl, ties = "breslow")
> print(summary(ag))
Call:
coxph(formula = Surv(start, stop, event) ~ arm, data = rl, ties = "breslow",
cluster = id)
n= 2813, number of events= 1313
coef exp(coef) se(coef) robust se z Pr(>|z|)
arm -0.31129 0.73250 0.05555 0.07376 -4.22 2.44e-05 ***
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
exp(coef) exp(-coef) lower .95 upper .95
arm 0.7325 1.365 0.6339 0.8464
Concordance= 0.539 (se = 0.009 )
Likelihood ratio test= 31.63 on 1 df, p=2e-08
Wald test = 17.81 on 1 df, p=2e-05
Score (logrank) test = 31.66 on 1 df, p=2e-08, Robust = 17.59 p=3e-05
(Note: the likelihood ratio and score tests assume independence of
observations within a cluster, the Wald and robust score tests do not).
Recurrent-event models compared
Choose by the question: Andersen-Gill for the admission rate among patients alive, PWP for the next admission given the number of earlier admissions, count models for admissions per month of follow-up. The table below lines up the models by clock, risk set and meaning. A published tutorial compares these models on epidemiological data [6], and a review applies them to recurrent heart-failure admissions in a large trial [7].
| Model | Clock | Risk set for the next event | What the ratio compares | Dependence handled by |
|---|---|---|---|---|
| Cox, first event | Total time | Patients alive, in follow-up and not yet admitted | First-admission rates | Nothing needed: one row per patient |
| Andersen-Gill | Total time | Everyone alive and in follow-up | Admission rates among patients alive, all admissions alike | Robust standard error clustered on patient |
| PWP total time | Total time | For admission k, patients with k - 1 admissions | Rates of the next admission, given the number so far | Strata by admission number and a clustered standard error |
| PWP gap time | Time since the last admission | As PWP total time, with the clock reset | Rates of the next admission by time since the last one, given the number so far | Strata by admission number and a clustered standard error |
| Wei-Lin-Weissfeld | Total time, per admission number | Everyone, for every admission number | Each admission number as its own outcome | Robust standard error |
| Shared frailty | Total time or gap time | As Andersen-Gill or PWP | Rates for patients with the same frailty | A random multiplier per patient |
| Negative binomial | None: counts per patient | Not applicable | Admissions per month of follow-up | A gamma multiplier per patient |
Common misreadings and their fixes
-
"Andersen-Gill with a robust standard error handles all dependence and death."
Clustering on patient changes only the standard error. The last row of a patient who died still ends with event = 0, the same code as a patient censored alive.
Fix: The robust standard error corrects the uncertainty for repeated events in one patient. It does not change what the hazard ratio estimates, and it still treats death as if it were ordinary censoring. Cluster on patient for the repeated admissions, then state that the hazard ratio describes admissions among patients still alive and report deaths by arm beside it.
-
"Robust standard errors fix a misspecified model."
In the simulated trial, clustering widened the Andersen-Gill interval and left the hazard ratio at 0.73.
Fix: A sandwich standard error repairs the standard error when the variance assumption is wrong. It does not repair a wrong mean model, and a biased coefficient keeps its bias.
-
"Counting only the first admission is the cautious choice."
Here it uses 683 of the 1,313 admissions, and its hazard ratio drifts toward 1 as the patients most prone to admission leave the risk set.
Fix: Keep a first-event analysis when the first admission is the question, and add a recurrent-event model, with deaths reported beside it, when the question is the burden of admissions.
-
"A PWP hazard ratio is a randomised comparison at every admission number."
Patients who reach a later admission are selected by their history, so within a later stratum the arms are no longer the randomised groups.
Fix: Read a PWP hazard ratio as conditional on the number of previous admissions, and report the rows and admissions in each stratum.
-
"Fewer admissions per randomised patient show that the drug prevents admissions."
A drug that shortens life also shortens the time in which patients can be admitted.
Fix: Report deaths by arm beside every admission result, the mean frequency function included, because it counts admissions per randomised patient and also falls when patients die sooner; consider a composite with death as a further view.
What to do in your own analysis
- Write the estimand, the precise quantity the analysis aims to estimate, before choosing a model: the admission rate among patients alive, the expected number of admissions per patient, or a composite with death.
- Build the start-stop rows from the admission dates, and check a few patients by hand against their records.
- When one patient contributes more than one event to Andersen-Gill, PWP or Wei-Lin-Weissfeld, cluster the standard error on patient; a frailty model instead handles the dependence through its random term.
- If the order of admissions matters, consider prespecifying PWP and its clock, and pooling sparse later strata.
- Consider deciding in advance whether days in hospital count as time at risk, since an inpatient cannot be admitted again.
- Report deaths by arm beside every admission result.
Glossary
- recurrent events (เหตุการณ์ซ้ำ)
- Events, such as admissions, that can occur more than once in the same patient.
- counting process (กระบวนการนับ)
- A view of follow-up in which each patient's event count steps up by one at every event.
- start-stop layout (รูปแบบ start-stop)
- One row per interval at risk, with a start time, a stop time, an event flag and the event number.
- hazard ratio (อัตราส่วนฮาซาร์ด)
- The ratio of event rates between arms among patients still at risk; it is not a ratio of risks.
- Andersen-Gill model (แบบจำลอง Andersen-Gill)
- A Cox model for recurrent events on total time, with one baseline rate for every event, usually reported with a standard error clustered on patient.
- robust (sandwich) standard error (ค่าคลาดเคลื่อนมาตรฐานแบบแซนด์วิช)
- A standard error built from each patient's total contribution, valid when events within a patient are correlated, provided the model for the mean is right.
- Prentice-Williams-Peterson model (แบบจำลอง Prentice-Williams-Peterson)
- A Cox model for recurrent events stratified by event number, so a patient is at risk for event k only after event k - 1.
- total time (เวลารวม)
- Time measured from randomisation for every event.
- gap time (เวลาระหว่างครั้ง)
- Time measured from the previous event, with the clock reset to 0 after each event.
- Wei-Lin-Weissfeld model (แบบจำลอง Wei-Lin-Weissfeld)
- A model that treats each event number as its own outcome, timed from randomisation, and averages over patients rather than conditioning on their history.
- frailty (ตัวคูณสุ่มรายผู้ป่วย)
- A patient-level random multiplier on all of that patient's event rates, capturing shared proneness to events.
- terminal event (เหตุการณ์ปลายทาง)
- An event, such as death, after which no further recurrent events can occur.
- negative binomial regression (การถดถอยทวินามลบ)
- A count model that adds a patient-level multiplier to Poisson regression to allow for overdispersion.
- mean frequency function (ฟังก์ชันความถี่เฉลี่ย)
- The expected number of events per patient by each time, with patients who die adding no further events.
References
- Andersen PK, Gill RD. Cox's regression model for counting processes: a large sample study. Ann Stat. 1982;10(4):1100-1120. doi:10.1214/aos/1176345976 https://doi.org/10.1214/aos/1176345976
- Lin DY, Wei LJ, Yang I, Ying Z. Semiparametric regression for the mean and rate functions of recurrent events. J R Stat Soc Series B Stat Methodol. 2000;62(4):711-730. doi:10.1111/1467-9868.00259 https://doi.org/10.1111/1467-9868.00259
- Prentice RL, Williams BJ, Peterson AV. On the regression analysis of multivariate failure time data. Biometrika. 1981;68(2):373-379. doi:10.1093/biomet/68.2.373 https://doi.org/10.1093/biomet/68.2.373
- Wei LJ, Lin DY, Weissfeld L. Regression analysis of multivariate incomplete failure time data by modeling marginal distributions. J Am Stat Assoc. 1989;84(408):1065-1073. doi:10.1080/01621459.1989.10478873 https://doi.org/10.1080/01621459.1989.10478873
- Ghosh D, Lin DY. Nonparametric analysis of recurrent events and death. Biometrics. 2000;56(2):554-562. doi:10.1111/j.0006-341X.2000.00554.x https://doi.org/10.1111/j.0006-341X.2000.00554.x
- Amorim LDAF, Cai J. Modelling recurrent events: a tutorial for analysis in epidemiology. Int J Epidemiol. 2015;44(1):324-333. doi:10.1093/ije/dyu222 https://doi.org/10.1093/ije/dyu222
- Rogers JK, Pocock SJ, McMurray JJV, Granger CB, Michelson EL, Ostergren J, et al. Analysing recurrent hospitalizations in heart failure: a review of statistical methodology, with application to CHARM-Preserved. Eur J Heart Fail. 2014;16(1):33-40. doi:10.1002/ejhf.29 https://doi.org/10.1002/ejhf.29
Key takeaways
- A first-event analysis keeps one admission per patient and ignores every later one.
- Andersen-Gill counts every admission on one clock, and its hazard ratio needs a standard error clustered on patient.
- Prentice-Williams-Peterson stratifies by admission number on total time or gap time, and its hazard ratio is conditional on admission history.
- A robust standard error changes the uncertainty, not the estimate, and it does not fix how death is handled.
- Death stops the count, so report deaths beside admissions or build them into the outcome.
Related in the wiki: [[time-to-event-survival-analysis]]