A Grammar for Choosing a Model: Five Questions Before Any Command

On this page
อ่านฉบับภาษาไทย (Thai version)
Abstract
A regression model is often chosen by name, but five features of a study narrow the choice before any command is typed. The grammar reads Outcome + Denominator/Time + Dependence + Estimand + Special structure = Analysis. The outcome may be continuous, binary, ordinal, a count or a time to an event. The denominator sets whether events are counted per person at a fixed horizon, per unit of person-time, or as time until the event. Dependence asks whether rows are independent or grouped in clusters, repeated visits or matched sets. The estimand names the contrast, the effect measure and the population. Special structure covers censoring, competing events, overdispersion, excess zeros, non-linearity and effect modification. In a hand example, 24 admissions over 1,200 person-years against 40 over 1,000 give a rate ratio of 0.50; counts alone give 0.60. This article concludes that the grammar narrows the choice to models that can answer the question, while the analyst still chooses covariates and checks assumptions.
A protocol meeting that starts with a model name
A team is drafting the protocol for a study of readmissions in heart failure. Before anyone has said what will be measured, a resident proposes logistic regression. The methodologist asks to settle five things first.
The outcome is the number of admissions per patient, with follow-up that differs between patients. Patients come from several hospitals, the team wants a ratio of admission rates, and follow-up stops at death.
These answers point to a count model with person-time, the total follow-up time of all patients, as its denominator, and an allowance for patients from the same hospital. Death ends a patient's person-time, and a programme that changes survival may change who remains at risk of admission, so deaths are usually reported beside the rate ratio. Logistic regression would answer another question, the odds of any admission by a fixed date.
The grammar in one line
Written as one line, the five requests become a grammar that narrows the choice of model.
One class behind many candidates
Many models that fill the slots are generalized linear models (GLMs) or extensions of them [1]. The GLM class includes linear and logistic regression, and Poisson regression, a model for counts. Each GLM pairs a distribution for the outcome, the family, with a scale on which covariates act, the link. How these two choices set the effect measure and the standard error is the subject of part 2.
Outcome: the kind of measurement
One heart-failure clinic can record five kinds of outcome. NT-proBNP, a blood marker of strain on the heart, is continuous. Admission within a fixed window after discharge is binary, and NYHA (New York Heart Association) class, a grading of breathlessness from I to IV, is ordinal. Admissions counted over follow-up are a count, and the time to the first admission is a time to an event.
- Continuous: linear regression, for a difference in means [2].
- Binary: logistic regression for an odds ratio, or binomial regression on the log scale for a risk ratio, or on the risk scale itself for a risk difference.
- Ordinal: ordinal logistic regression, which models the odds of being above each cut-point between grades, usually assuming one odds ratio shared by every cut-point (proportional odds) [3]; unordered categories call for multinomial logistic regression instead.
- Count: Poisson or negative binomial regression; the latter lets the variance exceed the mean.
- Time to an event: survival methods such as the Cox model, the usual time-to-event regression, whose hazard ratio compares event rates among patients still event-free.
The outcome settles the class of models, not the model; for a binary outcome, the estimand decides among the three effect measures.
Denominator or time: what the events are counted against
A heart-failure service compares a nurse-led follow-up programme with usual care. The programme started first, so its patients have been followed for longer.
Three answers are common. A risk is the proportion of patients with the event by a fixed horizon, such as admission within 7 days of discharge. A rate counts events per unit of person-time and suits counts when follow-up differs, provided the rate is roughly constant over follow-up. A time-to-event outcome uses the time until the event, and keeps patients whose follow-up ends early.
Hand example: two arms followed for different lengths of time
Hand example, with invented numbers. The programme arm had 24 admissions over 1,200 person-years, and the usual-care arm had 40 admissions over 1,000 person-years. As a dataset named hand, it has one row per arm and the columns arm (1 programme, 0 usual care), admissions and pyears (person-years).
-
Rate in the programme arm
\[ \frac{24}{1200} = 0.020 \text{ per person-year} = 20 \text{ per } 1000 \text{ person-years} \]
Admissions divided by person-years give the rate.
-
Rate in the usual-care arm
\[ \frac{40}{1000} = 0.040 \text{ per person-year} = 40 \text{ per } 1000 \text{ person-years} \]
The usual-care rate is twice the programme rate.
-
Rate ratio
\[ \frac{20}{40} = 0.50 \]
The programme arm has half the admission rate.
-
Rate difference
\[ 20 - 40 = -20 \text{ per } 1000 \text{ person-years} \]
The difference is negative because the programme arm has the lower rate.
-
Ratio of raw counts
\[ \frac{24}{40} = 0.60 \]
Counts ignore the extra person-years in the programme arm, so this ratio sits closer to 1.
Result: The rate ratio is 0.50 and the rate difference is -20 per 1,000 person-years. The ratio of raw counts, 0.60, ignores how long each arm was followed.
The offset: a rate inside a count model
Poisson regression handles unequal follow-up through an offset, a term whose coefficient is fixed at 1 [4]. With the log of person-years as the offset, the model is
$$\log \mathrm{E}(\text{admissions}) = \log(\text{person-years}) + \beta_0 + \beta_1\,\text{arm}$$Here $\mathrm{E}(\text{admissions})$ is the expected number of admissions, arm is 1 for the programme and 0 for usual care, and $\beta_0$ and $\beta_1$ are coefficients. Moving the offset to the left side gives the log of the rate, so $\exp(\beta_1)$ is the rate ratio. With arm as the only covariate, the fitted rate in each arm equals its observed rate, so the model returns 0.50 exactly, whether the data hold two summary rows or one row per patient.
Fixing the offset coefficient at 1 assumes a roughly constant rate over follow-up. If admissions cluster soon after discharge, the arm followed longer adds late, low-rate person-time and may look better for that alone. Comparing rates within the same periods since discharge, or a time-to-event method, may then be preferable. An arm that started first also differs in calendar time.
Its Stata and R commands are in the scenes table below, and a separate article covers Poisson regression for rates.
Dependence: which rows belong together
A trial randomises hospitals, not patients, to a new discharge checklist. Patients in one hospital share staff and routines, so their outcomes are more alike than outcomes from different hospitals.
Rows may be independent, or form clusters such as patients within hospitals, repeated measures of one patient over time, or matched sets of cases and controls grouped by design on age and sex.
Treating correlated rows as independent usually makes standard errors too small, especially for comparisons between clusters [2]. A mixed model adds random effects, meaning cluster-level deviations from the average. Generalized estimating equations (GEE) model the population-average outcome and allow for correlation within clusters. A third option keeps the model and uses a cluster-robust, or sandwich, standard error, estimated from the observed variation between clusters; part 4 explains what it repairs and what it cannot.
For repeated visits in a trial, a common choice is the mixed model for repeated measures (MMRM) [5]. It treats visit as a category and gives each visit its own variance and each pair of visits its own correlation. For matched sets, conditional logistic regression compares exposure within each set, so the matching factors drop out. Ordinary logistic regression with one indicator per set gives a biased odds ratio when sets are small, as with pairs.
Estimand: the quantity the study promises
At a trial steering meeting, a member asks for the effect of the programme on admissions. The statistician replies that this is not yet an estimand, because it names no scale, no comparison and no population.
The ICH (International Council for Harmonisation) E9(R1) guideline on clinical trials describes an estimand, the precise quantity a study aims to estimate, by five attributes: treatment, population, variable (the outcome measured), handling of intercurrent events, and population-level summary [6]. Intercurrent events, such as stopping treatment or death, occur after treatment starts and affect whether the outcome can be measured or how it is read.
For model choice, the slot settles a difference or a ratio; risks, rates, odds or hazards; a marginal or conditional contrast; and the population. A marginal contrast compares the whole population under one treatment with the whole population under the other. A conditional contrast compares patients who share the covariate values in the model.
A risk ratio is not read from a logistic coefficient, which is a log odds ratio. For a common outcome the log-binomial model, a binomial regression on the log scale, can fail to converge, meaning the software's step-by-step search cannot settle on coefficients that keep every patient's predicted risk below 1 and stops without a usable answer; part 3 gives a working sequence of alternatives. How a conditional and a marginal odds ratio from the same patients relate is the subject of part 5.
Special structure: features that bend a standard model
A reviewer reads a manuscript that models emergency visits with Poisson regression. The counts vary far more than their mean, more patients had none than a Poisson model predicts, some died and others moved away. Each is a special structure that a standard model does not handle by default.
- Censoring: follow-up ends before the event is seen, so the event is known only not to have happened by then.
- Competing events: events, such as death, that prevent the event of interest. Its probability by a stated time is the cumulative incidence function (CIF), which the Aalen-Johansen method estimates without a model; one minus the Kaplan-Meier estimate of the proportion still event-free, with deaths censored, overstates it (see cumulative incidence with competing risks). Fine-Gray regression (
stcrregin Stata) adds covariates and assumes proportional subdistribution hazards, in which patients who died still count as at risk (see cause-specific versus subdistribution hazards). - Recurrent events: admissions that recur in one patient, analysed as a count over person-time, as above, or with recurrent-event models that use the time of each admission.
- Overdispersion: count variance larger than the mean, which Poisson regression assumes equal; its standard errors are then too small, and negative binomial or overdispersed Poisson models are common [4].
- Excess zeros: more zeros than the count model predicts, where zero-inflated or hurdle models, which give the zeros their own part, may help.
- Non-linearity: an effect of age that bends; restricted cubic splines (smooth joined curves) can show it, while grouping age discards information [3].
- Effect modification: a treatment effect that differs across levels of another variable, modelled as an interaction; its presence depends on the scale, ratio or difference.
Often this slot adds a component rather than a new class: a variance model, a spline or an interaction term. Competing events and excess zeros can change the method itself.
Five research scenes through the grammar
The table parses five research scenes through the five slots.
| Research scene | Outcome | Denominator or time | Dependence | Estimand | Special structure | Model class | Stata command | R function |
|---|---|---|---|---|---|---|---|---|
| Delirium within 7 days of hip-fracture surgery, early versus later surgery | Binary | Per person, fixed horizon | Independent patients | Risk ratio, all operated patients | Common outcome; the fit may not converge (alternatives in part 3) | Log-binomial regression | glm delirium i.early, family(binomial) link(log) eform | glm(delirium ~ early, family = binomial(link = "log"), data = d) |
| Admissions under a nurse-led programme versus usual care | Count | Person-time, differing by patient | Independent patients | Rate ratio, all enrolled patients | Overdispersion to check | Poisson regression with a log person-time offset | poisson admissions i.arm, exposure(pyears) irr | glm(admissions ~ arm + offset(log(pyears)), family = poisson, data = hand) |
| Symptom score at four visits in a randomised trial | Continuous | Fixed visits | Repeated measures | Mean difference at the final visit, all randomised | Missing visits, assumed missing at random given the observed visits | MMRM | mixed score i.arm##i.visit || id:, noconstant residuals(unstructured, t(visit)) reml | mmrm::mmrm(score ~ arm * visit + us(visit | id), data = d) |
| Time to first admission after discharge, death competing | Time to an event | Time to admission, death or censoring | Independent patients | Cumulative incidence by a stated time, by arm | Censoring; death competing | Aalen-Johansen estimate of the cumulative incidence function (Fine-Gray regression if covariates are needed) | stset time, failure(status == 1), then stcompet cif = ci, compet1(2) by(arm) | cmprsk::cuminc(d$time, d$status, group = d$arm) |
| Matched case-control study of hip fracture and a prescribed sedative | Case or control | None: sampling is on the outcome | Sets matched on age and sex | Odds ratio within matched sets | Effects of the matching factors cannot be estimated | Conditional logistic regression | clogit case i.exposed, group(set) or | survival::clogit(case ~ exposed + strata(set), data = d) |
Reading the table
Two rows concern admissions, yet they land in different classes: a count over person-time, and a time to the first admission with death competing. The question, not the variable name, decides the slot.
Where the grammar stops
The grammar narrows the choice to models that can answer the question. It does not choose covariates, which follow from the causal question, often drawn as a DAG, or from the prediction task [3]. Nor does it check assumptions, such as a roughly constant rate or proportional hazards, a hazard ratio constant over follow-up, and it does not replace design thinking about randomisation, measurement and missing data.
Two analysts who fill the slots alike may still fit different models in one class, and both can be defensible when each states its estimand and assumptions.
Common misreadings and their fixes
-
"The grammar picks the model."
Filling the slots rules out models that cannot answer the question, but several candidates remain.
Fix: The grammar narrows the choice to models that can answer the question. The analyst still chooses covariates, checks assumptions and justifies the estimand.
-
"Comparing admission counts is enough."
In the hand example the counts give 0.60, but the rate ratio is 0.50, because the programme arm had 1,200 person-years against 1,000.
Fix: Put person-time in the denominator, as a rate or a log person-time offset, and check that the rate is roughly constant over follow-up.
-
"A binary outcome means logistic regression."
A logistic coefficient is a log odds ratio, and the odds ratio lies further from 1 than the risk ratio, more so for a common outcome.
Fix: Name the effect measure in the estimand, then choose a model that returns it.
-
"Time-to-event data means a Cox model."
A hazard ratio compares event rates among patients still event-free. When death competes, the probability of admission by a stated time comes from the cumulative incidence function.
Fix: Decide whether the estimand is a hazard ratio or a probability by a stated time.
What to do in your own analysis
- Write the five slots into the protocol before naming a model.
- Record each patient's follow-up time whenever it can differ, so a rate stays possible.
- Name the estimand, with its effect measure, contrast and population, and check that the planned model returns it.
- List what makes rows dependent, and state in advance how the model will allow for it.
- For each special structure, state how it will be handled; justify covariates and assumption checks separately.
Glossary
- generalized linear model (แบบจำลองเชิงเส้นวางนัยทั่วไป (generalized linear model, GLM))
- A regression class, including linear, logistic and Poisson regression, defined by a family and a link.
- person-time (เวลาติดตามสะสม (person-time))
- Follow-up time summed over all patients, such as person-years.
- offset (ออฟเซต (offset))
- A term with its coefficient fixed at 1; log person-time as the offset turns a count model into a rate model.
- estimand (ปริมาณเป้าหมายของการประมาณ (estimand))
- The precise quantity a study aims to estimate, with its population, contrast and summary measure.
- mixed model for repeated measures (แบบจำลองผสมสำหรับการวัดซ้ำ (MMRM))
- Visit as a category, treatment-by-visit terms and an unstructured covariance between visits.
- conditional logistic regression (การถดถอยลอจิสติกแบบมีเงื่อนไข (conditional logistic regression))
- Logistic regression that compares exposure within matched sets.
- overdispersion (ความแปรปรวนเกิน (overdispersion))
- Count variance larger than the mean, which Poisson regression assumes equal.
- censoring (การเซ็นเซอร์ (censoring))
- Follow-up that ends before the event is seen.
- competing event (เหตุการณ์แข่งขัน (competing event))
- An event, such as death, that prevents the event of interest.
- cumulative incidence function (อุบัติการณ์สะสม (cumulative incidence function, CIF))
- The probability of the event of interest by a stated time when competing events can occur.
References
- McCullagh P, Nelder JA. Generalized linear models. 2nd ed. London: Chapman and Hall; 1989. doi:10.1007/978-1-4899-3242-6 https://doi.org/10.1007/978-1-4899-3242-6
- Vittinghoff E, Glidden DV, Shiboski SC, McCulloch CE. Regression methods in biostatistics: linear, logistic, survival, and repeated measures models. 2nd ed. New York: Springer; 2012. doi:10.1007/978-1-4614-1353-0 https://doi.org/10.1007/978-1-4614-1353-0
- Harrell FE Jr. Regression modeling strategies: with applications to linear models, logistic and ordinal regression, and survival analysis. 2nd ed. Cham: Springer; 2015. doi:10.1007/978-3-319-19425-7 https://doi.org/10.1007/978-3-319-19425-7
- Gardner W, Mulvey EP, Shaw EC. Regression analyses of counts and rates: Poisson, overdispersed Poisson, and negative binomial models. Psychol Bull. 1995;118(3):392-404. doi:10.1037/0033-2909.118.3.392 https://doi.org/10.1037/0033-2909.118.3.392
- Fitzmaurice GM, Laird NM, Ware JH. Applied longitudinal analysis. 2nd ed. Wiley; 2011. doi:10.1002/9781119513469 https://doi.org/10.1002/9781119513469
- International Council for Harmonisation (ICH). ICH E9(R1): Addendum on estimands and sensitivity analysis in clinical trials to the guideline on statistical principles for clinical trials. ICH Harmonised Guideline, Step 4; 2019. https://database.ich.org/sites/default/files/E9-R1_Step4_Guideline_2019_1203.pdf
Key takeaways
- Five slots, from outcome to special structure, narrow the choice of model before any command is typed.
- When follow-up differs, count events over person-time and check that the rate is roughly constant: in the hand example the rate ratio is 0.50, while raw counts give 0.60.
- Rows that share a hospital, a patient or a matched set are not independent, and the analysis usually has to allow for it.
- The estimand decides which model in a class can deliver the number the study promised.
- The grammar narrows the choice; the analyst still chooses covariates, checks assumptions and justifies the estimand.
Related in the wiki: [[glm-link-family-guide]] [[poisson-regression-rates-lambda]]