A Grammar for Choosing a Model: Five Questions Before Any Command

Clinical Epidemiology ResearchMethodology and Research DesignUniqcret doctor knowledges
A Grammar for Choosing a Model: Five Questions Before Any Command
On this page

อ่านฉบับภาษาไทย (Thai version)

Abstract

A regression model is often chosen by name, but five features of a study narrow the choice before any command is typed. The grammar reads Outcome + Denominator/Time + Dependence + Estimand + Special structure = Analysis. The outcome may be continuous, binary, ordinal, a count or a time to an event. The denominator sets whether events are counted per person at a fixed horizon, per unit of person-time, or as time until the event. Dependence asks whether rows are independent or grouped in clusters, repeated visits or matched sets. The estimand names the contrast, the effect measure and the population. Special structure covers censoring, competing events, overdispersion, excess zeros, non-linearity and effect modification. In a hand example, 24 admissions over 1,200 person-years against 40 over 1,000 give a rate ratio of 0.50; counts alone give 0.60. This article concludes that the grammar narrows the choice to models that can answer the question, while the analyst still chooses covariates and checks assumptions.


Visual summary. Simulated data.

A protocol meeting that starts with a model name

A team is drafting the protocol for a study of readmissions in heart failure. Before anyone has said what will be measured, a resident proposes logistic regression. The methodologist asks to settle five things first.

The outcome is the number of admissions per patient, with follow-up that differs between patients. Patients come from several hospitals, the team wants a ratio of admission rates, and follow-up stops at death.

These answers point to a count model with person-time, the total follow-up time of all patients, as its denominator, and an allowance for patients from the same hospital. Death ends a patient's person-time, and a programme that changes survival may change who remains at risk of admission, so deaths are usually reported beside the rate ratio. Logistic regression would answer another question, the odds of any admission by a fixed date.

The grammar in one line

Written as one line, the five requests become a grammar that narrows the choice of model.

One class behind many candidates

Many models that fill the slots are generalized linear models (GLMs) or extensions of them [1]. The GLM class includes linear and logistic regression, and Poisson regression, a model for counts. Each GLM pairs a distribution for the outcome, the family, with a scale on which covariates act, the link. How these two choices set the effect measure and the standard error is the subject of part 2.

Outcome: the kind of measurement

One heart-failure clinic can record five kinds of outcome. NT-proBNP, a blood marker of strain on the heart, is continuous. Admission within a fixed window after discharge is binary, and NYHA (New York Heart Association) class, a grading of breathlessness from I to IV, is ordinal. Admissions counted over follow-up are a count, and the time to the first admission is a time to an event.

The outcome settles the class of models, not the model; for a binary outcome, the estimand decides among the three effect measures.

Denominator or time: what the events are counted against

A heart-failure service compares a nurse-led follow-up programme with usual care. The programme started first, so its patients have been followed for longer.

Three answers are common. A risk is the proportion of patients with the event by a fixed horizon, such as admission within 7 days of discharge. A rate counts events per unit of person-time and suits counts when follow-up differs, provided the rate is roughly constant over follow-up. A time-to-event outcome uses the time until the event, and keeps patients whose follow-up ends early.

Hand example: two arms followed for different lengths of time

Hand example, with invented numbers. The programme arm had 24 admissions over 1,200 person-years, and the usual-care arm had 40 admissions over 1,000 person-years. As a dataset named hand, it has one row per arm and the columns arm (1 programme, 0 usual care), admissions and pyears (person-years).

  1. Rate in the programme arm

    \[ \frac{24}{1200} = 0.020 \text{ per person-year} = 20 \text{ per } 1000 \text{ person-years} \]

    Admissions divided by person-years give the rate.

  2. Rate in the usual-care arm

    \[ \frac{40}{1000} = 0.040 \text{ per person-year} = 40 \text{ per } 1000 \text{ person-years} \]

    The usual-care rate is twice the programme rate.

  3. Rate ratio

    \[ \frac{20}{40} = 0.50 \]

    The programme arm has half the admission rate.

  4. Rate difference

    \[ 20 - 40 = -20 \text{ per } 1000 \text{ person-years} \]

    The difference is negative because the programme arm has the lower rate.

  5. Ratio of raw counts

    \[ \frac{24}{40} = 0.60 \]

    Counts ignore the extra person-years in the programme arm, so this ratio sits closer to 1.

Result: The rate ratio is 0.50 and the rate difference is -20 per 1,000 person-years. The ratio of raw counts, 0.60, ignores how long each arm was followed.

The offset: a rate inside a count model

Poisson regression handles unequal follow-up through an offset, a term whose coefficient is fixed at 1 [4]. With the log of person-years as the offset, the model is

$$\log \mathrm{E}(\text{admissions}) = \log(\text{person-years}) + \beta_0 + \beta_1\,\text{arm}$$

Here $\mathrm{E}(\text{admissions})$ is the expected number of admissions, arm is 1 for the programme and 0 for usual care, and $\beta_0$ and $\beta_1$ are coefficients. Moving the offset to the left side gives the log of the rate, so $\exp(\beta_1)$ is the rate ratio. With arm as the only covariate, the fitted rate in each arm equals its observed rate, so the model returns 0.50 exactly, whether the data hold two summary rows or one row per patient.

Fixing the offset coefficient at 1 assumes a roughly constant rate over follow-up. If admissions cluster soon after discharge, the arm followed longer adds late, low-rate person-time and may look better for that alone. Comparing rates within the same periods since discharge, or a time-to-event method, may then be preferable. An arm that started first also differs in calendar time.

Its Stata and R commands are in the scenes table below, and a separate article covers Poisson regression for rates.

Dependence: which rows belong together

A trial randomises hospitals, not patients, to a new discharge checklist. Patients in one hospital share staff and routines, so their outcomes are more alike than outcomes from different hospitals.

Rows may be independent, or form clusters such as patients within hospitals, repeated measures of one patient over time, or matched sets of cases and controls grouped by design on age and sex.

Treating correlated rows as independent usually makes standard errors too small, especially for comparisons between clusters [2]. A mixed model adds random effects, meaning cluster-level deviations from the average. Generalized estimating equations (GEE) model the population-average outcome and allow for correlation within clusters. A third option keeps the model and uses a cluster-robust, or sandwich, standard error, estimated from the observed variation between clusters; part 4 explains what it repairs and what it cannot.

For repeated visits in a trial, a common choice is the mixed model for repeated measures (MMRM) [5]. It treats visit as a category and gives each visit its own variance and each pair of visits its own correlation. For matched sets, conditional logistic regression compares exposure within each set, so the matching factors drop out. Ordinary logistic regression with one indicator per set gives a biased odds ratio when sets are small, as with pairs.

Estimand: the quantity the study promises

At a trial steering meeting, a member asks for the effect of the programme on admissions. The statistician replies that this is not yet an estimand, because it names no scale, no comparison and no population.

The ICH (International Council for Harmonisation) E9(R1) guideline on clinical trials describes an estimand, the precise quantity a study aims to estimate, by five attributes: treatment, population, variable (the outcome measured), handling of intercurrent events, and population-level summary [6]. Intercurrent events, such as stopping treatment or death, occur after treatment starts and affect whether the outcome can be measured or how it is read.

For model choice, the slot settles a difference or a ratio; risks, rates, odds or hazards; a marginal or conditional contrast; and the population. A marginal contrast compares the whole population under one treatment with the whole population under the other. A conditional contrast compares patients who share the covariate values in the model.

A risk ratio is not read from a logistic coefficient, which is a log odds ratio. For a common outcome the log-binomial model, a binomial regression on the log scale, can fail to converge, meaning the software's step-by-step search cannot settle on coefficients that keep every patient's predicted risk below 1 and stops without a usable answer; part 3 gives a working sequence of alternatives. How a conditional and a marginal odds ratio from the same patients relate is the subject of part 5.

Special structure: features that bend a standard model

A reviewer reads a manuscript that models emergency visits with Poisson regression. The counts vary far more than their mean, more patients had none than a Poisson model predicts, some died and others moved away. Each is a special structure that a standard model does not handle by default.

Often this slot adds a component rather than a new class: a variance model, a spline or an interaction term. Competing events and excess zeros can change the method itself.

Five research scenes through the grammar

The table parses five research scenes through the five slots.

No command here has been run, and none shows output. Names are placeholders; the admissions row uses the hand example's names (hand: arm, admissions, pyears). Covariates are left out: in a non-randomised comparison they follow from the causal question (see Where the grammar stops). In the competing-events row, status is 1 for admission, 2 for death and 0 for censoring; stcompet is user-written (ssc install stcompet).
Research sceneOutcomeDenominator or timeDependenceEstimandSpecial structureModel classStata commandR function
Delirium within 7 days of hip-fracture surgery, early versus later surgeryBinaryPer person, fixed horizonIndependent patientsRisk ratio, all operated patientsCommon outcome; the fit may not converge (alternatives in part 3)Log-binomial regressionglm delirium i.early, family(binomial) link(log) eformglm(delirium ~ early, family = binomial(link = "log"), data = d)
Admissions under a nurse-led programme versus usual careCountPerson-time, differing by patientIndependent patientsRate ratio, all enrolled patientsOverdispersion to checkPoisson regression with a log person-time offsetpoisson admissions i.arm, exposure(pyears) irrglm(admissions ~ arm + offset(log(pyears)), family = poisson, data = hand)
Symptom score at four visits in a randomised trialContinuousFixed visitsRepeated measuresMean difference at the final visit, all randomisedMissing visits, assumed missing at random given the observed visitsMMRMmixed score i.arm##i.visit || id:, noconstant residuals(unstructured, t(visit)) remlmmrm::mmrm(score ~ arm * visit + us(visit | id), data = d)
Time to first admission after discharge, death competingTime to an eventTime to admission, death or censoringIndependent patientsCumulative incidence by a stated time, by armCensoring; death competingAalen-Johansen estimate of the cumulative incidence function (Fine-Gray regression if covariates are needed)stset time, failure(status == 1), then stcompet cif = ci, compet1(2) by(arm)cmprsk::cuminc(d$time, d$status, group = d$arm)
Matched case-control study of hip fracture and a prescribed sedativeCase or controlNone: sampling is on the outcomeSets matched on age and sexOdds ratio within matched setsEffects of the matching factors cannot be estimatedConditional logistic regressionclogit case i.exposed, group(set) orsurvival::clogit(case ~ exposed + strata(set), data = d)

Reading the table

Two rows concern admissions, yet they land in different classes: a count over person-time, and a time to the first admission with death competing. The question, not the variable name, decides the slot.

Set each slot with its dropdown. The panel names the model family, the effect measure, a Stata command and an R function, and flags combinations the grammar cannot settle alone.

Where the grammar stops

The grammar narrows the choice to models that can answer the question. It does not choose covariates, which follow from the causal question, often drawn as a DAG, or from the prediction task [3]. Nor does it check assumptions, such as a roughly constant rate or proportional hazards, a hazard ratio constant over follow-up, and it does not replace design thinking about randomisation, measurement and missing data.

Two analysts who fill the slots alike may still fit different models in one class, and both can be defensible when each states its estimand and assumptions.

Common misreadings and their fixes

  • "The grammar picks the model."

    Filling the slots rules out models that cannot answer the question, but several candidates remain.

    Fix: The grammar narrows the choice to models that can answer the question. The analyst still chooses covariates, checks assumptions and justifies the estimand.

  • "Comparing admission counts is enough."

    In the hand example the counts give 0.60, but the rate ratio is 0.50, because the programme arm had 1,200 person-years against 1,000.

    Fix: Put person-time in the denominator, as a rate or a log person-time offset, and check that the rate is roughly constant over follow-up.

  • "A binary outcome means logistic regression."

    A logistic coefficient is a log odds ratio, and the odds ratio lies further from 1 than the risk ratio, more so for a common outcome.

    Fix: Name the effect measure in the estimand, then choose a model that returns it.

  • "Time-to-event data means a Cox model."

    A hazard ratio compares event rates among patients still event-free. When death competes, the probability of admission by a stated time comes from the cumulative incidence function.

    Fix: Decide whether the estimand is a hazard ratio or a probability by a stated time.

What to do in your own analysis

Glossary

generalized linear model (แบบจำลองเชิงเส้นวางนัยทั่วไป (generalized linear model, GLM))
A regression class, including linear, logistic and Poisson regression, defined by a family and a link.
person-time (เวลาติดตามสะสม (person-time))
Follow-up time summed over all patients, such as person-years.
offset (ออฟเซต (offset))
A term with its coefficient fixed at 1; log person-time as the offset turns a count model into a rate model.
estimand (ปริมาณเป้าหมายของการประมาณ (estimand))
The precise quantity a study aims to estimate, with its population, contrast and summary measure.
mixed model for repeated measures (แบบจำลองผสมสำหรับการวัดซ้ำ (MMRM))
Visit as a category, treatment-by-visit terms and an unstructured covariance between visits.
conditional logistic regression (การถดถอยลอจิสติกแบบมีเงื่อนไข (conditional logistic regression))
Logistic regression that compares exposure within matched sets.
overdispersion (ความแปรปรวนเกิน (overdispersion))
Count variance larger than the mean, which Poisson regression assumes equal.
censoring (การเซ็นเซอร์ (censoring))
Follow-up that ends before the event is seen.
competing event (เหตุการณ์แข่งขัน (competing event))
An event, such as death, that prevents the event of interest.
cumulative incidence function (อุบัติการณ์สะสม (cumulative incidence function, CIF))
The probability of the event of interest by a stated time when competing events can occur.

References

  1. McCullagh P, Nelder JA. Generalized linear models. 2nd ed. London: Chapman and Hall; 1989. doi:10.1007/978-1-4899-3242-6 https://doi.org/10.1007/978-1-4899-3242-6
  2. Vittinghoff E, Glidden DV, Shiboski SC, McCulloch CE. Regression methods in biostatistics: linear, logistic, survival, and repeated measures models. 2nd ed. New York: Springer; 2012. doi:10.1007/978-1-4614-1353-0 https://doi.org/10.1007/978-1-4614-1353-0
  3. Harrell FE Jr. Regression modeling strategies: with applications to linear models, logistic and ordinal regression, and survival analysis. 2nd ed. Cham: Springer; 2015. doi:10.1007/978-3-319-19425-7 https://doi.org/10.1007/978-3-319-19425-7
  4. Gardner W, Mulvey EP, Shaw EC. Regression analyses of counts and rates: Poisson, overdispersed Poisson, and negative binomial models. Psychol Bull. 1995;118(3):392-404. doi:10.1037/0033-2909.118.3.392 https://doi.org/10.1037/0033-2909.118.3.392
  5. Fitzmaurice GM, Laird NM, Ware JH. Applied longitudinal analysis. 2nd ed. Wiley; 2011. doi:10.1002/9781119513469 https://doi.org/10.1002/9781119513469
  6. International Council for Harmonisation (ICH). ICH E9(R1): Addendum on estimands and sensitivity analysis in clinical trials to the guideline on statistical principles for clinical trials. ICH Harmonised Guideline, Step 4; 2019. https://database.ich.org/sites/default/files/E9-R1_Step4_Guideline_2019_1203.pdf

Key takeaways

  • Five slots, from outcome to special structure, narrow the choice of model before any command is typed.
  • When follow-up differs, count events over person-time and check that the rate is roughly constant: in the hand example the rate ratio is 0.50, while raw counts give 0.60.
  • Rows that share a hospital, a patient or a matched set are not independent, and the analysis usually has to allow for it.
  • The estimand decides which model in a class can deliver the number the study promised.
  • The grammar narrows the choice; the analyst still chooses covariates, checks assumptions and justifies the estimand.

Related in the wiki: [[glm-link-family-guide]] [[poisson-regression-rates-lambda]]

0
Message for International and Thai ReadersUnderstanding My Medical Context in ThailandRead more →Message for International and Thai ReadersUnderstanding My Broader Content Beyond MedicineRead more →

Comments

No comments yet. Be the first to share your thoughts.

Sign in to comment