Sample Size Fundamentals: Turning an Effect Assumption Into a Trial
By Dr. Sakshi Bharti, COO Aomics GmbH
A useful habit for a new statistician is to ask what the number will mean before asking how to calculate it. Sample size reflects assumptions about effect, variability or event rates, alpha, power, allocation and attrition. Adaptive designs need simulation beyond a single closed-form number.
Sample Size Fundamentals: Turning an Effect Assumption Into a Trial: practical analytical flow.
The idea in plain language
Sample size reflects assumptions about effect, variability or event rates, alpha, power, allocation and attrition. Adaptive designs need simulation beyond a single closed-form number.
The practical question is always the same: what quantity or decision are we trying to support, and what assumptions connect the observed data to that target? This is why a protocol, estimand, analysis model and output should never be treated as separate paperwork exercises. They are parts of one argument.
The statistical core
For equal-allocation continuous outcomes, n_per_arm≈2σ²(z_(1−α/2)+z_power)²/δ².
The equation matters, but it should never float free of its clinical meaning. In production work, every symbol needs a source: an endpoint definition, a population rule, a planning assumption, a dataset variable, or a frozen configuration value.
Why this matters in a real trial
Planning should be shown as a sensitivity exercise, not a single magic N. Teams should see how N changes with effect, variability, event rate, dropout and allocation, and which inputs are sourced versus judgement-based.
A mature workflow separates design-time assumptions from analysis-time observations. Planning values retain their evidence source and approval; analysis values retain dataset, data cut, mapping, population and engine provenance.
A small worked example
Halving a target effect can roughly quadruple the required information.
The numbers in a worked example are illustrative. In a real study, Adaptrails should require user-approved parameters or versioned study data and retain the input source beside the result.
Inputs and provenance to capture
For this topic, the analysis record should capture at least: study and endpoint identifiers; analysis population; source of each planning assumption or input dataset; parameter values explicitly accepted by the user; engine/version; random seed where simulation is used; and the review state. For data-based analyses, the source dataset, immutable data cut and variable mapping belong in the same evidence record.
How this maps to Adaptrails
Relevant modules: M04, M05, M09. The platform should guide the user from inputs to analysis to explanation without duplicating statistical logic across modules. A method should have one authoritative implementation, while SAP Studio, dashboards and AI explanations consume the stored result.
Write the population, endpoint, treatment contrast and data source before selecting the method.
Make model, timing, missing-data and operational assumptions explicit.
Adaptrails provides evidence and recommendations; qualified professionals retain decision authority.
What a junior statistician should check
- Is the clinical question explicit before the calculation?
- Are all inputs sourced or user-approved?
- Does the analysis population match the estimand?
- Is uncertainty reported, not only a point estimate?
- Are thresholds and defaults visible rather than hidden?
- Can another statistician reproduce the result from the stored data cut and configuration?
Common mistakes
- Choosing a familiar statistical procedure before defining the target quantity.
- Treating a software default as if it were a study assumption.
- Reporting a technically correct number without explaining its clinical interpretation.
- Recalculating the same statistic in a document generator or UI instead of using the authoritative engine result.
Three questions to ask in a project meeting
1. What will change if the assumption is wrong?
Run sensitivity analysis rather than defending one planning value as truth.
2. Which data source will feed the analysis?
At design stage, record the evidence source for assumptions. At analysis stage, record dataset, data cut, mappings and population flags.
3. Who makes the decision?
The statistical system calculates and may recommend. A statistician, DMC, medical monitor, clinical team or other authorized body retains the decision appropriate to the context.
Where to go next
For broader clinical biometrics and data-science services, visit Aomics GmbH. To explore the adaptive-trial design and statistical workflow platform, visit Adaptrials.
A useful discipline is to treat every numerical output as the end of a documented argument rather than the beginning of one.
References and further reading
- ICH E9: Statistical Principles for Clinical Trials
- FDA: Adaptive Designs for Clinical Trials of Drugs and Biologics
Decision-support note: Adaptrails is designed to support qualified clinical and statistical professionals. Statistical results, simulations and AI-assisted explanations require appropriate human review before clinical, safety, operational or regulatory action.