P-Values, Confidence Intervals and Effect Sizes: How to Read Them Together
Effect estimates, confidence intervals and p-values answer related but different questions. The effect estimate tells you the direction and magnitude of the observed effect or association. The confidence interval shows the uncertainty around that estimate. The p-value summarizes compatibility with a specified null model. None of these alone tells you whether the result is scientifically, clinically or practically important.
Your results table says:
Risk difference: 7 percentage points
95% CI: 1 to 13 percentage points
p = 0.03
A surprisingly common interpretation is:
“The result was significant.”
That sentence has managed to ignore almost everything the result tells you.
The better question is:
What happened, how uncertain are we, and does it matter?
The three numbers are doing different jobs
A useful reading sequence is:
Estimate → uncertainty → compatibility with the null → practical importance → study credibility
Not:
p-value → celebration or despair → next table
That distinction matters because a statistically convincing result can be tiny, while an uncertain result can still be compatible with effects that would matter greatly.
1. The effect estimate asks: how much?
An effect estimate tells you the magnitude and usually the direction of the observed difference or association.
Depending on the study, that might be:
- a mean difference;
- a risk difference;
- a risk ratio;
- an odds ratio;
- a rate ratio;
- a hazard ratio;
- a regression coefficient;
- a correlation coefficient;
- a standardized mean difference.
Suppose an intervention increases contraceptive uptake from 51% to 58%.
One effect estimate is the absolute difference:
58% − 51% = 7 percentage points
Now the reader knows something useful.
The intervention group had an estimated uptake seven percentage points higher than the comparison group.
The p-value cannot tell you that magnitude.
2. The confidence interval asks: how uncertain is the estimate?
Now suppose the result is:
Risk difference = 7 percentage points
95% CI = 1 to 13 percentage points
The point estimate is seven.
The interval shows that there is still uncertainty around it.
Under the statistical model and assumptions used, the data are compatible with effects ranging from a small increase of around one percentage point to an increase of around 13 percentage points.
That range matters.
If a one-point change would be too small to justify a programme, but a 13-point change would alter national policy, the uncertainty is not a footnote.
It is part of the decision.
3. The p-value asks a narrower question
Now add:
p = 0.03
The p-value contributes information about how compatible the observed data are with the tested null model.
It does not tell you:
- that the intervention has a 97% chance of working;
- that the observed effect is large;
- that the result is important;
- that bias is absent;
- that the design supports a causal conclusion;
- that the result will replicate.
A small p-value does not upgrade a weak study into a strong one.
A p-value can contribute to the evidence. It cannot tell you whether the finding matters.
one result, five questions
Return to the contraceptive-uptake example:
- Comparison group: 51%
- Intervention group: 58%
- Risk difference: 7 percentage points
- 95% CI: 1 to 13 percentage points
- p = 0.03
Read it properly.
Question 1: What is the estimated effect?
Seven percentage points higher uptake.
Question 2: How uncertain is the estimate?
The interval spans roughly one to 13 percentage points.
Question 3: Is the interval compatible with the null?
For a risk difference, the usual null value is zero.
This interval does not include zero.
Question 4: Would the effect matter?
That cannot be determined from p = 0.03.
You need context.
Would a seven-point increase justify the intervention cost?
Would one point?
Would 13 points?
Would benefits be equitable across groups?
Are there adverse consequences?
Question 5: Do you trust the estimate?
Now leave the table and inspect the study.
- Was allocation appropriate?
- Was the outcome measured well?
- Was missing data substantial?
- Was the analysis prespecified?
- Was clustering handled?
- Were important confounding structures addressed, if observational?
- Does the design support the conclusion being made?
The numerical result is only as credible as the research process that produced it.
Four scenarios that all look different once you stop worshipping p < 0.05
Scenario A: substantial effect, reasonably precise
RR = 0.60
95% CI 0.48-0.75
p < 0.001
The estimate suggests a substantial reduction, and the interval is reasonably precise.
You still need to ask whether the design is credible and whether the effect is important in context.
But the p-value is not carrying the interpretation alone.
Scenario B: tiny effect, extremely precise
RR = 0.98
95% CI 0.97-0.99
p < 0.001
This may be highly statistically convincing.
It is still only an estimated 2% relative difference.
Whether that matters depends on:
- baseline risk;
- outcome severity;
- population size;
- intervention cost;
- potential harms;
- feasibility;
- equity.
“Highly significant” does not mean “highly important.”
Scenario C: potentially important effect, substantial uncertainty
RR = 0.70
95% CI 0.45-1.09
p = 0.11
The point estimate suggests a potentially meaningful reduction.
The confidence interval is wide and includes the null.
Calling this simply:
“There was no effect.”
throws away most of the information.
The appropriate conclusion is about uncertainty.
Scenario D: very small effect, reasonably precise
RR = 1.01
95% CI 0.99-1.03
p = 0.35
The result is not conventionally statistically significant.
But the narrow interval may tell you something important: large relative effects are not very compatible with these data under the model.
That can be scientifically useful.
Bench Check: reporting only the p-value
You wrote:
“There was a significant relationship between intervention exposure and contraceptive uptake (p = 0.03).”
The reader still does not know:
- how much uptake differed;
- in which direction;
- how uncertain the difference was;
- whether the difference matters.
Replace the significance label with an estimate-led sentence.
For example:
Contraceptive uptake was seven percentage points higher in the intervention group than in the comparison group (95% CI 1-13 percentage points; p = 0.03).
Now the reader can evaluate the result.
Effect size does not automatically mean “Cohen's d”
Another common problem is using “effect size” as though it refers to one special statistical number.
An effect estimate should match the research question.
If your outcome is binary, a risk ratio, risk difference or odds ratio may be relevant.
If your outcome is continuous, a mean difference may be more interpretable than a standardized effect.
If you need to compare effects across different measurement scales, a standardized effect measure may be useful.
The important question is:
What quantity best communicates the magnitude of the phenomenon I am studying?
Do not calculate Cohen's d merely because someone told you every study needs “an effect size.”
Relative and absolute effects can tell different stories
Suppose disease risk falls from:
2% to 1%
The relative risk is:
0.50
That is a 50% relative reduction.
The absolute risk difference is:
1 percentage point
Both are correct.
They answer different aspects of the question.
In public-health decision-making, absolute effects can be especially important because they connect directly to the number of events potentially prevented.
Whenever possible, choose effect measures that help the intended reader understand the actual consequences.
What do I actually write in my results section?
Avoid:
There was a statistically significant association between intervention exposure and contraceptive uptake (p < 0.05).
Prefer:
Contraceptive uptake was 58% in the intervention group and 51% in the comparison group, corresponding to an absolute difference of seven percentage points (95% CI 1-13; p = 0.03).
If the analysis is adjusted:
After adjustment for the prespecified covariates, intervention exposure was associated with [effect estimate] (95% CI [lower-upper]; p = [value]).
But the adjustment variables should have a substantive rationale.
Do not select covariates merely because they happened to produce p < 0.20 in a preliminary table.
The Methods Bench reading sequence
When you encounter a quantitative result, use this sequence:
1. What is the estimate?
What happened, and in which direction?
2. What is the uncertainty?
How wide is the confidence interval?
3. What does the interval include?
Null effects? Trivial effects? Important benefit? Potential harm?
4. What does the p-value contribute?
How compatible are the data with the tested null model?
5. Does the effect matter?
Clinically? Practically? Socially? Programmatically?
6. Is the study credible?
Design, measurement, analysis and bias still matter.
That is interpretation.
Everything else is number watching.
Open your main results table.
For your most important result, write five lines:
- Estimate:
- Confidence interval:
- Null value:
- Smallest effect I would care about:
- Main design limitation affecting interpretation:
If you can only fill in the p-value, the table is not yet doing enough scientific work.
Frequently asked questions
- Is the effect size more important than the p-value?
They answer different questions. The estimate tells you magnitude; the p-value does not. Interpretation should also consider uncertainty, practical importance and study credibility.
- Does a confidence interval replace a p-value?
A confidence interval often provides richer information because it shows the estimate and uncertainty. Some analyses still report p-values, but the p-value should not replace the interval or estimate.
- Can a result be statistically significant but not important?
Yes. With sufficient information, very small effects can produce small p-values.
- Can a result be important but not statistically significant?
Potentially. A point estimate may be substantively important while the confidence interval remains wide and includes the null. The appropriate conclusion is uncertainty, not proof of importance.
- What should I report first?
Usually lead with the effect estimate and confidence interval. Add the p-value where relevant to the reporting context and analysis.
Use the Methods Bench Manuscript Outline Planner to audit whether your results section reports:
- actual estimates;
- confidence intervals;
- p-values where useful;
- enough context to understand the effect.
References and further reading
- 1.Wasserstein RL, Lazar NA. The ASA's Statement on p-Values: Context, Process, and Purpose. The American Statistician. 2016.
- 2.American Statistical Association. Statement and task-force guidance on statistical significance and p-values.
- 3.Gardner MJ, Altman DG. Confidence intervals rather than P values: estimation rather than hypothesis testing. BMJ. 1986.
- 4.STROBE Statement. Reporting guidance for cohort, case-control and cross-sectional studies.
- 5.CONSORT reporting guidance for randomized trials.
Take it further
Better research, once a week.
Practical methods guidance, research tools and funding opportunities from Methods Bench.
Get practical research notes and new opportunities
Short, useful emails for researchers in Uganda and East Africa. No spam.