Statistics & Interpretation

P-Values, Confidence Intervals and Effect Sizes: How to Read Them Together

Methods BenchReviewed by Research Methods SpecialistPublished 19 August 2026Last reviewed 19 August 20268 min read
Direct answer

Effect estimates, confidence intervals and p-values answer related but different questions. The effect estimate tells you the direction and magnitude of the observed effect or association. The confidence interval shows the uncertainty around that estimate. The p-value summarizes compatibility with a specified null model. None of these alone tells you whether the result is scientifically, clinically or practically important.

Your results table says:

Risk difference: 7 percentage points
95% CI: 1 to 13 percentage points
p = 0.03

A surprisingly common interpretation is:

“The result was significant.”

That sentence has managed to ignore almost everything the result tells you.

The better question is:

What happened, how uncertain are we, and does it matter?

The three numbers are doing different jobs

A useful reading sequence is:

Estimate → uncertainty → compatibility with the null → practical importance → study credibility

Not:

p-value → celebration or despair → next table

That distinction matters because a statistically convincing result can be tiny, while an uncertain result can still be compatible with effects that would matter greatly.

1. The effect estimate asks: how much?

An effect estimate tells you the magnitude and usually the direction of the observed difference or association.

Depending on the study, that might be:

  • a mean difference;
  • a risk difference;
  • a risk ratio;
  • an odds ratio;
  • a rate ratio;
  • a hazard ratio;
  • a regression coefficient;
  • a correlation coefficient;
  • a standardized mean difference.

Suppose an intervention increases contraceptive uptake from 51% to 58%.

One effect estimate is the absolute difference:

58% − 51% = 7 percentage points

Now the reader knows something useful.

The intervention group had an estimated uptake seven percentage points higher than the comparison group.

The p-value cannot tell you that magnitude.

2. The confidence interval asks: how uncertain is the estimate?

Now suppose the result is:

Risk difference = 7 percentage points
95% CI = 1 to 13 percentage points

The point estimate is seven.

The interval shows that there is still uncertainty around it.

Under the statistical model and assumptions used, the data are compatible with effects ranging from a small increase of around one percentage point to an increase of around 13 percentage points.

That range matters.

If a one-point change would be too small to justify a programme, but a 13-point change would alter national policy, the uncertainty is not a footnote.

It is part of the decision.

3. The p-value asks a narrower question

Now add:

p = 0.03

The p-value contributes information about how compatible the observed data are with the tested null model.

It does not tell you:

  • that the intervention has a 97% chance of working;
  • that the observed effect is large;
  • that the result is important;
  • that bias is absent;
  • that the design supports a causal conclusion;
  • that the result will replicate.

A small p-value does not upgrade a weak study into a strong one.

Bench rule

A p-value can contribute to the evidence. It cannot tell you whether the finding matters.

On the bench

one result, five questions

Return to the contraceptive-uptake example:

  • Comparison group: 51%
  • Intervention group: 58%
  • Risk difference: 7 percentage points
  • 95% CI: 1 to 13 percentage points
  • p = 0.03

Read it properly.

Question 1: What is the estimated effect?

Seven percentage points higher uptake.

Question 2: How uncertain is the estimate?

The interval spans roughly one to 13 percentage points.

Question 3: Is the interval compatible with the null?

For a risk difference, the usual null value is zero.

This interval does not include zero.

Question 4: Would the effect matter?

That cannot be determined from p = 0.03.

You need context.

Would a seven-point increase justify the intervention cost?

Would one point?

Would 13 points?

Would benefits be equitable across groups?

Are there adverse consequences?

Question 5: Do you trust the estimate?

Now leave the table and inspect the study.

  • Was allocation appropriate?
  • Was the outcome measured well?
  • Was missing data substantial?
  • Was the analysis prespecified?
  • Was clustering handled?
  • Were important confounding structures addressed, if observational?
  • Does the design support the conclusion being made?

The numerical result is only as credible as the research process that produced it.

Four scenarios that all look different once you stop worshipping p < 0.05

Scenario A: substantial effect, reasonably precise

RR = 0.60
95% CI 0.48-0.75
p < 0.001

The estimate suggests a substantial reduction, and the interval is reasonably precise.

You still need to ask whether the design is credible and whether the effect is important in context.

But the p-value is not carrying the interpretation alone.

Scenario B: tiny effect, extremely precise

RR = 0.98
95% CI 0.97-0.99
p < 0.001

This may be highly statistically convincing.

It is still only an estimated 2% relative difference.

Whether that matters depends on:

  • baseline risk;
  • outcome severity;
  • population size;
  • intervention cost;
  • potential harms;
  • feasibility;
  • equity.

“Highly significant” does not mean “highly important.”

Scenario C: potentially important effect, substantial uncertainty

RR = 0.70
95% CI 0.45-1.09
p = 0.11

The point estimate suggests a potentially meaningful reduction.

The confidence interval is wide and includes the null.

Calling this simply:

“There was no effect.”

throws away most of the information.

The appropriate conclusion is about uncertainty.

Scenario D: very small effect, reasonably precise

RR = 1.01
95% CI 0.99-1.03
p = 0.35

The result is not conventionally statistically significant.

But the narrow interval may tell you something important: large relative effects are not very compatible with these data under the model.

That can be scientifically useful.

Bench check

Bench Check: reporting only the p-value

You wrote:

“There was a significant relationship between intervention exposure and contraceptive uptake (p = 0.03).”

The reader still does not know:

  • how much uptake differed;
  • in which direction;
  • how uncertain the difference was;
  • whether the difference matters.

Replace the significance label with an estimate-led sentence.

For example:

Contraceptive uptake was seven percentage points higher in the intervention group than in the comparison group (95% CI 1-13 percentage points; p = 0.03).

Now the reader can evaluate the result.

Effect size does not automatically mean “Cohen's d”

Another common problem is using “effect size” as though it refers to one special statistical number.

An effect estimate should match the research question.

If your outcome is binary, a risk ratio, risk difference or odds ratio may be relevant.

If your outcome is continuous, a mean difference may be more interpretable than a standardized effect.

If you need to compare effects across different measurement scales, a standardized effect measure may be useful.

The important question is:

What quantity best communicates the magnitude of the phenomenon I am studying?

Do not calculate Cohen's d merely because someone told you every study needs “an effect size.”

Relative and absolute effects can tell different stories

Suppose disease risk falls from:

2% to 1%

The relative risk is:

0.50

That is a 50% relative reduction.

The absolute risk difference is:

1 percentage point

Both are correct.

They answer different aspects of the question.

In public-health decision-making, absolute effects can be especially important because they connect directly to the number of events potentially prevented.

Whenever possible, choose effect measures that help the intended reader understand the actual consequences.

What do I actually write?

What do I actually write in my results section?

Avoid:

There was a statistically significant association between intervention exposure and contraceptive uptake (p < 0.05).

Prefer:

Contraceptive uptake was 58% in the intervention group and 51% in the comparison group, corresponding to an absolute difference of seven percentage points (95% CI 1-13; p = 0.03).

If the analysis is adjusted:

After adjustment for the prespecified covariates, intervention exposure was associated with [effect estimate] (95% CI [lower-upper]; p = [value]).

But the adjustment variables should have a substantive rationale.

Do not select covariates merely because they happened to produce p < 0.20 in a preliminary table.

The Methods Bench reading sequence

When you encounter a quantitative result, use this sequence:

1. What is the estimate?

What happened, and in which direction?

2. What is the uncertainty?

How wide is the confidence interval?

3. What does the interval include?

Null effects? Trivial effects? Important benefit? Potential harm?

4. What does the p-value contribute?

How compatible are the data with the tested null model?

5. Does the effect matter?

Clinically? Practically? Socially? Programmatically?

6. Is the study credible?

Design, measurement, analysis and bias still matter.

That is interpretation.

Everything else is number watching.

Do this now

Open your main results table.

For your most important result, write five lines:

  1. Estimate:
  2. Confidence interval:
  3. Null value:
  4. Smallest effect I would care about:
  5. Main design limitation affecting interpretation:

If you can only fill in the p-value, the table is not yet doing enough scientific work.

Frequently asked questions

Is the effect size more important than the p-value?

They answer different questions. The estimate tells you magnitude; the p-value does not. Interpretation should also consider uncertainty, practical importance and study credibility.

Does a confidence interval replace a p-value?

A confidence interval often provides richer information because it shows the estimate and uncertainty. Some analyses still report p-values, but the p-value should not replace the interval or estimate.

Can a result be statistically significant but not important?

Yes. With sufficient information, very small effects can produce small p-values.

Can a result be important but not statistically significant?

Potentially. A point estimate may be substantively important while the confidence interval remains wide and includes the null. The appropriate conclusion is uncertainty, not proof of importance.

What should I report first?

Usually lead with the effect estimate and confidence interval. Add the p-value where relevant to the reporting context and analysis.

Try it

Use the Methods Bench Manuscript Outline Planner to audit whether your results section reports:

  • actual estimates;
  • confidence intervals;
  • p-values where useful;
  • enough context to understand the effect.
Open the Manuscript Outline Planner

References and further reading

  1. 1.Wasserstein RL, Lazar NA. The ASA's Statement on p-Values: Context, Process, and Purpose. The American Statistician. 2016.
  2. 2.American Statistical Association. Statement and task-force guidance on statistical significance and p-values.
  3. 3.Gardner MJ, Altman DG. Confidence intervals rather than P values: estimation rather than hypothesis testing. BMJ. 1986.
  4. 4.STROBE Statement. Reporting guidance for cohort, case-control and cross-sectional studies.
  5. 5.CONSORT reporting guidance for randomized trials.

Take it further

Better research, once a week.

Practical methods guidance, research tools and funding opportunities from Methods Bench.

Get practical research notes and new opportunities

Short, useful emails for researchers in Uganda and East Africa. No spam.

You can unsubscribe at any time.

Written by Methods Bench. Reviewed by Research Methods Specialist.All research guides