Statistics & Interpretation

Statistically Significant Does Not Mean Important: How to Judge Whether a Result Matters

Methods BenchReviewed by Research Methods SpecialistPublished 29 August 2026Last reviewed 29 August 20266 min read
Direct answer

Statistical significance tells you whether data meet a statistical decision rule under a specified model. It does not tell you whether the estimated effect is large enough to matter. Clinical, practical or programmatic importance depends on the magnitude of the effect, uncertainty, baseline risk, consequences, costs, feasibility, equity and what difference would have been considered meaningful before seeing the result.

Your study has 18,000 participants.

The intervention increases a 100-point knowledge score by:

0.8 points

And the result is:

p < 0.001

Technically impressive.

Now ask the question the p-value cannot answer:

Would 0.8 points change anything that matters?

That is a different scientific question.

What is statistical significance?

Statistical significance usually refers to whether a result crosses a prespecified hypothesis-testing threshold, commonly p < 0.05.

It is influenced by:

  • effect magnitude;
  • sample size;
  • variability;
  • event frequency;
  • model assumptions;
  • the significance level.

A large study can therefore produce a very small p-value for a very small effect.

That does not make the effect practically large.

The American Statistical Association has explicitly cautioned against using statistical significance as a measure of scientific or practical importance.

What is practical significance?

Practical significance asks:

Is the magnitude large enough to matter in the real-world context of the study?

In clinical research, related concepts include:

  • clinical importance;
  • minimal important difference;
  • patient-important benefit;
  • clinically meaningful change.

In public health and social research, practical importance may involve:

  • number of people affected;
  • absolute rather than relative change;
  • feasibility;
  • cost;
  • equity;
  • implementation burden;
  • harms;
  • social consequences;
  • policy relevance.

There is no universal p-value threshold for “important.”

There cannot be.

Importance is substantive.

Bench rule

Statistical significance is a property of an analysis. Importance is a judgement about consequences.

Do not ask one number to do both jobs.

On the bench

tiny effect, tiny p-value

Suppose a school-based intervention produces:

  • Mean knowledge score in intervention group: 76.4
  • Mean knowledge score in comparison group: 75.6
  • Difference: 0.8 points on a 100-point scale
  • 95% CI: 0.5 to 1.1
  • p < 0.001
  • n = 18,000

The estimate is precise.

The statistical evidence against a zero-difference null may be strong.

But is 0.8 points educationally meaningful?

That requires context.

Ask:

  • What is a meaningful change on this scale?
  • Does the score predict a behaviour or outcome anyone cares about?
  • Did the programme cost 200?
  • Does the benefit reach disadvantaged students?
  • Is a 0.8-point gain sustained?
  • Are there other outcomes that matter more?

A tiny p-value cannot answer these questions.

Reverse the example: potentially important effect, uncertain estimate

Now suppose a programme increases contraceptive uptake by:

11 percentage points

but the confidence interval is wide:

95% CI: -2 to 24 percentage points
p = 0.09

The result is not conventionally statistically significant.

But the interval includes effects that a programme decision-maker might care about.

It also includes little or no benefit.

The correct conclusion is:

The estimate is potentially important but too uncertain to determine the magnitude reliably.

Not:

“There was no effect.”

And not:

“The programme works.”

Uncertainty is an answer.

What is clinical significance?

Clinical importance asks whether an effect is meaningful for patients, symptoms, function, quality of life, risk or another clinically relevant outcome.

For some outcomes, researchers use a minimal important difference or related threshold.

But be careful.

A minimal important difference is not one universally correct number that can be copied from the first paper you find.

Its value can depend on:

  • population;
  • instrument;
  • anchor;
  • estimation method;
  • direction of change;
  • context.

If an important-difference threshold is central to the study, define and justify it before interpreting the result.

What about public health, where tiny effects can matter?

This is where “small effect = unimportant” also fails.

Imagine a vaccine produces a small absolute reduction in a common serious outcome across millions of people.

The individual effect might look small.

Population impact may be substantial.

Or suppose a national intervention reduces mortality by one percentage point.

That may represent thousands of lives.

So practical importance should consider scale.

Relative effect

Risk falls from 2% to 1%.

Relative reduction:

50%

Sounds dramatic.

Absolute effect

Absolute reduction:

1 percentage point

Also correct.

Which is more useful depends on the decision.

Often you need both.

Bench check

Bench Check: “highly significant”

You wrote:

“The intervention produced a highly significant improvement (p < 0.001).”

Highly significant statistically?

Or highly important?

The adjective can blur the two.

Prefer:

“The intervention group scored 0.8 points higher than the comparison group (95% CI 0.5-1.1; p < 0.001).”

Then discuss whether 0.8 points matters.

Let the reader see the magnitude before you praise the p-value.

Decide what would matter before seeing the result

One of the strongest habits in study planning is to ask:

What magnitude would change my scientific or practical conclusion?

For a trial, that may inform sample-size planning.

For an observational study, it helps interpretation.

For programme evaluation, ask stakeholders:

  • What level of improvement would justify the cost?
  • What change would be too small to matter?
  • What harms would offset the benefit?
  • Which groups must benefit?
  • What level of uncertainty is acceptable?

This prevents a common post-hoc problem:

p < 0.05 → therefore whatever effect we found must be important.

No.

Statistical significance, practical significance and study credibility are three separate questions

A useful interpretation asks:

1. Is there statistical evidence?

P-value, confidence interval, model.

2. How large is the effect?

Absolute and relative magnitude as appropriate.

3. Does that magnitude matter?

Clinical, practical, social or programmatic context.

4. Is the estimate credible?

Bias, design, measurement, missing data, confounding, multiplicity, model assumptions.

A precise estimate of a biased effect is still biased.

A clinically important effect from a badly designed study is not automatically reliable.

What do I actually write?

Weak:

There was a highly significant improvement in wellbeing (p < 0.001).

Better:

Mean wellbeing scores were 0.8 points higher in the intervention group than in the comparison group (95% CI 0.5-1.1; p < 0.001).

Then in interpretation:

Although the estimate was precise, the magnitude was small relative to the prespecified threshold considered practically meaningful.

Or, if no established threshold exists:

The practical importance of the 0.8-point difference is uncertain because a meaningful-change threshold for this population and measure was not prespecified.

That is more scientifically honest than reverse-engineering importance from significance.

Do this now

Take the most statistically significant finding in your study.

Write:

If the p-value disappeared, would I still know whether this effect matters?

Then add:

  • effect magnitude;
  • confidence interval;
  • absolute effect where useful;
  • threshold or context for importance;
  • cost/feasibility implications if relevant.

If the scientific conclusion collapses without the p-value, the interpretation needs more work.

Frequently asked questions

What is the difference between statistical and clinical significance?

Statistical significance concerns the statistical evidence relative to a hypothesis-testing rule. Clinical significance concerns whether the magnitude is meaningful for patients or clinical decisions.

Can a small effect be important?

Yes. Population scale, severity of outcome, cost and feasibility can make a small effect important.

Can a statistically significant result be meaningless?

It can be too small to matter substantively, even if estimated precisely.

Does effect size tell me practical significance automatically?

No. Effect size quantifies magnitude. Whether that magnitude matters requires context.

Should I define a meaningful effect before analysis?

Where possible, yes. Prespecifying what difference would matter reduces the temptation to declare whatever result appears “important” after seeing it.

Try it

Read P-Values, Confidence Intervals and Effect Sizes: How to Read Them Together and apply the Methods Bench Results Interpretation Card.

Open the Manuscript Outline Planner

References and further reading

  1. 1.Wasserstein RL, Lazar NA. The ASA's Statement on p-Values: Context, Process, and Purpose. The American Statistician. 2016.
  2. 2.American Statistical Association President's Task Force. Statement on statistical significance and replicability.
  3. 3.Gardner MJ, Altman DG. Confidence intervals rather than P values. BMJ. 1986.
  4. 4.Relevant discipline-specific literature should be used to justify any minimal important difference or practical threshold used in a study.

Take it further

Better research, once a week.

Practical methods guidance, research tools and funding opportunities from Methods Bench.

Get practical research notes and new opportunities

Short, useful emails for researchers in Uganda and East Africa. No spam.

You can unsubscribe at any time.

Written by Methods Bench. Reviewed by Research Methods Specialist.All research guides