What Does “Statistically Significant” Actually Mean?

Research explained
Statistics
Research literacy
Why a statistically significant result is not automatically important, large, useful, or true.
Published

August 11, 2026

You have probably seen headlines like these:

“Researchers found a statistically significant improvement.”

“The difference between the two groups was statistically significant.”

Or perhaps:

“The treatment significantly reduced the risk.”

It sounds impressive. It may even sound like scientists have discovered something important.

But there is a problem.

“Statistically significant” does not mean “important”, “large”, or even “useful”.

It means something much more specific.

And understanding that difference can completely change how you read research.

Let’s start with a simple example

Imagine researchers want to test a new exercise programme.

They recruit two groups of people:

  • Group A follows the new exercise programme.
  • Group B continues with their usual routine.

After three months, Group A improves its fitness score by an average of 5 points, while Group B improves by 3 points.

So there is a difference of 2 points.

The obvious question is:

Did the exercise programme really make a difference, or could that difference have appeared just by chance?

This is where statistical significance comes in.

What does statistical significance mean?

Researchers usually begin with what is called a null hypothesis.

In simple terms, the null hypothesis says:

There is no real difference between the groups.

The researchers then look at their data and ask:

If there really were no difference, how surprising would these results be?

One common way of answering this involves something called a p-value.

You may have seen results written like this:

p < .05

Traditionally, researchers often describe a result as statistically significant when the p-value is below .05.

But what does that actually mean?

Very roughly, it means that the observed data would be relatively unusual if the statistical model assumed no real effect.

Importantly, it does not mean:

“There is a 95% chance that our theory is correct.”

And it does not mean:

“There is only a 5% chance that this result happened by chance.”

Those are common interpretations, but they are not what a p-value tells us.

A simple way to think about it

Imagine flipping a coin.

If the coin is fair, you expect roughly half heads and half tails.

Suppose you flip it 10 times and get 6 heads.

That does not seem particularly strange.

But suppose you flip it 100 times and get 95 heads.

Now you might start wondering whether the coin is actually fair.

Statistical testing works with a similar idea.

Researchers ask:

How compatible are the results we observed with a world in which there is no effect?

The more unusual the data would be under that assumption, the smaller the p-value tends to be.

So why is p < .05 so common?

The .05 threshold is mostly a convention.

If a researcher obtains:

p = .03

the result may be described as statistically significant.

If they obtain:

p = .06

it may be described as not statistically significant.

But notice how strange that can become.

There is very little practical difference between .049 and .051.

Yet one may receive the label “significant”, while the other does not.

Nature does not suddenly change at p = .05.

The threshold is a human decision, not a magical boundary separating truth from falsehood.

Statistical significance is not the same as importance

This is probably the most important point.

Imagine a study involving 100,000 people finds that a medication reduces average blood pressure by 0.5 mmHg.

Because the sample is enormous, the difference might be statistically significant.

But would half a millimetre of mercury meaningfully improve someone’s health?

Perhaps not.

Now imagine another small study finds an average reduction of 10 mmHg, but because it only includes 20 people, the result does not reach statistical significance.

Which finding is more important?

You cannot answer that question simply by looking at the p-value.

That is why researchers also need to consider things such as:

  • the size of the effect
  • the confidence interval
  • the sample size
  • the quality of the study
  • whether the result makes sense in the real world
  • whether other studies have found something similar

A statistically significant result can be tiny.

A non-significant result can still be potentially important.

Relative risk can also make findings sound bigger

Here is another common problem.

Imagine that the risk of a particular side effect increases from:

2 people in every 10,000

to:

3 people in every 10,000

Someone could describe this as a:

50% increase in risk.

That sounds dramatic.

And mathematically, it is correct.

But the absolute increase is only:

1 additional case per 10,000 people.

Both pieces of information matter.

Saying only “a 50% increase” can make a very small difference sound enormous.

This is why good research communication should ideally tell us both the relative and the absolute difference.

Sample size matters too

Imagine tossing a coin four times.

You get:

Heads
Heads
Heads
Heads

Does that prove the coin is biased?

Probably not.

Four flips simply do not give us much information.

Now imagine flipping the same coin 100,000 times and getting heads 80% of the time.

That would be much harder to dismiss.

Research works in a similar way.

Larger samples usually allow researchers to estimate effects more precisely and detect smaller differences.

But this creates an interesting consequence:

With a very large sample, even a tiny and practically meaningless difference can become statistically significant.

So “statistically significant” should never automatically be translated into:

“This effect is big.”

What about a result that is not statistically significant?

This is another area where research is often misunderstood.

Suppose a study reports:

p = .08

It is tempting to conclude:

The treatment does not work.

But that conclusion may be too strong.

Perhaps the treatment truly has no effect.

But perhaps the study was too small.

Perhaps the measurements were noisy.

Perhaps the effect exists but is smaller than researchers expected.

Perhaps the confidence interval includes several plausible possibilities.

A non-significant result does not automatically prove that there is no effect.

It means the study did not provide sufficiently strong statistical evidence, under the particular analysis used, to cross the chosen threshold.

That is a much more careful statement.

The difference between “statistically significant” and “clinically significant”

This distinction becomes especially important in medicine, psychology, rehabilitation, sport and health research.

Suppose a therapy improves a questionnaire score by 0.3 points.

The result might be statistically significant.

But would a patient actually notice a difference?

If not, the result may have limited clinical significance.

The same principle applies outside healthcare.

An educational programme might produce a statistically significant improvement in exam scores, but if the average improvement is only 0.2%, nobody is likely to redesign the education system because of it.

Researchers therefore need to ask two different questions:

Is there evidence that an effect exists?

and

Is the effect large enough to matter?

Those questions are related, but they are not the same.

What should you look for instead?

When you see the words “statistically significant”, do not stop reading.

Ask a few more questions.

How large was the effect?

A tiny difference can still be statistically significant.

Look for effect sizes, actual percentages, averages or differences between groups.

How many people were included?

Results from 20 participants and results from 20,000 participants should not automatically be interpreted in the same way.

What is the confidence interval?

Confidence intervals help show how precise an estimate is.

A very wide confidence interval usually suggests considerable uncertainty.

Was this the main outcome?

Large studies sometimes test dozens or even hundreds of relationships.

If enough tests are performed, some may appear statistically significant simply because of random variation.

Has the result been replicated?

One study rarely settles a scientific question.

Findings become much more convincing when independent researchers repeatedly observe similar results.

Does the difference matter in real life?

Perhaps the most important question of all.

Even if the result is statistically convincing, ask:

Would this difference actually matter to a person, patient, athlete, organisation or society?

Why does this matter?

Because the phrase “statistically significant” is incredibly easy to misuse.

Consider two headlines:

“Drug associated with a statistically significant increase in condition X.”

and

“Drug associated with one additional case of condition X per 100,000 people.”

Both could potentially describe the same study.

But they create completely different impressions.

Numbers do not speak for themselves.

How we present them matters.

And when statistical terminology leaves academic papers and enters newspapers, social media or political debates, important details can disappear very quickly.

Statistical significance is evidence, not a verdict

Science rarely gives us perfect certainty.

Instead, researchers gradually accumulate evidence.

A p-value can be one part of that evidence.

But it should not be treated as a scientific traffic light:

🟢 p < .05 = true

🔴 p > .05 = false

Real research is more complicated than that.

A good interpretation considers the size of the effect, uncertainty, methodology, sample, previous evidence and whether the finding actually matters outside the spreadsheet.

So the next time you read:

“The result was statistically significant…”

your next question should be:

“Okay. But how big was the difference?”

And after that:

“Does it actually matter?”

Those two questions will often tell you far more than the word significant ever could.

Back to top