Sampling & Study Design

How to Calculate Sample Size for a Cross-Sectional Study

Methods BenchReviewed by Research Methods SpecialistPublished 10 August 2026Last reviewed 10 August 20268 min read
Direct answer

For a cross-sectional study estimating a single proportion, sample size is commonly calculated from the expected proportion, desired confidence level and absolute precision. With 95% confidence, 5% precision and an expected proportion of 50%, the initial estimate is about 385 participants. But that number may need adjustment for finite populations, cluster sampling and non-response, and it is not appropriate for every cross-sectional research question.

“You need 384 participants. Use Cochran.”

That sentence has appeared in so many proposals that 384 has almost become a research tradition.

The problem is not the formula.

The problem is using a number without knowing what question it was calculated to answer.

First: what are you trying to estimate?

“Cross-sectional study” tells us something about the design.

It does not tell us enough to choose a sample-size formula.

A cross-sectional study might aim to:

  • estimate the prevalence of depression;
  • estimate the proportion of facilities offering a service;
  • compare contraceptive use between two groups;
  • assess an exposure-outcome association;
  • estimate a mean score;
  • fit a multivariable regression model.

Those are not necessarily the same sample-size problem.

This article focuses first on one of the most common cases:

A cross-sectional study designed to estimate a single population proportion with specified precision.

That is the situation where the familiar Cochran-style formula is often appropriate under a simple random sampling framework.

What is the basic formula for a prevalence estimate?

For a large population:

Where:

  • n₀ = initial required sample size
  • Z = critical value for the chosen confidence level
  • p = anticipated proportion or prevalence
  • 1 − p = the complementary proportion
  • d = desired absolute precision, often called the margin of error

For a 95% confidence level, Z ≈ 1.96.

On the bench

where does 384 come from?

Suppose you are planning a cross-sectional survey to estimate a prevalence.

You choose:

  • Confidence level: 95%
  • Precision: ±5 percentage points
  • Expected prevalence: unknown

When prevalence is genuinely unknown, p = 0.50 is often used because it produces the largest sample size for these other inputs.

So:

Round upward:

n = 385

That is the famous number.

Not magic.

Just assumptions.

Bench check

You wrote:

“A sample size of 384 was obtained using Cochran's formula.”

Bench check: That is incomplete.

A reviewer still cannot tell:

  • what you were estimating;
  • what confidence level you used;
  • what precision you wanted;
  • what proportion you assumed;
  • whether the sampling design was simple random;
  • whether you adjusted for non-response.

The number is not the justification. The assumptions are.

What if I know the expected prevalence?

Suppose previous credible research suggests that about 20% of the target population has the characteristic you are studying.

Use:

  • p = 0.20
  • 1 − p = 0.80
  • Z = 1.96
  • d = 0.05

This gives approximately:

n = 246

The estimate is smaller than when p = 0.50 because a binary proportion has its greatest variance around 50%, all else equal.

Do not automatically use 50% simply because it gives a familiar number. If a credible prior estimate exists and is relevant to your population and outcome, using it may be defensible.

What if the total population is small?

The basic large-population calculation assumes that the target population is sufficiently large that sampling a few hundred people does not substantially deplete it.

Suppose your total eligible population is only 1,000 people.

Using the initial estimate of about 385, you can apply a finite population correction:

Where N is the population size.

With n₀ ≈ 385 and N = 1,000:

n ≈ 278

That is because observing 278 of 1,000 eligible people gives you more information about that finite population than observing 278 people from a population of several million.

How do I adjust for non-response?

Suppose the required completed sample is 278, but you expect only 90% of selected participants to provide usable data.

Divide by the anticipated response proportion:

You would need to approach approximately:

309 people

The distinction matters:

  • required completed sample = 278
  • number to approach = 309, given a 90% response assumption

Do not simply “add 10%” without checking what that does mathematically. Dividing by the anticipated response rate is clearer.

What if I am using cluster sampling?

This is where many sample-size calculations become too optimistic.

The simple formula above corresponds to a simple random sampling framework. If you sample clusters, for example, villages, schools, facilities or enumeration areas, observations within clusters may be more similar to one another than observations selected independently across the whole population.

That reduces the effective information in the sample.

A design effect may therefore be needed when calculating the required sample size.

For example, if the simple-random-sample requirement were 385 and a justified design effect were 1.5:

You would then consider non-response and other design features as appropriate.

But do not choose a design effect because “1.5 is common.” It should be based on credible prior information or a defensible planning assumption.

And remember: if the design is clustered, the analysis must also account for the clustering. Inflating the sample size does not make a clustered design analytically simple random.

Does every cross-sectional study use Cochran's formula?

No.

This is one of the most important things to understand.

If your main objective is:

“To estimate the prevalence of hypertension with ±5% precision”

a single-proportion precision formula may be appropriate.

But if your main objective is:

“To compare hypertension prevalence between men and women”

you need a calculation suited to comparing groups.

If your objective is:

“To assess whether obesity is associated with hypertension”

you may need a power calculation based on an expected effect measure, exposure distribution and model.

If your outcome is a continuous mean rather than a proportion, a different calculation is needed.

Bench rule

Choose the sample-size method from the primary objective, not from the words “cross-sectional study.”

The design label does not choose the estimand for you.

A practical sample-size workflow

Before opening a calculator, write down these questions:

1. What is my primary objective?

Are you estimating one proportion, comparing groups, estimating a mean or testing an association?

2. What quantity do I want to estimate?

Prevalence? Mean? Risk difference? Odds ratio? Another parameter?

3. What confidence or power criterion am I using?

For precision-based calculations this may be a confidence level and margin of error. For hypothesis-testing calculations, power and effect size become central.

4. What prior information do I have?

Expected prevalence, standard deviation, anticipated effect or intracluster correlation may come from earlier studies or pilot data.

5. What is my sampling design?

Simple random, stratified, systematic, clustered or multistage?

6. What loss should I anticipate?

Non-response, missing records or attrition, depending on the design.

Only then should you calculate the number.

What do I actually write?

What do I actually write in my protocol?

If you are estimating a single proportion, a defensible example might be:

The sample size was calculated to estimate a single population proportion with 95% confidence and 5-percentage-point absolute precision. Because no reliable prior estimate of the prevalence was available, a proportion of 50% was used, producing an initial sample of 385 participants under a simple random sampling assumption. The sample was subsequently adjusted for [finite population/design effect/non-response, as applicable].

Do not copy the bracketed adjustments unless they actually apply to your design.

If your objective is comparison or association, use wording that reflects the calculation you actually performed.

Do this now

Look at the sample-size paragraph in your proposal.

Can you identify, without guessing:

  • the primary outcome or parameter;
  • the assumed prevalence/effect;
  • confidence level or power;
  • precision or effect size;
  • sampling design;
  • design effect, if relevant;
  • non-response allowance?

If not, your calculation may be numerically correct but methodologically under-explained.

Frequently asked questions

Why do so many studies get a sample size of 384?

Using 95% confidence, 5% absolute precision and p = 0.50 gives approximately 384.16 before rounding. Those assumptions, not the topic of the study, produce the number.

Should I always use 50% if prevalence is unknown?

It is a common conservative choice because it maximizes the sample requirement for the other inputs held constant. If relevant, credible prior evidence exists, another value may be more defensible.

Should I add 10% for non-response?

You can plan for non-response, but dividing the required completed sample by the expected response proportion is clearer than simply adding a percentage. For a required sample of 278 and a 90% expected response rate, 278/0.90 ≈ 309.

Does a larger sample fix a poor sampling method?

No. A very large convenience sample is still a convenience sample. Sample size and sampling quality are different issues.

Can I use this formula for a cluster randomized trial?

No. Cluster trials require calculations that account explicitly for the number and size of clusters, intracluster correlation, effect size and the trial design.

Try it

Use the Methods Bench Sample Size Calculator to enter your own confidence level, expected proportion, population size, design effect and anticipated response.

The calculator should support your reasoning, not replace it.

Open the Sample Size Calculator

References and further reading

  1. 1.OpenEpi. Sample Size for a Proportion or Descriptive Study. https://www.openepi.com/SampleSize/SSPropor.htm
  2. 2.OpenEpi documentation: sample size for proportions. https://www.openepi.com/PDFDocs/SSProporDoc.pdf
  3. 3.World Health Organization. WHO STEPS Surveillance Manual: Preparing the Sample. https://cdn.who.int/media/docs/default-source/ncds/ncd-surveillance/steps/part2-section2.pdf
  4. 4.CDC/NCHS. Guidance on complex survey design and analysis. https://wwwn.cdc.gov/nchs/nhanes/continuousnhanes/overviewbrief.aspx
  5. 5.Methods Bench. Sample Size Calculator. `/tools/sample-size-calculator`

Take it further

Better research, once a week.

Practical methods guidance, research tools and funding opportunities from Methods Bench.

Get practical research notes and new opportunities

Short, useful emails for researchers in Uganda and East Africa. No spam.

You can unsubscribe at any time.

Written by Methods Bench. Reviewed by Research Methods Specialist.All research guides