Proposal Development

How to Choose Variables for a Quantitative Study Before You Open SPSS

Methods BenchReviewed by Research Methods SpecialistPublished 26 August 2026Last reviewed 26 August 20268 min read
Direct answer

Choose variables by starting with the research question: define the outcome, the exposure or comparison of interest and the quantity you want to estimate. Then use theory, prior evidence and a plausible causal structure to identify variables needed for measurement, description, confounding control, mediation or effect modification. Do not automatically adjust for every available variable or select covariates only because they were statistically significant in bivariate analysis.

Age.

Sex.

Marital status.

Education.

Occupation.

Religion.

Income.

Then, somewhere after the demographic census:

“What was the research question again?”

If your variable list existed before the question, we may already have a problem.

A quantitative study should not begin by collecting every variable researchers traditionally collect.

It should begin with what you are trying to estimate.

Step 1: start with the research question

Suppose the question is:

Is distance to an adolescent-friendly clinic associated with use of the service?

Now two variables are immediately central:

Exposure: distance to the clinic.

Outcome: service use.

But even these need work.

What does “distance” mean?

  • straight-line kilometres?
  • road distance?
  • travel time?
  • transport cost?
  • perceived accessibility?

What does “service use” mean?

  • ever used?
  • used in the last 12 months?
  • number of visits?
  • used a particular service?

A concept is not yet a variable until you decide how it will be observed.

Bench rule

Start with the question. Then build the variables needed to answer it.

Not the other way around.

Step 2: decide what quantity you actually want to estimate

Are you trying to estimate:

  • prevalence?
  • a mean?
  • a difference between groups?
  • an exposure-outcome association?
  • a total causal effect?
  • a direct effect?
  • a prediction?

That choice changes what variables belong in the model.

Suppose your objective is:

To estimate the association between distance to the clinic and service use.

That is already different from:

To estimate the total causal effect of distance on service use.

And different again from:

To predict which adolescents will use the service.

Association, causal explanation and prediction are not interchangeable analytic goals.

Your variable-selection strategy should know which job it is doing.

Step 3: draw the plausible causal story before choosing adjustment variables

Suppose:

Exposure: distance to clinic
Outcome: service use

Potential variables include:

  • residential setting;
  • household income;
  • age;
  • perceived accessibility;
  • transport cost;
  • service awareness;
  • school attendance;
  • provider attitude.

Do not ask first:

“Which ones have p < 0.20?”

Ask:

How might these variables relate causally to the exposure and outcome?

For example:

Residential setting may influence where clinics are located and independently influence service use through transport, service availability or other pathways.

It could therefore be part of a confounding structure.

Perceived accessibility may partly result from actual distance and then influence service use.

It may be a mediator.

Those roles matter.

If your aim is to estimate the total effect of distance, automatically adjusting for perceived accessibility could remove part of the pathway you are trying to estimate.

On the bench

why “adjust for everything” is not a neutral strategy

Imagine this simplified causal story:

Residential setting → distance to clinic → perceived accessibility → service use

And:

Residential setting → service use

You want the total effect of distance on service use.

A reasonable causal analysis may need to address residential setting because it is a common cause of the exposure and outcome.

But if you also adjust for perceived accessibility, you may block part of the effect operating through accessibility.

Now you are no longer estimating the same quantity.

The model changed the question.

Schisterman and colleagues describe this type of problem as overadjustment when a variable on the causal pathway is controlled while estimating a total effect.

Bench check

Your model includes:

age, sex, marital status, education, income, knowledge, attitude, distance, waiting time, provider attitude, service availability...

Why?

“Because they were available.”

Availability is not a causal rationale.

A regression model is not improved automatically by collecting covariates like Pokémon.

Step 4: distinguish variables needed for different purposes

Your dataset may contain variables for several legitimate reasons.

Outcome variables

What is being estimated or explained?

Exposure or comparison variables

What condition, characteristic or group is central to the question?

Confounding-control variables

Variables needed to address confounding for a causal estimand, based on the assumed causal structure.

Mediators

Variables that may lie on a pathway from exposure to outcome.

These may be scientifically interesting.

But adjusting for them changes the estimand.

Effect modifiers

Variables across which the exposure-outcome relationship may genuinely differ.

This is a question about heterogeneity of effect, not confounding.

Descriptive variables

Characteristics useful for describing the sample or context.

They do not all need to enter the adjusted model.

Design variables

Strata, clusters, sampling weights, matching factors or randomization variables may need special analytic treatment because of the study design.

One variable can play different roles in different causal questions.

Step 5: operationalize every construct

Suppose your conceptual framework contains:

Socioeconomic status

That is not yet a variable.

Will you measure it using:

  • household income?
  • education?
  • occupation?
  • household assets?
  • an index?
  • area-level deprivation?

Each choice represents something slightly different.

Or suppose your framework contains:

Quality of care

Will you measure:

  • structural readiness?
  • client experience?
  • respectful care?
  • technical competence?
  • continuity?
  • method availability?
  • a validated multi-item measure?

A beautifully drawn conceptual box can hide a weak measurement decision.

Bench rule

A construct becomes useful only when you can explain how the data will represent it.

Step 6: use the literature for causal and measurement reasoning

The literature should help answer:

  • What variables consistently precede the exposure?
  • What variables cause both exposure and outcome?
  • What mechanisms are proposed?
  • Which measures have been validated?
  • Which variables are likely consequences of the exposure?
  • What effect modifiers are substantively plausible?
  • What design-specific factors need handling?

This is different from reading ten papers and copying every covariate they adjusted for.

Previous models can themselves be overadjusted or poorly justified.

Use prior studies as evidence.

Not as a shopping list.

Why selecting variables by bivariate p-values is risky

A common workflow is:

  1. Run bivariate analyses.
  2. Keep variables with p < 0.20.
  3. Put them all into multivariable regression.
  4. Call the remaining variables “independent predictors.”

This can fail because statistical association in your sample does not define causal role.

A true confounder could show a weak bivariate association because of sampling variability.

A mediator could show a strong association and be included even though controlling it changes the causal question.

A collider could be opened through inappropriate conditioning.

A variable-selection threshold is not a substitute for substantive knowledge.

For purely predictive modelling, variable selection raises a different set of issues.

Do not import causal language into a model whose purpose was prediction.

What about age and sex? Should I always adjust for them?

No variable is an automatic confounder in every analysis.

Age and sex can be important in many causal structures.

But their role depends on the question.

Ask:

Could this variable be a cause of the exposure and the outcome, directly or through relevant pathways?

Or:

Is it included because the design, estimand or substantive question requires it?

“Everyone adjusts for age and sex” is not a complete causal justification.

What do I actually write?

What do I actually write in my methods section?

Weak:

Variables with p < 0.20 at bivariate analysis were entered into the multivariable model.

If that was genuinely your procedure, report it.

But recognize its limitations.

A stronger causal-analysis description might be:

Potential confounders were identified a priori using subject-matter knowledge, previous evidence and an assumed causal structure linking clinic distance to service use. The primary adjusted model included [variables] because they were considered common causes of the exposure and outcome. Variables hypothesized to lie on the causal pathway, including perceived accessibility, were not included in the model estimating the total effect.

Only write that if you actually did the causal reasoning.

For a descriptive or predictive model, use language appropriate to that purpose.

Build a variable specification table before the questionnaire

Use:

ConceptRole in this questionOperational definitionSource/instrumentCodingPlanned analysis
Service useOutcomeUsed service in past 12 monthsSurvey0/1Logistic/binomial model as appropriate
Clinic distanceExposureTravel distance in kmGIS/self-reportContinuousPrimary exposure
Residential settingPotential confounderUrban/peri-urban/ruralSurveyCategoriesAdjustment
Perceived accessibilityPotential mediatorScale scoreValidated/adapted scaleContinuousSeparate mediation/descriptive analysis

This table exposes missing decisions early.

It is much cheaper to discover that “quality” has no measure before data collection.

Do this now

Take your planned adjusted model.

Write every covariate on a separate sticky note or line.

For each one, complete:

This variable is in the model because...

Acceptable answers might include:

  • plausible common cause of exposure and outcome;
  • design variable;
  • prespecified effect modifier;
  • part of a prediction model for a defined reason.

If the answer is:

“because p < 0.20”

add one more sentence explaining its substantive role.

If you cannot, reconsider the model.

Frequently asked questions

How do I identify independent and dependent variables?

Start with the research question. The outcome is what you are trying to estimate or explain; the exposure or comparison is the central factor whose relationship with that outcome you are examining.

Should every demographic variable go into regression?

No. Describing the sample and adjusting an effect estimate are different tasks.

Should I include variables with p < 0.20?

A threshold can be part of some modelling strategies, but it should not be treated as a general method for identifying confounders. Causal adjustment requires substantive reasoning.

What is a control variable?

The term is broad and often unclear. State why the variable is included: confounding control, design, precision, prediction, mediation, effect modification or another purpose.

Do I need a DAG?

Not every project needs a formal DAG, but drawing a causal diagram can force assumptions into the open and reveal inappropriate adjustment choices.

Try it

Use the Methods Bench Conceptual Framework Worksheet and add one new column:

Role in the analysis

The longer-term Methods Bench tool should turn this into a Variable & Adjustment Planner.

Open the Methods Bench tools

References and further reading

  1. 1.Schisterman EF, Cole SR, Platt RW. Overadjustment bias and unnecessary adjustment in epidemiologic studies. Epidemiology. 2009;20(4):488-495. PMCID: PMC2744485.
  2. 2.Greenland S, Pearl J, Robins JM. Causal diagrams for epidemiologic research. Epidemiology. 1999;10(1):37-48.
  3. 3.Hernán MA, Robins JM. Causal Inference: What If. Chapters on causal diagrams, exchangeability and adjustment.
  4. 4.World Health Organization. Recommended format for a research protocol: measurements, data management and statistical analysis.
  5. 5.Methods Bench. Conceptual Framework and research-objective guides.

Take it further

Better research, once a week.

Practical methods guidance, research tools and funding opportunities from Methods Bench.

Get practical research notes and new opportunities

Short, useful emails for researchers in Uganda and East Africa. No spam.

You can unsubscribe at any time.

Written by Methods Bench. Reviewed by Research Methods Specialist.All research guides