Step 1 of 5
Variables and Normality
Statistical selection depends on data scales and distribution shapes.
Choosing the right test is critical to draw valid conclusions. Learn the parameters of normality, independence, and variable scales that determine statistical matching.
Module complete. It is checked off in your course progress.
Step 1 of 5
Statistical selection depends on data scales and distribution shapes.
Goal
The nature of your research question determines the class of statistical test you need. This is the first and most important branch in any test selection framework (Bahr & Zacks, 2010).
Data scale
For correlation, the scale of measurement and whether normality assumptions are met determine whether to use Pearson's r (parametric) or Spearman's rho (non-parametric). Pearson's r exclusively tests linear relationships between continuous variables; Spearman's rho tests monotonic relationships and is appropriate for ordinal data (Bahr & Zacks, 2010).
Outcome type
The type of outcome (dependent) variable determines which regression model is appropriate. Linear regression requires a continuous outcome with normally distributed residuals; logistic regression is used when the outcome is binary (Field, 2018).
Sample relationship
Whether participants appear in one group only (independent) or are measured at multiple time points or matched with another participant (paired/repeated) is critical: these structures require different tests, and using the wrong one inflates or deflates the p-value (Field, 2018).
Data type
Parametric tests (t-test, ANOVA) assume approximately normal distributions and are more powerful when assumptions hold. Non-parametric alternatives (Mann-Whitney, Kruskal-Wallis) make no normality assumption and should be used for ordinal data or when normality cannot be assumed, especially with small samples (Dwivedi et al., 2021).
Data type
As with independent designs, the choice between parametric (paired t-test, repeated measures ANOVA) and non-parametric tests (Wilcoxon, Friedman) depends on whether the assumptions of normality are met. For paired designs, the normality of the differences between pairs (not raw scores) is the key assumption (Field, 2018).
Groups
The number of groups determines whether a t-test (two groups) or ANOVA (three or more groups) is required. Running multiple t-tests instead of ANOVA inflates the Type I error rate (the probability of a false positive), which is why ANOVA exists (Field, 2018).
Groups
The non-parametric equivalents follow the same group-count logic as parametric tests: Mann-Whitney U for two groups, Kruskal-Wallis for three or more (Bahr & Zacks, 2010).
Time points
For paired/repeated designs, two time points call for a paired t-test, while three or more require repeated measures ANOVA. This mirrors the independent group logic but within the same participants (Field, 2018).
Time points
For non-parametric paired designs: the Wilcoxon signed-rank test is used for two time points, and the Friedman test is used for three or more — the non-parametric counterparts of the paired t-test and repeated measures ANOVA respectively (Bahr & Zacks, 2010).
Sample size
The chi-square test relies on a large-sample approximation and becomes unreliable when expected cell frequencies fall below 5. In those cases, Fisher's Exact Test provides an exact p-value and should be used instead (Bahr & Zacks, 2010).
Correlation
Also known as: Pearson's r
Measures the strength and direction of a linear relationship between two continuous variables. Produces a coefficient (r) ranging from -1 (perfect negative) to +1 (perfect positive), with 0 indicating no linear relationship.
Data type
Continuous (interval/ratio)
Sample relationship
Independent
Normality required
Yes
Key assumptions
Strengths
Limitations
Example in practice
A researcher measures hours of sleep and exam scores in 80 students and uses Pearson's r to test whether more sleep is linearly associated with higher scores.
Sources: Bahr, S. & Zacks, J. (2010). Choosing Statistical Tests. Deutsches Arzteblatt International, 107(19), 343–348. PMC2881615. | Field, A. (2018). Discovering Statistics Using IBM SPSS Statistics (5th ed.). SAGE.
Correlation
Also known as: Spearman's rho (ρ)
A non-parametric measure of the monotonic relationship between two variables. Works by ranking the data and computing correlation on ranks. Suitable when data are ordinal or when the normality assumption is violated.
Data type
Ordinal or non-normal continuous
Sample relationship
Independent
Normality required
No
Key assumptions
Strengths
Limitations
Example in practice
A researcher asks 60 students to rate their research confidence on a 1-5 Likert scale and records their number of publications. Spearman's rho is used because confidence ratings are ordinal.
Sources: Bahr & Zacks (2010). PMC2881615. | Dwivedi, A.K. et al. (2021). Comparing statistical tests. Journal of Family Medicine and Primary Care. | Field, A. (2018). Discovering Statistics (5th ed.). SAGE.
Regression
Also known as: Ordinary Least Squares (OLS) regression
Models the relationship between one continuous predictor (independent variable) and one continuous outcome (dependent variable). Estimates how much the outcome changes per unit increase in the predictor, and allows prediction.
Data type
Continuous outcome
Sample relationship
Independent
Normality required
Yes (residuals)
Key assumptions
Strengths
Limitations
Example in practice
A researcher uses simple linear regression to predict a student's research output score from the number of hours per week spent in lab, producing a slope estimate and R² value.
Sources: Field, A. (2018). Discovering Statistics Using IBM SPSS Statistics (5th ed.). SAGE. | Bahr & Zacks (2010). Deutsches Arzteblatt International. PMC2881615.
Regression
Also known as: Binary logistic regression
Models the probability of a binary outcome (yes/no, success/failure) as a function of one or more predictor variables. Produces odds ratios indicating how much each predictor changes the odds of the outcome.
Data type
Binary categorical outcome
Sample relationship
Independent
Normality required
No
Key assumptions
Strengths
Limitations
Example in practice
A researcher uses logistic regression to model whether a student completes their research project (yes/no) based on number of mentorship hours, year of study, and prior research experience.
Sources: Field, A. (2018). Discovering Statistics (5th ed.). SAGE. | Hosmer, D.W. & Lemeshow, S. (2000). Applied Logistic Regression (2nd ed.). Wiley.
Categorical comparison
Also known as: Pearson’s χ² test of independence
Tests whether two categorical variables are statistically independent (i.e. whether the distribution of one variable differs across categories of another). Compares observed frequencies in a contingency table to expected frequencies under independence.
Data type
Categorical (nominal)
Sample relationship
Independent
Normality required
No
Key assumptions
Strengths
Limitations
Example in practice
A researcher tests whether the proportion of students who complete a research project differs between those who received mentoring and those who did not, using a 2×2 contingency table.
Sources: Bahr & Zacks (2010). PMC2881615. | Dwivedi, A.K. et al. (2021). Journal of Family Medicine and Primary Care. | Field, A. (2018). Discovering Statistics (5th ed.). SAGE.
Categorical comparison
Alternative to chi-square for small samples
An exact test of independence for 2×2 contingency tables. Calculates the exact probability of observing a table as extreme as the one obtained, given fixed marginal totals. Preferred over chi-square when expected cell counts are below 5.
Data type
Categorical (nominal), small samples
Sample relationship
Independent
Normality required
No
Key assumptions
Strengths
Limitations
Example in practice
A researcher has only 18 participants and tests whether those who attended a workshop differ from those who did not in whether they submitted a manuscript. Expected cells are below 5, so Fisher’s Exact Test is used.
Sources: Bahr & Zacks (2010). PMC2881615. | Mehta, C.R. & Patel, N.R. (1983). A network algorithm for performing Fisher's exact test. Journal of the American Statistical Association, 78, 427–434.
Group comparison
Also known as: Student’s t-test, unpaired t-test
Compares the means of a continuous outcome variable between two independent (unrelated) groups to determine whether the difference is statistically significant.
Data type
Continuous
Sample relationship
Independent
Normality required
Yes
Key assumptions
Strengths
Limitations
Example in practice
A researcher compares mean research anxiety scores between students in a mentorship programme and those not in one, using an independent samples t-test with Levene’s test for equality of variances.
Sources: Bahr & Zacks (2010). PMC2881615. | Field, A. (2018). Discovering Statistics (5th ed.). SAGE. | Dwivedi, A.K. et al. (2021). Journal of Family Medicine and Primary Care.
Group comparison
Also known as: Dependent t-test, repeated measures t-test
Compares the means of a continuous variable measured twice on the same participants (e.g. pre- and post-intervention), or on matched pairs. Tests whether the mean difference between paired measurements is significantly different from zero.
Data type
Continuous
Sample relationship
Paired/repeated
Normality required
Yes
Key assumptions
Strengths
Limitations
Example in practice
A researcher measures students’ self-reported research confidence before and after a 10-week mentorship programme, then applies a paired t-test to determine whether the intervention produced a significant change.
Sources: Bahr & Zacks (2010). PMC2881615. | Field, A. (2018). Discovering Statistics (5th ed.). SAGE.
Group comparison
Also known as: Wilcoxon rank-sum test
A non-parametric test that compares the distributions (medians) of a continuous or ordinal variable between two independent groups by ranking all observations and comparing rank sums. Used when the independent t-test assumptions are not met.
Data type
Ordinal or non-normal continuous
Sample relationship
Independent
Normality required
No
Key assumptions
Strengths
Limitations
Example in practice
A researcher compares student satisfaction scores (1–10 scale) between two lab environments. Because scores are unlikely to be normally distributed, the Mann-Whitney U test is used instead of a t-test.
Sources: Bahr & Zacks (2010). PMC2881615. | Field, A. (2018). Discovering Statistics (5th ed.). SAGE. | Dwivedi, A.K. et al. (2021). Journal of Family Medicine and Primary Care.
Group comparison
Non-parametric alternative to the paired t-test
A non-parametric test that compares two related samples or repeated measurements. Ranks the absolute differences between paired observations and tests whether the distribution of positive and negative ranks is symmetric around zero.
Data type
Ordinal or non-normal continuous
Sample relationship
Paired/repeated
Normality required
No
Key assumptions
Strengths
Limitations
Example in practice
A researcher measures lab skill ratings (ordinal scale) before and after training. Because the differences are unlikely to be normally distributed with a small sample, the Wilcoxon signed-rank test is used.
Sources: Bahr & Zacks (2010). PMC2881615. | Field, A. (2018). Discovering Statistics (5th ed.). SAGE.
Group comparison
Also known as: Analysis of Variance
Compares the means of a continuous outcome variable across three or more independent groups simultaneously. Tests whether at least one group mean differs significantly from the others. Post-hoc tests (e.g. Tukey’s HSD) are needed to identify which specific groups differ.
Data type
Continuous
Sample relationship
Independent
Normality required
Yes
Key assumptions
Strengths
Limitations
Example in practice
A researcher compares mean publication counts across three groups of students: those with no mentor, a near-peer mentor, or a faculty mentor. One-way ANOVA tests whether group means differ, followed by Tukey’s HSD.
Sources: Field, A. (2018). Discovering Statistics (5th ed.). SAGE. | Bahr & Zacks (2010). PMC2881615.
Group comparison
Non-parametric alternative to one-way ANOVA
A non-parametric test that compares the distributions of a continuous or ordinal variable across three or more independent groups using ranked data. The non-parametric counterpart of one-way ANOVA.
Data type
Ordinal or non-normal continuous
Sample relationship
Independent
Normality required
No
Key assumptions
Strengths
Limitations
Example in practice
A researcher compares research anxiety scores across four different lab environments. Scores are skewed and sample sizes are small, so the Kruskal-Wallis test is used, followed by Dunn’s post-hoc test.
Sources: Bahr & Zacks (2010). PMC2881615. | Field, A. (2018). Discovering Statistics (5th ed.). SAGE. | Dwivedi, A.K. et al. (2021). Journal of Family Medicine and Primary Care.
Group comparison
Also known as: Within-subjects ANOVA
Compares the means of a continuous outcome measured at three or more time points (or conditions) on the same participants. Accounts for the correlation between repeated measurements on the same individual.
Data type
Continuous
Sample relationship
Paired/repeated
Normality required
Yes
Key assumptions
Strengths
Limitations
Example in practice
A researcher measures student research self-efficacy at baseline, 4 weeks, and 8 weeks during a training programme. Repeated measures ANOVA tests whether scores change significantly across the three time points.
Sources: Field, A. (2018). Discovering Statistics (5th ed.). SAGE. | Girden, E.R. (1992). ANOVA: Repeated Measures. SAGE.
Group comparison
Non-parametric alternative to repeated measures ANOVA
A non-parametric test for comparing distributions of an ordinal or continuous variable across three or more related conditions or time points. The non-parametric counterpart of repeated measures ANOVA.
Data type
Ordinal or non-normal continuous
Sample relationship
Paired/repeated
Normality required
No
Key assumptions
Strengths
Limitations
Example in practice
A researcher asks 20 students to rate three different data analysis software packages on a 1–5 scale. The Friedman test compares whether median ratings differ significantly across the three packages.
Sources: Field, A. (2018). Discovering Statistics (5th ed.). SAGE. | Bahr & Zacks (2010). PMC2881615.