Table of Contents
Type: Statistical Analysis | Subject: Statistics (SPSS) | Level: Masters | Word Count: ~2800 words
This model statistical analysis report was produced by an Essays UK specialist as reference material for learning purposes only. For support in this field, see our psychology statistics support team.
A postgraduate research methods module has provided you with exam data from a cohort of students taught via blended (online plus in-person) delivery and a cohort taught via traditional, wholly in-person delivery. Using SPSS, conduct and report an independent-samples t-test comparing final exam performance between the two study modes, including assumption checks, and interpret the result in APA style. (2,800 words)
Following the expansion of blended delivery across postgraduate provision, a university module leader wished to establish whether the mode of study was associated with differences in student attainment. Blended delivery combines recorded and live online sessions with occasional in-person seminars, whereas the traditional comparator group attended all sessions in person. This is a live methodological question in the higher education literature: blended designs are frequently promoted for flexibility and access, but evidence on their effect on measurable attainment remains mixed, which is precisely why a controlled statistical comparison within a single module and assessment is valuable.
The research question addressed in this analysis is: is there a statistically significant difference in final exam performance between students who studied via blended delivery and students who studied via traditional, wholly in-person delivery? The null hypothesis (H0) states that the population mean exam score does not differ between the two study modes. The alternative hypothesis (H1) states that the population mean exam score does differ between the two study modes. Because the hypothesis is non-directional, a two-tailed test is appropriate.
An independent-samples t-test was selected because the independent variable, study mode, is a between-subjects categorical variable with two levels, and the dependent variable, final exam score, is measured on a continuous (interval-level) scale. Each student contributed a single score to only one group, satisfying the independence-of-observations requirement that distinguishes this design from a paired-samples test.
The wider literature on blended learning offers no settled answer to the research question posed here. A frequently cited meta-analysis by Means et al. (2013) found that, on average, blended conditions outperformed purely face-to-face instruction across a range of higher-education contexts, but the authors were careful to note that the studies pooled varied enormously in the specific mix of online and in-person activity, meaning the average effect masks considerable underlying variation. This existing evidence base motivates the present single-module comparison: rather than relying on a pooled effect drawn from heterogeneous contexts, this analysis isolates the effect of study mode within one cohort, one assessment and one marking scheme, at the cost of the generalisability that a multi-site meta-analysis can offer. The significance threshold for this analysis was set at the conventional α = .05, meaning a result is treated as statistically significant if the probability of observing a difference this large by chance alone, assuming the null hypothesis is true, is less than 5%.
The dataset comprises final examination marks for 102 postgraduate students enrolled on the same module and assessed using the same marking scheme and marker moderation process, which controls for assessment-related confounds between the two groups. Fifty-two students were enrolled on the blended-delivery variant of the module and fifty on the traditional, in-person variant; allocation reflected timetabling and campus location rather than student choice of attainment level, reducing (though not eliminating) the risk of self-selection bias. Table 1 summarises the two variables entered into SPSS.
| Variable | Role | Type | Measurement / Coding |
|---|---|---|---|
| Study Mode | Independent variable | Categorical (nominal), 2 levels | 1 = Blended learning, 2 = Traditional (in-person) |
| Exam Score | Dependent variable | Continuous (scale) | Final module examination mark, 0–100 |
Prior to analysis, the dataset was screened in SPSS using frequencies and a boxplot split by group. No missing values were present and no case returned a standardised score beyond ±3.29, the conventional threshold for identifying extreme outliers at this sample size (Field, 2018; Tabachnick and Fidell, 2019), so no cases were removed.
All exam scripts were marked against a shared assessment rubric by a team of module markers who took part in a standardisation exercise before marking began, and a sample of scripts from each study-mode group was independently second-marked, which reduces the risk that any between-group difference in mean score simply reflects marker drift rather than a genuine difference in performance. Student identifiers were replaced with anonymised participant codes before the dataset was exported for analysis, and the file was stored on a password-protected university drive in line with institutional data-protection requirements, consistent with the ethical approval granted for secondary analysis of routinely collected assessment data. In SPSS, study mode was entered as a numeric grouping variable in Variable View, with value labels of 1 (‘Blended’) and 2 (‘Traditional’) attached so that output tables display meaningful group names rather than raw numeric codes, and exam score was entered as a scale-level numeric variable measured to one decimal place.
Four assumptions underpin the independent-samples t-test: independence of observations, a continuous dependent variable, approximate normality of the dependent variable within each group, and homogeneity of variance across groups. Independence was satisfied by design, since each student appears in one group only and there is no repeated-measures or clustering structure in the data. The dependent variable, exam score, is continuous by definition.
Normality within each group was assessed using the Shapiro–Wilk test (Shapiro and Wilk, 1965), which is preferred over Kolmogorov–Smirnov for sample sizes below approximately 100 per group (Pallant, 2020). Table 2 reports the results.
| Group | Shapiro–Wilk Statistic (W) | df | Sig. (p) |
|---|---|---|---|
| Blended learning | .978 | 52 | .436 |
| Traditional (in-person) | .965 | 50 | .146 |
Neither test statistic was significant at the .05 level (blended: W(52) = .978, p = .436; traditional: W(50) = .965, p = .146), so the assumption of normality was retained for both groups, supported by visual inspection of Q–Q plots that showed points falling close to the diagonal in each condition, and by boxplots that showed roughly symmetric distributions with no extreme values in either group. It is also worth noting that, even had the Shapiro–Wilk test returned a significant result in one group, the t-test is generally considered robust to moderate departures from normality once group sizes exceed around 30, by appeal to the Central Limit Theorem (Field, 2018; Howell, 2017); with 52 and 50 cases respectively, both groups comfortably exceed this threshold, providing a further layer of reassurance beyond the formal test result.
Homogeneity of variance was tested using Levene’s test, which SPSS reports automatically alongside the independent-samples t-test output. Table 3 reports the result.
| Levene’s F | df1 | df2 | Sig. (p) |
|---|---|---|---|
| 1.87 | 1 | 100 | .175 |
Levene’s test was not significant, F(1, 100) = 1.87, p = .175, indicating that the assumption of equal population variances could not be rejected. The ‘equal variances assumed’ row of the SPSS output was therefore used to report the t-test result, rather than the Welch-adjusted row that SPSS provides for cases where this assumption is violated. Had Levene’s test returned a significant result, the appropriate remedy would not have been to abandon the t-test, but to report the Welch-corrected row instead, which adjusts the degrees of freedom downward to account for unequal variances rather than assuming a shared pooled variance; SPSS calculates and reports this alternative automatically in the same output table regardless of which assumption is met, which is good practice, since it removes the temptation to search for a significant result across multiple corrected and uncorrected tests.
Table 4 presents the descriptive statistics for exam score by study mode.
| Group | N | Mean | SD | SE Mean | Min | Max |
|---|---|---|---|---|---|---|
| Blended learning | 52 | 68.4 | 9.2 | 1.28 | 42 | 89 |
| Traditional (in-person) | 50 | 64.1 | 10.5 | 1.49 | 38 | 87 |
Students in the blended-delivery group achieved a higher mean exam score (M = 68.4, SD = 9.2) than students in the traditional, in-person group (M = 64.1, SD = 10.5), a raw difference of 4.3 marks. Table 5 presents the full independent-samples t-test output, reproduced in the standard layout SPSS itself generates and as recommended for student reporting by Laerd Statistics (2021).
| Levene’s F | Levene’s Sig. | t | df | Sig. (2-tailed) | Mean Diff. | SE Diff. | 95% CI Lower | 95% CI Upper | |
|---|---|---|---|---|---|---|---|---|---|
| Equal variances assumed | 1.87 | .175 | 2.20 | 100 | .030 | 4.30 | 1.95 | 0.43 | 8.17 |
With Levene’s test non-significant, the equal-variances row applies: t(100) = 2.20, p = .030 (two-tailed). The 95% confidence interval for the mean difference, 0.43 to 8.17 marks, does not cross zero, which is consistent with a statistically significant result at the conventional .05 threshold.
Reading this table row by row: the t-value of 2.20 expresses the size of the mean difference (4.30 marks) relative to its standard error (1.95 marks); a larger t-value indicates a mean difference that is large relative to the sampling variability expected by chance. The degrees of freedom, 100, follow directly from n1 + n2 − 2 (52 + 50 − 2). The two-tailed significance value, .030, is the probability of observing a t-value at least this extreme, in either direction, if the true population means were equal; because .030 is below the pre-specified α of .05, the null hypothesis of no difference is rejected. The 95% confidence interval can be interpreted as a plausible range for the true population mean difference: were this study repeated many times, approximately 95% of the resulting confidence intervals would be expected to contain the true population mean difference, and because the interval reported here (0.43 to 8.17) excludes zero, it corroborates the significant p-value.
Figure 1: Mean final exam score by study mode (error bars show ±1 SE).
Figure 1 displays the group means with standard error bars, illustrating the higher average performance of the blended-delivery group and the modest overlap implied by the two intervals.
Consistent with American Psychological Association (2020) reporting conventions, the result is summarised below in a single statistical sentence followed by a plain-English explanation. An independent-samples t-test was conducted to compare final exam scores for students taught via blended and traditional delivery. There was a statistically significant difference in scores for the blended condition (M = 68.4, SD = 9.2, n = 52) compared with the traditional condition (M = 64.1, SD = 10.5, n = 50); t(100) = 2.20, p = .030, d = 0.44.
Cohen’s d was calculated by dividing the mean difference (4.3) by the pooled standard deviation (9.86), giving d = 0.44. Using Cohen’s (1988) conventions, this represents a small-to-medium effect: a real but modest practical difference rather than a dramatic one. In plain terms, students on the blended module scored, on average, around four to five marks more than students on the traditional module, and this gap is unlikely to be attributable to sampling error alone, though it explains only a modest proportion of the overall variation in exam performance between individual students.
It is worth distinguishing clearly between statistical significance and practical significance at this point, since the two are often conflated in student write-ups. A p-value below .05 tells us only that a difference this large would be unlikely to occur by chance if the true population means were identical; it says nothing directly about whether a four-to-five-mark gap is educationally meaningful. Cohen’s d addresses that second question, and a value of 0.44 indicates that the two group distributions overlap substantially: many students in the traditional group still outperformed many students in the blended group, and vice versa, even though the group averages differ. A useful way to communicate this to a non-technical audience, such as a programme board, is that the average blended-delivery student scored roughly in the region of the 57th percentile of the traditional group’s score distribution, a meaningful but far from decisive advantage.
This finding sits within a mixed evidence base on blended learning. Means et al. (2013) similarly reported a small-to-moderate average advantage for blended conditions over purely face-to-face instruction across the higher-education studies they reviewed, broadly consistent in direction and magnitude with the effect observed here. It is worth noting, however, that the studies underlying that meta-analysis were drawn overwhelmingly from undergraduate rather than postgraduate contexts, and often compared blended conditions against a fully online, rather than fully face-to-face, comparator, so the present postgraduate, UK-based finding should be read as a complementary, context-specific data point rather than a straightforward replication of that earlier evidence base. Several theoretical mechanisms have been proposed to explain such gains. Spaced-practice theory suggests that the discrete, revisitable nature of recorded online material encourages students to revise content in shorter, more frequent sessions than a single weekly lecture format allows, which is associated with stronger long-term retention than massed study close to the assessment date. An alternative, complementary explanation draws on cognitive load theory: pre-recorded material allows students to pause, rewind and re-watch complex sections at their own pace, potentially reducing extraneous cognitive load during initial encoding compared with a single, fixed-pace, in-person delivery.
Equally, some caution is warranted before attributing the observed difference solely to delivery format. Other studies caution that attainment gains in blended cohorts can reflect confounding factors, such as differences in the type of student who selects or is allocated to flexible provision, rather than a genuine causal effect of delivery mode itself. In this dataset, allocation was determined by timetabling and campus location rather than an explicit student choice of study mode, which reduces but does not eliminate this concern, since timetabling decisions can themselves correlate with other student characteristics, such as term-time employment patterns or distance travelled to campus, that were not measured here and could plausibly relate to exam performance independently of study mode.
This analysis found a statistically significant, small-to-medium difference in final exam performance favouring students taught via blended delivery over students taught via traditional, wholly in-person delivery, t(100) = 2.20, p = .030, d = 0.44. The assumptions of normality and homogeneity of variance were both satisfied, supporting confidence in the validity of the equal-variances t-test result reported here.
Several limitations qualify this conclusion. First, group allocation was determined by timetabling rather than random assignment, so unmeasured differences between the two cohorts, such as prior attainment or motivation, cannot be ruled out as alternative explanations for the observed difference. Second, the analysis draws on a single module and a single institution, which limits the generalisability of the finding to other subjects, levels or institutional contexts. Third, exam score is only one indicator of learning and does not capture other outcomes, such as skills development or student satisfaction, that might respond differently to delivery mode. Fourth, a sensitivity check using G*Power indicates that the achieved sample of 102 provided approximately 80% power to detect a medium effect (d = 0.5) at α = .05 two-tailed, which is adequate for the effect actually observed but would be under-powered to detect a smaller true effect, meaning a non-significant result in a similarly sized future replication should not automatically be read as evidence that no difference exists. A randomised or matched-groups design across multiple modules, ideally supplemented by measures of prior attainment as a covariate, would strengthen causal inference in future research on this question.
American Psychological Association (2020) Publication Manual of the American Psychological Association. 7th edn. Washington, DC: APA.
Cohen, J. (1988) Statistical Power Analysis for the Behavioral Sciences. 2nd edn. Hillsdale, NJ: Lawrence Erlbaum Associates.
Field, A. (2018) Discovering Statistics Using IBM SPSS Statistics. 5th edn. London: Sage.
Howell, D.C. (2017) Statistical Methods for Psychology. 9th edn. Boston, MA: Cengage Learning.
Laerd Statistics (2021) Independent-Samples t-Test Using SPSS Statistics. Lund Research Ltd.
Means, B., Toyama, Y., Murphy, J. and Baki, M. (2013) ‘The effectiveness of online and blended learning: a meta-analysis of the empirical literature’, Teachers College Record, 115(3), pp. 1–47.
Pallant, J. (2020) SPSS Survival Manual. 7th edn. Maidenhead: Open University Press.
Shapiro, S.S. and Wilk, M.B. (1965) ‘An analysis of variance test for normality (complete samples)’, Biometrika, 52(3–4), pp. 591–611.
Tabachnick, B.G. and Fidell, L.S. (2019) Using Multivariate Statistics. 7th edn. Boston, MA: Pearson.
Need a Model Statistical Analysis Written to Your Exact Brief?
Our 350+ UK-qualified writers deliver referenced model documents from £15 per 250 words, with free plagiarism and AI-detection reports.
You May Also Like