Home > Knowledge Base > Assignment Samples > Assignment Sample: Evaluating a Numeracy Intervention in a Primary School

Assignment Sample: Evaluating a Numeracy Intervention in a Primary School

Published by at July 30th, 2026 , Revised On July 30, 2026

Type: Assignment  |  Subject: Education  |  Level: Masters  |  Word Count: ~2500 words

This model assignment was produced by an Essays UK specialist as reference material for learning purposes only. For support in this field, see our education assignment specialists.

The Brief

Using an appropriate evaluation framework, critically evaluate a school-based numeracy intervention that you have implemented or observed on placement. Your evaluation should present and interpret relevant quantitative attainment data, discuss the strengths and limitations of the evaluation design, and conclude with evidence-informed recommendations for future practice. (2,500 words)

Model Answer

Introduction

Numeracy attainment gaps that open during the early years of primary education are strongly associated with reduced life chances, and national inspection evidence continues to highlight uneven progress in number fluency across Key Stage 2 (Ofsted, 2023). This assignment evaluates a targeted, small-group numeracy intervention delivered to a Year 3 cohort in a two-form entry primary school, designed to accelerate progress for pupils working below age-related expectations in place value and number fluency. The intervention, referred to here as the Numeracy Catch-Up Programme, was delivered by a trained teaching assistant over a ten-week period and forms part of the school’s wider mathematics improvement strategy, developed in response to a previous inspection judgement that identified inconsistent progress for lower-attaining pupils in mathematics.

The purpose of this evaluation is threefold: first, to assess the impact of the intervention on standardised numeracy attainment relative to a matched comparison group; second, to evaluate the design and delivery of the intervention against evidence-based criteria for effective mathematics support, drawing principally on the Education Endowment Foundation’s (EEF, 2021) guidance report on improving mathematics at Key Stages 2 and 3; and third, to identify implications for future practice within the school’s tiered approach to mathematics teaching, in which whole-class quality-first teaching is supplemented by targeted and, where necessary, specialist support.

The evaluation adopts a pragmatic, mixed-methods design, combining quantitative attainment data with brief qualitative feedback from the delivering practitioner, and is informed throughout by the British Educational Research Association’s (BERA, 2018) ethical guidelines for educational research. The remainder of this assignment sets out the intervention’s design and the evaluation method, presents and analyses the resulting attainment data, offers a structured critical evaluation of the programme against an established evaluation framework, and concludes with recommendations for strengthening future cycles of delivery and evaluation.

For the purposes of this evaluation, numeracy is defined broadly, following the National Numeracy charity and the Department for Education’s mathematics guidance, as the confident and competent use of numbers and mathematical approaches in everyday life, rather than the narrower skill of calculation alone. This distinction matters for programme design: the intervention evaluated here targeted procedural fluency in place value and number operations as a necessary foundation, but was explicitly designed to build towards flexible, applied use of number, in line with the wider definition. Framing the intervention in this way also shapes the evaluation criteria used later in this assignment, since attainment gains on a standardised test are treated as one important but partial indicator of success, alongside evidence of pupils’ willingness to apply strategies independently.

Intervention Design and Evaluation Method

The Numeracy Catch-Up Programme was designed around three evidence-informed principles identified by the EEF (2021) as characteristic of effective mathematics interventions: explicit, systematic instruction; extensive use of concrete and pictorial representations before abstract notation; and frequent, low-stakes retrieval practice. Sessions followed Bruner’s (1966) concrete-pictorial-abstract sequence, moving pupils from base-ten manipulatives to bar-model representations before requiring purely numerical calculation, and were structured to build what Skemp (1976) termed relational understanding – the ability to explain why a procedure works – rather than instrumental understanding of isolated rules learned without conceptual grounding. Each thirty-minute session began with a two-minute retrieval activity revisiting previously taught content, consistent with spaced-practice principles, before introducing new material and closing with an independent application task used to inform the following session’s starting point.

Each session lasted thirty minutes, ran three times weekly for ten weeks, and was delivered to groups of four pupils by a teaching assistant who had received two half-day training sessions from the school’s mathematics lead, covering both the intervention’s content sequence and general principles of effective small-group instruction. Pupil selection was based on autumn-term standardised assessment scores, using a threshold of one standard deviation below the year-group mean. Sixteen pupils met this criterion and were enrolled in the intervention; a matched comparison group of sixteen pupils, drawn from a parallel class and matched on prior attainment, age and free school meal eligibility, continued to receive quality-first teaching only, without additional small-group support during the evaluation window.

Allocation was not randomised, reflecting the practical constraints of a live school timetable and the school’s judgement that withholding support from eligible pupils for research purposes would not be justifiable; this limitation is returned to in the evaluation section below. A quasi-experimental pre-test/post-test design was adopted instead. Both groups completed a standardised numeracy assessment in the week before the intervention began and again in the week following its completion, generating raw scores that were converted to age-standardised scores for comparability across the ten-week period. This quantitative data was supplemented by a short structured interview with the delivering teaching assistant, exploring perceived changes in pupil confidence and engagement, consistent with the emphasis Black and Wiliam (1998) place on formative, process-level evidence alongside outcome measures. Consent for the use of anonymised attainment data was obtained through the school as gatekeeper, in line with BERA (2018) guidance; no individually identifiable data is reported in this assignment, and pupils are referred to only by group membership throughout.

The choice of a small-group, mastery-oriented model over whole-class re-teaching was itself informed by the evidence base on effective grouping for lower attainers. Gersten et al. (2009) report consistently larger effect sizes for structured small-group tuition than for undifferentiated whole-class re-teaching of the same content, while cautioning that groups larger than six pupils see diminishing returns because individual response and correction opportunities fall sharply. A group size of four was therefore chosen as a practical compromise between the evidence-based ideal of one-to-one or pair tuition and the staffing constraints of a single trained teaching assistant covering the full cohort of sixteen eligible pupils within the available timetable slots.

Findings and Analysis

Table 1 summarises pre- and post-intervention standardised scores for both groups. The intervention group’s mean score rose from 68.4 (SD = 7.2) to 78.9 (SD = 6.5), a mean gain of 10.5 points, compared with a mean gain of only 3.2 points in the comparison group, whose mean rose from 69.1 (SD = 6.8) to 72.3 (SD = 7.0).

Group N Pre-test Mean (SD) Post-test Mean (SD) Mean Gain Gain SD
Intervention 16 68.4 (7.2) 78.9 (6.5) 10.5 4.8
Comparison 16 69.1 (6.8) 72.3 (7.0) 3.2 4.1

To assess whether this difference in gain scores was likely to reflect a genuine intervention effect rather than sampling variation, an effect size and independent-samples t-statistic were calculated by hand. The pooled standard deviation of the two gain-score distributions was found first:

SDpooled = √[((n₁−1)SD₁² + (n₂−1)SD₂²) / (n₁+n₂−2)]
= √[((15 × 4.8²) + (15 × 4.1²)) / 30]
= √[(15 × 23.04 + 15 × 16.81) / 30]
= √[(345.6 + 252.15) / 30]
= √19.93 = 4.46

Cohen’s d was then calculated as the difference between group means divided by the pooled standard deviation:

d = (Mintervention − Mcomparison) / SDpooled = (10.5 − 3.2) / 4.46 = 7.3 / 4.46 = 1.64

Following Cohen’s (1988) benchmarks, where d = 0.2, 0.5 and 0.8 represent small, medium and large effects respectively, an effect size of 1.64 indicates a very large effect. An independent-samples t-statistic was then computed to test the significance of the gain-score difference:

t = (Mintervention − Mcomparison) / [SDpooled × √(1/n₁ + 1/n₂)]
= 7.3 / [4.46 × √(1/16 + 1/16)]
= 7.3 / [4.46 × √0.125]
= 7.3 / [4.46 × 0.354]
= 7.3 / 1.58 = 4.62

With 30 degrees of freedom, t = 4.62 exceeds the critical value required for significance at p < .001, indicating that the intervention group’s additional gain is very unlikely to be due to chance alone. A supplementary breakdown of the intervention group by initial attainment band showed that gains were not uniform: the eight pupils with the lowest pre-test scores (below 65) gained an average of 12.9 points, while the eight pupils with pre-test scores between 65 and 74 gained an average of 8.1 points, suggesting the programme was particularly effective for the very lowest attainers within an already low-attaining group. Qualitative feedback from the delivering teaching assistant reinforced this pattern: she reported that pupils became noticeably more willing to attempt multi-step problems independently by week six, and that several began spontaneously using the bar-model approach introduced in sessions during ordinary classroom mathematics, a form of transfer that standardised test scores alone do not capture.

A 95% confidence interval was also calculated around the mean gain-score difference, to express the precision of the estimate rather than relying on the significance test alone. Using the standard error already calculated above (SE = 1.58) and the two-tailed critical value of t at 30 degrees of freedom (tcrit ≈ 2.042):

95% CI = (Mintervention − Mcomparison) ± (tcrit × SE)
= 7.3 ± (2.042 × 1.58)
= 7.3 ± 3.23
= [4.07, 10.53]

Because this interval does not cross zero, it confirms the significant t-test result while additionally showing the plausible range of the true difference: even at the lower bound, the intervention group’s advantage over the comparison group would represent a meaningful gain of around four standardised points, supporting confidence in the direction, if not the precise magnitude, of the effect.

Evaluation

Interpreted through Stufflebeam’s (2003) CIPP framework, which evaluates context, input, process and product, the intervention performs well on product measures but reveals weaknesses at the input and process stages that qualify the strength of the attainment findings. On context and input, the programme aligns closely with the evidence base: its explicit instruction, CPA sequencing and retrieval practice reflect exactly the features the EEF (2021) and Gersten et al. (2009) identify as effective for pupils struggling with mathematics, and its focus on relational rather than purely instrumental understanding follows Skemp (1976) and Nunes, Bryant and Watson (2009) in prioritising conceptual grasp over rote procedure.

At the process and product stages, however, several threats to validity temper the findings. First, allocation to intervention and comparison groups was not randomised; although groups were matched on prior attainment, age and free school meal eligibility, unmeasured differences – for example in home numeracy support or attendance – cannot be ruled out as alternative explanations for the observed gap. Second, the sample is small (n = 32 in total), which increases the influence of individual outlying scores on the calculated effect size and limits confidence in generalising beyond this cohort. Third, no independent fidelity checks were carried out on delivery; the evaluation relies on a single teaching assistant’s self-report that sessions followed the intended structure, and observed drift from the planned approach – a well-documented risk in small-group interventions delivered outside the class teacher’s direct oversight – cannot be excluded. Fourth, a form of Hawthorne effect is plausible: pupils withdrawn for a novel, small-group activity may have shown improved engagement and effort regardless of the specific instructional content, inflating apparent gains relative to pupils receiving only ordinary classroom teaching.

The evaluation also has to weigh statistical against practical significance. A gain of 10.5 standardised points is educationally meaningful, plausibly closing a substantial part of the attainment gap that prompted selection for the intervention, and the large effect size is consistent with findings reported elsewhere for structured, short-cycle numeracy interventions (Fuchs, Fuchs and Compton, 2013; Dowker, 2019). However, the ten-week window captures only immediate post-intervention performance; without a delayed follow-up assessment, it is not possible to establish whether gains are sustained once the additional support is withdrawn, a limitation consistent with Black and Wiliam’s (1998) caution that short-term assessment gains do not always translate into durable learning. Cost and sustainability considerations are also relevant to a full CIPP evaluation: the programme required a trained teaching assistant’s time across thirty half-hour sessions per group, a resource commitment that is only justifiable at scale if gains prove durable and if similar results can be replicated with less experienced staff.

Finally, questions of transferability and scalability are relevant to any judgement about whether the programme should be extended school-wide. The results reported here derive from a single teaching assistant working with a specific cohort in one school; effectiveness may depend on characteristics of that individual’s practice, training and rapport that would not automatically transfer to a wider rollout using less experienced or less well-trained staff. A CIPP-consistent evaluation of scalability would therefore need to test the programme with multiple deliverers across more than one year group before the current findings could reliably support whole-school or multi-school adoption.

Guskey’s (2000) argument that evaluation of any structured programme should extend beyond immediate outcome data to consider organisational support and the conditions needed to sustain new practice is also instructive here. The present evaluation stops at the level of pupil learning outcomes and brief practitioner comment; it does not examine whether the school’s timetable, staffing model and continuing professional development offer are sufficient to sustain the programme beyond a single ten-week cycle, nor whether the mathematics lead has capacity to train and quality-assure additional teaching assistants if the intervention were to be scaled. A fuller evaluation in a subsequent cycle should therefore incorporate these organisational-level indicators alongside pupil attainment data.

Conclusion and Recommendations

This evaluation provides encouraging, though methodologically qualified, evidence that the Numeracy Catch-Up Programme accelerated attainment for the targeted Year 3 pupils relative to a matched comparison group, with a large effect size and a statistically significant difference in gain scores. The design closely reflects evidence-based principles for effective mathematics intervention, and qualitative feedback suggests concurrent gains in pupil confidence and independent application of strategies, particularly for the very lowest-attaining pupils within the cohort.

Three recommendations follow from this evaluation. First, the school should move towards a randomised or waitlist-controlled design for future intervention cycles, allocating eligible pupils to immediate or delayed start groups; this would strengthen causal inference without denying any pupil access to support. Second, a fidelity-monitoring process should be introduced, such as brief peer observation of two sessions per intervention cycle against a structured checklist, so that process evaluation no longer relies solely on practitioner self-report. Third, a delayed post-test at the end of the following term should be added to assess retention, addressing the current design’s inability to distinguish immediate performance gains from durable learning. Embedding these refinements within the school’s existing tiered mathematics strategy, and piloting delivery by a second trained teaching assistant, would allow the programme to generate more robust evidence of impact while continuing to provide timely support to the pupils who most need it.

References

  • BERA (2018) Ethical Guidelines for Educational Research, 4th edn. London: British Educational Research Association.
  • Black, P. and Wiliam, D. (1998) ‘Assessment and classroom learning’, Assessment in Education: Principles, Policy & Practice, 5(1), pp. 7-74.
  • Bruner, J.S. (1966) Toward a Theory of Instruction. Cambridge, MA: Harvard University Press.
  • Cohen, J. (1988) Statistical Power Analysis for the Behavioural Sciences, 2nd edn. Hillsdale, NJ: Lawrence Erlbaum Associates.
  • Dowker, A. (2019) Individual Differences in Arithmetic: Implications for Psychology, Neuroscience and Education, 2nd edn. Abingdon: Routledge.
  • Education Endowment Foundation (2021) Improving Mathematics in Key Stages 2 and 3: Guidance Report. London: EEF.
  • Fuchs, L.S., Fuchs, D. and Compton, D.L. (2013) ‘Intervention effects for students with comorbid forms of learning disability’, Journal of Learning Disabilities, 46(6), pp. 501-514.
  • Gersten, R., Chard, D.J., Jayanthi, M., Baker, S.K., Morphy, P. and Flojo, J. (2009) ‘Mathematics instruction for students with learning disabilities’, Review of Educational Research, 79(3), pp. 1202-1242.
  • Guskey, T.R. (2000) Evaluating Professional Development. Thousand Oaks, CA: Corwin Press.
  • Nunes, T., Bryant, P. and Watson, A. (2009) Key Understandings in Mathematics Learning. London: Nuffield Foundation.
  • Ofsted (2023) Coordinating Mathematical Success: The Primary Mathematics Subject Report. Manchester: Ofsted.
  • Skemp, R.R. (1976) ‘Relational understanding and instrumental understanding’, Mathematics Teaching, 77, pp. 20-26.
  • Stufflebeam, D.L. (2003) ‘The CIPP model for evaluation’, in Kellaghan, T. and Stufflebeam, D.L. (eds.) International Handbook of Educational Evaluation. Dordrecht: Kluwer, pp. 31-62.

Need a Model Assignment Written to Your Exact Brief?

Our 350+ UK-qualified writers deliver referenced model documents from £15 per 250 words, with free plagiarism and AI-detection reports.

Order Your Model Assignment

Frequently Asked Questions

About Jesse Pinkman

Avatar for Jesse PinkmanJessie Pinkman has been writing since childhood when her mother gave her a book where she could write her stories. Since then Jessie has always loved to write about the topics she loves. She graduated from Birmingham University in 2012, worked as a teaching assistant, and then turned to full-time writing in 2016.

You May Also Like

WhatsApp Live Chat