Home > Knowledge Base > Dissertation Statistical Analysis Samples > Statistical Analysis Sample: Chi-Square Test of Association Between Age and Brand Preference

Statistical Analysis Sample: Chi-Square Test of Association Between Age and Brand Preference

Published by at August 13th, 2026 , Revised On August 13, 2026

Type: Statistical Analysis  |  Subject: Statistics  |  Level: Undergraduate  |  Word Count: ~2500 words

This model statistical analysis was produced by an Essays UK specialist as reference material for learning purposes only. For support in this field, see our marketing assignment specialists.

The Brief

Using SPSS, conduct a chi-square test of association to determine whether there is a significant relationship between age group and brand preference in the attached dataset (N = 300). Report the observed and expected frequencies, the test statistic, an appropriate effect size, and provide a full interpretation in APA style.

Model Answer

Introduction and Research Question

Understanding whether consumer preferences vary systematically across age groups is a routine but important question in applied marketing research, since it informs decisions about market segmentation, targeted advertising and product positioning (Solomon et al., 2019). This report presents a chi-square test of association conducted to establish whether there is a statistically significant relationship between a consumer’s age group and their stated preference among three competing soft drink brands sold in the UK market.

Targeted marketing spend that is not grounded in evidence risks both wasted budget and messaging that fails to resonate with its intended audience, making a preliminary statistical check of this kind good practice before a full campaign is greenlit by senior stakeholders. A chi-square test of association is a well-established and easily communicated tool for this purpose, since its output — a single test statistic, a p-value and an effect size — can be explained to a non-technical marketing audience far more readily than the output of a more complex multivariate model.

The analysis was requested by a marketing team preparing a segmentation strategy ahead of a rebranding campaign. Rather than assuming that brand preference is evenly distributed across age groups, the team wanted evidence-based confirmation of whether age is genuinely associated with brand choice before committing budget to age-targeted messaging. Because both variables of interest — age group and brand preference — are categorical rather than continuous, a chi-square test of association (also known as a chi-square test of independence) is the appropriate inferential statistic here, rather than a correlation or t-test, either of which requires at least one continuous variable (Field, 2018).

The research question was: is there a statistically significant association between age group (18–24, 25–40, 41+) and preferred brand (Brand A, Brand B, Brand C) among UK consumers? The null hypothesis (H0) stated that age group and brand preference are independent, meaning the distribution of brand preference does not differ across age groups. The alternative hypothesis (H1) stated that age group and brand preference are not independent, meaning the distribution of brand preference differs significantly by age group. An alternative approach, such as multinomial logistic regression, could in principle model the same relationship, but the chi-square test was selected as the most direct and widely reported test for a simple two-way categorical association of this kind (Howitt and Cramer, 2020). The sections that follow describe the sample, report the assumption checks required before a chi-square test can be trusted, present the test output, and interpret the findings for the marketing team.

Data and Variables

Data were collected via a structured intercept survey conducted across shopping centres in four English city centres over a two-week period, in which respondents were asked to select their age band and their most-preferred brand from the three available options. Fieldworkers were briefed to approach every fifth adult passing a fixed point in each location to reduce interviewer selection bias, a standard practice in intercept research. After removing incomplete responses and the small number of respondents who selected “no preference,” a final sample of 300 consumers was retained, comprising 100 respondents in each of the three age bands (18–24, 25–40 and 41+) by design, since a quota sampling approach was used to ensure adequate representation across all three age groups rather than relying on whichever age groups happened to be most present in each location at the time of fieldwork.

Both variables were categorical. Age group had three levels (18–24, 25–40, 41+) and brand preference had three levels (Brand A, Brand B, Brand C). Each respondent contributed exactly one response to one cell of the resulting 3 × 3 contingency table, satisfying the independence-of-observations requirement discussed in the following section. No other demographic variables, such as gender or income, were collected as part of this particular analysis, since the brief specified age as the segmentation variable of interest. Table 1 presents the observed frequencies — the raw counts — cross-tabulating age group against brand preference, together with the row and column totals.

Age Group Brand A Brand B Brand C Row Total
18–24 45 30 25 100
25–40 30 40 30 100
41+ 20 35 45 100
Column Total 95 105 100 300

As Table 1 shows, Brand A was most popular among the youngest age group, Brand B held a fairly even share across all three groups, and Brand C was clearly most popular among the oldest age group, with the reverse pattern true for Brand A. Whether this apparent pattern reflects a statistically significant association, or could plausibly have arisen through sampling chance alone, is addressed in the Test Results section below.

Preliminary Checks and Assumptions

Three assumptions must be satisfied for a chi-square test of association to produce a trustworthy result, and each was checked prior to running the test (Field, 2018; Pallant, 2020).

First, both variables must be categorical, which was satisfied by design: age group and brand preference were both measured as discrete, mutually exclusive categories rather than as continuous scores, and every respondent could belong to one and only one combination of age band and preferred brand.

Second, observations must be independent, meaning each respondent can appear in only one cell of the contingency table. This was ensured by the survey design, which asked each respondent to select exactly one age band and exactly one preferred brand, and by removing any duplicate respondent identifiers prior to analysis so that no individual was counted twice.

Third, and most commonly violated in small studies, the expected cell frequencies must be sufficiently large for the chi-square approximation to the sampling distribution to be valid. The conventional rule, following Cochran, is that no more than 20% of cells should have an expected count below 5, and no cell should have an expected count below 1 (Field, 2018). Table 2 reports the expected frequencies calculated for each cell of the contingency table, using the standard formula E = (row total × column total) / grand total.

Age Group Brand A (E) Brand B (E) Brand C (E)
18–24 31.67 35.00 33.33
25–40 31.67 35.00 33.33
41+ 31.67 35.00 33.33

As Table 2 shows, the smallest expected frequency across all nine cells was 31.67, comfortably above the minimum threshold of 5, and none of the nine cells fell anywhere near the 20% rule described above. This confirms that the sample size and cell distribution were adequate for the chi-square test to proceed with confidence, and that a Fisher’s exact test, typically reserved for small or sparse tables, was not required for this dataset (Agresti and Finlay, 2009).

Test Results

A Pearson chi-square test of association was conducted in SPSS (version 28) to test the relationship between age group and brand preference. Table 3 below shows the calculation of each cell’s contribution to the overall chi-square statistic, following the formula χ² = Σ[(O − E)² / E], where O is the observed frequency and E is the expected frequency for each cell.

Cell O E O − E (O − E)² (O − E)²/E
18–24 × Brand A 45 31.67 13.33 177.69 5.61
18–24 × Brand B 30 35.00 −5.00 25.00 0.71
18–24 × Brand C 25 33.33 −8.33 69.39 2.08
25–40 × Brand A 30 31.67 −1.67 2.79 0.09
25–40 × Brand B 40 35.00 5.00 25.00 0.71
25–40 × Brand C 30 33.33 −3.33 11.09 0.33
41+ × Brand A 20 31.67 −11.67 136.19 4.30
41+ × Brand B 35 35.00 0.00 0.00 0.00
41+ × Brand C 45 33.33 11.67 136.19 4.09
Total 300 300 17.92

Summing the final column across all nine cells gives the overall test statistic: χ²(4, N = 300) = 17.92, p = .001. The degrees of freedom were calculated as (rows − 1) × (columns − 1) = (3 − 1) × (3 − 1) = 4. Table 5 reproduces the “Chi-Square Tests” output block as it would appear directly in SPSS, including the Likelihood Ratio statistic (an alternative, asymptotically equivalent test) and the Linear-by-Linear Association test, which is not meaningful here given that brand preference is a nominal, not ordinal, category.

Test Value df Asymp. Sig. (2-sided)
Pearson Chi-Square 17.92 4 .001
Likelihood Ratio 17.72 4 .001
Linear-by-Linear Association 0.42 1 .517
N of Valid Cases 300

The Likelihood Ratio statistic (17.72) closely tracks the Pearson chi-square value (17.92), as would be expected when expected cell counts are all reasonably large, and leads to the same substantive conclusion. Because the overall test was statistically significant, an effect size was calculated to establish the strength, rather than merely the presence, of the association. Cramér’s V was used, as it is the appropriate effect size measure for chi-square tables larger than 2 × 2: V = √[χ² / (N × (k − 1))], where k is the smaller of the number of rows or columns. This gave V = √[17.92 / (300 × 2)] = √0.030 = .17, which falls in the small-to-moderate range using Cohen’s (1988) benchmarks for a 3 × 3 table (small = .07, medium = .21, large = .35).

To identify which cells contributed most strongly to the significant overall result, standardised residuals were calculated for each cell using the formula (O − E) / √E. As a rough guide, a standardised residual larger than ±1.96 in absolute value can be treated as a noteworthy deviation from independence at the .05 level (Field, 2018). Table 4 reports these values.

Cell Standardised Residual
18–24 × Brand A 2.37
18–24 × Brand B −0.85
18–24 × Brand C −1.44
25–40 × Brand A −0.30
25–40 × Brand B 0.85
25–40 × Brand C −0.58
41+ × Brand A −2.07
41+ × Brand B 0.00
41+ × Brand C 2.02

Figure 1 displays the observed brand preference counts broken down by age group, illustrating the pattern underlying the significant chi-square result.

Observed Brand Preference by Age Group Brand A Brand B Brand C 45 30 25 18–24 30 40 30 25–40 20 35 45 41+

Figure 1: Observed brand preference counts by age group (N = 300); χ²(4, N = 300) = 17.92, p = .001, Cramér’s V = .17.

Interpretation

The chi-square test of association was statistically significant, χ²(4, N = 300) = 17.92, p = .001, Cramér’s V = .17, indicating that brand preference is not independent of age group in this sample. The null hypothesis of independence is therefore rejected in favour of the alternative hypothesis that age group and brand preference are associated.

The standardised residuals reported in Table 4 help to pinpoint where this association is concentrated. Three cells exceeded the ±1.96 threshold: 18–24-year-olds who preferred Brand A had a standardised residual of 2.37, indicating considerably more respondents preferred Brand A in this age band than would be expected under independence; respondents aged 41+ who preferred Brand A had a standardised residual of −2.07, indicating considerably fewer than expected; and respondents aged 41+ who preferred Brand C had a standardised residual of 2.02, indicating considerably more than expected. All other cells, including every cell involving Brand B, fell within the normal range of sampling variation, suggesting that Brand B’s appeal is comparatively consistent across the age groups tested.

In plain terms for the marketing team, these results suggest that Brand A over-indexes with younger consumers while Brand C over-indexes with older consumers, and this pattern is unlikely to be due to sampling chance alone (p = .001). However, the effect size (Cramér’s V = .17) indicates that while the association is statistically reliable, it is modest in magnitude rather than a stark divide — a sizeable proportion of consumers in every age band still expressed a preference for each of the three brands. This distinction between statistical significance and practical significance matters directly for budget allocation: while age-targeted messaging is justified by this data, the team should not conclude that age group is the dominant driver of brand choice, and should continue to investigate other segmentation variables, such as usage occasion or price sensitivity, alongside age.

A further practical implication is that a rebranding campaign aimed at repositioning Brand A to appeal to a broader age range would need to address whatever is driving its comparatively weak performance among the 41+ segment, since the data suggest this is a genuine, non-random pattern in the market rather than statistical noise. Equally, any campaign built around Brand C should be cautious about assuming younger consumers will respond in the same way as older consumers, given the significant under-representation of Brand A preference reversing so clearly for Brand C at the older end of the age range.

For campaign planning purposes, these findings could reasonably support a simple two-tier segmentation approach rather than a fully bespoke message for each of the three age bands: one creative direction emphasising the attributes currently favoured by younger consumers, positioned primarily around Brand A, and a second direction emphasising attributes more associated with the 41+ segment, positioned primarily around Brand C. Brand B’s comparatively even appeal across all three age groups suggests it may be better served by an age-neutral campaign that instead differentiates on a variable not tested here, such as price point or usage occasion, rather than being folded into either age-specific message.

Conclusion and Limitations

This analysis found a statistically significant, small-to-moderate association between consumer age group and preferred soft drink brand among a sample of 300 UK consumers, χ²(4, N = 300) = 17.92, p = .001, Cramér’s V = .17. Brand A was disproportionately preferred by younger consumers and Brand C by older consumers, while Brand B showed relatively consistent appeal across the age range tested.

Several limitations should be borne in mind when applying these findings. First, the chi-square test establishes association, not causation; the data cannot determine why age is associated with brand preference, only that a relationship exists. Second, age was collected as a three-category band rather than as a continuous variable, which is standard practice for this type of test but does discard some information about within-band variation that a different analytical approach might have captured. Third, the sample was drawn from shopping-centre intercepts in a limited number of UK locations, which may not fully represent the national consumer population, particularly online-only shoppers who were not captured by the intercept method. Fourth, respondents who selected “no preference” were excluded from the analysis, which may have introduced a degree of selection bias if these respondents differed systematically, for instance by being less brand-loyal overall, from those who stated a clear preference. Fifth, the equal quota of 100 respondents per age band, while methodologically useful for ensuring adequate expected cell counts, does not reflect the true relative size of each age group within the wider UK population, and the findings should therefore be understood as describing the pattern of association between age and brand preference rather than the overall market share held by each brand.

Future research could usefully extend this analysis by including additional demographic and behavioural variables, such as income band, purchase frequency and channel, in a multivariate model, and by replicating the survey online to capture a broader and more representative sample. A follow-up study using a probability-proportional sample, rather than equal quotas, would also allow the findings to be more directly translated into estimates of overall market share by brand. Notwithstanding these limitations, the present findings offer the marketing team a statistically defensible basis for incorporating age into its segmentation strategy, while cautioning against treating age as the sole or dominant driver of brand choice.

References

Agresti, A. and Finlay, B. (2009) Statistical Methods for the Social Sciences. 4th edn. Upper Saddle River, NJ: Pearson.

Bryman, A. and Cramer, D. (2011) Quantitative Data Analysis with IBM SPSS 17, 18 and 19: A Guide for Social Scientists. Hove: Routledge.

Cohen, J. (1988) Statistical Power Analysis for the Behavioral Sciences. 2nd edn. Hillsdale, NJ: Lawrence Erlbaum.

Field, A. (2018) Discovering Statistics Using IBM SPSS Statistics. 5th edn. London: Sage.

Howitt, D. and Cramer, D. (2020) Introduction to Statistics in Psychology. 7th edn. Harlow: Pearson.

Pallant, J. (2020) SPSS Survival Manual. 7th edn. Maidenhead: Open University Press.

Solomon, M.R., Bamossy, G., Askegaard, S. and Hogg, M.K. (2019) Consumer Behaviour: A European Perspective. 7th edn. Harlow: Pearson.

Need a Model Statistical Analysis Written to Your Exact Brief?

Our 350+ UK-qualified writers deliver referenced model documents from £15 per 250 words, with free plagiarism and AI-detection reports.

Order Your Model Statistical Analysis

Frequently Asked Questions

About Jesse Pinkman

Avatar for Jesse PinkmanJessie Pinkman has been writing since childhood when her mother gave her a book where she could write her stories. Since then Jessie has always loved to write about the topics she loves. She graduated from Birmingham University in 2012, worked as a teaching assistant, and then turned to full-time writing in 2016.

You May Also Like

WhatsApp Live Chat