Skip to topic
    ← Back to course topics

    C4b — AQA GCSE Statistics

    Test yourself on C4b with AQA GCSE practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    Your focus

    1. Select appropriate form of representation.

    C4b exam tips

    Quick Revision Summary (Key Takeaway)

    C4b in AQA GCSE Statistics covers the comparison of two data sets using measures of location (mean, median, mode) and measures of spread (range, interquartile range, standard deviation). It requires students to calculate these statistics, interpret differences in context, and recognise the effect of outliers on each measure.

    Topic Overview

    C4b is a core topic in AQA GCSE Statistics that focuses on comparing two data sets using both measures of location and measures of spread. You will calculate the mean, median, mode, range, interquartile range, and standard deviation, and then use these to make meaningful comparisons in context. This topic is essential because it allows you to draw conclusions about differences between groups, such as comparing test scores, reaction times, or product lifespans.

    Understanding C4b is crucial for the wider subject because it underpins statistical inference and hypothesis testing. It also appears frequently in exam questions that require you to interpret data rather than just calculate. Mastering this topic will help you answer higher-mark questions that demand clear, contextual comparisons and an awareness of how outliers affect different statistics.

    Key Concepts
    • →Measures of location (mean, median, mode) summarise the centre of a data set, while measures of spread (range, IQR, standard deviation) describe how varied the data are.
    • →The mean is affected by outliers, but the median is resistant to outliers; the range is affected by outliers, but the IQR is resistant.
    • →Standard deviation measures the average distance of each data point from the mean; a larger standard deviation indicates greater spread.
    • →When comparing two data sets, you must compare both a measure of location and a measure of spread to draw a valid conclusion.
    • →The interquartile range (IQR) is the range of the middle 50% of the data and is calculated as upper quartile minus lower quartile.
    Examiner Tips
    • 💡Always write your comparisons in the context of the question. For example, say 'the girls' scores are more consistent' rather than 'the IQR is smaller'.
    • 💡When calculating standard deviation, show your working clearly, especially the squared differences from the mean, as method marks are often available.
    • 💡Use the correct statistical terminology: 'mean', 'median', 'interquartile range', 'standard deviation', and 'outlier'. Avoid vague words like 'average' without specifying which one.
    Common Mistakes
    • Students often think that a higher mean always means the data set is better, but this depends on the context; for example, a higher mean waiting time is worse.
    • Students sometimes believe that the median is always the middle value when data are listed, but for an even number of values, it is the mean of the two middle values.
    • Students may confuse standard deviation with range; standard deviation considers all data points, while range only uses the maximum and minimum.
    Revision Plan
    1. 1Day 1-2: Revise how to calculate each measure of location and spread from raw data and frequency tables. Practice with at least 10 problems.
    2. 2Day 3-4: Learn how to calculate standard deviation using the formula or a calculator. Focus on interpreting what the value means in context.
    3. 3Day 5-6: Practice comparing two data sets by writing full sentences that include both a measure of location and a measure of spread. Use past paper questions.
    4. 4Day 7-8: Review examiner reports and mark schemes to understand common pitfalls and how to structure answers for full marks.
    5. 5Day 9-10: Complete a timed past paper section on C4b and self-assess using the mark scheme, focusing on clarity and context.
    Exam Question Types
    • 📋Calculation questions: You may be asked to calculate the mean, median, mode, range, IQR, or standard deviation from a list or table. Show all steps clearly.
    • 📋Comparison questions: You will be given summary statistics for two groups and asked to compare them. Always comment on both location and spread, and relate to the context.
    • 📋Interpretation questions: You may be asked to explain the effect of an outlier on the mean and standard deviation, or to decide which average is most appropriate. Justify your answer.
    • 📋Graph-based questions: You may need to read data from a box plot or cumulative frequency graph to find quartiles and then compare distributions.
    Command Word Expectations (AQA)
    Calculate

    You must work out a numerical value using the given data. Show your method, especially for standard deviation, as method marks are awarded. Give your answer to an appropriate degree of accuracy.

    Compare

    You must state similarities and differences between two data sets, using both a measure of location and a measure of spread. Use comparative language and refer to the context. Typically worth 2-4 marks.

    Explain

    You must give reasons for your answer, often referring to the effect of outliers or the meaning of a statistic. Use full sentences and correct terminology. Usually worth 2-3 marks.

    How Students Lose Marks (Examiner Pitfalls)
    Pitfall: Students often compare two data sets by stating only the averages without mentioning spread, or they quote values without interpreting them in the context of the question.
    ❌ Weak Answer (Loses Marks):The mean for boys is 12.4 and the mean for girls is 14.1, so girls did better.
    Example improved answer:The mean score for girls (14.1) is higher than the mean score for boys (12.4), suggesting girls performed better on average. However, the standard deviation for boys (3.2) is greater than for girls (1.8), indicating that boys' scores are more spread out and less consistent than girls' scores.
    Examiner Tip: Always compare both an average and a measure of spread, and relate your comparison back to the real-life context given in the question. Use comparative language such as 'higher', 'more consistent', or 'less varied'.
    Pitfall: When calculating the interquartile range (IQR), students frequently forget to order the data first or misidentify the median when there is an even number of values.
    ❌ Weak Answer (Loses Marks):IQR = 18 - 7 = 11 (without ordering the data or correctly finding quartiles).
    Example improved answer:First, order the data: 3, 5, 7, 8, 10, 12, 15, 18, 20, 22. There are 10 values, so the median is the mean of the 5th and 6th values: (10 + 12) / 2 = 11. The lower quartile is the median of the lower half (3, 5, 7, 8, 10), which is 7. The upper quartile is the median of the upper half (12, 15, 18, 20, 22), which is 18. Therefore, IQR = 18 - 7 = 11.
    Examiner Tip: Always re-order the data before finding quartiles. For an even number of values, split the data into two equal halves and find the median of each half. If there is an odd number of values, exclude the median from both halves.
    Step-by-Step Worked Solutions

    Question: Two classes, 10A and 10B, sit the same test. The summary statistics are: 10A: mean = 65, median = 64, range = 40, IQR = 18. 10B: mean = 68, median = 70, range = 25, IQR = 10. Compare the performance of the two classes.

    1. 1.Step 1: Identify the given facts: 10A has mean 65, median 64, range 40, IQR 18. 10B has mean 68, median 70, range 25, IQR 10.
    2. 2.Step 2: Compare measures of location: The mean for 10B (68) is higher than for 10A (65), and the median for 10B (70) is also higher than for 10A (64). This suggests that, on average, class 10B performed better than class 10A.
    3. 3.Step 3: Compare measures of spread: The range for 10A (40) is larger than for 10B (25), and the IQR for 10A (18) is larger than for 10B (10). This indicates that the scores in 10A are more spread out and less consistent than the scores in 10B.
    4. 4.Step 4: State final conclusion: Class 10B achieved higher scores on average and had more consistent results, while class 10A had a wider variation in performance.
    Final Answer: Class 10B performed better on average (higher mean and median) and had more consistent scores (smaller range and IQR) compared to class 10A.

    Question: A data set of 8 values has a mean of 15 and a standard deviation of 2.5. A new value of 30 is added to the data set. Without calculating the new standard deviation, explain how the standard deviation is likely to change and why.

    1. 1.Step 1: Identify given facts: Original mean = 15, standard deviation = 2.5, new value = 30.
    2. 2.Step 2: Apply understanding of standard deviation: Standard deviation measures the spread of data around the mean. A value of 30 is much higher than the original mean of 15, so it is an outlier.
    3. 3.Step 3: Reason the effect: Adding an outlier increases the spread of the data, so the standard deviation will increase. The new mean will also increase, but the new value is still far from the new mean, so the overall spread increases.
    4. 4.Step 4: State final conclusion: The standard deviation will increase because the new value is far from the mean, increasing the overall spread of the data.
    Final Answer: The standard deviation will increase because the new value of 30 is an outlier that increases the spread of the data.
    Active Recall Memory Test
    What is the difference between the range and the interquartile range?
    Key Fact: The range is the difference between the maximum and minimum values, while the interquartile range is the difference between the upper quartile and lower quartile, covering the middle 50% of the data. The range is affected by outliers, but the IQR is not.
    How does an outlier affect the mean and the median?
    Key Fact: An outlier increases or decreases the mean significantly because it uses all values, but it has little effect on the median because the median depends only on the middle value(s).
    What does a standard deviation of 0 indicate?
    Key Fact: A standard deviation of 0 means all data values are identical; there is no spread.
    When comparing two data sets, what two types of measure should you always include?
    Key Fact: You should include a measure of location (such as the mean or median) and a measure of spread (such as the range, IQR, or standard deviation).
    Frequently Asked Questions
    What is the difference between standard deviation and interquartile range?
    Standard deviation measures the average distance of each data point from the mean, using all data values. The interquartile range measures the spread of the middle 50% of the data, ignoring the lowest and highest 25%. Standard deviation is more sensitive to outliers, while the IQR is resistant to them.
    How do I know which average to use when comparing data sets?
    Use the mean when the data are roughly symmetrical and there are no outliers, as it uses all values. Use the median when the data are skewed or contain outliers, as it is not affected by extreme values. Always consider the context and what you are trying to show.
    What does it mean if two data sets have the same mean but different standard deviations?
    It means that on average the values are the same, but the spread of the data is different. The set with the larger standard deviation has values that are more spread out from the mean, while the set with the smaller standard deviation has values that are more clustered around the mean.
    Can the standard deviation be negative?
    No, the standard deviation is always zero or positive because it is the square root of the variance, which is the average of squared differences. A standard deviation of zero means all data values are the same.
    How do I calculate the interquartile range from a list of data?
    First order the data from smallest to largest. Find the median (middle value). Then find the lower quartile (median of the lower half) and the upper quartile (median of the upper half). The IQR is upper quartile minus lower quartile. If there is an odd number of values, exclude the median from both halves.
    Why is it important to compare both averages and measures of spread?
    Comparing only averages can be misleading because two data sets can have the same average but very different spreads. For example, one set might be consistent while the other is highly variable. Including a measure of spread gives a fuller picture of the differences between the data sets.