Skip to topic
    ← Back to course topics

    E12b — AQA GCSE Statistics

    Test yourself on E12b with AQA GCSE practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    Your focus

    1. Use samples to estimate population mean.

    E12b exam tips

    Quick Revision Summary (Key Takeaway)

    E12b in AQA GCSE Statistics covers the interpretation and comparison of data distributions using measures of central tendency (mean, median, mode) and measures of dispersion (range, interquartile range, standard deviation). You must be able to calculate these measures, compare two or more distributions in context, and recognise the effect of outliers on each measure.

    Topic Overview

    This topic focuses on comparing data distributions using summary statistics. You need to calculate and interpret measures of central tendency (mean, median, mode) and measures of dispersion (range, interquartile range, standard deviation). You must also understand how outliers affect these measures and choose the most appropriate statistics for a given data set.

    Comparing distributions is a fundamental skill in statistics because it allows you to make informed decisions based on data. It appears frequently in AQA GCSE exam papers, often in the context of real-world scenarios such as comparing test scores, temperatures, or waiting times. Mastering this topic also prepares you for more advanced statistical analysis at A-level and beyond.

    Key Concepts
    • →Measures of central tendency: mean (average), median (middle value), mode (most frequent). Each has strengths and weaknesses depending on the data.
    • →Measures of dispersion: range (max - min), interquartile range (IQR = UQ - LQ), standard deviation (spread around the mean). Larger values indicate greater variability.
    • →Outliers: extreme values that can skew the mean and range. The median and IQR are resistant to outliers and often preferred when outliers are present.
    • →Comparing distributions: always compare both an average and a measure of spread, and relate your comparison back to the context of the data.
    • →Standard deviation: a measure of how far data values are from the mean. A small standard deviation means data is clustered around the mean; a large one means data is more spread out.
    Examiner Tips
    • 💡Always write your comparisons in the context of the question. For example, say 'the median test score for Class A is higher than for Class B' rather than 'the median is higher'.
    • 💡When comparing two distributions, aim to make two comparisons: one for average and one for spread. This is often worth two marks.
    • 💡If you are asked to calculate standard deviation, show your working clearly, including the mean and the squared deviations, as method marks are available even if the final answer is wrong.
    Common Mistakes
    • Students often think the mean is always the best measure of average. Correction: When outliers are present, the median is often more representative because it is not affected by extreme values.
    • Students confuse the interquartile range with the range. Correction: The range is the difference between the maximum and minimum values, while the IQR is the difference between the upper and lower quartiles and covers the middle 50% of data.
    • Students believe a larger standard deviation means the data is 'better' or 'higher'. Correction: Standard deviation measures spread, not the value of the data. A larger standard deviation simply means the data is more spread out from the mean.
    Revision Plan
    1. 1Day 1-2: Revise the definitions and calculations for mean, median, mode, range, and IQR. Practise with small data sets and frequency tables.
    2. 2Day 3-4: Learn how to calculate standard deviation step by step. Use the formula: standard deviation = sqrt(sum of squared deviations / n). Practise with at least five different data sets.
    3. 3Day 5-6: Focus on comparing distributions. Use past paper questions to practise writing comparison statements that include both an average and a measure of spread, in context.
    4. 4Day 7-8: Study the effect of outliers. Identify outliers in data sets and explain which measures are most appropriate. Practise justifying your choice.
    5. 5Day 9-10: Complete a full past paper question on this topic under timed conditions. Review your answers using the mark scheme and note any recurring mistakes.
    Exam Question Types
    • 📋Calculation questions: Calculate the mean, median, mode, range, IQR, or standard deviation from a list or table. Advice: Show all steps, especially for standard deviation, and double-check your arithmetic.
    • 📋Comparison questions: Compare two distributions using appropriate measures. Advice: Always comment on both average and spread, and use comparative language in context.
    • 📋Outlier questions: Identify outliers and explain their effect on summary statistics. Advice: State which measures are affected and which are resistant, and recommend the best measure to use.
    • 📋Interpretation questions: Given a set of statistics, interpret what they mean in context. Advice: Relate each statistic back to the real-world scenario and avoid simply restating numbers.
    Command Word Expectations (AQA)
    Calculate

    You must work out a numerical answer. Show your method clearly, as method marks are available. Give your answer to the required degree of accuracy, and include units if applicable.

    Compare

    You must describe similarities and differences between two or more distributions. Make at least two comparisons, typically one for average and one for spread, and always refer to the context. Use comparative language such as 'higher', 'lower', 'more consistent'.

    Explain

    You must give reasons for your answer. This often involves justifying why a particular measure is more appropriate, such as saying the median is better than the mean because it is not affected by outliers. Use because or as to link your reasoning.

    How Students Lose Marks (Examiner Pitfalls)
    Pitfall: Students calculate the mean and range correctly but fail to compare the two distributions in context, or they compare only one measure and ignore spread.
    ❌ Weak Answer (Loses Marks):The mean for Town A is 12.4 and the mean for Town B is 15.2. Town B has a higher mean.
    Example improved answer:On average, Town B has a higher rainfall (mean = 15.2 mm) than Town A (mean = 12.4 mm). However, Town B is more variable (IQR = 8.5 mm) compared with Town A (IQR = 4.2 mm), so Town B's rainfall is less consistent.
    Examiner Tip: Always compare both an average and a measure of spread, and always refer back to the original context (e.g. 'rainfall', 'test scores'). Use comparative language such as 'higher', 'more consistent', 'less variable'.
    Pitfall: Students use the wrong measure of spread when comparing distributions with outliers, or they fail to identify that the median and IQR are more appropriate than the mean and range in the presence of extreme values.
    ❌ Weak Answer (Loses Marks):The range for Shop A is 45 and the range for Shop B is 12, so Shop A is more spread out. The mean is also higher for Shop A.
    Example improved answer:Shop A has an outlier of 52, so the median and interquartile range are more appropriate measures. The median for Shop A is 18 (IQR = 9) and for Shop B is 20 (IQR = 6). Shop B has a slightly higher typical value and is more consistent, as its IQR is smaller.
    Examiner Tip: If a data set contains an outlier, state that the median and IQR are better measures because they are not affected by extreme values. Always justify your choice of measure.
    Step-by-Step Worked Solutions

    Question: The table shows the daily maximum temperatures (in degrees Celsius) for two cities over 10 days. City A: 12, 15, 14, 13, 16, 15, 14, 13, 15, 17. City B: 8, 22, 10, 25, 9, 20, 11, 24, 7, 26. Compare the two distributions using an appropriate measure of average and spread.

    1. 1.Step 1: Calculate the mean for City A: (12+15+14+13+16+15+14+13+15+17) / 10 = 144 / 10 = 14.4 degrees Celsius.
    2. 2.Step 2: Calculate the mean for City B: (8+22+10+25+9+20+11+24+7+26) / 10 = 162 / 10 = 16.2 degrees Celsius.
    3. 3.Step 3: Identify that City B has extreme values (7 and 26), so the median and IQR are more appropriate. Order City A: 12, 13, 13, 14, 14, 15, 15, 15, 16, 17. Median = (14+15)/2 = 14.5. Lower quartile = 13, Upper quartile = 15.5, IQR = 2.5. Order City B: 7, 8, 9, 10, 11, 20, 22, 24, 25, 26. Median = (11+20)/2 = 15.5. Lower quartile = 9, Upper quartile = 24, IQR = 15.
    4. 4.Step 4: Compare: City B has a slightly higher median temperature (15.5 vs 14.5) but a much larger IQR (15 vs 2.5), meaning City B's temperatures are far more variable and less consistent than City A's.
    Final Answer: City B has a higher median temperature (15.5 degrees Celsius) compared to City A (14.5 degrees Celsius), but City B is much more variable (IQR = 15) than City A (IQR = 2.5). City A has more consistent daily temperatures.

    Question: A student records the number of minutes spent revising per day for a sample of 8 students: 30, 45, 60, 25, 90, 40, 35, 55. Calculate the mean, median, and standard deviation (to 1 decimal place). Comment on the effect of the value 90 on these measures.

    1. 1.Step 1: Calculate the mean: (30+45+60+25+90+40+35+55) / 8 = 380 / 8 = 47.5 minutes.
    2. 2.Step 2: Order the data: 25, 30, 35, 40, 45, 55, 60, 90. Median = (40+45)/2 = 42.5 minutes.
    3. 3.Step 3: Calculate standard deviation: First find deviations from mean: -17.5, -2.5, 12.5, -22.5, 42.5, -7.5, -12.5, 7.5. Square these: 306.25, 6.25, 156.25, 506.25, 1806.25, 56.25, 156.25, 56.25. Sum of squares = 3050. Variance = 3050 / 8 = 381.25. Standard deviation = sqrt(381.25) = 19.5 minutes (1 d.p.).
    4. 4.Step 4: Comment: The value 90 is an outlier. It increases the mean (47.5) and standard deviation (19.5) substantially, but the median (42.5) is less affected. The median is a better measure of typical revision time.
    Final Answer: Mean = 47.5 minutes, Median = 42.5 minutes, Standard deviation = 19.5 minutes. The outlier of 90 inflates the mean and standard deviation, so the median is a more representative average.
    Active Recall Memory Test
    What is the formula for standard deviation?
    Key Fact: Standard deviation = sqrt( sum of (x - mean)^2 / n ), where n is the number of data values.
    When is the median a better measure of average than the mean?
    Key Fact: The median is better when the data contains outliers or is skewed, because it is not affected by extreme values.
    What does a large interquartile range tell you about a data set?
    Key Fact: A large interquartile range indicates that the middle 50% of the data is spread out over a wide range, meaning the data is more variable.
    How do you calculate the interquartile range?
    Key Fact: IQR = Upper Quartile - Lower Quartile. It measures the spread of the middle 50% of the data.
    Frequently Asked Questions
    What is the difference between standard deviation and range?
    The range is the difference between the maximum and minimum values and only considers the two most extreme values. Standard deviation measures how far, on average, each data value is from the mean, using all data points. Standard deviation is generally a more reliable measure of spread because it takes every value into account, whereas the range can be heavily influenced by outliers.
    How do I know which measure of average to use?
    Use the mean when the data is roughly symmetrical and has no outliers, as it uses all data values. Use the median when the data is skewed or has outliers, as it is not affected by extreme values. Use the mode when dealing with categorical data or when you need the most frequent value. Always consider the context and whether outliers are present.
    What is an outlier and how does it affect the mean?
    An outlier is a data value that is unusually far from the rest of the data. It can significantly increase or decrease the mean because the mean uses every value in its calculation. For example, in the data set 2, 3, 4, 5, 100, the mean is 22.8, which is not representative of the typical value. The median (4) is a better measure in this case.
    How do I compare two distributions in an exam?
    To compare two distributions, first calculate an appropriate measure of average (mean or median) and a measure of spread (range, IQR, or standard deviation) for each. Then write comparative statements in context, such as 'The median score for Group A is higher than for Group B, but Group B has a larger interquartile range, so Group A is more consistent.' Always make at least two comparisons: one for average and one for spread.
    What does a standard deviation of 0 mean?
    A standard deviation of 0 means that all data values are identical and equal to the mean. There is no variability in the data set. For example, if every student scored 70 on a test, the mean is 70 and the standard deviation is 0.
    Can I use the range to compare distributions if there are outliers?
    You can calculate the range, but it is not a reliable measure of spread when outliers are present because it only depends on the two most extreme values. In such cases, the interquartile range is a better choice because it focuses on the middle 50% of the data and is not affected by outliers. Always justify your choice of measure in your answer.