Skip to topic
    ← Back to course topics

    E1b — AQA GCSE Statistics

    Test yourself on E1b with AQA GCSE practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    Your focus

    1. Use probability values to calculate expected frequency of a specified characteristic within a sample or population.

    E1b exam tips

    Quick Revision Summary (Key Takeaway)

    E1b in AQA GCSE Statistics covers the interpretation and comparison of summary statistics (mean, median, mode, range, quartiles, interquartile range, and standard deviation) from tabulated data and graphs. You must calculate these measures accurately, compare distributions using them, and explain what they reveal about the data in context.

    Topic Overview

    E1b is a core topic in AQA GCSE Statistics that focuses on summarising and comparing data sets using numerical measures. You will learn to calculate and interpret the mean, median, mode, range, quartiles, interquartile range, and standard deviation. These measures help you describe the central tendency and spread of data, which is essential for making comparisons and drawing conclusions in real-world contexts.

    This topic is fundamental because it underpins many other areas of statistics, such as hypothesis testing and data analysis. In the exam, you will often be asked to compare two distributions using these statistics, and to explain which measure is most appropriate given the nature of the data. Mastering E1b will also help you in other subjects like science and geography, where data interpretation is key.

    Key Concepts
    • →Measures of central tendency: mean (average), median (middle value), and mode (most frequent value). The mean is affected by outliers, while the median is more robust.
    • →Measures of spread: range (max - min), interquartile range (IQR = Q3 - Q1), and standard deviation (a measure of how far data values are from the mean).
    • →Quartiles: lower quartile (Q1) is the median of the lower half of the data; upper quartile (Q3) is the median of the upper half. The IQR represents the spread of the middle 50% of the data.
    • →Standard deviation: a measure of spread that uses all data values. A small standard deviation indicates data is clustered around the mean; a large one indicates data is more spread out.
    • →When comparing distributions, always comment on both a measure of average and a measure of spread, and consider whether outliers affect the choice of measure.
    Examiner Tips
    • 💡Always show your working, especially for the mean and standard deviation. Even if your final answer is wrong, you can gain method marks for correct steps.
    • 💡When asked to compare distributions, use comparative language such as 'higher than', 'more consistent', 'less spread out'. Simply stating the values without comparison will not gain full marks.
    • 💡Interpret your results in the context of the question. For example, if the data is about waiting times, say 'the average waiting time is 12 minutes' rather than just 'the mean is 12'.
    Common Mistakes
    • Students often think the mean is always the best measure of average. However, when data contains outliers or is skewed, the median is often more representative. Always check the shape of the distribution.
    • Students sometimes confuse the interquartile range with the range. The range is the difference between the maximum and minimum values, while the IQR is the difference between the upper and lower quartiles, ignoring the extremes.
    • When calculating standard deviation, students may forget to square the deviations before summing, or they may divide by n instead of n-1. Remember that for a sample, you divide by n-1 to get an unbiased estimate.
    Revision Plan
    1. 1Day 1-2: Revise the definitions and formulas for mean, median, mode, range, quartiles, IQR, and standard deviation. Use flashcards to memorise them.
    2. 2Day 3-4: Practice calculating these statistics from small data sets (5-10 values). Check your answers using a calculator or online tool.
    3. 3Day 5-6: Work through exam-style questions that require comparing two distributions. Focus on writing clear comparative sentences and interpreting in context.
    4. 4Day 7-8: Tackle standard deviation calculations step by step. Ensure you understand the formula and can apply it accurately.
    5. 5Day 9-10: Complete a past paper section on E1b under timed conditions. Review your mistakes and revisit any weak areas.
    Exam Question Types
    • 📋Calculation questions: 'Calculate the mean and standard deviation for the following data.' Advice: Show all steps, use the correct formula, and round appropriately (usually 2 decimal places).
    • 📋Comparison questions: 'Compare the distribution of scores for Class A and Class B.' Advice: Use both a measure of average and a measure of spread, and make direct comparisons using comparative language.
    • 📋Interpretation questions: 'Explain why the median is a better measure of average than the mean in this context.' Advice: Refer to outliers or skewness and explain how they affect the mean but not the median.
    • 📋Graph-based questions: 'Using the box plot, compare the two distributions.' Advice: Read the median, quartiles, and extremes from the box plot, then compare the IQR and range.
    Command Word Expectations (AQA)
    Calculate

    You must work out a numerical answer using the correct method. Show all steps of your working, as method marks are available. Give your answer to an appropriate degree of accuracy (usually 2 decimal places for standard deviation).

    Compare

    You must describe similarities and differences between two or more sets of data. Use comparative language (e.g., 'higher', 'lower', 'more consistent') and refer to specific statistics such as the mean, median, or IQR. You must comment on both average and spread to gain full marks.

    Interpret

    You must explain what a calculated statistic means in the context of the problem. For example, 'The mean of 12.5 indicates that on average, there were 12.5 goals per match.' Do not just restate the number; relate it to the real-world scenario.

    How Students Lose Marks (Examiner Pitfalls)
    Pitfall: Students often calculate the mean correctly but fail to interpret it in the context of the question, or they confuse the mean and median when comparing distributions. Another common error is using the range instead of the interquartile range when the data contains outliers.
    ❌ Weak Answer (Loses Marks):The mean is 12.5 and the median is 12, so the mean is higher. The range is 20, so the data is spread out.
    Example improved answer:The mean number of goals scored per match is 12.5, while the median is 12. This suggests the distribution is slightly positively skewed, as the mean is greater than the median. The interquartile range is 6, indicating that the middle 50% of matches had goal totals within a spread of 6 goals. The range of 20 is affected by an outlier (a match with 30 goals), so the interquartile range is a more reliable measure of spread.
    Examiner Tip: Always relate your calculation back to the real-world context given in the question. When comparing two distributions, comment on both a measure of average (mean or median) and a measure of spread (IQR or standard deviation), and state which is more appropriate if outliers are present.
    Pitfall: When calculating standard deviation, students often forget to square the deviations before summing, or they divide by n instead of n-1 for a sample. In the exam, you must use the formula given in the AQA formula sheet, which uses n-1 for sample standard deviation.
    ❌ Weak Answer (Loses Marks):I added up the differences from the mean and divided by 5 to get the standard deviation.
    Example improved answer:First, calculate the mean. Then subtract the mean from each value, square each result, sum the squares, and divide by (n - 1). Finally, take the square root. For the data set 4, 8, 6, 5, 7: mean = 6. Deviations: -2, 2, 0, -1, 1. Squares: 4, 4, 0, 1, 1. Sum = 10. Divide by 4 (n-1) = 2.5. Standard deviation = sqrt(2.5) = 1.58 (2 d.p.).
    Examiner Tip: Write down each step clearly, especially the squared deviations. Use the exact formula from the AQA formula sheet. Check whether the data is a sample or a population; for GCSE Statistics, you will usually treat data as a sample unless told otherwise.
    Step-by-Step Worked Solutions

    Question: The table shows the number of cars sold by a dealership over 10 days: 5, 8, 6, 10, 7, 9, 4, 8, 7, 6. Calculate the mean, median, mode, range, and interquartile range. Interpret these statistics in the context of the data.

    1. 1.Step 1: Identify the data set and sort it in ascending order: 4, 5, 6, 6, 7, 7, 8, 8, 9, 10.
    2. 2.Step 2: Calculate the mean: sum = 4+5+6+6+7+7+8+8+9+10 = 70. Mean = 70 / 10 = 7 cars.
    3. 3.Step 3: Find the median: with 10 values, the median is the average of the 5th and 6th values: (7 + 7) / 2 = 7 cars.
    4. 4.Step 4: Identify the mode: the values 6, 7, and 8 each appear twice, so the data is trimodal with modes 6, 7, and 8 cars.
    5. 5.Step 5: Calculate the range: maximum - minimum = 10 - 4 = 6 cars.
    6. 6.Step 6: Calculate the interquartile range: lower quartile (Q1) is the median of the lower half (4,5,6,6,7) = 6. Upper quartile (Q3) is the median of the upper half (7,8,8,9,10) = 8. IQR = Q3 - Q1 = 8 - 6 = 2 cars.
    7. 7.Step 7: Interpret: The mean and median are both 7 cars, suggesting a symmetric distribution. The IQR of 2 cars indicates that the middle 50% of days had sales within a narrow range of 2 cars, showing consistent sales. The range of 6 cars shows the full spread from 4 to 10 cars.
    Final Answer: Mean = 7 cars, Median = 7 cars, Modes = 6, 7, and 8 cars, Range = 6 cars, IQR = 2 cars. The data is fairly symmetric and consistent, with typical sales around 7 cars per day.

    Question: Two football teams, Team A and Team B, have the following goals scored per match over 8 matches. Team A: 2, 3, 1, 4, 2, 3, 2, 1. Team B: 0, 5, 1, 6, 0, 4, 1, 3. Compare the distributions using appropriate summary statistics and comment on which team is more consistent.

    1. 1.Step 1: Calculate the mean for Team A: sum = 2+3+1+4+2+3+2+1 = 18. Mean = 18/8 = 2.25 goals.
    2. 2.Step 2: Calculate the mean for Team B: sum = 0+5+1+6+0+4+1+3 = 20. Mean = 20/8 = 2.5 goals.
    3. 3.Step 3: Calculate the median for Team A: sorted: 1,1,2,2,2,3,3,4. Median = (2+2)/2 = 2 goals.
    4. 4.Step 4: Calculate the median for Team B: sorted: 0,0,1,1,3,4,5,6. Median = (1+3)/2 = 2 goals.
    5. 5.Step 5: Calculate the interquartile range for Team A: Q1 = median of lower half (1,1,2,2) = 1.5; Q3 = median of upper half (2,3,3,4) = 3; IQR = 3 - 1.5 = 1.5 goals.
    6. 6.Step 6: Calculate the interquartile range for Team B: Q1 = median of lower half (0,0,1,1) = 0.5; Q3 = median of upper half (3,4,5,6) = 4.5; IQR = 4.5 - 0.5 = 4 goals.
    7. 7.Step 7: Compare: Team B has a slightly higher mean (2.5 vs 2.25) but the same median (2). Team A has a much smaller IQR (1.5 vs 4), indicating more consistent performance. Team B's scores are more spread out, with a higher range (6 vs 3).
    Final Answer: Team A is more consistent as it has a smaller interquartile range (1.5) compared to Team B (4). Although Team B has a slightly higher mean, its performance is more variable. Both teams have the same median of 2 goals.
    Active Recall Memory Test
    What is the formula for the interquartile range?
    Key Fact: IQR = Upper Quartile (Q3) - Lower Quartile (Q1).
    When should you use the median instead of the mean?
    Key Fact: Use the median when the data contains outliers or is skewed, as it is not affected by extreme values.
    What does a small standard deviation indicate about a data set?
    Key Fact: A small standard deviation indicates that the data values are clustered closely around the mean, showing low variability.
    How do you calculate the mean of a data set?
    Key Fact: Add up all the values and divide by the number of values.
    Frequently Asked Questions
    What is the difference between the range and the interquartile range?
    The range is the difference between the maximum and minimum values in a data set, so it is affected by outliers. The interquartile range (IQR) is the difference between the upper quartile (Q3) and lower quartile (Q1), representing the spread of the middle 50% of the data. The IQR is more robust because it ignores extreme values, making it a better measure of spread when outliers are present.
    How do I calculate standard deviation in AQA GCSE Statistics?
    To calculate standard deviation, first find the mean. Then subtract the mean from each value and square the result. Sum these squared deviations, divide by (n - 1) for a sample, and finally take the square root. The formula is: s = sqrt( sum of (x - mean)^2 / (n - 1) ). You will be given this formula in the exam, but you must know how to apply it step by step.
    What does it mean if the mean is greater than the median?
    If the mean is greater than the median, the distribution is positively skewed (skewed to the right). This means there are some high values pulling the mean upwards, while the median remains more central. In this case, the median is often a better measure of average because it is not affected by the extreme high values.
    How do I compare two distributions in an exam?
    To compare two distributions, comment on both a measure of average (mean or median) and a measure of spread (range, IQR, or standard deviation). Use comparative language such as 'higher than', 'lower than', 'more consistent', or 'more spread out'. Always relate your comparison back to the context of the question. For example, 'Team A has a lower median score than Team B, indicating they typically scored fewer points, but Team A also has a smaller interquartile range, showing their scores were more consistent.'
    Why is the median sometimes better than the mean?
    The median is better than the mean when the data contains outliers or is skewed. Outliers can significantly affect the mean, pulling it away from the typical value, while the median is based on the middle value and is not affected by extreme values. For example, in income data where a few people earn very high salaries, the mean income would be inflated, but the median income would better represent the typical person's earnings.
    What is the mode and when is it most useful?
    The mode is the value that appears most frequently in a data set. It is most useful for categorical data (e.g., favourite colour) or when you want to know the most common value. For numerical data, the mode is less commonly used than the mean or median, but it can be helpful in identifying the most popular item or the most frequent occurrence. A data set can have one mode, multiple modes, or no mode if all values appear equally often.