Skip to topic
    ← Back to course topics

    B3e — AQA GCSE Statistics

    Test yourself on B3e with AQA GCSE practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    Your focus

    1. Know the key features of a simple random sample.

    B3e exam tips

    Quick Revision Summary (Key Takeaway)

    B3e in AQA GCSE Statistics covers the interpretation and comparison of data distributions using measures of central tendency (mean, median, mode) and measures of dispersion (range, interquartile range, standard deviation). It also includes identifying outliers, constructing and interpreting box plots, and making informed comparisons between data sets in real-world contexts.

    Topic Overview

    B3e focuses on summarising and comparing data sets using numerical measures. You will learn to calculate and interpret measures of central tendency (mean, median, mode) and measures of dispersion (range, interquartile range, standard deviation). These tools allow you to describe the typical value and the spread of data, which is essential for making informed decisions in real-world contexts such as comparing schools, hospitals, or products.

    This topic also covers the identification of outliers and the construction and interpretation of box plots. Box plots provide a visual summary of the median, quartiles, and extremes, making it easy to compare distributions side by side. Understanding these concepts is crucial for the AQA GCSE Statistics exam, where you will be asked to perform calculations, interpret results, and write comparative statements about data sets.

    Key Concepts
    • →Measures of central tendency: mean (average), median (middle value), and mode (most frequent value) each summarise a data set with a single typical value.
    • →Measures of dispersion: range (max - min), interquartile range (IQR = Q3 - Q1), and standard deviation (spread around the mean) quantify how spread out the data are.
    • →Outliers: values that lie unusually far from the rest of the data, often defined as more than 1.5 * IQR below Q1 or above Q3, or more than 2 standard deviations from the mean.
    • →Box plots: a graphical representation showing median, lower quartile, upper quartile, and minimum/maximum values (or outliers), useful for comparing distributions.
    • →Comparing distributions: when comparing two or more data sets, you must comment on both a measure of central tendency and a measure of spread, using comparative language.
    Examiner Tips
    • 💡Always show your working for calculations, especially for standard deviation, as method marks are available even if the final answer is wrong.
    • 💡When asked to compare data sets, structure your answer with two clear points: one comparing averages and one comparing spread. Use words like 'higher', 'lower', 'more consistent', 'more varied'.
    • 💡For box plot questions, ensure you identify the correct values for median, quartiles, and extremes from the given data or graph. Label your box plot clearly.
    Common Mistakes
    • Students often think the mean is always the best measure of average. However, the mean is affected by outliers, so the median may be more representative for skewed data.
    • Students sometimes confuse standard deviation with range. Standard deviation measures the spread around the mean, while range only considers the extremes and ignores the middle values.
    • When comparing box plots, students may only compare medians and forget to compare interquartile ranges or overall spread, missing key differences in consistency.
    Revision Plan
    1. 1Day 1-2: Revise definitions and formulas for mean, median, mode, range, and interquartile range. Practice calculating these from small data sets.
    2. 2Day 3-4: Learn how to calculate standard deviation step by step. Use a table to organise your work and practice with at least five different data sets.
    3. 3Day 5-6: Study outliers: learn the 1.5 * IQR rule and the 2 standard deviations rule. Practice identifying outliers in given data.
    4. 4Day 7-8: Master box plots: learn how to draw them from raw data and how to interpret them. Practice comparing two box plots by writing comparative statements.
    5. 5Day 9-10: Complete past paper questions on B3e. Focus on comparison questions and standard deviation calculations. Review your mistakes and revisit weak areas.
    Exam Question Types
    • 📋Calculation questions: ask you to compute mean, median, mode, range, IQR, or standard deviation from a list or frequency table. Advice: show all steps, use a table for standard deviation, and double-check arithmetic.
    • 📋Comparison questions: present two data sets (often as box plots or summary statistics) and ask you to compare them. Advice: always comment on both an average and a measure of spread, using comparative language.
    • 📋Outlier identification: give a data set and ask whether a value is an outlier. Advice: state the rule you are using (e.g. 1.5 * IQR) and show the boundary calculations.
    • 📋Box plot construction: provide a data set and ask you to draw a box plot. Advice: accurately find median, quartiles, and extremes; use a ruler and label the axis clearly.
    Command Word Expectations (AQA)
    Calculate

    You must work out a numerical answer. Show all steps of your working, as method marks are awarded. Give your final answer with appropriate units if applicable.

    Compare

    You must describe similarities and differences between two or more data sets. Make at least two distinct points: one about an average (mean or median) and one about spread (range or IQR). Use comparative language such as 'higher than', 'lower than', 'more consistent'.

    Explain

    You must give reasons for your answer, often referring to the context. For example, explain why the median is a better measure than the mean when there are outliers. Use because or therefore to link your reasoning.

    How Students Lose Marks (Examiner Pitfalls)
    Pitfall: Students often compare distributions by stating only one measure (e.g. 'the mean is higher') without referencing both a measure of central tendency and a measure of spread, which limits marks to 1 out of 2 or 3.
    ❌ Weak Answer (Loses Marks):The mean of school A is higher than school B, so school A is better.
    Example improved answer:The median score for school A (72) is higher than for school B (65), indicating that students in school A generally performed better. Additionally, the interquartile range for school A (10) is lower than for school B (18), showing that the middle 50% of scores in school A are more consistent. Therefore, school A has both higher typical performance and greater consistency.
    Examiner Tip: Always make two distinct comparison points: one about an average (mean or median) and one about spread (range or IQR). Use comparative language such as 'higher than', 'lower than', 'more consistent', 'more varied'.
    Pitfall: When calculating standard deviation, students frequently forget to square the deviations before summing, or they divide by n instead of n-1 for a sample, leading to an incorrect value.
    ❌ Weak Answer (Loses Marks):Standard deviation = sum of (x - mean) divided by n.
    Example improved answer:Standard deviation = sqrt( sum of (x - mean)^2 / (n - 1) ) for a sample. First calculate the mean, then subtract the mean from each value and square the result. Sum these squared deviations, divide by (n - 1), and finally take the square root.
    Examiner Tip: Write out the formula clearly at the start of your calculation. Use a table to organise x, x - mean, and (x - mean)^2 to avoid arithmetic errors. Remember that for a sample you divide by n - 1, but for a population you divide by n.
    Step-by-Step Worked Solutions

    Question: The ages of 10 members of a gardening club are: 12, 15, 18, 20, 22, 25, 28, 30, 35, 40. Calculate the mean, median, and interquartile range.

    1. 1.Step 1: Calculate the mean by summing all values and dividing by the number of values. Sum = 12+15+18+20+22+25+28+30+35+40 = 245. Mean = 245 / 10 = 24.5.
    2. 2.Step 2: Find the median by ordering the data (already ordered) and finding the middle value. For 10 values, the median is the average of the 5th and 6th values: (22 + 25) / 2 = 23.5.
    3. 3.Step 3: Find the interquartile range. The lower quartile (Q1) is the median of the first 5 values: 12, 15, 18, 20, 22 -> Q1 = 18. The upper quartile (Q3) is the median of the last 5 values: 25, 28, 30, 35, 40 -> Q3 = 30. IQR = Q3 - Q1 = 30 - 18 = 12.
    Final Answer: Mean = 24.5 years, Median = 23.5 years, Interquartile range = 12 years.

    Question: A data set has a mean of 50 and a standard deviation of 5. A value of 62 is recorded. Determine whether this value is an outlier. Show your working.

    1. 1.Step 1: Recall that a common rule for outliers is that a value is an outlier if it is more than 2 standard deviations from the mean.
    2. 2.Step 2: Calculate the upper boundary: mean + 2 * standard deviation = 50 + 2*5 = 60. The lower boundary: mean - 2 * standard deviation = 50 - 2*5 = 40.
    3. 3.Step 3: Compare the value 62 to the boundaries. Since 62 > 60, it lies above the upper boundary.
    Final Answer: Yes, the value 62 is an outlier because it is more than 2 standard deviations above the mean.
    Active Recall Memory Test
    What is the formula for standard deviation?
    Key Fact: Standard deviation = sqrt( sum of (x - mean)^2 / (n - 1) ) for a sample, or sqrt( sum of (x - mean)^2 / n ) for a population.
    How do you identify an outlier using the interquartile range?
    Key Fact: A value is an outlier if it is less than Q1 - 1.5 * IQR or greater than Q3 + 1.5 * IQR.
    What two things must you compare when comparing two data sets?
    Key Fact: You must compare a measure of central tendency (mean or median) and a measure of spread (range, interquartile range, or standard deviation).
    What does the interquartile range represent?
    Key Fact: The interquartile range (IQR) is the range of the middle 50% of the data. It is calculated as Q3 - Q1 and is not affected by outliers.
    Frequently Asked Questions
    What is the difference between standard deviation and range?
    Range is the difference between the maximum and minimum values and only considers the extremes. Standard deviation measures how far, on average, each data value is from the mean, taking into account every value. Standard deviation is a more reliable measure of spread because it uses all data points, while range can be misleading if there are outliers.
    How do I know when to use the median instead of the mean?
    Use the median when the data set contains outliers or is skewed, because the median is not affected by extreme values. The mean is affected by outliers and can be pulled away from the typical value. For example, in income data where a few people earn very high salaries, the median is a better representation of typical income.
    What is an outlier and how do I calculate it?
    An outlier is a value that is unusually far from the rest of the data. A common method is to use the interquartile range (IQR): any value below Q1 - 1.5 * IQR or above Q3 + 1.5 * IQR is considered an outlier. Another method is to use standard deviation: values more than 2 standard deviations from the mean are often considered outliers.
    How do I compare two box plots in an exam?
    When comparing two box plots, first compare the medians to comment on typical values. Then compare the interquartile ranges (the length of the box) to comment on consistency or spread. Also mention the overall range if relevant. Always use comparative language such as 'higher median', 'lower IQR', 'more consistent'. Make at least two distinct points to secure full marks.
    What does a high standard deviation tell you about the data?
    A high standard deviation indicates that the data values are spread out over a wider range around the mean. This means there is more variability or inconsistency in the data. Conversely, a low standard deviation means the data values are clustered closely around the mean, indicating more consistency.
    Can the standard deviation be negative?
    No, the standard deviation cannot be negative because it is the square root of a variance, which is always non-negative. The smallest possible value is zero, which occurs when all data values are identical. A standard deviation of zero means there is no spread at all.