Skip to topic
    ← Back to course topics

    E5b — AQA GCSE Statistics

    Test yourself on E5b with AQA GCSE practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    Your focus

    1. Comment on outliers with reference to the original data.

    E5b exam tips

    Quick Revision Summary (Key Takeaway)

    Topic E5b covers the normal distribution in AQA GCSE Statistics, including its characteristic bell-shaped curve, symmetry about the mean, and the empirical rule. Students learn to calculate standardised scores (z-scores) and estimate proportions within one, two, and three standard deviations of the mean.

    Topic Overview

    Topic E5b focuses on the normal distribution, the most critical continuous probability model in statistics. Students explore the geometric and statistical properties of bell-shaped curves, understanding that continuous data often clusters symmetrically around a central mean.

    Mastering this topic enables students to apply the 68-95-99.7 empirical rule to estimate population proportions and to use standardised z-scores to compare values from datasets with different scales and units across real-world contexts.

    Key Concepts
    • →The normal distribution is continuous, unimodal, and perfectly symmetrical, with mean = median = mode at the line of symmetry.
    • →The total area beneath the normal distribution probability density curve is equal to 1 (or 100%).
    • →The empirical rule dictates that approximately 68% of data lies within mu +/- 1sigma, 95% within mu +/- 2sigma, and 99.7% within mu +/- 3sigma.
    • →A standardised score (z-score), calculated as z = (x - mu) / sigma, measures how many standard deviations a raw score x is above or below the population mean.
    Examiner Tips
    • 💡Always draw and annotate a normal distribution curve in your working, clearly marking mu, standard deviation intervals, and shading the target area.
    • 💡When comparing performances using z-scores, always state both calculated z-values clearly before writing your comparative conclusion in context.
    Common Mistakes
    • Assuming normal distribution properties apply to skewed data sets; the empirical rule and symmetrical properties only apply to bell-shaped, symmetrical data.
    • Forgetting to divide tail areas by 2; for example, knowing 5% lies outside mu +/- 2sigma but assigning the full 5% to just one upper tail instead of 2.5% per tail.
    • Treating the normal distribution as discrete rather than continuous, leading to confusion over inequalities such as P(X < a) versus P(X <= a).
    Revision Plan
    1. 1Day 1: Learn the key properties of the normal distribution curve and memorize the 68-95-99.7% rule values.
    2. 2Day 2: Practice sketching normal curves and calculating symmetric and asymmetric interval percentages using the empirical rule.
    3. 3Day 3: Master the standardised score formula z = (x - mu) / sigma and practice comparing different test scores or physical measurements.
    4. 4Day 4: Attempt past paper multi-mark questions combining normal distribution percentages with expected frequencies in real-world samples.
    Exam Question Types
    • 📋Empirical rule calculations: Questions requiring students to find percentages or counts of a population falling between specific boundary values.
    • 📋Standardised score comparisons: Questions presenting raw scores from two different subjects or groups, requiring z-score calculations to evaluate relative performance.
    • 📋Identifying distribution characteristics: Questions asking whether a given histogram or frequency polygon can be modeled adequately by a normal distribution.
    Command Word Expectations (AQA)
    Calculate

    Show clear mathematical working, including formula substitution (e.g. z = (x - mu) / sigma), arriving at an exact numerical answer with appropriate rounding.

    Compare

    Calculate both statistical measures (such as standardised scores), explicitly state which is higher or lower, and interpret the difference in context.

    Estimate

    Use the 68-95-99.7% approximation rule to find proportions or expected counts rather than trying to use advanced statistical tables.

    How Students Lose Marks (Examiner Pitfalls)
    Pitfall: Confusing the percentages associated with 1, 2, and 3 standard deviations, or misapplying them symmetrically.
    ❌ Weak Answer (Loses Marks):About 95% of data is within 1 standard deviation, so 5% is in the tails.
    Example improved answer:For a normal distribution, approximately 68% of the data lies within mu +/- 1sigma, 95% lies within mu +/- 2sigma, and 99.7% lies within mu +/- 3sigma. Therefore, approximately 2.5% lies in each tail beyond 2 standard deviations from the mean.
    Examiner Tip: Always sketch a bell curve and shade the required region before calculating percentages. Label mu, mu +/- sigma, and mu +/- 2sigma clearly on the horizontal axis.
    Pitfall: Calculating a standardised score (z-score) with the numerator reversed as (mu - x) instead of (x - mu), producing an incorrect sign.
    ❌ Weak Answer (Loses Marks):z = (70 - 78) / 4 = -2, so the student scored 2 standard deviations below the mean.
    Example improved answer:z = (x - mu) / sigma = (78 - 70) / 4 = +2.0. The positive sign indicates the score is 2 standard deviations above the population mean.
    Examiner Tip: Check that your z-score sign matches reality: values above the mean must yield positive z-scores, and values below the mean must yield negative z-scores.
    Step-by-Step Worked Solutions

    Question: The masses of adult hedgehogs in a nature reserve are normally distributed with a mean of 800 g and a standard deviation of 60 g. (a) Estimate the percentage of hedgehogs weighing between 740 g and 920 g. (b) A sample of 400 hedgehogs is examined. Estimate how many hedgehogs weigh more than 920 g.

    1. 1.Step 1: Calculate the boundary distances from the mean in terms of standard deviations. Lower boundary: 740 g = 800 - 60 = mu - 1sigma. Upper boundary: 920 g = 800 + (2 * 60) = mu + 2sigma.
    2. 2.Step 2: Use the empirical rule to find the probabilities between the mean and each boundary. Between mu - 1sigma and mu lies 68% / 2 = 34%. Between mu and mu + 2sigma lies 95% / 2 = 47.5%.
    3. 3.Step 3: Sum the two regions for part (a): 34% + 47.5% = 81.5%.
    4. 4.Step 4: For part (b), determine the percentage above mu + 2sigma. The upper tail beyond mu + 2sigma contains (100% - 95%) / 2 = 2.5%.
    5. 5.Step 5: Multiply this proportion by the total sample size: 2.5% of 400 = 0.025 * 400 = 10 hedgehogs.
    Final Answer: (a) 81.5% of hedgehogs weigh between 740 g and 920 g. (b) Approximately 10 hedgehogs are expected to weigh more than 920 g.

    Question: Maya scores 68 in a Biology exam where the mean is 60 and standard deviation is 5. In Chemistry, she scores 72 where the mean is 64 and standard deviation is 6. By calculating standardised scores, determine in which subject Maya performed relatively better.

    1. 1.Step 1: State the formula for the standardised score: z = (x - mu) / sigma.
    2. 2.Step 2: Calculate the standardised score for Biology: z_Biology = (68 - 60) / 5 = 8 / 5 = +1.6.
    3. 3.Step 3: Calculate the standardised score for Chemistry: z_Chemistry = (72 - 64) / 6 = 8 / 6 = +1.33 (to 2 d.p.).
    4. 4.Step 4: Compare the two standardised scores: 1.6 > 1.33.
    5. 5.Step 5: Interpret the comparison in context: Maya performed relatively better in Biology because her score is 1.6 standard deviations above the cohort mean, compared to only 1.33 standard deviations above the mean in Chemistry.
    Final Answer: Maya performed relatively better in Biology (z = +1.6) than in Chemistry (z = +1.33).
    Active Recall Memory Test
    What percentage of data lies within 1, 2, and 3 standard deviations of the mean in a normal distribution?
    Key Fact: Approximately 68% within 1 standard deviation, 95% within 2 standard deviations, and 99.7% within 3 standard deviations.
    What is the formula for calculating a standardised score (z-score)?
    Key Fact: z = (x - mu) / sigma, where x is the value, mu is the mean, and sigma is the standard deviation.
    What are the three measures of average that are identical in a perfect normal distribution?
    Key Fact: Mean, median, and mode are all equal and located at the vertical line of symmetry.
    What percentage of data in a normal distribution is expected to lie above mu + 1sigma?
    Key Fact: 16% (calculated as (100% - 68%) / 2 = 16%).
    Frequently Asked Questions
    What is a standardised score and why is it useful?
    A standardised score, or z-score, measures how many standard deviations an observation falls above or below the mean. It transforms data from different distributions onto a common scale with a mean of 0 and a standard deviation of 1. This allows fair comparisons between different variables, such as test results from exams with different difficulties and grade boundaries.
    Do I need to use statistical tables or calculus for normal distribution in GCSE Statistics?
    No, AQA GCSE Statistics does not require calculus or statistical lookup tables for the normal distribution. Exam questions rely strictly on the 68-95-99.7% empirical rule, symmetry properties, and standardised score calculations (z = (x - mu) / sigma). You only need to know boundaries that correspond to 1, 2, or 3 standard deviations from the mean.
    What does a negative z-score mean?
    A negative z-score simply indicates that the raw score is lower than the arithmetic mean of the dataset. For instance, a z-score of -1.5 means the value is 1.5 standard deviations below the mean. It does not mean the original data value itself was negative.
    How can I tell from a diagram if data is normally distributed?
    A normal distribution graph must be a continuous, single-peaked (unimodal) bell curve that is symmetrical about the centre. If the curve is skewed to the left or right, has multiple peaks (bimodal), or has a flat top, the data is not normally distributed and the empirical rule cannot be applied.
    Why is the normal distribution so important in statistics?
    The normal distribution naturally models many biological, physical, and social variables, such as human heights, blood pressure, and measurement errors. Additionally, under the central limit theorem, sample means of sufficiently large samples tend to follow a normal distribution, making it foundational for statistical inference.