Skip to topic
    ← Back to course topics

    E12d — AQA GCSE Statistics

    Test yourself on E12d with AQA GCSE practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    Your focus

    1. Apply Petersen capture/recapture formula to calculate an estimate of the size of a population.

    E12d exam tips

    Quick Revision Summary (Key Takeaway)

    E12d in AQA GCSE Statistics covers the interpretation and comparison of data distributions using measures of central tendency (mean, median, mode) and measures of dispersion (range, interquartile range, standard deviation). Students must calculate these statistics, construct and interpret box plots and cumulative frequency graphs, and use them to compare datasets in context.

    Topic Overview

    This topic, E12d, focuses on the calculation and interpretation of measures of central tendency and dispersion for individual and grouped data. You will learn to compute the mean, median, mode, range, interquartile range, and standard deviation, and understand when each measure is most appropriate. These skills are essential for summarising data effectively and form the foundation for comparing distributions.

    In the wider context of GCSE Statistics, this topic connects to data collection, representation (such as box plots and cumulative frequency graphs), and interpretation. Mastery of these measures allows you to make informed comparisons between datasets, identify outliers, and draw meaningful conclusions from real-world data. This is a core skill tested in both Paper 1 and Paper 2 of the AQA GCSE Statistics exams.

    Key Concepts
    • →Measures of central tendency (mean, median, mode) summarise a dataset with a single typical value, but each is affected differently by outliers and skewness.
    • →Measures of dispersion (range, interquartile range, standard deviation) describe the spread or variability of data, with the IQR being resistant to outliers and standard deviation using all data points.
    • →For grouped data, the mean is estimated using midpoints of class intervals, and the median and quartiles are found using cumulative frequency graphs or interpolation.
    • →Box plots visually display the median, quartiles, and extremes, allowing quick comparison of distributions, while cumulative frequency graphs are used to estimate medians and quartiles for grouped data.
    • →Standard deviation quantifies the average distance of each data point from the mean, with a larger value indicating greater spread.
    Examiner Tips
    • 💡Always show your working clearly, especially for the mean and standard deviation, as method marks are awarded even if the final answer is incorrect.
    • 💡When comparing distributions, use comparative statements that mention both the measure of central tendency and the measure of dispersion, and relate them to the context of the question.
    • 💡For cumulative frequency graphs, ensure you plot cumulative frequency against the upper class boundary, and use a smooth curve or straight lines as instructed. Label axes clearly and use a ruler for reading values.
    Common Mistakes
    • Students often think the mean is always the best measure of average, but the median is more appropriate when data contains outliers or is skewed, as it is not affected by extreme values.
    • When calculating the mean from a frequency table, students sometimes forget to multiply each value by its frequency before summing, or they divide by the number of distinct values instead of the total frequency.
    • Students may confuse the interquartile range with the range, or incorrectly calculate quartiles for small datasets by not using the correct position formula (e.g., for n values, Q1 is at position (n+1)/4).
    Revision Plan
    1. 1Day 1-2: Revise definitions and formulas for mean, median, mode, range, and interquartile range. Practice calculating these for small datasets and from frequency tables.
    2. 2Day 3-4: Learn to estimate the mean from grouped data using midpoints, and practice constructing and interpreting cumulative frequency graphs to find median and quartiles.
    3. 3Day 5-6: Study standard deviation: understand the formula and practice calculating it for small datasets. Compare its use with the range and IQR.
    4. 4Day 7-8: Work through exam-style questions that require comparing two distributions using appropriate statistics. Focus on writing clear comparative conclusions.
    5. 5Day 9-10: Complete a past paper or mock exam under timed conditions, then review your answers and target any weak areas.
    Exam Question Types
    • 📋Calculation questions: These ask you to compute specific statistics (e.g., mean, median, IQR) from a list of data or a frequency table. Advice: Show all steps and double-check calculations, especially when dealing with large numbers.
    • 📋Graph interpretation questions: You may be given a cumulative frequency graph or box plot and asked to estimate the median, quartiles, or compare distributions. Advice: Read values carefully from the graph, using a ruler, and always interpret in context.
    • 📋Comparison questions: These require you to compare two datasets using measures of central tendency and dispersion, and make a conclusion. Advice: Use comparative language and link back to the scenario, ensuring you mention both average and spread.
    • 📋Explain questions: You might be asked to explain why a particular measure is more appropriate, or what effect an outlier has. Advice: Refer to the properties of the measures (e.g., median is not affected by outliers) and apply to the context.
    Command Word Expectations (AQA)
    Calculate

    You must work out a numerical answer using the given data. Show all steps of your working, as method marks are available. The final answer should be clearly stated with units if applicable.

    Compare

    You must describe similarities and differences between two or more datasets, using appropriate statistics. Typically, you need to compare an average (mean or median) and a measure of spread (range or IQR), and make a concluding statement in context.

    Estimate

    You are expected to read values from a graph or use midpoints to approximate a value. Your answer should be reasonable and within an accepted range. Show how you obtained your estimate, e.g., by drawing lines on the graph.

    How Students Lose Marks (Examiner Pitfalls)
    Pitfall: Students often calculate the mean, median, and mode correctly but fail to compare the distributions in the context of the question, losing marks for not linking the statistics to the real-world scenario.
    ❌ Weak Answer (Loses Marks):The mean is 5.2 and the median is 5. The range is 8 and the interquartile range is 3.
    Example improved answer:The mean number of goals scored by Team A is 5.2 compared to 3.8 for Team B, showing Team A generally scores more goals. The interquartile range for Team A is 3, which is lower than Team B's IQR of 5, indicating Team A's scoring is more consistent. Therefore, Team A is both more prolific and more consistent.
    Examiner Tip: Always write a concluding sentence that directly compares the two datasets using the calculated statistics and refers back to the context of the problem. Use comparative language such as 'higher', 'lower', 'more consistent', 'more variable'.
    Pitfall: When constructing a cumulative frequency graph, students frequently plot the cumulative frequency against the upper class boundary instead of the midpoint or lower boundary, leading to an incorrect curve and wrong median/quartile readings.
    ❌ Weak Answer (Loses Marks):Plotting points at (10, 5), (20, 12), (30, 20) when the class intervals are 0-10, 10-20, 20-30.
    Example improved answer:For class intervals 0-10, 10-20, 20-30, plot cumulative frequency at the upper class boundaries: (10, 5), (20, 12), (30, 20). Then join with a smooth curve and read off the median at n/2, lower quartile at n/4, and upper quartile at 3n/4.
    Examiner Tip: Always label the x-axis with 'Upper class boundary' and remember that cumulative frequency is plotted at the upper boundary of each class. Use a sharp pencil and a ruler for straight-line segments if instructed, or a smooth curve for continuous data.
    Step-by-Step Worked Solutions

    Question: The table shows the daily temperatures (in °C) recorded in a city over 20 days: 12, 15, 14, 10, 18, 20, 22, 19, 16, 13, 11, 17, 21, 23, 24, 25, 26, 27, 28, 29. Calculate the mean, median, mode, range, and interquartile range.

    1. 1.Step 1: Order the data: 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29.
    2. 2.Step 2: Calculate the mean: sum all values = 400, divide by 20 = 20 °C.
    3. 3.Step 3: Find the median: with 20 values, the median is the average of the 10th and 11th values: (19 + 20) / 2 = 19.5 °C.
    4. 4.Step 4: Identify the mode: all values appear once, so there is no mode.
    5. 5.Step 5: Calculate the range: maximum - minimum = 29 - 10 = 19 °C.
    6. 6.Step 6: Find the interquartile range: lower quartile (Q1) is the 5.5th value = (14+15)/2 = 14.5 °C; upper quartile (Q3) is the 15.5th value = (24+25)/2 = 24.5 °C; IQR = 24.5 - 14.5 = 10 °C.
    Final Answer: Mean = 20 °C, Median = 19.5 °C, Mode = none, Range = 19 °C, IQR = 10 °C.

    Question: The cumulative frequency graph below shows the times taken (in minutes) by 80 students to complete a puzzle. Use the graph to estimate the median, lower quartile, and upper quartile times, and calculate the interquartile range.

    1. 1.Step 1: Identify the total number of students, n = 80.
    2. 2.Step 2: Locate the median at n/2 = 40 on the cumulative frequency axis, read across to the curve, and down to the time axis. Estimate: 25 minutes.
    3. 3.Step 3: Locate the lower quartile at n/4 = 20 on the cumulative frequency axis, read across to the curve, and down to the time axis. Estimate: 18 minutes.
    4. 4.Step 4: Locate the upper quartile at 3n/4 = 60 on the cumulative frequency axis, read across to the curve, and down to the time axis. Estimate: 32 minutes.
    5. 5.Step 5: Calculate the interquartile range: IQR = Q3 - Q1 = 32 - 18 = 14 minutes.
    Final Answer: Median ≈ 25 minutes, Q1 ≈ 18 minutes, Q3 ≈ 32 minutes, IQR ≈ 14 minutes.
    Active Recall Memory Test
    What is the difference between the range and the interquartile range?
    Key Fact: The range is the difference between the maximum and minimum values, while the interquartile range is the difference between the upper quartile (Q3) and lower quartile (Q1). The IQR is less affected by outliers.
    How do you estimate the mean from a grouped frequency table?
    Key Fact: Multiply the midpoint of each class interval by its frequency, sum these products, and divide by the total frequency.
    When is the median a better measure of central tendency than the mean?
    Key Fact: The median is better when the data contains outliers or is skewed, as it is not affected by extreme values, whereas the mean can be pulled towards the outliers.
    What does a standard deviation of 0 indicate?
    Key Fact: A standard deviation of 0 means all data values are identical; there is no spread.
    Frequently Asked Questions
    What is the difference between the mean, median, and mode?
    The mean is the sum of all values divided by the number of values. The median is the middle value when data is ordered. The mode is the most frequently occurring value. The mean uses all data and is affected by outliers, the median is resistant to outliers, and the mode is useful for categorical data.
    How do I calculate the interquartile range?
    First order the data. Find the lower quartile (Q1), which is the median of the lower half of the data, and the upper quartile (Q3), the median of the upper half. The interquartile range is Q3 minus Q1. For large datasets, use the position formulas: Q1 at (n+1)/4 and Q3 at 3(n+1)/4.
    Why is standard deviation important?
    Standard deviation measures the average spread of data around the mean. A low standard deviation indicates data points are close to the mean, while a high standard deviation indicates they are spread out. It is useful for comparing consistency between datasets.
    How do I compare two box plots?
    Compare the medians to see which dataset generally has higher values. Compare the interquartile ranges to see which dataset has more consistent middle 50% values. Also compare the overall ranges and mention any outliers. Always relate your comparison back to the context of the question.
    What is the formula for standard deviation in GCSE Statistics?
    The formula is: standard deviation = sqrt( sum of (x - mean)^2 / n ), where x represents each data value, mean is the arithmetic mean, and n is the number of data values. You may also use the alternative formula: sqrt( sum of x^2 / n - mean^2 ).
    Can the mean be a decimal even if the data are whole numbers?
    Yes, the mean can be a decimal even if all data values are whole numbers, because it is the sum divided by the count. For example, the mean of 2, 3, and 4 is 3, but the mean of 2, 3, and 5 is 10/3 ≈ 3.33.