Skip to topic
    ← Back to course topics

    E2a — AQA GCSE Statistics

    Test yourself on E2a with AQA GCSE practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    Your focus

    1. Compare experimental data with theoretical predictions to identify possible bias within the experimental design.

    E2a exam tips

    Quick Revision Summary (Key Takeaway)

    E2a in AQA GCSE Statistics covers the interpretation and comparison of data distributions using measures of central tendency (mean, median, mode) and measures of dispersion (range, interquartile range, standard deviation). Students must calculate these statistics, construct and interpret box plots and cumulative frequency diagrams, and use them to compare data sets in context.

    Topic Overview

    E2a is a core topic in AQA GCSE Statistics that focuses on summarising and comparing data sets using numerical measures. You will learn to calculate and interpret measures of central tendency (mean, median, mode) and measures of dispersion (range, interquartile range, standard deviation). These tools allow you to describe the typical value and the spread of data, which is essential for making informed comparisons and decisions.

    This topic is fundamental because it underpins much of statistical analysis. It connects to other areas such as data presentation (box plots, cumulative frequency diagrams) and probability. Understanding E2a helps you to critically evaluate data in real-world contexts, from comparing test scores to analysing scientific experiments, and is heavily examined in both foundation and higher tier papers.

    Key Concepts
    • →Measures of central tendency: mean (average), median (middle value), and mode (most frequent) summarise the typical value in a data set.
    • →Measures of dispersion: range (max - min), interquartile range (UQ - LQ), and standard deviation quantify the spread or variability of data.
    • →The interquartile range is often preferred over the range because it ignores outliers and focuses on the middle 50% of data.
    • →Standard deviation measures the average distance of each data point from the mean; a smaller standard deviation indicates greater consistency.
    • →Box plots visually display the median, quartiles, and extremes, allowing quick comparison of distributions.
    Examiner Tips
    • 💡Always show your working for calculations, especially for standard deviation, as method marks are available even if the final answer is wrong.
    • 💡When asked to compare data sets, use comparative language (e.g., 'higher than', 'more consistent') and refer back to the context of the question.
    • 💡For box plot comparisons, comment on median, interquartile range, and overall range to ensure you cover all aspects of the distribution.
    Common Mistakes
    • Students often think the mean is always the best measure of average. However, the median is better when data is skewed or has outliers, as it is not affected by extreme values.
    • Students sometimes confuse standard deviation with range. Standard deviation considers every data point's distance from the mean, while range only uses the maximum and minimum.
    • When comparing box plots, students may only compare medians and forget to compare spreads (IQR or range), missing key differences in consistency.
    Revision Plan
    1. 1Day 1-2: Revise definitions and calculations for mean, median, mode, range, and interquartile range. Practice with small data sets.
    2. 2Day 3-4: Learn to calculate standard deviation step by step. Use the formula and practice with at least 5 different data sets.
    3. 3Day 5-6: Study box plots and cumulative frequency diagrams. Practice interpreting and drawing them, and comparing two distributions.
    4. 4Day 7-8: Attempt past paper questions on E2a, focusing on comparison questions. Mark your work using the mark scheme and note common errors.
    5. 5Day 9-10: Review misconceptions and examiner tips. Create a summary sheet of key formulas and comparison phrases.
    Exam Question Types
    • 📋Calculation questions: Calculate the mean, median, mode, range, interquartile range, or standard deviation from a list or frequency table. Advice: Show all steps, especially for standard deviation.
    • 📋Comparison questions: Compare two data sets using box plots or summary statistics. Advice: Use comparative language and comment on both average and spread.
    • 📋Interpretation questions: Explain what a statistic means in context or which measure is most appropriate. Advice: Relate your answer to the specific scenario and justify your choice.
    • 📋Graph questions: Draw or interpret a box plot or cumulative frequency diagram. Advice: Label axes clearly and use a ruler for accuracy.
    Command Word Expectations (AQA)
    Calculate

    You must work out a numerical value using the given data. Show all steps of your working, as method marks are awarded. Give your final answer with appropriate units or rounding.

    Compare

    You must describe similarities and differences between two or more data sets. Use comparative language (e.g., 'higher', 'lower', 'more consistent') and refer to both measures of central tendency and dispersion.

    Explain

    You must give reasons or justify your answer. This often involves stating which measure is most appropriate and why, using the context of the data (e.g., presence of outliers, skewness).

    How Students Lose Marks (Examiner Pitfalls)
    Pitfall: Students often calculate the mean, median, and mode correctly but fail to interpret them in the context of the question, losing marks for not comparing the distributions or relating back to the real-world scenario.
    ❌ Weak Answer (Loses Marks):The mean is 12.5 and the median is 12. The range is 8.
    Example improved answer:The mean number of goals scored is 12.5, which is higher than the median of 12, suggesting the distribution is positively skewed. The range of 8 indicates the spread of goals is relatively small, so the team's performance is consistent.
    Examiner Tip: Always write a sentence that compares the two data sets or explains what the statistic means in the context of the problem. Use comparative language such as 'higher than', 'more consistent', or 'on average'.
    Pitfall: When comparing box plots, students often describe the median and interquartile range but forget to mention the overall range or the skewness, missing out on marks for a full comparison.
    ❌ Weak Answer (Loses Marks):The median for group A is higher than group B. The interquartile range for group A is smaller.
    Example improved answer:Group A has a higher median (15 compared to 12), indicating that on average, group A scored higher. Group A also has a smaller interquartile range (8 compared to 12), showing that the middle 50% of scores are more consistent. However, group B has a larger overall range (20 compared to 15), suggesting greater variability in the extremes.
    Examiner Tip: When comparing box plots, always comment on median (average), interquartile range (consistency of middle 50%), and overall range (spread of entire data). Use comparative statements and relate to the context.
    Step-by-Step Worked Solutions

    Question: The table shows the number of cars sold by two salespeople over 10 days. Calculate the mean and standard deviation for each and compare their performance. Salesperson A: 5, 7, 8, 6, 9, 10, 4, 7, 8, 6. Salesperson B: 2, 12, 3, 11, 4, 10, 5, 9, 6, 8.

    1. 1.Step 1: Calculate the mean for Salesperson A: sum = 5+7+8+6+9+10+4+7+8+6 = 70, mean = 70/10 = 7.
    2. 2.Step 2: Calculate the standard deviation for Salesperson A: first find squared deviations from mean, sum them, divide by n, then square root. Deviations: (5-7)^2=4, (7-7)^2=0, (8-7)^2=1, (6-7)^2=1, (9-7)^2=4, (10-7)^2=9, (4-7)^2=9, (7-7)^2=0, (8-7)^2=1, (6-7)^2=1. Sum = 30. Variance = 30/10 = 3. Standard deviation = sqrt(3) ≈ 1.73.
    3. 3.Step 3: Calculate the mean for Salesperson B: sum = 2+12+3+11+4+10+5+9+6+8 = 70, mean = 70/10 = 7.
    4. 4.Step 4: Calculate the standard deviation for Salesperson B: deviations: (2-7)^2=25, (12-7)^2=25, (3-7)^2=16, (11-7)^2=16, (4-7)^2=9, (10-7)^2=9, (5-7)^2=4, (9-7)^2=4, (6-7)^2=1, (8-7)^2=1. Sum = 110. Variance = 110/10 = 11. Standard deviation = sqrt(11) ≈ 3.32.
    5. 5.Step 5: Compare: Both have the same mean (7), but Salesperson A has a much smaller standard deviation (1.73 vs 3.32), indicating more consistent sales.
    Final Answer: Salesperson A: mean = 7, standard deviation ≈ 1.73. Salesperson B: mean = 7, standard deviation ≈ 3.32. Salesperson A is more consistent as their standard deviation is lower.

    Question: The cumulative frequency graph shows the heights of 80 plants. Estimate the median, lower quartile, and upper quartile heights, and draw a box plot to represent the data.

    1. 1.Step 1: Find the median: cumulative frequency = 80, so median is at 40 on the cumulative frequency axis. Read across to the curve and down to the height axis to estimate the median height (e.g., 25 cm).
    2. 2.Step 2: Find the lower quartile (LQ): at 20 on the cumulative frequency axis (80/4 = 20). Read across to the curve and down to estimate LQ (e.g., 18 cm).
    3. 3.Step 3: Find the upper quartile (UQ): at 60 on the cumulative frequency axis (3*80/4 = 60). Read across to the curve and down to estimate UQ (e.g., 32 cm).
    4. 4.Step 4: Draw a box plot: draw a box from LQ to UQ with a line at the median. Whiskers extend to the minimum and maximum values (from the graph, e.g., 10 cm and 45 cm).
    Final Answer: Median ≈ 25 cm, LQ ≈ 18 cm, UQ ≈ 32 cm. Box plot drawn with these values and whiskers to min and max.
    Active Recall Memory Test
    What is the difference between the range and the interquartile range?
    Key Fact: The range is the difference between the maximum and minimum values, while the interquartile range is the difference between the upper quartile and lower quartile, covering the middle 50% of data and ignoring outliers.
    How do you calculate standard deviation?
    Key Fact: First, find the mean. Then subtract the mean from each data point, square the result, and find the average of these squared differences (variance). Finally, take the square root of the variance.
    What does a small standard deviation indicate about a data set?
    Key Fact: A small standard deviation indicates that the data points are close to the mean, meaning the data is consistent and less spread out.
    When comparing two box plots, what three features should you compare?
    Key Fact: Compare the medians (average), the interquartile ranges (consistency of middle 50%), and the overall ranges (spread of all data).
    Frequently Asked Questions
    What is the difference between standard deviation and range?
    Standard deviation measures the average distance of each data point from the mean, using all data values. Range is the difference between the highest and lowest values and only considers two data points. Standard deviation is a more reliable measure of spread because it takes into account every value, while range can be heavily influenced by outliers.
    How do I know whether to use the mean or the median?
    Use the mean when the data is roughly symmetrical and there are no extreme outliers, as it uses all data values. Use the median when the data is skewed or has outliers, because the median is not affected by extreme values. In exam questions, if the data includes very high or low values, the median is often more appropriate.
    What does the interquartile range tell you about a data set?
    The interquartile range (IQR) tells you the spread of the middle 50% of the data. It is calculated as the upper quartile minus the lower quartile. A small IQR indicates that the middle half of the data is clustered closely together, suggesting consistency, while a large IQR indicates greater variability in the central portion of the data.
    How do I compare two sets of data in an exam?
    To compare two data sets, comment on both an average (usually the median or mean) and a measure of spread (usually the interquartile range or standard deviation). Use comparative language such as 'higher than', 'lower than', 'more consistent', or 'more spread out'. Always relate your comparison back to the context of the question, for example, 'Group A is more consistent because their interquartile range is smaller.'
    What is a box plot and what does it show?
    A box plot is a diagram that shows the median, lower quartile, upper quartile, and the minimum and maximum values of a data set. The box represents the interquartile range (middle 50% of data), the line inside the box is the median, and the whiskers extend to the minimum and maximum values. Box plots are useful for comparing distributions side by side.
    Why is standard deviation considered a better measure of spread than the range?
    Standard deviation is considered better because it uses every data value to calculate the average distance from the mean, giving a more representative measure of spread. The range only uses the two most extreme values, so it can be misleading if there are outliers. Standard deviation is less affected by outliers and provides a more stable measure of variability.