Skip to topic
    ← Back to course topics

    A2 — AQA GCSE Statistics

    Test yourself on A2 with AQA GCSE practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    Your focus

    1. Know the constraints that may be faced in designing an investigation to test a hypothesis: these may include factors such as time, costs, ethical issues, confidentiality and convenience etc.

    A2 exam tips

    Quick Revision Summary (Key Takeaway)

    A2 Statistics for AQA GCSE covers advanced data analysis techniques including probability distributions, hypothesis testing, correlation and regression, and the interpretation of statistical diagrams. It builds on foundational skills to develop critical evaluation of data and the ability to communicate statistical findings effectively.

    Topic Overview

    A2 Statistics for AQA GCSE extends the foundational concepts of data handling and probability to more advanced techniques such as hypothesis testing, correlation and regression, and probability distributions. It emphasises the interpretation and critical analysis of data, enabling students to draw meaningful conclusions and communicate findings effectively. This topic is essential for understanding how statistics is applied in real-world contexts, from scientific research to business decision-making.

    The unit builds on prior knowledge of averages, spread, and basic probability, and it prepares students for further study in statistics or related fields. Mastery of A2 Statistics involves not only computational skills but also the ability to evaluate statistical claims, recognise limitations, and use appropriate terminology. It is a core component of the GCSE Mathematics curriculum and is assessed through both calculation and written reasoning questions.

    Key Concepts
    • →Hypothesis testing: formulating null and alternative hypotheses, choosing significance levels, and interpreting p-values or critical regions.
    • →Probability distributions: understanding binomial and normal distributions, and calculating probabilities using them.
    • →Correlation and regression: calculating and interpreting the product moment correlation coefficient (PMCC) and the equation of the regression line.
    • →Data interpretation: analysing scatter diagrams, box plots, histograms, and cumulative frequency graphs to describe distributions and identify trends.
    • →Statistical reasoning: distinguishing between correlation and causation, understanding sampling methods, and evaluating the reliability of conclusions.
    Examiner Tips
    • 💡Always show your working clearly, especially in hypothesis testing: state hypotheses, calculate test statistic, compare with critical value, and write a conclusion in context.
    • 💡When interpreting statistical diagrams, comment on shape, centre, spread, and outliers, and use comparative language if comparing two sets of data.
    • 💡Use precise statistical terminology, such as 'significant', 'correlation', 'probability', and 'distribution', to demonstrate understanding and secure marks.
    Common Mistakes
    • Students often think that a high correlation coefficient proves causation. In fact, correlation only indicates a relationship; causation requires further evidence and control of confounding variables.
    • Many students believe that a non-significant result in a hypothesis test proves the null hypothesis is true. Actually, it only means there is insufficient evidence to reject it; the null could still be false.
    • When using the normal distribution, students sometimes forget to standardise values or apply continuity corrections when approximating a binomial distribution.
    Revision Plan
    1. 1Week 1: Review foundational concepts such as probability rules, averages, and spread. Practice calculating probabilities and interpreting basic diagrams.
    2. 2Week 1-2: Focus on hypothesis testing: learn the steps, practice with binomial and normal distributions, and complete past paper questions on this topic.
    3. 3Week 2: Study correlation and regression: calculate PMCC and regression lines, interpret scatter diagrams, and understand the difference between correlation and causation.
    4. 4Week 2: Revise probability distributions: binomial and normal, including calculations and approximations. Use flashcards for key formulas.
    5. 5Week 2: Complete a full past paper under timed conditions, then review mistakes and revisit weak areas.
    Exam Question Types
    • 📋Hypothesis testing question: often involves a scenario (e.g., coin bias, medical trial) where you must set up hypotheses, calculate a test statistic, and conclude in context. Advice: always define the parameter and significance level clearly.
    • 📋Correlation and regression question: given bivariate data, calculate PMCC, interpret its value, and find the regression line equation. Advice: remember that PMCC is unaffected by units and always lies between -1 and 1.
    • 📋Probability distribution question: may ask you to calculate probabilities using binomial or normal distribution, or to approximate binomial with normal. Advice: check conditions for approximation and apply continuity correction if needed.
    • 📋Data interpretation question: analyse a complex diagram (e.g., cumulative frequency graph) to estimate median, quartiles, and percentiles. Advice: read scales carefully and show construction lines.
    Command Word Expectations (AQA)
    Calculate

    You must show all steps of your working and give the final answer with appropriate units or rounding. Marks are awarded for correct method and accuracy.

    Interpret

    You must explain the meaning of a statistical result in the context of the problem, using correct terminology. For example, 'The PMCC of 0.85 indicates a strong positive correlation between variables X and Y.'

    Evaluate

    You must make a judgement based on the evidence, considering both strengths and limitations. For example, 'The conclusion is valid because the sample was random, but the small sample size may limit generalisability.'

    How Students Lose Marks (Examiner Pitfalls)
    Pitfall: Students often confuse correlation with causation when interpreting scatter diagrams, leading to incorrect conclusions about the relationship between variables.
    ❌ Weak Answer (Loses Marks):The scatter graph shows that as temperature increases, ice cream sales increase, so temperature causes more ice cream to be sold.
    Example improved answer:The scatter graph shows a strong positive correlation between temperature and ice cream sales. However, this does not necessarily imply causation; there may be a third variable, such as time of year, that influences both. Further investigation would be needed to establish a causal link.
    Examiner Tip: Always use the phrase 'correlation does not imply causation' and suggest possible confounding variables when interpreting relationships.
    Pitfall: When conducting a hypothesis test, students frequently fail to state the null and alternative hypotheses correctly or misinterpret the significance level.
    ❌ Weak Answer (Loses Marks):The null hypothesis is that the coin is fair. The alternative is that it is not fair. The significance level is 0.05, so if the probability is less than 0.05 we accept the null hypothesis.
    Example improved answer:The null hypothesis (H0) is that the coin is fair, so P(heads) = 0.5. The alternative hypothesis (H1) is that the coin is biased, so P(heads) ≠ 0.5. Using a 5% significance level, we calculate the probability of the observed result or more extreme, assuming H0 is true. If this probability is less than 0.05, we reject H0 in favour of H1; otherwise, we do not reject H0.
    Examiner Tip: Clearly define H0 and H1 in words and symbols, and remember that a small p-value leads to rejection of H0, not acceptance of H1.
    Step-by-Step Worked Solutions

    Question: A bag contains 5 red balls and 3 blue balls. Two balls are drawn at random without replacement. Calculate the probability that both balls are red.

    1. 1.Step 1: Identify the total number of balls initially: 5 + 3 = 8.
    2. 2.Step 2: Calculate the probability that the first ball is red: 5/8.
    3. 3.Step 3: After drawing one red ball, there are 4 red balls left and 7 balls in total. So the probability that the second ball is red is 4/7.
    4. 4.Step 4: Multiply the probabilities: (5/8) * (4/7) = 20/56 = 5/14.
    Final Answer: The probability that both balls are red is 5/14 or approximately 0.357.

    Question: A researcher claims that the mean height of adult males in a certain population is 175 cm. A random sample of 36 adult males has a mean height of 172 cm with a standard deviation of 8 cm. Test at the 5% significance level whether the mean height is different from 175 cm.

    1. 1.Step 1: State the null hypothesis H0: μ = 175 cm and the alternative hypothesis H1: μ ≠ 175 cm.
    2. 2.Step 2: Calculate the standard error: σ/√n = 8/√36 = 8/6 = 1.333 cm.
    3. 3.Step 3: Calculate the test statistic (z-score): z = (sample mean - population mean) / standard error = (172 - 175) / 1.333 = -3 / 1.333 ≈ -2.25.
    4. 4.Step 4: For a two-tailed test at 5% significance, the critical z-values are approximately ±1.96. Since -2.25 < -1.96, the test statistic falls in the critical region.
    5. 5.Step 5: Reject H0. There is sufficient evidence to suggest that the mean height is different from 175 cm.
    Final Answer: At the 5% significance level, there is sufficient evidence to reject the null hypothesis and conclude that the mean height of adult males is not 175 cm.
    Active Recall Memory Test
    What is the difference between the null and alternative hypotheses?
    Key Fact: The null hypothesis (H0) is a statement of no effect or no difference, assumed true until evidence suggests otherwise. The alternative hypothesis (H1) is the claim we are testing for, which contradicts H0.
    What does a correlation coefficient of -0.9 indicate?
    Key Fact: It indicates a strong negative correlation: as one variable increases, the other tends to decrease.
    When should you use a normal distribution to approximate a binomial distribution?
    Key Fact: When n is large and p is close to 0.5, typically when np > 5 and n(1-p) > 5, and you apply a continuity correction.
    What is the formula for the standard deviation of a sample?
    Key Fact: s = sqrt( Σ(x - x̄)² / (n - 1) ), where x̄ is the sample mean and n is the sample size.
    Frequently Asked Questions
    What is the difference between correlation and causation?
    Correlation means two variables are related, but causation means one variable directly affects the other. A correlation can be due to a third variable or coincidence. For example, ice cream sales and drowning incidents are correlated because both increase in summer, but ice cream does not cause drowning.
    How do I know whether to use a one-tailed or two-tailed test?
    Use a one-tailed test when the alternative hypothesis specifies a direction (e.g., 'greater than'). Use a two-tailed test when it does not specify a direction (e.g., 'not equal to'). The choice should be based on the research question before seeing the data.
    What does a p-value represent?
    The p-value is the probability of obtaining a result at least as extreme as the observed one, assuming the null hypothesis is true. A small p-value (typically < 0.05) suggests that the observed result is unlikely under H0, leading to rejection of H0.
    How do I calculate the interquartile range from a cumulative frequency graph?
    Find the median (50th percentile), lower quartile (25th percentile), and upper quartile (75th percentile) by reading across from the corresponding cumulative frequencies to the curve, then down to the variable axis. The interquartile range is the difference between the upper and lower quartiles.
    Why is standard deviation often preferred over range?
    Standard deviation uses all data values and is less affected by extreme outliers than the range. It provides a measure of spread around the mean, making it more reliable for comparing variability between datasets.
    What is a Type I error in hypothesis testing?
    A Type I error occurs when you reject a true null hypothesis. Its probability is equal to the significance level (α). For example, with α = 0.05, there is a 5% chance of incorrectly rejecting H0.