Skip to topic
    ← Back to course topics

    D7 — AQA GCSE Statistics

    Test yourself on D7 with AQA GCSE practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    Your focus

    1. Use collected data to calculate estimates of probabilities.

    D7 exam tips

    Quick Revision Summary (Key Takeaway)

    D7 in AQA GCSE Statistics is the controlled assessment (formerly coursework), worth 20% of the final grade, where students plan, collect, process and analyse data on a self-chosen hypothesis, then evaluate their findings. It assesses the full statistical enquiry cycle and is marked out of 40 across four strands: planning, data collection, processing and representation, and interpretation and evaluation.

    Topic Overview

    D7 is the controlled assessment component of AQA GCSE Statistics, accounting for 20% of the total marks. It requires students to independently plan and conduct a statistical investigation based on a hypothesis of their choice, following the full enquiry cycle: planning, data collection, processing and representation, and interpretation and evaluation. The assessment is marked out of 40 and is internally assessed then externally moderated.

    This topic is crucial because it tests your ability to apply statistical knowledge to a real-world context, rather than just answering abstract questions. It develops skills in designing investigations, handling data critically, and communicating findings effectively. Success in D7 demonstrates a deep understanding of the statistical problem-solving process, which is essential for further study and everyday data literacy.

    Key Concepts
    • →The statistical enquiry cycle: a continuous four-stage process of planning, data collection, processing and representation, and interpretation and evaluation.
    • →Hypothesis: a testable statement predicting a relationship or difference between two variables, which must be measurable and specific.
    • →Sampling methods: random, systematic, stratified, and quota sampling, each with advantages and disadvantages regarding bias and representativeness.
    • →Data types: primary vs secondary data, quantitative (discrete and continuous) vs qualitative, and the importance of consistency in collection.
    • →Statistical measures: measures of central tendency (mean, median, mode), measures of dispersion (range, interquartile range, standard deviation), and correlation (Pearson's r, Spearman's rank).
    Examiner Tips
    • 💡Always justify your choice of sampling method and sample size in the planning section. Explain why it is appropriate for your hypothesis and population.
    • 💡In the processing section, use appropriate graphs and calculations. For example, use a scatter graph and Pearson's r for correlation, or a box plot and standard deviation for comparing distributions.
    • 💡In the evaluation, be critical: discuss limitations of your data collection, possible sources of bias, and suggest specific improvements. Avoid generic statements like 'I could have collected more data' without explaining how that would improve reliability.
    Common Mistakes
    • Students often think that any correlation proves cause and effect. Correction: correlation does not imply causation; a third variable may explain the relationship.
    • Many believe that a larger sample size automatically removes bias. Correction: a large sample can still be biased if the sampling method is flawed (e.g. convenience sampling).
    • Students sometimes confuse the mean and median, or use the mean for skewed data. Correction: the median is more appropriate for skewed data or when outliers are present.
    Revision Plan
    1. 1Week 1: Review the statistical enquiry cycle and sampling methods. Choose a hypothesis and write a detailed plan, including variables, sampling method, and data collection strategy.
    2. 2Week 1-2: Collect your data, ensuring consistency and ethical considerations. Record data accurately in a table.
    3. 3Week 2: Process your data using appropriate calculations (e.g. means, standard deviation, correlation coefficient) and draw relevant graphs.
    4. 4Week 2: Interpret your findings, relate them back to your hypothesis, and evaluate the strengths and weaknesses of your investigation.
    5. 5Ongoing: Practice writing model answers for each strand using past D7 exemplars and mark schemes to understand the assessment criteria.
    Exam Question Types
    • 📋Planning question: 'Describe how you would collect data to test the hypothesis that...' Advice: Outline sampling method, sample size, variables, and how to ensure reliability.
    • 📋Processing question: 'Calculate the correlation coefficient and interpret it in context.' Advice: Show all steps, use correct formula, and relate the value to the hypothesis.
    • 📋Evaluation question: 'Discuss the limitations of your investigation and suggest improvements.' Advice: Be specific about sources of bias and how changes would improve validity.
    Command Word Expectations (AQA)
    Describe

    Give a detailed account of the steps or features. In D7, this often requires outlining your sampling method, data collection process, or how you would represent data. Marks are awarded for specific, relevant detail, not vague statements.

    Calculate

    Show all working and give your answer to an appropriate degree of accuracy (usually 2 or 3 significant figures). In D7, this may involve computing summary statistics or correlation coefficients. Marks are awarded for correct substitution and accurate arithmetic.

    Evaluate

    Make a judgement about the strengths and weaknesses of your investigation, supported by evidence. You must discuss limitations, possible improvements, and how these affect the validity and reliability of your conclusions. Marks are awarded for critical analysis, not just listing points.

    How Students Lose Marks (Examiner Pitfalls)
    Pitfall: Students often choose a hypothesis that is too vague or not statistically testable, such as 'Boys are better at maths than girls', which cannot be measured directly and leads to weak planning marks.
    ❌ Weak Answer (Loses Marks):My hypothesis is that boys are better at maths than girls. I will ask 20 people and make a bar chart.
    Example improved answer:Hypothesis: 'Students who spend more time on social media per day achieve lower scores in their GCSE Maths mock exam.' This is testable because both variables are measurable: time on social media (continuous, in hours) and mock exam score (discrete, out of 80). I will collect primary data from 50 Year 11 students using a stratified sample by gender, ensuring a representative spread. I will then calculate Pearson's product-moment correlation coefficient and draw a scatter graph with a line of best fit to test for negative correlation.
    Examiner Tip: Always ensure your hypothesis identifies two measurable variables and states the expected relationship. Use precise statistical language such as 'positive correlation', 'significant difference' or 'association' rather than vague comparative words like 'better'.
    Pitfall: In the interpretation and evaluation strand, students frequently describe what their graphs show but fail to relate findings back to the original hypothesis or discuss limitations and possible improvements.
    ❌ Weak Answer (Loses Marks):My scatter graph shows a negative correlation. The line of best fit goes down. This proves my hypothesis was correct.
    Example improved answer:The scatter graph shows a moderate negative correlation (r = -0.62), suggesting that as daily social media use increases, GCSE Maths mock scores tend to decrease. This supports my hypothesis. However, correlation does not imply causation; other factors such as revision time or prior attainment may confound the relationship. The sample was limited to 50 Year 11 students from one school, so findings may not be generalisable. To improve reliability, I would increase the sample size, collect data from multiple schools, and control for confounding variables by also recording weekly revision hours.
    Examiner Tip: Always link your conclusion back to the hypothesis, state whether it is supported or not, and include at least two specific limitations with corresponding improvements. Use the phrase 'correlation does not imply causation' where appropriate.
    Step-by-Step Worked Solutions

    Question: A student wants to investigate whether there is a relationship between the number of hours spent revising per week and the score achieved in a statistics test (out of 50). They collect data from 30 students. The summary statistics are: Σx = 210, Σy = 1050, Σx² = 1750, Σy² = 39500, Σxy = 8100, where x = hours revised and y = test score. Calculate Pearson's product-moment correlation coefficient, r, and interpret your result in context.

    1. 1.Step 1: Identify the given values: n = 30, Σx = 210, Σy = 1050, Σx² = 1750, Σy² = 39500, Σxy = 8100.
    2. 2.Step 2: Use the formula r = (nΣxy - ΣxΣy) / sqrt([nΣx² - (Σx)²][nΣy² - (Σy)²]).
    3. 3.Step 3: Substitute: numerator = (30 × 8100) - (210 × 1050) = 243000 - 220500 = 22500.
    4. 4.Step 4: Denominator part 1: 30 × 1750 - 210² = 52500 - 44100 = 8400.
    5. 5.Step 5: Denominator part 2: 30 × 39500 - 1050² = 1185000 - 1102500 = 82500.
    6. 6.Step 6: r = 22500 / sqrt(8400 × 82500) = 22500 / sqrt(693000000) = 22500 / 26325.8 = 0.854 (to 3 s.f.).
    7. 7.Step 7: Interpret: r = 0.854 indicates a strong positive correlation between hours revised and test score. Students who revise more tend to achieve higher test scores.
    Final Answer: r = 0.854 (3 s.f.), indicating a strong positive correlation. More revision hours are associated with higher test scores.

    Question: A student is planning their D7 investigation. They want to test the hypothesis: 'The older a car is, the lower its value.' Describe how they should collect a sample of 40 cars, what data to record, and how they would process and represent the data to test this hypothesis. Include one potential source of bias and how to reduce it.

    1. 1.Step 1: Define the population as all used cars advertised for sale in the local area. Use a sampling frame such as a local car sales website or newspaper listings.
    2. 2.Step 2: Select a sample of 40 cars using systematic sampling: number the listings and pick every 5th car until 40 are selected. This reduces selection bias compared to convenience sampling.
    3. 3.Step 3: Record two variables for each car: age (in years, discrete) and value (in pounds, continuous). Ensure data is collected consistently, e.g. from the same website on the same day.
    4. 4.Step 4: Process data by drawing a scatter graph with age on the x-axis and value on the y-axis. Calculate Pearson's r to measure the strength of negative correlation.
    5. 5.Step 5: Identify bias: sampling only from one website may exclude private sellers or certain car types. Reduce bias by using multiple sources and stratifying by car type (e.g. hatchback, SUV).
    Final Answer: Use systematic sampling from a defined sampling frame, record age and value for 40 cars, plot a scatter graph, calculate r, and interpret the negative correlation. Reduce bias by using multiple sources and stratified sampling.
    Active Recall Memory Test
    What are the four stages of the statistical enquiry cycle?
    Key Fact: Planning, data collection, processing and representation, and interpretation and evaluation.
    What is the difference between primary and secondary data?
    Key Fact: Primary data is collected first-hand by the researcher for the specific investigation; secondary data is collected by someone else for a different purpose and reused.
    When should you use the median instead of the mean?
    Key Fact: Use the median when the data is skewed or contains outliers, as it is not affected by extreme values.
    What does a correlation coefficient of r = 0.92 indicate?
    Key Fact: A very strong positive correlation between the two variables.
    Frequently Asked Questions
    What is D7 in AQA GCSE Statistics?
    D7 is the controlled assessment component of AQA GCSE Statistics, worth 20% of your final grade. It involves planning and carrying out a statistical investigation on a hypothesis of your choice, then analysing and evaluating your findings. It is marked out of 40 and assessed by your teacher, then moderated by AQA.
    How do I choose a good hypothesis for my D7 investigation?
    Choose a hypothesis that is testable and involves two measurable variables. For example, 'There is a positive correlation between hours spent on homework and test scores' is better than 'Homework helps students'. Ensure you can collect data for both variables and that the relationship can be statistically analysed.
    What sampling method should I use for my D7?
    The best method depends on your hypothesis and population. Random sampling is ideal for reducing bias, but stratified sampling ensures representation of subgroups. Systematic sampling is practical if you have a list. Always justify your choice in your plan and discuss limitations.
    How do I calculate Pearson's correlation coefficient?
    Use the formula r = (nΣxy - ΣxΣy) / sqrt([nΣx² - (Σx)²][nΣy² - (Σy)²]). You need the sum of x, sum of y, sum of x², sum of y², and sum of xy, plus the number of pairs, n. Show all steps and round to 2 or 3 significant figures.
    What do examiners look for in the evaluation section of D7?
    Examiners want a critical evaluation that links back to your hypothesis, discusses specific limitations (e.g. sample size, bias, confounding variables), and suggests realistic improvements. Avoid generic comments; be precise about how limitations affect your conclusions and how changes would improve reliability or validity.
    Can I use secondary data for my D7 investigation?
    Yes, you can use secondary data, but you must evaluate its reliability and relevance. Primary data is often preferred because you control the collection process, but secondary data can be useful for comparing or extending your investigation. Always cite your sources and discuss any potential bias.