D7 — AQA GCSE Statistics
Test yourself on D7 with AQA GCSE practice questions.
7 days Premium · Then free forever · No card, no charge
Your focus
- Use collected data to calculate estimates of probabilities.
D7 exam tips
Quick Revision Summary (Key Takeaway)
D7 in AQA GCSE Statistics is the controlled assessment (formerly coursework), worth 20% of the final grade, where students plan, collect, process and analyse data on a self-chosen hypothesis, then evaluate their findings. It assesses the full statistical enquiry cycle and is marked out of 40 across four strands: planning, data collection, processing and representation, and interpretation and evaluation.
Topic Overview
D7 is the controlled assessment component of AQA GCSE Statistics, accounting for 20% of the total marks. It requires students to independently plan and conduct a statistical investigation based on a hypothesis of their choice, following the full enquiry cycle: planning, data collection, processing and representation, and interpretation and evaluation. The assessment is marked out of 40 and is internally assessed then externally moderated.
This topic is crucial because it tests your ability to apply statistical knowledge to a real-world context, rather than just answering abstract questions. It develops skills in designing investigations, handling data critically, and communicating findings effectively. Success in D7 demonstrates a deep understanding of the statistical problem-solving process, which is essential for further study and everyday data literacy.
Key Concepts
- →The statistical enquiry cycle: a continuous four-stage process of planning, data collection, processing and representation, and interpretation and evaluation.
- →Hypothesis: a testable statement predicting a relationship or difference between two variables, which must be measurable and specific.
- →Sampling methods: random, systematic, stratified, and quota sampling, each with advantages and disadvantages regarding bias and representativeness.
- →Data types: primary vs secondary data, quantitative (discrete and continuous) vs qualitative, and the importance of consistency in collection.
- →Statistical measures: measures of central tendency (mean, median, mode), measures of dispersion (range, interquartile range, standard deviation), and correlation (Pearson's r, Spearman's rank).
Examiner Tips
- 💡Always justify your choice of sampling method and sample size in the planning section. Explain why it is appropriate for your hypothesis and population.
- 💡In the processing section, use appropriate graphs and calculations. For example, use a scatter graph and Pearson's r for correlation, or a box plot and standard deviation for comparing distributions.
- 💡In the evaluation, be critical: discuss limitations of your data collection, possible sources of bias, and suggest specific improvements. Avoid generic statements like 'I could have collected more data' without explaining how that would improve reliability.
Common Mistakes
- Students often think that any correlation proves cause and effect. Correction: correlation does not imply causation; a third variable may explain the relationship.
- Many believe that a larger sample size automatically removes bias. Correction: a large sample can still be biased if the sampling method is flawed (e.g. convenience sampling).
- Students sometimes confuse the mean and median, or use the mean for skewed data. Correction: the median is more appropriate for skewed data or when outliers are present.
Revision Plan
- 1Week 1: Review the statistical enquiry cycle and sampling methods. Choose a hypothesis and write a detailed plan, including variables, sampling method, and data collection strategy.
- 2Week 1-2: Collect your data, ensuring consistency and ethical considerations. Record data accurately in a table.
- 3Week 2: Process your data using appropriate calculations (e.g. means, standard deviation, correlation coefficient) and draw relevant graphs.
- 4Week 2: Interpret your findings, relate them back to your hypothesis, and evaluate the strengths and weaknesses of your investigation.
- 5Ongoing: Practice writing model answers for each strand using past D7 exemplars and mark schemes to understand the assessment criteria.
Exam Question Types
- 📋Planning question: 'Describe how you would collect data to test the hypothesis that...' Advice: Outline sampling method, sample size, variables, and how to ensure reliability.
- 📋Processing question: 'Calculate the correlation coefficient and interpret it in context.' Advice: Show all steps, use correct formula, and relate the value to the hypothesis.
- 📋Evaluation question: 'Discuss the limitations of your investigation and suggest improvements.' Advice: Be specific about sources of bias and how changes would improve validity.
Command Word Expectations (AQA)
Give a detailed account of the steps or features. In D7, this often requires outlining your sampling method, data collection process, or how you would represent data. Marks are awarded for specific, relevant detail, not vague statements.
Show all working and give your answer to an appropriate degree of accuracy (usually 2 or 3 significant figures). In D7, this may involve computing summary statistics or correlation coefficients. Marks are awarded for correct substitution and accurate arithmetic.
Make a judgement about the strengths and weaknesses of your investigation, supported by evidence. You must discuss limitations, possible improvements, and how these affect the validity and reliability of your conclusions. Marks are awarded for critical analysis, not just listing points.
How Students Lose Marks (Examiner Pitfalls)
Step-by-Step Worked Solutions
Question: A student wants to investigate whether there is a relationship between the number of hours spent revising per week and the score achieved in a statistics test (out of 50). They collect data from 30 students. The summary statistics are: Σx = 210, Σy = 1050, Σx² = 1750, Σy² = 39500, Σxy = 8100, where x = hours revised and y = test score. Calculate Pearson's product-moment correlation coefficient, r, and interpret your result in context.
- 1.Step 1: Identify the given values: n = 30, Σx = 210, Σy = 1050, Σx² = 1750, Σy² = 39500, Σxy = 8100.
- 2.Step 2: Use the formula r = (nΣxy - ΣxΣy) / sqrt([nΣx² - (Σx)²][nΣy² - (Σy)²]).
- 3.Step 3: Substitute: numerator = (30 × 8100) - (210 × 1050) = 243000 - 220500 = 22500.
- 4.Step 4: Denominator part 1: 30 × 1750 - 210² = 52500 - 44100 = 8400.
- 5.Step 5: Denominator part 2: 30 × 39500 - 1050² = 1185000 - 1102500 = 82500.
- 6.Step 6: r = 22500 / sqrt(8400 × 82500) = 22500 / sqrt(693000000) = 22500 / 26325.8 = 0.854 (to 3 s.f.).
- 7.Step 7: Interpret: r = 0.854 indicates a strong positive correlation between hours revised and test score. Students who revise more tend to achieve higher test scores.
Question: A student is planning their D7 investigation. They want to test the hypothesis: 'The older a car is, the lower its value.' Describe how they should collect a sample of 40 cars, what data to record, and how they would process and represent the data to test this hypothesis. Include one potential source of bias and how to reduce it.
- 1.Step 1: Define the population as all used cars advertised for sale in the local area. Use a sampling frame such as a local car sales website or newspaper listings.
- 2.Step 2: Select a sample of 40 cars using systematic sampling: number the listings and pick every 5th car until 40 are selected. This reduces selection bias compared to convenience sampling.
- 3.Step 3: Record two variables for each car: age (in years, discrete) and value (in pounds, continuous). Ensure data is collected consistently, e.g. from the same website on the same day.
- 4.Step 4: Process data by drawing a scatter graph with age on the x-axis and value on the y-axis. Calculate Pearson's r to measure the strength of negative correlation.
- 5.Step 5: Identify bias: sampling only from one website may exclude private sellers or certain car types. Reduce bias by using multiple sources and stratifying by car type (e.g. hatchback, SUV).