C3b — AQA GCSE Statistics
Test yourself on C3b with AQA GCSE practice questions.
7 days Premium · Then free forever · No card, no charge
Your focus
- Recognise where errors in construction lead to graphical misrepresentation including but not limited to incorrect scales, truncated axis, distorted sizing.
C3b exam tips
Quick Revision Summary (Key Takeaway)
C3b in AQA GCSE Statistics covers bivariate data analysis, focusing on scatter graphs, correlation, lines of best fit, and the interpretation of relationships between two variables. Students must calculate and interpret the product moment correlation coefficient (PMCC) and understand its limitations when describing real-world data.
Topic Overview
C3b focuses on bivariate data, where two variables are measured for each item in a sample. Students learn to represent this data on scatter graphs, describe the correlation (positive, negative, or none), and draw a line of best fit to make predictions. The topic also introduces the product moment correlation coefficient (PMCC) as a numerical measure of linear correlation, and emphasizes the difference between correlation and causation.
Understanding bivariate data is essential for analysing real-world relationships, such as the link between temperature and ice cream sales, or study time and exam performance. It builds on prior knowledge of representing data and calculating summary statistics, and it prepares students for more advanced statistical techniques like regression and hypothesis testing. This topic is heavily examined in AQA GCSE Statistics, often appearing in both calculator and non-calculator papers.
Key Concepts
- →Scatter graphs visually display the relationship between two continuous variables, with each point representing an individual data pair.
- →Correlation describes the strength and direction of a linear relationship: positive, negative, or no correlation; strong, moderate, or weak.
- →The line of best fit is a straight line that best represents the trend on a scatter graph, used to make predictions within the data range (interpolation) or outside it (extrapolation).
- →The product moment correlation coefficient (PMCC), denoted r, quantifies linear correlation on a scale from -1 to 1, where -1 is perfect negative, 0 is no linear correlation, and 1 is perfect positive.
- →Correlation does not imply causation; a strong correlation between two variables may be due to a confounding variable or coincidence.
Examiner Tips
- 💡When describing correlation, always comment on the strength (strong, moderate, weak), direction (positive, negative), and context (the variables involved). For example, 'There is a moderate positive correlation between hours of exercise and heart rate.'
- 💡When making predictions from a line of best fit, state whether the prediction is interpolation or extrapolation and comment on reliability. Interpolation within the data range is more reliable than extrapolation.
- 💡For PMCC questions, show your working clearly, especially the substitution into the formula. Round your final answer to an appropriate degree of accuracy (usually 3 significant figures) and always interpret the value in context.
Common Mistakes
- Students often think that a strong correlation means one variable causes the other to change. Correction: Correlation only indicates a relationship; causation requires experimental evidence or a plausible mechanism.
- Students may believe that the PMCC measures any relationship, including non-linear ones. Correction: PMCC only measures the strength of a linear relationship; a curved relationship could have r near 0 even if there is a strong non-linear association.
- Students sometimes interpret the y-intercept of a line of best fit as a meaningful value even when x=0 is not within the data range or is unrealistic. Correction: The y-intercept should only be interpreted if it makes sense in context and is within the range of observed data.
Revision Plan
- 1Day 1-2: Review the basics of scatter graphs: plotting points, identifying outliers, and describing correlation in words. Practice with past paper questions.
- 2Day 3-4: Learn to draw a line of best fit by eye and use it to make predictions. Understand the concepts of interpolation and extrapolation, and when predictions are reliable.
- 3Day 5-6: Study the PMCC formula and practice calculating r from summary statistics. Interpret the value of r in context, including the coefficient of determination (r²).
- 4Day 7-8: Focus on the distinction between correlation and causation. Work through examples where a confounding variable explains the relationship.
- 5Day 9-10: Complete a mixed set of exam-style questions on C3b, including those requiring explanation and interpretation. Review mark schemes to understand how marks are awarded.
Exam Question Types
- 📋Describe the correlation shown in a scatter graph and comment on any outliers. Advice: Use the words 'positive', 'negative', 'strong', 'weak', and refer to the variables in context.
- 📋Calculate the product moment correlation coefficient from summary statistics and interpret it. Advice: Show all steps in the formula, round to 3 s.f., and comment on strength and direction in context.
- 📋Use a line of best fit to estimate a value and comment on the reliability of the estimate. Advice: State whether it is interpolation or extrapolation, and mention any limitations.
- 📋Explain why correlation does not imply causation in a given scenario. Advice: Identify a possible confounding variable and explain how it could affect both variables.
Command Word Expectations (AQA)
Give a detailed account of the correlation, including strength, direction, and context. For example, 'There is a strong positive correlation between temperature and ice cream sales.'
Show all steps in the calculation, substitute correctly into the formula, and give the final answer to an appropriate degree of accuracy (usually 3 significant figures).
Explain what the calculated value or graph means in the context of the problem. For PMCC, comment on strength and direction; for predictions, comment on reliability.
How Students Lose Marks (Examiner Pitfalls)
Step-by-Step Worked Solutions
Question: A researcher collects data on the number of hours spent revising (x) and the test scores (y) for 10 students. The summary statistics are: Σx = 50, Σy = 600, Σx² = 300, Σy² = 38000, Σxy = 3200. Calculate the product moment correlation coefficient (PMCC) and interpret your result in context.
- 1.Step 1: Identify the formula for PMCC: r = (nΣxy - ΣxΣy) / sqrt((nΣx² - (Σx)²)(nΣy² - (Σy)²)).
- 2.Step 2: Substitute the given values with n = 10: numerator = 10*3200 - 50*600 = 32000 - 30000 = 2000.
- 3.Step 3: Calculate the denominator: sqrt((10*300 - 50²)(10*38000 - 600²)) = sqrt((3000 - 2500)(380000 - 360000)) = sqrt(500 * 20000) = sqrt(10000000) = 3162.27766.
- 4.Step 4: Compute r = 2000 / 3162.27766 = 0.632455532.
- 5.Step 5: Interpret: r = 0.632 (to 3 s.f.), indicating a moderate positive correlation between hours spent revising and test scores. This suggests that as revision hours increase, test scores tend to increase, but the relationship is not very strong.
Question: A scatter graph shows a strong negative correlation between the age of a car (in years) and its resale value (in thousands of pounds). The line of best fit is given by y = -1.5x + 20. Interpret the gradient and y-intercept in context. Predict the resale value of a 6-year-old car and comment on the reliability of this prediction.
- 1.Step 1: Interpret the gradient: The gradient is -1.5, meaning that for each additional year of age, the resale value decreases by approximately £1,500 on average.
- 2.Step 2: Interpret the y-intercept: The y-intercept is 20, which suggests that a brand new car (age 0 years) would have a resale value of £20,000. However, this may not be realistic as cars depreciate immediately after purchase.
- 3.Step 3: Predict for a 6-year-old car: Substitute x = 6 into the equation: y = -1.5(6) + 20 = -9 + 20 = 11. So the predicted resale value is £11,000.
- 4.Step 4: Comment on reliability: The prediction is within the range of the data (assuming the data included cars up to at least 6 years old), so it is likely reliable. However, if the data did not include cars of that age, the prediction would be an extrapolation and less reliable.