B5b — AQA GCSE Statistics
Test yourself on B5b with AQA GCSE practice questions.
7 days Premium · Then free forever · No card, no charge
Your focus
- Know why data may need to be 'cleaned' before further processing, including issues that arise on spreadsheets and apply techniques to clean data in context.
B5b exam tips
Quick Revision Summary (Key Takeaway)
B5b in AQA GCSE Statistics covers bivariate data analysis, specifically scatter graphs, correlation, and lines of best fit. You must be able to interpret and calculate the correlation coefficient (Pearson's r) and use the line of best fit to make predictions, understanding the difference between correlation and causation.
Topic Overview
B5b focuses on bivariate data, where two variables are measured for each item in a sample. You will learn to display this data on scatter graphs, describe the correlation (positive, negative, or none), and identify outliers. Understanding correlation is essential for analysing relationships in real-world contexts, from science to economics.
You will also calculate Pearson's correlation coefficient (r) to quantify the strength and direction of a linear relationship. This topic builds on your knowledge of summary statistics and graphs, and it prepares you for more advanced statistical analysis, such as regression lines and hypothesis testing.
Key Concepts
- →Scatter graphs display bivariate data, with one variable on the x-axis and the other on the y-axis. Each point represents an individual data pair.
- →Correlation describes the relationship between two variables: positive (both increase together), negative (one increases as the other decreases), or no correlation. Strength ranges from weak to strong.
- →Pearson's correlation coefficient (r) is a number between -1 and 1 that measures the strength and direction of a linear relationship. r = 1 is perfect positive, r = -1 is perfect negative, r = 0 is no linear correlation.
- →The line of best fit is a straight line drawn through the centre of the data points on a scatter graph. It can be used to make predictions within the range of the data (interpolation) but predictions outside the range (extrapolation) are less reliable.
- →Correlation does not imply causation: just because two variables are correlated does not mean one causes the other. There may be a lurking variable or coincidence.
Examiner Tips
- 💡When describing correlation, always mention both strength (e.g., strong, moderate, weak) and direction (positive or negative), and relate it to the context of the variables.
- 💡For calculations of r, show all your working, including the sums and the formula. This allows you to gain method marks even if you make an arithmetic error.
- 💡When interpreting the line of best fit, always comment on the reliability of predictions, especially if extrapolating beyond the data range.
Common Mistakes
- Students often think that a strong correlation means one variable causes the other. Correction: Correlation only shows a relationship; causation requires further evidence, often from controlled experiments.
- Students may believe that if r is close to 0, there is no relationship at all. Correction: r measures only linear relationships; there could be a non-linear relationship that r does not detect.
- Students sometimes confuse the sign of r with the strength. Correction: The sign indicates direction (positive or negative), while the magnitude (closeness to 1 or -1) indicates strength.
Revision Plan
- 1Day 1-2: Revise the basics of scatter graphs: plotting points, describing correlation, and identifying outliers. Practice with past paper questions.
- 2Day 3-4: Learn the formula for Pearson's r and practice calculating it with given summary statistics. Check your answers using a calculator or spreadsheet.
- 3Day 5-6: Understand the line of best fit: how to draw it by eye and how to use it for predictions. Practice interpreting the meaning of the gradient and intercept in context.
- 4Day 7-8: Work through exam-style questions on correlation and causation, focusing on explaining why correlation does not imply causation.
- 5Day 9-10: Complete a full past paper section on B5b under timed conditions, then review your mistakes and revisit weak areas.
Exam Question Types
- 📋Describe the correlation shown in a scatter graph and comment on any outliers. Advice: Use the words 'positive', 'negative', 'strong', 'weak', and relate to the variables.
- 📋Calculate Pearson's correlation coefficient from summary statistics. Advice: Write down the formula, substitute carefully, and show all steps.
- 📋Interpret the line of best fit and make a prediction. Advice: Always state the units and comment on reliability if the prediction is outside the data range.
- 📋Explain why correlation does not imply causation in a given context. Advice: Suggest a possible lurking variable and explain how it could affect both variables.
Command Word Expectations (AQA)
Give a detailed account of the correlation, including strength and direction, and mention any outliers or unusual features.
Show all steps of the calculation, including the formula and substitution, and give the final answer to an appropriate degree of accuracy (usually 2 decimal places for r).
Explain what the value of r or the line of best fit means in the context of the problem, including the direction and strength of the relationship, and the reliability of any predictions.
How Students Lose Marks (Examiner Pitfalls)
Step-by-Step Worked Solutions
Question: A researcher records the number of hours studied (x) and the exam score (y) for 8 students. The summary statistics are: Σx = 40, Σy = 480, Σx² = 250, Σy² = 32000, Σxy = 2800. Calculate Pearson's correlation coefficient (r) and interpret it.
- 1.Step 1: Identify the formula for Pearson's r: r = (nΣxy - ΣxΣy) / sqrt((nΣx² - (Σx)²)(nΣy² - (Σy)²)).
- 2.Step 2: Substitute the given values: n = 8, Σx = 40, Σy = 480, Σx² = 250, Σy² = 32000, Σxy = 2800.
- 3.Step 3: Calculate the numerator: 8*2800 - 40*480 = 22400 - 19200 = 3200.
- 4.Step 4: Calculate the denominator: sqrt((8*250 - 40²)(8*32000 - 480²)) = sqrt((2000 - 1600)(256000 - 230400)) = sqrt(400 * 25600) = sqrt(10240000) = 3200.
- 5.Step 5: Divide to find r: r = 3200 / 3200 = 1.
- 6.Step 6: Interpret: r = 1 indicates a perfect positive correlation between hours studied and exam score.
Question: A scatter graph shows a moderate negative correlation between the age of a car (in years) and its value (in £1000s). The line of best fit is y = -1.5x + 20. Predict the value of a car that is 6 years old and comment on the reliability of this prediction.
- 1.Step 1: Identify the line of best fit equation: y = -1.5x + 20, where y is value in £1000s and x is age in years.
- 2.Step 2: Substitute x = 6 into the equation: y = -1.5(6) + 20 = -9 + 20 = 11.
- 3.Step 3: Interpret the result: y = 11 means £11,000.
- 4.Step 4: Comment on reliability: The prediction is within the range of the data (assuming ages up to at least 6 years were plotted), so it is likely reliable. However, if 6 years is outside the observed range, it would be an extrapolation and less reliable.