E11c — AQA GCSE Statistics
Test yourself on E11c with AQA GCSE practice questions.
7 days Premium · Then free forever · No card, no charge
Your focus
- Use action and warning lines in quality assurance sampling applications.
E11c exam tips
Quick Revision Summary (Key Takeaway)
E11c in AQA GCSE Statistics refers to the statistical analysis and interpretation of bivariate data using scatter diagrams, correlation, and lines of best fit. Students must calculate and interpret the product moment correlation coefficient (PMCC) and understand its limitations when describing relationships between two quantitative variables.
Topic Overview
E11c is a key topic in AQA GCSE Statistics that focuses on bivariate data analysis. You will learn how to construct and interpret scatter diagrams, identify different types of correlation (positive, negative, none), and draw a line of best fit to make predictions. The topic also covers the calculation and interpretation of the product moment correlation coefficient (PMCC), a numerical measure of linear correlation.
Understanding E11c is essential for analysing relationships between two variables in real-world contexts, such as height and weight, temperature and ice cream sales, or study time and exam performance. It builds on your knowledge of representing data and calculating summary statistics, and it prepares you for more advanced statistical techniques like regression. Mastery of this topic will enable you to critically evaluate claims of correlation and avoid common pitfalls such as confusing correlation with causation.
Key Concepts
- →Scatter diagrams visually display the relationship between two quantitative variables, with one variable on the x-axis and the other on the y-axis.
- →Correlation describes the strength and direction of a linear relationship: positive correlation (both variables increase together), negative correlation (one increases as the other decreases), or no correlation.
- →The product moment correlation coefficient (PMCC), denoted r, is a number between -1 and 1 that quantifies the strength and direction of linear correlation. Values close to 1 or -1 indicate strong correlation; values close to 0 indicate weak or no linear correlation.
- →A line of best fit is a straight line drawn through the centre of the data points on a scatter diagram, used to make predictions. It should pass through the mean point (x̄, ȳ).
- →Correlation does not imply causation; a strong correlation between two variables does not mean that one causes the other. There may be a confounding variable or the relationship may be coincidental.
Examiner Tips
- 💡When interpreting the PMCC, always comment on both the strength (e.g., strong, moderate, weak) and the direction (positive or negative) of the correlation, and relate it to the context of the question.
- 💡For scatter diagram questions, ensure you label axes correctly with units and use a sensible scale. When drawing a line of best fit, use a ruler and aim for roughly equal numbers of points above and below the line.
- 💡In questions asking about predictions, always state whether the prediction is interpolation (within the data range) or extrapolation (outside the data range) and comment on reliability accordingly.
Common Mistakes
- Students often think that a strong correlation means one variable causes the other. Correction: Correlation only indicates a relationship; causation requires further evidence, often from controlled experiments.
- Students may believe that the PMCC can only be positive. Correction: The PMCC can be negative, indicating a negative linear correlation, and ranges from -1 to 1.
- Students sometimes draw a line of best fit that connects the first and last points rather than balancing points above and below the line. Correction: The line of best fit should minimise the distances to all points and pass through the mean point.
Revision Plan
- 1Day 1-2: Revise the basics of scatter diagrams: how to plot points, label axes, and describe correlation in words. Practice identifying positive, negative, and no correlation from given diagrams.
- 2Day 3-4: Learn the formula for PMCC and practice calculating it using summary statistics. Check your answers using a calculator or spreadsheet. Focus on interpreting the value in context.
- 3Day 5-6: Practice drawing lines of best fit on scatter diagrams and using them to make predictions. Understand the difference between interpolation and extrapolation and when predictions are reliable.
- 4Day 7-8: Work through exam-style questions on E11c, including those that require interpretation and comments on correlation vs causation. Review mark schemes to understand what examiners expect.
- 5Day 9-10: Complete a timed practice paper or set of questions under exam conditions. Review any mistakes and revisit weak areas. Create a summary sheet of key points and formulas.
Exam Question Types
- 📋Description and interpretation of scatter diagrams: You may be asked to describe the correlation shown in a scatter diagram and comment on any outliers. Advice: Use precise language such as 'strong positive correlation' and mention the context.
- 📋Calculation of the PMCC: You will be given summary statistics and asked to calculate r. Advice: Show all steps clearly, use the formula correctly, and round your answer to 2 or 3 decimal places as required.
- 📋Drawing and using a line of best fit: You may need to draw a line of best fit on a scatter diagram and use it to estimate a value. Advice: Use a ruler, ensure the line passes through the mean point, and read values accurately from the graph.
- 📋Commenting on correlation vs causation: You may be asked to evaluate a statement that implies causation from correlation. Advice: State that correlation does not imply causation and suggest possible confounding variables.
Command Word Expectations (AQA)
You must use the given data or formula to work out a numerical answer. Show all steps of your working, and round appropriately if required. For PMCC, use the formula and substitute correctly.
You must explain what a result means in the context of the question. For example, interpret the PMCC by stating the strength and direction of correlation and what it implies about the variables.
You must give a brief statement or opinion based on the data, often about reliability or causation. For example, comment on whether a prediction is reliable or whether a correlation implies causation.
How Students Lose Marks (Examiner Pitfalls)
Step-by-Step Worked Solutions
Question: A researcher collects data on the number of hours studied (x) and exam scores (y) for 10 students. The summary statistics are: Σx = 50, Σy = 600, Σx² = 300, Σy² = 38000, Σxy = 3200. Calculate the product moment correlation coefficient (PMCC) and interpret your result.
- 1.Step 1: Identify given facts: n = 10, Σx = 50, Σy = 600, Σx² = 300, Σy² = 38000, Σxy = 3200.
- 2.Step 2: Apply the PMCC formula: r = (nΣxy - ΣxΣy) / sqrt((nΣx² - (Σx)²)(nΣy² - (Σy)²)).
- 3.Step 3: Substitute values: numerator = 10*3200 - 50*600 = 32000 - 30000 = 2000. Denominator part 1: 10*300 - 50² = 3000 - 2500 = 500. Denominator part 2: 10*38000 - 600² = 380000 - 360000 = 20000. So r = 2000 / sqrt(500*20000) = 2000 / sqrt(10000000) = 2000 / 3162.27766 ≈ 0.632.
- 4.Step 4: Interpret: r ≈ 0.632, indicating a moderate positive linear correlation between hours studied and exam scores.
Question: A scatter diagram shows a strong negative correlation between the age of a car (in years) and its resale value (in thousands of pounds). The line of best fit is given by y = -1.5x + 20. Predict the resale value of a 6-year-old car and comment on the reliability of this prediction.
- 1.Step 1: Identify the equation of the line of best fit: y = -1.5x + 20, where x is age in years and y is resale value in thousands.
- 2.Step 2: Substitute x = 6 into the equation: y = -1.5(6) + 20 = -9 + 20 = 11.
- 3.Step 3: Interpret the result: The predicted resale value is 11 thousand pounds, i.e., £11,000.
- 4.Step 4: Comment on reliability: Since the data shows a strong negative correlation and the prediction is within the range of the data (assuming ages up to around 10 years), the prediction is likely reliable. However, extrapolation beyond the data range would be unreliable.