Data presentation and interpretation (AS Unit 2: Applied Mathematics A) — WJEC A-Level Mathematics
Test yourself on Data presentation and interpretation (AS Unit 2: Applied Mathematics A) with WJEC A-Level practice questions.
7 days Premium · Then free forever · No card, no charge
Data presentation and interpretation (AS Unit 2: Applied Mathematics A) explained
Empirical frequency presentations approximate theoretical probability models.
Read the full explanation
As sample size grows and class intervals narrow, a relative frequency histogram smooths towards a continuous probability density curve, where total area equals 1 and the area between two values gives P(a ≤ X ≤ b). Discrete relative frequency charts similarly approach theoretical probability mass functions. Recognising symmetry, unimodality or skewness in sample diagrams helps statisticians choose candidate models, but shape alone is not decisive: a positively skewed histogram is consistent with several distributions, such as Poisson or exponential, and further evidence is needed. Comparing observed proportions with theoretical probabilities assesses model fit.
Your focus
- Demonstrate how relative frequency histograms approximate continuous probability density functions.
- Identify candidate theoretical probability models from the shape and skewness of sample data, recognising that shape alone is not decisive.
- Compare observed sample proportions against theoretical model probabilities to assess goodness-of-fit.
Data presentation and interpretation (AS Unit 2: Applied Mathematics A) exam tips
Quick Revision Summary (Key Takeaway)
Data presentation and interpretation in WJEC AS Unit 2 covers measures of central tendency and dispersion, along with visual representations like histograms and box plots. Mastering linear interpolation, coding, standard deviation calculations, and outlier identification is essential for statistical problem solving.
Topic Overview
Data presentation and interpretation forms the foundation of the statistics component in WJEC AS Unit 2 Applied Mathematics A. It equips students with the statistical tools required to summarise, model, and critique real-world numerical datasets accurately.
This unit covers central tendency (mean, median, mode), dispersion (variance, standard deviation, interquartile range), visual representations (histograms, cumulative frequency diagrams, box plots), and data cleaning through outlier identification. These methods underpin hypothesis testing and probability modelling in later units.
Key Concepts
- →Measures of central tendency: calculating mean from raw or grouped data, estimating median and quartiles via linear interpolation.
- →Measures of spread: calculating variance and standard deviation using Sigma x and Sigma x^2, and understanding their sensitivity to extreme values.
- →Data transformations and coding: applying linear transformations y = a * x + b and understanding the differing effects on location versus dispersion.
- →Visual representation: plotting histograms where frequency is proportional to area (frequency density = frequency / class width) and constructing box plots.
- →Outlier identification: applying statistical rules such as Q1 - 1.5*IQR / Q3 + 1.5*IQR or mean +/- 2*standard deviation.
Marking Points
- Explaining that relative frequency areas represent probabilities and that the total area under a probability density curve equals 1.
- Identifying that a symmetrical bell-shaped histogram suggests a Normal probability distribution as a candidate model.
- Recognising that skewness alone does not determine a distribution: positive skew is consistent with several models, so further evidence is needed.
- Comparing empirical sample proportions with theoretical probabilities to evaluate model fit.
Examiner Tips
- 💡Link shape to candidate distribution: a bell-shaped histogram suggests a Normal model, but state that shape alone is not conclusive.
- 💡State clearly that the area under a continuous probability density curve corresponds to probability, and that total area equals 1.
- 💡Write down the coding substitution explicitly (e.g. y = (x - 100)/5) and check your decoded answers intuitively against original raw values.
- 💡Always state units in final interpretation questions, and provide direct contextual comparisons (e.g. comparing both average and consistency) when comparing two distributions.
- 💡Use your scientific calculator's statistics mode to verify values for Sigma x, Sigma x^2, mean, and standard deviation to eliminate arithmetic errors.
Common Mistakes
- Confusing empirical sample histograms with exact theoretical probability distributions; a histogram only approximates a model.
- Forgetting that the total area under a probability density curve must equal 1.
- Claiming that positive skew alone implies a Poisson model; skewness is consistent with several distributions and is not decisive.
- Plotting frequency on the vertical axis of a histogram instead of frequency density when class widths are unequal.
- Applying addition/subtraction adjustments to standard deviation during decoding, forgetting that shift transformations do not alter spread.
- Failing to identify true continuous boundaries for grouped discrete data when estimating medians or drawing histograms.
Revision Plan
- 1Day 1-3: Master calculation of mean, variance, and standard deviation from summary statistics and frequency tables.
- 2Day 4-6: Practice linear coding extensively, ensuring complete clarity on which measures change under addition versus multiplication.
- 3Day 7-9: Work through grouped data problems involving linear interpolation and drawing histograms with frequency density.
- 4Day 10-12: Solve mixed past-paper questions focusing on comparing distributions using mean/standard deviation versus median/IQR.
Exam Question Types
- 📋Summary statistics and coding questions: Evaluating mean and standard deviation from Sigma(x - a) forms.
- 📋Grouped frequency interpolation: Estimating quartiles, percentiles, or the median from continuous data intervals.
- 📋Histogram and box plot analysis: Interpreting areas to find unknown frequencies and checking for outliers to plot modified box-and-whisker plots.
- 📋Comparative commentary questions: Writing two-part comparative sentences evaluating average performance and spread in practical contexts.
Command Word Expectations (WJEC)
Requires a numerical answer supported by relevant working steps, formulas, and appropriate rounding.
Explain the statistical result in the specific context of the problem, referring to real-world units and trends rather than abstract numbers.
Provide two distinct comparison statements: one evaluating central tendency (which group has higher average) and one evaluating dispersion (which group is more consistent), both in context.
How Students Lose Marks (Examiner Pitfalls)
Step-by-Step Worked Solutions
Question: The masses, w kg, of 60 newborn calves are recorded. A coded summary gives Sigma(w - 30) = 240 and Sigma(w - 30)^2 = 4320. Calculate the mean and standard deviation of the original masses w.
- 1.Step 1: Define the coded variable y = w - 30, where n = 60, Sigma y = 240, and Sigma y^2 = 4320.
- 2.Step 2: Calculate the mean of the coded variable: y_bar = (Sigma y) / n = 240 / 60 = 4.
- 3.Step 3: Decode to find the mean of w: w_bar = y_bar + 30 = 4 + 30 = 34 kg.
- 4.Step 4: Calculate the variance of y: Var(y) = (Sigma y^2 / n) - (y_bar)^2 = (4320 / 60) - (4)^2 = 72 - 16 = 56.
- 5.Step 5: Calculate the standard deviation of y: sigma_y = sqrt(56) = 7.4833... kg.
- 6.Step 6: Since w = y + 30, the scale factor is 1, so the standard deviation of w equals sigma_y: sigma_w = sqrt(56) = 7.48 kg (to 3 s.f.).
Question: In a study of journey times, Q1 = 18 minutes and Q3 = 42 minutes. An outlier is defined as any value lying more than 1.5 * IQR outside the interquartile range. Determine whether a journey time of 80 minutes is an outlier, showing all boundary calculations.
- 1.Step 1: Calculate the interquartile range: IQR = Q3 - Q1 = 42 - 18 = 24 minutes.
- 2.Step 2: Calculate 1.5 * IQR = 1.5 * 24 = 36 minutes.
- 3.Step 3: Find the upper outlier boundary: Upper fence = Q3 + (1.5 * IQR) = 42 + 36 = 78 minutes.
- 4.Step 4: Find the lower outlier boundary: Lower fence = Q1 - (1.5 * IQR) = 18 - 36 = -18 minutes (practically 0 minutes).
- 5.Step 5: Compare the value 80 minutes to the upper boundary: 80 > 78.
- 6.Step 6: State the final conclusion clearly referencing the threshold.