Skip to topic
    ← Back to course topics

    Data presentation and interpretation (AS Unit 2: Applied Mathematics A) — WJEC A-Level Mathematics

    Test yourself on Data presentation and interpretation (AS Unit 2: Applied Mathematics A) with WJEC A-Level practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    Data presentation and interpretation (AS Unit 2: Applied Mathematics A) explained

    Empirical frequency presentations approximate theoretical probability models.

    Read the full explanation

    As sample size grows and class intervals narrow, a relative frequency histogram smooths towards a continuous probability density curve, where total area equals 1 and the area between two values gives P(a ≤ X ≤ b). Discrete relative frequency charts similarly approach theoretical probability mass functions. Recognising symmetry, unimodality or skewness in sample diagrams helps statisticians choose candidate models, but shape alone is not decisive: a positively skewed histogram is consistent with several distributions, such as Poisson or exponential, and further evidence is needed. Comparing observed proportions with theoretical probabilities assesses model fit.

    Your focus

    1. Demonstrate how relative frequency histograms approximate continuous probability density functions.
    2. Identify candidate theoretical probability models from the shape and skewness of sample data, recognising that shape alone is not decisive.
    3. Compare observed sample proportions against theoretical model probabilities to assess goodness-of-fit.

    Data presentation and interpretation (AS Unit 2: Applied Mathematics A) exam tips

    Quick Revision Summary (Key Takeaway)

    Data presentation and interpretation in WJEC AS Unit 2 covers measures of central tendency and dispersion, along with visual representations like histograms and box plots. Mastering linear interpolation, coding, standard deviation calculations, and outlier identification is essential for statistical problem solving.

    Topic Overview

    Data presentation and interpretation forms the foundation of the statistics component in WJEC AS Unit 2 Applied Mathematics A. It equips students with the statistical tools required to summarise, model, and critique real-world numerical datasets accurately.

    This unit covers central tendency (mean, median, mode), dispersion (variance, standard deviation, interquartile range), visual representations (histograms, cumulative frequency diagrams, box plots), and data cleaning through outlier identification. These methods underpin hypothesis testing and probability modelling in later units.

    Key Concepts
    • →Measures of central tendency: calculating mean from raw or grouped data, estimating median and quartiles via linear interpolation.
    • →Measures of spread: calculating variance and standard deviation using Sigma x and Sigma x^2, and understanding their sensitivity to extreme values.
    • →Data transformations and coding: applying linear transformations y = a * x + b and understanding the differing effects on location versus dispersion.
    • →Visual representation: plotting histograms where frequency is proportional to area (frequency density = frequency / class width) and constructing box plots.
    • →Outlier identification: applying statistical rules such as Q1 - 1.5*IQR / Q3 + 1.5*IQR or mean +/- 2*standard deviation.
    Marking Points
    • Explaining that relative frequency areas represent probabilities and that the total area under a probability density curve equals 1.
    • Identifying that a symmetrical bell-shaped histogram suggests a Normal probability distribution as a candidate model.
    • Recognising that skewness alone does not determine a distribution: positive skew is consistent with several models, so further evidence is needed.
    • Comparing empirical sample proportions with theoretical probabilities to evaluate model fit.
    Examiner Tips
    • 💡Link shape to candidate distribution: a bell-shaped histogram suggests a Normal model, but state that shape alone is not conclusive.
    • 💡State clearly that the area under a continuous probability density curve corresponds to probability, and that total area equals 1.
    • 💡Write down the coding substitution explicitly (e.g. y = (x - 100)/5) and check your decoded answers intuitively against original raw values.
    • 💡Always state units in final interpretation questions, and provide direct contextual comparisons (e.g. comparing both average and consistency) when comparing two distributions.
    • 💡Use your scientific calculator's statistics mode to verify values for Sigma x, Sigma x^2, mean, and standard deviation to eliminate arithmetic errors.
    Common Mistakes
    • Confusing empirical sample histograms with exact theoretical probability distributions; a histogram only approximates a model.
    • Forgetting that the total area under a probability density curve must equal 1.
    • Claiming that positive skew alone implies a Poisson model; skewness is consistent with several distributions and is not decisive.
    • Plotting frequency on the vertical axis of a histogram instead of frequency density when class widths are unequal.
    • Applying addition/subtraction adjustments to standard deviation during decoding, forgetting that shift transformations do not alter spread.
    • Failing to identify true continuous boundaries for grouped discrete data when estimating medians or drawing histograms.
    Revision Plan
    1. 1Day 1-3: Master calculation of mean, variance, and standard deviation from summary statistics and frequency tables.
    2. 2Day 4-6: Practice linear coding extensively, ensuring complete clarity on which measures change under addition versus multiplication.
    3. 3Day 7-9: Work through grouped data problems involving linear interpolation and drawing histograms with frequency density.
    4. 4Day 10-12: Solve mixed past-paper questions focusing on comparing distributions using mean/standard deviation versus median/IQR.
    Exam Question Types
    • 📋Summary statistics and coding questions: Evaluating mean and standard deviation from Sigma(x - a) forms.
    • 📋Grouped frequency interpolation: Estimating quartiles, percentiles, or the median from continuous data intervals.
    • 📋Histogram and box plot analysis: Interpreting areas to find unknown frequencies and checking for outliers to plot modified box-and-whisker plots.
    • 📋Comparative commentary questions: Writing two-part comparative sentences evaluating average performance and spread in practical contexts.
    Command Word Expectations (WJEC)
    Calculate

    Requires a numerical answer supported by relevant working steps, formulas, and appropriate rounding.

    Interpret

    Explain the statistical result in the specific context of the problem, referring to real-world units and trends rather than abstract numbers.

    Compare

    Provide two distinct comparison statements: one evaluating central tendency (which group has higher average) and one evaluating dispersion (which group is more consistent), both in context.

    How Students Lose Marks (Examiner Pitfalls)
    Pitfall: Confusing class boundaries when calculating linear interpolation from grouped frequency tables, especially with discrete data rounded to integers.
    ❌ Weak Answer (Loses Marks):Lower boundary is taken directly as the written class limit, e.g. using 20 instead of 19.5 for the class 20-29.
    Example improved answer:For the class interval 20-29 representing continuous time data rounded to the nearest integer, the class boundaries are 19.5 and 29.5, giving a class width of 10. The interpolation formula must use L = 19.5.
    Examiner Tip: Always check whether the data is continuous or rounded discrete values, and clearly write down the true lower and upper class boundaries before substituting into the linear interpolation formula.
    Pitfall: Failing to decode the standard deviation correctly after applying a linear coding transformation y = (x - a) / b.
    ❌ Weak Answer (Loses Marks):Multiplying the standard deviation of y by b and then adding a.
    Example improved answer:Standard deviation is a measure of dispersion and is unaffected by addition or subtraction. Therefore, if y = (x - a) / b, then sigma_y = sigma_x / b, meaning sigma_x = b * sigma_y.
    Examiner Tip: Remember that location measures (mean, median, mode) are affected by both addition and multiplication, but dispersion measures (standard deviation, variance, IQR) are only affected by multiplication/division.
    Step-by-Step Worked Solutions

    Question: The masses, w kg, of 60 newborn calves are recorded. A coded summary gives Sigma(w - 30) = 240 and Sigma(w - 30)^2 = 4320. Calculate the mean and standard deviation of the original masses w.

    1. 1.Step 1: Define the coded variable y = w - 30, where n = 60, Sigma y = 240, and Sigma y^2 = 4320.
    2. 2.Step 2: Calculate the mean of the coded variable: y_bar = (Sigma y) / n = 240 / 60 = 4.
    3. 3.Step 3: Decode to find the mean of w: w_bar = y_bar + 30 = 4 + 30 = 34 kg.
    4. 4.Step 4: Calculate the variance of y: Var(y) = (Sigma y^2 / n) - (y_bar)^2 = (4320 / 60) - (4)^2 = 72 - 16 = 56.
    5. 5.Step 5: Calculate the standard deviation of y: sigma_y = sqrt(56) = 7.4833... kg.
    6. 6.Step 6: Since w = y + 30, the scale factor is 1, so the standard deviation of w equals sigma_y: sigma_w = sqrt(56) = 7.48 kg (to 3 s.f.).
    Final Answer: Mean mass = 34 kg, Standard deviation = 7.48 kg (3 s.f.)

    Question: In a study of journey times, Q1 = 18 minutes and Q3 = 42 minutes. An outlier is defined as any value lying more than 1.5 * IQR outside the interquartile range. Determine whether a journey time of 80 minutes is an outlier, showing all boundary calculations.

    1. 1.Step 1: Calculate the interquartile range: IQR = Q3 - Q1 = 42 - 18 = 24 minutes.
    2. 2.Step 2: Calculate 1.5 * IQR = 1.5 * 24 = 36 minutes.
    3. 3.Step 3: Find the upper outlier boundary: Upper fence = Q3 + (1.5 * IQR) = 42 + 36 = 78 minutes.
    4. 4.Step 4: Find the lower outlier boundary: Lower fence = Q1 - (1.5 * IQR) = 18 - 36 = -18 minutes (practically 0 minutes).
    5. 5.Step 5: Compare the value 80 minutes to the upper boundary: 80 > 78.
    6. 6.Step 6: State the final conclusion clearly referencing the threshold.
    Final Answer: Since 80 > 78 minutes, a journey time of 80 minutes is classified as an outlier.
    Active Recall Memory Test
    What is the formula for frequency density on a histogram?
    Key Fact: Frequency density = Frequency / Class width.
    If variable x is coded as y = (x - a) / b, how do the mean and standard deviation of x relate to y?
    Key Fact: Mean(x) = b * Mean(y) + a; Standard Deviation(x) = b * Standard Deviation(y).
    When is the median and IQR preferred over the mean and standard deviation to summarise data?
    Key Fact: When the data distribution is heavily skewed or contains extreme outlier values, as median and IQR are resistant to outliers.
    What are the two standard mathematical definitions for an outlier?
    Key Fact: Values smaller than Q1 - 1.5*IQR or greater than Q3 + 1.5*IQR; alternatively, values outside Mean +/- 2*Standard Deviation (or 3*SD depending on specification instruction).
    Frequently Asked Questions
    Why does frequency density matter on a histogram instead of just plotting frequency?
    In a histogram, the area of each bar is proportional to the frequency, not the height. If classes have unequal widths, plotting frequency on the vertical axis distorts the visual representation, making wider classes appear artificially over-represented. Frequency density ensures that area accurately reflects frequency across variable intervals.
    How do I choose whether to use (n/2) or ((n+1)/2) for the median position?
    For discrete, un-grouped raw data, use (n + 1) / 2 to find the exact middle position. For continuous or grouped data represented in a frequency table where linear interpolation is used, standard WJEC convention uses n / 2, n / 4, and 3n / 4 to identify median and quartile positions.
    Does standard deviation change if I add a constant to every data value?
    No, standard deviation measures how spread out the data points are relative to the mean. Adding a constant shifts every value and the mean by the same amount, leaving the distances between data points and the mean unchanged, so standard deviation remains identical.
    What must I include when an exam question asks to 'compare two distributions'?
    You must make two separate contextual comparisons: one comparing a measure of average (e.g. median or mean) and one comparing a measure of spread (e.g. IQR or standard deviation). Never just state numbers; explicitly use comparative terms such as 'higher on average' and 'more consistent / less varied'.
    How do I spot whether data is discrete or continuous in grouped frequency questions?
    Look at the class intervals in the question table. If intervals are written as 10-19, 20-29 and describe integer counts (e.g. number of books), the data is discrete and the true class boundaries are 9.5 to 19.5, 19.5 to 29.5. If intervals are written with inequalities like 10 <= t < 20, the data is continuous and the boundaries are exactly 10 and 20.