Skip to topic
    ← Back to course topics

    Statistics — AQA GCSE Mathematics

    Test yourself on Statistics with AQA GCSE practice questions.

    Start free

    7 days Premium · Then free forever · No card, no charge

    Statistics explained

    A population is every member of the group a question is about.

    Read the full explanation

    A sample is the smaller set you actually measure, and you use what the sample shows to say something about the whole population. Scale up with proportion: if 12 out of 80 sampled trees are diseased, that is a proportion of 0.15, so in a population of 2000 trees you would expect about 300. An estimate is only trustworthy if the sample is large enough and chosen without bias, so asking people leaving a gym about exercise tells you about gym users, not about the town. Larger samples give more reliable estimates but never certainty. Questions ask you to estimate a population figure from a sample, or to say why a sample may not represent the population.

    interpret and construct tables, charts and diagrams, including frequency tables, bar charts, pie charts and pictograms for categorical data, vertical line charts for ungrouped discrete numerical data, and know their appropriate use

    Data is often first organised in a frequency table, which lists categories and their counts (frequencies), sometimes using a tally column to help count raw data. Categorical data, which sorts into named groups, can then be displayed in several ways. A bar chart uses bars of equal width with gaps between them. A pictogram uses symbols, so a key is essential to show what each symbol represents. A pie chart shows proportions of a whole. To calculate the angle for each sector, use the formula: (frequency ÷ total frequency) × 360°. For example, if a category has a frequency of 5 out of a total of 40, its angle is (5 ÷ 40) × 360° = 45°. Ungrouped discrete numerical data, such as shoe size, is best shown with a vertical line chart, where the height of each thin line above the value shows its frequency.

    including tables and line graphs for time series data

    Time series data records measurements taken at successive time intervals, often regular (e.g. monthly sales) but sometimes uneven. Organise the data in a table with a time column (e.g. Quarter 1, Quarter 2) and a value column, then plot a line graph. Time goes on the horizontal (x) axis and the measured quantity on the vertical (y) axis. Plot each point from the table and join consecutive points with straight lines in time order. These lines show change over time but do not represent actual readings between plotted points. Describe the trend by its overall direction across the whole period, not short-term ups and downs. For example, sales may rise overall despite a winter dip. Read axis scales carefully, as they may not start at zero.

    construct and interpret diagrams for grouped discrete data and continuous data, ie histograms with equal and unequal class intervals and cumulative frequency graphs, and know their appropriate use (Higher tier only)

    Higher tier only. For grouped data, histograms are used, and the bars must touch. For grouped discrete data (e.g., number of items 10-19), class boundaries are used to close the gaps (e.g., the bar spans 9.5 to 19.5). If class intervals are all equal, the vertical axis shows frequency. If class intervals are unequal, the vertical axis must be frequency density, calculated as frequency ÷ class width. This ensures the area of each bar (height × width) represents the frequency. For example, a class of width 10 holding 20 values has a frequency density of 20 ÷ 10 = 2. A cumulative frequency graph plots a running total against the upper boundary of each class. Points are joined with a smooth curve, which is used to estimate the median (at half the total frequency) and quartiles.

    interpret, analyse and compare the distributions of data sets from univariate empirical distributions through: •• appropriate graphical representation involving discrete, continuous and grouped data •• appropriate measures of central tendency (median, mean, mode and modal class) and spread (range, including consideration of outliers)

    To compare two data sets, give one statement about a typical value and one about spread, each backed by a figure. The mean is the total divided by how many values there are; the median is the middle value of the ordered list, or halfway between the two middle values when the count is even; the mode is the most common value. Grouped data hides individual values, so estimate the mean using midpoints: multiply each midpoint by its frequency, total the products, then divide by the total frequency, and name the modal class rather than a single mode. Range is largest minus smallest. A value far from the rest is an outlier: it pulls the mean towards itself but barely moves the median, so the median often describes skewed data better. Distributions can also be compared from graphs such as back-to-back stem-and-leaf diagrams or comparative bar charts, reading off typical values, spread and shape.

    •• including box plots •• including quartiles and inter-quartile range (Higher tier only)

    This is Higher tier only. A box plot shows five figures: the smallest value, the lower quartile, the median, the upper quartile and the largest value. Put the data in order first. The lower quartile is the value a quarter of the way along the ordered list and the upper quartile is the value three quarters of the way along. Subtract the lower quartile from the upper quartile to get the inter-quartile range, which measures the spread of the middle half of the data, so a single extreme value at either end cannot inflate it. To set two box plots against each other, say which median is larger and which inter-quartile range is smaller, reading both off the same scale. Cumulative frequency curves are a common source for the five figures.

    apply statistics to describe a population

    Describing a group of people or objects with numbers means choosing figures that answer the question asked, then saying what they mean in context. Pick an average that suits the data: the mean uses every value, the median resists extreme values, and the mode is the only one that works for categories such as favourite sport. Add a measure of spread so the description is more than a single number, and keep the units. For example, saying that the mean mass of the apples is 112 g and that their range is 38 g describes both a typical apple and how varied the crop is. Where the figures come from a sample, say that the description is an estimate for the whole group. Questions give raw data or a table and ask for a description in words.

    use and interpret scatter graphs of bivariate data recognise correlation

    Bivariate data pairs two measurements taken from the same person or object, such as height and arm span. Each pair becomes one point on a scatter graph, with one quantity on each axis. The pattern of the points is what you read. Points rising to the right show positive correlation, meaning that as one quantity increases so does the other. Points falling to the right show negative correlation. A shapeless cloud shows no correlation. Give strength as well as direction, so "strong positive correlation" when the points lie close to a line and "weak" when they are widely scattered. A point well away from the pattern is an outlier and may be a recording error, so name it rather than pretending it is not there.

    know that it does not indicate causation draw estimated lines of best fit make predictions interpolate and extrapolate apparent trends whilst knowing the dangers of so doing

    A line of best fit is one straight ruled line through the pattern on a scatter graph, with roughly as many points above it as below, and it need not pass through the origin. Use it to predict: go up from a value on one axis, across to the line, then read the matching value on the other axis. A prediction inside the range of the plotted data is interpolation and is reasonably safe. A prediction beyond that range is extrapolation, and the pattern may not carry on, so the answer is unreliable. A pattern only shows that two quantities move together. Ice cream sales and sunburn cases rise together because both depend on hot weather, not because one produces the other.

    Your focus

    1. Describe the difference between a population and a sample, and state why a figure taken from a sample is only an estimate.
    2. Estimate a population figure by working out the sample proportion and multiplying it by the population size, as 0.15 of 2000 trees.
    3. Give a reason why a sample may not represent its population, naming the flaw such as its size or where it was taken.
    Show all 35 objectives
    1. Infer a property of a distribution from a sample, for example using the sample mean to estimate the population mean or the sample range to estimate spread.
    2. Construct a frequency table from raw data using a tally chart.
    3. Calculate pie chart angles by dividing the category frequency by the total frequency and multiplying by 360°.
    4. Draw or complete a bar chart, pictogram or vertical line chart with labelled axes and a key where needed.
    5. Explain which type of chart is most appropriate for a given set of categorical or ungrouped discrete data.
    6. Complete and interpret a table of time series data.
    7. Construct a time series graph by plotting points from a table and joining consecutive points with straight lines.
    8. Describe the overall trend shown by a time series graph, supporting the description with values.
    9. Interpret a time series graph, paying careful attention to the scale on the vertical axis.
    10. Distinguish between histograms with equal class widths (frequency axis) and unequal class widths (frequency density axis).
    11. Calculate frequency density as frequency divided by class width and use it to draw a histogram.
    12. Determine correct class boundaries for grouped discrete and continuous data.
    13. Draw a cumulative frequency graph by plotting running totals at each upper class boundary and use it to estimate the median and quartiles.
    14. Work out the mean, median, mode and range of a list of values, taking the median as the middle of the ordered list.
    15. Estimate the mean of grouped data from midpoints multiplied by frequencies divided by the total frequency, and name the modal class.
    16. Interpret and compare distributions shown in graphs such as back-to-back stem-and-leaf diagrams and comparative bar charts.
    17. Compare two distributions with one statement about a typical value and one about spread, each quoting a figure.
    18. Explain why an outlier drags the mean towards itself but barely moves the median, so the median suits skewed data better.
    19. Find the median and both quartiles from ordered data or by reading across a cumulative frequency curve.
    20. Draw a box plot against a scale, with the box from lower to upper quartile and whiskers out to the smallest and largest values.
    21. Work out the inter-quartile range as upper quartile minus lower quartile and state that it measures the middle half of the data.
    22. Compare two box plots by saying which median is larger and which inter-quartile range is smaller, in the context of the question.
    23. Justify the choice of average for a data set, noting the mode is the only one that works for categories such as favourite sport.
    24. Describe a population with one average and one measure of spread, both with units, such as mean mass 112 g and range 38 g.
    25. Explain why a description built from a sample must be given as an estimate for the whole group rather than an exact value.
    26. Complete a scatter graph by plotting each pair of measurements from the same person or object as one point.
    27. Describe correlation by direction and strength, such as strong positive when the points lie close to a line.
    28. Interpret the correlation in the words of the question and identify any outlier, saying why it does not fit the pattern.
    29. Draw a line of best fit as one ruled straight line following the trend, with roughly as many points above it as below.
    30. Estimate a value from a line of best fit by going up from one axis, across to the line and off the other axis.
    31. Justify whether a prediction is reliable by deciding if it is interpolation inside the data range or extrapolation beyond it.
    32. Explain why correlation does not show cause, naming a factor such as hot weather that could drive both quantities.

    Statistics exam tips

    Quick Revision Summary (Key Takeaway)

    Statistics in AQA GCSE Mathematics covers collecting, representing, analysing and interpreting data using tables, graphs, averages and measures of spread. You must be able to calculate mean, median, mode and range, construct and read charts such as bar charts, histograms and cumulative frequency graphs, and describe correlation and distributions in context.

    Topic Overview

    Statistics is the branch of mathematics concerned with collecting, organising, summarising and interpreting data. In AQA GCSE Mathematics, you will learn to construct and interpret a wide range of statistical diagrams including bar charts, pie charts, histograms, box plots and cumulative frequency graphs. You will also calculate and compare averages (mean, median, mode) and measures of spread (range and interquartile range) to describe data sets.

    This topic is essential for understanding real-world data in fields such as science, business and social research. It also connects to probability and algebra, and often appears in crossover questions with other areas of the GCSE. Mastering statistics helps you make informed decisions and spot misleading representations of data.

    Key Concepts
    • →Averages: mean, median and mode, and when each is most appropriate to use.
    • →Measures of spread: range and interquartile range, and how they describe consistency.
    • →Constructing and interpreting statistical diagrams: bar charts, pie charts, histograms, cumulative frequency graphs and box plots.
    • →Scatter graphs and correlation: positive, negative and no correlation, and drawing a line of best fit.
    • →Sampling: understanding populations, samples, and the importance of random sampling to avoid bias.
    Marking Points
    • the correct sample proportion, written as a fraction or a decimal, even if the scaling that follows is wrong
    • multiplying that proportion by the population size, and one for the resulting estimate
    • a reason that names the sampling flaw, such as the sample being too small or drawn from only one place
    • credit for presenting the figure as an estimate rather than an exact population count
    • inferring a distributional property from a sample, such as using the sample mean to estimate the population mean, or the sample range or interquartile range to estimate spread
    • A correctly constructed frequency table from raw data, including a tally column if required.
    • A correct method for calculating a pie chart angle, such as (frequency ÷ total frequency) × 360°.
    • Accurately drawn charts (e.g., bars of correct height, sectors of correct angle within tolerance) with labelled axes or sectors, and a key for a pictogram.
    • A correct value read from a chart or table, or a correct interpretation of what the chart shows.
    • Correctly complete a table of time series data from given information, matching each time period to its value.
    • Plot all points at the correct coordinates from the table, within a small tolerance.
    • Join consecutive points with straight line segments in the correct time order.
    • Describe the overall trend by its direction over time, supporting the description with values read from the graph or table.
    • Interpret a time series graph correctly when the vertical axis does not start at zero, avoiding false proportional comparisons.
    • Correctly identifying class boundaries for grouped discrete or continuous data.
    • A correct calculation of frequency density for at least one class where intervals are unequal.
    • Bars drawn with no gaps, at the correct heights (frequency or frequency density) and spanning the correct class boundaries.
    • Cumulative frequencies correctly calculated and plotted against the upper class boundary, with a sensible curve drawn through the points.
    • a correct method for the mean, namely the total of the values divided by how many values there are
    • using midpoints on grouped data, one for the total of midpoint multiplied by frequency, and one for dividing by the total frequency
    • a comparison of averages and a separate mark for a comparison of spread, each quoting a value
    • naming the modal class, or the class that contains the median, from a running total
    • interpreting a back-to-back stem-and-leaf diagram or comparative bar chart, for example reading the median and range of each data set from the display
    • comparing the shape of two distributions from their graphs, for example noting which is more symmetric or which has the longer tail
    • the median and one for both quartiles, taken from ordered data or read off a cumulative frequency graph
    • a box running from the lower quartile to the upper quartile with the median line inside it, drawn against a scale
    • whiskers reaching the smallest and the largest values
    • the inter-quartile range as upper quartile minus lower quartile, even if one quartile is wrong
    • comparing medians and one for comparing spread, each in the context of the question
    • choosing an appropriate average and a correct value for it
    • a measure of spread quoted alongside that average
    • a statement written in the context of the question, with units
    • making clear that a figure based on a sample is an estimate rather than an exact value
    • all points plotted correctly, within half a small square
    • naming the type of correlation, and a further mark for stating its strength
    • interpreting the correlation in the words of the question rather than in general terms
    • identifying an outlier and saying why it does not fit the pattern
    • a ruled straight line following the trend with points balanced on both sides of it
    • the read-off drawn on the graph, and one for a predicted value inside the accepted range
    • judging a prediction reliable or unreliable, with a reason that refers to the range of the data
    • stating that a pattern does not show cause, supported by another factor that could explain both quantities
    Examiner Tips
    • 💡Write the sample proportion as a fraction of the sample size first, then multiply by the population size, so you avoid rounding early.
    • 💡When a question asks for a reason a sample is unreliable, name the group that is missing or over-represented, rather than only saying that the sample is small.
    • 💡For distributional inference, quote the sample statistic (mean, median, range or interquartile range) and state that it estimates the corresponding population value.
    • 💡Check that your pie chart angles total 360° before you start drawing. If they are a degree out due to rounding, adjust the largest sector.
    • 💡When constructing a frequency table from a long list of data, use a tally system (e.g., groups of five) to avoid errors in counting.
    • 💡Before plotting, work out what one small square is worth on each axis and write it down.
    • 💡When describing a trend, state both the direction and the period, e.g. 'a steady rise across the four years'.
    • 💡For a histogram with unequal intervals, always add columns for class width and frequency density to the table before you start drawing.
    • 💡On a cumulative frequency graph, mark the median (1/2 total), lower quartile (1/4 total) and upper quartile (3/4 total) positions on the vertical axis first, then read across to the curve and down to the answer.
    • 💡Write each comparison as one sentence containing a figure, for example that the mean height is larger, so this group is taller on average.
    • 💡A smaller range means the values are more consistent, so use the word consistent when you compare spread.
    • 💡When a question shows a graph, quote values read from the graph in your comparison rather than recalculating from raw data.
    • 💡Mark the five values lightly on the scale before drawing the box, so the median line lands in the right place.
    • 💡Write one comparison sentence about the medians and one about the inter-quartile ranges, and say what each means for the situation described.
    • 💡Answer in a full sentence that names what is being described, the statistic and the units.
    • 💡If one value sits far from the rest, choose the median and say that the extreme value would distort the mean.
    • 💡Give three things when you describe a scatter graph: strength, direction, and what it means for the two named quantities.
    • 💡Plot with a sharp pencil and a small cross, because a thick dot can sit on the wrong side of a gridline.
    • 💡Check whether the value you are asked about lies inside the plotted range, because that decides whether you call the prediction reliable.
    • 💡Draw the read-off lines in pencil and leave them on the graph, since they can earn a method mark on their own.
    • 💡Always show your working for calculations, especially for the mean from a frequency table. Method marks are often available even if the final answer is wrong.
    • 💡When interpreting graphs, refer to the context of the question. For example, say 'the median time is 25 seconds' rather than just 'the median is 25'.
    • 💡Check the scale on axes carefully. A common mistake is misreading the scale, leading to incorrect values for quartiles or medians.
    Common Mistakes
    • Dividing by the population size instead of multiplying, so the estimate comes out smaller than the sample itself.
    • Treating the scaled-up figure as an exact count of the population rather than an estimate.
    • Answering "the sample was biased" without saying which group was left out or over-represented.
    • Treating the sample as though it were the whole population, so the answer describes only the people who were asked.
    • Using the sample mean as if it were the exact population mean, ignoring that a different sample would give a different value.
    • Drawing bars touching for categorical data. This is the convention for a histogram, not a bar chart, where gaps are required.
    • Calculating pie chart angles incorrectly, for example as percentages of 100 instead of proportions of 360°.
    • Using a bar chart for ungrouped discrete numerical data where a vertical line chart is more appropriate.
    • In a pictogram, drawing a partial symbol that does not accurately represent the required fraction of the key's value.
    • Swapping the axes, putting the measured quantity on the horizontal axis instead of time. Correction: time always goes on the x-axis.
    • Describing a short-term fluctuation between two points as the overall trend for the whole graph. Correction: describe the direction across the full period, noting fluctuations separately.
    • Misreading a vertical axis that does not start at zero, leading to incorrect calculations or comparisons. Correction: check the scale and starting value before reading or comparing.
    • Plotting points at even spacing on the x-axis when the time intervals in the table are uneven. Correction: space points according to the actual time intervals.
    • Plotting frequency instead of frequency density on the vertical axis when class widths are unequal.
    • Leaving gaps between histogram bars or using incorrect class boundaries for discrete data (e.g., a bar for 10-19 drawn from 10 to 19).
    • Plotting cumulative frequency at the class midpoint or lower boundary instead of the upper boundary.
    • Forgetting that the area of a histogram bar represents the frequency, especially when asked to find the frequency of part of a group.
    • Dividing by the number of classes instead of the total frequency when estimating a mean from a grouped table.
    • Giving the position of the median rather than the value that sits in that position.
    • Comparing averages only and saying nothing about spread, which loses the second comparison mark.
    • Picking the middle of the list before putting the values in order.
    • Reading a back-to-back stem-and-leaf diagram as if both sides shared one ordered list, instead of treating each side as its own data set.
    • Reading quartiles straight from an unordered list, so the box is built from the wrong values.
    • Drawing the whiskers to the quartiles and the box edges to the extremes, which turns the plot inside out.
    • Subtracting the wrong way round and giving a negative inter-quartile range.
    • Saying one group did better without quoting either the median or the inter-quartile range.
    • Quoting a mean to several decimal places when the data are whole counts, which claims more precision than the data carry.
    • Giving an average on its own, with nothing about how spread out the values are.
    • Using the mean on categorical data, where only the mode has any meaning.
    • Dropping the units, so the figure describes nothing in particular.
    • Joining the points up dot to dot, which turns a scatter graph into a line graph.
    • Writing "correlation" with no direction, so the answer never says whether the pattern rises or falls.
    • Calling a clear downward pattern no correlation because the values are getting smaller.
    • Judging the strength of the correlation with an obvious outlier included in the pattern.
    • Drawing the line through the first and last points instead of through the whole pattern.
    • Predicting well beyond the plotted values and still calling the answer reliable.
    • Claiming that one quantity makes the other happen because the pattern is strong.
    • Letting an outlier drag the line of best fit away from the bulk of the points.
    • Students often think the mean is always the best average to use. In fact, the median is better when there are extreme values (outliers) because it is not affected by them.
    • Many students confuse the interquartile range with the range. The range is the difference between the highest and lowest values, while the interquartile range is the spread of the middle 50% of data.
    • When drawing histograms, students frequently use frequency on the vertical axis instead of frequency density. For unequal class widths, you must calculate frequency density = frequency ÷ class width.
    Revision Plan
    1. 1Week 1: Revise the definitions and calculations for mean, median, mode and range. Practice with simple data sets and frequency tables.
    2. 2Week 1: Learn to construct and interpret bar charts, pie charts and scatter graphs. Focus on labelling axes and understanding correlation.
    3. 3Week 2: Master cumulative frequency graphs and box plots. Practice finding median, quartiles and interquartile range.
    4. 4Week 2: Study histograms with unequal class widths. Learn to calculate frequency density and draw histograms accurately.
    5. 5Week 2: Complete past paper questions on statistics, timing yourself to build exam confidence. Review mistakes carefully.
    Exam Question Types
    • 📋Calculation questions: Find the mean, median, mode or range from a list or frequency table. Always show your method.
    • 📋Graph interpretation: Read values from a cumulative frequency graph or box plot to find median, quartiles and interquartile range. Use a ruler and state units.
    • 📋Comparison questions: Compare two data sets using averages and measures of spread. Comment on consistency and typical values in context.
    • 📋Construction questions: Draw a bar chart, pie chart, histogram or box plot from given data. Ensure correct scales and labels.
    Command Word Expectations (AQA)
    Calculate

    You must work out a numerical answer. Show all steps of your working, as method marks are awarded even if the final answer is incorrect.

    Compare

    You must comment on similarities and differences between two sets of data, usually using averages and measures of spread. Refer to the context and give numerical values.

    Interpret

    You must explain what a result or graph means in the context of the question. Use the correct statistical terminology and relate it to the real-world situation.

    How Students Lose Marks (Examiner Pitfalls)
    Pitfall: Confusing the mean, median and mode when data is presented in a frequency table, especially when the mean is required from grouped data.
    ❌ Weak Answer (Loses Marks):The mean is 4 because 4 appears most often in the table.
    Example improved answer:The mode is 4 because it has the highest frequency. To find the mean from a frequency table, multiply each value by its frequency, sum these products, then divide by the total frequency. For grouped data, use midpoints and state that the mean is an estimate.
    Examiner Tip: Always label which average you are calculating. For grouped data, write 'estimated mean' and show the midpoint column clearly. Check that your total frequency matches the sum of the frequency column.
    Pitfall: Misinterpreting cumulative frequency graphs, particularly reading the median, quartiles and interquartile range incorrectly or not using the correct axis scale.
    ❌ Weak Answer (Loses Marks):The median is 10 because that is where the curve starts.
    Example improved answer:To find the median from a cumulative frequency graph, locate half of the total frequency on the cumulative frequency axis, draw a horizontal line to the curve, then read down to the variable axis. The lower quartile is at one quarter of the total frequency and the upper quartile at three quarters. The interquartile range is upper quartile minus lower quartile.
    Examiner Tip: Always write down the total frequency first. Use a ruler to draw accurate lines. State the quartiles clearly and show the subtraction for the interquartile range. Check the scale on both axes carefully.
    Step-by-Step Worked Solutions

    Question: The table shows the number of goals scored by a football team in 20 matches. Calculate the mean number of goals per match. Goals scored: 0 (frequency 3), 1 (frequency 6), 2 (frequency 7), 3 (frequency 3), 4 (frequency 1).

    1. 1.Step 1: Multiply each number of goals by its frequency: 0×3=0, 1×6=6, 2×7=14, 3×3=9, 4×1=4.
    2. 2.Step 2: Sum these products: 0+6+14+9+4 = 33.
    3. 3.Step 3: Sum the frequencies to check total: 3+6+7+3+1 = 20.
    4. 4.Step 4: Divide the total goals by total frequency: 33 ÷ 20 = 1.65.
    Final Answer: The mean number of goals per match is 1.65.

    Question: A cumulative frequency graph shows the times taken by 80 students to complete a puzzle. The median time is 25 seconds, the lower quartile is 18 seconds and the upper quartile is 34 seconds. Calculate the interquartile range and interpret it in context.

    1. 1.Step 1: Identify the lower quartile (LQ) = 18 seconds and upper quartile (UQ) = 34 seconds.
    2. 2.Step 2: Calculate the interquartile range: IQR = UQ - LQ = 34 - 18 = 16 seconds.
    3. 3.Step 3: Interpret: The middle 50% of students took between 18 and 34 seconds, a spread of 16 seconds.
    Final Answer: The interquartile range is 16 seconds. This means the middle 50% of students' times are spread over 16 seconds.
    Active Recall Memory Test
    What is the difference between the range and the interquartile range?
    Key Fact: The range is the difference between the highest and lowest values in a data set. The interquartile range is the difference between the upper quartile and lower quartile, representing the spread of the middle 50% of the data.
    How do you calculate the mean from a frequency table?
    Key Fact: Multiply each value by its frequency, sum these products, then divide by the total frequency (the sum of the frequencies).
    What is frequency density and when is it used?
    Key Fact: Frequency density is calculated as frequency divided by class width. It is used on the vertical axis of a histogram when class widths are unequal, so that the area of each bar represents the frequency.
    What does a correlation of -0.8 indicate on a scatter graph?
    Key Fact: It indicates a strong negative correlation, meaning that as one variable increases, the other tends to decrease.
    Frequently Asked Questions
    What is the difference between the mean, median and mode?
    The mean is the sum of all values divided by the number of values. The median is the middle value when data is ordered. The mode is the most frequent value. The mean uses all data but is affected by outliers; the median is resistant to outliers; the mode is useful for categorical data.
    How do I find the median from a cumulative frequency graph?
    First, find the total frequency (the highest value on the cumulative frequency axis). The median is at half of this total. Draw a horizontal line from that value on the cumulative frequency axis to the curve, then draw a vertical line down to the variable axis and read off the median. For quartiles, use one quarter and three quarters of the total frequency.
    What is the interquartile range and why is it useful?
    The interquartile range (IQR) is the difference between the upper quartile (UQ) and lower quartile (LQ). It measures the spread of the middle 50% of the data. It is useful because it ignores extreme values (outliers) and gives a better measure of spread for skewed data.
    How do I draw a histogram with unequal class widths?
    For each class, calculate the frequency density by dividing the frequency by the class width. Plot frequency density on the vertical axis and the variable on the horizontal axis. The area of each bar represents the frequency. Ensure the bars are drawn with no gaps and the correct widths.
    What does it mean if two data sets have the same mean but different ranges?
    It means that on average the values are similar, but one data set is more spread out than the other. The set with the larger range has more variability, so it is less consistent. When comparing, mention both the mean and the range (or interquartile range) to give a full picture.
    How do I know whether to use the mean or the median?
    Use the mean when the data is fairly symmetrical and you want to include all values. Use the median when there are outliers or skewed data, as it is not affected by extreme values. For example, house prices often use median because a few very expensive houses would skew the mean.