Anomalous Results
Spotting and handling outliers fairly
Lesson 118 of 4,500 · Measurement, Units and SI
Learning objectives
- Identify anomalous results in tables and on graphs
- Suggest possible causes of an anomalous result
- Decide fairly whether to exclude, repeat or keep an unusual reading
Introduction
Occasionally one reading in a set sticks out: four titres agree to within 0.05 cm³ and the fifth is almost a whole cubic centimetre higher. Such a reading is called an anomalous result or outlier . Handling it well is an important scientific skill. Including it blindly can distort the mean; deleting it without thought can hide a genuine discovery or give a false impression of precision.
Core explanation
What makes a result anomalous? An anomalous result is one that does not fit the pattern of the other data. In a table of repeats, it is a value much further from the mean than the others. On a graph, it is a point well away from the line or curve of best fit that the rest of the points follow.
Spotting anomalies in tables. Look along each set of repeat readings. If most values cluster closely and one is clearly separate, it is probably anomalous. For example, in titres of 22.45, 22.50, 23.30 and 22.40 cm³, the value 23.30 cm³ is anomalous.
Spotting anomalies on graphs. Plot all the points first, then draw a line of best fit through the trend shown by most of them. Any point far from the line is an outlier. The line should not be forced through an anomalous point, and the point should be circled or labelled rather than rubbed out.
Possible causes. Anomalies usually come from a one-off mistake rather than steady random scatter, for example:
- Misreading a scale or writing digits in the wrong order (24.1 recorded as 21.4). - Overshooting the end point in a titration. - Using the wrong concentration or volume of a reagent in one run. - An air bubble in the tip of a burette or pipette. - A sudden change in conditions, such as a draught or temperature change.
Handling anomalies fairly.
1. Check the working and the recording — a simple transcription error can sometimes be corrected. 2. Repeat the measurement if possible. If the repeat agrees with the others, the anomaly was a one-off mistake. 3. Exclude the anomalous value from the mean only when there is good reason to think it is a mistake, and state clearly in your write-up that you have excluded it and why. 4. Keep and investigate a result that cannot be explained, especially if it repeats. An unexpected result can reveal a real effect.
Honesty matters. Removing inconvenient data just because it spoils the conclusion is scientific misconduct. Good practice is to record every reading, mark anomalies clearly, and be transparent about decisions.
Step-by-step reasoning
When you meet an unusual reading:
1. Compare it with the other repeats or with the trend on a graph. 2. Decide whether it is clearly out of line, not just at the edge of normal scatter. 3. Look for a likely cause in your notes or procedure. 4. Repeat the reading if time allows. 5. Exclude it from calculations only with a reason, and record that decision.
Visual explanation
Imagine a scatter graph where eight points lie neatly along a rising straight line. A ninth point sits well below the line. The line of best fit passes through the eight, ignoring the ninth, which is ringed and labelled "anomalous — possible misreading".
Real-world analogy
If your daily journey to school usually takes 20 to 23 minutes but one day it took 55 minutes because of a road accident, you would not count that day when planning how early to leave. You would, however, note what happened in case accidents on that road turn out to be common.
Real-world example
In 1985, scientists monitoring the atmosphere over Antarctica reported unexpectedly low ozone readings. Some satellite processing had earlier flagged very low values as suspect, but ground measurements confirmed them. The "anomalies" turned out to be the ozone hole, showing why unusual data must be investigated rather than simply discarded.
Why?
Why should an anomalous result be excluded from a mean? The mean is meant to estimate the true value using readings affected only by small random errors. A value caused by a blunder is not part of that scatter; including it would drag the mean away from the true value and exaggerate the uncertainty.
Common misconception
"Any reading that is different from the rest is anomalous and should be deleted." Every data set has natural scatter. A reading slightly further from the mean may be perfectly valid. Only values clearly outside the pattern, preferably with an identifiable cause, should be treated as anomalous.
Worked example
Question: Mass losses of 1.52 g, 1.49 g, 1.87 g and 1.51 g are recorded in repeat runs. Identify any anomaly and calculate the mean.
Reasoning: Three values lie between 1.49 g and 1.52 g; 1.87 g is far from them, so it is anomalous. Mean of the rest = (1.52 + 1.49 + 1.51) ÷ 3 = 4.52 ÷ 3 = 1.51 g (to two decimal places).
Answer: 1.87 g is anomalous; mean = 1.51 g.
Quick check
1. On a graph, should the line of best fit be drawn through an anomalous point? Answer: No — draw the line through the trend of the other points and circle or label the anomalous one.
Exam focus
Examiners often give a table with one odd value. Identify it by name and value, exclude it from the mean, and say so in your answer. If asked to improve an investigation, suggest repeating readings so that anomalies can be identified. Never rub out an anomalous point on a graph — circle it.
Advanced insight
Statisticians have formal tests for outliers, such as the Grubbs test or Dixon's Q test, which compare how far a suspect value lies from the rest with the overall spread of the data. These tests set clear, objective criteria for rejection, reducing the risk that scientists discard data simply because it is inconvenient.
Summary
An anomalous result is a reading that does not fit the pattern of the others. Spot it in tables of repeats or as a point far from a line of best fit. Check for mistakes, repeat the measurement, and exclude it from calculations only with a clear reason that you record. Unexplained anomalies should be investigated, because they can reveal real effects.
Practice questions
1. Define an anomalous result. Answer: A reading that does not fit the pattern or trend shown by the other results. 2. Titres of 26.10, 26.15, 26.95 and 26.05 cm³ are recorded. Which is anomalous, and what is the mean of the others? Answer: 26.95 cm³ is anomalous; mean = (26.10 + 26.15 + 26.05) ÷ 3 = 78.30 ÷ 3 = 26.10 cm³. 3. Suggest two possible causes of an anomalous titre. Answer: Overshooting the end point, and misreading or wrongly recording a burette reading (an air bubble in the burette tip is another). 4. Why is it poor practice to delete an anomalous result without mentioning it? Answer: It hides data from others, may remove a genuine effect, and makes results look more precise than they are, which is dishonest.