Outliers and Data Exclusion

Investigating unusual observations without convenient post-hoc deletion

Lesson 4387 of 4,500 · Research Methods, Data Analysis and Literature

Learning objectives

Introduction

An unexpected point can result from a spilled sample, mislabeling, instrument carryover or a real chemical event. Deleting it because it weakens a preferred claim turns scientific judgment into selection bias. Keeping every value without investigation can also conceal a known mistake. A defensible workflow preserves the raw record, checks plausible causes and reports how the conclusion depends on the decision.

Core explanation

An outlier is defined relative to an expected distribution or fitted model; it is not automatically invalid. A high product yield might reveal a side reaction, a local hotspot or a weighing error. A low battery capacity might signal a defective tab weld rather than the active material's intrinsic behavior. The observation's unusualness is a reason to investigate, not a reason to delete. Plot raw data with sample identity, run order and batch; inspect associated metadata before applying a statistical rule.

Predefine exclusion criteria where possible. Examples include instrument failure documented by a failed calibration check, lost sample, known wrong reagent or measurement outside the validated calibration range. Criteria should apply regardless of whether a result supports or opposes the hypothesis. A point should not be excluded solely because its residual exceeds a convenient threshold, especially with small sample sizes. Statistical outlier tests themselves require assumptions about the distribution and can be misused when many tests are performed.

Trace the observation through the process. Check labels, raw chromatograms or spectra, integration settings, preparation notebook, blank sequence, standards, instrument alarms and environmental logs. Repeat the measurement of the same prepared sample to test readout problems, then prepare a fresh independent sample to test reproducibility. These repeats answer different questions. If a known error is found, record it and follow the prespecified exclusion rule. If no error is found, the point may be genuine variation or a missing mechanism.

An observation can be influential without being far from the fitted line. In regression, a point at an extreme predictor value can strongly affect slope. Conversely, a large residual near the center may have less effect on parameters. Inspect influence diagnostics and compare fitted estimates with and without the point. This is a sensitivity analysis, not an automatic license to select the preferred result. State both analyses when the decision materially changes the conclusion.

Robust statistical methods can reduce sensitivity to occasional extreme values, but they do not replace chemical diagnosis. A robust fit may reveal the central trend while preserving the unusual point for further investigation. If the unusual observation reflects a real second population, such as a minority defective manufacturing batch, averaging it away can be practically dangerous. NIST's guidance on residual examination treats structured departures as information about assumptions and process, not merely numerical inconveniences.

Transparent reporting means describing the number of attempts, exclusions and reasons. Preserve raw data and original images; show primary results under the planned rule and sensitivity results under reasonable alternatives. A claim that depends entirely on one contested deletion is weaker than one robust to it. Honest uncertainty may be more valuable than an artificially neat plot.

Step-by-step reasoning

Before analysis, define objective failure criteria and document all runs. When an unusual value appears, inspect raw records and instrument quality controls without changing the value. Test for a documented error and repeat if feasible at the appropriate experimental-unit level. Apply the predeclared rule consistently. Calculate the result with the point and, when relevant, without it; compare practical conclusions. Report the decision and remaining uncertainty.

Visual explanation

Draw a scatterplot with one point far above the fitted line. Next to it show a decision tree: documented instrument or preparation failure → exclude under a stated rule, preserving the record; no identified failure → retain and investigate mechanism; uncertain status → report sensitivity. A second panel shows a far-right point close to the line that changes the slope strongly, illustrating influence distinct from residual size.

Real-world analogy

If one thermometer reports a much higher room temperature, do not discard it immediately. Check whether it was placed beside a heater or has a dead battery; compare with an independent calibrated thermometer. The reading might reveal a real hotspot. A chemical outlier deserves the same combination of skepticism and openness.

Real-world example

Five nominally identical catalytic reactions give yields of 71%, 73%, 72%, 74% and 40%. The 40% vessel's notebook shows a stopped stirrer for half the reaction, and its temperature log records a hotspot. If mixing was a predefined validity condition, the run can be excluded from the primary catalyst estimate with the failure documented. If there is no such evidence, deleting it merely to raise the mean would be improper; repeat independent vessels and study whether the chemistry is unstable.

Why?

Why report results both with and without an uncertain point? The comparison shows whether a scientific conclusion depends on a discretionary choice. If the claimed catalyst advantage disappears when one point is included, the evidence is fragile and more data are needed. If the conclusion remains similar, readers can see that the inference is robust even while the unusual observation is investigated.

Common misconception

“A point beyond two standard deviations is automatically an error.” Valid observations can be rare, and the standard deviation itself may be distorted by small samples. “Never exclude data” is also too rigid when a documented failed measurement occurs. The key is a consistent rule, preserved raw record and a transparent explanation, not a blanket slogan.

Worked example

Reaction yields are 68, 70, 69 and 95%. The mean with all four is (68+70+69+95)/4 = 75.5% ; the mean of the first three is 69%. Deleting 95% changes the estimate by 6.5 percentage points. Inspection finds no mislabeled vial or calibration failure, but 95% came from a different reagent lot. The correct next step is not automatic deletion: test additional reactions from both lots and report that lot may modify the result. If the stated claim covers all available lots, the unusual value is relevant evidence of variation.

Quick check

1. Is a large residual enough by itself to justify deleting a point from a chemical calibration? Answer: No. Investigate raw data, standard preparation, instrument checks and the model; apply a prespecified failure criterion if a documented problem is found, and show sensitivity if uncertain.

Exam focus

Separate detection of an unusual observation from the decision to exclude it. Name a chemical or instrumental follow-up. Distinguish residual size from influence on fitted parameters. State objective criteria and preserve raw records. If exclusion changes the conclusion, explicitly report that sensitivity and gather new independent data rather than hiding it.

Advanced insight

Outlier procedures interact with model choice. A point can look extreme because a linear model is wrong, while a curved model may explain it; a flexible model may in turn hide genuine contamination. Influence and leave-one-out analyses reveal sensitivity but not truth. Preregistered rules and independent confirmation protect against choosing a model and exclusions together to manufacture a preferred answer.

Summary

Unusual observations are prompts for investigation. They may be errors, real heterogeneity or evidence that a model is incomplete. Predefined criteria, raw-data inspection, independent repeats and sensitivity analyses make exclusion decisions defensible. Report what was omitted and why; never quietly remove an inconvenient result.

Practice questions

1. A chromatogram has a large peak, but the instrument log shows a failed calibration check for that run. What should be recorded? Answer: Preserve the raw chromatogram, note the failed check and apply the method's predefined invalid-run rule. Recalibrate and remeasure if possible; report the exclusion and reason.

2. A point has a small residual but is the only standard at very high concentration. Why might it still need special attention? Answer: Its extreme predictor value can give it high leverage and strongly affect the fitted slope. Compare fits and validate the upper range even though its residual is small.

3. An unusual yield has no identified error and the treatment advantage vanishes when it is included. What is the honest conclusion? Answer: The evidence for the advantage is fragile. Report both analyses, retain the observation under the planned rule and collect further independent data to resolve the variation.