Detecting Outliers: Q and Grubbs Tests
When a suspect result may be rejected
Lesson 3464 of 4,500 · Analytical Chemistry
Learning objectives
- Calculate conceptual Q and Grubbs statistics for a single suspect extreme
- Explain why a statistical flag requires investigation and documented treatment
Introduction
One result in a replicate set may differ strongly from the others. It could reflect a spilled aliquot, a transcription mistake, an instrumental glitch, a genuinely unusual sample or ordinary chance. Outlier tests compare the gap with what a statistical model expects, but they do not identify the cause. The most important first step is examining records and sample context, not deleting the result to improve precision.
Core explanation
For a small ordered set with one suspect end value, a commonly taught Dixon Q statistic is the gap between the suspect extreme and its nearest neighbour divided by the full sample range. For sorted x₁ ≤ … ≤ xₙ and suspected high xₙ, Q = (xₙ − xₙ₋₁)/(xₙ − x₁). Compare Q with a table value appropriate to n and significance level. The exact version of Dixon's test and its critical table must match the sample size and hypothesis; it is not a general permission to remove whichever result is inconvenient.
Grubbs' test uses G = max xᵢ − x̄ /s for a single suspected outlier in approximately normal univariate data. It asks whether the most extreme standardised deviation is inconsistent with a no-outlier model. The critical value depends on n, significance level and whether the alternative is one- or two-sided.
Neither test identifies an error mechanism. A high reading from a certified standard run may indicate contamination, while an extreme result among different field samples may be real environmental heterogeneity. Review notebooks, labels, blank values, instrument logs and preparation records. If a documented procedural failure occurred, correct or repeat the affected measurement when possible. If no cause is found, a statistical rule chosen in advance may guide treatment, but report the excluded value and criterion rather than silently removing it.
Repeatedly testing and deleting extremes until a set looks tidy inflates the chance of falsely rejecting valid data. Grubbs' single-outlier test also can be masked by multiple extremes; NIST notes that tests designed for multiple potential outliers may be preferable when more than one is suspected. An outlier can signal that a normal model is wrong or that the measured population contains subgroups. Robust summaries may be more informative than forced deletion.
An outlier decision changes the mean and standard deviation used in the final result, sometimes substantially for small n. Recalculate transparently with and without the suspect point and explain the consequence. A statistical flag is evidence to investigate, not a chemical diagnosis.
Step-by-step reasoning
1. Plot or inspect data in collection order and check transcription and procedural records. 2. Decide whether one extreme result is a plausible suspect and whether observations are comparable. 3. Choose a predeclared test and check its distribution and sample-size assumptions. 4. Calculate Q or G and compare with the correct critical value. 5. Document the result, cause investigation, treatment and effect on the reported estimate.
Visual explanation
Draw four clustered points and one distant point on a concentration axis. For Q, bracket the small gap from the extreme to its neighbour and the full range. For G, draw the distance from extreme to mean and scale it by s. Below, draw a time-order chart: a jump coincident with a reagent change suggests a process shift rather than one random bad value.
Real-world analogy
An unusual exam score might be a marking error, an exceptional performance or a student who sat a different paper. A statistical rule can identify that it is unusual but cannot tell which story is true. Checking the original script and conditions is the analogue of reviewing analytical records.
Real-world example
Four replicate chloride assays cluster near 10 mg L⁻¹ and one reads 14 mg L⁻¹. The analyst discovers that the 14 mg L⁻¹ vial was mislabeled and belongs to another sample. That documented error justifies removing it from this sample set. Without such evidence, the analyst would test assumptions and report how including the observation affects the conclusion.
Why?
Why is a visual plot useful before a formal test? It can reveal trends, batches or multiple subgroups that violate the test's single stable population assumption. A single computed Q or G ignores the order and context in which values were produced.
Common misconception
“Any result failing an outlier test must be deleted” is too strong. A test's false-positive rate permits some valid extremes to be flagged, and a real unusual sample may be scientifically important. Another mistake is applying the same single-outlier test repeatedly until no unusual points remain.
Worked example
Sorted replicate values are 9.8, 9.9, 10.0, 10.1 and 11.2 mg L⁻¹. For a suspected high extreme, Q = (11.2 − 10.1)/(11.2 − 9.8) = 1.1/1.4 ≈ 0.786. This is only the test statistic; the decision requires the correct critical value for n = 5 and the preselected significance level. Before any decision, check whether the 11.2 result has a procedural explanation or represents a different sample.
Quick check
1. Does a large Grubbs G value tell you why a result is extreme? Answer: No. It measures extremeness relative to mean and sample standard deviation under a model. Laboratory records and chemical reasoning are needed to investigate the cause.
Exam focus
Show the correct numerator and range for Q and the mean-deviation-over-s form for G. State the assumption of one suspected extreme and appropriate critical tables. Treat statistical rejection as a documented decision, not an automatic editing step, and distinguish measurement mistakes from genuine heterogeneity.
Advanced insight
Selective outlier removal after seeing which decision it changes is a form of data-dependent analysis that can bias conclusions. A better method plan specifies acceptance rules and rerun criteria before sample results are known. If many samples show extremes, investigate the process or use a model allowing heavy tails rather than applying one-outlier tests indefinitely.
Summary
Q and Grubbs tests quantify how unusual one extreme result is under stated assumptions. They do not reveal its chemical cause or make deletion mandatory. Investigate records, check data structure, apply a predefined critical rule and report any exclusion and its effect on the final result.
Practice questions
1. What is the numerator for a high-end Dixon Q statistic? Answer: The gap between the largest observation and the next-largest observation in the sorted data.
2. What assumption is central to the standard Grubbs test? Answer: The observations form an approximately normal univariate data set with at most one suspected outlier for the single-outlier form.
3. Why is deleting points until RSD looks small invalid? Answer: Repeated data-dependent removal can discard valid variation and make precision appear better than it is, inflating false-positive rejection and concealing process problems.