The F-Test: Comparing Precision

Testing whether two variances differ

Lesson 3463 of 4,500 · Analytical Chemistry

Learning objectives

Introduction

Two analytical methods may give similar means but different scatter. An F test compares two variance estimates to ask whether a spread difference is larger than expected from finite sampling under a stated model. It is sensitive to departures from normality, so the calculation should follow an inspection of the data rather than replace it. Precision also depends on what was repeated: injections, preparations or field samples.

Core explanation

For two independent replicate sets, calculate sample standard deviations s₁ and s₂, then variances s₁² and s₂². A convenient upper-tail statistic is F = larger sample variance/smaller sample variance, so F ≥ 1. Its numerator degrees of freedom are n for the larger-variance group minus one; denominator degrees of freedom are n for the smaller-variance group minus one. Compare with a critical F value selected for the desired tail and significance level. A two-sided equality test requires the appropriate two-sided critical convention; simply comparing to an arbitrary table value can be wrong.

The null hypothesis is equal population variances. A large observed ratio may lead to rejection if it exceeds the correct critical value. A ratio near one is consistent with similar precision but does not prove equality; small samples may lack power to reveal a difference. The test assumes independent observations from approximately normal populations. One outlier can inflate a variance dramatically and drive rejection for the wrong reason.

Before comparing variance, ensure the measurement levels are comparable. A method with s = 1 mg L⁻¹ at 100 mg L⁻¹ and another with s = 0.5 mg L⁻¹ at 1 mg L⁻¹ have different absolute and relative precision; an F test on raw variances may not answer the practical comparison intended. Methods should generally be assessed on the same type of sample, concentration range and replicate level. A variance comparison across mixed concentrations may require a model of how variance changes with concentration.

Precision is not trueness. One method could have a smaller variance but a larger systematic bias. A laboratory deciding which method is preferable weighs detection range, accuracy, selectivity, cost and uncertainty as well as precision. Statistical rejection of equal variances says the model of identical spreads is implausible under the assumptions, not which method is chemically valid.

Step-by-step reasoning

1. Verify independent replicate design, comparable sample level and approximately normal distributions. 2. Calculate sample s for each group, then square to get variances. 3. Put the larger variance above the smaller for the selected upper-tail convention. 4. Assign numerator and denominator degrees of freedom to the correct groups. 5. Compare with the appropriate critical value and discuss practical magnitude and bias separately.

Visual explanation

Draw two horizontal dot plots with equal centres but different widths. Beneath each, show s and s²; connect the larger variance to the numerator of F. Add a second pair with the same width but shifted centres to show that an F test would not detect a mean bias difference.

Real-world analogy

Two archers may centre their arrows near the same target point, yet one groups arrows much more tightly. The F test compares group spread, not aim location. If one archer fires at a different distance or in gusty wind, comparing their spreads without accounting for conditions may be unfair.

Real-world example

A laboratory tests two sample-preparation methods using independent aliquots of the same homogeneous material. Method A has replicate s = 0.20 mg L⁻¹; method B has s = 0.40 mg L⁻¹. A appears more repeatable in absolute units. An F test can assess whether the observed fourfold variance ratio is surprising for the sample counts, while a reference material checks whether either method has biased recovery.

Why?

Why compare squared standard deviations? Variance is the quantity with a known ratio distribution under the normal, independent equal-variance model. The F distribution describes the ratio after accounting for degrees of freedom. Taking a ratio of standard deviations instead would require a different critical rule.

Common misconception

“A larger F proves method B is worse” ignores whether both methods are unbiased and suitable. Another mistake is reporting n₁ and n₂ themselves as degrees of freedom; for sample variance estimates, they are n₁ − 1 and n₂ − 1.

Worked example

Two methods each have six independent replicates. Their sample standard deviations are 0.20 and 0.40 mg L⁻¹. Variances are 0.040 and 0.160 (mg L⁻¹)². F = 0.160/0.040 = 4.00, with numerator df = 5 and denominator df = 5 under the larger-over-smaller convention. The numerical ratio alone is not a significance conclusion; the analyst must specify one- or two-sided testing and consult the matching critical F value.

Quick check

1. If two methods have identical sample standard deviations, what is the larger-over-smaller F ratio? Answer: F = 1 because their variances are equal. That observation is consistent with equal precision in the sample but does not establish exact equality of population variances.

Exam focus

Square s values before forming F, label numerator and denominator degrees of freedom, and state the test direction. Mention normality sensitivity and outliers. Keep variance comparison separate from mean comparison and method trueness.

Advanced insight

When normality is doubtful, robust variance comparisons or resampling approaches may be more appropriate than a classical F test. A method-comparison study can also model precision as a function of concentration, using replicates at several levels rather than one single F ratio. This better reflects analytical methods whose absolute error changes with signal size.

Summary

The F test compares sample variances under independent approximately normal data. Its statistic is a variance ratio with degrees of freedom from both replicate counts. It addresses precision, not bias, and its value needs a correctly chosen critical rule and a fair measurement design.

Practice questions

1. Standard deviations are 2 and 3 units. What larger-over-smaller variance ratio is F? Answer: F = 3²/2² = 9/4 = 2.25.

2. If each group has eight results, what are the F test degrees of freedom? Answer: Seven for each variance estimate, so numerator df = 7 and denominator df = 7 when both sample sizes equal eight.

3. Why inspect outliers before an F test? Answer: A single extreme value can inflate sample variance and produce a large ratio, while the standard F reference distribution assumes approximately normal observations.