Confidence Intervals and Prediction Intervals

Distinguishing parameter uncertainty from future-observation variation

Lesson 4386 of 4,500 · Research Methods, Data Analysis and Literature

Learning objectives

Introduction

A regression line predicts an average response at a concentration, but a new individual sample can deviate from that average. These are different uncertainties. A confidence interval for the mean calibration signal may be narrow after many standards; a prediction interval for one future measurement must also include ordinary sample-to-sample or instrument scatter. Confusing them can make a chemical measurement appear more certain than it is.

Core explanation

A confidence interval around a model's mean response represents uncertainty in the estimated average at a chosen predictor value, under specified assumptions. More informative data tend to narrow it. A prediction interval for a future individual response includes uncertainty in the mean and residual variation of a new observation, so it is generally wider. In a simple regression, both intervals often widen away from the center of observed predictor values because the fitted line is less constrained there. NIST's regression interval guidance distinguishes estimation of average response from prediction of individual outcomes.

Coverage language requires care. A conventional 95% confidence-interval procedure is constructed so that, under its model and repeated sampling process, about 95% of such intervals cover the fixed true target. Once one interval is computed, the parameter is not randomly moving within it under the classical interpretation. A 95% prediction interval describes a new observation under the fitted data-generating process and its assumptions. Neither interval automatically incorporates every systematic uncertainty, such as a wrong calibration stock or unmodeled matrix effect.

In chemical calibration, the direction can be reversed: an unknown signal is used to estimate concentration. A mean-response confidence band is not by itself the correct concentration interval. Uncertainty in the measured unknown signal, regression coefficients, blank and sample dilution must be included. Measurements near the ends of the calibration range are often less certain, and an unknown beyond that range is extrapolation. For very small or near-zero signals, a symmetric interval may be inappropriate; detection-limit conventions and constrained models may matter.

Replicates change the target. The mean of ten independent future assays varies less than one future assay, so a predictive interval for their average can be narrower than one for a single result. Ten injections of the same sample do not eliminate sample-preparation uncertainty, however. State whether the interval concerns an instrument reading, a new independently prepared sample, a batch mean or a population of field samples. The source of variation included depends on the level of independent replication.

Model validation remains essential. If residuals are curved, heteroscedastic or time-correlated, textbook intervals from ordinary linear regression may miss their advertised coverage. A narrow interval from a wrong model is not reassuring. Check residual patterns, calibration controls and relevant independent samples before treating interval endpoints as decision boundaries. For a safety threshold, consider whether the interval covers the uncertainty in the complete measurement process.

Step-by-step reasoning

Define the target: parameter, mean response, new measurement or new sample. Fit and check an appropriate model. Identify which uncertainty components and replication levels are included. Choose a confidence or prediction procedure matching the target and report its stated level. Examine whether the target predictor lies in the observed range. Interpret the interval conditionally on model and sampling assumptions, and mention important uncertainty sources outside it.

Visual explanation

Plot a fitted calibration line with two shaded bands. The narrow inner band is a confidence band for mean response; the wider outer band is a prediction band for a future individual response. Both widen at the edges. Place one new measured point inside the outer band but outside the inner band to show that a plausible individual result need not sit within the mean band. Label what each band covers.

Real-world analogy

A teacher may know the average height of students in a class quite precisely, but that does not mean the next student will have a height close to the average. Uncertainty in the class mean is different from natural variation among individuals. Chemistry has an analogous difference between the average instrument response at a concentration and one new sample's response.

Real-world example

A laboratory fits absorbance versus concentration with many standards. The estimated mean absorbance at 2 mg/L is 0.200 with a narrow 95% mean interval. A new 2 mg/L sample produces 0.215, outside that narrow band but inside the wider single-reading prediction interval. That does not automatically indicate contamination or method failure; individual readings fluctuate. A sequence of new samples consistently above the mean, however, would suggest bias, drift or matrix interference.

Why?

Why is a prediction interval wider? It must account for uncertainty in where the true mean curve lies and for the fresh random deviation of the next observation from that curve. A confidence interval for the mean includes the first component but not the full second component. As sample size grows, uncertainty in the mean can shrink substantially, while irreducible single-measurement scatter remains.

Common misconception

“A 95% confidence interval means there is a 95% chance the fixed true value lies in this particular computed interval.” That is not the usual frequentist interpretation. “A future reading must fall inside the confidence band for the mean” confuses targets. “A narrow statistical interval includes all laboratory error” ignores unmodeled bias and matrix effects.

Worked example

Suppose a fitted method predicts an average response of 10.0 units at a given concentration. Its estimated uncertainty in that mean is 0.2 units, while typical individual residual standard deviation is 1.0 unit. Using a rough factor of two only for illustration, a mean-response interval might have half-width about 2 × 0.2 = 0.4 units. An individual prediction must combine both scales: approximate standard uncertainty sqrt(0.2²+1.0²) ≈ 1.02 , giving a rough half-width about 2.0 units. Exact intervals use appropriate degrees of freedom and model formulas, but the example shows why a new reading of 11.0 can be ordinary even though it is far outside 9.6–10.4.

Quick check

1. A lab wants a range for one new independently prepared sample. Should it use a mean-response confidence interval or a prediction interval? Answer: A prediction interval matched to an independently prepared sample, because the new sample adds preparation and measurement variation beyond uncertainty in the fitted mean.

Exam focus

Name the target and the sources of variation included. State why prediction intervals are generally wider than confidence intervals for a mean. Interpret 95% as long-run coverage under model assumptions, not a guarantee about one interval. Check whether a result is interpolation and whether residual assumptions are plausible. Do not use a fitted-mean band as an uncertainty interval for a new unknown concentration without inverse-calibration analysis.

Advanced insight

Calibration uncertainty is often asymmetric after inverting the response equation, especially near a shallow slope or a detection boundary. If the denominator slope is uncertain, simple symmetric formulas can fail. Prediction for a new laboratory or batch may also require additional random-effect variance absent from a within-day calibration. Interval quality depends more on the measurement model and design than on the label printed by software.

Summary

Confidence intervals describe uncertainty in a parameter or mean relationship; prediction intervals include variation in a future observation and are typically wider. Both depend on model assumptions, replication level and observed range. State the target, coverage and omitted uncertainties. A well-calibrated interval is a tool for reasoning, not a substitute for checking chemistry and measurement quality.

Practice questions

1. Why can an individual future response fall outside a narrow confidence interval for the mean without contradicting the regression model? Answer: The mean interval does not include the full residual scatter of individual observations. A prediction interval is the appropriate comparison for one future response.

2. A fitted curve's interval widens near the highest standard. Give one reason. Answer: Fewer observations constrain the fitted mean near the edge, so uncertainty in coefficients translates into greater response uncertainty there. Prediction intervals include this plus residual variation.

3. Does a 95% prediction interval guarantee that exactly 95 of the next 100 chemical readings will fall inside it? Answer: No. Coverage is a long-run property under the model; finite sequences vary, and drift or model failure can change the actual coverage.