Calibration Curves and Linear Regression
Least squares, slope, intercept and correlation
Lesson 3465 of 4,500 · Analytical Chemistry
Learning objectives
- Use a linear calibration model to estimate unknown concentration
- Interpret slope, intercept and residuals without treating correlation as validation
Introduction
Many analytical instruments produce a signal, not concentration directly. Calibration standards establish a relation between known analyte levels and measured response. A common model is a straight line, but fitting one is only the start: blanks, residuals, measurement range and matrix effects determine whether its inverse gives a trustworthy unknown concentration.
Core explanation
Write the model y = a + bx, where x is analyte concentration and y is measured signal. The slope b expresses sensitivity in signal units per concentration unit. The intercept a is the predicted response at zero analyte, which may include background or instrument offset. An unknown with response yu gives xu = (yu − a)/b if the line is valid in that range. A negative intercept does not automatically prove negative physical signal; it may reflect fitting noise or an unsuitable model.
Ordinary least squares chooses a and b to minimise Σ(yi − a − bxi)², the sum of squared vertical response residuals. This fitting assumes standard concentrations are known much more accurately than their response measurements and that response errors are suitably independent with a model of variance. When response variance grows with concentration, weighted regression may be better. Do not force the line through zero solely because the ideal chemistry seems proportional; evaluate blank behaviour and residuals first.
Use several standards spanning the unknowns, with independent preparation when possible. A high correlation coefficient or R² can occur even when the calibration bends systematically, especially over a wide concentration range. Plot residuals against concentration: random scatter around zero supports the straight-line model, while curvature, a fan shape or time drift suggests problems. A calibration check standard measured during the sample sequence detects changing response after the initial fit.
Calibration uncertainty comes from standard preparation, scatter around the fit, slope and intercept estimates, blank correction and the unknown's measured signal. One unknown result inverted from a line is not equivalent to an exact algebra exercise. Matrix effects can cause standards in pure solvent to respond differently from real samples. Matrix-matched calibration, standard addition or an internal standard may be appropriate, but each needs validation.
The intended reporting interval matters. A method can be linear over one range and saturated above it. Diluting a high-response unknown into the calibrated interval is preferable to extrapolating far outside the standards, provided dilution is quantitative and recorded. An unknown below the lowest reliable standard may be detected but not accurately quantified.
Step-by-step reasoning
1. Prepare standards and blanks spanning the relevant unknown range. 2. Measure responses with controls interleaved where drift is possible. 3. Fit a stated model and inspect residuals, not just R². 4. Invert the validated equation for unknowns inside the range. 5. Apply dilution factors and report uncertainty and matrix limitations.
Visual explanation
Draw response versus concentration with five standards, a fitted line and an unknown horizontal response intersecting it. Beneath, plot residuals around zero. A second residual sketch curves upward, showing that a visually impressive straight-line fit can still be systematically wrong.
Real-world analogy
Calibrating an instrument resembles marking a ruler by measuring known lengths. A straight ruler line is useful only if the markings behave consistently across the range. A ruler that bends slightly may still give a high correlation with length yet misread objects in the middle; residual checks reveal the bend.
Real-world example
A lab measures absorbance for nitrate standards from 0 to 10 mg L⁻¹, fits a line and tests a mid-range check solution. An unknown gives a response above the highest standard, so the lab dilutes it fivefold and repeats the measurement. It multiplies the calibrated diluted concentration by five, while adding pipette and flask contributions to uncertainty.
Why?
Why examine residuals after least-squares fitting? The fitting algorithm always returns some best line, even if the true response curves or drifts. Residual patterns show where the model consistently over- or underpredicts, providing evidence that inverse concentrations would be biased.
Common misconception
“R² close to one proves accurate calibration” is false. R² measures how much observed response variation the fitted line explains, not whether standards are correct, residuals random or unknown matrix comparable. Another error is treating the slope as a universal detector property independent of wavelength, flow or day.
Worked example
A validated calibration is y = 0.020 + 0.100x, where y is absorbance and x is mg L⁻¹. An unknown gives y = 0.370. Then x = (0.370 − 0.020)/0.100 = 3.50 mg L⁻¹ in the measured solution. If the original sample was diluted 1:4, meaning one volume brought to four final volumes, original concentration is 14.0 mg L⁻¹. The result assumes 3.50 lies within the validated standard range.
Quick check
1. Why should an unknown above the highest reliable calibrant usually be diluted and remeasured? Answer: Extrapolation may leave the validated linear range and produce a biased estimate. A documented dilution can bring the response within standards, then be reversed in calculation.
Exam focus
Identify axes and units, interpret slope and intercept, and invert y = a + bx correctly. Distinguish one-to-four final dilution from adding four volumes of solvent. State why residuals and a check standard matter and why high R² alone is insufficient.
Advanced insight
Inverse calibration uncertainty can become large when slope is small or response noise is high. Near a detection limit, the concentration interval can be asymmetric or even include values below zero under a naïve model. Reporting rules should follow a validated detection and quantification framework rather than forcing a confident positive concentration from every numerical signal.
Summary
Linear calibration maps known concentration to response through slope and intercept. Least squares provides an estimate, but residuals, standards, blanks and matrix checks test whether it is valid. Unknowns should be measured within the calibrated range, converted back through dilution and reported with appropriate uncertainty.
Practice questions
1. With y = 5 + 20x and yunknown = 65, what is x? Answer: x = (65 − 5)/20 = 3 in the concentration units used for the standards.
2. What does a curved residual plot suggest? Answer: The straight-line model systematically misses the response pattern; a different range, transformation or model may be needed before unknowns are quantified.
3. Can a high R² compensate for contaminated calibration standards? Answer: No. Contaminated standards can lie neatly on a line while assigning biased concentration values to unknowns.