Proof of the Variation Theorem
Expanding a trial function in exact eigenfunctions
Lesson 2958 of 4,500 · Quantum Chemistry I
Learning objectives
- Expand an arbitrary trial function in the complete set of eigenfunctions of the Hamiltonian
- Use orthonormality to express the trial energy as a weighted average of eigenvalues
- Prove that the trial energy cannot lie below the ground-state energy
- Show that the energy error is second order in the admixture of excited states
Introduction
The variation principle is used so widely that it is worth seeing exactly why it is true. The proof is short and elegant. It relies on three facts established earlier in the course: the eigenfunctions of a Hermitian operator such as the Hamiltonian form a complete set, they can be chosen orthonormal, and their eigenvalues are real. With these tools any trial function becomes a mixture of exact states, and its energy becomes an average that can never dip below the lowest eigenvalue.
Core explanation
Setting up. Suppose the exact eigenfunctions and eigenvalues of Ĥ are ψ₀, ψ₁, ψ₂, … and E₀ ≤ E₁ ≤ E₂ ≤ …, so that
Ĥψₙ = Eₙψₙ and ∫ψₘ ψₙ dτ = δₘₙ,
where δₘₙ is 1 if m = n and 0 otherwise. We do not need to know these functions explicitly; we only need their existence and properties.
Expanding the trial function. Because the eigenfunctions form a complete set, any acceptable trial function ψ can be written as
ψ = Σₙ cₙψₙ,
with coefficients cₙ = ∫ψₙ ψ dτ. This is analogous to writing any vector as a combination of perpendicular unit vectors.
The denominator. Using orthonormality,
∫ψ ψ dτ = Σₘ Σₙ cₘ cₙ ∫ψₘ ψₙ dτ = Σₙ cₙ ².
All cross terms vanish because different eigenfunctions are orthogonal.
The numerator. Acting with Ĥ on ψ gives Ĥψ = Σₙ cₙEₙψₙ. Then
∫ψ Ĥψ dτ = Σₘ Σₙ cₘ cₙEₙ ∫ψₘ ψₙ dτ = Σₙ cₙ ²Eₙ.
The key step. The trial energy is therefore
E trial = Σₙ cₙ ²Eₙ / Σₙ cₙ ².
This is a weighted average of the eigenvalues, with non-negative weights cₙ ². Now subtract E₀:
E trial − E₀ = Σₙ cₙ ²(Eₙ − E₀) / Σₙ cₙ ².
Every cₙ ² is non-negative and every Eₙ − E₀ is non-negative because E₀ is the lowest eigenvalue. The numerator is a sum of non-negative terms and the denominator is positive, so
E trial − E₀ ≥ 0, that is, E trial ≥ E₀.
The equality case. The difference is zero only if every term with Eₙ > E₀ has cₙ = 0. For a non-degenerate ground state, this means ψ = c₀ψ₀: the trial function is the exact ground state.
Step-by-step reasoning
The proof in five moves:
1. Expand ψ in the exact eigenfunctions of Ĥ. 2. Use orthonormality to reduce the norm to Σ cₙ ². 3. Use Ĥψₙ = Eₙψₙ to reduce ⟨Ĥ⟩ to Σ cₙ ²Eₙ. 4. Write E trial − E₀ as a ratio of non-negative sums. 5. Conclude E trial ≥ E₀, with equality only when ψ ∝ ψ₀.
Visual explanation
Draw the energy levels as rungs on a ladder. Hang weights cₙ ² on each rung. The trial energy is the height of the balance point of all the weights. However the weights are distributed, the balance point cannot lie below the bottom rung.
Real-world analogy
A class's average exam mark cannot be lower than the lowest individual mark. If every student scored at least 40, the average is at least 40, and it equals 40 only if every student scored exactly 40. The trial energy is an average in just the same sense.
Real-world example
The same argument underpins the convergence checks in computational chemistry. When a basis set is enlarged, the space of possible trial functions grows to include the old one, so the minimum energy can only stay the same or fall. Software packages rely on this to guarantee that larger calculations never give variationally worse energies.
Why?
Why does the proof need a Hermitian Hamiltonian? Hermiticity guarantees that the eigenvalues are real, so they can be ordered from lowest to highest, and that the eigenfunctions are orthogonal, which removes all cross terms. Without these properties the weighted-average picture would collapse.
Common misconception
"The proof requires knowing the exact eigenfunctions." It uses only their existence and general properties. We never evaluate any ψₙ; the argument works precisely because it holds for a system we cannot solve.
Worked example
Question: A normalised trial function for some system is ψ = 0.98ψ₀ + 0.20ψ₁, where E₀ = −10.0 eV and E₁ = −4.0 eV. Find E trial and its error.
Reasoning: The weights are c₀ ² = 0.9604 and c₁ ² = 0.0400, summing to 1.0004 (normalised to rounding). E trial = (0.9604 × −10.0 + 0.0400 × −4.0)/1.0004 = (−9.604 − 0.160)/1.0004 = −9.760 eV.
Answer: E trial ≈ −9.76 eV, which is 0.24 eV above E₀. The error equals c₁ ²(E₁ − E₀) = 0.04 × 6.0 eV, showing it depends on the square of the admixture coefficient.
Quick check
1. In the expansion ψ = Σcₙψₙ, why does the denominator ∫ψ ψ dτ reduce to Σ cₙ ²? Answer: Because the eigenfunctions are orthonormal, so every cross term with m ≠ n integrates to zero.
Exam focus
You may be asked to reproduce the proof. Write the expansion, the norm and the energy expectation as sums, then subtract E₀ and argue from signs. State clearly where completeness, orthonormality and the ordering of eigenvalues are used.
Advanced insight
The worked example illustrates a general result: if ψ = ψ₀ + εφ, with φ orthogonal to ψ₀, the energy error is of order ε². A related theorem due to Hylleraas, Undheim and MacDonald shows that in a linear variation calculation the k-th lowest root is an upper bound to the k-th exact energy, extending the principle to excited states.
Summary
Expanding a trial function as ψ = Σcₙψₙ in the exact, orthonormal eigenfunctions of Ĥ turns its energy into a weighted average Σ cₙ ²Eₙ / Σ cₙ ². Subtracting E₀ gives a sum of non-negative terms, proving E trial ≥ E₀, with equality only for the exact ground state. The error depends on the squares of the excited-state coefficients.
Practice questions
1. Which property of the Hamiltonian's eigenfunctions allows any trial function to be expanded in them? Answer: Completeness: the eigenfunctions of a Hermitian operator form a complete set. 2. Show that E trial − E₀ = Σ cₙ ²(Eₙ − E₀)/Σ cₙ ². Answer: Subtract E₀ = E₀Σ cₙ ²/Σ cₙ ² from Σ cₙ ²Eₙ/Σ cₙ ² and combine the two sums over the same denominator. 3. A normalised trial function has c₀ ² = 0.90 and c₂ ² = 0.10, with E₀ = 1.0 and E₂ = 9.0 in some unit. Find E trial. Answer: E trial = 0.90 × 1.0 + 0.10 × 9.0 = 1.8 in that unit. 4. Why is the ordering E₀ ≤ E₁ ≤ E₂ essential to the proof? Answer: It guarantees each term Eₙ − E₀ is non-negative, so the weighted sum cannot be negative.