Sampling, Equilibration and Autocorrelation
Why trajectory length does not equal independent-sample count
Lesson 4152 of 4,500 · Computational Chemistry
Learning objectives
- Distinguish equilibration from production sampling
- Explain why nearby frames are correlated
- Estimate effective sample size and recognize unvisited metastable states
Introduction
A molecular-dynamics file may contain millions of frames, but its statistical information can be far smaller. A water molecule's orientation at one picosecond is related to its orientation moments earlier, and a protein loop may remain in one conformation for microseconds. The first part of a run may still be relaxing from an artificial initial condition. Interpreting the trajectory requires asking both whether the sampled region represents the target equilibrium and how many effectively independent observations it contains. File size and simulated time alone cannot answer either question.
Core explanation
Equilibration addresses initial-condition bias. A constructed structure may have strained contacts, incorrect solvent density or an unusual conformation. Energy minimization removes severe overlaps, but a minimized state is not a thermal ensemble. During early MD, temperature, volume, hydration and slow structural variables may relax. Analysts commonly discard a burn-in interval before estimating equilibrium observables. The cutoff should be justified by time-series evidence and sensitivity analysis rather than chosen only as a fixed percentage of every run. Primary research on automated equilibration detection uses the tradeoff between retaining data and reducing correlation to select an initial discard interval.
Even after equilibration, consecutive frames are not independent. For an observable A(t), its normalized autocorrelation C(τ) measures how strongly A at time t predicts A at t + τ. A rapidly decorrelating bond vibration and a slowly switching protein conformation can have very different correlation times in the same trajectory. A common statistical inefficiency g expresses how many recorded correlated observations amount to one independent observation, giving an approximate effective sample size Neff ≈ N/g. The precise formula and factors depend on sampling interval and autocorrelation convention, so report the method. The key point is that writing frames more frequently increases N without necessarily increasing Neff. Original sampling-assessment research developed effective-sample measures for molecular ensembles.
Standard errors estimated as sample standard deviation divided by √N are too optimistic when N counts correlated frames. Block averaging, autocorrelation-based estimators or multiple independent trajectories can better quantify uncertainty, provided blocks exceed relevant correlation times. Independent starting seeds are valuable because one trajectory may be stuck in a metastable basin. But several runs initialized from the same structure can all remain trapped, so their apparent agreement is not definitive evidence of global equilibrium. Best-practice work by molecular-simulation researchers emphasizes checks for insufficient sampling and uncertainty estimation.
Different observables equilibrate at different rates. Temperature may stabilize quickly while a conformational population changes slowly. A flat potential-energy trace does not prove that the slow coordinate has crossed between relevant states. For a two-state system, many transitions in both directions and stable relative populations across independent runs provide stronger evidence. If no transitions occur, one can report a conditional within-basin average, but not a well-converged global equilibrium fraction. Enhanced sampling may be needed to overcome barriers, yet its reweighting and overlap must also be checked.
Time-series diagnostics include plotting the observable, comparing early and late block averages, examining running means and estimating autocorrelation. Sensitivity to the chosen equilibration cutoff is a warning. A very long single run can reduce random error in one basin without resolving systematic undersampling of other basins. Conversely, a shorter run that visits all relevant states several times can give better equilibrium information for some observables. The sampling question must be tied to the physical property—energy, density, diffusion, binding pose or conformational equilibrium—rather than assessed by one universal scalar.
Step-by-step reasoning
1. Select observables that represent both fast environmental relaxation and the slow process central to the scientific question. 2. Plot raw and running-average traces, checking for initial drift or state changes. 3. Choose and document an equilibration interval, then test sensitivity to reasonable alternatives. 4. Estimate autocorrelation or block-size behavior for each important production observable. 5. Convert frame count to an effective independent-sample estimate and uncertainty. 6. Compare independent trajectories and assess whether metastable states have been visited adequately.
Visual explanation
Draw two time traces of the same length. One fluctuates rapidly around a stable mean; the other remains low for most of the run before one rare jump to a high state. Mark the first trace's early drift as equilibration, then show broad blocks used for independent estimates. In the second trace, counting every saved frame as independent would be absurd because only one state transition occurred. A separate autocorrelation curve decays quickly for the first observable and slowly for the second.
Real-world analogy
Taking one hundred photographs of the same parked car from one fixed camera every second does not provide one hundred independent examples of traffic patterns. The photos may be numerous but highly related. MD frames similarly inherit memory of neighboring frames. To learn traffic behavior, one needs observations over varied times and states; to learn molecular equilibrium, one needs state exploration and statistically independent information.
Real-world example
A team simulates a ligand bound to a protein for 200 ns and saves a frame every picosecond, producing 200,000 frames. The ligand makes one major pose change at 180 ns. An average over all frames appears precise if treated as 200,000 independent samples, but the pose population is supported by only one observed transition and is not converged. The team starts independent runs from both poses and compares transitions and free-energy estimates. Meanwhile, water-density averages may have many more effective samples in the same trajectory because their correlation time is shorter.
Why?
Why does saving frames ten times more frequently usually fail to reduce statistical uncertainty tenfold? The physical trajectory is unchanged. Extra frames between old save points often record nearly the same molecular state, so autocorrelation rises in the stored series and statistical inefficiency grows. The independent information is governed mainly by simulated physical time relative to relaxation times, not the storage interval. More frequent output can improve time-resolution analysis but does not create new equilibrium transitions.
Common misconception
“A flat temperature plot proves the protein is equilibrated.” Temperature can stabilize far faster than conformational populations. “Every trajectory frame is a new independent observation.” Adjacent frames are correlated. “Discarding the first 10% always fixes equilibration.” The necessary interval is property and system dependent. “Many independent seeds prove convergence if they all start in one trapped basin.” They may all miss the same alternative state.
Worked example
Suppose a production trajectory contains N = 100,000 saved values of an observable. An autocorrelation analysis gives statistical inefficiency g ≈ 50, so Neff ≈ 100,000/50 = 2,000. If the standard deviation across frames is 10 units, a naive independent-frame standard error would be 10/√100,000 ≈ 0.032, whereas a rough correlation-adjusted estimate is 10/√2,000 ≈ 0.224 units. This sevenfold difference matters when reporting precision. The estimate still assumes the production segment samples the target stationary distribution; if an unvisited basin exists, neither calculation captures that systematic uncertainty.
Quick check
1. If adjacent frames are strongly correlated, is Neff close to the raw frame count? Answer: No. Effective independent information can be much smaller. 2. Does a stable mean potential energy prove all slow conformations have equilibrated? Answer: No. Distinct conformational basins can have similar energy averages and rare transitions.
Exam focus
Explain equilibration, production and autocorrelation in different sentences. Compute Neff from a supplied N and g under a stated convention. Compare naive and correlation-adjusted standard errors. For a two-state trajectory with one transition, identify why a population estimate is weak despite many frames. Propose independent starting conditions or enhanced sampling when slow states are unvisited.
Advanced insight
An autocorrelation estimate for one scalar observable can miss slow modes to which that observable is insensitive. Structural-state analysis and multiple coordinates can reveal metastability. Effective sample size may also be hard to estimate when the trajectory is shorter than the true correlation time; an apparently finite g then reflects truncation, not reliable convergence. Enhanced-sampling data require proper reweighting to the target ensemble, and uncertainty depends on overlap between sampled windows or states. Report both statistical error within the sampled model and unresolved sampling limitations.
Summary
MD trajectory length and frame count do not equal equilibrium information. Initial relaxation should be separated from production, and correlated frames should be converted into an effective independent-sample count before estimating precision. Different observables have different timescales, while unvisited metastable states create systematic errors that ordinary standard errors cannot capture. Time traces, autocorrelation, block analysis and independent state exploration make simulation conclusions more defensible.
Practice questions
1. A run saves 50,000 frames and has statistical inefficiency g = 25. What is approximate Neff? Answer: 2,000 under the stated convention. 2. Why might a fixed 5-ns equilibration discard be insufficient for one protein but excessive for a simple liquid? Answer: Their slow relaxation times differ, so equilibration is system and observable dependent. 3. Can a narrow standard error from one trapped trajectory justify a global equilibrium population? Answer: No. It describes within-basin variation and misses unsampled states. 4. What does block averaging attempt to reduce in an uncertainty estimate? Answer: The effect of time correlation by combining observations into blocks longer than relevant correlation times.