VectorCertain
n_eff 1.007 → 13.000 (THEORETICAL MAX)
PAIRWISE CORRELATION 0.993 → 0.000
DIVERGENCE AMPLIFIED 8.5×
Research · Engineering study · Companion to the Convergence Trap

Variance-Weighted Decorrelation: n_eff 1.007 → 13.000

Measuring correlated failure is the easy half. This is the half where we did something about it — including discarding our own prior calibration work.

154 OF 256 DIMENSIONS ZERO-VARIANCE ON LIVE DATA
CALIBRATION DOCTRINE: LIVE DATA ONLY · ZERO-VARIANCE DIMS FORBIDDEN

The state of the evidence, in one paragraph: effective sample size (n_eff) measures how many genuinely independent decision sources an ensemble provides — for 13 models, the theoretical maximum is 13.000. Live measurement of VectorCertain's own governance masks returned 1.007: thirteen models delivering the independence of one [VC-B]. Root-cause analysis traced the collapse to a calibration artifact — mask weights derived from synthetic profiles — and the redesign derived weights from the per-dimension variance of live model behavior instead. Result, on live data: n_eff = 13.000, mean pairwise correlation reduced from 0.993 to 0.000, and genuine model divergence amplified 8.5× [VC-B]. This page documents the failure, the redesign, the two engineering rules promoted to permanent doctrine, and the reasons a result at the theoretical maximum should itself be treated with suspicion.

Request the methodology walkthrough

§ 1 · THE FAILUREThe Calibration Artifact: 154 Dead Dimensions

The original governance mask weights were derived from synthetic behavioral profiles produced by a mock client — standard practice when live traffic is not yet available, and exactly the practice this study now forbids. Run against live model traffic, 154 of 256 behavioral dimensions showed zero variance across all 13 models: every model behaved identically on them [VC-B].

A dimension with zero variance carries no independence signal — but the synthetic masks were weighting those dead dimensions anyway, which amplified measurement noise instead of behavioral signal and drove mean pairwise correlation to 0.993, near-total redundancy. The instrument built to measure model independence was, through its own calibration, erasing it — and nothing in its dashboards looked wrong, because a confidently miscalibrated instrument produces clean, consistent, plausible numbers. Finding this required discarding prior work rather than defending it — the study is documented as a correction, not an achievement, in the sealed validation record [VC-B].

Before and after: dimension weight distribution and the resulting effective sample size bars MASK WEIGHT ACROSS 256 DIMENSIONS400 synthetic: dead dims weighted variance-weighted: zero weight on dead dims 410 EFFECTIVE SOURCES (n_eff)420 theoretical max = 13 1.007SYNTHETIC 13.000VARIANCE-WEIGHTED
FIG. 1The 256-dimension mask (400): under synthetic calibration, dead dimensions (striped) carry weight; under variance weighting (410), weight concentrates on live-divergent dimensions and n_eff rises to the 13.000 ceiling (420).
Table 1 · Before and after the redesign (live data, 13 models)
MetricSynthetic-derived masksVariance-weighted masks
Effective independent sources (n_eff)1.00713.000 — theoretical max
Mean pairwise correlation0.9930.000 (100% reduction)
Zero-variance dimensions weighted154 of 2560 — forbidden by doctrine
Genuine divergence amplification8.5×
Calibration ground truthsynthetic mock-client profileslive model traffic only

§ 2 · THE REDESIGNVariance-Weighted Masking on Live Data

The redesign replaces synthetic-derived weights with weights derived from the per-dimension variance of live model projections: measure where the 13 models actually diverge, weight each dimension in proportion to that measured divergence, and assign exactly zero weight to zero-variance dimensions. No dimension earns weight by assumption; every weight traces to observed behavior [VC-B].

Operationally, live calibration means the weights are an output of the deployment, not an input to it. The 13 models run against real traffic; per-dimension variance is computed from their actual projections; weights follow. Nothing in the loop requires a synthetic profile, a mock client, or an engineer's estimate of where models "should" differ — the three places the original artifact entered. Recalibration against fresh live traffic is repeatable on demand, which is what lets the correction survive model updates that shift where genuine divergence lives [VC-B].

Applied to live traffic, the redesign moved n_eff from 1.007 to 13.000 — the theoretical maximum for a 13-model ensemble — reduced mean pairwise correlation from 0.993 to 0.000 (a 100% Pearson reduction), and amplified genuine model divergence 8.5×. The mechanism matters as much as the numbers: this is not a new model, a new prompt, or a vendor change. It is a measurement-layer correction, which means it composes with any underlying model set — the property that keeps the HCF2-SG framework architecture-agnostic [VC-C].

§ 3 · THE ARITHMETICWhy Correlation Destroys Effective Sample Size

The collapse from 13 models to 1.007 effective sources is not a mystery once the arithmetic is visible. Under the standard equicorrelation design-effect model, an ensemble of N sources with mean pairwise correlation ρ provides an effective sample size of n_eff = N / (1 + (N − 1)ρ). The formula's behavior at the extremes is the whole story: at ρ = 0 the denominator is 1 and all N sources count; as ρ → 1 the denominator approaches N and the ensemble collapses toward a single effective source, no matter how many models it contains [VC-B].

The measured values sit exactly where that model predicts. With N = 13 and the synthetic-mask correlation of ρ = 0.993, the formula gives 13 / (1 + 12 × 0.993) ≈ 1.007 — the measured figure. With the variance-weighted masks driving measured correlation to 0.000, it gives 13 / 1 = 13.000 — again the measured figure [VC-B]. The agreement is worth stating because it means the headline numbers are not artifacts of an exotic metric: they follow from textbook statistics applied to measured correlations, and anyone can re-derive them from the two inputs.

The arithmetic also explains why the problem hides. Adding a fourteenth model to a ρ = 0.993 ensemble moves n_eff by less than a hundredth — scale is powerless against correlation. Every unit of engineering effort spent reducing ρ buys more effective independence than any unit spent adding models, which is why the redesign targeted the measurement layer and left the model set untouched [VC-B].

The same arithmetic carries a warning for vendor-diversification strategies. "We use three different providers" is an independence claim stated as a procurement fact — but the formula consumes measured correlation, not vendor count, and the Convergence Trap measurement found frontier models correlating at 0.95 across vendor lines. Diversification that is not accompanied by a correlation measurement is a redundancy budget spent on an unverified assumption; the constructive version measures first, then diversifies along the dimensions where divergence actually exists [VC-A][VC-B].

§ 4 · THE DOCTRINETwo Rules Promoted to Permanent Doctrine

Two engineering rules were promoted from this study into permanent doctrine for every production governance parameter [VC-B]. First: zero-variance dimensions are forbidden in mask configurations. A dimension on which all models agree is not evidence of anything except itself; weighting it converts noise into apparent signal. Second: live data — never synthetic fixtures — is the calibration ground truth. Synthetic profiles answer the question "what did we imagine models would do?"; only live traffic answers "what do they do?" The gap between those two questions was, in this study, the entire difference between n_eff = 1.007 and n_eff = 13.000. "Forbidden" is meant structurally, not aspirationally: a mask configuration carrying a zero-variance dimension is an invalid configuration under the doctrine, in the same category as a failing test — not a style preference an engineer may weigh against a deadline. The doctrine exists precisely because the original artifact was not a careless mistake; it was reasonable practice that produced a wrong instrument, and only a rule with no discretion in it prevents reasonable practice from doing so again [VC-B]. Both rules align with the measurement discipline the NIST AI RMF's MEASURE function asks of AI risk quantification [2][3]: measured properties, on real behavior, re-verifiable on demand.

The transferable lesson is larger than masks. Any governance instrument calibrated on fixtures — synthetic profiles, mocked clients, imagined distributions — should be presumed miscalibrated until live behavior says otherwise, because the failure mode demonstrated here produced maximum-confidence wrong numbers, not obviously broken ones. The instrument reported an ensemble; it was measuring an echo [VC-B].

§ 5 · LIMITATIONSWhy a Perfect Score Warrants Suspicion

The intellectual-honesty note carried in the study's own handoff, reproduced in substance here [VC-B]: a result at the theoretical maximum is exactly the kind of result that demands a re-run. n_eff = 13.000 is consistent with complete decorrelation; it is also consistent with a new artifact that amplifies noise into apparent divergence — the mirror image of the failure just corrected. The follow-on sprint was scoped specifically to re-measure and determine whether 13.000 is stable signal. Both outcomes are informative, and it is worth stating in advance what each would mean: a stable 13.000 across re-runs and traffic windows would confirm that the dead-dimension artifact was the whole story; a regression toward some intermediate value would indicate the variance-weighted instrument has its own bias, and would hand the program a second correction to make — publicly, the same way this one was made. Pre-registering both readings here is deliberate: it removes the room to reinterpret whichever number arrives. Further boundaries: the result is measured on one live environment and one 13-model set; generalization to other model cohorts is asserted by mechanism, not yet by measurement. And the 8.5× divergence amplification is a property of the measurement layer, not a claim that the underlying models became more diverse. One further boundary: the equicorrelation formula assumes a uniform pairwise correlation, while real model pairs vary around the mean — the formula is used here as the interpretive frame the measured means slot into, not as a substitute for the full 78-pair matrix, which the Convergence Trap page documents. Sealed artifacts [VC-B] are hash-verified and reviewable under NDA; public URLs will replace artifact names at publication.

CONTACTTalk to VectorCertain

Every figure on this page traces to sealed, hash-verified validation artifacts — including findings that contradicted our own published estimates, which we recorded rather than reconciled away. Technical briefings are available for enterprises, evaluators, and standards bodies.

Request the technical briefing

REFERENCES

  1. Lakshminarayanan, B., Pritzel, A., & Blundell, C. (2017). Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles. NeurIPS 2017. arxiv.org/abs/1612.01474
  2. NIST. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1. nvlpubs.nist.gov
  3. NIST. AI Risk Management Framework (program page). nist.gov/itl/ai-risk-management-framework
First-party validation artifacts (VectorCertain, sealed and hash-verified):

[VC-B] Variance-weighted decorrelation study — n_eff 1.007 → 13.000, live-data calibration doctrine; Sprint 39 validation record. [VC-C] Patent portfolio and platform engineering baseline — 77-claim hub filing, stack integration claims, 36,181-test regression suite; portfolio documentation, January 2026. Public URLs will replace artifact names when the corresponding research pages publish.

Join the waitlist

Our signup form is temporarily offline while we perform maintenance. Nothing is lost — reach us directly and we'll add you by hand.

PLACEHOLDER · Tally.so form returns here · engineering ticket open

Email us to join