AreaP06
PartThe measurement engine
StageObserved in real data
Results3
●   How we do it / P06

Confidence that means what it says

Objective

State confidence that follows the actual evidence and conditions, so that over many comparisons the errors behave as the stated uncertainty says they should.

Observed in real data Current focus See the results How stages are defined

Checking stated confidence against real results as more real data comes in.

§ 01
The challenge

Uncertainty is easy to state and hard to get right. Calibration errors are often shared across pixels and over time rather than independent, reference measurements carry uncertainty of their own, and poorly constrained conditions are exactly where a fixed figure misleads most.

§ 02
Why it matters

A confidence figure is only useful if it can be relied on. Stated too tight, it invites decisions the data cannot support. Stated too loose, it throws information away.

§ 03
Our approach

We compare stated uncertainty with actual errors many times over, across clean, degraded and poorly constrained conditions, taking the uncertainty of the reference measurements into account. The centre, the spread and the extremes of the errors are assessed separately.

§ 04
What success looks like

Across conditions, observed errors match the stated uncertainty in centre, spread and extremes, and weak evidence carries correspondingly little weight.

§ 05
Results

Each result is dated and published on its own page with its data, conditions and limits. Results that went against the objective are listed here with the rest.

  1. Evidence dated 22 Sept 2026 Observed in real data

    Confidence coverage on a like-for-like surface basis

    The earlier under-coverage was primarily a reference-basis mismatch, not evidence that RefCal's surface-confidence intervals were uniformly too narrow. On the corrected basis the overall intervals are slightly conservative in this real-data set. Because the register still marks the result preliminary and several support limits remain, this set records evidence seen in real data rather than claiming validation under all stated conditions.

    Full result, data and credits ->
  2. Evidence dated 30 Aug 2026 Observed in real data Superseded

    Historical diagnosis of the first confidence gap

    This analysis is useful history but no longer establishes why current surface-confidence coverage behaves as it does. The corrected surface-to-surface proof set supersedes both this diagnosis and the first failed run.

    Full result, data and credits ->
  3. Evidence dated 7 Aug 2026 Did not hold Superseded

    First claim-grade confidence check did not meet its target

    This run failed on the basis used at the time and remains part of the dated record. It taught a more fundamental lesson than the original verdict: confidence coverage is not interpretable when the product and reference describe different measurement levels. The corrected surface-to-surface result now carries the current evidence.

    Full result, data and credits ->
§ 06
Background reading

Established work this area builds on. Results from the programme are published separately, with their data and limits.

§ Contact

Questions about this work?

If you work on this problem, or hold reference data that bears on it, contact us.

Your message is emailed to info@spectraworks.nl. See the privacy policy.