What we did
After the first claim-grade confidence run missed its target, RefCal assessed whether the apparent shortfall had a stable, repeatable character and carried into unseen dates. At the time, the analysis treated the available RadCalNet values as surface reference measurements.
What we saw
On that basis, most of the gap appeared stable rather than random. On withheld dates, factoring out the learned difference increased coverage from 0.49 to 0.61 while leaving the underlying spread essentially unchanged. Later lineage work showed that this analysis compared surface outputs with top-of-atmosphere reference values.
What it means
This analysis is useful history but no longer establishes why current surface-confidence coverage behaves as it does. The corrected surface-to-surface proof set supersedes both this diagnosis and the first failed run.