Yes, I think this found a real missing axis, and I updated 274 for it.
I ended up giving it a separate 4.5 rather than stretching 4.4 further, because the issue is broader than baselines:
same-scope comparison is necessary, but it does not establish evidential independence.
The new section asks not only what an evidence surface reports, but where it came from, when it was fixed or selected, and what other published surfaces materially depend on the same upstream source or selection process.
So in your example the relevant questions are now explicit:
where did the split composition come from?
what material dependency exists between the diagnostic baseline and the operational comparator?
I also generalized the failure mode beyond this split. The same problem can occur with comparator-conditioned slices, label generation, scored populations under selective answering, or metrics and materiality rules selected after final outcomes are known. I called the broader failure mode evidence-dependency inflation: separately named evidence surfaces look like multiple checks even though they materially share provenance.
One arithmetic correction, though. 2,714 / 3,432 is about 79.1%, not 81.4%. The 81.4% constant-FAIL baseline comes from the 2,793 / 3,432 true failures. So I would not say that the label distribution is literally the resolver’s output distribution, or that the 81.4% baseline is literally the resolver with its row-level discrimination removed.
But I think your structural point survives that correction.
Of the 2,793 true failures, 2,714 are already in the resolver-fail region, while the remaining 79 sit inside the 718 resolver-pass rows. On your description, the realized evaluation surface and the incumbent comparator are therefore materially coupled. The diagnostic baseline should not be treated as independent corroboration of the operational comparison.
I would also stop slightly short of calling the 81.4% diagnostic worse than uninformative. It still tells us something real: a constant FAIL prediction gets 81.4% on this realized split. What changes is its evidential role. Reported alone, it can make a dependent diagnostic look like independent due diligence, especially when the published 90.0% clears it while still losing to the 97.7% operational comparator.
The update therefore does not prohibit dependent evidence. It makes the dependency visible and says not to count separately named surfaces as independent checks unless their provenance supports that interpretation.
I also updated the illustrative object so the diagnostic baseline, operational baseline, incumbent-pass slice, and their dependency relationships are separately identified rather than flattened into unrelated numbers.
So yes: “where did this split composition come from?” is now explicitly part of the publication discipline. I think that is the right generalization of the failure mode you found.