Selection Evidence Expansion V1: MEASURED
Transport grade, profile/evidence grade, and selection evidence grade are separate.
Constructed schema-derived cases are not external public gold; S2 is not eligible.
External Gold V2 Batches 1–6 are exposed as separate immutable benchmark records; no combined headline score is used.
Batch 3 observed an early exploratory boundary-risk signal; Batch 4 remained promising but incomplete.
External Gold V2 Batch 4: 46/48 strict, benchmark-level S2; scanner status remained PROMISING_CROSS_SUBJECT_SIGNAL_MORE_DATA_REQUIRED.
External Gold V2 Batch 5: 74/77 strict, benchmark-level S2; all three misses occurred in DeepWiki, so the pre-frozen cross-subject dispersion criterion C2 failed. BOUNDARY_RISK_SCAN_V1 is not validated.
External Gold V2 Batch 6: 80/80 strict, benchmark-level S2; DeepWiki was excluded prospectively, no semantic failures occurred, HIGH-risk cases did not underperform, and pre-frozen C1/C2/C3 failed. Final BOUNDARY_RISK_SCAN_V1 status: WEAK_PREDICTIVE_SIGNAL. The scanner is archived experimental research evidence, not a validated predictor, production risk score, or reliable pre-measurement forecast.
Empirical Selection Reliability: measured historical selection reliability from frozen benchmarks. It is separate from transport health, profile trust, availability, current-live status, and promotion; it makes no prediction about future failures and does not pool batches into a headline score.
Reliability Explorer · Evaluate My MCP · Profile JSON · Empirical Reliability JSON · Comparison JSON · Batch 3 JSON · Batch 4 JSON · Batch 5 JSON · Batch 6 JSON