← Back to Reliability Explorer

Measured selection evidence

Context7

context7

Measured historical tool-selection reliability from frozen benchmark evidence.

Single-batchCurrent live
Current live tools-list snapshot. Historical evidence does not establish current endpoint availability or current production behavior.
1independent benchmark batches
24measured frozen cases
S1evidence level(s)
0observed semantic misses

First measurement: 2026-09-10 13:23:12.077611 UTC · Last measurement: 2026-09-10 13:23:12.077611 UTC · Observed strict accuracy range: 100.0%–100.0%

Batch history

Each immutable benchmark batch is shown separately. No pooled accuracy is calculated here.

Benchmark / batchTypeFrozen casesStrict scoreWilson 95% CIProvider-valid coverageInfrastructure failuresEvidence / measured
selection-evidence-expansion-v1
selection-evidence-expansion-v1
CONSTRUCTED_FROZEN_SELECTION_BENCHMARK24 24/24
100.0%
Not recorded–Not recorded 24 / 24 (100.0%)0 S1
2026-09-10 13:23:12.077611 UTC

Observed semantic misses

These are preserved adjudicated historical observations, attributable to a benchmark, batch, and case.

Benchmark / batchCase IDExpected toolSelected toolBoundary familyAdjudication
No observed semantic misses in current evidence. This is not proof that no failures occurred.

Evidence provenance

Technical identities are secondary detail for auditability.

Context7
Canonical subject ID
context7
Provenance type
CONSTRUCTED_FROZEN_SELECTION_BENCHMARK
Identity mode
LIVE_TOOLS_LIST_SNAPSHOT
Identity source
LIVE_TOOLS_LIST
selection-evidence-expansion-v1 immutable identities
case_freeze_sha256
5d248e845fc1b52ef53585682a97a592fecd0d3c1b8c9e561c758a4222273aed
freeze_commit_sha
d2937086b419ba0bf77517da772540de1975c45f
measurement_artifact_sha256
27f281b49b50fe5218102274c57db0110f3119091ca5dd53ebe954997f4fe826
model_lock_sha256
a8df78457b76e0dea73f66e2a75907feab3c42885db23658caada6973c5cb478
promotion_policy_sha256
6714ca504b2c5679c0f1ba0c7da2e5ee52f61b70b0e814ecaf9071f080f9f61d