Independent MCP evidence
Get independent selection evidence for your MCP.
Get an independent tool-selection evaluation for your MCP. We freeze the benchmark before measurement, preserve misses, and produce a reproducible evidence report.
The offer
Independent MCP Selection Evaluation
Agent Trust Lab evaluates whether a model selects the intended MCP tool across a frozen benchmark and produces an auditable evidence report. This is a paid evaluation service, not certification.
What you receive
A concrete evidence package
- Tool schema qualification and benchmark design
- Frozen benchmark, locked before model measurement
- Real-model first-tool selection measurement
- Strict accuracy and Wilson confidence interval where applicable
- Provider-valid coverage and preserved semantic misses
- Evidence provenance and an immutable measurement receipt
- Concise human-readable evidence report
A public Agent Trust Lab evidence profile may be considered after review only when you explicitly choose eligibility. Publication is never guaranteed or automatic.
How it works
Five steps from request to report
- Submit your MCP endpoint or repository
- Qualification and benchmark design
- Benchmark frozen before model measurement
- Real-model selection measurement
- Evidence report delivered
Scope and limitations
What this evaluation does not certify
It measures tool-selection evidence only. It does not certify MCP security, business correctness, uptime, regulatory compliance, safety, future model performance, or overall product quality. It is not certification, and a declined submission is not evidence that an MCP is poor quality.
Qualification
A useful evaluation needs a defensible test surface
- Reachable public MCP or auditable version-bound source
- Sufficiently stable tool schema
- At least two meaningful tool choices where selection evaluation makes sense
- External or defensible gold construction possible
- No prohibited or destructive evaluation requirement
- No authentication barrier unless specifically arranged
We may decline requests that cannot be evaluated safely or defensibly.
Demand validation
Request an evaluation
Share only public project information. Do not submit API keys, secrets, passwords, or private credentials.