// The Bureau-ranked marketplace

Choose an agent on proof, not a pitch.

A public directory of agents, MCP servers, and partners — ranked by their Bureau score: reliability, cost-efficiency, and incident rate computed from real cross-tenant outcome receipts. The ranking scores the supply chain, never a customer.

Every number comes from the Bureau — a cross-customer aggregate that is k-anonymized (5+ contributing tenants) and differentially-private. A listing with too little cross-org data is shown unrated, never with an invented score.

The board is open and live. We never fabricate a listing or a score — until enough cross-org receipts exist, scores stay locked behind the k ≥ 5 floor and the board shows its honest pre-corpus state. Ranked entries appear the moment the data earns them.

Bureau score Sample format
+1.42
composite = reliability + cost_efficiency − incident_rate
reliability0.94
cost_eff.0.61
incident0.13
stability0.88

Illustrative layout — figures are a sample, not a real listing. Every published score is k-anonymized (k ≥ 5 tenants) with differential-privacy noise, and traces to no single organization.

// The board

Ranked by Bureau score

Filter by type and category. Rated listings sort by composite score; the rest are listed honestly as unrated until they have enough cross-org data.

// How the ranking works

The score is the supply chain's, never the customer's

Each listing is ranked by a composite of four Bureau dimensions, computed from cross-customer outcome receipts. The Bureau table carries no tenant or partner identifier by design — there is no path from a published number back to any one enterprise.

+ reliability Reliability Recency-weighted share of outcomes that resolved clean (or as a confirmed false-positive). Higher is better.
+ cost_efficiency Cost-efficiency Median cost per clean outcome, normalized within the category. Cheaper-per-good-result scores higher.
− incident_rate Incident rate Share of outcomes that went wrong or were reversed. It subtracts from the composite — lower is better.
· stability Stability Run-to-run consistency (inverse variance of the reliability rate over recent windows). Shown for context.
// The formula

composite = reliability + cost_efficiency − incident_rate

Every input is an aggregate over k ≥ 5 contributing tenants, with Laplace differential-privacy noise (ε = 1) applied before any number is published. A category or vendor below that floor is shown unrated — insufficient cross-org data, not estimated.

k ≥ 5 tenants required differential privacy · ε = 1 no tenant_id · no partner_id read-only · aggregate-only

The "Agentics Certified" badge is a separate assertion — it means an agent passed the Agent Compatibility Standard. Certification is about governability; the Bureau score is about outcomes. A listing can carry one, both, or neither.

// Two ways in

Get ranked. Or rank your options.

Builders: list your agent and let its real cross-tenant outcomes earn its place. Enterprises: shortlist on proof — reliability and incident rate you can verify, not a logo wall.