Numbers you can reproduce yourself.
A performance claim that cannot be reproduced is marketing. In-process enforcement adds work to every protected call, and the honest way to discuss that is to publish the method, the workload and the harness alongside the result.
1. What this enforces
Nothing. This page exists to make a claim falsifiable.
2. Where it enforces
Not applicable — but the measurement point matters and will be stated. A figure measured at the gateway is not a figure measured at the method.
3. What evidence it produces
The harness, the workload definition, the hardware and runtime, the distribution rather than a single number, and the behaviour under failure — not only under load.
4. Why you cannot assemble this from what you already own
A prospective buyer’s real question is not ‘is it fast’ but ‘what does it cost me per protected call, at my percentile, with my policy shape’. No competitor’s published number answers that, and neither does an unreproducible one of ours.
5. What Q4 2026 adds
Public Reproducible Performance Benchmarks
Publish p50/p95/p99 latency, throughput, memory, startup, cache, batch, multi-tenant and degraded-mode benchmarks.
6. Acceptance criteria
These are the conditions the capability must satisfy to be considered complete. They are quoted from the specification rather than paraphrased, because an acceptance criterion that has been reworded is no longer the criterion.
- Benchmark harness and source are public or customer-reproducible. MF-022
- Results include p50, p95 and p99 under declared hardware and workload conditions. MF-022
- Regression thresholds fail release CI when exceeded. MF-022
7. What is still open
The following are genuinely undecided rather than merely undocumented, and each one changes what the capability is:
- Whether results are published per policy shape, since a relationship-heavy policy and a scope check are not the same workload.
- Who runs the benchmark for publication, and whether an independent run is required before the Independently Verified badge applies.
- Which failure modes are measured: cold start, cache miss, control-plane unavailable, partition.
Every Q4 2026 extension is also held to four platform-wide requirements, six test classes and four release gates. They are published once, on the Q4 2026 roadmap, rather than repeated on every page.