Reliability
What the last run measured.
We publish the numbers because most tools in this category do not. Every figure below comes from running 39 agent scenarios against the deployed servers, including the failure cases.
100%scenarios passed
39/39scenario results
5degraded upstream
2.80sslowest p95 across servers
30 Septmeasured 10:32 UTC
verifydesk
Measured 30 Sept, 10:32 UTC, against the deployed server.
| Tool | Scenarios | Passed | p95 |
|---|---|---|---|
| verify_company | 4 | 100% | 2.80s |
| verify_vat | 2 | 100% | 1.20s |
| validate_iban | 4 | 100% | 664ms |
| screen_sanctions | 3 | 100% | 1.62s |
| enrich_company | 2 | 100% | 1.21s |
| get_usage | 1 | 100% | 583ms |
jobsradar
Measured 30 Sept, 10:32 UTC, against the deployed server.
| Tool | Scenarios | Passed | p95 |
|---|---|---|---|
| list_supported_companies | 2 | 100% | 1.32s |
| search_jobs | 5 | 100% | 1.34s |
| get_company_jobs | 3 | 100% | 892ms |
| get_usage | 1 | 100% | 504ms |
invoiceforge
Measured 30 Sept, 10:32 UTC, against the deployed server.
| Tool | Scenarios | Passed | p95 |
|---|---|---|---|
| generate_invoice | 5 | 100% | 1.43s |
| validate_invoice | 3 | 100% | 583ms |
| extract_invoice | 2 | 100% | 582ms |
| describe_coverage | 1 | 100% | 563ms |
| get_usage | 1 | 100% | 548ms |
How this is measured
- The suite runs real calls against the deployed endpoint, not a mock: registry lookups, VAT checks, IBAN validation, sanctions screening.
- A scenario passes only if the answer is usable by an agent: right shape, no leaked internals, and errors that say what to do next.
- Failure cases count. Unknown companies, malformed input and unsupported countries are part of the run, because that is where tools usually break.
- p95 is the slowest response in the top 5 percent, measured end to end from outside Cloudflare, so it includes the upstream registries and job boards. Searching many job boards at once means the slowest board sets the number, and it moves between runs.
- Some sources throttle or go down without warning: the European VAT service does it regularly. When that happens the scenario is marked degraded upstream and still counts as met, because what we promise in that case is to fail well: a clean, actionable, retryable error rather than a break or a wrong answer. It is counted separately and shown in its own figure above, so a run where an upstream was down never reads as a run where everything worked. If the tool had broken or leaked internals instead, it would count as our failure.
This page shows the most recent published run rather than a rolling uptime average. The suite runs every day against the deployed servers; this page carries the run published with the site, and the dates above say exactly when it was measured. Read it as a dated measurement rather than as live uptime. We would rather publish one honest run than an average we do not actually compute.