How we validate decision quality

    Every claim on this page is backed by a file on disk. Agent behaviour is measured, recorded and reproducible, across multiple independent evidence types and two independent model families.

    Evidence current at 27 August 2026. Figures are measured, not modelled.

    Evaluation method

    Agent behaviour is evaluated on four independent evidence types, each producing files rather than assertions.

    Offline tool assertions

    418 at today's baseline, covering 361 tool invocations. They exercise every deterministic function with no language model involved, confirming that a tool returns exactly what it claims, including returning null where data is absent. They run in seconds and gate every merge.

    Verification records

    These measure behaviour under a language model. Every case runs three times against each of two independent model families, so a result is never the property of a single model. The fleet total is 318 verification cases across thirteen agents, 21 to 33 per agent, and 1,908 model-bearing runs per full fleet pass, representing 23.5 hours of sequential wall-clock, roughly one to four hours per agent. Each record states the pass rate of every case with the commit, the agent version, both model configurations and the serving endpoint. Verification is strictly sequential, so every measurement is reproducible. 24 records are on disk.

    Structural gates

    These are enforced by the build rather than by review: 484 engine and conformance tests, including 8 conformance tests covering package isolation between agents, unique identifiers, endpoint presence and register completeness, inside a suite of 1,287 test functions across 81 test files, 29 agent-side and 52 platform-side.

    Mutation checks

    Each test is re-run against a deliberately broken implementation to confirm it fails when the behaviour is absent. This proves the test measures the thing it claims to measure.

    Measured throughput for context: a single-agent answer comes back in well under a minute; the same agent working as one step inside a five-agent chain averages a little under a minute and a half, measured 27 August over 100 chained answers.

    Further testing in place

    The four evidence types above sit within a wider standing programme. Additional suites run continuously alongside them:

    • 1,287 test functions across 81 test files, run on every build
    • 484 engine and conformance tests, including 8 conformance tests enforcing structural guarantees across the agent estate
    • 418 offline tool assertions covering 361 tool invocations
    • 318 verification cases across thirteen agents, 21 to 33 per agent, each run three times against each of two model families
    • 1,908 model-bearing runs per full fleet pass
    • 24 verification records held on disk, each naming its commit, agent version and serving endpoint
    • Mutation coverage confirming that every test fails when the behaviour under test is removed
    • Chained-execution measurement across 100 multi-agent answers, recorded 27 August

    Reproducibility and provenance

    Numerical calculation, policy evaluation and source-record retrieval are separated from language-model generation. Current-state figures are computed by deterministic services from customer-held source data and trace to a source record. Forecasts are produced by approved analytical models, labelled as forecasts, and carry their inputs. A language model interprets the task, structures the reasoning and explains the result; it does not originate the numbers.

    Every answer is written to the hash-chained decision ledger before a commander sees it, recording the commit, component versions and the configuration that served it, so any recommendation can be reconstructed after the fact. Where a signing key is configured, each ledger record carries an ML-DSA-87 (FIPS 204) signature.

    What this evidence establishes.

    That every figure came from a named function that actually executed, that unavailable data is reported as unavailable, that provenance survives to the output, and that the same question behaves the same way across two independent model families. Every output is a recommendation, and the deciding authority remains the officer the system reports to.

    Adversarial assurance

    An adversarial assurance battery of 13 tests in 5 series was run against release v2.1.1: provenance forgery (A1 to A4), determinism and forged markings (B5 to B7), fragment fusion and classification downgrade (C8a to C8f, C9), canary flooding and supervisor-override injection (D10 to D11), and non-English rule bypass plus homoglyph tool spoofing (E12 to E13).

    The battery recorded zero failures. The strongest result was C9, in which the system rejected a forged classification marking, redacted an untrusted token and caught a phantom tool call in a single exchange. Under homoglyph spoofing, a tool name carrying 3 Cyrillic bytes resolved correctly to the genuine tool and no fabricated figure passed, because tool selection proceeds by intent rather than spelling. The model stayed warm and the application healthy throughout, with no crash and no cold start.

    13
    tests across 5 series
    0
    failures

    Results are recorded against release v2.1.1 and the deployment they were run on. A battery result names its version and deployment.

    Confidence handling

    Where confidence values are reported, they are computed by a deterministic function from a feed's own uncertainty product or a published skill table, and never authored by a language model. Where neither source exists, the value is null rather than estimated. Confidence arrives with the planning capability, and calibration becomes measurable at that point.

    Tool interfaces referenced elsewhere in this documentation set: forecast_at, forecast_window, feed_status.

    This page is revised when the evidence changes, not when the product does. The date of currency at the top of the page records the last revision.

    Request the underlying verification records under NDA