Measured Evidence & Benchmark Datasets
Agent reliability cannot be established through anecdotal demos or self-grading LLM judges. Every number on this site originates from preregistered experimental designs, evaluated by deterministic test oracles across 18,000 multi-agent missions and 1,980 single-agent stress sessions.
Inline Interception & Circuit Breaker in Action
Interact with the live runtime monitor below to see how AgentAssert inspects streamed tokens character-by-character, halting generation in under 10ms upon detecting an invariant breach.
Single-Agent Domain Adherence Across 7 Frontier LLMs
Evaluated across 200 distinct scenarios spanning Financial Services, Healthcare, Support, Code Generation, Research, and Governance. Uncontracted baselines fail soft clauses 5.2–6.8 times per session, while AgentAssert maintains $\Theta = 0.9541$ aggregate reliability.
| Evaluation Domain | Scenarios | Sessions | Uncontracted Violations | Contracted Theta (Θ) | Hard Compliance |
|---|---|---|---|---|---|
| Financial Wealth Management | 30 | 300 | 5.8 / sess | 0.9812 | 100% |
| Healthcare Clinical Triage | 30 | 300 | 6.4 / sess | 0.9840 | 100% |
| Enterprise Customer Support | 30 | 300 | 5.2 / sess | 0.9805 | 100% |
| Autonomous Code Generation | 30 | 300 | 6.8 / sess | 0.9790 | 98.2% |
| Scientific Research Synthesis | 25 | 240 | 5.4 / sess | 0.9710 | 100% |
| Policy & Regulatory Governance | 25 | 240 | 5.1 / sess | 0.9725 | 100% |
| Multi-Agent Composition (Baseline) | 30 | 300 | 9.2 / sess | 0.8920 | 88.0% |
The 90.0% Multi-Agent Co-Failure Phenomenon
When two instances of the same model family are composed in a handoff pipeline, they do not fail independently. They fail together on 90.0% of missions where either fails, creating massive signed bias if independence is assumed.
Substituting a distinct model architecture (e.g. GPT-4o + Claude 3.5 Sonnet) dramatically reduced co-failure correlation across 6 of 6 registered contrasts.
Switching hosting vendors (e.g. OpenAI Direct vs Azure OpenAI) for the same underlying foundation model yielded zero statistical reduction in co-failure rates.