Star on GitHub ★
Empirical Evidence · AgentContract-Bench · ABC II §3

Measured Evidence & Benchmark Datasets

Agent reliability cannot be established through anecdotal demos or self-grading LLM judges. Every number on this site originates from preregistered experimental designs, evaluated by deterministic test oracles across 18,000 multi-agent missions and 1,980 single-agent stress sessions.

18,000
PREREGISTERED MISSIONS
90.0%
SHARED CO-FAILURE
6.66
LOG ODDS RATIO (OR)
6 / 6
CONTRASTS REDUCED
Live Simulation Testbed Sub-10ms Inline Circuit Breaker

Inline Interception & Circuit Breaker in Action

Interact with the live runtime monitor below to see how AgentAssert inspects streamed tokens character-by-character, halting generation in under 10ms upon detecting an invariant breach.

Live Runtime Invariant Interceptor < 10ms
LATENCY: 4.2 ms STATUS: MONITORING
Simulate Attack / Breach:
Clause: during: invariant_I2
Tokens Audited: 0
Interception Verdict: PASS
Dataset 1 AgentContract-Bench v1 (arXiv:2602.22302)

Single-Agent Domain Adherence Across 7 Frontier LLMs

Evaluated across 200 distinct scenarios spanning Financial Services, Healthcare, Support, Code Generation, Research, and Governance. Uncontracted baselines fail soft clauses 5.2–6.8 times per session, while AgentAssert maintains $\Theta = 0.9541$ aggregate reliability.

Evaluation Domain Scenarios Sessions Uncontracted Violations Contracted Theta (Θ) Hard Compliance
Financial Wealth Management 30 300 5.8 / sess 0.9812 100%
Healthcare Clinical Triage 30 300 6.4 / sess 0.9840 100%
Enterprise Customer Support 30 300 5.2 / sess 0.9805 100%
Autonomous Code Generation 30 300 6.8 / sess 0.9790 98.2%
Scientific Research Synthesis 25 240 5.4 / sess 0.9710 100%
Policy & Regulatory Governance 25 240 5.1 / sess 0.9725 100%
Multi-Agent Composition (Baseline) 30 300 9.2 / sess 0.8920 88.0%
Domain Breakdown · Reliability Θ Across 1,980 Sessions
0.9812
Finance
0.9840
Healthcare
0.9805
Support
0.9790
Code Gen
0.9710
Research
0.9725
Governance
0.8920
Composition
Simulated Evaluation Stream · 100 Live Runs Under Contract
Figure 3 · 240 sessions — drag horizontally to set k
k 3 0.950
Each square is one session. Green: clause held clean or recovery fixed it within k attempts. Red: needed more than k attempts — a genuine violation counted against the guarantee.
Dataset 2 18,000 Preregistered Missions Study (arXiv:2608.12895)

The 90.0% Multi-Agent Co-Failure Phenomenon

When two instances of the same model family are composed in a handoff pipeline, they do not fail independently. They fail together on 90.0% of missions where either fails, creating massive signed bias if independence is assumed.

Figure 4 · Co-failure lattice — drag left/right to change φ
φ 0.916 log OR 6.66 co-fail 90.0%
Model Architecture Diversity (Significant Effect)

Substituting a distinct model architecture (e.g. GPT-4o + Claude 3.5 Sonnet) dramatically reduced co-failure correlation across 6 of 6 registered contrasts.

Cloud Vendor Substitution (Null Result)

Switching hosting vendors (e.g. OpenAI Direct vs Azure OpenAI) for the same underlying foundation model yielded zero statistical reduction in co-failure rates.

Empirical Coverage Decay Under Unchecked Independence Assumption
Figure 5 · Drag horizontally for more sample data (n)
n 2,000 FITTED: MISSES TRUTH