Three open-source benchmarks, one thesis: a system is reliable only if its decision holds when an input that shouldn't matter changes. Each turns a failure mode into a single number you can run against any model.
Benchmark
Measures
Type
Status
๐Ÿ“Š
BreakBench
Does a single agent keep its constraints and stay honest under pressure that shouldn't change its behavior? Consistency, not skill.
BreakRate
Open source
Live
BreakBench · dimensions
๐Ÿ“‹
Policy compliance
Keeps its stated policy under multi-turn pressure.
BreakRate
๐Ÿ”ข
Metric fabrication
Refuses to invent numbers it cannot support.
BreakRate
๐Ÿšง
Guardrail bypass
Holds its guardrails when pressed to drop them.
BreakRate
๐Ÿ“Š
CaptureBench
How do agents behave as economic actors? Measures the share of available value a model wins in negotiation, bargaining, and competition.
CaptureRate
Open source
Live
๐Ÿ“Š
TriageBench
Does a clinical model give the same triage decision when you change a patient's gender, language, or ZIP code? One command against any model.
Counterfactual consistency
Open source
Live
Source experiments
๐Ÿงช
ai-behavioral-experiments
The open harness behind the benches. Herd behavior: 995 of 1,000 agents sold on one bad headline. Coordination cost: 5 agents burned 12x the tokens to underperform one.
Experiments
Open source
Repo