Measurement-first studies of LLM robustness in clinical triage. Each asks one question: does the decision change when an input that should not matter changes? Framed as consistency, not correctness.
Paper
Venue
Focus
Status
Gender-Dependent Diagnostic Substitution in LLM Triage
Does the leading diagnosis flip when only the patient's stated gender changes, and nothing clinically relevant.
arXiv:2606.03641
Robustness
Published
Implicit Geographic and Language Inference in LLM Triage
Does the model infer location or language from phrasing alone, then let it move the triage decision.
arXiv:2606.01204
Robustness
Published
Socioeconomic Inference from ZIP Code in LLM Triage
Does the model read socioeconomic status off a ZIP code and let it shift a decision the clinical facts do not.
In preparation
Robustness
Draft
Three open-source benchmarks, one thesis: a system is reliable only if its decision holds when an input that shouldn't matter changes. Each turns a failure mode into a single number you can run against any model.
Benchmark
Measures
Type
Status
BreakBench
Does a single agent keep its constraints and stay honest under pressure that shouldn't change its behavior? Consistency, not skill.
BreakRate
Open source
Live
BreakBench · dimensions
Policy compliance
Keeps its stated policy under multi-turn pressure.
BreakRate
Metric fabrication
Refuses to invent numbers it cannot support.
BreakRate
Guardrail bypass
Holds its guardrails when pressed to drop them.
BreakRate
CaptureBench
How do agents behave as economic actors? Measures the share of available value a model wins in negotiation, bargaining, and competition.
CaptureRate
Open source
Live
CaptureBench · live demos
Negotiation
Hotel-room and salary negotiations played out between models.
Live demo
Run
Bargaining
The ultimatum game: how much one model will give up to close.
Live demo
Run
Competition
Price war and prisoner's dilemma between competing agents.
Live demo
Run
Collusion
Two AI pricing agents quietly form a cartel โ and the one prompt that breaks it.
Live demo
Run
TriageBench
Does a clinical model give the same triage decision when you change a patient's gender, language, or ZIP code? One command against any model.
Counterfactual consistency
Open source
Live
TriageBench · dimensions
Gender
Same diagnosis when only the patient's stated gender changes.
arXiv:2606.03641
Paper
Language
Same triage when only the inferred language changes.
arXiv:2606.01204
Paper
ZIP code
Same decision when only the ZIP's socioeconomic signal changes.
in preparation
Draft
Source experiments
ai-behavioral-experiments
The open harness behind the benches. Herd behavior: 995 of 1,000 agents sold on one bad headline. Coordination cost: 5 agents burned 12x the tokens to underperform one.
Experiments
Open source
Repo
Build
Category
Stack
Status
HawkerSense
7,000+ users. #27 on the App Store Health and Fitness chart. Featured by AsiaOne.
AI / Mobile
Swift, Gemini
Live
ELIXIR
Maps personal taste into a 22-dimensional space and recommends by position via UMAP, not by tags.
AI / Web
Python, UMAP
4,180 beverages
NameWise
Pronunciation, cultural context, and addressing tips for any name. Same language-inference thesis as the triage research.
AI / Web
Next.js, Gemini
Live
More builds
GroupPulse
WhatsApp group analysis
React
Beside
AI relationship translator
Swift, Gemini
Shiok Scout
ML residual analysis
Streamlit, sklearn
YouTube Vibe Matcher
Multimodal similarity
Next.js, Gemini
Log Cake Protocol
Computer vision fairness
Streamlit, Gemini
AlwaysLive
Autonomous AI streamer
Next.js, HeyGen
Solara
AR sun visualizer
SwiftUI, ARKit
Wrap Me Up
Digital-footprint recap
Next.js, Gemini
Autonomous Content Creator
Story-to-Instagram agent
Vertex AI
AI Workout Form Corrector
Real-time pose estimation
MediaPipe, OpenCV