CandorCheck.
An open-source, local-first pressure test for the moment an AI stops knowing and starts guessing.

CandorCheck asks a specific question: when an AI model doesn't know something, does it say so — or does it start guessing?
It works like a pressure test. A deterministic 12-task naturalistic prompt invites guessing: fabricated citations, API traps, false premises. Every trap ships with dated, hashed evidence, so scoring is checkable rather than arguable. A guided scorer then separates response mode, resolution and factual claims instead of collapsing everything into one number.
Reports export as portable JSON and compare locally in the browser — no submissions, no central leaderboard. And CandorCheck scopes its own claims honestly: it reports behavior on one named form under recorded conditions. It does not estimate a universal hallucination rate or certify that one model is universally better.