AI evaluation · 2026 · Age 15

CandorCheck.

An open-source, local-first pressure test for the moment an AI stops knowing and starts guessing.

candorcheck.vercel.app
CandorCheck's landing page — Test how an AI behaves when guessing is tempting, with the 12-task pressure test
12Deterministic tasks per form
0Central leaderboards or submissions
100%Local — reports never leave your machine

CandorCheck asks a specific question: when an AI model doesn't know something, does it say so — or does it start guessing?

It works like a pressure test. A deterministic 12-task naturalistic prompt invites guessing: fabricated citations, API traps, false premises. Every trap ships with dated, hashed evidence, so scoring is checkable rather than arguable. A guided scorer then separates response mode, resolution and factual claims instead of collapsing everything into one number.

Reports export as portable JSON and compare locally in the browser — no submissions, no central leaderboard. And CandorCheck scopes its own claims honestly: it reports behavior on one named form under recorded conditions. It does not estimate a universal hallucination rate or certify that one model is universally better.

← PreviousFixMapNext →SoundWise