Evidence
Benchmarks
We publish methodology before theater. This page states what we measure, what we do not claim, and how we will grow the public eval set.
What “good” means for CiteSafe
- Recall of citation strings in legal drafts (extract quality)
- Precision of fabrication / void-cite flags on known traps
- Coverage of public primary indexes (CourtListener, CAP, GovInfo) when configured
- Time-to-status for a typical memo-length paste
- Honesty: offline Scout admits offline - never pretends live primary when API is down
What we do not publish as marketing math
- “Hallucination-free” percentages
- Win-rate or case-outcome predictions
- Shepard’s / KeyCite parity scores without licensed treatment data
- Inflated corpus counts before bulk ingest is production-true
Published accuracy (FP / FN)
Trap recall, false-verify, and false-fail rates live on /accuracy/. Each home verify also reports live primary-index hit % vs corpus-only vs not found.
Public trap demos
One-screen demo: Catch a plausible fake on the homepage - known-good cites plus a real-looking fabricated 9th Cir. line, then export JSON/HTML. Open demo →
Roadmap for formal evals
Public “cite hallucination” benchmark set (generators fail, CiteSafe flags) is on the Trust OS roadmap - the competitive moat is an open eval culture, not secret green badges. Roadmap →
Ready to verify a draft?
Free beta · 3 checks/mo · Confidential Mode on by default.