Evidence

Benchmarks

We publish methodology before theater. This page states what we measure, what we do not claim, and how we will grow the public eval set.

What “good” means for CiteSafe

  • Recall of citation strings in legal drafts (extract quality)
  • Precision of fabrication / void-cite flags on known traps
  • Coverage of public primary indexes (CourtListener, CAP, GovInfo) when configured
  • Time-to-status for a typical memo-length paste
  • Honesty: offline Scout admits offline — never pretends live primary when API is down

What we do not publish as marketing math

  • “Hallucination-free” percentages
  • Win-rate or case-outcome predictions
  • Shepard’s / KeyCite parity scores without licensed treatment data
  • Inflated corpus counts before bulk ingest is production-true

Public trap demos

The site includes a trap sample with invented citations so anyone can see Could Not Verify behavior. Use Try sample and Trap on the home box, or the product demo at /#demo.

Roadmap for formal evals

Public “cite hallucination” benchmark set (generators fail, CiteSafe flags) is on the Trust OS roadmap — the competitive moat is an open eval culture, not secret green badges. Roadmap →

Ready to verify a draft?

Free beta · 3 checks/mo · Confidential Mode on by default.