Blog
Engineering notes
How we build and measure an auditable multi-model fact-checker.
Who checks the fact-checker? The verification layer behind every report
A fact-checker's worst errors aren't wrong verdicts. They're the quiet ones underneath: a claim restated slightly wrong, a source cited for something it doesn't say, evidence that only looks relevant. We added an independent layer whose only job is to check our own pipeline's work.
Where checkable claims hide (and how we taught our extractor to find them)
AI extractors catch a sentence's headline fact and miss the claims tucked into appositives, asides, and subordinate clauses. How we measured that gap, closed it, and made every claim traceable to its exact source text. With an interactive demo.
What multi-model consensus catches that single models miss
Any one AI model fails quietly and confidently. A panel of rivals, an adversarial challenger, and an independent judge fail loudly. Here is the mechanism, failure mode by failure mode.
How we benchmark a fact-checker (and why our numbers aren't up yet)
Accuracy claims from AI products are usually unfalsifiable. Here's the evaluation we designed instead: the datasets, the metrics, and why the full suite runs against our October 2026 quality milestone.