Skip to content

Benchmarks

How we measure accuracy

A fact-checker's accuracy is itself a factual claim. So we apply our own standard to it: we publish the numbers, the datasets, and the metric definitions, or we do not make the claim.

What we measure

  • Verdict accuracy. Agreement with ground-truth labels on public claim-verification datasets and on curated documents seeded with a known mix of true, false, and misleading claims. We collapse the 7-point scale to supported, refuted, or unverifiable so datasets are comparable.
  • Calibration. Whether an 80%-confidence verdict is right about 80% of the time. A well-calibrated "moderate" is more useful than an overconfident "high."
  • Stability. The same document, checked twice, should produce the same claims, the same price, and the same verdicts. The methodology page describes how we engineer for this.

Results and timeline

The full benchmarking suite runs after our first major quality milestone, planned for October 2026. The pipeline is changing week to week right now: panels, judges, and adjudication logic. A benchmark is only meaningful against a configuration that holds still. A full run is also expensive. Every panel model votes on every claim, a judge and a challenger weigh in, and the run repeats enough times to measure stability. The inference bill for that can reach tens of thousands of dollars, and possibly hundreds of thousands. We would rather spend that once, on a build worth measuring.

The table below shows the layout the results will use. Every number will come with its dataset, date, and the exact panel configuration that produced it.

MetricModel Fact-Check (T2)Live Research (T3)Definition
Verdict accuracy (supported / refuted / unverifiable)PendingPendingAgreement with ground-truth labels
Calibration errorPendingPendingGap between stated confidence and observed accuracy
Run-to-run verdict stabilityPendingPendingIdentical document, repeated runs

There are no numbers yet because we publish measurements, not projections. The write-up will go on the blog when the results are in.

What you can check today

Two design decisions are checkable without any benchmark. Every report exposes its full audit trail: each model's vote, the challenger's refutation attempt, the judge's reasoning, and citations. And we compute confidence scores from observable signals rather than taking them from a model. The methodology page documents the formula. You can audit any verdict we produce, starting with your own documents.