Skip to content

Benchmarks

Measured, not asserted

A fact-checker's accuracy is itself a factual claim, so we hold ourselves to the standard we apply to everyone else: publish the numbers, the datasets, and the metric definitions — or don't make the claim.

What we measure

  • Verdict accuracy against ground-truth labels, on public claim-verification datasets and on curated documents seeded with a known mix of true, false, and misleading claims. For comparability across datasets, the 7-point scale is collapsed to supported / refuted / unverifiable.
  • Calibration — whether an 80%-confidence verdict is right about 80% of the time. A well-calibrated "moderate" is more useful than an overconfident "high."
  • Stability — the same document, checked twice, should produce the same claims, the same price, and the same verdicts. We already publish how we engineer for this on the methodology page.

Results and timeline

The full benchmarking suite runs after our first major quality milestone, planned for October 2026. We're iterating on the pipeline rapidly right now — panels, judges, and adjudication logic improve week to week — and a benchmark is only meaningful against a configuration that holds still. Benchmarking a multi-model pipeline honestly is also genuinely expensive: a single full run — every claim voted on by every panel model, judged, challenged, and repeated enough times to measure stability — carries an inference bill that can reach into the tens, potentially hundreds, of thousands of dollars. We'd rather spend that once, on a build worth measuring, than on a moving target.

The table below shows the shape of what will publish here. Every published number will come with its dataset, date, and the exact panel configuration that produced it.

MetricModel Fact-Check (T2)Live Research (T3)Definition
Verdict accuracy (supported / refuted / unverifiable)Agreement with ground-truth labels
Calibration errorGap between stated confidence and observed accuracy
Run-to-run verdict stabilityIdentical document, repeated runs

No numbers yet is deliberate: we publish measurements, not projections. Follow the blog for the write-up when the results land.

Why trust the process before the numbers land

Two design decisions are checkable today without any benchmark: every report exposes its full audit trail (each model's vote, the challenger's refutation attempt, the judge's reasoning, and citations), and confidence scores are computed from observable signals rather than self-reported — the formula is documented on the methodology page. You can audit any verdict we produce, starting with your own documents.