The runs behind the figures
Our product pages carry numbers. The runs behind them are published here, with the day each was run and what came back.
Quorum Technologies · newest run 13 September 2026
01
The latest
Kerf Online against TinyFish, on 76 addresses
Two products that turn a URL into Markdown, called the same way on one frozen set of pages and scored against each page’s own HTML.
76/76
pages came back
28/29
tables stayed tables
90%
of headings kept
A pricing page and a deck, as Kerf read them
A pricing page with its whole comparison table, and a deck read by what the slide draws rather than the order the shapes were saved in.
14,696
characters returned
1
table stayed a table
8
cuts in the deck
Attest on 190 hand-labelled claims
A labelled set of claims: multi-hop reasoning, arithmetic, distractors, scope traps, conflicting sources. The run also found a defect, published on the same page.
95.9%
verdict accuracy
1.6%
overclaim rate
190
hand-labelled examples
02
Everything we publish
| Date | Subject | What we published | What it is | Where |
|---|---|---|---|---|
| 13 Sept 2026 | Kerf | Kerf Online against TinyFish, on 76 addresses | Study | Full run |
| 12 Sept 2026 | Kerf | A pricing page and a deck, as Kerf read them | Saved call | On /kerf |
| 06 Sept 2026 | Attest | Attest on 190 hand-labelled claims | Study | On /attest |
Study: a frozen set, a method fixed before the run, every per-item result published. Saved call: one real answer, replayed on the page it illustrates.
Dates are the day the work ran, not the day the page shipped. A repeated run replaces its figures and its date together.
The savings calculators are not in this list. They are assumptions about how long a job takes by hand, and they say so where they sit.
03
How we publish
01
Freeze
The set is written to a file with a timestamp before anything is called. The runner reads only that file.
02
Run
Each provider is called once on each item. The answer is saved verbatim, next to its source. No retries.
03
Score
The saved answers are measured against that source, or against a label written beforehand.
No model marks our work. Extraction is scored against each page's own HTML, claim verification against labels written before the run.
Runs that go badly are published as well. Attest's own measurement turned up a robustness defect, and it sits beside the accuracy figure on the Attest page.
They are also how we decide what to fix. What a run turns up in our own service becomes work, and the next run is where you see whether the work landed.
Where only one column of a comparison has been run again, both dates are printed. In the extraction comparison, Kerf's column was re-run after changes to how it reads pages; TinyFish's answers are the ones saved on the day, and every page is still scored against the HTML saved then.
The extraction run recorded three failures which turned out to be our own edge reporting a site's refusal as our error. They stay in the run as they happened, with the re-check under them rather than in place of them: a run is a record of a day, not a claim about today.
04
Ask us to measure something
If a figure you need is missing, or you would rather see the comparison run against your own pages than against ours, say so. We will run it and publish what comes back.
Attest's labelled set is one author's work on one distribution, and hosted inference is not bit-reproducible, so small differences between runs are noise. Both caveats are stated on the Attest page. If your documents look nothing like that set, the useful number is one measured on yours.
Figures are read from the run that produced them rather than typed onto a page. Naming another product is not a claim about anything except what it returned on the addresses we tested, on the day we tested them.