We measure whether the answers are actually right.

Most document AI ships on vibes. ChatPDF is graded against a fixed set of 60 questions written from 5 real public documents - 237 pages of federal frameworks, financial reports and policy circulars - with every answer scored for correctness, grounding and citation accuracy.

91.5%

answered end to end

54 of 59 questions cleared every gate

90.9%

citations verified

100 of 110 quotes matched the source text exactly

1.6s

median answer time

194ms to find the evidence

$0.53

to run the suite

2.34M tokens across 353 API calls

How we score

Four independent gates, not one blended number

A single accuracy score hides which part of the system failed. We grade retrieval, content, citations and behaviour separately, so a weakness in one never disguises itself as a weakness in another. End to end means all four passed on the same question.

Pass rate by gate

Hover any bar for the 95% confidence interval.

Retrieval

98.3%

The evidence needed to answer reached the model's context window.

Content

96.6%

The answer is correct, complete, and faithful to the retrieved sources.

Citations

93.2%

Every claim is attributable to a quote that exists in the document.

Behaviour

96.6%

The system abstains when it should and respects page scoping.

Where it falls short

Every miss is classified by root cause

When a question fails, it lands in exactly one bucket - attributed to the earliest stage responsible, so a retrieval problem is never mislabelled as a writing problem. This is the list we work from.

Outcome distribution

60 questions in the set

Passed every gate

Correct, grounded, cited and behaving as intended.

54

System error

The case never produced a judgeable answer.

1

Abstention failure

Answered a question the document cannot support, or refused an answerable one.

2

Citation failure

The answer itself is right; the citations are missing, wrong, or unverifiable.

3

Answer quality

Graded on seven dimensions, not a thumbs up

Two independent judge passes score every answer from 0 to 4 on each dimension. When they disagree, a third adjudication pass settles it. Further from the centre is better.

Quality profile

Mean score across every evaluated question, out of 4.

Citations that hold up

Quotes matched to source text

90.9%

Every citation is checked deterministically against the document text, not taken on the model's word. 100 of 110 quotes were found verbatim in the source.

Context precision

How much of what we retrieve is actually on-topic.

Irrelevant context in the window

2.40 / 4

Scored 0 to 4 where lower is better, so a wider bar means a tighter, less padded context window.

Speed

Median and 95th percentile.

Find the evidence
194ms / 459ms p95
Write the answer
1.6s / 3.3s p95

The test set

Real documents, not toy examples

237 pages across 5 public documents - dense tables, appendices, footnotes and multi-column layouts, the material that actually breaks document AI.

End-to-end pass rate by document

Every document carries the same number of questions.

What's in the corpus

  • Artificial Intelligence Risk Management Framework (AI RMF 1.0)

    Published 2023-01-26

    48 pp

  • The NIST Cybersecurity Framework (CSF) 2.0

    Published 2024-02-26

    32 pp

  • Zero Trust Maturity Model Version 2.0

    Published 2023-04-01

    32 pp

  • Economic Well-Being of U.S. Households in 2024

    Published 2025-05-28

    90 pp

  • OMB Circular No. A-123: Management's Responsibility for Internal Control

    Published 2026-03-10

    35 pp

91%

single turn

34 questions

100%

page scoped

15 questions

80%

unanswerable

10 questions

What it costs

Every token is accounted for

The harness keeps a ledger of each billable API call, split by pipeline stage, so cost per answered question is a measured number rather than an estimate.

Cost by pipeline stage

Grading an answer costs more than producing one, deliberately.

$0.0089

per question

$0.0183 at p95

2.34M

tokens used

353 API calls

899.6k

tokens served from cache

42% of all input

127.7k

reasoning tokens

hidden thinking, billed at output rate

Method

How the grading works

The harness runs against the real production pipeline on real models. Nothing is mocked.

  1. 01

    Fixed corpus

    5 public documents are checksummed and pinned, so a change in results can only come from a change in the system.

  2. 02

    Authored questions

    60 questions with reference answers and exact supporting quotes, each verified against the source text before entering the set.

  3. 03

    Real pipeline

    Each question runs through the same retrieval and generation path the product uses, capturing the full trace.

  4. 04

    Independent judging

    2 independent gradings at high reasoning effort - well above what answering uses - with a third adjudication pass whenever they disagree.

Answer model gpt-5.6-lunaJudge gpt-5.6-lunaEmbeddings text-embedding-3-smallDataset v1.1.0Run 2026-08-30

Ask your own documents the hard questions.

Every answer arrives with citations you can click straight through to the page it came from.